Image processing method and device, electronic equipment and readable storage medium

By acquiring the hardware noise characteristics and spatial distribution characteristics of image acquisition devices through deep learning, and combining them with image denoising and texture enhancement, the problem of neglecting details and wasting resources in existing image quality enhancement methods is solved, achieving a more efficient comprehensive improvement in image quality.

CN113674159BActive Publication Date: 2025-12-19BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202011185859.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-29
Filing Date
2020-10-29
Publication Date
2025-12-19
Estimated Expiration
2040-10-29

AI Technical Summary

Technical Problem

Existing image quality enhancement methods focus too much on the overall image and neglect details, failing to effectively denoise and enhance texture. Furthermore, the cascading of multiple tasks leads to resource waste and poor results.

Method used

A deep learning-based image denoising model is adopted. By acquiring the noise intensity characteristics of the hardware environment of the image acquisition device and combining them with the spatial distribution characteristics of the noise, targeted denoising processing is performed. Combined with image tone and texture enhancement, an ideal super model is constructed to achieve comprehensive quality enhancement.

Benefits of technology

It improves image denoising, enhances image details, optimizes resource utilization, and achieves a more efficient overall improvement in image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113674159B_ABST
    Figure CN113674159B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image processing method and device, electronic equipment and readable storage medium, and belong to the technical field of image processing and artificial intelligence. The method comprises: acquiring a to-be-processed image; performing quality enhancement on the to-be-processed image by using at least one image quality enhancement manner to obtain a processed image. Based on the scheme provided in the embodiments of the present application, the quality of the to-be-processed image can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing and artificial intelligence, in particular, the present application relates to an image processing method and device, electronic equipment and readable storage medium. BACKGROUND

[0002] At present, the market of mobile terminal (including smart phone) is very hot, and the photographing performance of smart terminal becomes one of the focuses of fierce competition of major smart phone manufacturers. Various mobile terminal manufacturers from hardware, software, application and other aspects, constantly refresh the new height of image quality taken by smart terminal, greatly improve the user's photographing experience. Image quality enhancement is a broad concept, although the existing image quality enhancement scheme has achieved very good technical effect, but there is still a large improvement space. SUMMARY

[0003] The purpose of the embodiments of the present application is to at least solve one of the above technical defects, improve the image quality, and the specific technical scheme provided by the embodiments of the present application is as follows:

[0004] In a first aspect, the present application provides an image processing method, which comprises:

[0005] obtaining a to-be-processed image;

[0006] adopting at least one image quality enhancement method to enhance the quality of the to-be-processed image, and obtaining a processed image.

[0007] In a second aspect, the present application provides an image processing device, which comprises:

[0008] an image acquisition module, configured to acquire a to-be-processed image;

[0009] an image processing module, configured to adopt at least one image quality enhancement method to enhance the quality of the to-be-processed image, and obtain a processed image.

[0010] In a third aspect, the present application provides an electronic device, which comprises a memory and a processor, wherein the memory stores a computer program, and the processor executes the method provided by the embodiments of the present application when running the computer program stored in the memory.

[0011] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program executes the method provided by the embodiments of the present application when running by a processor.

[0012] The beneficial effects brought by the technical scheme provided by the embodiments of the present application will be described in the specific implementation manner part of the following text in combination with various optional embodiments, which will not be expanded here. BRIEF DESCRIPTION OF DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced.

[0014] Figure 1 A schematic diagram of the principle of a super model is shown;

[0015] Figure 2 A schematic diagram of the flow of an image processing method provided by an embodiment of the present application is shown;

[0016] Figure 3a A schematic diagram of the principle of image denoising by an image denoising model provided by an embodiment of the present application is shown;

[0017] Figure 3b A schematic diagram of the flow of an image denoising processing method provided by an embodiment of the present application is shown;

[0018] Figure 3c A schematic diagram of the principle of image denoising by an image denoising model provided by another embodiment of the present application is shown;

[0019] Figure 3d A schematic diagram of the structure of a denoising model provided by an example of the present application is shown;

[0020] Figure 3e A schematic diagram of the training principle of an image denoising model provided by an example of the present application is shown;

[0021] Figure 4a A schematic diagram of the principle of image tone adjustment by an image tone enhancement model provided by an embodiment of the present application is shown;

[0022] Figure 4b A schematic diagram of the principle of image tone adjustment by an image tone enhancement model provided by another embodiment of the present application is shown;

[0023] Figure 4c A schematic diagram of the principle of a guide map subnetwork provided by an example of the present application is shown;

[0024] Figure 4d A schematic diagram of the structure of a guide map subnetwork provided by an example of the present application is shown;

[0025] Figure 5a 、 Figure 5b and Figure 5c A schematic diagram of the principle of a texture enhancement model provided by three examples of the present application is shown respectively;

[0026] Figure 5dA structural schematic diagram of an image texture enhancement network provided in an example of the present application is shown.

[0027] Figure 5e A flow schematic diagram of a texture enhancement processing method provided in an example of the present application is shown.

[0028] Figure 5f A data processing principle schematic diagram of a double convolution module provided in an example of the present application is shown.

[0029] Figure 5g An amplification schematic diagram of the output in Figure 5f

[0030] Figure 6a A flow schematic diagram of an image processing method provided in an example of the present application is shown.

[0031] Figure 6b A principle schematic diagram of determining a processing sequence of multiple enhancement modes by a processing sequence prediction network provided in an example of the present application is shown.

[0032] Figure 6c A structural schematic diagram of an image processing model provided in an example of the present application is shown.

[0033] Figure 7a 、 7b and Figure 7c A training principle schematic diagram of several image processing models provided in an example of the present application is shown.

[0034] Figure 8 A schematic diagram of an image processing device provided in an example of the present application is shown.

[0035] Figure 9 A structural schematic diagram of an electronic device provided in an example of the present application is shown. DETAILED DESCRIPTION

[0036] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary, only for explaining the present application, and cannot be interpreted as a limitation of the present application.

[0037] ​It is to be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It is further understood that the terms "comprise" (or comprise), "comprises" (or comprises) and "comprising" (or comprising) when used in this specification, specify the presence of stated features, integers, steps, operations, elements, or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof. It is further understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element or intervening elements can be present. In addition, the word "connected" or "coupled" as used herein can include wirelessly connected or wirelessly coupled. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0038] In order to better understand and illustrate the various optional solutions provided by the present application, first, the related art involved in the present application and the problems existing in the prior image quality enhancement solutions will be described.

[0039] Image quality enhancement in the general sense is a broad concept, including image denoising, deblurring, image inpainting, super-resolution, texture enhancement and other sub-tasks of low-level image understanding. Each task is used to solve a specific sub-problem. For example, image denoising is mainly used to remove unnecessary noise information in the image, image inpainting is mainly used to repair and reconstruct damaged images or remove redundant objects in the image, image super-resolution refers to recovering a high-resolution image from a low-resolution image or image sequence, and image deblurring mainly involves how to eliminate the image blur phenomenon caused by hand shaking or defocus during shooting. In addition, there are also some image quality improvement solutions that focus on brightness, hue, contrast and other aspects of the image to make the image look more vivid.

[0040] For image quality enhancement, the most common approach in the industry is to concatenate various sub-tasks into a workflow and sequentially execute each sub-task. For example, the ISP (Image Signal Processor) in the camera adopts this typical mode. Although the ISPs of major terminal device manufacturers are not the same, they usually cannot cover sub-tasks such as image deblurring, inpainting, and texture enhancement. When there is such a task requirement, additional processing modules need to be added.

[0041] With the development of artificial intelligence technology, current image processing technology based on deep learning has made great progress. Using deep learning technology has greatly improved image processing. However, current image processing technology based on deep learning is usually developed for a specific task and usually only involves one aspect of image quality enhancement.

[0042] The inventors of the present application find through research on the prior art that the existing image quality enhancement methods at least have the following to be improved:

[0043] (1) The existing image quality enhancement methods are usually overall image quality enhancement methods, paying too much attention to the overall image and paying little attention to the image details. Such overall image quality enhancement methods mainly focus on image brightness, hue, contrast, etc., but do not pay attention to the detailed information of the image (such as image texture detail enhancement, noise information elimination, etc.). With this scheme, one possible case is that when the brightness and hue of the image are well improved, the noise in the dark part of the image becomes more obvious. The overall image quality enhancement method cannot cover image denoising, texture enhancement, etc.

[0044] (2) In some existing schemes, some specific methods for specific tasks are simply concatenated to achieve multi-aspect enhancement of image quality, but such simple concatenation does not take into account the characteristics of the tasks themselves, for example, image denoising tends to remove information, while texture enhancement tends to increase information, and it is impossible to determine their topological relationship in the overall concatenation process.

[0045] (3) Simple multi-task concatenation inevitably causes poor real-time performance, because no matter the quality of the image, the image needs to go through the preset serial processing flow, for example, one possible case is that a high-quality photo does not need additional processing, but still needs to go through all the processing flow, causing unnecessary waste of space-time resources.

[0046] For image quality enhancement, the most ideal solution is to establish an ideal supermodel (ISM) that can simultaneously solve all subtasks of image quality enhancement, such as Figure 1 As shown in FIG. 1, an input image that needs to be enhanced in quality is input into the ideal supermodel, so as to obtain an output image that has been processed in terms of quality. However, such an ideal model is difficult to achieve optimal improvement of image quality. Because such a model does not consider the nature of each task and the relationship between them, for example, a possible and relatively common task combination is image denoising, color / brightness adjustment (i.e., hue adjustment), and texture enhancement, and among them, the denoising model tends to remove information in the input image, while the texture enhancement tends to add additional information to the input image, and these two subtasks with opposite image space operation characteristics make it difficult for the model to achieve both, and the final result is to find a balance between noise elimination and texture enhancement, neither of which can achieve optimal denoising effect nor good detail enhancement goal, and more likely, some noise is enhanced as texture, even reducing the image quality.

[0047] The present application aims to solve at least one of the above technical problems in the prior art. The technical solutions of the present application and how the technical solutions solve the above technical problems will be described in detail below with specific embodiments. The following specific optional embodiments can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments. The various optional embodiments of the present application will be described below with reference to the accompanying drawings.

[0048] Figure 2 A flowchart of an image processing method provided by an embodiment of the present application is shown, which can be performed by any electronic device, such as a mobile terminal of a user, such as a smart phone. Based on the method, the user can perform real-time enhancement processing on a photo taken by the smart phone, or can process a photo already stored in the phone to obtain a photo with higher quality. For example, Figure 2 The method provided by the embodiment of the present application can include the following steps, as shown in the method.

[0049] Step S110: obtaining a to-be-processed image;

[0050] Step S120: performing quality enhancement on the to-be-processed image by using at least one image quality enhancement manner to obtain a processed image.

[0051] The to-be-processed image can be any image that needs to be enhanced in image quality, such as an image taken by a user's mobile phone in real time, an image stored in the mobile phone, or an image obtained from other devices or storage spaces.

[0052] The at least one image quality enhancement manner refers to a scheme for performing image quality enhancement processing from at least one dimension. In the embodiment of the present application, the at least one image quality enhancement manner can include but is not limited to one or more of image denoising, image tone adjustment, and image texture enhancement, wherein the image tone adjustment includes image brightness adjustment and / or image color adjustment.

[0053] The image processing method provided by the present application will be described in detail below in combination with various optional embodiments.

[0054] In an optional embodiment of the present application, the at least one image quality enhancement manner includes image denoising.

[0055] Performing quality enhancement on the to-be-processed image by using at least one image quality enhancement manner includes:

[0056] Obtaining a noise intensity feature of the to-be-processed image;

[0057] Performing denoising processing on the to-be-processed image according to the noise intensity feature.

[0058] The noise intensity feature of the to-be-processed image (which can also be referred to as a noise intensity distribution feature) represents a noise distribution feature corresponding to a device hardware environment of an image acquisition device that acquires the to-be-processed image. The noise intensity feature of the to-be-processed image can be predicted by a neural network model (which can be referred to as a noise feature estimation model, or a noise intensity feature network, or a noise intensity feature estimation network, or a noise intensity feature prediction network, or a noise intensity prediction network), that is, the to-be-processed image is input into a pre-trained neural network model to obtain the noise intensity feature of the to-be-processed image.

[0059] Blind denoising of real image denoising is very challenging and has always been a research difficulty in the academic field. Although there is a strong demand in the industry, the denoising effect of the prior art still needs to be improved. The purpose of image denoising is to eliminate noise information such as color noise, compression noise and pattern noise. The main difficulty of denoising processing is that it is very difficult to model real image noise, which is greatly different from the Gaussian white noise that the academic field is keen to model. Different sensors and environments will lead to different noise distributions. This noise can come from the camera sensor or the image processing algorithm, or even from the compression storage process of the image. The real noise distribution depends not only on the software but also on the hardware, and the noise distribution related to the hardware environment is difficult to achieve high-quality denoising with a general deep learning model.

[0060] To solve the technical problem, the optional solution provided by the embodiments of the present application considers the noise level of the input data (i.e., the input image) and the fact that the noise distribution has both intensity and spatial features (for example, for low-light images, the noise in dark areas can be significantly higher than the noise in bright areas) at the same time. Noise level assessment is performed before noise elimination.

[0061] As an optional solution, the noise intensity feature corresponding to the device hardware environment of the image acquisition device used to acquire the to-be-processed image can be obtained to improve the denoising effect based on the noise intensity feature corresponding to the device hardware environment. The device hardware environment can refer to one or more of the hardware configuration information of the image acquisition device, which can include but is not limited to camera sensor configuration information, processor configuration information, available storage space information, and the like.

[0062] In an optional embodiment of the present application, the to-be-processed image is denoised according to the noise intensity feature, comprising:

[0063] According to the noise intensity feature and the to-be-processed image, a noise residual of the to-be-processed image is obtained;

[0064] According to the noise residual and the to-be-processed image, a denoised image is obtained.

[0065] Specifically, the noise residual image (i.e., noise residual, which can also be referred to as a noise image) can be obtained according to the noise intensity feature and the to-be-processed image. For example, the noise intensity feature and the to-be-processed image are input into a neural network model (which can be referred to as a denoising network or a denoising model) to obtain the noise residual image. Then, the to-be-processed image and the noise residual image are fused to obtain the denoised image. For example, the to-be-processed image is added to the noise residual image to obtain the denoised image. It can be understood that the size of the noise residual image is the same as that of the to-be-processed image. Adding the to-be-processed image and the noise residual image means adding the element values of the same pixel points in the two images, that is, performing a pointwise add operation on the two images to obtain the pixel value at the same position in the denoised image.

[0066] In optional embodiments of the present application, the noise intensity feature of the to-be-processed image includes noise intensity features corresponding to respective channel images of the to-be-processed image.

[0067] In optional embodiments of the present application, obtaining the noise intensity feature of the to-be-processed image includes:

[0068] Obtaining respective channel images of the to-be-processed image;

[0069] Respectively obtaining noise intensity features of the respective channel images;

[0070] Splicing the noise intensity features of the respective channel images to obtain the noise intensity feature of the to-be-processed image.

[0071] That is, when obtaining the noise intensity feature of the to-be-processed image, the noise intensity feature of each channel image of the to-be-processed image can be obtained according to the channel, and the noise intensity features corresponding to the respective channels are spliced to obtain the noise intensity feature of the image.

[0072] For an image, the noise distribution of different channels is usually different. Research has found that the noise distribution of each channel has some rules, which is close to a Gaussian distribution, and the Gaussian distribution parameters of different channels are usually different, such as different variances. Based on the rules, the image can be split according to the channel, and the noise intensity feature of each channel image can be estimated respectively, so that the noise intensity of each channel of the image can be more accurately evaluated, the predicted noise intensity feature is more consistent with the actual noise distribution of the channel image, and thus the denoising process can be more targeted and accurate, improving the denoising performance.

[0073] It can be understood that for different color space modes, the channel mode of the image will also be different. For example, for an R(red, red) G(Green, green) B(Blue, blue) color mode, the respective channel images of the to-be-processed image include an image of a red channel, an image of a green channel, and an image of a blue channel.

[0074] In optional embodiments of the present application, the noise intensity feature of each channel image can include:

[0075] Based on each channel image, the noise intensity feature of the corresponding channel image is obtained by using the noise feature estimation network corresponding to each channel image.

[0076] Since the noise distribution features of different channel images are different, in order to more accurately estimate the noise intensity feature corresponding to each channel, a noise feature estimation network (also referred to as a noise intensity feature network, or a noise intensity feature estimation network, or a noise intensity feature prediction network, or a noise intensity prediction network) corresponding to each channel can be pre-trained, so that when predicting the noise intensity feature of the image to be processed, the noise intensity feature corresponding to each channel can be obtained by using the noise feature estimation network corresponding to each channel.

[0077] In optional embodiments of the present application, the noise intensity feature is used to perform denoising processing on the image to be processed, which can include:

[0078] Obtaining a luminance channel image of the image to be processed;

[0079] Obtaining a noise spatial distribution feature of the image to be processed according to the luminance channel image;

[0080] Performing denoising processing on the image to be processed according to the noise intensity feature and the noise spatial distribution feature of the image to be processed.

[0081] Optionally, the noise spatial distribution feature of the image to be processed is obtained according to the luminance channel image, which can include:

[0082] Using a noise spatial feature estimation network to determine the noise spatial distribution feature of the image to be processed according to the luminance channel image and the noise intensity feature.

[0083] Optionally, the noise intensity feature and the noise spatial distribution feature are used to perform denoising processing on the image to be processed, which can include:

[0084] Obtaining a noise residual of the image to be processed according to the noise intensity feature and the image to be processed;

[0085] Performing weighted processing on the noise residual according to the noise spatial distribution feature to obtain a weighted noise residual;

[0086] Obtaining a denoised image according to the weighted noise residual and the image to be processed.

[0087] As can be known from the foregoing description, in actual application, for the same image, the noise distribution of the image has not only the intensity distribution characteristic, but also the spatial distribution characteristic. The brightness information of different regions in the image is generally different, some regions are brighter, and some regions are darker, and the noise size of the regions with different brightness is also different. For example, a typical scene is that for a low-light image, the noise of the darker part of the image is obviously higher than that of the brighter part of the image. Therefore, in order to better realize better denoising effect for different regions, the scheme of the embodiment of the present application can predict the noise spatial distribution characteristic of the image by extracting the brightness channel (such as the L channel) of the image. When denoising the image to be processed, by simultaneously considering the noise intensity characteristic and the noise spatial characteristic in the image, the denoising of the image can be more accurately and effectively realized, for example, for the region with lower brightness in the image, a greater degree of denoising is performed according to the image intensity characteristic, for the region with higher brightness in the image, a relatively smaller degree of denoising is performed, so that the denoising is more targeted, and different denoising processing is realized for different spatial regions in the image.

[0088] In one example, the value of the noise spatial distribution characteristic of a pixel point is between 0 and 1, 1 is the maximum weight, which means that the denoising ability of the pixel point is greater, and 0.2 is a relatively small weight, which means that the denoising ability of the pixel point is relatively light.

[0089] As an optional embodiment of the present application, the noise intensity characteristic of each channel image of the image to be processed can be obtained by using the noise characteristic estimation network corresponding to each channel image, and then the noise residual corresponding to each channel can be obtained by using the image denoising model according to the noise intensity characteristic of each channel and the image to be processed. For the noise residual of each channel, the noise spatial distribution characteristic can be used to weight the noise residual of each channel to obtain the weighted noise residual corresponding to each channel. Then, when the weighted noise residual and the image to be processed are used, the channel image of each channel of the image to be processed and the noise residual of the channel can be added to obtain the denoised image corresponding to each channel.

[0090] Optionally, when the prediction of the noise spatial distribution feature of the image is implemented based on the luminance channel image of the image and the noise intensity feature of the image, the prediction can be implemented by using a noise spatial distribution feature estimation network (which can also be referred to as a noise spatial feature prediction network, or a noise spatial distribution feature prediction network, or a noise spatial distribution feature estimation network, or a noise spatial feature network).

[0091] Optionally, the noise spatial distribution feature can include the noise spatial distribution feature corresponding to each pixel point in the to-be-processed image, the noise intensity distribution information of each point in the to-be-processed image is included in the noise residual image, the noise spatial distribution feature is used as a noise processing weight feature map of the to-be-processed image, the element value of each element point in the feature map is the noise weight of the corresponding position point in the noise residual image, the noise residual image can be weighted based on the noise processing weight feature map, and a noise residual image (that is, a noise feature fused with the noise intensity distribution feature and the noise spatial distribution feature) associated with the luminance information of the image and more in line with the actual situation is obtained, and a better denoising effect image is obtained based on the weighted result and the to-be-processed image.

[0092] Optionally, if the to-be-processed image is an RGB image, the to-be-processed image can be converted from the RGB color space to the LAB color space first, and the image of the L channel after the conversion is the luminance channel image of the to-be-processed image.

[0093] In an optional embodiment of the present application, the noise intensity feature of the to-be-processed image can be obtained by a noise feature estimation model, and specifically, the noise intensity feature can be obtained by the noise intensity feature estimation model. Similarly, the noise processing weight feature of the to-be-processed image can also be obtained by a neural network model, and optionally, the luminance channel image of the to-be-processed image and the noise intensity feature are concatenated and input into the noise spatial feature estimation model to obtain the noise processing weight feature (which can also be referred to as a noise spatial feature distribution).

[0094] The noise feature estimation model is a model corresponding to a device hardware environment. That is, different noise feature estimation models can correspond to different device hardware environments. In actual application, noise feature estimation models corresponding to various device hardware environments can be trained respectively, and when image enhancement processing is needed, the noise feature information can be estimated by using the corresponding model according to the device hardware environment corresponding to the image to be processed. For example, for a smart phone, corresponding noise estimation models can be pre-trained according to different mobile phone brands, mobile phone models and the like, and then when image enhancement is performed, the noise feature can be estimated by using a model corresponding to the brand, model and the like of the mobile phone used to capture the image to be processed. Of course, it can be understood that if the hardware environments of mobile phones of different brands or models are the same or substantially the same, the same model can also correspond to them.

[0095] The noise intensity feature estimation method provided by the embodiment of the present application is mainly a prior estimation of noise intensity features based on a deep learning method, which can provide prior information for the subsequent denoising part, thereby helping to remove image noise conforming to a true distribution.

[0096] The specific model architecture of the noise feature estimation model is not limited by the embodiment of the present application, for example, a convolutional neural network-based estimation model can be used, and when the noise feature information of the image to be processed is predicted by the model, the image to be processed can be directly input into the estimation model to obtain the noise feature information corresponding to the image. Optionally, the estimation model can include noise feature estimation modules corresponding to each channel of the image, and when the image is processed, the images of each channel of the image to be processed can be input into the estimation modules corresponding to each channel respectively to obtain the noise feature information corresponding to each channel. Specifically, for example, for an image to be processed in an RGB color mode, the R, G and B channels of the input image (i.e., the image to be processed) can be input into convolutional neural networks with the same or different network structures respectively, and the noise intensity feature maps of the three channels can be output respectively. The noise intensity feature maps corresponding to the three channels can be spliced to obtain the noise intensity feature map of the image to be processed. Then, the noise intensity feature map and the image to be processed can be input into a denoising model (denoising network) to obtain the noise residual map (also referred to as noise map) of the image to be processed, that is, the noise feature.

[0097] In an optional embodiment of the present application, the noise feature estimation model can be obtained by the following method: obtaining each training sample image, wherein the training sample image carries a labeled label, and the labeled label represents the labeled noise intensity feature of the training sample image;

[0098] training the initial neural network model based on each training sample image until the loss function of the model converges, and taking the model at the end of the training as the noise feature estimation model;

[0099] The input of the initial neural network model is the training sample image, the output is the predicted noise intensity feature of the training sample image, and the value of the loss function represents the difference between the predicted noise intensity feature and the labeled noise intensity feature of the training sample image.

[0100] The above training sample image used to train the noise feature estimation model is obtained by the following method:

[0101] An initial sample image is obtained, and a reference image containing noise is collected under a device hardware environment;

[0102] Based on the reference image, the reference noise intensity feature is determined;

[0103] The reference noise intensity feature and the initial sample image are fused to obtain a training sample image, wherein the labeled noise intensity feature of the training sample image is the reference noise intensity feature.

[0104] It can be understood that the device hardware environment is a hardware environment corresponding to (the same or substantially the same as) the device hardware environment of the image collection device used to collect the image to be processed, that is, the same type of hardware environment.

[0105] The specific way of obtaining the initial sample image is not limited in the embodiments of the application, and can be an image in a public data set or a generated sample image. For example, a Gaussian white noise data set can be generated, and each image in the data set is used as an initial sample image.

[0106] The reference image can be one or more, and optionally, the reference image can usually be multiple, and the multiple reference images can be images containing different intensity noise. Based on each reference image, one or more reference noise intensity features can be obtained. When the reference noise intensity feature is fused with the initial sample image, one reference noise intensity feature can be randomly selected to be fused with the initial sample image, and the selected reference noise intensity feature is taken as the labeled noise intensity feature of the training sample image.

[0107] In an optional embodiment of the application, based on the reference image, the reference noise feature information is determined, comprising:

[0108] The reference image is filtered to obtain a filtered image;

[0109] Based on the reference image and the filtered image, a noise image is obtained;

[0110] The total variation of each image region in the noise image is used to determine at least one target image region from each image region, and pixel information of the at least one target image region is used as a reference noise intensity feature.

[0111] It can be understood that the size of the target image region is usually the same as the size of the initial sample image, so that the training sample image is obtained by fusing the pixel information of the target image region and the initial sample image, and the specific fusion manner is not limited in the embodiment of the application, for example, it can be superimposed.

[0112] When the selected target image region is multiple, the pixel information of each target image region can be used as a reference noise intensity feature, and a reference noise intensity feature can be randomly selected when the initial sample image and the reference noise intensity feature are fused.

[0113] In order to use the noise feature estimation model to learn the estimation ability of the noise intensity feature, the model needs to be supervised and trained by using the true value (that is, the labeled noise intensity feature) of the noise intensity feature related to the hardware environment system. In the above scheme provided by the embodiment of the application, the true value data (that is, the training sample image) can be obtained by using the data degradation model, that is, the initial sample image is subjected to data degradation processing to obtain the training sample image.

[0114] The above image denoising scheme provided by the embodiment of the application provides a deep learning denoising method related to the hardware environment of the image acquisition device. Here, the hardware environment related does not mean that the method is only applicable to a specific hardware environment, but means that the method can be applicable to any hardware environment. In actual application, for different hardware environments, the noise feature estimation model can be previously associated with the specific hardware environment, and once the association is established, targeted high-quality denoising can be performed. That is, the noise feature estimation model corresponding to the hardware environment to be applied can be previously trained, and when the image acquired under the hardware environment is subjected to image enhancement processing, the noise intensity feature estimation can be performed by using the model corresponding to the hardware environment.

[0115] In order to realize the noise estimation based on deep learning, two key points can be included: one is to previously evaluate the hardware related noise intensity feature (that is, the noise intensity feature reflecting the image noise distribution characteristics under the hardware environment), and the other is to establish a data degradation model related to the hardware environment system to obtain the training sample image for training the model. In order to better illustrate the denoising scheme provided by the embodiment of the application, two specific examples are further described in detail below.

[0116] For the convenience of describing the neural network structure provided in each optional example of the present application, first, some parameters that may be involved in each example are uniformly described. For the convolution layer in the network structure, it can be represented as [convk x k, c = C, s = S], where convk x k represents that the kernel (i.e., the convolution kernel) size of the convolution layer is k x k, the number of channels c is C, and the stride s is S. Optionally, a ReLu layer (activation function layer) can be followed after all convolution layers, and batch normalization is not performed. As an example, assuming that a convolution layer is represented as [conv3 x 3, c = 32, s = 1], the convolution kernel size of the convolution layer is 3 x 5, the number of output channels is 32, and the convolution stride is 1. If it is represented as [conv3 x 3, c = 32, s = 1] x n, it means that n such convolution layers are cascaded.

[0117] As an example, Figure 3a The structure diagram of an image denoising model provided by an embodiment of the present application is shown in FIG. 1. As shown in the figure, the image denoising model can include a cascaded noise feature estimation model (i.e., a noise intensity feature estimation model / network) and a denoising model (i.e., a denoising network). In this example, the noise feature estimation model includes network structures (which can be convolutional neural networks) corresponding to the R, G, and B channels of an image, respectively. The input of the noise feature estimation model is a to-be-processed image (i.e., an input image). Specifically, the R, G, and B channels of the input image are respectively input into convolutional neural networks with the same network structure (the network parameters corresponding to each channel are different, i.e., the convolution processing parameters are different), and the noise intensity features of the three channels are respectively output. The noise intensity features of the three channels and the input image are input into the denoising model, and the denoising model obtains a noise residual image based on the noise intensity features of the input image and the input image. Then, based on the noise residual image and the input image, an output image (i.e., a denoised image) is obtained. The specific network structure of the denoising model is not limited by the embodiments of the present application. Optionally, the denoising model can generally include an encoding network and a decoding network. The encoding network is used to encode the input information of the model to obtain an encoding result, and the decoding network obtains an output result (i.e., a noise residual image) based on the encoding result.

[0118] It can be understood that, in actual application, for the noise feature estimation model Figure 3a In the example shown in FIG. 1, the output image obtained based on the noise residual image and the input image can be implemented inside the denoising model shown in the figure, or can be implemented outside the denoising model shown in the figure.

[0119] As an example, Figure 3b The flow diagram of an image denoising method provided by an embodiment of the present application is shown in FIG. 2. As shown in FIG. 2, in this example, the input image (i.e., the to-be-processed image) is an RGB image, Figure 3b Figure 3b ​In the network A in the figure is a noise feature estimation network, the R, G, B three channels each correspond to a respective noise feature estimation network, the network B is a denoising model, and the network C is a noise spatial distribution feature estimation network. When denoising the image to be processed, the R, G, B channel images are respectively input into the respective network A corresponding to each channel to obtain the noise intensity features of the three channels. Then, the noise intensity features of the three channels and the input image after splicing are input into the network B to obtain a noise map. In order to obtain the noise spatial distribution feature of the input image, first, the RGB image can be converted into an LAB image (RGB→LAB shown in the figure), the luminance channel image, i.e., the L channel image, and the spliced noise intensity features are input into the network C to obtain a noise spatial distribution feature map (noise spatial feature shown in the figure). Then, a pointwise production operation is performed on the noise spatial feature map and the noise map to obtain a weighted noise map. By performing an element-wise addition operation on the weighted noise map and the input image, a denoised image, i.e., an output image, is obtained.

[0120] As another example, Figure 3c A structure diagram of an image denoising model provided by an embodiment of the present application is shown in the figure. As shown in the figure, the image denoising model can include a noise intensity feature estimation model, a noise spatial estimation model, and a denoising model. The parts of the image denoising model in this example are described below.

[0121] Noise intensity feature estimation model: optionally, in order to reduce the complexity of the model and realize lightweight of the model, the part can be independently processed by the R channel, the G channel, and the B channel through the same model composed of five convolutional layers, that is, the same model structure can be used to extract the noise intensity features corresponding to the three channels. Optionally, the five convolutional layers can be represented as: [conv3×3, c=32, s=1]×4; [conv3×3, c=1, s=1]. The noise intensity features of the R channel, the G channel, and the B channel are extracted through the five convolutional layers respectively to obtain the noise intensity feature maps of the three channels. The three independent results (i.e., three noise intensity feature maps) are concatenated as the output of the noise intensity feature estimation model.

[0122] Noise spatial feature estimation model, i.e., noise spatial distribution feature estimation network: this part concatenates the L channel (converted from RGB to LAB space) of the input image and the output of the noise intensity feature estimation model above as input. Optionally, the concatenation result can also be processed by a five-layer convolutional network: such as: [conv3×3, c=32, s=1]×4 and [conv3×3, c=1, s=1], and the final output is represented as w, w is the noise spatial distribution feature, that is, the noise processing weight feature.

[0123] Optionally, the network structure of the noise intensity feature estimation model and the noise spatial feature estimation model provided in the embodiments of the present application can be implemented by a simple convolutional network, so as to realize the lightweight of the model structure, and make the scheme of the embodiments of the present application better applicable to mobile terminal devices such as smart phones.

[0124] Denoising model (i.e. noise removing model, or denoising net): This part takes the output of the noise intensity feature estimation model as input. Optionally, the encoder part of the model can be composed of a series of interleaved convolution-pooling layers: [conv3x3, c=16, s=1]x2; [maxpooling]; [conv3x3, c=32, s=1]x2; [maxpooling]; [conv3x3, c=64, s=1]x3; [maxpooling]; [conv3x3, c=128, s=1]x6. In the representation of the above convolution-pooling layers, the description order of the convolution layers and the pooling layer (maxpooling) represents the cascade order between the layers, for example, in the case of [conv3x3, c=16, s=1]x2; [maxpooling], it represents two cascaded [conv3x3, c=16, s=1] convolution layers followed by a maximum pooling layer (other pooling layer structures such as mean pooling can also be used).

[0125] For the decoder part of the denoising model, the decoder part can be composed of the following order layers: [upsample] (i.e. upsample, sampling rate is 2); [conv3x3, c=64, s=1]x4; [upsample] (sampling rate is 2); [conv3x3, c=32, s=1]x3; [upsample] (sampling rate is 2); [conv3x3, c=16, s=1]x2; [conv3x3, c=3, s=1].

[0126] In addition, in order to accelerate the model training, the model structure of this part can adopt a structure similar to UNet to fuse the output feature maps of the corresponding encoder and the output feature maps of the decoder, and take the fused feature maps as the extraction of the decoding features of the next level (corresponding to the arrow shown in the figure).

[0127] As an optional solution, Figure 3d A structure diagram of a denoising model provided in the embodiments of the present application is shown in FIG. 2. The network parameters shown in the figure are the parameters of each hidden layer of the model, which can be referred to in Figure 3cThe parameters in the code, taking 3×3,conv,16 as an example, represent a convolutional layer with a kernel size of 3×3 and 16 output channels. ×2upsample indicates upsampling with a sampling rate of 2. Figure 3d As can be seen, the network structure of this denoising model is similar to that of UNet, but the difference lies in that it uses fewer channels and smaller convolutional kernels in each convolutional layer. In the feature fusion process between the encoder and decoder, addition is used instead of concatenation, as shown by the arrows with plus signs in the figure, indicating that the corresponding encoded and decoded features are added. Using addition instead of concatenation speeds up forward inference while still achieving the fusion of shallow and deep features, thus improving the model's data processing efficiency. The encoder output consists of three noise maps corresponding to the RGB channels. Figure 3c (The noise residual is shown in the figure).

[0128] Let the final output of the decoder be R. Therefore, the final image denoising result can be expressed as:

[0129] y=x+R*w (1)

[0130] Where x and y represent the input image and the output image, respectively, i.e., the image to be processed and the image after denoising, and R and w are the noise residual and the noise processing weight feature, respectively.

[0131] The following examples further illustrate the optional schemes for obtaining training sample images provided in the embodiments of this application.

[0132] In this example, a degradation model can be built to obtain training sample images for training the noise feature estimation model. The main purpose of building the data degradation model is to simulate the distribution of hardware-related noise data. The specific data degradation process is described as follows:

[0133] (1) Obtain the image noise pattern under a defined hardware environment. The noise pattern, also known as the noise distribution feature, can be understood as a sampling of the noise distribution that depends on the hardware environment. The sampling has a certain degree of typicality and universality. Determining the hardware environment means that the image denoising process must be applied to the specific hardware environment corresponding to the image. For example, if the image enhancement processing method is applied to the processing of an image taken by a certain model of mobile phone, then the specific hardware environment can be the hardware environment of that model of mobile phone.

[0134] In a specific hardware environment, an image I (i.e., a reference image) including different noise levels can be obtained, and a mean filtering process can be performed on the image I, such as a mean convolution operation f with a fixed kernel size. The kernel size of the convolution kernel can be configured according to actual requirements, such as a typical kernel size of 11x11. The convolution result of the image I is denoted as f(I) (i.e., a filtered image). Based on the image I and f(I), a noise image P can be obtained. Optionally, P = I-f(I)+128, that is, the pixel value of a pixel point at a corresponding position on the noise image P is obtained by subtracting the pixel values of the pixel points at the corresponding positions on the image I and the image f(I) and then adding 128.

[0135] After obtaining the noise image P, a plurality of image regions can be cropped from P, and a plurality of regions with a total variation (also referred to as a total variation) less than a certain threshold t in the plurality of cropped image regions are taken as target image regions (denoted as P'), which can be used as candidate noise patterns. The total variation is defined as the integral of the gradient amplitude and can be represented as:

[0136]

[0137] wherein P' satisfies:

[0138]

[0139]

[0140] wherein, D u is a support domain of the image, that is, the cropped image region. u represents the image region (the image can be understood as a two-dimensional (x direction and y direction) function), and u x represents the gradient of any pixel point in D u in the x direction, and u y represents the gradient of any pixel point in D u in the y direction.

[0141] tb(P') represents the total variation of P', and P' x is the gradient of any pixel point in the target image region in the x direction, and P' y is the gradient of any pixel point in the target image region in the y direction.

[0142] (2) Generate a Gaussian white noise data set or directly introduce a public data set. The generated noise pattern in step (1) can be randomly superimposed on the noise data set to obtain a training sample image data set. Specifically, as a scheme, the training sample image can be obtained in the following manner:

[0143] I new_noise = n · I noise + I noise

[0144] n = δ · P' / 255

[0145] Wherein, n is the noise level, δ is the superimposed noise intensity, which is a random quantity in a specified range, I noise is the introduced noise data set, I new_noise is the degraded noise data set. Specifically, I noise can be understood as any image in the noise data set (that is, the initial sample image), I new_noise is the corresponding training sample image, based on the pixel value of each pixel point in P', through n = δ · P' / 255, the noise level corresponding to each pixel point can be calculated, P' and I noise have the same image size, by multiplying the noise level corresponding to each pixel point in P' with the pixel value of the pixel point at the corresponding position in I noise , and then adding the pixel value of the pixel point in I noise , the pixel value of the pixel point at the position in I new_noise can be obtained.

[0146] In this example, for each training sample image I new_noise , the corresponding true noise intensity feature information is the true noise feature distribution P' (that is, the label noise intensity feature), based on each training sample image, the initial neural network model can be trained to make the noise intensity feature of the training sample image output by the model approximate to its true noise intensity feature, and obtain a noise feature estimation model.

[0147] For image denoising, it can be specifically implemented by a pre-trained image denoising model. In an optional embodiment of the present application, for training of the noise denoising model, in order to obtain good training data, the above noise distribution features of different channels in the image can be used to synthesize training data similar to the actual noise distribution, for example, for a clean image not containing noise, the parameters (such as mean and variance, etc.) of the Gaussian distribution of each channel can be simulated to perform noise processing on the clean image to obtain an image containing noise distribution, and the model is trained based on the image containing noise and the corresponding clean image. Based on this method, the training process of the model can be more controllable, and the model is more likely to converge.

[0148] Specifically, the training data can include pairs of sample images, each pair of sample images including a clean image not containing noise and a noisy image obtained by adding noise to the clean image, wherein the noisy image is obtained by adding noise to each channel of the clean image according to the noise distribution characteristics of each channel, for example, if the clean image is an RGB image, for the R channel, the parameters of the Gaussian noise distribution of this channel can be fitted to synthesize the noisy image corresponding to the R channel.

[0149] When training the image denoising model based on the training data, as an optional manner, a training manner based on multi-scale Gram loss can be used to better maintain the texture details of the image (the texture details can be maintained at different scales). During training, the noisy image in the pair of sample images is input into the image denoising model to obtain a denoised image (referred to as an output image). For the clean image and the output image, different sizes can be cropped to obtain a plurality of size pairs of cropped images (cropped images of the clean image and the output image), and each size of the cropped image can be adjusted to the resolution of the original image, that is, the image is resized to obtain an image with the same size as the original size (the size of the clean image and the noisy image). For each size of the cropped image pair, the corresponding Gram loss is calculated. For a pair of sample images, the corresponding Gram loss, i.e., Loss, can be L1 loss, which can be expressed as follows:

[0150]

[0151] wherein n represents the number of different sizes of the cropped images, i represents the i-th size of the image pair, real_scale_i represents the cropped image of the clean image of the i-th size, predict_scale_i represents the cropped image of the output image of the i-th size, and |Gram real_scale_i -Gram predict_scale_i represents the L1 loss between real_scale_i and predict_scale_i.

[0152] By adding the corresponding loss of each pair of sample images, the total loss corresponding to the model is obtained, and the model is iteratively trained based on the total loss and the training data, so that the denoised image output by the model approaches the clean image.

[0153] As an example, Figure 3eAn optional training principle diagram of an image denoising model is shown in FIG. 6. As shown in the figure, during training, a clean image without noise is input, and after the image degradation processing described in the foregoing is performed on the clean image by using a degradation model, a noisy image (i.e., a training sample image) is obtained. The image is input into the image denoising model to obtain an output image after denoising. Based on the input image and the output image, a training loss (loss shown in the figure) can be calculated. Optionally, the training loss can be the multi-scale Gram loss described in the foregoing. The image denoising model is iteratively trained based on the training loss, and the model parameters are optimized until the training loss converges, and a trained image denoising model is obtained. It can be understood that the manner of obtaining the noisy image can include but is not limited to the manner of processing by using the degradation model, and can also be to perform noise processing on the clean image according to the noise distribution characteristics (such as Gaussian distribution parameters) of each channel of the image.

[0154] As an example, the loss function during training of the denoising model can be an L2 loss, i.e., the value of the loss function = minimize (ground truth-output) 2 . Wherein, minimize means minimization, ground truth refers to the clean image without noise of the training sample image, output refers to the result image obtained after the image containing noise in the training sample image is processed by using the denoising model, i.e., the image after denoising, and (ground truth-output) 2 is the L2 loss between the clean image in the sample image and the image after denoising.

[0155] The image denoising method provided in the embodiments of the present application can predict the noise intensity characteristics corresponding to each channel respectively according to the different characteristics of each channel of the image, so that the denoising model can predict the noise residual corresponding to each channel according to the noise intensity characteristics corresponding to each channel. In addition, considering the different spatial distribution of image noise (the noise distribution of the relatively bright pixel points in the image (such as the points near the light source in the image) and the relatively dark pixel points in the image is different), the noise spatial distribution characteristics of the image are predicted by using a noise spatial feature estimation model, the noise residuals of each channel are weighted by using the spatial distribution characteristics, the noise residuals more consistent with the actual noise distribution are obtained, and the output image with obviously improved denoising effect is obtained based on the weighted residuals.

[0156] As an illustrative description, assuming that the image to be processed is an image taken at night, there is a light source in the image, A point is a point near the light source in the image, and B point is a point with weak light in the input image. Based on the scheme provided in the embodiment of the present application, in the predicted noise residual image (i.e. noise image), the noise value of the point corresponding to A is a small value, and the noise value of the point corresponding to B is a large value. In the noise space feature map, the noise space feature value of the point corresponding to A has a small value, which tends to 0, and the noise space feature value of the point corresponding to B has a large value, which tends to 1. Therefore, after the noise image is weighted by the noise space feature map, the weighted value corresponding to A point is very small, indicating that the noise corresponding to A point is weak, and the weighted value corresponding to B point is relatively large, indicating that the noise corresponding to B point is strong. Therefore, the scheme of the embodiment of the present application can realize targeted denoising processing of different parts of the image, and has good denoising effect.

[0157] In an optional embodiment of the present application, the at least one image quality enhancement manner described above includes image brightness adjustment, and the quality of the image to be processed is enhanced by using at least one image quality enhancement manner, including:

[0158] The brightness enhancement parameter of the image to be processed is determined.

[0159] The brightness channel image is adjusted in brightness based on the brightness enhancement parameter.

[0160] It can be understood that for different color space modes, the name of the brightness channel of the image may also be different. For example, for YUV mode, the brightness channel of the image is Y channel, and for LAB mode, the brightness channel of the image is L channel. Of course, conversion can also be performed between different modes.

[0161] For an image, improving the color tone (including brightness and / or color) of the image, especially for the dark area of the image, can most directly improve the user's shooting experience, because the visual impact of brightness and / or color is much greater than that of noise and texture details. Therefore, the adjustment of the color tone of the image is also one of the common ways to improve the image quality, and is very important for image enhancement.

[0162] In the existing image tone adjustment manner, a main trend is to provide a global color enhancement effect, which can make the image more informative, but at the same time, can also introduce the effects of over-saturation and over-exposure. In order to alleviate these problems, the scheme provided by the embodiments of the present application makes a more in-depth study on the color (i.e. color) details and the brightness details of the image, and proposes a two-branch network that can respond to brightness enhancement and color enhancement respectively. That is, in the optional embodiments of the present application, when the image tone is adjusted, the image brightness adjustment and the image color adjustment are taken as two independent processing manners, that is, the brightness adjustment and the color adjustment are separated into two independent tasks instead of being taken as one task. By splitting the two, a better processing effect can be achieved, because: when the brightness adjustment and the color adjustment are not separated, it is a complex composite task, and by separating the brightness information and the color information, only a single task needs to be considered when the brightness information or the color information is processed, which simplifies the task. By separating the brightness information and the color information, the processing result can be more natural when the color adjustment is performed, and the problem of over-enhancement of brightness can be reduced to a certain extent when the brightness adjustment is performed. Furthermore, by separating the two, the brightness adjustment and the color adjustment are taken as a task respectively, and then the processing manner of each task can be adjusted according to the characteristics and actual requirements of the task, so as to achieve a better processing effect.

[0163] In actual application, for different to-be-processed images, the brightness information of different images can be different. If the same brightness adjustment processing manner is used for all images, over-brightness phenomenon can exist for some images. In addition, some images can not need to be processed by brightness enhancement. If the brightness enhancement is still performed, not only resources and time can be wasted, but also the image quality can be reduced. In order to solve these problems, in the brightness adjustment scheme provided by the embodiments of the present application, a brightness enhancement parameter is introduced as an enhancement strength control parameter. For to-be-processed images with different brightness information, different brightness enhancement parameters can be correspondingly set, so as to achieve the purpose of adjusting the brightness of different images to different degrees, so as to better meet the actual requirements and improve the image processing effect.

[0164] In the embodiments of the present application, the specific value form of the brightness enhancement parameter is not limited. Optionally, the parameter can be a value in a set range, such as an integer not less than 1. The greater the value of the parameter, the greater the brightness enhancement strength can be. When the value of the parameter is 1, the brightness enhancement processing can not be performed.

[0165] In the optional embodiments of the present application, determining the brightness enhancement parameter of the to-be-processed image can include at least one of the following:

[0166] Obtaining luminance information of the image to be processed, and determining a luminance enhancement parameter based on the luminance information;

[0167] Obtaining luminance adjustment indication information input by a user, and determining the luminance enhancement parameter based on the indication information.

[0168] Optionally, the luminance enhancement parameter determined based on the indication information can be a luminance enhancement parameter of each pixel point in the image to be processed.

[0169] That is, the luminance enhancement parameter can be determined according to the luminance information of the image to be processed, or can be determined according to the indication information of the user, for example, the user can input the indication information of the luminance adjustment according to his own needs, and the final luminance enhancement parameter can be the parameter value corresponding to the indication information.

[0170] Optionally, a mapping relationship between the luminance information and the luminance enhancement coefficient can be configured, and for the image to be processed, the corresponding luminance enhancement parameter can be determined based on the luminance information of the image and the mapping relationship. When the luminance parameter is determined based on the indication information of the user, a corresponding determination strategy can be configured according to different needs, for example, the user can be provided with a range of values of the luminance enhancement parameter, and the user can directly determine a value in the range as the luminance enhancement parameter of the image to be processed, that is, the indication information can be directly the value of the parameter.

[0171] In actual application, if the user inputs the indication information, the luminance enhancement parameter can be determined based on the user indication information, or the luminance enhancement parameter can be determined based on the user indication information and the luminance information of the image to be processed; if the user does not give the indication, the luminance enhancement parameter can be determined by the device according to the luminance information of the image to be processed.

[0172] In addition, it should be noted that for the image to be processed, the image can correspond to a value of the luminance enhancement parameter, or each pixel point in the image can correspond to a value of the luminance enhancement parameter, or the image can be divided into regions, and each region can correspond to a value of the luminance enhancement parameter. That is, the luminance enhancement parameter can be one value or multiple values. For example, one luminance enhancement parameter of the image to be processed can be determined according to the average luminance information of the image to be processed, or the value of the luminance enhancement parameter of each pixel point can be determined according to the luminance value of the pixel point.

[0173] As an optional mode of the present application, the luminance enhancement parameter map of the image to be processed can be predicted through a luminance adjustment intensity prediction network according to the image to be processed, and the element value of each element point in the luminance parameter map is the luminance enhancement parameter of the corresponding pixel point in the image to be processed.

[0174] That is, the luminance enhancement parameters corresponding to each pixel in the to-be-processed image can be predicted by the neural network model, so as to realize more detailed luminance adjustment processing of the image.

[0175] Optionally, determining the luminance enhancement parameters of the to-be-processed image can include:

[0176] Obtaining a luminance channel image of the to-be-processed image;

[0177] Based on the luminance channel image, obtaining global luminance information of the to-be-processed image and local luminance information of the to-be-processed image;

[0178] Based on the global luminance information and the local luminance information, determining the luminance enhancement parameters of each pixel of the to-be-processed image.

[0179] Optionally, based on the luminance channel image, obtaining the local luminance information of the to-be-processed image includes:

[0180] Based on the luminance channel image, using a local luminance estimation network to estimate the semantic-related local luminance information.

[0181] Optionally, based on the luminance enhancement parameters, performing luminance adjustment on the to-be-processed image includes:

[0182] Based on the luminance enhancement parameters and the luminance channel image of the to-be-processed image, using a luminance enhancement network to perform luminance adjustment on the to-be-processed image.

[0183] The optional scheme provided in the present application considers the global luminance information and the local luminance information of the image. The global luminance information can guide whether to perform luminance enhancement on the image and the overall enhancement strength, that is, the luminance adjustment mode is considered from a coarse granularity (the entire image), and the local luminance information is more detailed consideration of the local luminance information of each region and each pixel in the image, which guides the luminance adjustment of each region and each pixel in the image from a more detailed granularity.

[0184] The global luminance information can be a global luminance statistical value of the to-be-processed image, such as a global luminance mean value, that is, an average value of the luminance values of all pixels in the image, or a luminance value obtained by further processing the luminance mean value, such as processing the mean value by a pre-configured function and taking the processed value as the global luminance information. The specific processing mode can be processed according to experimental results or experience, so that the luminance enhanced image is more consistent with the actual perception of people.

[0185] For the local brightness information of the to-be-processed image, optionally, the brightness value of each pixel point in the to-be-processed image can be the local brightness information of the to-be-processed image. As another optional solution, considering the semantic correlation between each pixel point in the image (such as whether the pixel points are pixel points of the same object or whether the pixel points have other semantic correlations), the local brightness information of the to-be-processed image can be obtained through a neural network (i.e., a local brightness estimation network), such as inputting the brightness channel image of the to-be-processed image into the encoder part of the neural network, extracting the encoding features of the image through the part, and then decoding the encoding features through the decoder part of the neural network to obtain a feature map with the same size as the to-be-processed image. The element value of each element point in the feature map is the local brightness information of the pixel point at the corresponding position of the to-be-processed image. Optionally, the global brightness information and the local brightness information can be fused, such as multiplying the global brightness information and the local brightness information pixel by pixel to obtain a brightness enhancement parameter map (i.e., the brightness enhancement parameter of each pixel point). Then, the brightness channel image is subjected to brightness enhancement processing based on the brightness enhancement parameter map, for example, the brightness enhancement parameter map and the brightness channel image are input into a brightness enhancement network, and the brightness channel image is enhanced based on the brightness enhancement parameter map through the network to obtain an image after brightness enhancement.

[0186] Based on the optional solution, the global brightness information can control the overall brightness adjustment through statistical and predefined functions, and the local brightness information can control the local adjustment. According to the global brightness information and the local brightness information, the brightness adjustment can be better controlled to avoid overexposure.

[0187] In an optional embodiment of the present application, the at least one image quality enhancement manner described above includes image color adjustment, and the at least one image quality enhancement manner is used to enhance the quality of the to-be-processed image, including:

[0188] Obtaining a color channel image of the to-be-processed image;

[0189] Performing resolution reduction processing on the color channel image;

[0190] Performing color adjustment on the color channel image after the resolution reduction.

[0191] In actual application, considering that the sensitivity of the human eye to color details is weak, in order to reduce the resource consumption of the device and improve the image processing efficiency, when the image is subjected to color adjustment, the resolution of the image can be reduced first, the image after the resolution reduction is subjected to color adjustment processing, and then the resolution of the processed image is increased to the resolution before the reduction.

[0192] Understandably, when adjusting the brightness and color of an image, one can first separate the brightness component (the image of the brightness channel (Y channel)) and the color component (the image of the color channel (UV channel)). The brightness component is adjusted, and the color component is adjusted. After the processing is complete, the two parts of the image are merged to obtain the image after brightness and color adjustment.

[0193] It will be clear to those skilled in the art that brightness adjustment and / or color adjustment can be implemented directly using a neural network model. Specifically, brightness adjustment can be implemented using a brightness adjustment model, and color adjustment can be implemented using a color adjustment model.

[0194] To better illustrate the brightness adjustment and color adjustment schemes and effects provided in the embodiments of this application, the optional schemes will be further explained below with reference to examples.

[0195] As an example, Figure 4a The diagram illustrates the principle of an image tone (including brightness and color information) adjustment process provided in this application embodiment. The image tone enhancement process in this example can be implemented using an image tone adjustment model, such as... Figure 4a As shown, the image tone adjustment model includes two branches: a brightness adjustment branch (i.e., a brightness enhancement network) and a color adjustment branch, namely, a brightness adjustment model and a color adjustment model. The specific model architecture of the brightness adjustment model and the color adjustment model is not limited in this embodiment.

[0196] Optionally, the two branches can use the same or different encoder-decoder structures, with the UV (color information) channel and Y (luminance information) channel as inputs, respectively. Specifically, for the image to be processed (i.e., the input image), the UV channel image can be input into the color adjustment model to obtain the color-adjusted image, and the Y channel image can be input into the luminance adjustment model to obtain the luminance-adjusted image. Then, the images processed by these two branches are fused to obtain the output image. As described above, separating the luminance and color components has the following advantages:

[0197] 1. The complex task of brightness and color adjustment has been split into two separate tasks, which simplifies the task, maintains contrast and saturation, makes the colors of the processed result more natural, and reduces the problem of excessive brightness enhancement to a certain extent.

[0198] 2. The dual-branch strategy allows for the adjustment of sub-modules according to task characteristics and actual needs. That is, the brightness adjustment model and the color adjustment model can be adjusted separately according to their characteristics.

[0199] For example, the luminance branch adds a parameter for adjusting the enhancement strength, i.e., a luminance enhancement parameter. According to the actual requirements and characteristics of luminance enhancement processing, when training the luminance adjustment model, both dark-bright image pairs (i.e., an image pair composed of one image with lower luminance and one image with higher luminance) and bright-bright image pairs can be used as samples for training, which effectively reduces the problem of over-enhancement of bright areas. It should be noted that the two images contained in the above image pair have the same image content, and the difference lies in the luminance of the image. When training the model, for a dark-bright image pair, the image with lower luminance is input into the model, and the model outputs the image after luminance enhancement. The training loss corresponding to the image pair is calculated by the difference between the output image and the image with higher luminance. The color branch considers that the sensitivity of the human eye to color details is weak, so the UV channel can be processed with reduced resolution to improve speed and reduce memory consumption.

[0200] For the dual-branch strategy, the luminance adjustment part and the color adjustment part can be trained separately using different data sets to reduce the training cost and improve the training efficiency. For example, the low-light image data set (such as the SID (See-in-the-Dark, low-light) data set) includes many bright-dark image pairs, but often lacks color saturation. The luminance branch can be trained using this data set. Many existing image enhancement data sets include high-definition images collected during the day with strong color saturation, but lack night images. The color branch can be trained using this type of data set.

[0201] For the luminance adjustment branch, an enhancement strength control parameter, i.e., a luminance enhancement parameter, is introduced. This parameter has the following two functions: first, the model can adaptively control the luminance enhancement strength according to the scene luminance of the input image, i.e., determine the value of the parameter according to the luminance information of the image to be processed; second, the parameter can be manually adjusted to meet the specific needs of the user for luminance, i.e., determine the value of the parameter according to the user input indication information.

[0202] To achieve the above-mentioned variable luminance enhancement parameter, the following mechanisms can be adopted when training the luminance adjustment model and processing the image to be processed using the model:

[0203] In training, the average brightness ratio of the bright image and the low-light image (i.e. the image with higher brightness and the image with lower brightness contained in the above-mentioned dark-bright image pair) can be taken as the value of the parameter (of course, the value of the parameter corresponding to the image pair can also be set artificially), which enables the network model to make different enhancement responses to the size of the parameter, but at this time, the over-brightening phenomenon may still exist. Therefore, a bright-bright image pair can also be introduced for training, at this time, the value of the parameter corresponding to the bright-bright image pair can be set to 1, and this strategy adds an explicit constraint to the network model, that is, when the value of the parameter is 1, the model should not enhance the brightness.

[0204] In inference, that is, when the trained model is used for processing the to-be-processed image, in cooperation with the training mechanism, a segmented function of the average brightness of the input image can be designed to determine the value of the enhancement parameter of the current image (i.e. the to-be-processed image), that is, the mapping relationship between the brightness and the parameter value can be configured, different brightness values or different brightness ranges can correspond to different parameter values, and for a to-be-processed image, the current value of the brightness enhancement parameter in image processing can be determined based on the average brightness of the image and the mapping relationship.

[0205] As another example, Figure 4b The results and working principle of an image tone adjustment model provided in an optional embodiment of the present application are shown in the figure, as shown in the figure, the image tone adjustment model includes a brightness adjustment branch (the brightness branch shown in the figure) and a color adjustment branch (the color branch shown in the figure), the brightness branch aims to make the image brighter, and the color branch is committed to adjusting the color saturation. As shown in the figure, in the optional embodiment, the input part of the brightness adjustment branch includes not only the brightness channel image (i.e. the Y channel image), but also the guide map sub-network shown in the figure, the input of the network is the brightness channel image, and the output is the brightness enhancement parameter feature map of the image. Among them, the specific network structure of the brightness adjustment branch and the color adjustment branch is not limited in the embodiments of the present application, for example, the structures of the two branches can both adopt the structure similar to UNet (which can be referred to the description in the foregoing text).

[0206] In order to achieve a good balance between speed and accuracy, as an optional solution, the network structure parts in the example of the present application can adopt the network structure described below. The parts of the image tone adjustment model are described below respectively.

[0207] Color branch (UV channel): The model of this branch can include an encoder and a decoder. Optionally, the encoder for feature extraction can be composed of six convolutional layers: [conv3x3, c=4, s=2]x4; [conv3x3, c=8, s=1]x2. Therefore, the encoder generates features with a stride of 16. The decoder adopts a structure similar to UNet and can include three up-sampling convolutional layers [conv3x3, c=12, s=1]x3 and a pointwise convolutional layer [conv1x1, c=2, s=1]. It should be noted that each up-sampling convolutional layer in the decoder is followed by an up-sampling operation. Then, the up-sampling output connected together with the encoder feature maps of the same spatial size (i.e., consistent with the feature map size) will constitute the input of the next decoding layer. The final output of the encoder is the image after color adjustment.

[0208] Guidance map subnetwork: A guidance map subnetwork can be adopted to adaptively obtain a pixel-by-pixel luminance control parameter, i.e., a luminance enhancement parameter, by using two branches of local luminance semantics and global luminance prior.

[0209] Luminance branch (Y channel): The luminance branch can have the same network structure as the color branch, except that the color branch uses [conv3x3, c=4, s=2] in the first convolutional layer, while the luminance branch uses [conv3x3, c=4, s=1] with a stride of 1, which can achieve more refined extraction of image luminance features. The input of this branch includes the output of the guidance map subnetwork and the Y channel image, and the output is an image after luminance enhancement processing.

[0210] Then, the output images of the luminance branch and the color branch are fused to obtain an image after hue adjustment.

[0211] As an example, Figure 4c A structure diagram of a guidance map subnetwork provided by an embodiment of the present application is shown in FIG. 2. As shown in FIG. 2, the guidance map subnetwork can include two branches of local luminance semantics and global luminance prior. Figure 4cAs shown in the figure, the sub-network includes two branches, the upper branch is a global brightness information processing branch (may be referred to as a global branch), and the lower branch is a local brightness information processing branch (may be referred to as a local branch, i.e., a local brightness estimation network). Specifically, when the sub-network is used to obtain the brightness enhancement parameter feature map of the image, the inputs of the two branches of the sub-network are both Y channel images. For the global branch, the global brightness mean of the image can be obtained first, which corresponds to the global statistics shown in the figure. After the global brightness mean is processed by a pre-configured function, the global guidance (i.e., the processed global brightness information, which is essentially a global brightness adjustment indication value) is obtained. As shown in the figure, it is a global brightness information map, and the size of the map is the size of the image to be processed. Each element value in the map is the same value, i.e., the value of the mean after being processed according to the function. For the local branch, the local guidance (i.e., the local brightness information, which is essentially a brightness adjustment indication value of each pixel) can be obtained by a neural network. Then, the global guidance and the local guidance are multiplied to obtain the brightness adjustment guidance map of the image, i.e., the brightness enhancement parameter map.

[0212] As an optional solution, the pre-configured parameter can be as follows:

[0213] f(x) = 1 + Relu(0.4 - mean) x 0.25

[0214] wherein mean represents the brightness mean (the normalized brightness mean) of all pixel points of the Y channel image, Relu() represents an activation function, the value range of Relu(0.4 - mean) is [0-1], and f(x) represents the global guidance. The greater the value of f(x) is, the greater the overall brightness enhancement intensity of the image will be.

[0215] As an example, Figure 4d As shown in the figure, the structure of the sub-network of the guidance map provided in the embodiment of the present application is shown. As shown in the figure, the local brightness estimation network of the sub-network can adopt a neural network based on a convolution structure. Through the network, brightness adjustment information related to the semantics of the image can be predicted, for example, Figure 4cIn the image shown in the figure, the image area of the light source can not be adjusted in brightness or slightly adjusted in brightness, and the input of the network can also be used for local guidance of image brightness adjustment. The neural network can also adopt a structure similar to UNet, as shown in the figure. The encoder part of the network can include six convolutional layers: [conv3x3, c=4, s=1] (i.e., 3x3, conv4, s=1 shown in the figure), [conv3x3, c=4, s=2]x3, and [conv3x3, c=8, s=1]x2; the decoder can include three up-sampling convolutional layers [conv3x3, c=12]x3 and one pointwise convolutional layer [conv1x1, c=1], and the output of the decoder is the local brightness information map shown in the figure. The figure shows that the encoded features and the decoded features of the corresponding layer are spliced. For the global branch, the global brightness information can be obtained by global statistics and a preconfigured function, and the global information is used for brightness adjustment of the entire image and is a global guide for brightness adjustment. Based on the global guide and the local guide, the brightness of each pixel in the image can be accurately controlled, thereby solving or reducing the problem of overexposure or underexposure and providing the performance of brightness adjustment.

[0216] Since the attention points of the brightness channel image (such as the Y channel image) and the color channel image (such as the UV channel image) of the image are different, the brightness channel focuses on the brightness information of the image, and the color channel focuses on the color information of the image, the brightness channel and the color channel can be processed separately to avoid the influence between the color channel and the brightness channel, and the brightness adjustment and the color adjustment can be more targeted.

[0217] In addition, the semantic information of the brightness branch and the color branch is different. The color adjustment can only need local semantic information and can not need to pay attention to global semantic information. For example, the red of the flower in the image has nothing to do with the blue of the sky in the image, so the color branch can be implemented by using a shallow neural network (such as a shallow type Unet structure). On the contrary, the brightness adjustment needs to pay attention to the global semantic information (such as the image is taken at night and the brightness is low) and the local semantic information (such as there is a light source in the image with low brightness), so the brightness adjustment needs to be more careful. The brightness adjustment scheme based on the global brightness information and the local brightness information provided in the embodiments of the present application can achieve a good image brightness adjustment effect.

[0218] In an optional embodiment of the present application, the at least one image quality enhancement manner described above includes image texture enhancement, and the quality of the to-be-processed image is enhanced by using at least one image quality enhancement manner, including:

[0219] The texture enhancement residual and the noise suppression residual of the image to be processed are obtained by using the image texture enhancement network, and the texture enhancement residual and the noise suppression residual are fused to obtain a texture residual.

[0220] According to the texture residual and the image to be processed, an image with enhanced texture is obtained.

[0221] The target of image texture enhancement is to enhance the texture details of the input image (i.e. the image to be processed). The main difficulties of image texture enhancement are as follows. First, how to correctly distinguish noise and small texture, and how to ensure or avoid the amplification of noise while enhancing the texture. Second, overshoot and noise boosting are prone to occur in the process of texture enhancement, because it is difficult to enhance the strong edges and weak edges in different degrees in the same model when a neural network model is used for image texture enhancement.

[0222] To solve the above problems in the image texture enhancement process, the optional solution provided in the embodiments of the present application can first obtain the texture enhancement residual image corresponding to the image to be processed based on the image to be processed, and obtain the noise suppression residual image corresponding to the image to be processed based on the image to be processed. Specifically, the texture enhancement residual image and the noise suppression residual image of the image can be obtained by using the pre-trained neural network (i.e. the image texture enhancement network) based on the image to be processed. For example, the texture enhancement residual image can be obtained by using the first convolution processing module based on the image to be processed, and the noise suppression residual image can be obtained by using the second convolution processing module based on the image to be processed. That is, the image to be processed is processed from two aspects of texture enhancement and noise suppression (noise suppression). By obtaining the noise suppression residual image and the texture enhancement residual image of the image to be processed, the noise suppression residual image can be subtracted from the texture enhancement residual image, and the processed image can be obtained by superimposing the difference value (i.e. the texture residual image used for texture enhancement processing) and the image to be processed. Since the noise suppression residual is removed based on the texture enhancement residual, the amplification of noise can be avoided while the small texture is enhanced, and better texture enhancement effect can be obtained. In addition, in actual application, the noise information of the image is usually different for different texture regions of the image, such as the strong edge region and the weak edge region. Therefore, by obtaining the difference value of the texture enhancement residual image and the noise suppression residual image of the image, the texture residual of the strong edge region can be weakened, and the texture residual of the weak edge region can be enhanced, so that the problems of excessive texture enhancement of the strong edge region and weak texture enhancement of the weak edge region can be effectively avoided, and the texture enhancement processing effect is improved.

[0223] In optional embodiments of the present application, the image texture enhancement network can include at least one double convolution module, one double convolution module including a first branch for obtaining a texture enhancement residual of the image to be processed, a second branch for obtaining a noise suppression residual of the image to be processed, and a residual fusion module for fusing the texture enhancement residual and the noise suppression residual to obtain a texture residual. The network parameters of the first branch and the second branch are different.

[0224] For one double convolution module, fusing the texture enhancement residual and the noise suppression residual to obtain a texture residual includes:

[0225] Subtracting the texture enhancement residual and the noise suppression residual to obtain a texture residual.

[0226] Correspondingly, obtaining an image after texture enhancement according to the texture residual and the image to be processed includes:

[0227] Superimposing the texture residual corresponding to each double convolution module and the image to be processed to obtain an image after texture enhancement.

[0228] The image texture enhancement processing can be realized through a texture enhancement model, i.e., a texture enhancement network. To realize the above-mentioned embodiments of the present application, the texture enhancement model can specifically include at least two branches, one enhancement branch (first branch) and one suppression branch (second branch). The enhancement branch is used to predict an enhancement residual, i.e., the texture enhancement residual (enhancement of useful texture and useless noise), and the suppression branch is used to predict a suppression residual, i.e., the noise suppression residual. The branch can suppress or reduce amplified noise and reasonably adjust the problem of over-enhancement. Through the combination of the enhancement branch and the suppression branch, the texture after enhancement processing can be more realistic and natural, thereby improving the performance of texture enhancement.

[0229] In optional embodiments of the present application, for one double convolution module, the first branch includes a first convolution module for obtaining a texture enhancement residual of the image to be processed, and a first nonlinear activation function layer for performing nonlinear processing on the texture residual output by the first convolution module; the second branch includes a second convolution module for obtaining a noise suppression residual of the image to be processed, and a second nonlinear activation function layer for performing nonlinear processing on the noise suppression residual output by the second convolution module; and the convolution processing parameters of the first convolution module and the second convolution module are different.

[0230] The specific network structure of the texture enhancement model is not limited by the embodiments of the present application. As an optional solution, the texture enhancement model can be a neural network model based on a convolution structure. The model can include two convolution branches, and the convolution processing parameters of each branch are different. When processing an image, different convolution branches have different effects on different texture regions. Specifically, for one of the branches, the branch can obtain a texture enhancement residual image of the to-be-processed image based on the to-be-processed image through convolution processing. Another branch obtains a noise suppression residual image of the to-be-processed image based on the to-be-processed image through convolution processing. Then, a difference image of the texture enhancement residual image and the noise suppression residual image is used as a texture residual image used by the to-be-processed image finally. The texture-enhanced image is obtained by superimposing the texture residual image and the to-be-processed image.

[0231] Further, for different texture regions in the to-be-processed image, the texture information and noise information of different regions can be different. Therefore, for different texture regions, some convolution branches can have an effect, and some convolution branches can not have an effect on the enhancement effect. The texture enhancement model can distinguish noise points and small textures, and strong edges and weak edges in the image by selecting different branches. That is, texture enhancement focuses on local details of an image and is very sensitive to image noise. Therefore, there are two main challenges in the texture enhancement task: small texture problem, which requires accurate separation of noise and small texture regions; and overshoot problem, which requires processing of strong and weak edges with different intensities. Therefore, the texture enhancement model with a multi-branch structure can well handle the two difficult problems.

[0232] As an optional solution of the present application, the texture enhancement model can specifically adopt a residual network-based model, that is, a multi-branch residual network. The network parameters of each branch are different. For different texture regions in the to-be-processed image, the multi-branch residual network can play a role of "branch selection", that is, according to the characteristics of the texture region, different branches play different roles, or some branches play a role and some branches do not play a role.

[0233] As an example, Figure 5aA structure diagram of a multi-branch residual network provided in an embodiment of the present application is shown. In this example, the multi-branch residual network includes a double convolution module, which includes a network having two convolution branches (convolution processing modules shown in the figure). The convolution processing parameters of the two convolution branches are different. Each convolution branch can include, but is not limited to, a convolution layer, and can also include an excitation function layer (i.e., a nonlinear activation layer such as a relu excitation function), etc. After the input image is processed by the convolution layer, the convolution processing result can be further processed by the excitation function layer for nonlinear processing. The convolution processing results of each convolution branch and the input image are superimposed to obtain the final texture-enhanced image. In the figure, one branch obtains a texture residual image based on the image to be processed, and the other branch obtains a noise residual image based on the image to be processed. Then, the difference image of the texture residual image and the noise residual image is taken as the final texture residual image, which is superimposed with the image to be processed to obtain the final image.

[0234] It can be understood that the subtraction or superposition of the graph (image or feature map) and the graph (image or feature map) in the embodiments of the present application refers to the subtraction or addition of the element values of the corresponding position points in the two graphs.

[0235] Optionally, at least two texture-enhanced residuals of the image to be processed can be obtained based on the image to be processed by using at least two first convolution processing parameters, and at least two noise suppression residuals corresponding to the texture residuals can be obtained based on the at least two second convolution processing residuals.

[0236] Based on the corresponding texture-enhanced residuals and noise suppression residuals, at least two difference results, i.e., two texture residuals, are obtained.

[0237] According to the at least two difference results and the image to be processed, a texture-enhanced image is obtained.

[0238] That is, the image texture enhancement model, i.e., the texture enhancement network, can include multiple texture enhancement branches (i.e., the double convolution module described above), and the convolution network types and / or convolution processing parameters (including but not limited to convolution kernel size, etc.) of each texture enhancement branch are different. Each texture enhancement branch includes two branches of the branch. For each texture enhancement branch, one branch of the branch is used to obtain a texture-enhanced residual image, and the other branch is used to obtain a noise suppression residual image. The texture-enhanced residual image and the noise suppression residual image of the branch are subtracted to obtain a difference result corresponding to the branch. Finally, the texture-enhanced residual results corresponding to each texture enhancement branch and the image to be processed are superimposed to obtain a texture-enhanced image. By processing through multiple texture enhancement branches, better processing results can be obtained.

[0239] Optionally, the image texture enhancement network can include at least two double convolution modules based on a cavity convolution network, wherein the dilatation rates of the cavity convolution networks of the double convolution modules based on the cavity convolution network are different.

[0240] As an example, Figure 5b A structure diagram of a texture enhancement model provided by an embodiment of the present application is shown in FIG. 1, two convolution processing modules (Conv 3x3 shown in the figure) in the figure are a texture enhancement branch, each texture enhancement branch includes two convolution branches, and the convolution processing parameters of the two convolution branches can be different, such as Figure 5b As shown in FIG. 2, the convolution kernel size of each convolution processing parameter in the example can be 3*3, but the model parameters of different convolution kernels are different, for each texture enhancement branch, the processing results (i.e., texture enhancement residual and noise suppression residual) corresponding to the two convolution processing parameters of the branch can be subtracted to obtain the processing result corresponding to the texture enhancement branch, in an optional embodiment of the present application, the texture enhancement model can include the processing results of N (N≥1) texture enhancement branches, and the processing results of the N branches and the input image are superimposed to obtain an output image, i.e., a texture enhanced image. As an example, Figure 5c A structure diagram of a texture enhancement model (i.e., an image texture enhancement network including two double convolution modules) when N=2 in FIG. 3 is shown in FIG. 4. Figure 5b A structure diagram of a texture enhancement model (i.e., an image texture enhancement network including two double convolution modules) when N=2 in FIG. 3 is shown in FIG. 4.

[0241] It is clear to those skilled in the art that for each convolution branch, for different application requirements, a convolution layer (Conv shown in the figure) can be further connected with an excitation function layer, and a pooling layer can be further arranged between the convolution layer and the excitation function layer.

[0242] As an optional solution, the pixel value of a pixel point in an input image is x (0≤x≤1, i.e., a normalized pixel value), the pixel value of a pixel point in an output image (i.e., an image obtained after texture enhancement processing) is y (0≤y≤1), when a multi-branch residual network is used for texture enhancement processing, the residual res corresponding to the output image and the input image can be represented as: res=y-x, if the value range of res is -1≤res≤1, then according to the value range, the following multi-branch residual network can be designed:

[0243]

[0244] Wherein, N represents the number of branches of the residual network branch (i.e., the texture enhancement branch), i1 and i2 respectively represent two convolution processing parameters of the i-th branch, conv(·) represents a convolution layer, represents convolution processing using the first convolution processing parameter of the i-th branch (i.e., the first branch), denotes the second kind of convolution processing parameter (i.e., the second branch) is adopted to perform the convolution processing, and Relu(·) denotes an activation function.

[0245] When the image texture enhancement processing is performed by using the multi-branch residual network based on the principle of the expression, when , the first branch (i.e., the branch using the first kind of convolution processing parameter to perform the convolution processing) of the branch i (i.e., the double convolution module i) does not work in the texture enhancement; only when , the first branch of the branch i can work in the texture enhancement process. By using the network, for different texture regions, the branches that work in the texture enhancement process are different. In this way, the multi-branch structure model can distinguish the noise and the small texture, and the strong edge and the weak edge by using different combinations of branches, so that a better texture enhancement effect is obtained.

[0246] As another example, Figure 5d a structure schematic diagram of a texture enhancement model provided by another optional embodiment of the present application is shown, which is a texture enhancement model when N=4 (i.e., the number of residual branches / double convolution modules is 4), each dashed box in the figure corresponds to a texture enhancement branch, and the convolution processing modules in each branch can use a hollow convolution module. The expansion rates (i.e., dilation rates) of the hollow convolutions corresponding to the four branches can be different, that is, four double convolution modules (one double convolution module is one branch) with different expansion rates are applied to the input image (for example, the expansion rates of the convolution modules of the four branches can be set to 1, 2, 5, and 7), and the two hollow convolutions in the same module can have the same expansion rate. Different branches use different expansion rates, which can realize different enhancement processing of texture details of different scales in the image. This is because a small expansion rate pays attention to close-range information, and a large expansion rate pays attention to long-distance information. The convolution modules with different expansion rates can reflect texture details of different scales. Texture has different proportions, noise enhancement and overshoot also occur at different proportions, and different expansion rates can extract features of textures of different proportions. The branch structure of the double convolution module can effectively overcome the problems of noise amplification and texture overshoot in the implementation of texture enhancement. Each convolution layer in the model can be cascaded with a ReLu layer, and can have no batch normalization layer (of course, it can also have a batch normalization layer). Based on the scheme of the embodiments of the present application, by interleaving a group of double convolution blocks with different expansion rates, the texture enhancement module can capture short-distance context information (guiding the enhancement of small texture) and long-distance context information (guiding the enhancement of strong texture), so that a better texture enhancement effect is obtained.

[0247] Wherein, the specific network structure of each convolution processing model is not limited in the embodiments of the present application, for example, each convolution processing module can adopt a dilated convolution structure with a convolution kernel size of 3*3, and the dilation rates of the four branches shown in the figure can be set to 1, 2, 5, and 7 respectively.

[0248] In the image texture enhancement network provided by the embodiments of the present application and containing at least one double convolution module, each double convolution module can contain two branches with the same structure, i.e., the first branch and the second branch described above, and the convolution processing parameters of the two branches are different, such as the parameters of the convolution kernels, different convolution processing parameters can predict different features, and the texture enhancement network can be obtained through training. Wherein, the first branch of a double convolution module predicts a global texture residual (i.e., a texture enhancement residual) used for texture enhancement processing of the image to be processed, but when performing global texture enhancement of the image, it is very likely to cause problems of increased noise and over-enhanced image texture. To solve this problem, the second branch of the double convolution module is used to predict a noise suppression residual used for adjusting the global texture residual, which is used for local adjustment of the global texture residual and plays a role in suppressing or reducing amplified noise and over-enhanced texture. That is, the global texture residual plays a role in enhancing the overall texture of the image, and the noise suppression residual is used for adjusting the local enhancement amount, one is a coarse-grained overall adjustment, and the other is a fine-grained local correction, thereby avoiding or reducing the problems of over-enhancement and noise amplification. Wherein, the output of a double convolution module is the difference between the first branch and the second branch, and the difference is added to the image to be processed as the final texture residual.

[0249] It is clear to those skilled in the art that the "residual" appearing in the optional embodiments of the present application is a concept of an adjustment amount, i.e., how much to increase or decrease, for example, an image pixel value is enhanced from 150 to 166, and 16 is the residual of the enhancement.

[0250] In the optional embodiments of the present application, the image texture enhancement network is used to obtain the texture enhancement residual and the noise suppression residual of the image to be processed, comprising:

[0251] Obtaining the luminance channel image and the non-luminance channel image of the image to be processed;

[0252] Using the image texture enhancement network based on the luminance channel image to obtain the texture enhancement residual and the noise suppression residual of the image to be processed;

[0253] Correspondingly, obtaining the image after texture enhancement according to the texture residual and the image to be processed, comprising:

[0254] Obtaining the luminance channel image after texture enhancement according to the texture residual and the luminance channel image;

[0255] The texture-enhanced luminance channel image and the non-luminance channel image are fused together to obtain the texture-enhanced image.

[0256] Since the color channel images do not affect the image texture, only the luminance channel image can be used when performing image texture enhancement. Therefore, when predicting the texture enhancement residual and noise suppression residual of an image, only the luminance channel image can be used. The input of the dual convolution module is the luminance channel image. By adding the output of each dual convolution module to the luminance channel image, the texture-enhanced luminance channel image is obtained. Then, this image and the non-luminance channel images (i.e., the color channel images) are fused again to obtain the texture-enhanced image.

[0257] As an example, Figure 5e The figure shows a schematic diagram of the principle of an image texture enhancement processing scheme provided by an embodiment of this application. As shown in the figure, the input image, i.e. the image to be processed, is an RGB image. The image texture enhancement network includes four dual convolutional modules. Each dual convolutional module includes two convolutional branches. Each convolutional branch may include cascaded convolutional layers and ReLU layers. The convolution in this example adopts dilated convolution. The kernel size of each convolutional layer in the four dual convolutional modules can be 3×3. The dilation rates of the four dual convolutional modules are 1, 2, 5, and 7, respectively. By using different dilation rates, it is possible to predict the texture enhancement residual and noise suppression residual at different scales and ranges. Small dilation rates focus on information in a smaller range, while large dilation rates focus on information in a larger range. Therefore, by using dual convolutional modules corresponding to multiple dilation rates, texture detail information corresponding to multiple scales can be obtained.

[0258] Specifically, based on Figure 5e The network structure shown first converts the input image from RGB to YUV (RGB→YUV as shown in the figure). The Y channel image is the luminance channel image, and the UV channel images are the non-luminance channel images. The Y channel image is then input into each dual convolutional module. Taking the first dual convolutional module on the left in the figure as an example, the texture enhancement residual can be obtained through its first branch. Its second branch leads to the noise suppression residual. The texture residual corresponding to the first dual convolutional module is Similarly, the texture residuals corresponding to the other three dual convolutional modules can be obtained. Then, the texture residuals corresponding to the four dual convolutional modules and the luminance channel image are added pixel by pixel to obtain the texture-enhanced luminance channel image. The texture-enhanced luminance channel image and the non-luminance channel image are then merged and converted into an RGB image (YUV→RGB as shown in the figure). This RGB image is the texture-enhanced output image.

[0259] To better understand the principle of the dual convolution module provided in the embodiments of this application, the following will be combined with... Figure 5f The text explains the data processing principle of a double convolution module. For example... Figure 5f As shown in the figure, branch 1 is the first branch used to obtain the texture enhancement residual of the processed image (the enhancement residual shown in the figure), and branch 2 is the second branch used to obtain the noise suppression residual of the image to be processed (the suppression residual shown in the figure). The input in the figure corresponds to the luminance channel image of the image to be processed, and the output is the texture-enhanced luminance channel image. It can be understood that... Figure 5f The input and output diagrams in the figure are provided for ease of understanding the principle of the dual convolution module. Each bar in the bar chart can be understood as the value of a pixel in the image. As shown in the figure, for the input image, branch 1 predicts the enhancement residual, which enhances both useful textures and useless noise in the image. Branch 2 predicts the suppression residual, which suppresses or reduces amplified (i.e., enhanced) noise and adjusts over-enhanced textures. Subtracting the enhancement residual (as shown in the figure) The difference obtained by subtracting the enhancement residual and the suppression residual pixel by pixel is used for texture enhancement of the image. This difference is then added to the image (as shown in the figure). That is, the difference is added to the image pixel by pixel to obtain the output. The texture in the output image is reasonably enhanced and the noise in the image is effectively suppressed. Figure 5g It shows Figure 5f The diagram shows an enlarged view of the output section. The bars filled with pure black represent pixels where texture enhancement is applied, and the pure black areas indicate the intensity of enhancement. The bars corresponding to dashed lines represent pixels where noise is suppressed, and the space between the dashed lines and the corresponding bars indicates the intensity of suppression. As can be seen from this diagram, the texture enhancement network provided in this application can achieve adaptive enhancement of textures in an image, that is, it can achieve reasonable texture enhancement while effectively suppressing noise in the image.

[0260] In the training of the image texture enhancement network based on the training sample images, the training sample images can include sample image pairs, the image pairs can include an image with clear texture and an image corresponding to the image with clear texture and requiring texture enhancement processing, that is, the contents of the images in the image pairs are the same, but the textures are different. In the training of the texture enhancement network, the loss function can be L1 loss, and the value of the loss function is minimize (|ground truth-output|). Wherein, minimize refers to minimization, ground truth refers to the image with clear texture of the training sample image, output is the result image obtained after the image requiring texture enhancement processing in the training sample image is processed by the image texture enhancement network, that is, the image after texture enhancement, and |ground truth-output| represents the L1 loss between the image after texture enhancement processed by the network and the corresponding image with clear texture.

[0261] In an optional embodiment of the present application, the at least one image quality enhancement manner includes at least two enhancement manners, and the quality of the to-be-processed image is enhanced by using at least one image quality enhancement manner to obtain a processed image, including:

[0262] The to-be-processed image is enhanced by using at least two enhancement manners;

[0263] The processed image is obtained based on the processing results corresponding to the enhancement manners.

[0264] That is, when multiple image enhancement processing manners are used to process the to-be-processed image, each manner can be used to perform corresponding enhancement processing on the to-be-processed image, and then the final enhanced processing result is obtained based on the processing results corresponding to each manner, for example, the processing results corresponding to each enhancement manner can be fused to obtain the final image. For example, the to-be-processed image can be processed by using image denoising and image texture enhancement, and then the processing results corresponding to the two processing manners are fused to obtain the final image.

[0265] In an optional embodiment of the present application, the to-be-processed image is enhanced by using at least two enhancement manners, including:

[0266] The to-be-processed image is enhanced in sequence according to the processing order of the at least two enhancement manners.

[0267] Any enhancement manner except the first enhancement manner can process the to-be-processed image based on the processing result of at least one enhancement processing manner located before the enhancement manner.

[0268] That is, when multiple enhancement manners are used to process the to-be-processed image, there can be a certain processing order between the multiple enhancement manners, and the input information of the first enhancement manner is the to-be-processed image. For the other enhancement manners except the first enhancement manner, the input information of the enhancement manner can include the output information of at least one enhancement manner before the enhancement manner.

[0269] The processing order corresponding to the multiple enhancement manners can be determined according to the characteristics and advantages of each enhancement manner. Specifically, the processing order can be set in advance, for example, can be determined according to experimental data and / or experience summary, or can be predicted, for example, by using a neural network model to predict the processing order.

[0270] In an optional embodiment of the present application, the method further includes:

[0271] determining scene information corresponding to the to-be-processed image;

[0272] determining, according to the scene information, an enhancement manner corresponding to the to-be-processed image and a processing order between different enhancement manners.

[0273] In actual applications, the quality of an image is directly related to the scene characteristics of the image. Images acquired in different scenes usually have different image characteristics, and different targeted processing manners need to be used. For example, an image acquired in a scene with good light conditions can only need to be subjected to image texture enhancement processing, while an image acquired in a scene with poor light conditions can need to be subjected to brightness adjustment and / or noise reduction processing.

[0274] In order to achieve more targeted image enhancement processing and better meet different actual application requirements, the optional scheme of the embodiment of the present application can determine the enhancement manner and the enhancement processing order of the to-be-processed image according to the scene information (for example, the scene type) of the to-be-processed image, so as to perform image enhancement processing according to the manner matched with the scene information of the image, so as to achieve the purpose of targeted enhancement and improve the image processing effect.

[0275] As an optional manner, Figure 6a A flowchart of an image processing method provided by the present application is shown in the figure. As shown in the figure, for a to-be-processed image, scene detection can be first performed on the image to determine the scene information corresponding to the image. Then, the corresponding enhancement module and the connection manner between the modules can be selected according to the determined scene information, and image enhancement processing is performed according to the selected manner to obtain an enhanced image.

[0276] Optionally, the scene information of the to-be-processed image is determined by:

[0277] The scene information of the to-be-processed image is determined by a scene detection network based on the to-be-processed image.

[0278] As an option, the prediction of the scene information of the image can be implemented by a pre-trained neural network.

[0279] In actual applications, the scene information of various images can be classified, and in this case, the prediction of the scene information of the image can be converted into a classification problem of the scene information. For an image to be processed, the scene category (i.e., type) corresponding to the image can be predicted by a neural network, that is, the scene information of the image.

[0280] The division of the scene type is not limited in the embodiments of the present application, and can be divided according to actual requirements and various characteristics of the image. Optionally, the illumination condition and the noise level of the image are two classic descriptors of the image, which are helpful for the recognition of the image scene. Therefore, the scene type can be divided according to the illumination condition, the noise level, and the like of the image. As an example, the scene type of the image can be divided into four typical scenes: normal illumination, backlight, weak light level 1, and weak light level 2. Taking the above four scene types as an example, the recognition of the scene type of the image can be converted into a four-classification task of the image, and the recognition of the scene type of the image to be processed can be implemented by a pre-trained neural network model (such as a convolutional neural network model).

[0281] In addition, when the enhancement manner of the image includes two or more than two, the processing order of the multiple enhancement manners is different, and the image processing effect produced is also different. In order to further improve the processing effect of the image in different scenes, for different scene types, the corresponding enhancement manner and the processing order between different enhancement processing manners can be pre-configured. After the scene type of the image to be processed is determined, the corresponding enhancement manner and the processing order can be used for the enhancement processing of the image according to the scene type.

[0282] It can be understood that the enhancement manner corresponding to a scene type can also be an enhancement manner. The correspondence between different scene types and the enhancement manner and the processing order can be configured according to empirical values and / or experimental values. For example, the enhancement manner corresponding to each scene type can be configured according to empirical values. When the enhancement manner is two or more than two, the processing order between the multiple enhancement manners can be determined according to experience or by experiment.

[0283] As an example, taking the above four scene types as an example, an optional solution of an image enhancement processing manner provided by an embodiment of the present application for different scene types is shown in the following table. As shown in the table, different scene types correspond to different image characteristics, and different targeted image enhancement manners can be used. In the table, ③ represents texture enhancement processing, ② represents brightness and color enhancement processing, and ① represents denoising processing. The connection order between the processing manners represents the connection order between the processing manners, and the arc line between the processing manners represents that the models of different processing manners can use the dense connection (see the description below) manner. Taking the weak light level 1 scene type as an example, the corresponding enhancement manner can include image denoising processing and texture enhancement processing. The input of the image denoising model is the to-be-processed image, the input of the texture enhancement model includes the to-be-processed image and the output of the denoising model, and the image after enhancement processing can be obtained according to the input image, the output of the denoising model and the output of the texture enhancement model.

[0284]

[0285] In an optional embodiment of the present application, any enhancement manner except the first enhancement manner is based on the processing result of at least one enhancement processing manner located before the enhancement manner and the to-be-processed image to process the to-be-processed image.

[0286] That is, the various different enhancement processing manners can use the dense connection manner. For any enhancement manner except the first enhancement manner, the input information of the manner can include the processing result of at least one enhancement manner located before the manner and the to-be-processed image, that is, the input information of each enhancement manner includes at least two, so that better processing effect can be obtained based on more diversified input information.

[0287] In an optional embodiment of the present application, any enhancement manner except the first enhancement manner is based on the processing result of all enhancement processing manners located before the enhancement manner to process the to-be-processed image. At this time, the processed image is obtained based on the processing result corresponding to each enhancement manner, including:

[0288] The processing result corresponding to each enhancement manner is fused to obtain the processed image, or the processing result of the last enhancement manner is taken as the processed image.

[0289] Specifically, for the last enhancement manner, since the input information thereof contains the processing results of various enhancement manners before it, the processing result thereof can be taken as the final processing result, of course, the processing results corresponding to various processing manners can also be fused to obtain the final processing result. The specific fusion manner is not limited in the embodiments of the present application, for example, it can be superposition or other fusion manners, for example, various enhancement manners can correspond to different weights, and the processing results corresponding to various enhancement manners can be weightedly fused based on the weights corresponding to various enhancement manners to obtain the final processing result.

[0290] In the optional embodiments of the present application, the at least one image quality enhancement manner includes image denoising, image tone enhancement and image texture enhancement, and the processing order of various enhancement manners is: image denoising, image tone adjustment and image texture enhancement.

[0291] Optionally, the image tone adjustment includes image brightness adjustment and / or image color adjustment.

[0292] As an optional solution, the image quality enhancement method provided by the embodiments of the present application can simultaneously achieve the purposes of image denoising, tone adjustment and texture enhancement. For the three specific image quality enhancement sub-tasks of image denoising, color / brightness adjustment and texture enhancement, the present application proposes a prior topological information, which is derived from a large number of experimental examples and experience summaries. In an image processing workflow involving the above three sub-tasks, the priority of image denoising is the highest. Otherwise, if image denoising is after color adjustment, the noise distribution may be changed to affect the subsequent denoising, and if image denoising is after texture enhancement, the noise will be amplified and enhanced while the texture is enhanced, which is not conducive to subsequent denoising. Therefore, image denoising can be performed first. In comparison, the coupling between the tone adjustment task and the texture enhancement task is much weaker, but the brightness and color information is helpful to restore the texture details, especially for dark areas in the image. Therefore, through a large number of experimental examples and experience summaries, the prior task topological information determined by the embodiments of the present application can be described as: image denoising—>color enhancement (i.e. tone adjustment)—>texture enhancement. In order to further explore the potential relationship between the sub-tasks and optimize the combination relationship.

[0293] In an optional implementation of the embodiments of the present application, when the enhancement processing is performed on the image in multiple aspects, that is, the image quality enhancement includes multiple sub-tasks, dense connection can be introduced between the sub-tasks, that is, the input information of a subsequent sub-task includes the output results of each previous sub-task and the image to be processed. For example, the image quality enhancement processing mode includes the image denoising, the image tone enhancement and the image texture enhancement described above. According to the prior task topology information, the processing order among the three can be the image denoising, the image tone adjustment and the image texture enhancement in sequence. When the image is processed, the models of the three sub-tasks are densely connected.

[0294] Specifically, the image to be processed is input into the image denoising model, the output of the image denoising model and the image to be processed are input into the image tone adjustment model, the image tone adjustment model further processes according to the output of the image denoising model and the image to be processed, the image to be processed and the output of the image tone adjustment model are input into the texture enhancement model, and finally, the result image after quality enhancement can be obtained based on the image to be processed, the output of the image denoising model, the output of the image tone adjustment model and the output of the texture enhancement model. The image denoising model, the image tone adjustment model and the texture enhancement model can be implemented by using existing models, or at least one of the three can be implemented by using the model provided in the foregoing embodiments of the present application.

[0295] In an optional embodiment of the present application, the at least one image quality enhancement mode described above can be at least one of the candidate quality enhancement modes, and before the at least one image quality enhancement mode is used to perform quality enhancement on the image to be processed, the method can further include:

[0296] According to the image to be processed, at least one image quality enhancement mode corresponding to the image to be processed is determined from the candidate quality enhancement modes.

[0297] That is, in actual application, a plurality of candidate quality enhancement processing modes can be pre-configured. For different images to be processed, the image information such as noise and brightness can be different, and in order to obtain better processing effect, different images also need to be processed by different processing modes. For example, the brightness of some images is already good, so it can not be necessary to adjust the brightness again. The texture of some images is clear, so it can not be necessary to perform texture enhancement processing. Therefore, in order to realize individualized processing for different images, the scheme provided in the embodiments of the present application can first determine which processing mode or which processing modes are used to process the image to be processed before the image to be processed is processed, and then the image to be processed is processed based on the determined image enhancement mode.

[0298] The at least one image quality enhancement manner corresponding to the to-be-processed image is determined from each candidate image quality enhancement manner according to the to-be-processed image, and can also be implemented by using a neural network model. Specifically, the to-be-processed image can be input into a pre-trained image processing manner screening model, and the image quality enhancement manner actually used is determined based on the output of the model.

[0299] In an optional embodiment of the present application, when the determined at least one image quality enhancement manner is at least two, the method can further include:

[0300] determining the processing order between each of the at least one image quality enhancement manner.

[0301] Optionally, when the processing order between the at least two enhancement manners is determined, the processing order between the at least two enhancement manners can be determined based on the to-be-processed image by using a processing order prediction network.

[0302] Optionally, the manner of determining the processing order between each of the at least one image quality enhancement manner can be determined based on a preconfigured mapping relationship, that is, a mapping relationship between different combinations of processing manners and corresponding processing orders can be pre-set. For the to-be-processed image, when the corresponding multiple processing manners are determined, the corresponding processing order can be determined based on the mapping relationship.

[0303] As another optional manner, the above processing order can also be determined by using a neural network model, for example, the determined multiple image quality enhancement manners can be input into a sequence determination model (i.e., a processing order prediction network), or the to-be-processed image and the multiple image quality enhancement manners can be input into the model, and the corresponding processing order is obtained based on the output of the model.

[0304] It can be understood that the image processing manner screening model and the sequence determination model can be one model, that is, the two models can be sequentially cascaded to form one model, and the output of the image processing manner screening model in the cascaded model is the input of the sequence determination model. Of course, the image processing manner screening model and the sequence determination model can also be two independent models.

[0305] In an optional embodiment of the present application, the processing order prediction network includes a decision branch for selecting a current candidate enhancement manner from the at least two enhancement manners based on input information, and an inference branch for determining whether the current candidate enhancement manner is a target enhancement manner, wherein the input information is the to-be-processed image or an enhancement processing result of the to-be-processed image by the enhancement manner with a determined sequence.

[0306] The specific network structure of the prediction network is not limited in the embodiments of the present application. Optionally, when the neural network model is used to determine the processing sequence of the at least two enhancement manners, the neural network model can be a model based on a recurrent neural network, for example, a model based on LSTM (Long Short Term Memory). The recurrent neural network outputs one of the at least two enhancement manners at each step, and the output sequence of each enhancement manner is the processing sequence corresponding to the enhancement manner.

[0307] The scheme of determining the processing sequence among the multiple enhancement manners by using the neural network model is described in detail below with reference to an example.

[0308] As shown in FIG. 6b, a schematic diagram of determining the processing sequence of the multiple enhancement manners by using the LSTM-based neural network model in this example is shown. In this example, it is assumed that the enhancement manners include N kinds. Each enhancement manner can be implemented by using a neural network model to perform enhancement processing, for example, an image denoising model can be used to perform denoising processing, a tone adjustment model can be used to perform tone adjustment, and a texture enhancement model can be used to perform texture enhancement. The processing sequence is determined, that is, the data processing sequence among the models is determined. One module in the figure corresponds to one processing model, that is, one enhancement manner. The processing flow of this example is as follows:

[0309] Figure 6b The input of the model shown in FIG. 6b includes an image to be processed. As shown in the figure, at the first processing step of the LSTM, that is, the first step, the image to be processed is input into the LSTM to obtain a one-dimensional vector with a length of N. Each value in the vector represents the probability of being selected as a module, that is, the probability of being selected as a module for the N modules. Optionally, in order to determine the first module (module 1 shown in the figure) for performing enhancement processing, it can be determined whether there is a probability greater than a set threshold value in the N values. If there is, the module 1 can be determined from the modules corresponding to the probabilities greater than the set threshold value. For example, the module corresponding to the maximum value in the probabilities greater than the threshold value in the N values can be determined as the module 1. If there is no probability greater than the set threshold value in the N values, the flow ends.

[0310] For each processing step except the first processing step, the input of the LSTM includes the output image of the image to be processed after the modules in the determined order are processed and the hidden state vector of the LSTM of the previous processing step, and the output is still a feature vector with a length of N. The module corresponding to the step can be determined by using the same processing method as described above. Taking the second step as an example, the input is the image processed by the module 1 and the hidden state vector of the LSTM of the first step (as shown by the arrow between the two LSTMs in the figure). It should be noted that for the N probabilities corresponding to each step except the first step, if the module corresponding to the maximum probability in the probabilities greater than the set threshold in the N probabilities corresponding to the step is the module that has been selected before, the module selection result corresponding to the step can be the module corresponding to the second largest probability in the probabilities greater than the set threshold.

[0311] It should be noted that, as an example, the figure only shows the process of determining the first module (module 1) and the second module (module 2), and the ellipsis in the figure is omitted for other processing steps. The principle of the omitted processing steps is the same as that of the second supplementary step.

[0312] In addition, if the processing order between the enhancement modes is not determined at the end of the process, the processing order can be determined according to experience or the method of determining the order according to the scene type provided in the foregoing or other methods. Of course, the set threshold value described above can not be set, and for each processing, the module corresponding to the maximum value in the N probabilities corresponding to the current processing can be directly determined as the target module determined by the current processing. For example, without judging whether to set the threshold value, the maximum value in the N probabilities corresponding to the first processing time can be directly determined as the module 1.

[0313] The specific model structure of the LSTM can be selected according to actual needs. For example, it can include a convolution structure and a fully connected layer structure that are sequentially cascaded. The convolution structure is used to extract the image features of the input (the image to be processed or the output image of the previous module and the hidden state features of the LSTM corresponding to the previous processing step), and the fully connected layer structure is used to fully connect the features output by the convolution structure and then output a feature vector with a length of N. Each feature value of the vector corresponds to an enhancement mode, i.e., a module, and each feature value can represent the confidence of the module corresponding to the feature value being placed in the processing order corresponding to the current step. Then, the feature vector can be converted into N probabilities by using the Sigmod function (corresponding to the arrow in the figure). ) The N probabilities are the probabilities used to represent the selection of the modules.

[0314] It will be clear to those skilled in the art that the solution provided in the embodiments of this application can be implemented by designing an end-to-end image quality enhancement model. The end-to-end image quality enhancement model provided in the embodiments of this application will be described in detail below with reference to an example.

[0315] As an optional embodiment, Figure 6c The figure shows a schematic diagram of the structure of an image quality enhancement model provided in an embodiment of this application. As shown in the figure, the image quality enhancement model may include an image denoising model, a tone adjustment model, and a texture enhancement model. The three models can specifically adopt the model structures provided in the previous examples of this application. For example, the image denoising model can adopt... Figure 3a or Figure 3b The structure shown in the figure, the tone adjustment model can be adopted Figure 4a or Figure 4b The structure shown, the texture enhancement model can adopt Figure 5a , Figure 5b , Figure 5c or Figure 5d The structure shown.

[0316] Depend on Figure 6c As can be seen from the embodiments of this application, after establishing lightweight models for each subtask, an overall quality enhancement model can be built end-to-end based on the established prior topology information. However, the overall model is not simply a series of subtasks, because there are potential coupling relationships between the tasks. Dense connections effectively tap into the advantages and characteristics of each subtask, as well as the optimal combination relationships between tasks. Through this image quality enhancement model, image quality enhancement tasks can be completed in one step from multiple aspects such as image noise, image brightness and color, and image texture. It fully taps into the characteristics and advantages of image denoising, image brightness and color adjustment, and texture enhancement subtasks, ensuring the enhancement effect and real-time processing of individual subtasks. Dense connections effectively tap into the advantages of each subtask and the optimal combination relationships between tasks, achieving end-to-end image quality enhancement.

[0317] The core idea of ​​dense connections is to establish short connections between tasks. If there are L tasks, then there will be L(L+1) / 2 connections. The input of each subtask comes from the outputs of all preceding tasks. Dense connections can leverage the strengths of each task and achieve optimized combinations between them. For example, one possible scenario is that while denoising an image, some texture details may be lost. However, the texture enhancement model can take on both the output of the denoising task and the input of the original image (i.e., the image to be processed). This helps to maintain the denoising effect while also aiding in the restoration and enhancement of texture details.

[0318] For the end-to-end image quality enhancement model provided by the embodiments of the present application, in the training stage, the multi-stage joint debugging mechanism is introduced, as shown in Figure 7a and Figure 7b :

[0319] Specifically, in the training, the individual training of each subtask can be performed first, that is, the image denoising model, the tone adjustment model and the texture enhancement model can be trained individually first, and after the single task precision reaches a certain precision, the corresponding subtasks are connected in stages, that is, the denoising subtask and the tone enhancement subtask can be connected first, that is, the image denoising model and the tone adjustment model are densely connected, and the connected image denoising model and tone adjustment model are trained as a whole, and after a certain precision is reached, that is, a certain convergence result is obtained, the texture enhancement subtask can be connected (that is, densely connected) on this basis, and the entire image quality enhancement model is trained.

[0320] In the stage joint debugging stage, that is, the model training stage after connecting two or three subtasks, the model mainly learns the dense connection weights between subtasks. In this stage, the main model weights of part of the subtask models can be fixed during training, because the subtask model reaches a certain training precision during single subtask training. Therefore, this stage can only focus on learning the weights corresponding to the dense connection part. Among them, which model weights of each subtask are fixed can be configured according to experience or experimental results, for example, for the image denoising model shown in Figure 3a or Figure 3b , the model parameters of the denoising model part can be fixed, and for the tone adjustment model shown in Figure 4a or Figure 4b , the model weights of the color adjustment model and the brightness adjustment model can be fixed, that is, the core weight parameters of each subtask model can be fixed, and this stage mainly learns the model parameters of the input and output parts of each subtask model, that is, the weight parameters of the connection part between the subtask models.

[0321] Optionally, in the training, the loss function involved can include a subtask loss function, a stage loss function and a global loss function, wherein the subtask loss function is the loss function corresponding to each subtask model, such as Figure 7aThe loss for Task 1 shown is the loss function of the sub-task in the image denoising model, representing the difference between the labeled output image corresponding to the input image and the processing result output by the image denoising model. Similarly, the loss for Task 2 is the loss function of the sub-task in the tone adjustment model, representing the difference between the labeled output image corresponding to the input image and the processing result output by the tone adjustment model. The loss for Task 3 is the loss function of the sub-task in the tone adjustment model. The stage loss function is the loss function used when training by concatenating the sub-task models. After concatenating different sub-task models, the model also needs to learn the weight parameters of the connection parts between the sub-task models, such as... Figure 7b and Figure 7c The connection curves between the different sub-task models shown in the figure represent the content, while the global loss function (the total loss shown in the figure) is the loss function of the model when training the entire end-to-end image quality enhancement model.

[0322] Given the three subtasks in the image quality enhancement task, there should be potential correlations among them. To leverage the strengths of each subtask and enhance the deep interactions between them, the aforementioned task interaction mechanism based on dense connections (dense connections) can be employed, as shown in Figure 6. Figure 7b and Figure 7c As shown in the diagram. Furthermore, to properly supervise the subtasks and accelerate training, a training task-specific loss Li(x, xi*) can be assigned to each subtask (optionally, this loss uses L1 loss). Optionally, the final training loss can be expressed as:

[0323] L(x,xi*)=Σ i α i L i (x,xi*)

[0324] Where i = 1, 2, 3, corresponding to the three subtasks, α i L represents the weight of the i-th subtask. i (x, xi*) represents the training loss corresponding to the i-th subtask.

[0325] Among them, the model proposed in the optional embodiment of the application can adopt a multi-stage training strategy to accelerate the convergence of the model. First, we train the image denoising sub-task network, and then perform tone enhancement and texture enhancement training. In order to strengthen the interaction between sub-tasks, finally, we use dense connection operation to promote information fusion from different levels, which can alleviate some conflicts existing in these sub-tasks. For example, denoising processing tends to remove some information in the image, while texture enhancement tends to add some details. Therefore, the dense connection operation can provide some image details that may be deleted in the denoising module. In addition, dense connection can also accelerate model training.

[0326] For the training of the model, the learning of the sub-tasks must be supervised, and the true value must be generated for each sub-task. Optionally, a degradation algorithm can be used to establish the training true value for each sub-task in the model. For example, for a given high-quality image GT3 (i.e. the true value for texture enhancement): first, the texture degradation can be applied to obtain GT2 to enhance the tone; then, the brightness and color degradation of GT2 can be processed using image processing tools to obtain GT1 for image denoising; finally, multi-level noise is added to GT1 to obtain the low-quality input image of the proposed model. Optionally, the above degradation processing can be performed in the following way:

[0327] Texture degradation: for the image GT3, bilinear down-sampling and up-sampling with different scale factors can be applied in turn (such as setting the scaling ratio to 2x, 3x and 4x) to obtain GT2, which has the same resolution as GT3 but has degraded texture.

[0328] Brightness and color degradation: GT2 images can be processed using image processing tools to obtain GT1, so that these images look like they were taken in a weak light environment.

[0329] Noise degradation: the noise points in the real noise image are related to the camera sensor. In order to generate real noise, the degradation algorithm provided in the optional embodiment of the foregoing (the method described in the foregoing for obtaining training images) can be used to add noise in the real noise image to the clean image.

[0330] The image quality enhancement method based on deep learning provided in the embodiments of the present application can complete the image quality enhancement task in one step from one or more aspects, for example, for the typical processing combination described above, the image is simultaneously denoised, overall tone is adjusted, and texture is enhanced, based on the scheme, the advantages of each subtask and the optimization combination relationship between tasks can be effectively tapped on the basis of ensuring the enhancement effect of the single subtask and the real-time processing, so that the input image can be fully improved whether the overall tone or the detail information. The image quality enhancement task is decomposed into denoising, brightness enhancement, and texture enhancement from the global and local aspects, and the brightness enhancement corresponds to the global enhancement, and the other parts contribute to the local enhancement.

[0331] The image quality enhancement task is decomposed into denoising, brightness enhancement, and texture enhancement from the global and local aspects, and the brightness enhancement corresponds to the global enhancement, and the other parts contribute to the local enhancement.

[0332] The scheme provided in the embodiments of the present application can improve the image quality from one or more aspects such as denoising, tone enhancement, and texture enhancement, and can significantly improve the problem of image quality degradation caused by the hardware limitation of the image acquisition device (such as a mobile phone), so that the image noise is reduced, has bright colors and tone, and has rich texture detail information, and can effectively meet the requirements of people for image quality.

[0333] In addition, based on the scheme provided in the embodiments of the present application, a lightweight image processing model can be designed to make the image processing model better applicable to a mobile terminal, and specifically, the lightweight model can be realized from one or more of the following aspects:

[0334] In the image denoising model, since more practical prior information (noise intensity features, noise spatial features, etc.) is predicted as the input of the denoising network, the burden of the denoising network can be effectively reduced, and the network structure can be reduced.

[0335] In the tone adjustment aspect, the color branch of the tone adjustment model can be inferred with a smaller size to reduce the inference time, and in addition, with the aid of prior brightness distribution information (such as global brightness information, local brightness information, etc.), a smaller number of channels and layers can also be designed.

[0336] In the texture enhancement aspect, the dual-branch texture enhancement model provides more optimization combination space and gives the model stronger spatial fitting capability, and in each branch, a very light structure can be used.

[0337] In the connection of the processing models, the dense connection mode can be used, and through feature reuse and bypass setting, the network parameters can be greatly reduced.

[0338] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application also provide an image processing device, as shown in the figure, the image processing device 100 can include an image acquisition module 110 and an image processing module 120. Wherein: Figure 8 The image acquisition module 110 is used to acquire a to-be-processed image.

[0339] The image acquisition module 110 is used to acquire a to-be-processed image.

[0340] The image processing module 120 is used to enhance the quality of the to-be-processed image by using at least one image quality enhancement method to obtain a processed image.

[0341] Optionally, the at least one image quality enhancement method includes image denoising, and the image processing module 120 can be used to:

[0342] Obtain the noise intensity feature of the to-be-processed image.

[0343] According to the noise intensity feature, the to-be-processed image is denoised.

[0344] Optionally, the image processing module 120 can be used to:

[0345] According to the noise intensity feature, obtain the noise residual of the to-be-processed image.

[0346] According to the noise residual and the to-be-processed image, obtain the denoised image.

[0347] Optionally, the noise intensity feature of the to-be-processed image includes the noise intensity feature corresponding to each channel image of the to-be-processed image.

[0348] Optionally, the image processing module 120 can be used to:

[0349] Obtain each channel image of the to-be-processed image.

[0350] Respectively obtain the noise intensity feature of each channel image.

[0351] Splice the noise intensity features of each channel image to obtain the noise intensity feature of the to-be-processed image.

[0352] Optionally, the image processing module 120 can be used to: based on each channel image, respectively use a corresponding noise feature estimation network to obtain the noise intensity feature of the corresponding channel image.

[0353] Optionally, the image processing module 120 can be used to:

[0354] obtain a luminance channel image of the to-be-processed image;

[0355] obtain a noise spatial distribution feature of the to-be-processed image according to the luminance channel image;

[0356] perform denoising processing on the to-be-processed image according to the luminance channel image and the noise spatial distribution feature.

[0357] Optionally, the image processing module 120 can be configured to:

[0358] estimate the noise spatial distribution feature of the to-be-processed image using the noise spatial distribution feature estimation network according to the luminance channel image and the noise intensity feature.

[0359] Optionally, the image processing module 120 can be configured to:

[0360] obtain a noise residual of the to-be-processed image according to the noise intensity feature and the to-be-processed image;

[0361] perform weighted processing on the noise residual according to the noise spatial distribution feature to obtain a weighted noise residual;

[0362] obtain a denoised image according to the weighted noise residual and the to-be-processed image.

[0363] Optionally, the at least one image quality enhancement manner includes image brightness adjustment, and when the image processing module 120 performs quality enhancement on the to-be-processed image using the at least one image quality enhancement manner, the image processing module 120 can be configured to:

[0364] determine a brightness enhancement parameter of the to-be-processed image;

[0365] perform brightness adjustment on the to-be-processed image based on the brightness enhancement parameter.

[0366] Optionally, when the image processing module 120 determines the brightness enhancement parameter of the to-be-processed image, the image processing module 120 can be configured to:

[0367] obtain brightness information of the to-be-processed image, and determine the brightness enhancement parameter based on the brightness information;

[0368] obtain brightness adjustment instruction information input by a user, and determine the brightness enhancement parameter of each pixel point of the to-be-processed image based on the instruction information.

[0369] Optionally, the image processing module 120 can be configured to:

[0370] obtain a luminance channel image of the to-be-processed image;

[0371] obtain global brightness information of the to-be-processed image and local brightness information of the to-be-processed image based on the luminance channel image;

[0372] determine the luminance enhancement parameter of each pixel point of the to-be-processed image based on the global luminance information and the local luminance information.

[0373] Optionally, the image processing module 120 can be configured to estimate the semantic-related local luminance information using a local luminance estimation network based on the luminance channel image.

[0374] Optionally, the image processing module 120 can be configured to perform luminance adjustment on the to-be-processed image using a luminance enhancement network according to the luminance enhancement parameter and the luminance channel image of the to-be-processed image.

[0375] Optionally, the at least one image quality enhancement manner includes image color adjustment, and when the image processing module 120 performs quality enhancement on the to-be-processed image using the at least one image quality enhancement manner, the image processing module 120 can be configured to:

[0376] obtain a color channel image of the to-be-processed image;

[0377] perform resolution reduction processing on the color channel image;

[0378] perform color adjustment on the color channel image after the resolution reduction.

[0379] Optionally, the at least one image quality enhancement manner includes image texture enhancement, and when the image processing module 120 performs quality enhancement on the to-be-processed image using the at least one image quality enhancement manner, the image processing module 120 can be configured to:

[0380] obtain a texture enhancement residual and a noise suppression residual of the to-be-processed image using an image texture enhancement network, and fuse the texture enhancement residual and the noise suppression residual to obtain a texture residual;

[0381] obtain a texture-enhanced image according to the texture residual and the to-be-processed image.

[0382] Optionally, the image texture enhancement network includes at least one double convolution module, and one double convolution module includes a first branch for obtaining the texture enhancement residual of the to-be-processed image, a second branch for obtaining the noise suppression residual of the to-be-processed image, and a residual fusion module for fusing the texture enhancement residual and the noise suppression residual to obtain the texture residual.

[0383] Optionally, for a double convolution module, the double convolution module subtracts the texture enhancement residual and the noise suppression residual to obtain the texture residual when fusing the texture enhancement residual and the noise suppression residual;

[0384] Correspondingly, when the image processing module obtains the texture-enhanced image according to the texture residual and the to-be-processed image, the image processing module is configured to:

[0385] Superimpose the texture residual corresponding to each double convolution module and the to-be-processed image to obtain the image after texture enhancement.

[0386] Optionally, for a double convolution module, the first branch includes a first convolution module used to obtain the texture enhancement residual of the to-be-processed image, and a first nonlinear activation function layer used to perform nonlinear processing on the texture residual output by the first convolution module; the second branch includes a second convolution module used to obtain the noise suppression residual of the to-be-processed image, and a second nonlinear activation function layer used to perform nonlinear processing on the noise suppression residual output by the second convolution module; wherein the convolution processing parameters of the first convolution module and the second convolution module are different.

[0387] Optionally, the image texture enhancement network includes at least two double convolution modules, and the convolution network types and / or convolution processing parameters of different double convolution modules are different.

[0388] Optionally, the image texture enhancement network includes at least two double convolution modules based on a dilated convolution network, wherein the dilation rates of the dilated convolution networks of different double convolution modules based on the dilated convolution network are different.

[0389] Optionally, when the image processing module uses the image texture enhancement network to obtain the texture enhancement residual and the noise suppression residual of the to-be-processed image, the image processing module can be used for:

[0390] obtaining a luminance channel image and a non-luminance channel image of the to-be-processed image;

[0391] using the image texture enhancement network to obtain the texture enhancement residual and the noise suppression residual of the to-be-processed image based on the luminance channel image;

[0392] When the image processing module obtains the image after texture enhancement according to the texture residual and the to-be-processed image, the image processing module can be used for:

[0393] obtaining a luminance channel image after texture enhancement according to the texture residual and the luminance channel image;

[0394] fusing the luminance channel image after texture enhancement and the non-luminance channel image to obtain the image after texture enhancement.

[0395] Optionally, the at least one image quality enhancement manner includes at least two enhancement manners, and the image processing module 120 can be used for:

[0396] adopting at least two enhancement manners to perform enhancement processing on the to-be-processed image respectively;

[0397] obtaining the processed image based on the processing results corresponding to each enhancement manner.

[0398] Optionally, the image processing module 120 can be configured to sequentially perform the enhancement processing on the to-be-processed image according to the processing order of the at least two enhancement manners.

[0399] Optionally, the image processing module 120 can be configured to determine the scene information corresponding to the to-be-processed image, and determine the enhancement manner corresponding to the to-be-processed image and the processing order between different enhancement manners according to the scene information.

[0400] Optionally, the image processing module 120 can be configured to determine the scene information of the to-be-processed image based on the to-be-processed image through a scene detection network.

[0401] Optionally, the image processing module 120 can be configured to determine the processing order of the at least two enhancement manners based on the to-be-processed image through a processing order prediction network.

[0402] Optionally, the processing order prediction network comprises a decision branch configured to select a current candidate enhancement manner from the at least two enhancement manners based on input information, and an inference branch configured to determine whether the current candidate enhancement manner is a target enhancement manner, wherein the input information is the to-be-processed image or an enhancement processing result of the to-be-processed image by the enhancement manner with a determined order.

[0403] Optionally, when the image processing module 120 adopts the at least two enhancement manners to perform the enhancement processing on the to-be-processed image, the image processing module 120 can be configured to:

[0404] sequentially perform the enhancement processing on the to-be-processed image according to the processing order of the at least two enhancement manners.

[0405] Optionally, any enhancement manner other than the first enhancement manner is based on the processing result of at least one enhancement processing manner located before the enhancement manner and the to-be-processed image to process the to-be-processed image.

[0406] Optionally, the image quality enhancement manner comprises image denoising, image tone adjustment and image texture enhancement, and the processing order of each enhancement manner is in sequence: image denoising, image tone adjustment and image texture enhancement; wherein the image tone enhancement comprises image brightness adjustment and / or image color enhancement.

[0407] Optionally, the image tone enhancement comprises image brightness adjustment and / or image color enhancement.

[0408] It can be understood that each module of the image quality enhancement apparatus provided in the embodiments of the present application can have the function of implementing the corresponding steps in the image quality enhancement method provided in the embodiments of the present application. The function can be implemented by hardware, or the function can be implemented by hardware executing corresponding software. Each module described above can be software and / or hardware, and each module can be implemented individually or multiple modules can be integrated to implement. The function description of each module of the image quality enhancement apparatus can refer to the corresponding description in the image quality enhancement method in the embodiments described above, and will not be described here again.

[0409] Based on the same principle as the method and apparatus provided in the embodiments of the present application, the embodiments of the present application further provide an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor, when running the computer program, can execute the method provided in any optional solution of the present application.

[0410] Optionally, the electronic device can be a mobile terminal device (such as a smart phone), which can further include an image acquisition apparatus (such as a camera), the image acquisition apparatus being configured to acquire an image and send the acquired image to the processor, and the processor, by running the computer program stored in the memory, can enhance the quality of the image to be processed by using at least one image quality enhancement manner to obtain a processed image.

[0411] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program, when run by a processor, can execute the method provided in any optional solution of the present application.

[0412] In the embodiments provided in the present application, the image processing method executed by the electronic device can be executed by using an artificial intelligence model.

[0413] According to the embodiments of the present application, in the image processing method in the electronic device, the processing method for enhancing the image quality can obtain output data of recognizing an image or image content features in the image by using image data as input data of an artificial intelligence model. The artificial intelligence model can be obtained by training. Here, “obtained by training” means that a basic artificial intelligence model configured to perform a predefined operation rule or an artificial intelligence model of a desired feature (or purpose) is obtained by training an algorithm with multiple training data. The artificial intelligence model can include multiple neural network layers. Each layer of the multiple neural network layers includes multiple weight values, and performs neural network calculation by calculation between the calculation result of the previous layer and the multiple weight values.

[0414] Visual understanding is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.

[0415] In the embodiments provided in the present application, at least one of the plurality of modules can be implemented by an AI model. The functions associated with AI can be performed by a non-volatile memory, a volatile memory, and a processor.

[0416] The processor can include one or more processors. At this time, the one or more processors can be a general-purpose processor (e.g., a central processing unit (CPU), an application processor (AP), etc.), or a pure graphics processing unit (e.g., a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI dedicated processor (e.g., a neural processing unit (NPU))).

[0417] The one or more processors control the processing of input data according to a predefined operation rule or an artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operation rule or the artificial intelligence model is provided by training or learning.

[0418] Here, the provision by learning means that a predefined operation rule or an AI model having a desired characteristic is obtained by applying a learning algorithm to a plurality of learning data. The learning can be performed in the device itself according to the embodiments, and / or can be implemented by a separate server / system.

[0419] The AI model can be composed of a plurality of neural network layers. Each layer has a plurality of weight values, and the calculation of one layer is performed by the calculation result of the previous layer and the plurality of weights of the current layer. Examples of neural networks include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and a deep Q-network.

[0420] The learning algorithm is a method of training a predetermined target device (e.g., a robot) using a plurality of learning data so as to allow or control the target device to determine or predict. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0421] As an example, Figure 9 A structural schematic diagram of an electronic device to which a scheme provided by an embodiment of the present application is applicable is shown in FIG. 1. Figure 9As shown, the electronic device 4000 can include a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, through a bus 4002. Optionally, the electronic device 4000 can further include a transceiver 4004. It should be noted that the transceiver 4004 is not limited to one in actual applications, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0422] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure. The processor 4001 can also be a combination of computing functions, such as one or more microprocessor combinations, combinations of DSP and microprocessor, etc.

[0423] The bus 4002 can include a path for transmitting information between the above-mentioned components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience, Figure 9 Only one thick line is used in the middle, but it does not mean that there is only one bus or one type of bus.

[0424] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing computer instructions or data structures and that can be accessed by a computer, but is not limited thereto.

[0425] The memory 4003 is configured to store computer programs for implementing the solutions of the present application, and the processor 4001 is configured to control the execution of the computer programs stored in the memory 4003. The processor 4001 is configured to execute the computer programs stored in the memory 4003 to implement the content shown in any of the optional method embodiments described above.

[0426] It should be understood that, although each step in the flowchart of the accompanying drawings is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other sequences. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.

[0427] The above only describes some embodiments of the present application, and it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. An image processing method, characterized by, The method comprises: obtaining a to-be-processed image; performing quality enhancement on the to-be-processed image by using at least one image quality enhancement manner to obtain a processed image; wherein the at least one image quality enhancement manner comprises image denoising, and the performing quality enhancement on the to-be-processed image to obtain the processed image comprises: obtaining a plurality of channel images of the to-be-processed image; obtaining noise intensity corresponding to each channel image in the plurality of channel images by using a noise intensity network corresponding to each channel image in the plurality of channel images; performing denoising on the to-be-processed image based on the noise intensity corresponding to each channel image in the plurality of channel images to obtain a denoised image.

2. The method of claim 1, wherein, The performing denoising on the to-be-processed image based on the noise intensity corresponding to each channel image in the plurality of channel images comprises: obtaining noise residuals of the to-be-processed image according to the noise intensity corresponding to each channel image in the plurality of channel images; performing denoising on the to-be-processed image according to the noise residuals.

3. The method of claim 1, wherein, The noise intensity network is a convolutional neural network.

4. The method of claim 1, wherein, The method further comprises: obtaining a luminance channel image of the to-be-processed image; obtaining a noise spatial distribution of the to-be-processed image based on the luminance channel image; The performing denoising on the to-be-processed image based on the noise intensity corresponding to each channel image in the plurality of channel images comprises: performing denoising on the to-be-processed image according to the noise intensity corresponding to each channel image in the plurality of channel images and the noise spatial distribution.

5. The method of claim 4, wherein, The obtaining a noise spatial distribution of the to-be-processed image based on the luminance channel image comprises: obtaining a noise spatial distribution of the to-be-processed image by using a noise spatial feature network based on the luminance channel image.

6. The method according to any one of claims 1 to 4, characterized in that, The obtaining noise intensity corresponding to each channel image in the plurality of channel images comprises: obtaining noise intensity corresponding to each channel image in the plurality of channel images based on a cascaded noise intensity network.

7. The method according to any one of claims 1 to 4, characterized in that, The performing denoising on the to-be-processed image based on the noise intensity corresponding to each channel image in the plurality of channel images comprises: splicing the noise intensity of each channel image to obtain spliced noise intensity; performing denoising on the to-be-processed image according to the spliced noise intensity.

8. The method of claim 7, wherein, The performing denoising on the to-be-processed image according to the spliced noise intensity comprises: obtaining noise residuals by using a denoising network according to the spliced noise intensity; performing denoising on the to-be-processed image according to the noise residuals.

9. The method of claim 8, wherein, The method further comprises: obtaining a luminance channel image of the to-be-processed image; obtaining a noise spatial distribution of the to-be-processed image according to the luminance channel image and the spliced noise intensity; The performing denoising on the to-be-processed image according to the noise residuals comprises: performing denoising on the to-be-processed image according to the noise residuals and the noise spatial distribution.

10. The method according to claim 4 or 9, characterized in that, The obtaining a noise spatial distribution of the to-be-processed image comprises: obtaining a noise spatial distribution of the to-be-processed image by using a noise spatial feature network based on the luminance channel image.

11. The method of claim 9, wherein, The performing denoising on the to-be-processed image according to the noise residuals and the noise spatial distribution comprises: The noise residuals are weighted according to the noise spatial distribution to obtain weighted noise residuals; The image to be processed is denoised according to the weighted noise residuals.

12. The method of claim 11, wherein, The denoised image is obtained, including: The denoised image is obtained by fusing the weighted noise residuals and the image to be processed.

13. The method of claim 8, wherein, The structure of the denoising network is a UNet-like structure.

14. The method of claim 1, wherein, The at least one image quality enhancement manner further includes image brightness adjustment, and the quality of the image to be processed is enhanced by using at least one image quality enhancement manner, including: The brightness enhancement parameter of the image to be processed is determined; The brightness of the image to be processed is adjusted based on the brightness enhancement parameter.

15. The method of claim 14, wherein, The brightness enhancement parameter of the image to be processed is determined, including: The brightness adjustment instruction information input by the user is obtained, and the brightness enhancement parameter of each pixel point of the image to be processed is determined based on the instruction information.

16. The method of claim 14, wherein, The brightness enhancement parameter of the image to be processed is determined, including: The brightness channel image of the image to be processed is obtained; The global brightness information of the image to be processed and the local brightness information of the image to be processed are obtained based on the brightness channel image; The brightness enhancement parameter of each pixel point of the image to be processed is determined based on the global brightness information and the local brightness information.

17. The method of claim 16, wherein, The local brightness information of the image to be processed is obtained based on the brightness channel image, including: The semantic-related local brightness information is estimated using a local brightness estimation network based on the brightness channel image.

18. The method according to any one of claims 14 to 17, characterized in that, The brightness of the image to be processed is adjusted based on the brightness enhancement parameter, including: The brightness of the image to be processed is adjusted using a brightness enhancement network according to the brightness enhancement parameter and the brightness channel image of the image to be processed.

19. The method of claim 1, wherein, The at least one image quality enhancement manner further includes image texture enhancement, and the quality of the image to be processed is enhanced by using at least one image quality enhancement manner, including: The texture enhancement residual and the noise suppression residual of the image to be processed are obtained using an image texture enhancement network, and the texture residual is obtained by fusing the texture enhancement residual and the noise suppression residual. The image after texture enhancement is obtained according to the texture residual and the image to be processed.

20. The method of claim 19, wherein, The image texture enhancement network includes at least one double convolution module, and one double convolution module includes a first branch for obtaining the texture enhancement residual of the image to be processed, a second branch for obtaining the noise suppression residual of the image to be processed, and a residual fusion module for fusing the texture enhancement residual and the noise suppression residual to obtain the texture residual.

21. The method of claim 20, wherein, For a double convolution module, the texture residual is obtained by fusing the texture enhancement residual and the noise suppression residual, including: The texture enhancement residual and the noise suppression residual are subtracted to obtain the texture residual. The image after texture enhancement is obtained according to the texture residual and the image to be processed, including: The texture residual corresponding to each double convolution module and the image to be processed are superimposed to obtain the image after texture enhancement.

22. The method of claim 20, wherein, For a double convolution module, the first branch includes a first convolution module for obtaining a texture enhancement residual of the to-be-processed image, and a first nonlinear activation function layer for performing nonlinear processing on the texture residual output by the first convolution module; the second branch includes a second convolution module for obtaining a noise suppression residual of the to-be-processed image, and a second nonlinear activation function layer for performing nonlinear processing on the noise suppression residual output by the second convolution module; wherein the convolution processing parameters of the first convolution module and the second convolution module are different.

23. The method of any one of claims 20-22, wherein, The convolution network types and / or convolution processing parameters of different double convolution modules in the image texture enhancement network are different.

24. The method of claim 23, wherein, The image texture enhancement network includes at least two double convolution modules based on a cavity convolution network, wherein the dilation rates of the cavity convolution networks of the double convolution modules based on the cavity convolution network are different.

25. The method of any one of claims 20-22, wherein, Using the image texture enhancement network, the texture enhancement residual and the noise suppression residual of the to-be-processed image are obtained, including: Obtaining a luminance channel image and a non-luminance channel image of the to-be-processed image; Using the image texture enhancement network based on the luminance channel image, obtaining the texture enhancement residual and the noise suppression residual of the to-be-processed image; According to the texture residual and the to-be-processed image, obtaining a texture-enhanced image, including: According to the texture residual and the luminance channel image, obtaining a texture-enhanced luminance channel image; Fusing the texture-enhanced luminance channel image and the non-luminance channel image to obtain a texture-enhanced image.

26. The method of claim 1, wherein, The at least one image quality enhancement manner includes at least two enhancement manners, and the at least one image quality enhancement manner is used to enhance the quality of the to-be-processed image to obtain a processed image, including: According to the processing order of the at least two enhancement manners, sequentially performing enhancement processing on the to-be-processed image; Based on the processing result corresponding to each enhancement manner, obtaining a processed image.

27. The method of claim 1, wherein, Further comprising: Determining the scene information corresponding to the to-be-processed image; According to the scene information, determining the enhancement manner corresponding to the to-be-processed image and the processing order between different enhancement manners.

28. The method of claim 27, wherein, Determining the scene information of the to-be-processed image, including: Based on the to-be-processed image, determining the scene information of the to-be-processed image through a scene detection network.

29. The method of claim 26, wherein, Further comprising: Based on the to-be-processed image, determining the processing order of the at least two enhancement manners through a processing order prediction network.

30. The method of claim 29, wherein, The processing order prediction network includes a decision branch for selecting a current candidate enhancement manner from the at least two enhancement manners based on input information, and an inference branch for determining whether the current candidate enhancement manner is a target enhancement manner, wherein the input information is the to-be-processed image or the enhancement processing result of the to-be-processed image by the enhancement manner with a determined order.

31. The method of any one of claims 26-30, wherein, Any enhancement manner except the first enhancement manner is based on the processing result of at least one enhancement processing manner located before the enhancement manner and the to-be-processed image to process the to-be-processed image.

32. The method of any one of claims 26-30, wherein, The image quality enhancement mode includes image denoising, image tone adjustment and image texture enhancement, and the processing sequence of each enhancement mode is in turn: image denoising, image tone adjustment and image texture enhancement.

33. An electronic device, comprising: A computer program product comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the method of any one of claims 1-32 by running the computer program.

34. A computer-readable storage medium, characterized in that, A storage medium stores a computer program, and the computer program, when executed by a processor, executes the method of any one of claims 1-32.

Citation Information

Patent Citations

  • Image processing method, apparatus, electronic device, and computer-readable storage medium

    CN109242794A

  • Image denoising method and device based on deep learning, equipment and storage medium

    CN109658344A

  • Joint noise estimation and image denoising method based on deep learning

    CN109658348A

  • Image denoising method, system and device based on transfer learning and medium

    CN110738605A