Training Method of Image Denoising Model

By using real noise information and image quality scores in the training of image noise reduction model, the problem of poor training effect of image noise reduction model in the prior art is solved, and more efficient image noise reduction and recovery effects are achieved.

CN119722512BActive Publication Date: 2025-06-03HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510228113.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-03
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the prior art, the training effect of the image noise reduction model is poor and cannot be effectively applied to the image noise reduction task of actual scenes. It is mainly because the noise images in the training set are not real, and the model only focuses on image features and ignores noise information.

Method used

A training method for an image noise reduction model is proposed. By acquiring a plurality of first-class sample pairs in the target training set, each first-class sample pair includes a target image and its corresponding noise image. The noise image is generated by specifying noise processing to use noise information of the target noise image. This method combines the initial denoising module and the image content restoration module to use the image quality score of the noisy image to perform feature stitching processing to generate more accurate denoising results.

Benefits of technology

By using real noise information and image quality scores, the noise reduction effect of the image noise reduction model in actual scenes is improved, ensuring that the model can more effectively denoising and recovering image details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722512B_ABST
    Figure CN119722512B_ABST
Patent Text Reader

Abstract

The present application provides a training method for an image denoising model, which relates to the technical field of image processing and includes: obtaining a plurality of first-class sample pairs in a target training set; inputting the noisy image and the corresponding image quality score in each first-class sample pair into the target image denoising model to be trained to obtain the denoising result corresponding to the noisy image in each sample pair; determining the target loss value of the target image denoising model based on the denoising result corresponding to the noisy image in each first-class sample pair and the target image in each first-class sample pair; in response to determining that the target image denoising model has not converged according to the target loss value, adjusting the parameters of the target image denoising model, and returning to the step of obtaining a plurality of first-class sample pairs in the target training set. Through this solution, the denoising effect of the trained image denoising model in the actual scenario can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular, to a method for training an image denoising model. Background Art

[0002] Image denoising is a common processing task in the field of image processing. Among them, image denoising can remove noise in an image, improve the quality and clarity of the image, and thus restore the true information of the image.

[0003] In the related art, usually, a sample pair in a training set is used to train an image denoising model, so as to obtain a trained image denoising model for subsequent use in image denoising. Among them, the sample pair includes a target image and a noise image corresponding to the target image. The training purpose of the image denoising model is that after denoising the noise image in the sample pair, the target image in the sample pair can be obtained. The noise image corresponding to the target image can be considered as an image obtained by adding noise to the target image, and the target image is usually considered as an image without noise.

[0004] However, in the related art, due to the fact that the noise image in the sample pair in the training set is not an image obtained by adding real noise to the target image of the sample pair, or the image denoising model only focuses on the image features of the input image itself during training, etc., the trained image denoising model has a poor denoising effect on images and is not suitable for the image denoising task in the actual scenario. Therefore, there is an urgent need for a method for training an image denoising model to improve the denoising effect of the trained image denoising model in the actual scenario. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide a method for training an image denoising model to improve the denoising effect of the trained image denoising model in the actual scenario. The specific technical solution is as follows:

[0006] In a first aspect, the embodiments of the present application provide a method for training an image denoising model, and the method includes:

[0007] Obtain multiple first - type sample pairs in the target training set; where each first - type sample pair includes a target image and a noise image corresponding to the target image, and the noise image in any first - type sample pair is generated for the target image in the first - type sample pair according to a predetermined noise - image generation method; the predetermined noise - image generation method includes: selecting a spare noise image from multiple spare noise images as the target noise image; determining the noise information to be utilized on each color channel of each pixel point in the target noise image as the noise information to be utilized of the target noise image; performing specified noise addition processing on the target image based on the noise information to be utilized of the target noise image; determining the noise image corresponding to the target image based on the image after the specified noise addition processing; each spare noise image is generated based on the target area in one of the initial images among the initial images; each initial image is an image obtained by photographing a target color card; the target color card includes multiple color patches, and the target area is the image area of one of the multiple color patches;

[0008] Input the noise image and the corresponding image quality score in each first - type sample pair into the target image denoising model to be trained, so as to obtain the denoising result corresponding to the noise image in each sample pair; where the target image denoising model includes an initial denoising module and an image content restoration module connected in series. The initial denoising module is used to process the input image to obtain a feature map of the same size as the input image; the image content restoration module is used to perform feature splicing processing on the feature map obtained by the initial denoising module and the image quality score input into the image content restoration module to obtain the spliced image data; process the spliced image data to generate the denoising result of the image input into the initial denoising module;

[0009] Based on the denoising result corresponding to the noise image in each first - type sample pair and the target image in each first - type sample pair, determine the target loss value of the target image denoising model;

[0010] In response to determining that the target image denoising model has not converged according to the target loss value, adjust the parameters of the target image denoising model and return to the step of obtaining multiple first - type sample pairs in the target training set.

[0011] Advantageous effects of the embodiments of this application:

[0012] A training method for an image denoising model provided by an embodiment of the present application. The noise image in each pair of first-class samples used is an image obtained by performing a specified noise addition process on a target image based on the noise information to be utilized in the target noise image, and the target noise image is generated based on the color block regions in the initial image obtained by photographing a target color card, so that the information of the noise added to the noise image is more consistent with the real noise information. Since the target training set is applicable to the image denoising task (i.e., the data authenticity of the sample pairs is good); and, when training the target image denoising model, the present application takes into account the noise image and the corresponding image quality score. The target image denoising model can utilize more features to adaptively denoise the noise image and can achieve reasonable denoising of the noise image based on the image quality score of the noise image. It can be seen that through this solution, the denoising effect of the trained image denoising model in the actual scenario can be improved.

[0013] Of course, it is not necessary for any product or method implementing the present application to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.

[0015] Figure 1 It is a schematic flowchart of a training method for an image denoising model provided by an embodiment of the present application;

[0016] Figure 2 It is a schematic flowchart of a generation method for any spare noise image provided by an embodiment of the present application;

[0017] Figure 3 It is a schematic overall flowchart of a training method for an image denoising model provided by an embodiment of the present application;

[0018] Figure 4 It is another schematic flowchart of a training method for an image denoising model provided by an embodiment of the present application;

[0019] Figure 5 It is a schematic diagram of a processing flow of a target image denoising model provided by an embodiment of the present application;

[0020] Figure 6 It is a schematic diagram of a processing flow of an image restoration attention module provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.

[0022] In the image acquisition scenario, the noise and clarity of the images captured by the image acquisition device are greatly affected by the module settings in the device ISP (Image Signal Processor) process. The captured images are usually images containing noise and low clarity, and image noise reduction is required for such images. For the image noise reduction algorithm, it is necessary to distinguish the details and noise in the image, remove the noise in the image, restore the details in the image, and obtain clear and natural images.

[0023] In the actual application process, it is very difficult to obtain the sample pairs for model training. This difficulty makes it impossible to effectively train the deep learning model (image noise reduction model), resulting in the deep learning model being unable to achieve an ideal image noise reduction processing effect on the noise images.

[0024] Among them, the sample pairs in the training set usually come from public datasets, manually captured datasets, or image degradation datasets. The types of target images and corresponding noise images in the sample pairs in the training set constructed from public datasets are not applicable to real image denoising tasks. Although the sample pairs in the training set constructed from manually captured datasets are applicable to image denoising tasks, however, due to the requirement for the richness of sample pairs (that is, the more the number and types of sample pairs, the better), the construction efficiency of its training set is very low, that is, the images in the manually captured dataset are difficult to obtain, and the operation is difficult, the complexity of image acquisition is high, the acquisition efficiency is low, and it is time-consuming and laborious (for example: an image of a target scene can be captured by a camera. In some scenes (such as daytime scenes), this image is a noise-free target image. At this time, the camera parameters still need to be adjusted to capture an image with noise of the target scene as the noise image corresponding to the target image; in some other scenes (such as night scenes), this image is a noise image. At this time, it is very difficult to capture a noise-free target image for the target scene by adjusting the camera parameters). The sample pairs in the training set constructed by other image degradation methods are also not applicable to all image denoising tasks. Moreover, the difficulty of image degradation lies in how to generate a real noise image corresponding to the target image (that is, the training set constructed from the image degradation dataset is not applicable to various image denoising tasks. Although the sample pairs in the training set constructed from the image degradation dataset include the target image and the noise image of the target image generated by other image degradation methods, however, the generated noise image is not just an image with noise added to the target image, and the added noise is not real).

[0025] In the image acquisition scenario, the image noise levels of the images captured by the image acquisition device under different gains vary greatly, and are also quite different from the images in the public dataset. For the images with noise captured by the image acquisition device, the image noise levels and noise manifestations obtained under different gains may vary significantly. It is very difficult for the deep learning model to achieve a real and natural processing effect on the noise images in the full gain range (the sample pairs in the public dataset can be used to train the deep learning model, but there are significant differences between the images in the public dataset and the images in the image acquisition scenario, resulting in the deep learning model still being unable to obtain effective training; and the existing deep learning models cannot perform reasonable image denoising on images with various gains and various noise levels).

[0026] Moreover, due to reasons such as the image denoising model only focusing on the image features of the input image itself during training, the trained image denoising model has a poor denoising effect on images, is not applicable to the image denoising tasks in actual scenarios, and cannot perform reasonable image denoising on images with various gains and various noise levels.

[0027] Based on this, the present application provides a method for training an image denoising model to improve the denoising effect of the trained image denoising model in an actual scenario.

[0028] First, a method for training an image denoising model provided by the present application will be introduced below.

[0029] Among them, a method for training an image denoising model provided by the present application can be applied to an electronic device, and this electronic device is used to train the image denoising model. In a specific application, this electronic device can be a mobile phone, a computer, etc. The present application does not limit the specific form of the electronic device. A method for training an image denoising model provided by the present application can be applied to a scenario of image denoising for a noisy image. Based on a first type of sample pair of a target image and a noisy image corresponding to the target image, the image denoising model is trained; moreover, the target image can be an image in an image acquisition scenario, so as to train the image denoising model for the image acquisition scenario.

[0030] Next, with reference to the accompanying drawings, a method for training an image denoising model provided by an embodiment of the present application will be introduced.

[0031] As Figure 1 shown, a method for training an image denoising model provided by an embodiment of the present application may include the following steps:

[0032] S101: Obtain multiple first type of sample pairs in a target training set;

[0033] Among them, each first type of sample pair includes a target image and a noisy image corresponding to the target image. The noisy image in any first type of sample pair is generated for the target image in this first type of sample pair according to a predetermined noisy image generation method; the predetermined noisy image generation method includes: selecting a spare noisy image from multiple spare noisy images as the target noisy image; determining the noise information to be utilized on each color channel of each pixel point in the target noisy image as the noise information to be utilized of the target noisy image; based on the noise information to be utilized of the target noisy image, performing specified noise addition processing on the target image; based on the image after the specified noise addition processing, determining the noisy image corresponding to the target image; each spare noisy image is generated based on a target region in an initial image among various initial images; each initial image is an image obtained by photographing a target color card; the target color card includes multiple color blocks, and the target region is the image region of one color block among the multiple color blocks;

[0034] When training the target image denoising model, multiple first type of sample pairs in the target training set can be obtained first. Each first type of sample pair includes a target image and the noisy image corresponding to the target image; thus, rich first type of sample pairs can be obtained for model training.

[0035] For each target image, the present application can generate one or more noise images of the target image according to a predetermined noise image generation method, so as to construct a target training set and use the target training set to train a target image denoising model; in order to generate a noise image corresponding to the target image, multiple spare noise images are also provided in the present application. One spare noise image can be selected from the multiple spare noise images as the target noise image to be used for adding noise to the target image to generate a noise image corresponding to the target image. Among them, the target image can be an image selected from each noise-free image (which can be an image captured in an image acquisition scenario). An image can be randomly selected from each noise-free image as the target image, and, one spare noise image can be randomly selected from the multiple spare noise images as the target noise image. Different target noise images can be respectively used for adding noise to the same target image to generate multiple noise images of the target image, so as to enrich the number of first-type sample pairs. It should be emphasized that the method of selecting the target image from each clean image and the method of selecting the target noise image from the multiple spare noise images can be the same or different; and, the selection method can include: random selection, sequential selection, etc., and the present application does not limit this.

[0036] It should be noted that each spare noise image is generated based on the target area in one of the initial images among the respective initial images. Each initial image is an image obtained by photographing a target color card. Exemplarily, in an optional implementation manner, the parameter values of the target shooting parameters are different when different initial images are photographed. The target shooting parameters include environmental parameters and / or camera parameters; that is, each spare noise image is generated based on the image area of a color block in one initial image under a target shooting parameter; Exemplarily, each spare noise image can be understood as an image obtained by extracting image data from the image area of a color block in one initial image under a target shooting parameter. By different parameter values of the target shooting parameters and different colors of color blocks, the diversity of the spare noise images is enriched. The present application does not limit the number of color blocks included in the target color card, and any color card existing in the related art can be used as the target color card of the present application. Exemplarily, the target color card can be a 24-color card containing 24 color blocks of different colors. The target color card can be placed in an area (such as an area in a light box), and then parameters such as the light source, imaging size, and focal length are adjusted, that is, the environmental parameters and / or camera parameters are adjusted, to obtain different initial images.

[0037] In addition, the target shooting parameters include environmental parameters and / or camera parameters. The environmental parameters may include lighting conditions, such as the type and intensity of the light source. The type of the light source may include natural light sources, artificial light sources, etc.; the camera parameters may include the internal parameters and external parameters of the camera, etc., such as aperture parameters, focal length parameters, parameters of each module in the ISP process (that is, the image processing process may include multiple processes, each process may be implemented by a module, each module may involve the processing parameters corresponding to the module, and the processing parameters corresponding to each module can be adjusted), etc. The present application does not limit this. Among them, the initial images captured under the parameters of each module in the same ISP process can be understood as the initial images of the same ISP style. Moreover, when the parameters of each module in the ISP process are the same, other parameters (such as aperture parameters, focal length parameters, or the type and intensity of the light source, etc.) may be the same or different.

[0038] It should be noted that since a spare noise image is generated based on the target area in one of the initial images among all the initial images, it does not care whether the initial image contains a background area other than the target color card; that is, the initial image can be captured in any scene (or background) under a stable light source. When shooting, different target shooting parameters can be adjusted to obtain different initial images.

[0039] Among them, the generation method of the spare noise image will be introduced in detail in the subsequent embodiments and will not be elaborated here; moreover, any method capable of generating a spare noise image is applicable to the present application, and the present application does not limit this.

[0040] After determining the target noise image, the noise information to be utilized on each color channel of each pixel point in the target noise image can be further determined as the noise information to be utilized of the target noise image, so as to adapt to the target image for subsequent noise addition processing. Among them, the target image is usually an image in RGB format. The noise information to be utilized on the R (red) channel, G (green) channel, and B (blue) channel of each pixel point in the target noise image can be determined as the noise information to be utilized of the target noise image, which can be used to be added to the target image.

[0041] After determining the noise information to be utilized in the target noise image, based on the noise information to be utilized in the target noise image, a specified noise addition process can be performed on the target image. Exemplarily, in one implementation, the specified noise addition process includes a first type of noise addition process and a second type of noise addition process. The first type of noise addition process is an image processing for adding noise to all pixel points of the image to be processed; the second type of noise addition process includes at least one of a first image processing, a second image processing, and a third image processing. The first image processing is an image processing for adding noise to the edge pixel points of the image to be processed. The second image processing is an image processing for adding noise to the image to be processed according to the brightness and darkness of the pixel points. The third image processing is an image processing for adding noise to the image region of the specified object in the image to be processed.

[0042] In this application, the second type of noise addition process includes a first image processing, a second image processing, and a third image processing (only as an example, the second type of noise addition process may also only include the first image processing, or include the first image processing and the second image processing, and this application is not limited thereto), that is, noise is added through various dimensions to add real noise to the image to be processed. Additionally, the specified noise addition process may also include directly adding the target noise image to the target image (such as adding the pixel values of the pixel points on the target noise image to the corresponding pixel points on the target image), and this application is not limited thereto.

[0043] The target image is a noise-free image, which can be understood as a high-quality image. The noise image corresponding to the target image is the same as the object, the position of the object, the background, etc. included in the target image. The target image is a noise-free high-quality image, and the noise image of the target image is a low-quality noisy image in the same scene (there is a difference in clarity between the target image and the noise image corresponding to the target image (for example, the clarity of the target image is greater than the clarity of the noise image corresponding to the target image), but usually the comparison is not made through clarity; the difference between the two is that the target image is a noise-free high-quality image, and the noise image corresponding to the target image is a low-quality image with added noise to the target image). Among them, any noise image can be used as the target image to generate the noise image corresponding to the image, and this application is not limited thereto. Moreover, after generating the noise image corresponding to the target image, a first type of sample pair can be constructed and stored in the target training set.

[0044] For each target image, through the above steps, one or more noise images of the target image can be generated. Any noise image of the target image can be a noise image obtained through cropping, blurring, adding noise (by performing first-class noise addition processing, first image processing, second image processing, and third image processing on a target noise image), and quality compression (or just adding noise, etc., which is not limited). The target noise images used for adding noise to multiple noise images of the target image are different.

[0045] After that, based on multiple target images and the corresponding noise images, a target training set containing multiple sample pairs can be obtained. The target training set can be used for training an image denoising model; where each sample pair includes a target image and a noise image corresponding to the target image. In different sample pairs, the included target image can be the same target image, and the corresponding noise images are different. For example: in sample pair 1, it includes target image 1 and the noise image 1 corresponding to target image 1; in sample pair 2, it includes target image 1 and the noise image 2 corresponding to target image 1. This application does not make any limitations in this regard.

[0046] S102: Input the noise image and the corresponding image quality score in each first-class sample pair into the target image denoising model to be trained to obtain the denoising result corresponding to the noise image in each sample pair;

[0047] Among them, the target image denoising model includes an initial denoising module and an image content restoration module connected in series. The initial denoising module is used to process the input image to obtain a feature map with the same size as the input image; the image content restoration module is used to perform feature splicing processing on the feature map obtained by the initial denoising module and the image quality score input into the image content restoration module to obtain the spliced image data; process the spliced image data to generate the denoising result of the image input into the initial denoising module.

[0048] The target image in each first-class sample pair can be used as the ground truth for subsequent loss calculation; the noise image in each first-class sample pair can be input into the target image denoising model to be trained. The target image denoising model includes an initial denoising module and an image content restoration module connected in series. The initial denoising module is used to remove extremely large noise (noise blocks) in the image and generate a feature map with the same size as the input image for subsequent denoising processing; the image content restoration module performs feature splicing processing on the feature map obtained by the initial denoising module and the input image quality score, and processes the spliced image data to generate the denoising result corresponding to the image input into the initial denoising module, that is, to obtain the denoising result corresponding to the noise image in each first-class sample pair.

[0049] In this application, a quality score of an image can also be generated. The predetermined noise image generation method further includes:

[0050] In response to the end of any kind of image processing, based on the image quality score of the image processed by this kind of image processing, the predetermined processing parameters used in this kind of image processing, the current image quality degradation parameter corresponding to this kind of image processing, and the current weight parameter set for this predetermined processing parameter, calculate the quality score of the image obtained after this kind of image processing; wherein, the image quality degradation parameter is used to characterize the degradation degree of the image;

[0051] And, after obtaining the quality score of the image obtained by this kind of image processing, based on the quality score of the image obtained by this kind of image processing and the preset prior score, adjust the current image quality degradation parameter and the current weight parameter set for this predetermined processing parameter.

[0052] After the end of any kind of image processing process (any image processing process among the first type of noise addition processing, the first image processing, the second image processing, the third image processing, and the blur processing, etc.), this application can be based on the image quality score of the image processed by this kind of image processing, the current image quality degradation parameter corresponding to this kind of image processing (each image processing process has a corresponding current image quality degradation parameter, that is, the image quality score that this image processing process can currently reduce, for example: the blur processing can currently reduce the image quality score by 0.1, such as: before the blur processing, the image quality score is 1, and after the blur processing, the image quality score is 0.9), the predetermined processing parameters used in this kind of image processing (which can be understood as the degradation parameters used in this image processing process), and the current weight parameter set for this predetermined processing parameter, calculate the quality score of the image obtained after this kind of image processing to characterize the image quality after the end of this kind of image processing. Among them, the image quality degradation parameter can characterize the degradation degree of the image. For any processing process among the first type of noise addition processing, the first image processing, the second image processing, and the third image processing, the image quality degradation parameter of this processing process is used to characterize the noise addition degree of the image; and, the quality score of the image is used to characterize the quality level of the image, and the value range of the quality score is (0, 1), and the image quality score that each processing process can currently reduce is related to the predetermined processing parameter of this processing process.

[0053] Exemplarily, a difference can be calculated for the product of the quality score of the image obtained in the previous image processing process of this image processing process (i.e., the image quality score of the image processed by this kind of image processing, which can be understood as the quality score that the image had before this kind of image processing), the current image quality degradation parameter corresponding to this kind of image processing, the predetermined processing parameter used in this kind of image processing, and the current weight parameter set for this predetermined processing parameter, so as to obtain the quality score of the image after the end of this image processing process. Additionally, in the actual application process, for each degraded processing process, the image quality score that can be reduced by this processing process (i.e., the product of the current image quality degradation parameter corresponding to this kind of image processing, the predetermined processing parameter used in this kind of image processing, and the current weight parameter set for this predetermined processing parameter) can be calculated, and the quality score of the target image is subtracted from the image quality scores that can be reduced by each degraded processing process, so as to obtain the quality score of the noise image corresponding to the target image. For example: setting the quality score of the target image to 1, the quality score of the noise image of the obtained target image can be: 1 - 0.1 - 0.05 - 0.2 - 0.05 = 0.6. It should be noted that any method capable of calculating the quality score of the noise image is applicable to this application, and the above calculation method should not constitute a limitation to this application.

[0054] Exemplarily, the quality score of the image after the end of this image processing process can be calculated according to the following formula:

[0055] ; (1)

[0056] Wherein, and are respectively the image quality scores corresponding to the low-quality degradation data obtained in the degradation step S - 1 and the degradation step S ( i.e., the quality score of the image after the end of this image processing process), is the image quality degradation parameter of the S-th step, is the degradation parameter value corresponding to this degradation step (i.e., the predetermined processing parameter), is the weight coefficient (i.e., the weight parameter of the predetermined processing parameter).

[0057] Moreover, in order to ensure that for images of the same quality, their corresponding quality scores are matched (i.e., to ensure the accuracy of the image quality scores), this application can also, after obtaining the image quality score obtained by this kind of image processing, adjust the current image quality degradation parameter and the current weight parameter set for this predetermined processing parameter based on the image quality score obtained by this kind of image processing and the preset prior score (i.e., for and Adjust (if necessary) so that the quality score of the adjusted image matches the preset prior score. The preset prior score can be understood as the true quality score of the image (prior information representing the quality level of the image data, such as noise level, image visibility information, etc.).

[0058] It can be seen that in this application, for any image processing process, after it ends, the quality score of the image obtained after such image processing can also be generated to characterize the image quality; and, to ensure the accuracy of the image quality score, the image quality degradation parameter and the current weight coefficient set for the predetermined processing parameter can be adjusted; thus, the noise image corresponding to the obtained target image has an accurate, degraded quality score, so as to use more features for training the target image denoising model.

[0059] S103: Determine the target loss value of the target image denoising model based on the denoising result corresponding to the noise image in each first-type sample pair and the target image in each first-type sample pair;

[0060] After obtaining the denoising result of the noise image, the target loss value of the target image denoising model can be determined based on the denoising result corresponding to the noise image in each first-type sample pair and the target image (i.e., the ground truth) in each first-type sample pair.

[0061] Exemplarily, determining the target loss value of the target image denoising model based on the denoising result corresponding to the noise image in each first-type sample pair and the target image in each first-type sample pair includes: calculating a first-type loss value, a second-type loss value, and a third-type loss value based on the denoising result corresponding to the noise image in each first-type sample pair and the target image in each first-type sample pair; determining the target loss value of the target image denoising model based on the first-type loss value, the second-type loss value, and the third-type loss value; where the first-type loss value is the loss value obtained based on the difference in the color dimension between the target image in each first-type sample pair and the denoising result corresponding to the corresponding noise image; the second-type loss value is the loss value obtained based on the absolute value of the difference in pixel values between the target image in each first-type sample pair and the denoising result corresponding to the corresponding noise image; the third-type loss value is the loss value obtained based on the feature difference in the specified feature space between the target image in each first-type sample pair and the denoising result corresponding to the corresponding noise image.

[0062] In this application, considering the differences in the color dimension, pixel value dimension, and specified feature space dimension, based on the denoising results corresponding to the noisy images in each first-class sample and the target images in each pair of first-class samples, the first-class loss value representing the difference in the color dimension, the second-class loss value regarding the absolute value of the difference in pixel values, and the third-class loss value of the feature difference in the specified feature control can be calculated; and based on the first-class loss value, the second-class loss value, and the second-class loss value, the target loss value of the target image denoising model is determined. For example: directly summing the first-class loss value, the second-class loss value, and the second-class loss value, or assigning corresponding weight coefficients to the first-class loss value, the second-class loss value, and the second-class loss value and then summing, etc. This application does not make any limitations in this regard.

[0063] Exemplarily, the target loss value can be calculated according to the following formula:

[0064] ; (2)

[0065] Among them, is the total loss function of the model (i.e., the calculation function of the target loss value), represents the loss function of the entire image, represents the loss function of the image mask region, is the weight occupied by the mask region loss function.

[0066] ; (3)

[0067] ; (4)

[0068] Among them, mask is the mask image of the facial feature region at the edge of the image obtained by the segmentation model Seg_net, is the weight coefficient of the perceptual loss, is the calculation function of the first-class loss value, is the calculation function of the second-class loss value, is the calculation function of the third-class loss value, represents the prediction result (denoising result) of the model, and gt is the target image.

[0069] ; (5)

[0070] Among them, N is the number of samples, represents the prediction result (denoising result) of the model, and gt represents the HQ image, that is, the target image.

[0071] ; (6)

[0072] ; (7)

[0073] Among them, and respectively represent the noise reduction result of the noise image and the feature representation of the target image in the i-th layer of the pre-trained neural network, and M is the number of feature layers. It should be noted that the above method for calculating the target loss value is only an example, and any method capable of calculating the target loss value of the target image denoising model is applicable to this application (for example: calculating the cosine similarity or Euclidean distance between the denoising result corresponding to the noise image and the target image, etc.), and this application does not make any limitations in this regard.

[0074] S104: In response to determining that the target image denoising model has not converged based on the target loss value, adjust the parameters of the target image denoising model, and return to the step of obtaining multiple first-type sample pairs in the target training set.

[0075] After determining the target loss value of the target image denoising model, it is possible to determine whether the target image denoising model has converged based on the target loss value. If it has converged, for example, if the target loss value is less than a preset threshold, the trained target image denoising model can be directly obtained; if it has not converged, for example, if the target loss value is not less than the preset threshold, the target image denoising model needs to be continuously trained. First, the model parameters of the target image denoising model can be adjusted, for example, the model parameters of the initial denoising module and the image content restoration module can be adjusted respectively, and multiple first-type sample pairs in the target training set can be re-obtained for training.

[0076] Among them, for each module, the parameters of the module can be adjusted so that the output result based on the adjusted parameters is closer to the target image in the first-type sample pair. The specific adjustment method of the parameters can be similar to the prior art and will not be elaborated here.

[0077] In the technical solution of this application, operations such as obtaining, storing, using, processing, transmitting, providing, and disclosing the target image, the standby noise image, the noise information, and the noise image corresponding to the target image are all carried out under the premise of obtaining user authorization.

[0078] A training method for an image denoising model provided by an embodiment of the present application. The noise image in each pair of first-class samples used is an image obtained by performing a specified noise addition process on a target image based on the noise information to be utilized of the target noise image, and the target noise image is generated based on the color block regions in the initial image obtained by photographing a target color card. In this way, the information of the noise added to the noise image is more consistent with the real noise information. Since the target training set is suitable for the image denoising task (i.e., the data authenticity of the sample pairs is good); and when training the target image denoising model, the present application takes into account the noise image and the corresponding image quality score. The target image denoising model can use more features to adaptively denoise the noise image and can achieve reasonable denoising of the noise image based on the image quality score of the noise image. It can be seen that through this solution, the denoising effect of the trained image denoising model in the actual scenario can be improved. The information of the noise added to the noise image is more consistent with the real noise information, the authenticity and data richness of the sample pairs in the training set are good, and the data volume is large; and the present application can combine the image quality score corresponding to the noise image to denoise the noise image and can achieve reasonable denoising of images with various noise levels based on the image quality score of the noise image.

[0079] The target training set provided by the present application, compared with the training set constructed from the publicly available dataset, the present application can use various target images and target noise images to generate one or more noise images for each target image, and can generate a training set suitable for adapting to the real image denoising scenario; compared with the training set constructed from the manually photographed dataset, the present application can automatically generate one or more noise images of the target image according to the target image and the target noise image, and the number and types of sample pairs included in the constructed training set are rich, and the construction efficiency of the training set is greatly improved; compared with the training set constructed from the image degradation dataset of other image degradation methods, the present application adds noise to all pixel points of the image to be processed, and adds noise to the edge pixel points, adds noise according to the brightness of the pixel points, and / or adds noise to the image region of the specified object to obtain the real noise image corresponding to the target image; the noise image generated by the present application is a noise image with real noise added to the target image, and the constructed target training set is suitable for the image denoising task. It can be seen that through the present application, a real training set suitable for the image denoising task can be generated, thereby improving the denoising effect of the trained image denoising model in the actual scenario.

[0080] The process of generating the noise image corresponding to the target image, that is, the process of degrading the target image. In addition to adding noise, this process may also include other steps; optionally, in another embodiment of the present application, before performing the specified noise addition process on the target image based on the noise information to be utilized of the target noise image, the predetermined noise image generation method further includes:

[0081] Perform specified preprocessing on the target image to obtain the preprocessed target image; wherein, the specified preprocessing includes: image processing for cropping to obtain an image of a target size and / or image processing for blurring the image;

[0082] Based on the noise information to be utilized in the target noise image, perform specified noise addition processing on the target image, including:

[0083] Based on the noise information to be utilized in the target noise image, perform specified noise addition processing on the preprocessed target image;

[0084] Before performing the specified noise addition processing on the target image, it is also possible to perform specified preprocessing on the target image, that is, crop and / or blur the target image to obtain the preprocessed target image; Exemplarily, based on a predetermined cropping function, such as the crop function (a function for lossless cropping of pictures), the target image can be cropped to obtain an image of a target size, where the target size can be 256, that is, the image of the target size is an image with both length and width of 256; and, based on a blur kernel, the target image or the cropped target image can be blurred to obtain the preprocessed target image. Adaptively, based on the noise information to be utilized in the target noise image, perform specified noise addition processing on the preprocessed target image.

[0085] In one implementation, the target image can be processed for cropping, blurring, noise addition, and subsequent quality compression respectively according to the following formula to obtain the noise image corresponding to the target image:

[0086] ; (8)

[0087] Wherein, is the cropping process, is the convolution operation, K represents the blur kernel for the blurring process, represents the noise addition process (which can include noise addition of the first type of noise addition process, the first image processing, the second image processing, and the third image processing), represents the quality compression in the JPEG format with a quality factor of q (only as an example and not limited), is the high-quality target image, is the noise image corresponding to the degraded target image.

[0088] In this application, considering different target noise images, their noise intensities may be different. The processing methods for performing the first type of noise addition processing based on the noise information to be utilized in the target noise image include:

[0089] Obtain the first noise addition coefficient corresponding to the target noise image; based on the noise information to be utilized in the target noise image, and in accordance with the first noise addition coefficient corresponding to the target noise image, perform noise addition on all pixel points of the image to be processed.

[0090] Among them, the determination method of the first noise addition coefficient corresponding to any alternative noise image includes: based on the pixel value information of each pixel point of the alternative noise image, determine the evaluation value for characterizing the noise intensity of the alternative noise image as the target evaluation value; based on the pre-set corresponding relationship between each evaluation value range and the value range of the noise addition coefficient, determine the value range corresponding to the evaluation value range to which the target evaluation value belongs, and select the value in the determined value range as the first noise addition coefficient corresponding to the alternative noise image; among them, the evaluation value range and the value range in the corresponding relationship are negatively correlated.

[0091] When performing the first type of noise addition processing, the first noise addition coefficient corresponding to the target noise image can be obtained first, and then based on the noise information to be utilized in the target noise image, and in accordance with the first noise addition coefficient corresponding to the target noise image, perform noise addition on all pixel points of the image to be processed (the target image, or the pre-processed target image, that is, the image obtained in the previous processing process of the noise addition process).

[0092] Among them, in this application, based on the noise intensity of any alternative noise image, a corresponding first noise addition coefficient is set. The first noise addition coefficient can characterize the overall addition ratio of the alternative noise image when added to a target image. The value range of the first noise addition coefficient of any alternative noise image can be (0, 1); the evaluation value for characterizing the noise intensity of the alternative noise image can be determined first based on the pixel value information of each pixel point of the alternative noise image. For example, the value of the variance of each pixel point of the alternative noise image is used as the target evaluation value; based on the pre-set corresponding relationship between each evaluation value range and the value range of the noise addition coefficient, determine the value range corresponding to the evaluation value range to which the target evaluation value belongs, and select the value in the determined value range as the first noise addition coefficient corresponding to the alternative noise image; among them, considering the authenticity of noise addition, the higher the target evaluation value of any alternative noise image, that is, the higher the noise intensity, the lower the noise addition coefficient should be, that is, the evaluation value range and the value range are negatively correlated. For example: each evaluation value range includes: (0, T 1 , (T 1 , T 2 , (T 2 , T 3 ……, and the corresponding value range of the noise addition coefficient is [k 1 , k 2 , [k 3 , k 4 , [k5 , k 6 ..., where k1 > k3 > k5 and k2 > k4 > k6.

[0093] In addition, when performing the first type of noise addition processing, the noise information of each pixel in the noise information to be utilized of the target noise image (the size of the target noise image can be the same as the size of the target image or the image of the target size, and the pixel points correspond one by one) can be multiplied by the first noise addition coefficient corresponding to the target noise image to obtain the first type of noise to be added to each pixel point, and then the first type of noise of each pixel point can be added to the pixel point corresponding to the image to be processed, so as to achieve noise addition for all pixel points.

[0094] Exemplarily, the first type of noise addition processing can be performed according to the following formula:

[0095] ; (9)

[0096] Wherein, and respectively represent the images obtained after the degradation step S - 1 and the degradation step S, represents the noise information to be utilized of the target noise image, k represents the proportion of noise addition, that is, the first noise addition coefficient, and i represents the value of each pixel point in the color channel of the target image.

[0097] In this application, in order to add noise more realistically, in addition to adding noise to all pixel points, noise addition can also be performed according to at least one of the first image processing, the second image processing, and the third image processing. Wherein, the specified object can be any object in the image, for example: a facial object, a vehicle object, a tree object, etc. The image area of the facial object can include: the eye area, the nose area, the mouth area, etc. The image area of the vehicle object can include: the headlight area, the window area, etc. It should be noted that for any processing process, the image to be processed is the image obtained from the previous processing process of this processing process. For example: for the first image processing, the image to be processed is the image obtained from the first type of noise addition processing; for the second image processing, the image to be processed can be the image obtained from the first type of noise addition processing or the image obtained from the first image processing (the previous processing process of the processing process of the second image processing can be the processing process of the first noise addition processing or the processing process of the first image processing); and, in this application, the order of the processing processes of the first image processing, the second image processing, and the third image processing is not limited.

[0098] In this application, the noise addition requirements of the edge pixel points, the brightness and darkness, and the image area of the specified object of the image are also considered, and noise is added to the image to obtain a more realistic noise image corresponding to the target image.

[0099] Exemplarily, the present application can perform differential noise addition on an image for edge pixel points, brightness levels, and image regions of a specified object in the image, and perform a first image processing method based on the noise information to be utilized in the target noise image, including:

[0100] Perform edge detection on the target image to obtain the target positions of each pixel point at the edge; wherein, the currently to-be-processed image is an image to be subjected to noise addition processing based on edge information.

[0101] According to the first method or the second method, perform noise addition processing on the to-be-processed image based on the noise information to be utilized in the target noise image; wherein, the first method is: for each pixel point at the target position of the to-be-processed image, based on the noise information to be utilized in each color channel of the pixel point located at the target position in the noise information to be utilized in the target noise image, and according to the second noise addition coefficient, perform noise addition on the pixel point at the target position of the to-be-processed image; the second method is: for each pixel point at a specified position of the to-be-processed image, based on the noise information to be utilized in each color channel of the pixel point located at the specified position in the noise information to be utilized in the target noise image, and according to the second noise addition coefficient, perform noise addition on the pixel point at the specified position of the to-be-processed image, wherein each specified position is a target position among the target positions where the pixel value representing the edge intensity is greater than a predetermined threshold.

[0102] And / or, a processing method for performing second image processing based on the noise information to be utilized in the target noise image, including:

[0103] For each pixel point in the to-be-processed image, based on the noise information to be utilized in each color channel of the specified pixel point in the noise information to be utilized in the target noise image, and according to the third noise addition coefficient and the noise addition principle, perform noise addition on the pixel point, wherein the specified pixel point is the pixel point in the target noise image with the same position as this pixel point, and the noise addition principle is that the noise addition degrees of pixel points with different brightness levels in the target image are different.

[0104] And / or, a processing method for performing third image processing based on the noise information to be utilized in the target noise image, including:

[0105] Perform region detection on the target image for the specified object to obtain a mask image of the specified object in the target image;

[0106] Based on the noise information to be utilized in the target noise image and according to the fourth noise addition coefficient and the obtained mask image, perform noise addition on the pixel points regarding the specified object in the to-be-processed image.

[0107] When performing the first image processing, edge detection can be first performed on the target image (any edge detection algorithm can be used for edge detection) to obtain the target positions of each pixel point on the edge. At this time, the currently to-be-processed image is the image to be noise-added based on the edge information. Then, the to-be-processed image can be noise-added based on the noise information to be utilized of the target noise image in the first manner or the second manner. The first manner is that for each pixel point at each target position of the to-be-processed image, based on the noise information to be utilized of each color channel of the pixel point at the target position in the noise information to be processed of the target noise image, the pixel point at the target position is noise-added according to the second noise-adding coefficient. The second manner is to noise-add the pixel points at the specified positions in the above-mentioned manner, that is, to noise-add the target positions among the various target positions where the pixel points representing the edge intensity are greater than the predetermined threshold. The second manner can achieve differential noise addition for the edge pixel points.

[0108] Among them, the second noise-adding coefficient can represent the addition ratio of the noise information to be utilized of the target noise image regarding the edge when added to the to-be-processed image. When performing the first image processing, regarding the first manner, for each pixel point at each target position, the noise information to be utilized of each color channel of the pixel point at the target position in the noise information to be utilized of the target noise image can be multiplied by the second noise-adding coefficient to obtain the first image noise to be added to each pixel point at the target position, and then the first image noise at each target position is added to the corresponding pixel point of the to-be-processed image. Regarding the second manner, for each pixel point at each specified position, the noise information to be utilized of each color channel of the pixel point at the specified position in the noise information to be utilized of the target noise image can be multiplied by the second noise-adding coefficient to obtain the second image noise to be added to each pixel point at the specified position, and then the second image noise at each specified position is added to the corresponding pixel point of the to-be-processed image.

[0109] Exemplarily, in one implementation manner, the following formula can be used to perform edge detection on the target image:

[0110] ; (10)

[0111] Among them, is an edge detection function, is the edge detection result, is the target image.

[0112] In this application, the following specific manner is adopted to obtain the edge information of the target image:

[0113] ; (11)

[0114] Among them, , ; is a horizontal edge detection operator, is a vertical edge detection operator, is the grayscale image information corresponding to the target image.

[0115] And, the following formula can be adopted to add noise to the first image processing in the first manner:

[0116] ; (12)

[0117] Or, the following formula is adopted to perform differential noise addition on the first image processing in the second manner:

[0118] ; (13)

[0119] ; (14)

[0120] Among them, is the low-quality noise image obtained in this degradation step, is the image obtained from the previous degradation, is the weight coefficient (the second noise addition coefficient) for adding noise to the image edge region, is the calculated image edge information, is the maximum value of the image edge information, used for normalizing the edge signal, is the noise information to be utilized of the target noise image, is the set threshold parameter, is for the judgment result, taking values of 0 or 1, c is the color channel, and (x, y) is the coordinate of the pixel point.

[0121] When performing the second image processing, the image to be processed at this time is the image to be added noise based on the brightness of the pixel points. For each pixel point in the image to be processed, based on the noise information to be utilized of the specified pixel point in each color channel of the target noise image, the pixel point is added noise according to the third noise addition coefficient and the noise addition principle; among them, in this application, the third noise addition coefficient is determined based on the brightness of the target image (or understood as the magnitude of the brightness), and is negatively correlated with the brightness of the target image. That is, the brighter the target image, the lower the third noise addition coefficient, and the darker the target image, the higher the third noise addition coefficient; and, the noise addition degrees of the pixel points with different brightnesses in the target image are different. The higher the brightness, the lower the noise addition degree, and the lower the brightness, the higher the noise addition degree, so as to reasonably add noise to the second image processing.

[0122] Among them, the third noise addition coefficient can characterize the addition ratio of the noise information to be utilized in the target noise image regarding the light and dark levels when added to the image to be processed. When performing the second image processing, for each pixel point in the image to be processed, the noise information to be utilized in each color channel of the specified pixel point in the noise information to be utilized in the target noise image can be multiplied by the third noise addition coefficient respectively to obtain the third image noise of the specified pixel point, and according to the third principle, the third image noise of the specified pixel point is added to the corresponding pixel point in the image to be processed.

[0123] Exemplarily, in one implementation, the noise addition for the second image processing can be performed according to the following formula:

[0124] ; (15)

[0125] Among them, is the low-quality noise image obtained from this degradation step, is the image obtained from the previous degradation step, is the target image, is the weight coefficient for adding noise to the light and dark regions of the image, that is, the third noise addition coefficient.

[0126] When performing the third image processing, at this time, the image to be processed is the image to be added with noise based on the image region of the specified object. First, the region detection of the specified object in the target image can be performed (for example: the region detection of the specified object is performed through a fully convolutional image segmentation neural network) to obtain the mask image (i.e., the mask image, which can be a binary image) of the specified object in the target image; then, based on the noise information to be utilized in the target noise image and according to the fourth noise addition coefficient and the obtained mask image, the pixel points regarding the specified object in the image to be processed are added with noise.

[0127] Among them, the fourth noise addition coefficient can characterize the addition ratio of the noise information to be utilized in the target noise image regarding the image region of the specified object when added to the image to be processed. When performing the third image processing, the noise information of the pixel points belonging to the specified object (determine whether the pixel point belongs to the specified object according to the mask image) in the noise information to be utilized in the target noise image can be multiplied by the fourth noise addition coefficient respectively to obtain the fourth image noise to be added to each pixel point belonging to the specified object, and then the fourth image noise of each pixel point belonging to the specified object is added to the corresponding pixel point in the image to be processed.

[0128] Exemplarily, in one implementation, the noise addition for the third image processing can be performed according to the following formula:

[0129] ; (16)

[0130] ; (17)

[0131] Wherein, is the low-quality noisy image obtained by degrading this step, is the image obtained by degrading the previous step, is the target image (i.e., the high-quality HQ image), is the weight coefficient for adding noise to the facial mask area in the target image, i.e., the fourth noise addition coefficient, is the binary image showing the eye, nose, and mouth areas in the target image (the specified object can be the facial features of a human body, only as an example, not limited), is a deep fully convolutional image segmentation neural network.

[0132] It should be noted that the above-mentioned second noise addition coefficient, third noise addition coefficient, and fourth noise addition coefficient are respectively the noise addition coefficients for adding noise to edge pixel points, brightness and darkness levels, and the image area of the specified object (other pixel points that are not edge pixel points are not added with noise, which can be understood as its second noise addition coefficient being 0; pixel points in the image area that do not belong to the specified object are not added with noise, which can be understood as its fourth noise addition coefficient being 0). The third noise addition coefficient is determined based on the brightness and darkness level of the target image. The second noise addition coefficient, third noise addition coefficient, and fourth noise addition coefficient may be the same or different, and their value ranges can be (0.1, 0.5). Through each noise addition coefficient, reasonable edge noise addition, brightness and darkness noise addition, and / or noise addition to the image area of the specified object can be performed; and any method that can implement the above-mentioned first type of noise addition processing and second type of noise addition processing is applicable to this application, and this application does not limit it.

[0133] After the specified noise addition processing, based on the image after the specified noise addition processing, the noise image corresponding to the target image can be determined.

[0134] Exemplarily, determining the noise image corresponding to the target image based on the image after the specified noise addition processing includes:

[0135] Performing quality compression on the image after the specified noise addition processing to obtain the noise image corresponding to the target image.

[0136] The specified image after noise addition processing can also be subjected to quality compression to obtain the noise image corresponding to the target image. Among them, during quality compression, format conversion and the like can also be performed. The quality compression can be quality compression for a specified format, and the specified format can include: JPEG (Joint Photographic Experts Group, a common image format), PNG (Portable Network Graphics, a lossless compressed bitmap graphic format); quality compression can be performed according to the quality factor of the specified format, and the present application does not limit this.

[0137] Optionally, in another embodiment of the present application, as Figure 2 shown, the generation method of any one of the alternative noise images includes:

[0138] S201: Determine the image area of a color block from an initial image to obtain a target area;

[0139] S202: For each pixel point in the target area, extract the target noise information on each color channel corresponding to the pixel point according to a predetermined extraction method;

[0140] Among them, the predetermined extraction method includes, for each color channel, determining the target noise information of the color channel corresponding to the pixel point based on the value of the color channel of the pixel point and the color mean value corresponding to the color channel; the color mean value corresponding to any color channel is: the value obtained by averaging the values of the color channel of each pixel point in the target area;

[0141] S203: Generate an alternative noise image corresponding to the target area of the initial image based on the target noise information of each color channel corresponding to each pixel point in the target area;

[0142] Among them, the value of each color channel of each pixel point in the alternative noise image is determined based on the target noise information of each color channel corresponding to a pixel point in the target area;

[0143] When generating the alternative noise image, first determine the image area of a color block from an initial image to obtain a target area; for each pixel point in the target area, for each color channel, determine the target noise information of the color channel corresponding to the pixel point based on the value of the color channel of the pixel point and the color mean value corresponding to the color channel, and extract the target noise information on each color channel corresponding to the pixel point; generate an alternative noise image corresponding to the target area of the initial image based on the target noise information of each color channel corresponding to each pixel point in the target area.

[0144] It should be noted that the values of each color channel of each pixel point in the standby noise image are determined based on the target noise information of the corresponding color channels of a corresponding pixel point in the target area. For example, the target noise information of the corresponding color channels of a pixel point in the target area is directly determined as the values of the corresponding color channels of a pixel point in the standby noise image, that is, the target noise information of the R channel, G channel, and B channel corresponding to the pixel point in the target area is respectively determined as the values of the corresponding pixel point in the standby noise image in the R channel, G channel, and B channel.

[0145] Exemplarily, for each color channel, determining the target noise information of the corresponding color channel of the pixel point based on the value of the color channel of the pixel point and the color mean value corresponding to the color channel includes:

[0146] For each color channel, subtracting the color mean value corresponding to the color channel from the value on the color channel of the pixel point to obtain an intermediate value, and summing the intermediate value with a predetermined bias value to obtain the target noise information of the corresponding color channel of the pixel point;

[0147] Determining the noise information to be utilized on each color channel of each pixel point in the target noise image as the noise information to be utilized in the target noise image includes: subtracting the predetermined bias value from the values of each color channel of each pixel point in the target noise image to obtain the noise information to be utilized on each color channel of each pixel point; and determining the noise information to be utilized on each color channel of each pixel point as the noise information to be utilized in the target noise image.

[0148] When determining the target noise information of each pixel point for each color channel, for each color channel, first subtract the color mean value corresponding to the color channel from the value on the color channel of the pixel to obtain an intermediate value. Since the intermediate value may be negative, for convenient storage, the intermediate value can be summed with a predetermined bias value to obtain the target noise information of the corresponding color channel of the pixel point, and store the target noise information of the corresponding color channel of the pixel point. The predetermined bias value is used to correct the intermediate value so that the value of the target noise information of the corresponding color channel of the pixel point after correction conforms to a predetermined range for the preservation of the target noise information.

[0149] Adapted. Since the target noise information of each color channel corresponding to the pixel is noise information added with a predetermined bias value, when adding noise, if it is necessary to utilize the noise information to be utilized on each color channel of the pixel, the value of each color channel of each pixel in the target noise image can be subtracted by the predetermined bias value first to obtain the noise information to be utilized on each color channel of each pixel, and it is determined as the noise information to be utilized of the target noise image (i.e., the real noise information to be added), so as to add the real noise information to be utilized of the target noise image to the target image. Exemplarily, the predetermined bias value can be 128, which can be adjusted according to actual requirements or experience, and the present application does not limit this.

[0150] Exemplarily, in one implementation, the following formula can be used to extract the target noise information:

[0151] ; (18)

[0152] Wherein, represents the value of the pixel of color channel i, represents the target noise information of the pixel in the color channel, i is the data of color channels R / G / B, represents the corresponding color mean value of the color channel, and 128 is the predetermined bias value.

[0153] And, when adding noise, the following formula can be used to calculate the real noise information to be utilized:

[0154] ; (19)

[0155] represents the noise information to be utilized extracted from the color channel of the target noise image, represents the value of the color channel of the target noise image, and 128 is the bias value.

[0156] It can be seen that in the present application, an image area of a color block, that is, a target area, can be determined from an initial image, and the target noise information of each color channel corresponding to each pixel can be extracted to generate a spare noise image of the initial image; and, through the predetermined bias value, the target noise information of each color channel corresponding to each pixel can be stored, and subsequently the bias value can be removed to add the real noise information to be utilized of the target noise image to the target image.

[0157] Optionally, in another embodiment of the present application, the initial denoising module in the target image denoising model is a pre-trained module; the method for pre-training the initial denoising module includes:

[0158] Obtain multiple second sample pairs; input the noisy images in each second sample pair into the initial denoising module to obtain the feature maps output by the initial denoising module corresponding to the noisy images in each second type of sample pair; determine the loss value of the initial denoising module based on the feature maps corresponding to the noisy images in each second type of sample pair and the auxiliary images in each second type of sample pair; in response to determining that the initial denoising module has not converged according to the loss value of the initial denoising module, adjust the parameters of the initial denoising module, and return to the step of obtaining multiple second sample pairs; wherein, each second type of sample pair includes an auxiliary image and the noisy image corresponding to the auxiliary image, and the noisy image corresponding to the auxiliary image is an image obtained by adding noise to the auxiliary image according to the target noise addition method, and the target noise addition method is a method for adding noise with specified particles, and the specified particles can be extra-large particles, that is, the target noise addition method can be a method for adding noise with extra-large particles; and, the method of adding noise with specified particles to the auxiliary image in the target noise addition method can be similar to the method of the first type of noise addition processing, except that the noise addition coefficient of the target noise addition method can be larger, and the present application does not limit this. And, there is a difference in clarity between the auxiliary image and the noisy image corresponding to the auxiliary image (for example, the clarity of the auxiliary image is greater than the clarity of the noisy image corresponding to the auxiliary image), but usually it is not compared by clarity. The auxiliary image is a high-quality image without noise, and the noisy image corresponding to the auxiliary image is a low-quality image with noise added to the auxiliary image. The auxiliary image and the noisy image corresponding to the auxiliary image can be similar to the target image and the noisy image corresponding to the target image. The difference is that the noisy image corresponding to the target image is a noisy image obtained by performing the first type of noise addition processing and the second type of noise addition processing on the target image, and the noisy image corresponding to the auxiliary image is a noisy image obtained by adding noise to the auxiliary image according to the target noise addition method (which can be similar to the first type of noise addition processing).

[0159] In the present application, the initial denoising module can be a pre-trained module. Multiple second sample pairs can be obtained, that is, sample pairs of an auxiliary image and the noisy image corresponding to the auxiliary image. The noisy image corresponding to the auxiliary image is an image obtained by adding noise to the auxiliary image according to the target noise addition method. The target noise addition method is used to add noise with specified particles, and the specified particles can be extra-large particles, that is, the target noise addition method can be a method for adding noise with extra-large particles; and, the method of adding noise with specified particles to the auxiliary image in the target noise addition method can be similar to the method of the first type of noise addition processing. The difference is that the noise addition coefficient of the target noise addition method can be larger, and the present application does not limit this. And, there is a difference in clarity between the auxiliary image and the noisy image corresponding to the auxiliary image (for example, the clarity of the auxiliary image is greater than the clarity of the noisy image corresponding to the auxiliary image), but usually it is not compared by clarity. The auxiliary image is a high-quality image without noise, and the noisy image corresponding to the auxiliary image is a low-quality image with noise added to the auxiliary image. The auxiliary image and the noisy image corresponding to the auxiliary image can be similar to the target image and the noisy image corresponding to the target image. The difference is that the noisy image corresponding to the target image is a noisy image obtained by performing the first type of noise addition processing and the second type of noise addition processing on the target image, and the noisy image corresponding to the auxiliary image is a noisy image obtained by adding noise to the auxiliary image according to the target noise addition method (which can be similar to the first type of noise addition processing).

[0160] Then, input the noisy images in each second sample pair into the initial denoising module. The initial denoising module denoises the input noisy images for extremely large particles and generates corresponding feature maps, that is, obtains the feature maps corresponding to the noisy images in each second type of sample pair; based on the feature maps of the noisy images in each second type of sample pair and the auxiliary images in each second type of sample pair (specifically, the feature maps of the auxiliary images), the loss value of the initial denoising module can be determined; according to this loss value, it can be judged whether the initial denoising module converges. If it does not converge, the parameters of the initial denoising module can be adjusted, and multiple second sample pairs can be obtained again for continued training. Among them, the noisy image corresponding to the auxiliary image in the second type of sample pair can be understood as the noisy image obtained by performing the first type of noise addition processing only on the auxiliary image.

[0161] In addition, the image content restoration module can splice the feature maps obtained by the initial denoising module and the image quality score input to the image content restoration module (this image quality score is the image quality score of the noisy image input to the initial denoising module; that is, the noisy image is input to the initial denoising module, and the image quality score corresponding to the noisy image is input to the image content restoration module) to obtain the spliced image data; the spliced image data can be understood as the feature map containing the features of the image quality score. The image content restoration module performs denoising processing on the noisy image on the premise of considering the image quality score and generates the denoising result of the noisy image.

[0162] It can be seen that for the target image denoising module trained in this application, the image content restoration module can also utilize the feature of the image quality score of the noisy image to restore the content of the noisy image, that is, perform denoising processing, and can reasonably generate the denoising result of the noisy image for noisy images with different image quality scores.

[0163] Optionally, the image content restoration module includes a predetermined attention mechanism module. In the image content restoration module of this application, a predetermined attention mechanism module is included. This attention mechanism module can be an image restoration attention module (Restoration Attention Block, RAB module). The image restoration attention module can include: an up-convolution module, a down-convolution module, a convolution module, and a CBAM module (Convolutional Block Attention Module, the attention mechanism module of the convolution module, is a module that combines channel attention and spatial attention). Its specific structure will be introduced in detail in subsequent embodiments and will not be elaborated here.

[0164] This application designs a target image denoising model based on image score prior information. The target image denoising model includes: a specific noise removal module (initial denoising module) and an image restoration module (image content restoration module). The specific noise removal module can eliminate obvious large granular high-level noise (which can be extremely large noise) in the input noise image, assist the model in removing extremely large noise in the noise image, and generate a feature map corresponding to the noise image after removing the extremely large noise; the image restoration module performs adaptive denoising and restoration of the noise image based on the image quality score corresponding to the input noise image and on the basis of the feature map of the noise image output by the specific noise removal module. Among them, the image quality score (which can be understood as an accurate quality score aligned with the prior information) helps the target image denoising model determine the degree of denoising and restoration of the input noise image, and helps achieve a natural and realistic denoising effect for noise images with various gains and various noise levels.

[0165] As Figure 3 shown, during the training process, first, a dataset 1 (the second type of sample pair) is generated for the target image only by adding extremely large noise (that is, only adding noise of specified particles) to train the specific noise removal module, that is, training part of the module; after the training is completed, for the target image, a dataset 2 (the first type of sample pair) is generated based on a differential degradation scheme (that is, a process including differential noise addition). The LQ image (Low Quality, low-quality image data used to train and evaluate the model, that is, the noise image) in this dataset 2 contains rich degradations and is labeled with corresponding image quality scores. Based on this dataset 2, the entire target image denoising model including the specific noise removal module and the image restoration module is trained, that is, training the entire model, and supervised based on the constructed loss function. The dataset degradation scheme can provide a training dataset containing rich degradations, improve the performance of the target image denoising model, and the calculation of the image quality score in the degradation scheme enables the image quality score to act on the target image denoising model, achieving a natural and realistic denoising effect for noise images with various gains and various noise levels.

[0166] Differential degradation scheme: In order to enable the target image denoising model to better process real noise images and generate real denoised images; in the training dataset (that is, each sample pair) corresponding to the target image denoising model, the LQ image should be as close as possible to the real noise image data, and the HQ image (High Quality, high-quality image data used to train and evaluate the model, that is, the target image) should be as close as possible to the desired effect (that is, as clean as possible and not containing noise). Only in this way can the trained target image denoising model perform ideal image denoising and restoration on real noise images.

[0167] As Figure 4As shown, during the production of the dataset 2, the HQ image (target image) is subjected to operations such as cropping, blurring, adding noise, and quality compression to obtain the corresponding low-quality noisy image (noisy image). At the same time, the quality score of the noisy image is calculated (for each process, the corresponding quality score can be calculated).

[0168] For the HQ image, first, perform size cropping to crop the image to the appropriate size of 256×256 (i.e., the target size). Then, blur the image. Next, add noise to the entire image and perform differential noise addition based on brightness and darkness distinction, edge flat area distinction, specified area, etc. Finally, perform image quality compression to obtain the low-quality noisy image LQ image corresponding to the HQ image.

[0169] The entire process can be expressed by the degradation process as shown in formula (8).

[0170] The specific operation process of the noise addition step in the degradation process is as follows:

[0171] 1. Obtain noise image data: Place the 24-color card in the light box, adjust the stable lighting conditions, place the 24-color card under stable light source conditions, ensure that the pixel size of a single color block in the obtained color card image (initial image) is at least 100×100, and adjust the focal length of the device (such as a camera and other shooting devices) to ensure that the obtained color card image is in clear focus.

[0172] During the capture process, to ensure the diversity of noise image data collection, the light source can be switched to capture noise images under different light source illumination conditions. At the same time, adjust the parameters of each module of the ISP module to capture noise images under different gains, different ISP parameters, and modes, so as to obtain a rich variety of ISP-style noise information covering the full gain level, that is, obtain the initial images of various shooting parameters, environmental parameters, and image signal processing parameters.

[0173] For the captured color card noise image data with different gains, extract the single color block area in the image and extract the noise information on the R / G / B channels respectively (that is, determine the target noise information corresponding to the pixel point in each color channel). Specifically, the target noise information can be extracted according to formula (18). Based on the extracted noise information, large-sized spare noise images can be generated (for example: by splicing, large-sized spare noise images are generated). The obtained spare noise images are the real noise images of various ISP styles, and noise addition processing can be performed on the clean HQ image to generate the corresponding noise images.

[0174] When adding noise to the HQ image, it is necessary to extract the noise information to be utilized from the target noise image, and the real and to-be-utilized noise information can be specifically extracted according to formula (19).

[0175] 2. Adding noise to produce low-quality noise images: Since each spare noise image is extracted and generated under different illuminations, different gains, different ISP parameters, and different color blocks, the noise levels and morphologies of each spare noise image are different. The spare noise images can be grouped according to their noise levels (noise intensities), and different degradation parameters (the first noise addition coefficient) can be set for different groups of spare noise images, so as to avoid the noise images generated by different groups of spare noise images using the same set of degradation parameters from being too different from each other and having too large a distribution gap from the real noise image, resulting in the inability to train the target image denoising model well.

[0176] In order to classify each spare noise image more accurately, the variance of the noise information of the pixel points of each spare noise image data is calculated: ; where N represents the noise information of the pixel points of the spare noise image, represents the variance value of the spare noise image. The larger the value, the higher the noise level of the spare noise image.

[0177] Set the noise level thresholds [T1, T2, T3,...], compare with the values of each threshold, determine the evaluation value range to which the target evaluation value belongs, and according to the corresponding relationship between each evaluation value range and the value range of the noise addition coefficient, determine the value range corresponding to the evaluation value range to which the target evaluation value belongs, and select the value in the determined value range as the first noise addition coefficient corresponding to this spare noise image; in order to group the noise levels of the real noise image data for each spare noise image, the value ranges of the first noise addition coefficient k corresponding to different groups of spare noise images are as follows:

[0178] ;

[0179] After determining the first noise addition coefficient, first, use formula (9) to add the noise information of the whole image to the image.

[0180] By observing and analyzing the real noise image, it can also be found that the noise information in the edge area of the noise image is richer, and the noise in the flat area is relatively weak. Therefore, in the process of degrading the low-quality noise image, in order to make it more conform to the actual low-quality noise image, the present application also distinguishes between the edge and flat areas of the face and adds noise to different degrees.

[0181] First, use the edge detection algorithm shown in formula (10) to obtain the edge information of the target image, and then, based on the judgment of the image edge and non-edge regions, add different levels of noise.

[0182] In addition, edge detection operators such as sobel and canny operators can also be used for calculation. In the actual use of this application, the sobel operator in formula (11) is used to obtain the edge information of the HQ image, and noise is added to the first image processing according to formula (12).

[0183] This application can also perform differential noise addition on images whose pixel values of the image edge information exceed a certain threshold according to formula (13) - formula (14).

[0184] Similarly, the noise in the bright and dark areas of the noisy image will also be different. The noise level in the bright area will be weaker, and the noise in the dark area will be more obvious (i.e., the noise addition principle). Therefore, during the degradation operation, it will also be distinguished according to the brightness and darkness of the image, and different levels of noise will be added according to formula (15). In order to make the noisy image more realistic and natural, the noise addition coefficient is associated with the brightness and darkness of the image itself.

[0185] In order to enable the model to better learn the mapping relationship of the image information in the facial feature area (a specified area, taken as an example here, not limited), find the mask images of the areas such as eyes, nose, and mouth in the face in the image, and perform additional differential noise addition on the facial feature areas of the facial image according to the facial feature mask images. For the clean facial HQ image data, use the facial detection algorithm shown in formula (16) to obtain the corresponding facial feature mask images, and then perform differential noise addition on the facial feature areas according to formula (17). The mask image is a binary image showing the areas of eyes, nose, and mouth in the face in the image.

[0186] Through the above steps, based on the clean and noise-free high-resolution image HQ image, perform degradation to obtain a low-resolution noisy image as the LQ image, and form an image data pair (sample pair) of HQ and LQ as the training data set of the neural model to train the target image denoising model. During the entire noisy image degradation process, at each degradation step, according to the degree of influence of the degradation step on the image quality and the size of the parameters used in the degradation step, calculate the quality score of the low-quality image obtained by this degradation according to formula (1), and align this quality score with the preset score prior.

[0187] As Figure 5As shown in the figure, the neural model structure of this application consists of two parts. One part is the SpecialNoiseReduceBlock, that is, the specific noise removal module (also known as the extremely large noise removal module), and the other part is the image restoration module ImageRestorationBlock based on the prior input of image quality.

[0188] Specific noise removal module: It mainly realizes the removal of extremely large noise in the input noisy image. By inputting the low-quality noisy image into the model, a feature map with extremely large noise particles removed from the input noisy image can be obtained. The specific noise removal module consists of two convolutional blocks (ConvBlock, conv) and two residual blocks (Residualblock, RB) in total; among them, the convolutional kernel sizes in the convolutional block and the residual block are both 3×3. For a noisy image with the size of H×W×C, an output feature map with the same size of H×W×C can be obtained through this module, and the pixel points in the feature map correspond one by one to each pixel point in the input noisy image. The output of this module is used as the direct input of the subsequent image restoration module, which is beneficial for the entire target image denoising model to achieve better denoising and restoration effects for low-quality noisy images with large noise, such as image data with low illuminance and high gain at night, and can remove the extremely large noise in the noisy image completely.

[0189] Image restoration module: Its function is to adaptively output the corresponding image restoration effect according to the input noisy image (feature map) and the quality score of the noisy image, that is, the denoising effect corresponding to the noisy image. The image restoration module inputs by splicing the output feature map of the specific noise removal module and the image quality score, with the input size of H×W×(C + 1) and the output size of H×W×C, that is, the size of the denoising result finally output by the target image denoising model is also H×W×C. The input feature map is first processed by a convolutional module for feature dimension expansion, and then processed by three convolutional downsampling modules (Conv_D) in the encoder part of the image restoration module to extract a feature map of H / 4×W / 4×256. This feature map is sent to the RAB module (RestorationAttentionBlock, image restoration attention module), and the RAB module can help the model better extract and represent features, guiding the model to pay special attention to some feature information in the image. Then, it is processed by three convolutional upsampling modules (Conv_U) in the decoder part for image information restoration, and through the last convolutional module of the model (converting the image data from 64 channels to 3 channels), a denoising result of H×W×C is obtained for output.

[0190] The structure of the RAB module is as Figure 6As shown, the input of this module is the output feature of the last Conv_D convolutional downsampling module in the image restoration module Figure 1 FeatureMap1 (input feature). The feature Figure 1 After entering the RAB module, it is processed in two paths. The first path is first processed by the convolutional module, and then the result obtained and the result of this feature processed by the attention mechanism module CBAM are added together. The second path continues to be processed by the convolutional module and the convolutional downsampling module to obtain a feature map of H / 2×W / 2×2c, and this feature map is also added to the result of its processing by the attention mechanism module CBAM; the feature map obtained by adding the second path is processed by the convolutional module and the convolutional upsampling module to restore the feature map to the size of H×W×C, and then it is concatenated with the feature map added in the first path in the channel dimension to obtain the feature Figure 2 FeatureMap2 (intermediate feature). For this feature Figure 2 It is processed by the convolutional module for output, and the output feature Figure 3 FeatureMap3 (output feature) has a size of H×W×2C. This module can perform cross-scale feature extraction in combination with attention, increasing the ability to extract and express key features, and helping to improve the image restoration effect of the model.

[0191] Training method of the target image denoising model:

[0192] When training the model, split training is required. First, on the clean HQ image, only add extremely large noise with obvious grains to construct clean-noise image pairs (the second type of sample pairs) for the specific noise removal module, and supervise and train the model parameters of the specific noise removal module part based on the L1 loss function.

[0193] After completing the training of the specific noise removal module, import the parameter information of the pre-trained specific noise removal module, and use the clean-noise image dataset (the first type of sample pairs) obtained by the above-mentioned differential degradation to train the parameters of the entire model structure, and design a loss function for supervised learning.

[0194] The total loss function of the entire model is composed of the full-image loss function and the image facial feature mask area loss function, and is trained using the Adam (Adaptive Moment Estimation) optimizer. The total loss function is shown in formula (2).

[0195] It is composed of three parts: the L1 loss function (the calculation function of the second type of loss value), the color loss function (the calculation function of the first type of loss value), and the perceptual loss function (the calculation function of the third type of loss value). It consists of two parts: the L1 loss function and the perceptual loss function. That is, they can be calculated separately according to Formula (3) and Formula (4). and .

[0196] Among them, the color loss function is realized by calculating the cosine similarity between the predicted image and the HQ image, and can be expressed by the following Formula (5).

[0197] The L1 loss is the sum of the absolute values of the differences between the predicted values and the true values, and can be expressed by Formula (6).

[0198] The perceptual loss calculates the difference between two images through a pre-trained neural model. Usually, pre-trained convolutional neural models are used, which have been trained on large-scale datasets and can extract high-level features of images. For example, the convolutional layers in the VGG-19 network (Visual Geometry Group, a convolutional neural network architecture) can extract the texture and structure information of images, and the fully connected layers can extract the semantic information of images.

[0199] The calculation method of the perceptual loss is usually to pass the denoising results of the target image and the noisy image through a pre-trained neural network respectively to obtain their feature representations in the network. Then, these feature representations are used as the input of the loss function, and the Euclidean distance or Manhattan distance between them is calculated. As shown in Formula (7), the perceptual loss can minimize the distance between the denoising result of the noisy image and the target image in the feature space.

[0200] When this application uses a deep learning model for image denoising, ① design a degradation strategy, based on high-quality image data, make low-quality degraded data (the noisy image corresponding to the target image) close to the actual noisy image, and form a dataset for neural network training to solve the problem of lack of training data under specific denoising tasks; ② build a neural network, input the noisy image and the corresponding image quality score of the noisy image at the same time, design a specific noise removal module and an image restoration attention module to help the target image denoising model achieve full-gain segment adaptive denoising and achieve a real and natural restoration effect on various noisy images; ③ construct a loss function to guide the training of the model and achieve the denoising and restoration task of the facial noisy image captured by video acquisition.

[0201] Noisy image pair degradation strategy: ① For the extracted real noise information, group it according to the noise level and set different degradation parameters to ensure that the degraded data contains real noisy images of various noise levels and forms;

[0202] ② Distinguish and add noise to the edge and smooth areas in the image for differential degradation;

[0203] ③Differentiate and add noise according to the bright and dark regions of the image, perform differential degradation, so that the degraded data is more consistent with the real data;

[0204] ④Detect the facial feature region in the image to obtain a mask, and further process the degraded data and the HQ image in the facial feature mask region, so that the model can learn better facial image processing effects on this training data.

[0205] ⑤During the degradation process, calculate the image quality scores corresponding to the low-quality image data obtained in each degradation step, so that the model can learn the association between the image and its quality score, and achieve a natural and realistic noise reduction processing effect for images of different qualities.

[0206] Construction of the neural network: ①Design a specific noise removal module to better remove extremely large flat noises in the image, better remove large particle noises in the image, and improve the image quality; ②Introduce an image restoration attention module in the image restoration module to extract and express features, which helps to improve the processing effect of the model; ③Input the image quality score corresponding to the noisy image into the model to help the model establish a connection between the noisy image and the quality score, and achieve a natural and realistic noise reduction and restoration effect for noisy images at different levels;

[0207] Construction of the loss function: Calculate the full-image loss function and the facial mask region loss function for the denoising result predicted by the model and the HQ image respectively, and weight them to obtain the total loss function of the network. The full-image loss function is calculated by weighting the L1 loss, color loss, and perceptual loss; the facial mask region loss function is calculated by weighting the L1 loss and perceptual loss functions of the denoising result and the mask region of the HQ image. Using this loss function to supervise the learning process of the model helps the model to obtain a better image noise reduction and restoration effect, especially in the facial region of the image.

[0208] A training method for an image denoising model provided by this application includes dataset production, network structure design, and loss function construction. It can perform different degrees of noise reduction and restoration according to the quality level of the noisy image itself, achieve effective image restoration in the full gain range, restore image details, and improve image quality.

[0209] This application presents an object image denoising model based on image score prior information and noisy image degradation, which can achieve a better restoration effect in the full gain range for real noisy images (for the image acquisition scenario, this noisy image is the noisy image in the image acquisition scenario), and the restored image is natural and realistic.

[0210] The foregoing are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.

Claims

1. A training method for an image denoising model, characterized in that: The method comprises: Acquire multiple first-category sample pairs in a target training set; wherein each first-category sample pair includes a target image and a noise image corresponding to the target image, and the noise image in any first-category sample pair is generated for the target image in the first-category sample pair according to a predetermined noise image generation method; the predetermined noise image generation method includes: selecting a spare noise image from multiple spare noise images as a target noise image; determining noise information to be used on each color channel of each pixel in the target noise image as noise information to be used for the target noise image; performing a specified noise addition process on the target image based on the noise information to be used for the target noise image; determining the noise image corresponding to the target image based on the image after the specified noise addition process; each spare noise image is generated based on a target area in an initial image among the initial images; each initial image is an image obtained by photographing a target color card; the target color card includes multiple color blocks, and the target area is an image area of ​​a color block among the multiple color blocks; Inputting the noisy image and the corresponding image quality score in each first-category sample pair into the target image denoising model to be trained to obtain the denoising result corresponding to the noisy image in each sample pair; wherein the target image denoising model comprises an initial denoising module and an image content restoration module connected in series, the initial denoising module is used to process the input image to obtain a feature map of the same size as the input image; the image content restoration module is used to perform feature splicing processing on the feature map obtained by the initial denoising module and the image quality score input to the image content restoration module to obtain spliced ​​image data; the spliced ​​image data is processed to generate the denoising result of the image input to the initial denoising module; Determine a target loss value of the target image denoising model based on a denoising result corresponding to the noise image in each first-category sample pair and a target image in each first-category sample pair; In response to determining that the target image denoising model has not converged according to the target loss value, adjusting parameters of the target image denoising model, and returning to the step of obtaining multiple first category sample pairs in the target training set.

2. The method according to claim 1, characterized in that The target shooting parameters used when shooting different initial images have different parameter values, and the target shooting parameters include environment parameters and / or camera parameters; The designated noise addition processing includes a first type of noise addition processing and a second type of noise addition processing. The first type of noise addition processing is image processing for performing noise addition on all pixels of the image to be processed; the second type of noise addition processing includes at least one of a first image processing, a second image processing and a third image processing; the first image processing is image processing for performing noise addition on edge pixels of the image to be processed, the second image processing is image processing for performing noise addition on the image to be processed according to the brightness of the pixels, and the third image processing is image processing for performing noise addition on an image area of ​​a designated object of the image to be processed.

3. The method according to claim 2, characterized in that Based on the noise information to be utilized of the target noise image, a first type of noise addition processing method includes: Acquire a first noise addition coefficient corresponding to the target noise image; Based on the noise information to be utilized of the target noise image and according to the first noise adding coefficient corresponding to the target noise image, adding noise to all the pixel points of the image to be processed; The method for determining the first noise addition coefficient corresponding to any spare noise image includes: Based on the pixel value information of each pixel point of the backup noise image, determining an evaluation value for characterizing the noise intensity of the backup noise image as a target evaluation value; Based on a preset correspondence between each evaluation value range and a value range of the noise addition coefficient, determine a value range corresponding to the evaluation value range to which the target evaluation value belongs, and select a value in the determined value range as a first noise addition coefficient corresponding to the backup noise image; Among them, the evaluation value range and the value range in the corresponding relationship are negatively correlated.

4. The method according to claim 2, characterized in that: The method of performing first image processing based on the noise information to be utilized of the target noise image includes: Performing edge detection on the target image to obtain the target position of each pixel point at the edge; wherein the current image to be processed is an image to be subjected to noise addition processing based on edge information; According to the first method or the second method, based on the noise information to be used of the target noise image, the image to be processed is subjected to noise processing; wherein the first method is: for each pixel point at the target position of the image to be processed, based on the noise information to be used of each color channel of the pixel point located at the target position in the noise information to be used of the target noise image, and according to the second noise addition coefficient, the pixel point at the target position of the image to be processed is subjected to noise processing; the second method is: for each pixel point at the designated position of the image to be processed, based on the noise information to be used of each color channel of the pixel point located at the designated position in the noise information to be used of the target noise image, and according to the second noise addition coefficient, the pixel point at the designated position of the image to be processed is subjected to noise processing, wherein each designated position is a target position among the target positions, the pixel value representing the edge strength of which is greater than a predetermined threshold value; and / or, The method of performing second image processing based on the noise information to be utilized of the target noise image includes: For each pixel in the image to be processed, based on the noise information to be used of each color channel of the designated pixel in the noise information to be used of the target noise image, and according to the third noise addition coefficient and the noise addition principle, the pixel is noised, wherein the designated pixel is a pixel in the target noise image at the same position as the pixel, and the noise addition principle is that the noise addition degrees of pixels with different brightness in the target image are different; and / or, The method of performing the third image processing based on the noise information to be utilized of the target noise image includes: Performing region detection on the target image with respect to a designated object to obtain a mask image of the designated object in the target image; Based on the noise information to be utilized of the target noise image, and according to the fourth noise addition coefficient and the obtained mask image, the pixel points related to the designated object in the image to be processed are denoised.

5. The method according to claim 2, characterized in that: The predetermined noise image generation method also includes: In response to the completion of any image processing, a quality score of an image obtained after the image processing is calculated based on an image quality score of the image processed by the image processing, a predetermined processing parameter used by the image processing, a current image quality degradation parameter corresponding to the image processing, and a current weight parameter set for the predetermined processing parameter; wherein the image quality degradation parameter is used to characterize a degree of image degradation; as well as, After obtaining the quality score of the image obtained by the image processing, the current image quality degradation parameter and the current weight parameter set for the predetermined processing parameter are adjusted based on the quality score of the image obtained by the image processing and a preset priori score.

6. The method according to claim 1 or 2, characterized in that: Any alternate noise image can be generated by: From an initial image, determine an image area of ​​a color block to obtain a target area; For each pixel point in the target area, extract the target noise information on each color channel corresponding to the pixel point according to a predetermined extraction method; wherein the predetermined extraction method includes, for each color channel, determining the target noise information of the color channel corresponding to the pixel point based on the value of the color channel of the pixel point and the color mean value corresponding to the color channel; the color mean value corresponding to any color channel is: the value obtained by averaging the values ​​of the color channel of each pixel point in the target area; Based on the target noise information of each color channel corresponding to each pixel point of the target area, a backup noise image corresponding to the target area of ​​the initial image is generated; wherein the value of each color channel of each pixel point in the backup noise image is determined based on the target noise information of each color channel corresponding to a pixel point of the target area.

7. The method according to claim 6, characterized in that The step of determining, for each color channel, target noise information of the color channel corresponding to the pixel point based on the value of the color channel of the pixel point and the color mean value corresponding to the color channel, includes: For each color channel, subtract the color mean value corresponding to the color channel from the value of the pixel on the color channel to obtain an intermediate value, and sum the intermediate value with a predetermined offset value to obtain the target noise information of the color channel corresponding to the pixel; The step of determining the noise information to be utilized on each color channel of each pixel in the target noise image as the noise information to be utilized of the target noise image includes: Subtracting the predetermined offset value from the value of each color channel of each pixel in the target noise image to obtain the noise information to be used on each color channel of each pixel; The noise information to be utilized on each color channel of each pixel is determined as the noise information to be utilized of the target noise image.

8. The method according to any one of claims 1 to 5, characterized in that: The initial denoising module in the target image denoising model is a pre-trained module; The method of pre-training the initial denoising module includes: Acquire multiple second sample pairs; wherein each second type sample pair includes an auxiliary image and a noise image corresponding to the auxiliary image, wherein the noise image corresponding to the auxiliary image is an image after the auxiliary image is noised according to a target noise adding method, wherein the target noise adding method is a method for adding noise of specified particles; Inputting the noise image in each second sample pair into the initial denoising module to obtain a feature map corresponding to the noise image in each second type sample pair output by the initial denoising module; Determine a loss value of an initial denoising module based on a feature map corresponding to the noise image in each second-class sample pair and an auxiliary image in each second-class sample pair; In response to determining that the initial denoising module has not converged according to the loss value of the initial denoising module, adjusting parameters of the initial denoising module, and returning to the step of obtaining a plurality of second sample pairs.

9. The method according to any one of claims 1 to 5, characterized in that: The step of determining a target loss value of the target image denoising model based on a denoising result corresponding to the noise image in each first-category sample pair and a target image in each first-category sample pair includes: Based on the denoising result corresponding to the noise image in each first-category sample pair and the target image in each first-category sample pair, a first-category loss value, a second-category loss value, and a third-category loss value are calculated; Determining a target loss value of the target image denoising model based on the first type of loss value, the second type of loss value and the third loss value; The first type of loss value is a loss value obtained based on the difference in color dimension between the denoising results corresponding to the target image and the corresponding noise image in each first type sample pair; The second type of loss value is a loss value obtained based on the absolute value of the difference between the pixel values ​​of the denoising results corresponding to the target image and the corresponding noise image in each first type of sample pair; The third type of loss value is a loss value obtained based on the feature difference in a specified feature space between the denoising results corresponding to the target image and the corresponding noise image in each first type sample pair.

10. The method according to any one of claims 1 to 5, characterized in that: Before performing the designated noise adding process on the target image based on the noise information to be utilized of the target noise image, the predetermined noise image generating method further includes: Performing a specified preprocessing on the target image to obtain a preprocessed target image; wherein the specified preprocessing includes: image processing for cropping an image of a target size and / or image processing for blurring the image; The step of performing a specified noise adding process on the target image based on the noise information to be utilized of the target noise image comprises: Based on the noise information to be utilized of the target noise image, performing a specified noise adding process on the preprocessed target image; The step of determining the noise image corresponding to the target image based on the specified image after noise addition processing includes: The quality of the image after the specified noise processing is compressed to obtain a noise image corresponding to the target image.

Citation Information

Patent Citations

  • Image enhancement method, device and equipment

    CN113436112A

  • Training method of image denoising model, image denoising method and related device

    CN118657677A