An image noise reduction method and an image noise reduction module
Through deep learning model and dual subnet training, the adaptability problem of image noise under different illumination is solved, and efficient image noise reduction effect and delicate image quality protection are achieved.
Patent Information
- Application Number
- CN202110067919.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-01-19
AI Technical Summary
Existing image noise reduction algorithms are difficult to take into account image noise at different illuminances, and deep learning noise reduction algorithms require high adaptability to scenes, so they need to design network structures and training parameters specifically.
Image noise reduction is performed using deep learning models. The sample data includes target image data under different illuminances, and the model is trained with a dual sub network and a multi-dimensional mixed loss function, and the target image and the area of interest are respectively reduced.
It improves the adaptability of image noise reduction, can take into account noise at different illumination, enhances image quality, protects image boundary details, and improves training efficiency and convenience of parameter adjustment.
Smart Images

Figure CN114862685B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular, to an image noise reduction method. Background Art
[0002] In the field of video surveillance, higher and higher requirements are imposed on the image effect of imaging devices. Severe image noise will affect the visual intuitive experience.
[0003] For the currently widely used traditional image noise reduction algorithms, such as the BM3D (Block-Matching and 3D filtering) algorithm, its principle is as follows:
[0004] First, a frame of image is segmented into smaller pixel patches. After selecting a reference patch, search for pixel patches similar to the reference patch, and form the first similar block with the similar pixel patches;
[0005] Then perform a 3D transform on all the first similar blocks to obtain 3D blocks;
[0006] Screen the 3D blocks according to a set threshold to remove image noise and obtain the denoised 3D blocks;
[0007] Perform an inverse 3D transform on the denoised 3D blocks to obtain the second similar block;
[0008] Finally, perform weighted averaging on all the second similar blocks and restore them to the image.
[0009] The core of this algorithm lies in the adoption of different noise reduction strategies, that is, by searching for similar blocks and filtering in the transform domain, weighting each pixel point in the second similar block, so as to obtain the final noise reduction effect.
[0010] Whether it is the BM3D algorithm or other various traditional image noise reduction algorithms, it is very difficult to take into account the image noise at various different illuminations.
[0011] With the development of deep learning technology, deep learning noise reduction algorithms have been greatly developed in recent years. However, deep learning noise reduction algorithms have high requirements for scene adaptability, and it is necessary to specifically design a reasonable network structure according to the actual usage situation and obtain appropriate parameters through certain training means. Summary of the Invention
[0012] The present invention provides an image noise reduction method to reduce the image noise in images under different illuminations.
[0013] An image noise reduction method provided by the present invention includes:
[0014] Performing j times of convolutional encoding on the image to be noise-reduced in sequence to obtain a feature encoding;
[0015] Perform j - times of convolutional decoding on the feature encoding in sequence to obtain a denoised image with the same pixel size as the image to be denoised;
[0016] Wherein,
[0017] Perform the first - time convolutional decoding on the feature encoding to obtain the result of this convolutional decoding,
[0018] Perform the next convolutional decoding adjacent to this convolutional decoding on the convolutional encoding result having the same pixel size as the result of this convolutional decoding and the result of this convolutional decoding, and recursively perform in sequence until the j - th convolutional decoding is performed,
[0019] j is a natural number greater than or equal to 1.
[0020] Preferably, the method further includes,
[0021] Extract the region of interest in the image based on the image to be denoised to obtain a region - of - interest image,
[0022] Use the region - of - interest image as the image to be denoised, and execute the steps of performing j - times of convolutional encoding on the image to be denoised in sequence to obtain a feature encoding and performing j - times of convolutional decoding on the feature encoding in sequence to obtain a denoised image of the region - of - interest image;
[0023] Fuse the denoised image of the image to be denoised and the denoised image of the region - of - interest image in the image to be denoised;
[0024] Before performing j - times of convolutional encoding on the first input image to be denoised in sequence, it further includes,
[0025] Perform downsampling on the image to obtain m downsampled images,
[0026] The step of performing j - times of convolutional encoding on the first input image to be denoised in sequence to obtain a feature encoding includes,
[0027] Perform convolutional encoding on the image to be denoised and m downsampled images in parallel respectively to obtain m + 1 feature encodings, wherein each image performs j - times of convolutional encoding in sequence according to the set convolutional encoding step length, and the feature encoding is the feature encoding after j - times of convolutional encoding,
[0028] The step of performing j - times of convolutional decoding on the feature encoding in sequence to obtain a denoised image with the same pixel size as the image to be denoised includes,
[0029] Perform convolutional decoding on m + 1 feature encodings in parallel respectively to obtain the images after convolutional decoding, wherein each feature encoding performs j - times of convolutional decoding in sequence according to the same convolutional decoding step length,
[0030] Upsample the image obtained by convolving and decoding each downsampled image to obtain m upsampled images.
[0031] Fuse the image obtained by convolving and decoding the image to be denoised and the m upsampled images to obtain a denoised image with the same pixel size as the image to be denoised.
[0032] Where m is a natural number greater than or equal to 1.
[0033] Preferably, performing the next convolution decoding adjacent to the current convolution decoding on the convolution encoding result having the same pixel size as the current convolution decoding result and the current convolution decoding result, and recursively performing the operation until the j-th convolution decoding is performed, includes:
[0034] Performing a second convolution decoding on the first convolution decoding result obtained from the first convolution decoding and the (j - 1)-th convolution encoding result obtained from the (j - 1)-th convolution encoding to obtain a second convolution decoding result.
[0035] Performing a third convolution decoding on the second convolution decoding result and the (j - 2)-th convolution encoding result obtained from the (j - 2)-th convolution encoding to obtain a third convolution decoding result.
[0036] Recursively performing the operation
[0037] Until performing the j-th convolution decoding on the (j - 1)-th convolution decoding result and the first convolution encoding result obtained from the first convolution encoding to obtain the j-th convolution decoding result.
[0038] Preferably, the sequentially performing j times of convolution encoding includes: at each time of convolution encoding, sequentially performing convolution and enhanced residual processing.
[0039] The sequentially performing j times of convolution decoding includes: at each time of convolution decoding, sequentially performing enhanced residual and deconvolution processing.
[0040] Where the enhanced residual processing includes q times of convolution processing, and q is a natural number greater than or equal to 1.
[0041] Preferably, q is an even number, and the q times of convolution processing includes:
[0042] Sequentially performing q times of convolution on the first input data for enhanced residual processing.
[0043] Where
[0044] The first input data sequentially performs the first convolution and the second convolution to obtain a second convolution result.
[0045] The second convolution result and the first input data sequentially perform the third convolution and the fourth convolution to obtain a fourth convolution result.
[0046] The fourth convolution result, the second convolution result, and the first input data are sequentially subjected to a fifth convolution and a sixth convolution to obtain a sixth convolution result.
[0047] By successive recursion,
[0048] The result of the p-th convolution, the result of the p - 2k-th convolution, and the first input data are sequentially subjected to a (p + 1)-th convolution and a (p + 2)-th convolution to obtain a (p + 2)-th convolution result.
[0049] …
[0050] Until p + 2 equals q, a q-th convolution result is obtained;
[0051] p is an even number, p - 2k is greater than 1, and k is a natural number greater than or equal to 1.
[0052] Preferably, m is 3 or 4, j is a natural number from 2 to 6, and q is a natural number from 4 to 8;
[0053] The size of the convolution encoding step is such that: the pixel size of the convolution encoding result obtained by this convolution encoding is half of the pixel size of the convolution encoding result obtained by the previous convolution encoding.
[0054] The size of the (j - 1)-th convolution decoding step is such that: the pixel size of the convolution decoding result obtained by this convolution decoding is twice the pixel size of the convolution decoding result obtained by the previous convolution decoding.
[0055] The size of the j-th convolution decoding step is such that: the pixel size of the convolution decoding result obtained by this convolution decoding is the same as the pixel size of the input image.
[0056] The present invention also provides a noise reduction module for image noise reduction. This noise reduction module is used for:
[0057] Subjecting the image to be noise-reduced to j times of convolution encoding in sequence to obtain a feature encoding;
[0058] Subjecting the feature encoding to j times of convolution decoding in sequence to obtain a noise-reduced image having the same pixel size as the image to be noise-reduced;
[0059] Wherein,
[0060] Performing a first convolution decoding on the feature encoding to obtain a result of this convolution decoding,
[0061] Performing the next convolution decoding adjacent to this convolution decoding on the convolution encoding result having the same pixel size as the result of this convolution decoding and the result of this convolution decoding, and by successive recursion until the j-th convolution decoding is performed.
[0062] j is a natural number greater than or equal to 1.
[0063] Preferably, the noise reduction module is further configured to:
[0064] Extract a region of interest in the image based on the image to be noise-reduced to obtain a region-of-interest image,
[0065] Use the region-of-interest image as the image to be noise-reduced, perform noise reduction processing to obtain a noise-reduced image of the region-of-interest image;
[0066] Fuse the noise-reduced image of the image to be noise-reduced with the noise-reduced image of the region-of-interest image in the image to be noise-reduced;
[0067] The noise reduction module is a trained deep learning model module, and the deep learning model module includes a first sub-network for denoising a target image and a second sub-network for denoising a region of interest in the target image.
[0068] Wherein,
[0069] The first sub-network processes the image to be noise-reduced into a noise-reduced image of the image to be noise-reduced;
[0070] The second sub-network processes the region-of-interest image extracted based on the image to be noise-reduced into a noise-reduced image of the region-of-interest image,
[0071] The first sub-network and the second sub-network have the same network structure,
[0072] Each sub-network includes m channels, and each channel includes j convolutional encoding modules and j convolutional decoding modules.
[0073] In each channel: the j convolutional encoding modules are connected in sequence, and the j convolutional decoding modules are connected in sequence. Among them, the j-th convolutional encoding result of the j-th convolutional encoding module is output to the first convolutional decoding module, and any one of the remaining convolutional decoding modules except the first convolutional decoding module is connected to any one of the remaining convolutional encoding modules except the j-th convolutional encoding module, and the convolutional encoding result of this convolutional encoding module has the same pixel size as the other input of this convolutional decoding module;
[0074] m is a natural number greater than or equal to 1.
[0075] Preferably, the noise reduction module further includes,
[0076] A region-of-interest extraction module for extracting a region-of-interest image in the image to be noise-reduced,
[0077] A sub-network image fusion module for fusing the noise-reduced image from the first sub-network with the noise-reduced image from the second sub-network;
[0078] Any convolutional decoding module among the remaining convolutional decoding modules except the first convolutional decoding module is connected to any convolutional encoding module among the remaining convolutional encoding modules except the j-th convolutional encoding module, including,
[0079] The j-1 convolutional encoding result of the j-1 convolutional encoding module is input into the second convolutional decoding module,
[0080] The j-2 convolutional encoding result of the j-2 convolutional encoding module is input into the third convolutional decoding module,
[0081] By successive recursion,
[0082] until the first convolutional encoding result of the first convolutional encoding module is input into the j-th convolutional decoding module.
[0083] Preferably, the convolutional encoding module includes a convolutional sub-module and an enhanced residual sub-module connected in sequence, and the convolutional decoding module includes an enhanced residual sub-module and a deconvolutional sub-module connected in sequence,
[0084] wherein,
[0085] The enhanced residual sub-module includes q convolutional layers connected in sequence,
[0086] wherein,
[0087] The first input data input into the enhanced residual sub-module is successively input into the first convolutional layer and the second convolutional layer to obtain a second convolutional result,
[0088] The second convolutional result and the first input data are successively input into the third convolutional layer and the fourth convolutional layer to obtain a fourth convolutional result,
[0089] The fourth convolutional result, the second convolutional result, and the first input data are successively input into the fifth convolutional layer and the sixth convolutional layer to obtain a sixth convolutional result,
[0090] By successive recursion,
[0091] The result of the p-th convolution, the result of the p-2k-th convolution, and the first input data are successively input into the p+1 convolutional layer and the p+2 convolutional layer to obtain a p+2 convolutional result,
[0092] …
[0093] until p+2 is equal to q to obtain a q convolutional result;
[0094] p is an even number, p-2k is greater than 1, and k is a natural number greater than or equal to 1.
[0095] The present invention also provides an image denoising method based on a deep learning model, which includes, on the side of the trained deep learning model,
[0096] The trained deep learning model receives the image to be denoised and processes it to obtain a denoised image;
[0097] Wherein,
[0098] The deep learning model is trained using training sample data that at least includes target image data collected under different illuminations.
[0099] Preferably, the training sample data includes target image pairs composed of target input images and target reference images, and color card image pairs composed of color card input images and color card reference images.
[0100] Wherein,
[0101] The input image is a short-frame image, and the reference image is a long-frame image. The short-frame image and the long-frame image are collected according to an exposure cycle including a first exposure time and a second exposure time, and the first exposure time is not equal to the second exposure time. The image with a longer exposure time is the long-frame image, and the image with a shorter exposure time is the short-frame image.
[0102] The pixel size of the target input image is the same as that of the color card input image.
[0103] The pixel size of the target reference image is the same as that of the color card reference image.
[0104] Preferably, within two adjacent exposure cycles, the exposure alternation order of the first exposure time and the second exposure time is the same.
[0105] The image pairs are intercepted from the collected video data. The intercepted long-frame images and the intercepted short-frame images have the same time information, and the displacement difference is less than or equal to a set displacement difference threshold.
[0106] The input images and reference images in the image pairs respectively have set target frames. Among them, the positions of the target frames of the input images and the positions of the target frames of the reference images are respectively offset according to the same perturbation parameters. The perturbation parameters include a random direction in one of the four directions of up, down, left, and right of the image and a random offset amount.
[0107] Preferably, the positions of the target frames in the target input image and the positions of the target frames in the target reference image are respectively offset according to the first perturbation parameter;
[0108] The positions of the target frames in the color card input image and the positions of the target frames in the color card reference image are respectively offset according to the second perturbation parameter.
[0109] Preferably, among the training sample data, the number of image pairs under different illuminations and different image exposure gains tends to be the same. Among them, the image exposure gain starts from the initial image exposure gain and increases in equal steps until it reaches the set image exposure gain threshold, and the image exposure gain threshold is determined according to the image quality;
[0110] The images included in the target image pair are single target images, and the single target images are respectively the images in the target front view direction, the left view direction, the right view direction, the top view direction, and the bottom view direction; the target moves according to the set moving speed threshold.
[0111] Preferably, the target is a human face, and the image acquisition device for collecting the human face is installed at a height position with a certain height from the ground. There is a set distance between the collected target and the projection position of the image acquisition device on the ground, and the distance between the pupils in the collected human face image reaches the set pixel number threshold.
[0112] The image acquisition device has different device models.
[0113] Preferably, the deep learning model includes
[0114] A first sub-network for denoising the target image and a second sub-network for denoising the region of interest in the target image.
[0115] Among them,
[0116] The image to be denoised is input into the first sub-network, and the first sub-network outputs the first denoised image;
[0117] The deep learning model further includes
[0118] The region of interest image extracted from the image to be denoised is input into the second sub-network, and the second sub-network outputs the second denoised image.
[0119] Preferably, the first sub-network and the second sub-network have the same network structure, and each sub-network includes m channels.
[0120] The image to be denoised is downsampled into m - 1 downsampled images. The image to be denoised and the m - 1 downsampled images are respectively input into the m channels in the first sub-network. The output images of each downsampled image from the channels are upsampled to obtain m - 1 upsampled images. The output image of the image to be denoised from the channel and the m - 1 upsampled images are fused to obtain the first denoised image;
[0121] The image of the region of interest is downsampled into m - 1 downsampled images. The image of the region of interest and these m - 1 downsampled images are respectively input into m channels in the second sub - network. The output images of each downsampled image from the channels are upsampled to obtain m - 1 upsampled images. The output image of the region - of - interest image from the channel and these m - 1 upsampled images are fused to obtain a second denoised image;
[0122] m is a natural number greater than or equal to 1;
[0123] The first denoised image and the second denoised image are fused, and the fused image is the final denoising result.
[0124] Preferably, each channel includes j convolutional encoding modules and j convolutional decoding modules,
[0125] In each channel: the j convolutional encoding modules are connected in sequence, and the j convolutional decoding modules are connected in sequence. Among them, the j - th convolutional encoding result of the j - th convolutional encoding module is output to the first convolutional decoding module, and for any convolutional decoding module among the remaining convolutional decoding modules except the first convolutional decoding module, it is connected to any convolutional encoding module among the remaining convolutional encoding modules except the j - th convolutional encoding module, and the convolutional encoding result of this convolutional encoding module has the same pixel size as the other input of this convolutional decoding module.
[0126] Preferably, for any convolutional decoding module among the remaining convolutional decoding modules except the first convolutional decoding module, being connected to any convolutional encoding module among the remaining convolutional encoding modules except the j - th convolutional encoding module includes,
[0127] The (j - 1)-th convolutional encoding result of the (j - 1)-th convolutional encoding module is input to the second convolutional decoding module,
[0128] The (j - 2)-th convolutional encoding result of the (j - 2)-th convolutional encoding module is input to the third convolutional decoding module,
[0129] By successive recursion,
[0130] until the first convolutional encoding result of the first convolutional encoding module is input to the j - th convolutional decoding module;
[0131] j is a natural number greater than or equal to 1.
[0132] Preferably, the convolutional encoding module includes a convolutional sub - module and an enhanced residual sub - module connected in sequence, and the convolutional decoding module includes an enhanced residual sub - module and a transposed convolutional sub - module connected in sequence,
[0133] The enhanced residual sub - module includes q convolutional layers connected in sequence,
[0134] Among them,
[0135] The first input data of the input enhancement residual sub-module is sequentially input into the first convolutional layer and the second convolutional layer to obtain a second convolutional result.
[0136] The second convolutional result and the first input data are sequentially input into the third convolutional layer and the fourth convolutional layer to obtain a fourth convolutional result.
[0137] The fourth convolutional result, the second convolutional result, and the first input data are sequentially passed through the fifth convolutional layer and the sixth convolutional layer to obtain a sixth convolutional result.
[0138] By successive recursion,
[0139] The result of the p-th convolution, the result of the (p - 2k)-th convolution, and the first input data are sequentially passed through the (p + 1)-th convolutional layer and the (p + 2)-th convolutional layer to obtain a (p + 2)-th convolutional result.
[0140] …
[0141] Until p + 2 equals q, a q-th convolutional result is obtained;
[0142] p is an even number, p - 2k > 1, and k is a natural number greater than or equal to 1.
[0143] Preferably, the deep learning model is trained in the following manner:
[0144] Receive the current batch of training sample data input externally.
[0145] Calculate the multi-dimensional mixed loss function of the current batch of training sample data.
[0146] When the current multi-dimensional mixed loss function reaches the set loss threshold, end the training, save the network parameters in the current deep learning model, and obtain the trained deep learning model. Otherwise, receive the next batch of training sample data and return to execute the step of calculating the multi-dimensional mixed loss function of the current batch of training sample data.
[0147] Among them,
[0148] The number of target image pairs and the number of color card image pairs included in each batch of training sample data are random numbers.
[0149] The multi-dimensional mixed loss function is the sum of the loss function of the first sub-network and the loss function of the second sub-network.
[0150] The loss function of the first sub-network is: the first accumulation result of the difference between the value 1 and the structural similarity index of each channel in the first sub-network, and the sum of the minimum absolute deviation losses of each channel in the first sub-network.
[0151] The loss function of the second sub-network is: the second accumulation result of the difference between the value 1 and the structural similarity index of each channel in the second sub-network, plus the sum of the minimum absolute deviation losses of each channel in the second sub-network;
[0152] The structural similarity index is obtained by calculating the similarity mean of the training sample data of the current batch based on the denoised image and the image without noise points.
[0153] The minimum absolute deviation loss is obtained by calculating the mean of the minimum absolute deviation losses of the training sample data of the current batch based on the denoised image and the image without noise points.
[0154] Preferably, the first accumulation result is the first weighted result of accumulation, and the first weighted result is: the weighted result of the difference between the value 1 and the structural similarity index of each channel in the first sub-network weighted by the first weighting coefficient of each channel in the first sub-network;
[0155] The sum of the minimum absolute deviation losses of each channel in the first sub-network is the second weighted result of accumulation, and the second weighted result is: the weighted result of the minimum absolute deviation loss of each channel in the first sub-network weighted by the second weighting coefficient of each channel in the first sub-network;
[0156] The second accumulation result is the third weighted result of accumulation, and the third weighted result is: the weighted result of the difference between the value 1 and the structural similarity index of each channel in the second sub-network weighted by the third weighting coefficient of each channel in the second sub-network,
[0157] The sum of the minimum absolute deviation losses of each channel in the second sub-network is the fourth weighted result of accumulation, and the fourth weighted result is: the weighted result of the minimum absolute deviation loss of each channel in the second sub-network weighted by the fourth weighting coefficient of each channel in the second sub-network.
[0158] The present invention also provides a deep learning model module for image denoising. The deep learning model module is used to process the image to be denoised through the trained deep learning model to obtain a denoised image;
[0159] Among them,
[0160] The deep learning model is trained using training sample data that at least includes target image data collected under different illuminations.
[0161] The image denoising method provided by the present invention can perform denoising processing on the image to be denoised in a manner based on a deep learning model, by using a deep learning model trained with training sample data including at least target image data collected under different illuminations. It has high adaptability to the scene, can take into account the denoising of image noise under different illuminations, improves the denoising ability of differential noise, and improves the image quality. Further, the dual sub-networks adopted by the deep learning model enable training according to the multi-dimensional mixed loss function of each channel of each sub-network during the training process, which is beneficial to improving the training efficiency and facilitating the parameter adjustment of the deep learning model parameters. The trained deep learning model performs denoising on the whole and the region of interest of the target image respectively. In this way, not only the overall denoising effect of the target image is improved, but also the image in the key region of the target image is denoised, avoiding taking the boundary details in the image as noise and making the image denoising process more delicate, which is beneficial to the protection of the boundary details of the image. The fused denoised image is beneficial to further improving the image quality. Description of the Drawings
[0162] Figure 1 It is a schematic diagram of the network structure of the deep learning model in an embodiment of the present application.
[0163] Figure 2 It is a schematic diagram of a general enhanced residual processing.
[0164] Figure 3 It is a schematic diagram of the network structure of a deep learning model including multiple channels.
[0165] Figure 4 It is a schematic diagram of image acquisition.
[0166] Figure 5 It is a schematic diagram of adding perturbation to the target box in the image.
[0167] Figure 6 It is a schematic diagram of the structure of the deep learning network model of the present application.
[0168] Figure 7 It is a schematic diagram of the region of interest of the human face.
[0169] Figure 8 It is a schematic diagram of the network architecture of a sub-network.
[0170] Figure 9 It is a schematic diagram of a convolutional encoding module.
[0171] Figure 10 It is a schematic diagram of a convolutional decoding module.
[0172] Figure 11 It is a schematic diagram of the enhanced residual sub-module in this embodiment.
[0173] Figure 12 It is a schematic diagram of a device for training a deep learning model.
[0174] Figure 13 It is a schematic diagram of a deep learning model module for image denoising. Specific implementation manners
[0175] In order to make the objectives, technical means and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings.
[0176] Under different illuminances, the noise points in the captured images have different degrees of noise, and the morphological distributions of the noise points are also different, showing great differences, which not only affect the visual effect but also reduce the accuracy of image recognition. The applicant found that when using traditional image processing methods for image denoising, it is easy to remove the boundary details of the targets in the images as noise points, resulting in poor protection of details. In view of the fact that the deep learning model for image denoising does not require manual feature design, and the large number of training samples it depends on can also be easily obtained through monitoring devices, it has advantages in solving the problem of image denoising.
[0177] For the deep learning model for image denoising in the present application, training sample data including at least the target image data captured under different illuminances is used for training, and the network parameters of the trained deep learning model are saved. The image to be denoised is input into the trained deep learning model for processing to obtain a denoised image, thereby being able to reduce the image noise in images under different illuminances.
[0178] In order to obtain better denoising effects, the embodiment of the present application also improves the network structure of the deep learning model itself. Refer to Figure 1 as shown Figure 1 It is a schematic diagram of the network structure of the deep learning model in the embodiment of the present application. In this network structure, the image to be denoised is sequentially subjected to j times of convolutional encoding to obtain a feature encoding; the feature encoding is sequentially subjected to j times of convolutional decoding to obtain an output image having the same pixel size as the image to be denoised;
[0179] wherein,
[0180] The first convolutional decoding is performed on the feature encoding to obtain the result of this convolutional decoding,
[0181] The convolutional encoding result having the same pixel size as the result of this convolutional decoding and the result of this convolutional decoding are subjected to the next convolutional decoding adjacent to this convolutional decoding, and so on until the jth convolutional decoding is performed,
[0182] j is a natural number greater than or equal to 1.
[0183] In each convolutional coding process, convolution and enhanced residual processing are performed sequentially. In each convolutional decoding, enhanced residual and deconvolution processing are performed sequentially. Refer to Figure 2 as shown in Figure 2 It is a schematic diagram of enhanced residual processing. The first input data for enhanced residual processing is sequentially convolved q times, where
[0184] the first input data is sequentially convolved for the first time and the second time to obtain a second convolution result,
[0185] the second convolution result and the first input data are sequentially convolved for the third time and the fourth time to obtain a fourth convolution result,
[0186] the fourth convolution result, the second convolution result, and the first input data are sequentially convolved for the fifth time and the sixth time to obtain a sixth convolution result,
[0187] and so on by recursion,
[0188] the result of the p-th convolution, the result of the p - 2k-th convolution, and the first input data are sequentially convolved for the p + 1-th time and the p + 2-th time to obtain a p + 2 convolution result,
[0189] …
[0190] until p + 2 equals q to obtain a q convolution result,
[0191] p is an even number, p - 2k is greater than 1, and k is a natural number greater than or equal to 1.
[0192] Preferably, in order to obtain a high-dimensional feature vector of a to-be-denoised image with low-dimensional information, the length and width dimensions of the input image can be reduced and the thickness can be increased through channel stretching, so that the to-be-denoised image can be downsampled to obtain m downsampled images.
[0193] Refer to Figure 3 as shown in Figure 3 It is a schematic diagram of the network structure of a deep learning model including multiple channels. The to-be-denoised image is downsampled by a downsampling module into m downsampled images. The to-be-denoised image and the m downsampled images are respectively and concurrently subjected to convolutional coding to respectively obtain m + 1 feature codings. The m + 1 feature codings are respectively and concurrently subjected to convolutional decoding to respectively obtain images after convolutional decoding. Among them, each feature coding is sequentially subjected to j times of convolutional decoding according to the same convolutional decoding step size. The images after convolutional decoding of each downsampled image can be upsampled by an upsampling module to obtain m upsampled images. The image after convolutional decoding of the to-be-denoised image and the m upsampled images are fused by a channel image fusion module to obtain an output image with the same pixel size as the to-be-denoised image;
[0194] Wherein, m is a natural number greater than or equal to 1.
[0195] It should be understood that the number of downsampling modules and upsampling modules can be m respectively. In this way, each sampling module performs corresponding processing respectively; alternatively, the number of downsampling modules and upsampling modules can be 1 respectively.
[0196] The image denoising method provided by this application can adapt to the denoising processing of images under different illuminations. The improvement of the network structure of the deep learning model itself not only helps to improve the training efficiency but also improves the image denoising performance of the deep learning model itself.
[0197] The following takes the denoising of face images as an example for illustration. It should be understood that this application is not limited to face images, and any image including other content can also be applicable, such as license plate images, fingerprint images, and so on.
[0198] In order to obtain high-quality training sample data for the deep learning model used for denoising, face image data is collected in a near-real scenario. Since the target to be collected is a face, the image acquisition device is fixed at a position at a set height from the ground, and a face at a set distance from the projection position of the image acquisition device on the ground is collected. The focal length of the image acquisition device is adjusted so that the distance between the pupils in the face image reaches a set pixel number threshold. See Figure 4 As shown, the image acquisition device is installed at a height position of 2 to 3 meters from the ground, and the distance between the target to be collected and the projection position of the image acquisition device on the ground is 10 - 15 meters. The focal length of the image acquisition device is adjusted so that the distance between the pupils in the face image reaches 30 - 40 pixels.
[0199] Preferably, the image acquisition device is adjusted to work in a wide dynamic intensity mode. For example, the wide dynamic intensity can be set to 50 db; image data is collected according to the exposure cycle. Among them, within the same exposure cycle, it includes a first exposure time and a second exposure time. Preferably, two adjacent exposure cycles alternate according to the first exposure time and the second exposure time. For example, within the first exposure cycle, image data is collected in the order of the first exposure time and the second exposure time, and within the second exposure cycle adjacent to the first exposure cycle, image data is also collected in the order of the first exposure time and the second exposure time. In other words, the entire collection process is carried out in an alternating manner according to the first exposure time and the second exposure time. In this way, within the same exposure cycle, a first image collected at the first exposure time and a second image collected at the second exposure time are obtained respectively. Among them, the first image and the second image have the same time information, and the image brightness is as close as possible.
[0200] The first exposure time is not equal to the second exposure time. For ease of description, the image with a longer exposure time is called a long-frame image, and the image with a shorter exposure time is called a short-frame image. Since the sum of the exposure times of the long- and short-frame images cannot exceed a set exposure time threshold, for example, 1 / 25 s, in actual image data acquisition, the exposure time of the long-frame image is 2 to 4 times that of the short-frame image. For example, the exposure time of the long-frame image is 1 / 50 s, and the exposure time of the short-frame image is less than or equal to 1 / 100 s. In this way, the difference in image exposure gain between the long-frame image and the short-frame image can be maintained between 6 - 12 db.
[0201] To simulate the angles of various human faces in a real scene, image data of 5 view directions of a single human face in a single scene are collected, which are the images in the front view direction, left view direction, right view direction, top view direction, and bottom view direction of the human face, respectively, so as to obtain a front face image, a left face image, a right face image, a top face image, and a bottom face image respectively.
[0202] During the data acquisition process, if the human face moves too fast, motion blur will occur in the long-frame image of the human face. In the short-frame image, due to a certain displacement difference of moving objects, a slowly moving human face is beneficial to minimizing the displacement difference as much as possible. Therefore, it is necessary to keep the moving speed of the human face reach a set moving speed threshold, for example, the moving speed of the human face is about 0.01 m / s, so as to ensure the clarity of the human face in the long-frame image, reduce the displacement difference in the short-frame image, and improve the learning efficiency of the deep learning model. Among them, the displacement difference refers to the difference in pixel coordinate positions between the pixel points of the object contour in the long-frame image and the pixel points of the object contour in the short-frame image of the moving object.
[0203] Since the noise manifestation forms under different illuminances are affected by the ambient light intensity, therefore, long- and short-frame image data under different illuminances are collected. For example, long- and short-frame image data under different illuminances are collected by adjusting the light source in the scene. During the actual acquisition process, starting from the initial image exposure gain, the image exposure gain can be collected in equal step increments until the set image exposure gain threshold is reached. Among them, the image exposure gain threshold is determined according to the image quality. This is because when the image exposure gain is large, the face in the short frame is severely damaged. Especially in the case of human body movement, the facial features are less recognizable, and there is a serious trailing effect. Such training samples will result in limited noise reduction effect of the model trained later.
[0204] To enrich the sample data, preferably, face image data is collected using image acquisition devices of different models. To prevent the network parameters from being biased towards the illuminance with a larger quantity during the training process and having a negative impact on the effects of other illuminances, the quantity ratios of the collected images are made to approach consistency under different illuminances and different image gains. For example, the following table shows that the quantities of the collected images are the same under different illuminances and different image exposure gains, where the quantities of the long-frame and short-frame images are the same:
[0205]
[0206] To enable color fidelity of the face images during the noise reduction process, color card image data needs to be added to the training sample data. The same device as that used for collecting face data is adopted to collect color card images in the same manner as the face image collection, that is, the image acquisition device operates in the wide dynamic intensity mode and collects color card images under different illuminances and different image exposure gains according to the exposure cycle composed of the first exposure time and the second exposure time.
[0207] Preprocess the collected image data:
[0208] For the collected video data, intercept the video frames corresponding to the long and short frames. Before interception, the video frames of the long and short frames can be adjusted to the same moment. Since the displacement difference between the long and short frames will cause the trained model to learn the displacement difference feature, and thus the image to be noise-reduced (short-frame image) may have local area deformation, it is necessary to avoid intercepting the image data with a displacement difference greater than the set displacement difference threshold, that is, the displacement difference between the intercepted long and short frames is less than or equal to the set displacement difference threshold;
[0209] For the noise reduction processing task of face images, the collected long-frame face images and short-frame face images correspond one by one according to the time information, that is, the long-frame image at a certain time corresponds to the short-frame image at that time. Mark the position of the target box (face box). For example, the range of the face box can be set to 300 * 300 pixels. To increase the richness of the training sample data, add random perturbations to the position of the face box. The rule is: the position of the face box randomly selects one of the four directions of up, down, left, and right for offset, and the offset amount is a random value in the pre-set array. See Figure 5 shown Figure 5 A schematic diagram of adding perturbations to the target box in the image. The same first perturbation parameter is used for the face boxes in each long-frame face image and its corresponding short-frame face image, that is, the first perturbation parameter is simultaneously assigned to the face boxes in the long-frame face image and its corresponding short-frame face image. The advantage of such enhanced data is that the background information corresponding to each face is non-overlapping, reducing the possibility of the deep learning model quickly overfitting during the training process;
[0210] For the collected color card images, the same processing is performed as for the collected face images, that is, the collected long-frame color card images and short-frame color card images correspond one by one according to the time information, and the positions of the target boxes (color card boxes) are marked according to the face frame. The same second perturbation parameter is used for the target boxes in each long-frame color card image and its corresponding short-frame color card image, so that the positions of the target boxes are offset in a random direction and by a random offset amount in one of the four directions of up, down, left, and right.
[0211] In this way, after preprocessing, training sample data is obtained. The training sample data includes: long-frame face images and their corresponding short-frame face images, as well as long-frame color card images and their corresponding short-frame color card images. Among them, the pixel size of the color card images is the same as that of the face images, and the target boxes in the images are perturbed using random perturbation parameters, and the perturbation parameters include a random direction in one of the four directions of up, down, left, and right, and a random offset amount.
[0212] Based on the training sample data, the long-frame face image is used as the target reference image, its corresponding short-frame face image is used as the target input image, the long-frame color card image is used as the color card reference image, and its corresponding short-frame color card image is used as the color card input image. The target input image and the target reference image are a target image pair, and the color card input image and the color card reference image are a color card image pair. A batch of image data is input into the deep learning model to train the deep learning model. The number of target image pairs and color card image pairs included in the same batch of image data is random, and the number of image pairs included in each batch of image data can be different. For example, the number of image pairs in batch 1 is 16, and the number of image pairs in batch 2 is 8.
[0213] See Figure 6 shown Figure 6 is a schematic structural diagram of a deep learning network model of the present application. The deep learning network model includes two sub-networks, namely the second sub-network (FACE-ROI-Net) for noise reduction of the face region of interest and the first sub-network (FACE-Net) for face noise reduction. Among them, the input of FACE-Net is the image to be denoised, that is, the original face image data (including face image pairs and color card image pairs), and the output is the face image after noise reduction by the FACE-Net network; the input of FACE-ROI-Net is the ROI region image extracted based on the image to be denoised, that is, the original face region of interest image data, and the output is the face region of interest image after noise reduction by the FACE-ROI-Net network.
[0214] Preferably, the denoised face image (the first denoised image) and the denoised face region of interest image (the second denoised image) are fused to further improve the denoising effect. As Figure 6As shown, the first denoised image and the second denoised image are fused through the sub-network image fusion module, and the fused image is the final denoised image.
[0215] See Figure 7 As shown, Figure 7 It is a schematic diagram of a face region of interest. The face region of interest can be located according to facial feature points. The dots in the figure are face detection feature points, and the boxes are ROI regions. The ROI regions can select the eye, nose, and mouth regions according to the feature points of the facial features, and different weights can be given to different ROI regions. Or the whole face can be selected according to the feature points of the whole face contour, or other selection strategies can be adopted.
[0216] The FACE-ROI-Net sub-network and the FACE-Net sub-network adopt the same network architecture. See Figure 8 As shown, Figure 8 It is a schematic diagram of the network architecture of a sub-network. There are m channels in the sub-network, and each channel includes a convolutional encoding module and a convolutional decoding module. The number of convolutional encoding modules and convolutional decoding modules is equal, and the connection structures of the convolutional encoding modules and convolutional decoding modules in each channel are the same. To balance the denoising effect and overall complexity, preferably, the number of convolutional encoding modules and convolutional decoding modules is 2-4, and the number of channels is 3-4.
[0217] For the convenience of explanation, a sub-network with 3 channels is taken as an example for elaboration. See Figure 8 As shown, it is a schematic diagram of the sub-network structure with 3 channels and each channel having 3 convolutional encoding modules and convolutional decoding modules. The input image 1 (original input image) is downsampled twice to obtain two downsampled input images. In this way, there are 3 input images including input image 1, the downsampled input image 2, and input image 3, which are respectively input into their respective channels. After being processed by each channel, the image data of each channel is fused to obtain a denoised image; each channel has 3 convolutional encoding modules and 3 convolutional encoding modules. The convolutional encoding module performs convolutional encoding on its input according to the convolutional encoding step length, which is equivalent to filtering according to the filtering step length. The convolutional decoding module performs convolutional decoding on its input according to the convolutional decoding step length, which is equivalent to restoring the data. Through the processing of convolutional encoding - convolutional decoding, the denoising effect is achieved.
[0218] Three convolutional encoding modules are connected in sequence, and three convolutional decoding modules are connected in sequence. The output of the last convolutional encoding module is connected to the input of the first convolutional decoding module. The input of the second convolutional encoding module (the output of the first convolutional encoding module) is connected to the output of the second convolutional decoding module (the input of the third convolutional encoding module), the input of the third convolutional encoding module (the output of the second convolutional encoding module), and the input of the second convolutional decoding module (the output of the first convolutional encoding module). That is to say, every other convolutional encoding layer is connected to the symmetric convolutional decoding layer with a jumper wire, so forward and backward propagation can be directly performed. For example:
[0219] The pixel size of the input image 1 is 256×256. Through two downsamplings, two input-size images of 128×128 (input image 2) and 64×64 (input image 3) are obtained. Given that a too large stride will cause information loss, preferably, the convolutional encoding module performs convolutional encoding according to a 1 / 2 times encoding stride, and the convolutional decoding module performs convolutional decoding according to a 2 times convolutional decoding stride. In this way,
[0220] The first convolutional encoding module, the second convolutional encoding module, and the third convolutional encoding module in the first channel convolutional-encode the input image 1 into a feature encoding 1. Among them, the first convolutional encoding module is a 128×128×32 convolutional network, and the pixel size of the obtained convolutional encoding result is 128×128×32. The second convolutional encoding module is a 64×64×64 convolutional network, and the pixel size of the obtained convolutional encoding result is 64×64×64. The third convolutional encoding module is a 32×32×128 convolutional network, and the pixel size of the obtained convolutional encoding result (feature encoding 1) is 32×32×128. The feature encoding 1 is input to the first convolutional decoding module, and the first feature convolutional encoding is sequentially convolutional-decoded through the first convolutional decoding module, the second convolutional decoding module, and the third convolutional decoding module to obtain the output image 1. Among them, the first convolutional decoding module is a 64×64×64 convolutional network, and the pixel size of the obtained convolutional decoding result is 64×64×64. The second convolutional decoding module is a 128×128×32 convolutional network, and the pixel size of the obtained convolutional decoding result is 128×128×32. The third convolutional decoding module is a 256×256×1 convolutional network, and the pixel size of the obtained convolutional decoding result (the output image 1 of the input image 1) is 256×256×1. In this way, the pixel size of the output image 1 is the same as that of the input image 1.
[0221] Similarly,
[0222] The first convolutional encoding module, the second convolutional encoding module, and the third convolutional encoding module in the second channel perform convolutional encoding on the input image 2 to obtain the second feature convolutional encoding. Among them, the first convolutional encoding module is a convolutional network of 64×64×32, and the pixel size of the obtained convolutional encoding result is 64×64×32. The second convolutional encoding module is a convolutional network of 32×32×64, and the pixel size of the obtained convolutional encoding result is 32×32×64. The third convolutional encoding module is a convolutional network of 16×16×128, and the pixel size of the obtained convolutional encoding result (feature encoding 2) is 16×16×128. The feature encoding 2 is input to the first convolutional decoding module, and the feature encoding 2 is successively subjected to convolutional decoding through the first convolutional decoding module, the second convolutional decoding module, and the third convolutional decoding module to obtain the output image 2. Among them, the first convolutional decoding module is a convolutional network of 32×32×64, and the pixel size of the obtained convolutional decoding result is 32×32×64. The second convolutional decoding module is a convolutional network of 64×64×32, and the pixel size of the obtained convolutional decoding result is 64×64×32. The third convolutional decoding module is a convolutional network of 128×128×1, and the pixel size of the obtained convolutional decoding result is 128×128×1. In this way, the pixel size of the output image 2 is the same as that of the input image 2.
[0223] The first convolutional encoding module, the second convolutional encoding module, and the third convolutional encoding module in the third channel perform convolutional encoding on the third input image to obtain the feature encoding 3. Among them, the first convolutional encoding module is a convolutional network of 32×32×32, and the pixel size of the obtained convolutional encoding result is 32×32×32. The second convolutional encoding module is a convolutional network of 16×16×64, and the pixel size of the obtained convolutional encoding result is 16×16×64. The third convolutional encoding module is a convolutional network of 8×8×128, and the pixel size of the obtained convolutional encoding result (feature encoding 3) is 8×8×128. The feature encoding 3 is input to the first convolutional decoding module, and the feature encoding 3 is successively subjected to convolutional decoding through the first convolutional decoding module, the second convolutional decoding module, and the third convolutional decoding module to obtain the output image 3. Among them, the first convolutional decoding module is a convolutional network of 16×16×64, and the pixel size of the obtained convolutional decoding result is 16×16×64. The second convolutional decoding module is a convolutional network of 32×32×32, and the pixel size of the obtained convolutional decoding result is 32×32×32. The third convolutional decoding module is a convolutional network of 64×64×1, and the pixel size of the obtained convolutional decoding result is 64×64×1. In this way, the pixel size of the output image 3 is the same as that of the input image 3.
[0224] The weights of each channel of the sub-network are shared to reduce the training difficulty of the network. The output image 2 and the output image 3 are upsampled to obtain images with the same pixel size as the output image 1, and finally the output images of the three channels are fused to obtain the denoised image.
[0225] See Figure 9 as shown Figure 9 is a schematic diagram of a convolutional encoding module. The convolutional encoding module is composed of 1 convolutional layer and 1 enhanced residual sub-module (EN-Res Block). The input data of the convolutional encoding module is processed by the convolutional layer and the enhanced residual sub-module in sequence and then output.
[0226] See Figure 10 as shown Figure 10 is a schematic diagram of a convolutional decoding module. The convolutional decoding module is composed of 1 enhanced residual sub-module (EN-Res Block) and 1 deconvolutional layer. The input data of the convolutional decoding module is processed by the enhanced residual sub-module and the deconvolutional layer in sequence and then output.
[0227] The enhanced residual sub-module is a module designed for the face denoising task and is used to enhance the local residuals in the image. See Figure 11 as shown Figure 11 is a schematic diagram of the enhanced residual sub-module. According to the ablation experiment data, the number of convolutional layers is 4-8. In this embodiment, 6 convolutional layers are taken as an example. The 6 convolutional layers are connected in sequence, and the input of the first convolutional layer is connected to the output of the second convolutional layer (the input of the third convolutional layer), the input of the third convolutional layer is connected to the output of the fourth convolutional layer (the input of the fifth convolutional layer), and the input of the fifth convolutional layer is connected to the output of the sixth convolutional layer, that is, every 2 convolutional layers are connected by a jumper to perform the addition of the feature matrices. In addition, in order to enhance the feature transitivity of the entire enhanced residual sub-module, the input of the first convolutional layer is connected to the output of the fourth convolutional layer (the input of the fifth convolutional layer), the output of the second convolutional layer (the input of the third convolutional layer) is connected to the output of the sixth convolutional layer, and the input of the first convolutional layer is connected to the output of the sixth convolutional layer, that is, among the convolutional layers connected by a jumper every 2 convolutional layers: a jumper is connected every 4 convolutional layers, a jumper is connected every 6 convolutional layers, and so on, to perform the addition of the feature matrices.
[0228] Based on the deep learning model, the target image pairs and color card image pairs in the training sample data are input into the deep learning model to train the deep learning model. For each batch of training sample data input, the multi-dimensional mixed loss function of this batch of training samples is calculated, and different batches of training sample data are repeatedly input until the multi-dimensional mixed loss function reaches the set loss threshold.
[0229] The multi-dimensional hybrid loss function is: the sum of the loss function of the face sub-network and the loss function of the face region of interest sub-network, which is expressed mathematically as:
[0230] loss = loss face + loss face-ROI (1)
[0231] Wherein,
[0232] loss face is the loss function of the face sub-network;
[0233] loss face-ROI is the loss function of the face region of interest sub-network.
[0234] loss face is: the first cumulative result of the difference between the value 1 in each channel of all channels of the face sub-network and the structural similarity index of that channel in the face sub-network, and the sum of the minimum absolute deviation losses of each channel of the face sub-network, which is expressed mathematically as:
[0235]
[0236] Preferably, loss face is: the first weighted result of the difference between the first weighted coefficient weighted value 1 of channel i in all channels of the face sub-network and the structural similarity index of that channel, and the second weighted result of the second weighted coefficient weighted minimum absolute deviation loss of channel i, which is expressed mathematically as:
[0237]
[0238] loss face-ROI is: the second cumulative result of the difference between the value 1 in each channel of all channels of the face region of interest sub-network and the structural similarity index of that channel, and the sum of the minimum absolute deviation losses of each channel of the face region of interest sub-network, which is expressed mathematically as:
[0239]
[0240] Preferably, loss face-ROI is the sum of the third weighted result of the difference between the third weighted coefficient weighted value 1 of channel i in all channels of the face region of interest sub-network and the structural similarity index of the face region of interest sub-network, and the fourth weighted result of the fourth weighted coefficient weighted minimum absolute deviation loss of channel i, which is expressed mathematically as:
[0241]
[0242] Wherein,
[0243] i is the number of the i-th channel;
[0244] m is the total number of sub-network channels;
[0245] SSIM is the structural similarity index, and its value range is between 0 and 1. According to the denoised images of each channel and the noiseless images, the similarity mean value of the training sample data of the current batch is calculated. L1 is the minimum absolute deviation loss. According to the denoised images of each channel and the noiseless images, the mean value of the minimum absolute deviation loss of the training sample data of the current batch is calculated. Among them, for the downsampled image, the denoised image is the upsampled image. a, b, c, and d are the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, and the fourth weighting coefficient respectively.
[0246] See Figure 12 as shown in Figure 12 It is a schematic diagram of a device for deep learning model training. The training sample data is input into the first sub-network. The ROI region image is extracted through the ROI extraction module based on the currently input training sample data and input into the second sub-network. The first sub-network loss function calculation module calculates the first sub-network loss function according to the denoised image output by the first sub-network and the input noiseless image. The second sub-network loss function calculation module calculates the second sub-network loss function according to the denoised ROI image output by the second sub-network and the input noiseless ROI image. The loss function calculation module accumulates the first sub-network loss function from the first sub-network loss function calculation module and the second sub-network loss function from the second sub-network loss function calculation module to obtain a multi-dimensional mixed loss function. When the multi-dimensional mixed loss function reaches the set loss threshold, the training stops. Otherwise, the training sample data of the next batch is read.
[0247] When the trained model needs to be used in a monitoring device, the model needs to be transplanted into the device for simulation and corresponding integration tests.
[0248] In this embodiment, the training sample data including target images under different illuminations is used to train the deep learning model, which is beneficial to improving the noise reduction ability of the deep learning model for various differential noises; the deep learning model adopts a dual sub-network and is trained according to the multi-dimensional mixed loss function of each channel of each sub-network during the training process, which is beneficial to improving the training efficiency and facilitating the parameter adjustment of the deep learning model parameters; the trained deep learning model performs noise reduction on the whole and ROI regions of the target image respectively. In this way, not only the overall noise reduction effect of the target image is improved, but also the key region images in the target image are subjected to noise reduction processing, making the image noise reduction processing more delicate, and the fused denoised image is beneficial to further improving the image quality.
[0249] SeeFigure 13 As shown Figure 13 It is a schematic diagram of a deep learning model module for image denoising. The module includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to implement the image denoising process.
[0250] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0251] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0252] The embodiment of the present invention also provides a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the image denoising process as described above.
[0253] For the device / network-side device / storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.
[0254] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0255] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An image noise reduction method, characterized in that, The method includes downsampling the image to be denoised to obtain m downsampled images performing j - time convolutional encodings on the image to be denoised and the m downsampled images respectively in sequence to obtain m + 1 feature encodings performing j - time convolutional decodings on the m + 1 feature encodings respectively in sequence to obtain m + 1 images after convolutional decoding upsampling the image after convolutional decoding of each downsampled image to obtain m upsampled images fusing the image after convolutional decoding of the image to be denoised and the m upsampled images to obtain a denoised image with the same pixel size as the image to be denoised; wherein when j is a natural number greater than 1 performing a first convolutional decoding on the feature encoding to obtain the result of this convolutional decoding performing the next convolutional decoding adjacent to this convolutional decoding on the convolutional encoding result with the same pixel size as the result of this convolutional decoding and the result of this convolutional decoding, and recursively performing in sequence until the j - th convolutional decoding is performed; when j is equal to 1, performing one - time convolutional decoding on the feature encoding to obtain the result of convolutional decoding; m is a natural number greater than or equal to 1.
2. The method according to claim 1, characterized in that, The method further includes extracting the region of interest in the image based on the image to be denoised to obtain an image of the region of interest using the image of the region of interest as the image to be denoised, and performing the step of downsampling the image to be denoised to obtain m downsampled images to obtain a denoised image of the image of the region of interest; fusing the denoised image of the image to be denoised and the denoised image of the image of the region of interest in the image to be denoised; The step of performing j - time convolutional encodings on the image to be denoised and the m downsampled images respectively in sequence includes performing convolutional encodings on the image to be denoised and the m downsampled images in parallel respectively, wherein each image performs j - time convolutional encodings in sequence according to the set convolutional encoding step length, and the feature encoding is the feature encoding after j - time convolutional encodings; The step of performing j - time convolutional decodings on the m + 1 feature encodings respectively in sequence includes performing convolutional decodings on the m + 1 feature encodings in parallel respectively, wherein each feature encoding performs j - time convolutional decodings in sequence according to the same convolutional decoding step length.
3. The method according to claim 2, wherein The step of performing the next convolutional decoding adjacent to this convolutional decoding on the convolutional encoding result with the same pixel size as the result of this convolutional decoding and the result of this convolutional decoding, and recursively performing in sequence until the j - th convolutional decoding is performed, includes performing a second convolutional decoding on the first convolutional decoding result obtained from the first convolutional decoding and the (j - 1)-th convolutional encoding result obtained from the (j - 1)-th convolutional encoding to obtain a second convolutional decoding result performing a third convolutional decoding on the second convolutional decoding result and the (j - 2)-th convolutional encoding result obtained from the (j - 2)-th convolutional encoding to obtain a third convolutional decoding result; recursively performing in sequence until performing the j - th convolutional decoding on the (j - 1)-th convolutional decoding result and the first convolutional encoding result obtained from the first convolutional encoding to obtain the j - th convolutional decoding result.
4. The method according to claim 2, wherein The step of performing j - time convolutional encodings in sequence includes: in each convolutional encoding, performing convolution and enhanced residual processing in sequence; The sequential j - time convolutional decoding includes: during each convolutional decoding, enhanced residual and deconvolution processing are sequentially performed; Among them, the enhanced residual processing includes q convolutional processings, where q is a natural number greater than or equal to 1.
5. The method according to claim 4, characterized in that, The q is an even number, and the q convolutional processings include sequentially performing q convolutional operations on the first input data for enhanced residual processing, where the first input data is sequentially subjected to the first convolution and the second convolution to obtain a second convolution result, the second convolution result and the first input data are sequentially subjected to the third convolution and the fourth convolution to obtain a fourth convolution result, the fourth convolution result, the second convolution result, and the first input data are sequentially subjected to the fifth convolution and the sixth convolution to obtain a sixth convolution result, sequentially recursing, the result of the p - th convolution, the result of the p - 2k - th convolution, and the first input data are sequentially subjected to the (p + 1)-th convolution and the (p + 2)-th convolution to obtain the (p + 2) - th convolution result, … until p + 2 equals q, obtaining the q - th convolution result; p is an even number, p - 2k is greater than 1, and k is a natural number greater than or equal to 1.
6. The method according to claim 5, wherein The m is 3 or 4, the j is a natural number from 2 to 6, and the q is a natural number from 4 to 8; The size of the convolutional coding step is such that: the pixel size of the convolutional coding result obtained by this convolutional coding is half of the pixel size of the convolutional coding result obtained by the previous convolutional coding, The size of the (j - 1)-th convolutional decoding step is such that: the pixel size of the convolutional decoding result obtained by this convolutional decoding is twice the pixel size of the convolutional decoding result obtained by the previous convolutional decoding; The size of the j - th convolutional decoding step is such that: the pixel size of the convolutional decoding result obtained by this convolutional decoding is the same as the pixel size of the input image.
7. A noise reduction module for image noise reduction, characterized in that, This noise reduction module includes: a downsampling module for downsampling the image to be denoised to obtain m downsampled images, an encoding module for sequentially performing j - time convolutional encodings on the image to be denoised and the m downsampled images respectively to obtain m + 1 feature encodings; a decoding module for sequentially performing j - time convolutional decodings on the m + 1 feature encodings respectively to obtain m + 1 images after convolutional decoding, an upsampling module for upsampling the image after convolutional decoding of each downsampled image to obtain m upsampled images, a channel image fusion module for fusing the image after convolutional decoding of the image to be denoised and the m upsampled images to obtain a denoised image with the same pixel size as the image to be denoised; where in the case where j is a natural number greater than 1, the decoding module performs the first convolutional decoding on the feature encoding to obtain the result of this convolutional decoding, the decoding module performs the next convolutional decoding adjacent to this convolutional decoding on the convolutional coding result having the same pixel size as the result of this convolutional decoding and the result of this convolutional decoding, and sequentially recurses until the j - th convolutional decoding; in the case where j equals 1, the decoding module performs one - time convolutional decoding on the feature encoding to obtain the convolutional decoding result; m is a natural number greater than or equal to 1.
8. The noise reduction module according to claim 7, wherein, The encoding module and the decoding module are trained deep learning model modules. The deep learning model module includes a first sub-network for denoising the target image and a second sub-network for denoising the region of interest in the target image. Wherein, The first sub-network processes the image to be denoised into a denoised image of the image to be denoised; The second sub-network processes the image of the region of interest extracted based on the image to be denoised into a denoised image of the image of the region of interest. The first sub-network and the second sub-network have the same network structure. Each sub-network includes m channels, and each channel includes j convolutional encoding modules and j convolutional decoding modules. In each channel: the j convolutional encoding modules are connected in sequence, and the j convolutional decoding modules are connected in sequence. Among them, the jth convolutional encoding result of the jth convolutional encoding module is output to the first convolutional decoding module, and any one of the remaining convolutional decoding modules except the first convolutional decoding module is connected to any one of the remaining convolutional encoding modules except the jth convolutional encoding module. The convolutional encoding result of this convolutional encoding module has the same pixel size as the other input of this convolutional decoding module.
9. The noise reduction module according to claim 8, wherein, The denoising module further includes, A region of interest extraction module for extracting the image of the region of interest in the image to be denoised. A sub-network image fusion module for fusing the denoised image from the first sub-network and the denoised image from the second sub-network. Any one of the remaining convolutional decoding modules except the first convolutional decoding module is connected to any one of the remaining convolutional encoding modules except the jth convolutional encoding module, including, The (j - 1)th convolutional encoding result of the (j - 1)th convolutional encoding module is input to the second convolutional decoding module. The (j - 2)th convolutional encoding result of the (j - 2)th convolutional encoding module is input to the third convolutional decoding module. By successive recursion, Until the first convolutional encoding result of the first convolutional encoding module is input to the jth convolutional decoding module.
10. The noise reduction module according to claim 8, wherein, The convolutional encoding module is used to perform convolution and enhanced residual processing in sequence during each convolutional encoding. The convolutional decoding module is used to perform enhanced residual and deconvolution processing in sequence during each convolutional decoding. Among them, the enhanced residual processing includes q convolutional processes, and q is a natural number greater than or equal to 1.
11. The noise reduction module according to claim 10, wherein, When q is an even number, the q convolutional processes include, Successively performing q convolutional processes on the first input data for enhanced residual processing. Wherein, The first input data is successively subjected to the first convolution and the second convolution to obtain a second convolution result. The second convolution result and the first input data are successively subjected to the third convolution and the fourth convolution to obtain a fourth convolution result. The fourth convolution result, the second convolution result, and the first input data are successively subjected to the fifth convolution and the sixth convolution to obtain a sixth convolution result. By successive recursion, The result of the pth convolution, the result of the (p - 2k)th convolution, and the first input data are successively subjected to the (p + 1)th convolution and the (p + 2)th convolution to obtain a (p + 2) convolution result. … Until p + 2 is equal to q, the q convolution result is obtained. p is an even number, p - 2k is greater than 1, and k is a natural number greater than or equal to 1.
12. The noise reduction module according to claim 11, characterized in that, m is 3 or 4, j is a natural number from 2 to 6, and q is a natural number from 4 to 8; The size of the convolutional coding step is such that: the pixel size of the convolutional coding result obtained from this convolutional coding is half of the pixel size of the convolutional coding result obtained from the previous convolutional coding. The size of the (j - 1)-th convolutional decoding step is such that: the pixel size of the convolutional decoding result obtained from this convolutional decoding is twice the pixel size of the convolutional decoding result obtained from the previous convolutional decoding. The size of the j-th convolutional decoding step is such that: the pixel size of the convolutional decoding result obtained from this convolutional decoding is the same as the pixel size of the input image.
Citation Information
Patent Citations
Image processing method and device based on multiple frames of images
CN110072051A
Image denoising method based on adversarial generative network
CN110223254A