Image denoising method and device and storage medium
Through the combination of wavelet decomposition and multi-scale convolution kernel attention module, the fine features of the image are extracted and denoised, which solves the problem of poor denoising effect of rich details and achieves efficient image denoising and detail retention.
Patent Information
- Application Number
- CN202510180199.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively remove noise in images with rich details, especially in near-infrared image processing, resulting in poor denoising effect.
Wavelet decomposition technology is used to obtain multi-level target wavelet subbands, and these wavelet subbands are processed using a first network containing attention modules of multi-scale convolution kernels, refined image features are extracted, and image denoising is finally performed based on these features.
The fine denoising effect of rich details is achieved, so that the denoised image can retain more image details, and improve the accuracy and effect of denoising.
Smart Images

Figure CN119991493A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an image denoising method, device and storage medium. Background Art
[0002] The goal of image denoising is to remove noise from an image and retain image details as much as possible. In recent years, many deep learning-based image denoising algorithms have been proposed to better solve the image denoising problem.
[0003] However, some images are rich in details, and compared with ordinary images, the difficulty and complexity of image denoising are higher. For example, compared with far-infrared images, near-infrared images have richer image details, which correspondingly increases the difficulty and complexity of near-infrared image denoising. Therefore, how to achieve good denoising effects for images rich in details is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The main technical problem solved by the present application is to provide an image denoising method, device and storage medium, which can achieve a refined denoising effect so that the denoised clean image can retain more image details.
[0005] To solve the above technical problems, a technical solution adopted in the present application is: to provide an image denoising method, the method comprising: obtaining target wavelet subbands of multiple levels obtained after wavelet decomposition of the target object, the target object comprising an image to be denoised acquired by an image acquisition unit; using a first network to process target wavelet subbands of different levels to obtain a target feature map; the first network comprises an attention module composed of multiple first convolution kernels of different scales; denoising the image to be denoised based on the target feature map to obtain a target clean image corresponding to the image to be denoised.
[0006] To solve the above technical problems, another technical solution adopted in the present application is: to provide an electronic device, comprising a memory and a processor coupled to each other, the memory storing program instructions; the processor is used to execute the program instructions stored in the memory to implement the above method.
[0007] In order to solve the above technical problem, another technical solution adopted by the present application is: providing a computer-readable storage medium for storing program instructions, which can be executed to implement the above method.
[0008] In the above scheme, after obtaining the multi-level target wavelet subbands obtained after wavelet decomposition of the target object, the first network of the attention module including the multi-scale first convolution kernel is used to process the target wavelet subbands of different levels to obtain the target feature map. Since convolution kernels of different scales can perceive the features of different scales of the image, compared with the method of using the same scale convolution kernel for image denoising, the method of using the first network including the above attention module in the present application can extract refined image features and obtain a target feature map including refined features, and then perform denoising on the denoised image based on the target feature map, which can achieve a refined denoising effect and enable the denoised clean image to retain more image details. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is a flowchart of an embodiment of an image denoising method provided by the present application;
[0010] Figure 2 It is a schematic diagram of the basic principle of two-dimensional discrete wavelet transform provided by this application;
[0011] Figure 3 yes Figure 1 The flowchart of step S12 is shown as an embodiment;
[0012] Figure 4 The target feature is obtained by extracting the first network provided by this application Figure 1 A schematic diagram of a framework of an embodiment;
[0013] Figure 5 is a processing flow chart of an embodiment of an attention module provided by the present application;
[0014] Figure 6 yes Figure 3 The flowchart of step S32 is shown as an embodiment;
[0015] Figure 7 It is a schematic diagram of the framework of an embodiment of training the third network provided by the present application;
[0016] Figure 8 It is a schematic diagram of the framework of an embodiment of an image denoising device provided by the present application;
[0017] Fig. 9 It is a schematic diagram of a framework of an embodiment of an electronic device provided by the present application;
[0018] Fig.10 It is a schematic diagram of the framework of the computer-readable storage medium provided by this application. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail below with reference to the accompanying drawings and examples.
[0020] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0021] See also Figure 1 , Figure 1 is a flow chart of an embodiment of image denoising provided by the present application. It should be noted that if there are substantially the same results, this embodiment does not use Figure 1 The process sequence shown is limited. Figure 1 As shown, this embodiment includes:
[0022] S11: Acquire target wavelet subbands of multiple levels obtained after wavelet decomposition of the target object, where the target object includes the image to be denoised acquired by the image acquisition unit.
[0023] First, the wavelet decomposition is briefly described as follows:
[0024] Wavelet analysis (wavelet decomposition) is a powerful signal processing tool. Its basic principle is to use the oscillating waveform of a "mother wavelet" of finite length or fast decay to match the input signal by scaling and translating. This analysis method shows excellent feature characterization capabilities in both time domain and frequency domain, and can meet the needs of information analysis and high-frequency area information analysis. Specifically, when wavelet analysis performs multi-layer (level) wavelet decomposition on a signal, it first decomposes the signal to be decomposed into two parts, low frequency and high frequency, and then decomposes the low frequency part again when decomposing it again, and so on, and finally realizes multi-layer (level) wavelet decomposition of the signal. Among them, when performing multi-layer wavelet decomposition on a signal, only the low frequency part is decomposed each time, and the high frequency part is no longer decomposed, and the image can also be regarded as a signal.
[0025] The basic principle of two-dimensional discrete wavelet transform is as follows Figure 2As shown. x[m,n] represents a discrete input signal, i.e., an image signal, with a width of m and a height of n. g[n] is a low-pass filter, h[m] is a high-pass filter, and ↓2 represents a downsampling filter. If x[n] is used as input, the output y[n] = x[2n]. The image x will first pass through a low-pass filter and a high-pass filter and then be downsampled to obtain v 1,L and v 1,H , and then these two components will pass through the high-pass filter and low-pass filter again, and then downsample again. Finally, four wavelet coefficients x are obtained 1,A ,x 1,V ,x 1,H ,x 1,D .
[0026] In one implementation scenario, the discrete two-dimensional wavelet transform of the Haar wavelet can be used for image processing. The Haar wavelet is a classic wavelet basis. The discrete two-dimensional wavelet transform of the Haar wavelet only requires addition and subtraction operations but not multiplication operations. It is easy to implement and fast, and is more commonly used in image analysis.
[0027] The four wavelet coefficients obtained after the first-level wavelet transform of the image can be regarded as a low-frequency subband and three high-frequency subbands obtained by decomposing the original image. The three high-frequency subbands contain high-frequency information in the horizontal, vertical and diagonal directions of the image, respectively, while the low-frequency subband contains the low-frequency information of the original image. These subbands can be further transformed by wavelet to obtain subbands of subbands, which is the multi-scale characteristic of wavelet transform. For example, a second-level wavelet transform can be performed on the basis of the low-frequency subband in the first-level wavelet subband to obtain a second-level wavelet subband, and a third-level wavelet transform can be performed on the basis of the low-frequency subband in the second-level wavelet subband to obtain a third-level wavelet subband.
[0028] The decomposition of the image by wavelet transform can be expressed as follows:
[0029] x i,A , x i,V , x i,H , x i,D =DWT i (Y)
[0030] Among them, DWT represents the wavelet transform process, x i,A , x i,V , x i,H , x i,D Respectively represent low-frequency component, high-frequency vertical component, high-frequency horizontal component and high-frequency diagonal component, and i represents the i-th level wavelet decomposition. It is worth noting that wavelet transform and its inverse operation are reversible and will not cause information loss. Therefore, the complete original image can be obtained by inverse transforming the wavelet subband image.
[0031] In this embodiment, the target object may include only the image to be denoised and may also include a stripe noise prior and / or at least one adjacent frame image of the image to be denoised.
[0032] In the video denoising scenario, at least one adjacent frame image is an adjacent frame image acquired before the image to be denoised, and / or an adjacent frame image acquired after the image to be denoised. For example, each of time t1, t2 and t3 corresponds to an acquired image, wherein the acquired image at time t2 is the image to be denoised, and the images acquired at time t1 and t3 are adjacent frame images of the image to be denoised.
[0033] Stripe noise usually refers to fixed pattern noise that appears in an image or video. The stripe noise prior is an image or data that represents the inherent stripe noise pattern of the camera. The stripe noise prior is used in subsequent image denoising to identify or remove stripe noise in the image. These noises may be caused by sensor defects, circuit problems, or environmental factors of the camera. These noises appear in the image as a series of fixed, usually periodic brightness changes, similar to stripes.
[0034] Optionally, a method of using an image acquisition unit to shoot images or videos in a completely dark environment and averaging each frame of the image can be used to obtain the stripe noise prior. For example, cover the camera lens to completely block external light to ensure that any content in the captured image or video is only generated by internal factors of the camera (such as sensor noise); then turn off the bias correction switch, shoot multiple frames of video, and add each frame to get the average value to obtain the stripe noise prior. Among them, turning off the bias correction switch can prevent the camera from automatically performing certain types of noise correction, so that purer stripe noise can be captured. It should be noted that the above-mentioned method of shooting multiple frames of video and averaging is to reduce the influence of random noise in order to retain a fixed stripe noise pattern; of course, other statistical methods that can reduce the influence of random noise can also be used to obtain the stripe noise prior.
[0035] In one implementation scenario, taking the example that the level of the target wavelet subband includes three levels (i.e., the first-level target wavelet subband, the second-level target wavelet subband and the third-level target wavelet subband), if the target object only includes the image to be denoised, the first-level wavelet subband of the image to be denoised is the first-level target wavelet subband, the second-level wavelet subband of the image to be denoised is the second-level target wavelet subband, and the third-level wavelet subband of the image to be denoised is the third-level target wavelet subband.
[0036] If the target object includes not only the image to be denoised but also the stripe noise prior and / or at least one adjacent frame image of the image to be denoised, the stripe noise prior and / or at least one adjacent frame image of the image to be denoised can be combined with the image to be denoised to obtain at least one image group, wherein each image group includes not only the image to be denoised but also the stripe noise prior and / or at least one adjacent frame image of the image to be denoised.
[0037] Wherein, each image group includes a stripe noise prior and / or at least one adjacent frame image, and an image to be denoised; each image group corresponds to a group of target wavelet subbands of different levels, and for each image group, the target wavelet subbands of each level corresponding to the image group are obtained by splicing the wavelet subbands of the corresponding levels of each image in the image group. For example, if an image group includes an image to be denoised and a stripe noise prior, the wavelet subband obtained by splicing the first-level wavelet subband of the image to be denoised and the first-level wavelet subband of the stripe noise prior is the first-level target wavelet subband, and the second-level target wavelet subband and the third-level target wavelet subband are obtained by splicing the same first-level target wavelet subband.
[0038] It should be noted that the number of target feature maps obtained by processing target wavelet subbands of different levels using the first network in step S12 is the same as the number of image groups. That is, one image group corresponds to one target feature map, and for each image group, the target feature map corresponding to the image group is obtained by processing target wavelet subbands of various levels corresponding to the image group using the first network. For example, the image group includes image group 1, image group 2, and image group 3. Then, step S12 uses the first network to process target wavelet subbands of different levels to obtain target feature maps, including: using the first network to process target wavelet subbands of various levels of image group 1 to obtain target feature maps. Figure 1 ; Use the first network to process the target wavelet subbands at all levels of image group 2 to obtain the target features Figure 2 ; Use the first network to process the target wavelet subbands at all levels of image group 3 to obtain the target features Figure 3 .
[0039] In addition, it should be noted that, considering that the stripe noise in the image is generally relatively uniform and does not have many image details, there is no need to extract multiple levels of wavelet subbands. Therefore, when the target object includes a stripe noise prior, only the stripe noise prior can be subjected to a first-level wavelet decomposition to obtain a first-level wavelet subband of the stripe noise prior, without the need for a second-level and a third-level wavelet decomposition. That is, only the first-level target wavelet subband is obtained by concatenating the first-level wavelet subband of the image to be denoised and the first-level wavelet subband of the stripe noise prior, the second-level target wavelet subband is the second-level wavelet subband of the image to be denoised, and the third-level target wavelet subband is the third-level wavelet subband of the image to be denoised.
[0040] The specific number of image groups can be selected according to actual needs. If the denoising efficiency of the image is to be improved, the target object can be set to include only one image group. However, if the image is to be avoided from being blurred, the target object can also be set to include multiple image groups.
[0041] In one implementation scenario, multiple image groups may be set. For example, the target object includes the image t to be denoised, the two adjacent frames of the image to be denoised (image t-2, image t-1, image t+1 and image t+2) before and after the image to be denoised, and the stripe noise prior. First, the images of the three adjacent time points of the image t to be denoised and the stripe noise prior are combined into an image group according to the time sequence of the images, and the three image groups obtained are: image group 1 (stripe noise prior, image t to be denoised, image t-2 and image t-1); image group 2 (stripe noise prior, image t to be denoised, image t-1 and image t+1); image group 3 (stripe noise prior, image t to be denoised, image t+1 and image t+2). For each image group in the multiple image groups, the following step S12 is performed, and after using the first network to process the target wavelet subbands of different levels corresponding to each image group to obtain the corresponding target feature map, the target feature maps corresponding to each image group are spliced to obtain a spliced feature map, and the spliced feature map is used as a new target feature map to perform step S13.
[0042] Among them, if the target object does not contain the stripe noise prior in the above example, at least one adjacent frame image of the image to be denoised can be combined with the image to be denoised to obtain at least one image group, and the specific number of image groups can be determined according to actual needs; of course, if the target object does not contain at least one adjacent frame image in the above example, the stripe noise prior and the image to be denoised can only be combined into one image group. This is because the stripe noise prior is generally a fixed pattern of noise, and the stripe noise in the image captured by the image acquisition unit is generally relatively uniform, so there is no need to combine it with the image to be denoised into multiple image groups.
[0043] It should be noted that the setting of the target object to include the adjacent frame images of the image to be denoised is based on the consideration that in the denoising scenario of the image frames in the video, maintaining the temporal consistency between frames and eliminating flicker are key factors affecting the perceived quality of the result. In order to achieve these goals, when denoising a given frame (the image to be denoised) of the image sequence, the temporal information existing in the adjacent frames can be used. When denoising a given pixel or image block, there are two benefits of finding similar pixels or image blocks in the adjacent frames of the noisy image. First, the adjacent frames provide more additional information for denoising; second, the use of temporally adjacent frames helps to reduce jitter because the residual errors between adjacent image frames are correlated. The setting of the target object to include the stripe noise prior is to better distinguish the actual information and stripe noise in the image in the subsequent image denoising, so as to better remove the stripe noise in the image.
[0044] The image to be denoised in this embodiment may be an image captured by any type of image acquisition unit or a frame of image in a video, and may be, but not limited to, a near-infrared image, a far-infrared image, a radar image, or the like.
[0045] S12: Using the first network to process target wavelet subbands of different levels to obtain a target feature map; the first network includes an attention module composed of multiple first convolution kernels of different scales.
[0046] It should be noted that after the image is decomposed by multi-level wavelet, target wavelet subbands of different scales can be obtained, that is, the scales of target wavelet subbands of different levels are different, and the target wavelet subbands of each scale include the features of different frequencies of the original image. By extracting these features in a targeted manner, the first network can capture the multi-frequency features of the image in a finer granularity, so as to better distinguish noise from image edge details. However, the ratio of noise and details in different frequency features is different, the features are different, and the importance of denoising is different. In order to extract useful information more targetedly and reduce the interference of useless information, the present application sets an attention module containing multiple first convolution kernels of different scales.
[0047] Among them, multiple first convolution kernels of different scales can be two square convolution kernels with equal sides, or two strip convolution kernels with unequal sides, or some of the first convolution kernels can be square convolution kernels and some of the first convolution kernels can be strip convolution kernels.
[0048] Compared with the square convolution kernel, the strip convolution kernel can not only reduce the amount of calculation, but also better capture the strip features, which is conducive to distinguishing the column stripe noise and the detail features in the column direction in the image. Therefore, in one implementation scenario, some or all of the convolution kernels in the multiple convolution kernels can be set as strip convolution kernels.
[0049] See also Figure 3, Figure 3 yes Figure 1 The flowchart of an embodiment of step S12 is shown. In this embodiment, step S12 further includes:
[0050] S31: Using the attention module to perform attention processing based on target wavelet subbands at different levels, to obtain first attention features of target wavelet subbands at different levels.
[0051] In this embodiment, the first attention feature includes at least one second attention feature, the number of the second attention features is n-1, n represents the number of levels of the target wavelet subband, and different second attention features are obtained by using the attention module to perform attention processing on two target wavelet subbands of different adjacent levels. That is, the input of the attention module is two target wavelet subbands of two adjacent levels.
[0052] In one embodiment, before using the attention module to perform attention processing based on target wavelet subbands at different levels, it also includes using at least one feature extraction module to perform preliminary feature extraction on target wavelet subbands at different levels to obtain initial feature maps of target wavelet subbands at different levels. Then, the attention module is used to perform attention processing based on target wavelet subbands at different levels to obtain first attention features of target wavelet subbands at different levels.
[0053] Specifically, the attention module is used to perform attention processing based on target wavelet subbands of different levels to obtain the first attention features of target wavelet subbands of different levels, including: according to the first order of multiple levels from small to large, two adjacent target wavelet subbands are used as two first current subbands in turn; the first feature map of the front target wavelet subband in the two first current subbands and the second feature map of the rear target wavelet subband are obtained; the first feature map and the second feature map of the two first current subbands are fused to obtain the first fused feature map; the attention module is used to perform attention processing on the first fused feature map to obtain the second attention features corresponding to the two first current subbands; and the second attention feature is used as the first feature map of the front target wavelet subband in the next two first current subbands; the above steps are repeated until the second attention features corresponding to the last two first current subbands are obtained; wherein the first attention feature includes each second attention feature. wherein the rear target wavelet subband is used as the front target wavelet subband in the next two adjacent target wavelet subbands; the first feature map and the second feature map of the first two target wavelet subbands in the first order are both the initial feature maps of the corresponding target wavelet subbands.
[0054] For example, see Figure 4 , Figure 4 The target feature is obtained by extracting the first network provided by this application Figure 1 A schematic diagram of the framework of an embodiment. Figure 4The MSAM in represents an attention module. The noise image represents the image to be denoised, or the image to be denoised and at least one adjacent frame image, the stripe prior represents the stripe noise prior mentioned above, Figure 4 The input module and basic module in are both feature extraction modules. Figure 4 The example uses the first network to process three different levels of target wavelet subbands.
[0055] like Figure 4 As shown, before using MSAM (attention module) to perform attention processing based on target wavelet subbands of different levels, at least one feature extraction module is first used to perform preliminary feature extraction on target wavelet subbands of various levels to obtain initial feature maps of target wavelet subbands of various levels. For example, the input module and the basic module are used in sequence to perform feature extraction on the first-level target wavelet subband to obtain an extracted feature map, and then the extracted feature map is downsampled to obtain an initial feature map of the first-level target wavelet subband. For another example, the input module is used to perform feature extraction on the second-level target wavelet subband to obtain an initial feature map of the second-level target wavelet subband, and the input module is used to perform feature extraction on the third-level target wavelet subband to obtain an initial feature map of the third-level target wavelet subband.
[0056] Then, according to the order of the three levels of the target wavelet subbands from small to large, the first-level target wavelet subband and the second-level target wavelet subband are first used as the two first current subbands, and the second-level and third-level target wavelet subbands are the next two first current subbands of the first-level and second-level target wavelet subbands.
[0057] For the primary and secondary target wavelet subbands, obtain the first feature map of the primary target wavelet subband (which is the initial feature map of the primary target wavelet subband) and the second feature map of the secondary target wavelet subband (which is the initial feature map of the secondary target wavelet subband), and fuse the first feature map and the second feature map of the primary and secondary target wavelet subbands to obtain the first fused feature map, then use MSAM (attention module) to perform attention processing on the first fused feature maps of the primary and secondary target wavelet subbands to obtain the attention features corresponding to the primary and secondary target wavelet subbands, and use the attention features corresponding to the primary and secondary target wavelet subbands as the new first feature map of the secondary target wavelet subband, and repeat the above steps until the second attention features corresponding to the secondary target wavelet subband and the tertiary target wavelet subband are obtained. Specifically, the secondary and tertiary target wavelet subbands are taken as the two new first current subbands. For the secondary and tertiary target wavelet subbands, a new first feature map of the secondary target wavelet subband and a second feature map of the tertiary target wavelet subband (which is the initial feature map of the tertiary target wavelet subband) are obtained. The first feature maps and the second feature maps of the secondary and tertiary target wavelet subbands are fused to obtain the first fused feature maps of the secondary and tertiary target wavelet subbands. The attention module is used to perform attention processing on the first fused feature maps of the secondary and tertiary target wavelet subbands to obtain the second attention features of the secondary and tertiary target wavelet subbands.
[0058] In one embodiment, the attention module includes a channel attention submodule and a spatial attention submodule composed of a plurality of first convolution kernels of different scales. The channel attention submodule is used to estimate the weights of each channel of target wavelet subbands of different levels, and the spatial attention submodule is used to adjust the weights of different spatial regions of the target wavelet subband to maximize the extraction of useful feature information.
[0059] Specifically, an attention module is used to perform attention processing based on target wavelet subbands of different levels to obtain first attention features of target wavelet subbands of different levels, including: using a channel attention submodule to perform channel attention processing based on target wavelet subbands of different levels to obtain corresponding channel attention weights, and using a spatial attention submodule to perform spatial attention processing based on target wavelet subbands of different levels to obtain corresponding spatial attention weights; then the channel attention weights and the spatial attention weights are combined to obtain a comprehensive attention weight; and then based on the comprehensive attention weights and target wavelet subbands of different levels, the first attention features of target wavelet subbands of different levels are obtained.
[0060] In one implementation scenario, the spatial attention submodule also includes a second convolution kernel and a third convolution kernel. The spatial attention submodule is used to perform spatial attention processing based on target wavelet subbands of different levels to obtain corresponding spatial attention weights, including: using the second convolution kernel to aggregate local information of target wavelet subbands of different levels to obtain aggregated information; using each first convolution kernel to perform convolution processing of the aggregated information at a corresponding scale to obtain a corresponding convolution result; using the third convolution kernel to perform attention processing on the fusion result of different convolution results to obtain spatial attention weights. For details, please refer to the following about Figure 5 Description of the part.
[0061] It should be noted that the input of the attention module includes at least feature maps corresponding to target wavelet subbands of two different levels. The feature maps of target wavelet subbands of each level are the first feature maps or the second feature maps of the target wavelet subbands of the corresponding level mentioned above. The specific details can be combined with the relevant descriptions of the first feature map and the second feature map above, and no further details will be given here.
[0062] See also Figure 5 , Figure 5 It is a processing flow chart of an embodiment of the attention module provided in this application.
[0063] like Figure 5 In the attention module shown in (a), input 1 and input 2 of the attention module are respectively the first feature map and the second feature map in the two first current sub-bands, and the output of the attention module is the second attention features corresponding to the two first current sub-bands extracted, for example, the second attention features corresponding to the primary and secondary target wavelet sub-bands, and for example, the second attention features corresponding to the secondary and tertiary target wavelet sub-bands. Figure 5 As shown, the spatial attention submodule outputs the spatial attention weight, and the channel attention submodule outputs the channel attention weight. The channel attention weight and the spatial attention weight are first integrated to obtain the comprehensive attention weight, and then the sigmoid activation function is solved for the comprehensive attention weight to obtain the normalized comprehensive attention weight. After that, the comprehensive attention weight is multiplied by the first feature map and the first fused feature map of the second feature map to obtain the corresponding product result, and then the product result is fused with the first fused feature map to obtain the second attention feature corresponding to the two adjacent level target wavelet subbands.
[0064] like Figure 5The channel attention submodule shown in (b) has the first fused features of two adjacent target wavelet subbands as input. First, the channel attention module compresses the first fused features of the input through average pooling and maximum pooling to obtain a one-dimensional vector, the length of which is the number of channels 2C, thereby encoding the global features of each channel, and then obtains the attention weights between channels through two fully connected layers. Among them, the output of the channel attention submodule is the channel attention weights of two adjacent target wavelet subbands. Optionally, in order to reduce overhead, the first fully connected layer compresses the number of vector channels to C / r, where r is the compression coefficient, and the second fully connected layer expands the number of compressed one-dimensional vector channels to C to obtain the channel attention weights corresponding to the two adjacent target wavelet subbands.
[0065] like Figure 5 The spatial attention submodule shown in (c) above has the first fusion features of two adjacent target wavelet subbands as input. Figure 5 As shown, the second convolution kernel is a 1×1 convolution kernel, which is used to aggregate the local information in the first fusion features of two adjacent target wavelet subbands to obtain aggregated information. The first convolution kernels of multiple different scales include 1×3 and 3×1 first convolution kernels, 1×5 and 5×1 first convolution kernels, 1×7 and 7×1 first convolution kernels, and each first convolution kernel is used to perform convolution processing of the corresponding scale on the aggregated information to obtain the corresponding convolution result, for example, the aggregated information is processed by 1×3 and 3×1 convolution kernels to obtain the corresponding convolution result. The third convolution kernel is a 1×1 convolution kernel, which is used to perform attention processing on the fusion result of the convolution results of the first convolution kernels of different scales to obtain the spatial attention weights of the two adjacent target wavelet subbands.
[0066] S32: Based on the first attention features of the target wavelet subbands at different levels, a target feature map is obtained.
[0067] Combining the above process, after obtaining the spatial attention weights of each two adjacent target wavelet subbands, the target feature map can be obtained in the following way. Specifically:
[0068] See also Figure 6 , Figure 6 yes Figure 3 The flowchart of step S32 of an embodiment is shown. In this embodiment, step S32 further includes:
[0069] S61: according to the second order of the multiple levels from large to small, two adjacent target wavelet sub-bands are sequentially used as two second current sub-bands.
[0070] Taking the target wavelet subband as level 3 as an example, the second order is the third-level target wavelet subband, the second-level target wavelet subband and the first-level target wavelet subband. After obtaining the second attention features corresponding to the first-level and second-level target wavelet subbands, and the second attention features corresponding to the second-level and third-level target wavelet subbands, first use the second-level and third-level target wavelet subbands as two second current subbands, and execute steps S62-S64, and then use the first-level and second-level target wavelet subbands as two second current subbands, and execute steps S62-S65.
[0071] S62: Obtain a third characteristic map of the subsequent target wavelet subband and a fourth characteristic map of the preceding target wavelet subband of the two second current subbands.
[0072] In this embodiment, if the two second current subbands are located in the first two target wavelet subbands in the second order, the second attention features corresponding to the two second current subbands are used as the fourth feature map of the front target wavelet subband, and the second attention features corresponding to the rear target wavelet subband and the next adjacent target wavelet subband of the rear target wavelet subband are used as the third feature map of the rear target wavelet subband.
[0073] If the two second current subbands include a target wavelet subband located at the last position in the second order, the fourth feature map of the non-last target wavelet subband in the two second current subbands is the second fused feature map of the first two second current subbands of the two second current subbands, and the third feature map of the last target wavelet subband is the initial feature map of the last target wavelet subband.
[0074] Continuing with the example of the target wavelet subband level being level 3, please continue to refer to Figure 4 ,like Figure 4 As shown in the figure, if the two second current subbands are the secondary and tertiary target wavelet subbands (the first two target wavelet subbands in the second order), the second attention features corresponding to the secondary and tertiary target wavelet subbands are used as the fourth feature map of the tertiary target wavelet subband, and the second attention features corresponding to the primary and secondary target wavelet subbands are used as the third feature map of the secondary target wavelet subband. If the two second current subbands are the primary and secondary target wavelet subbands (i.e., including the target wavelet subband at the end of the second order), the fourth feature map of the non-last secondary target wavelet subband is the second fused feature map of the secondary and tertiary target wavelet subbands, and the third feature map of the last primary target wavelet subband is the initial feature map of the primary target wavelet subband.
[0075] like Figure 4 As shown, the first-level residual obtained after the feature extraction of the input module and the convolution processing of the basic module on the first-level target wavelet subband can be used as the initial feature map of the first-level target wavelet subband.
[0076] S63: Fusing the third feature map and the fourth feature map of the two second current sub-bands to obtain a second fused feature map; and using the second fused feature map as the fourth feature map of the previous target wavelet sub-band in the next two second current sub-bands.
[0077] Please continue reading Figure 4 , the fourth feature map of the third-level target wavelet subband is the second attention feature corresponding to the second-level and third-level target wavelet subbands, that is, Figure 4 The medium attention module (MSAM) is based on the attention features (denoted as A) output by the secondary and tertiary wavelet subbands of the noise image. The third feature map of the secondary target wavelet subband is the second attention feature corresponding to the primary and secondary target wavelet subbands, i.e. Figure 4 The medium attention module (MSAM) is based on the attention features (referred to as B) output by the primary and secondary wavelet subbands of the noise image. It takes the attention feature A processed by the convolution of the basic module and the upsampling module as the updated attention feature A, and takes the attention feature B processed by the convolution of the basic module (i.e. Figure 4 The second-level residual in is used as the updated attention feature B, and the updated attention feature A and the updated attention feature B are fused to obtain the second fused feature map of the secondary and tertiary target wavelet subbands.
[0078] The second fusion feature map of the secondary and tertiary target wavelet subbands is used as the fourth feature map of the secondary target wavelet subband, and the third feature map of the primary target wavelet subband is used as the initial feature map of the primary target wavelet subband. Figure 4 As shown, the initial feature map of the primary target wavelet subband is the first-level residual obtained after the feature extraction of the input module and the convolution processing of the basic module are performed on the primary target wavelet subband. The fourth feature map of the secondary target wavelet subband is processed by the basic module and the upsampling module in sequence to obtain the feature map as the updated fourth feature map, and the fourth feature map is fused with the first-level residual to obtain the second fused feature map of the primary and secondary target wavelet subbands.
[0079] S64: Repeat the above steps until the second fusion features corresponding to the last two second current sub-bands are obtained.
[0080] S65: Obtain a target feature map based on the second fusion features corresponding to the last two second current sub-bands.
[0081] In one embodiment, the target wavelet subband at the last position and the second fused features corresponding to the last two second current subbands may be fused to obtain a target feature map.
[0082] For example, please refer to Figure 4 ,like Figure 4As shown, the second fusion feature corresponding to the last primary target wavelet subband and the last primary and secondary target wavelet subbands is fused to obtain the corresponding target feature map.
[0083] In one implementation scenario, if the target object includes not only the image to be denoised, but also stripe noise prior and / or adjacent frame images of the image to be denoised, step S65 also includes: fusing the first-level wavelet subband of the image to be denoised and the second fusion features corresponding to the first-level and second-level target wavelet subbands to obtain a corresponding target feature map.
[0084] S13: De-noising the image to be denoised based on the target feature map to obtain a target clean image corresponding to the image to be denoised.
[0085] It should be noted that the number of target feature maps obtained by processing target wavelet subbands of different levels using the first network in step S12 is the same as the number of image groups. That is, one image group corresponds to one target feature map, and for each image group, the target feature map corresponding to the image group is obtained by processing target wavelet subbands of various levels corresponding to the image group using the first network.
[0086] If there are multiple image groups, multiple target feature maps can be spliced to obtain a spliced feature map, and then the spliced feature map is used as a new target feature map to perform denoising on the image to be denoised based on the spliced target feature map to obtain a target clean image corresponding to the image to be denoised.
[0087] In this embodiment, the image to be denoised is denoised based on the target feature map to obtain at least one initial clean image, and then the initial clean images are fused to obtain a target clean image.
[0088] In one embodiment, the step of performing denoising on the image to be denoised based on the target feature map to obtain at least one initial clean image is performed by a second network, wherein the input of the second network is the target feature map or the concatenated target feature map.
[0089] In one implementation scenario, the second network includes at least one denoising module group, each denoising module group includes several denoising sub-modules, and different denoising module groups share at least two denoising sub-modules; the several denoising sub-modules in each denoising module group include a first denoising sub-module for outputting an initial clean image, and at least two second denoising sub-modules for providing feature representations of different scales for the first denoising sub-module.
[0090] Each denoising submodule is composed of at least one of a downsampling layer, an upsampling layer, and a convolutional layer. Figure 7 About the functions of various arrows.
[0091] See also Figure 7 , the second network includes Figure 7 At least one of the three denoising module groups shown.
[0092] The second network Figure 7 Taking the network shown in FIG. 1 as an example, the second network includes three denoising module groups. Denoising module group 1 includes three denoising submodules, x0,0, x1,0 and x0,1; denoising module group 2 includes six denoising submodules, x0,0, x1,0, x0,1, x2,0, x1,1 and x0,2; and denoising module group 3 includes ten denoising submodules, x0,0, x1,0, x0,1, x2,0, x1,1, x0,2, x3,0, x2,1, x1,2 and x0,3. Denoising module group 1 shares the denoising submodules in denoising module group 1 with denoising module group 2 and denoising module group 3 respectively; denoising module group 2 shares the denoising submodules in denoising module group 2 with denoising module group 3.
[0093] The first denoising submodules for outputting an initial clean image in denoising module group 1, denoising module group 2 and denoising module group 3 are x0,1, x0,2 and x0,3 respectively. For denoising module group 1, at least two second denoising submodules for providing feature representations of different scales for the first denoising submodule are x0,0 and x1,0; for denoising module group 2, at least two second denoising submodules for providing feature representations of different scales for the first denoising submodule include two of x0,0, x0,1 and x1,1; for denoising module group 3, at least two second denoising submodules for providing feature representations of different scales for the first denoising submodule include two of x0,0, x0,1, x0,2 and x1,2.
[0094] The number of initial clean images is the same as the number of denoising module groups included in the second network. If the second network only includes one denoising module group, the number of initial clean images is only one, so the one initial clean image can be directly used as the target clean image corresponding to the image to be denoised. However, if the second network includes at least two denoising module groups, the number of initial clean images is two, and the two initial clean images can be fused to obtain the final target clean image.
[0095] In the above scheme, after obtaining the multi-level target wavelet subbands obtained after wavelet decomposition of the target object, the first network of the attention module including the multi-scale first convolution kernel is used to process the target wavelet subbands of different levels to obtain the target feature map. Since convolution kernels of different scales can perceive the features of different scales of the image, compared with the method of using the same scale convolution kernel for image denoising, the method of using the first network including the above attention module in the present application can extract refined image features and obtain a target feature map including refined features, and then perform denoising on the denoised image based on the target feature map, which can achieve a refined denoising effect and enable the denoised clean image to retain more image details.
[0096] In one embodiment, before using the first network and the second network to perform the above processing, the first network and the second network need to be trained. Specifically, the training steps include:
[0097] First, obtain sample objects corresponding to a number of sample clean images respectively; wherein the sample objects include sample noise images corresponding to the sample clean images.
[0098] Before each network is trained, a number of sample clean images and sample objects corresponding to each sample clean object are first obtained. The sample object may include only a sample noise image, or may include a stripe noise prior and / or at least one adjacent frame image of the sample noise image. The method for obtaining the stripe noise prior may refer to the above description, and will not be elaborated here.
[0099] The sample clean image can be obtained in the following manner. For example, for a certain shooting scene, the bias correction function of the camera is first turned on, and M images are continuously taken, and these images only contain random noise. In order to obtain a clean image without random noise, these M frames of images are added and averaged during post-processing to obtain a sample clean image.
[0100] Then, a noise image including stripe noise, Poisson noise and Gaussian noise is generated by simulation. Then, the noise image is fused with the clean image at a certain ratio to obtain a sample noise image corresponding to the clean image. Wherein, fusing the noise image with the sample clean image at a certain ratio means that the pixel value of the pixel point in the noise image is multiplied by a first coefficient to obtain a first product; the pixel value of the pixel point in the sample clean image is multiplied by a second coefficient to obtain a second product, and then the first product and the second product are summed to obtain the sample noise image. The method for obtaining at least one adjacent frame image of the sample noise image can refer to the sample noise image.
[0101] Second, obtain multiple levels of sample wavelet subbands obtained after wavelet decomposition of each sample object.
[0102] For details, please refer to the relevant description of step S11 above, and no further details will be given here.
[0103] Third, the first network is used to process multiple levels of sample wavelet subbands to obtain a sample feature map.
[0104] For details, please refer to the relevant description of step S12 above, and no further details will be given here.
[0105] Fourth, the sample feature map is processed using the third network to obtain a sample denoised image corresponding to the sample noisy image.
[0106] It should be noted that the network used to process the sample feature map during the training process to obtain the sample denoised image corresponding to the sample noise image is the third network. Figure 7 , Figure 7 In the third network, three denoising module groups are included, each of which is used to output a sample denoised image (i.e. Figure 7 Output image 1, output image 2 and output image 3 are shown). Each output image is a sample denoised image.
[0107] Fifth, based on the difference between the sample denoised image and the corresponding sample clean image, the network parameters of the first network and the third network are adjusted.
[0108] Continue reading Figure 7 , respectively determine the output image 1 and the corresponding sample clean image ( Figure 7 a first difference between the output image 2 and the clean image, a second difference between the output image 2 and the clean image, a third difference between the output image 3 and the clean image, combining the first difference, the second difference and the third difference to obtain a comprehensive difference, and adjusting the network parameters of the first network and the third network based on the comprehensive difference until the first network and the third network are trained to converge.
[0109] The comprehensive difference may be, but is not limited to, the mean of the first difference, the second difference, and the third difference, and may also be a weighted mean, a median, or the like.
[0110] In one implementation scenario, the difference between each output image of the third network and the clean image is the difference between the primary wavelet decomposition subband of each output image and the primary wavelet subband of the clean image, that is, the difference between the sample denoised image and the corresponding sample clean image is substantially the difference between the primary wavelet subbands corresponding to the sample denoised image and the corresponding sample clean image, so that the determined difference includes the difference between the low-frequency wavelet subbands and the difference between the high-frequency wavelet subbands.
[0111] It should be noted that, in one embodiment, in order to reduce the computational complexity of the second network, the plurality of denoising module groups in the trained third network may be pruned, and the pruned third network is used as the final second network. Figure 7 The third network includes three denoising module groups, wherein at least one of the denoising module group 2 for obtaining the output image 2 and the denoising module group 3 for obtaining the output image 3 can be pruned, and the pruned third network is used as the second network.
[0112] The specific denoising module group to be pruned can be determined by comprehensively considering the performance and efficiency of the network after pruning.
[0113] It should be noted that if the clean image output by the second network is the corresponding first-level wavelet subband, the first-level wavelet subband can be inversely transformed to obtain the final target clean image.
[0114] See also Figure 8 , Figure 8 It is a schematic diagram of the framework of an embodiment of an image denoising device provided by the present application. In this embodiment, the image denoising device 80 includes: an acquisition module 81, a first processing module 82, and a second processing module 83. The acquisition module 81 is used to obtain multiple levels of target wavelet subbands obtained after the target object is decomposed by wavelet, and the target object includes the image to be denoised collected by the image acquisition unit; the first processing module 82 is used to use the first network to process the target wavelet subbands of different levels to obtain a target feature map; the first network includes an attention module composed of multiple first convolution kernels of different scales; the second processing module 83 is used to perform denoising on the image to be denoised based on the target feature map to obtain a target clean image corresponding to the image to be denoised.
[0115] In some embodiments, the plurality of first convolution kernels of different scales include at least one strip convolution kernel, and two sides of the strip convolution kernel have different lengths; and / or, the image to be denoised is a near-infrared image.
[0116] In some embodiments, the first processing module 82 uses the first network to process target wavelet subbands of different levels to obtain a target feature map, including: using the attention module to perform attention processing based on target wavelet subbands of different levels to obtain first attention features of target wavelet subbands of different levels; based on the first attention features of target wavelet subbands of different levels, obtaining a target feature map.
[0117] In some embodiments, the attention module includes a channel attention submodule and a spatial attention submodule composed of multiple first convolution kernels of different scales; the attention module is used to perform attention processing based on target wavelet subbands of different levels to obtain first attention features of target wavelet subbands of different levels, including: using the channel attention submodule to perform channel attention processing based on target wavelet subbands of different levels to obtain corresponding channel attention weights, and using the spatial attention submodule to perform spatial attention processing based on target wavelet subbands of different levels to obtain corresponding spatial attention weights; combining the channel attention weights and the spatial attention weights to obtain the comprehensive attention weights; based on the comprehensive attention weights and target wavelet subbands of different levels, the first attention features of target wavelet subbands of different levels are obtained.
[0118] In some embodiments, the spatial attention submodule also includes a second convolution kernel and a third convolution kernel; the spatial attention submodule is used to perform spatial attention processing based on target wavelet subbands of different levels to obtain corresponding spatial attention weights, including: using the second convolution kernel to aggregate local information of target wavelet subbands of different levels to obtain aggregated information; using each first convolution kernel to perform convolution processing of the aggregated information at a corresponding scale to obtain a corresponding convolution result; using the third convolution kernel to perform attention processing on the fusion result of different convolution results to obtain the spatial attention weight.
[0119] In some embodiments, the first attention feature includes at least one second attention feature, the number of second attention features is n-1, n represents the number of levels of the target wavelet subband, and different second attention features are obtained by using the attention module to pay attention to two target wavelet subbands of different adjacent levels.
[0120] In some embodiments, before using the attention module to perform attention processing based on target wavelet subbands at different levels to obtain the first attention features of target wavelet subbands at different levels, it also includes: using at least one feature extraction module to perform preliminary feature extraction on the target wavelet subbands at each level to obtain the initial feature maps of the target wavelet subbands at each level.
[0121] In some embodiments, an attention module is used to perform attention processing based on target wavelet subbands of different levels to obtain first attention features of target wavelet subbands of different levels, including: in a first order of multiple levels from small to large, two adjacent target wavelet subbands are used as two first current subbands in turn; a first feature map of the front target wavelet subband in the two first current subbands and a second feature map of the rear target wavelet subband are obtained; wherein the rear target wavelet subband is used as the front target wavelet subband in the next two adjacent target wavelet subbands; the first feature map and the second feature map of the first two target wavelet subbands in the first order are both initial feature maps of the corresponding target wavelet subbands; the first feature map and the second feature map of the two first current subbands are fused to obtain a first fused feature map; the attention module is used to perform attention processing on the first fused feature map to obtain second attention features corresponding to the two first current subbands; and the second attention feature is used as the first feature map of the front target wavelet subband in the next two first current subbands; the aforementioned steps are repeated until the second attention features corresponding to the last two first current subbands are obtained; wherein the first attention feature includes each second attention feature.
[0122] In some embodiments, a target feature map is obtained based on the first attention features of target wavelet subbands at different levels, including: according to a second order of multiple levels from large to small, two adjacent target wavelet subbands are taken as two second current subbands in turn; a third feature map of the target wavelet subband at the rear of the two second current subbands and a fourth feature map of the target wavelet subband at the front are obtained; the third feature map and the fourth feature map of the two second current subbands are fused to obtain a second fused feature map; and the second fused feature map is taken as the fourth feature map of the target wavelet subband at the front of the next two second current subbands; the aforementioned steps are repeated until the second fused features corresponding to the last two second current subbands are obtained; and the target feature map is obtained based on the second fused features corresponding to the last two second current subbands.
[0123] In some embodiments, obtaining a third feature map of a subsequent target wavelet subband in two second current subbands and a fourth feature map of a preceding target wavelet subband includes: in response to the two second current subbands being located in the first two target wavelet subbands in the second order, taking the second attention features corresponding to the two second current subbands as the fourth feature map of the preceding target wavelet subband, and taking the second attention features corresponding to the subsequent target wavelet subband and the next adjacent target wavelet subband of the subsequent target wavelet subband as the third feature map of the subsequent target wavelet subband; in response to the two second current subbands including a target wavelet subband located at the last position in the second order, the fourth feature map of a non-last target wavelet subband in the two second current subbands is the second fused feature map of the first two second current subbands of the two second current subbands, and the third feature map of the last target wavelet subband is the initial feature map of the last target wavelet subband.
[0124] In some embodiments, based on the second fusion features corresponding to the last two second current sub-bands, a target feature map is obtained, including: fusing the last target wavelet sub-band and the second fusion features corresponding to the last two second current sub-bands to obtain the target feature map.
[0125] In some embodiments, the second processing module 83 performs denoising on the image to be denoised based on the target feature map to obtain a target clean image corresponding to the image to be denoised, including: performing denoising on the image to be denoised based on the target feature map to obtain at least one initial clean image; and fusing the initial clean images to obtain a target clean image.
[0126] In some embodiments, the second processing module 83 performs denoising on the image to be denoised based on the target feature map, and the step of obtaining at least one initial clean image is executed by the second network; and / or, the second network includes at least one denoising module group, each denoising module group includes a number of denoising sub-modules, and different denoising module groups share at least two denoising sub-modules; the several denoising sub-modules in each denoising module group include a first denoising sub-module for outputting an initial clean image, and at least two second denoising sub-modules for providing feature representations of different scales for the first denoising sub-module.
[0127] In some embodiments, each denoising submodule is composed of at least one of a downsampling layer, an upsampling layer and a convolutional layer; and / or, the method also includes: obtaining sample objects corresponding to several sample clean images respectively; wherein the sample objects include sample noise images corresponding to the sample clean images; obtaining multiple levels of sample wavelet subbands obtained after wavelet decomposition of each sample object; using the first network to process the multiple levels of sample wavelet subbands to obtain a sample feature map; using the third network to process the sample feature map to obtain a sample denoised image corresponding to the sample noise image; based on the difference between the sample denoised image and the corresponding sample clean image, adjusting the network parameters of the first network and the third network.
[0128] In some embodiments, the third network includes several denoising module groups; the method further includes: pruning the several denoising module groups in the third network, and using the pruned third network as the second network.
[0129] In some embodiments, the target object also includes a stripe noise prior and / or at least one adjacent frame image of the image to be denoised; the stripe noise prior and / or at least one adjacent frame image are used to be combined with the image to be denoised to obtain at least one image group, each image group includes a stripe noise prior and / or at least one adjacent frame image, and the image to be denoised; each image group corresponds to a group of target wavelet subbands of different levels, and for each image group, the target wavelet subbands of each level corresponding to the image group are obtained by splicing the wavelet subbands of each image in the image group at the corresponding level; the number of target feature maps is the same as the number of image groups, and each target feature map is obtained by processing the target wavelet subbands of each level corresponding to each image group using the first network.
[0130] See also Fig. 9 , Fig. 9 1 is a schematic diagram of a framework of an electronic device according to an embodiment of the present application. In this embodiment, the electronic device 90 includes a memory 91 and a processor 92 coupled to each other.
[0131] The memory 91 stores program instructions, and the processor 92 is used to execute the program instructions stored in the memory 91 to implement the steps of any of the above-mentioned method implementation methods. In a specific implementation scenario, the electronic device 90 may include, but is not limited to: a microcomputer, a server, and in addition, the electronic device 90 may also include a mobile device such as a laptop computer and a tablet computer, which is not limited here.
[0132] Specifically, the processor 92 is used to control itself and the memory 91 to implement the steps of any of the above-mentioned embodiments. The processor 92 can also be called a CPU (Central Processing Unit). The processor 92 may be an integrated circuit chip with signal processing capabilities. The processor 92 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 92 can be implemented by an integrated circuit chip.
[0133] See also Fig.10 , Fig.10It is a schematic diagram of the framework of the computer-readable storage medium provided by the present application. The computer-readable storage medium 100 of the embodiment of the present application stores a program instruction 101, and the program instruction 101 is executed to implement the method provided by any embodiment of the above method and any non-conflicting combination. Among them, the program instruction 101 can form a program file and be stored in the above-mentioned computer-readable storage medium 100 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) executes all or part of the steps of each implementation method of the present application. The aforementioned computer-readable storage medium 100 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, or a terminal device such as a computer, a server, a mobile phone, and a tablet.
[0134] In the above scheme, after obtaining the multi-level target wavelet subbands obtained after wavelet decomposition of the target object, the first network of the attention module including the multi-scale first convolution kernel is used to process the target wavelet subbands of different levels to obtain the target feature map. Since convolution kernels of different scales can perceive the features of different scales of the image, compared with the method of using the same scale convolution kernel for image denoising, the method of using the first network including the above attention module in the present application can extract refined image features and obtain a target feature map including refined features, and then perform denoising on the denoised image based on the target feature map, which can achieve a refined denoising effect and enable the denoised clean image to retain more image details.
[0135] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0136] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0137] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0138] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0139] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0140] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0141] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An image denoising method, characterized in that: The method comprises: Acquire target wavelet subbands of multiple levels obtained after wavelet decomposition of the target object, wherein the target object includes the image to be denoised acquired by the image acquisition unit; Using a first network to process target wavelet subbands of different levels to obtain a target feature map; the first network includes an attention module composed of a plurality of first convolution kernels of different scales; The image to be denoised is denoised based on the target feature map to obtain a target clean image corresponding to the image to be denoised.
2. The method according to claim 1, characterized in that The multiple first convolution kernels of different scales include at least one strip convolution kernel, and two sides of the strip convolution kernel have different lengths; And / or, the image to be denoised is a near-infrared image.
3. The method according to claim 1, characterized in that The method of using the first network to process target wavelet subbands of different levels to obtain a target feature map includes: Using the attention module to perform attention processing based on target wavelet subbands at different levels, the first attention features of target wavelet subbands at different levels are obtained; Based on the first attention features of the target wavelet subbands at different levels, the target feature map is obtained.
4. The method according to claim 3, characterized in that The attention module includes a channel attention submodule and a spatial attention submodule composed of a plurality of first convolution kernels of different scales; The attention module is used to perform attention processing based on target wavelet subbands at different levels to obtain first attention features of target wavelet subbands at different levels, including: Using the channel attention submodule to perform channel attention processing based on target wavelet subbands of different levels to obtain corresponding channel attention weights, and using the spatial attention submodule to perform spatial attention processing based on target wavelet subbands of different levels to obtain corresponding spatial attention weights; The channel attention weight and the spatial attention weight are combined to obtain a comprehensive attention weight; Based on the attention comprehensive weight and the target wavelet subbands at different levels, the first attention features of the target wavelet subbands at different levels are obtained.
5. The method according to claim 4, characterized in that The spatial attention submodule also includes a second convolution kernel and a third convolution kernel; The spatial attention submodule is used to perform spatial attention processing based on target wavelet subbands of different levels to obtain corresponding spatial attention weights, including: Using the second convolution kernel to aggregate local information of target wavelet subbands of different levels to obtain aggregated information; Using each of the first convolution kernels to perform convolution processing of a corresponding scale on the aggregation information to obtain a corresponding convolution result; The third convolution kernel is used to perform attention processing on the fusion result of different convolution results to obtain the spatial attention weight.
6. The method according to claim 3, characterized in that The first attention feature includes at least one second attention feature, the number of the second attention features is n-1, and n represents the number of levels of the target wavelet subband. The second attention feature is obtained by using the attention module to pay attention to two target wavelet subbands of different adjacent levels.
7. The method according to claim 3, characterized in that Before using the attention module to perform attention processing based on target wavelet subbands of different levels to obtain first attention features of target wavelet subbands of different levels, the method further includes: At least one feature extraction module is used to perform preliminary feature extraction on target wavelet subbands at each level to obtain initial feature maps of target wavelet subbands at each level.
8. The method according to claim 7, characterized in that The attention module is used to perform attention processing based on target wavelet subbands at different levels to obtain first attention features of target wavelet subbands at different levels, including: According to the first order of the multiple levels from small to large, two adjacent target wavelet subbands are sequentially used as two first current subbands; Acquire a first feature map of a front target wavelet subband and a second feature map of a rear target wavelet subband in the two first current subbands; wherein the rear target wavelet subband is used as the front target wavelet subband in the next two adjacent target wavelet subbands; the first feature map and the second feature map of the first two target wavelet subbands in the first sequence are both initial feature maps of the corresponding target wavelet subbands; Fusing the first feature map and the second feature map of the two first current sub-bands to obtain a first fused feature map; Using the attention module to perform attention processing on the first fusion feature map to obtain a second attention feature corresponding to the two first current sub-bands; and using the second attention feature as a first feature map of the previous target wavelet sub-band in the next two first current sub-bands; Repeat the above steps until the second attention features corresponding to the last two first current sub-bands are obtained; wherein the first attention features include each of the second attention features.
9. The method according to claim 8, characterized in that The first attention features based on the target wavelet subbands at different levels, obtaining the target feature map, includes: According to the second order of the multiple levels from large to small, two adjacent target wavelet subbands are sequentially used as two second current subbands; Acquire a third characteristic map of the latter target wavelet subband and a fourth characteristic map of the former target wavelet subband of the two second current subbands; Fusing the third feature map and the fourth feature map of the two second current sub-bands to obtain a second fused feature map; and using the second fused feature map as the fourth feature map of the previous target wavelet sub-band in the next of the two second current sub-bands; Repeat the above steps until the second fusion features corresponding to the last two second current sub-bands are obtained; Based on the second fusion features corresponding to the last two second current sub-bands, the target feature map is obtained.
10. The method according to claim 9, characterized in that The step of obtaining the third characteristic map of the target wavelet subband that is behind the two second current subbands and the fourth characteristic map of the target wavelet subband that is ahead of the two second current subbands comprises: In response to the two second current subbands being located in the first two target wavelet subbands in the second order, using the second attention features corresponding to the two second current subbands as the fourth feature map of the preceding target wavelet subband, and using the second attention features corresponding to the succeeding target wavelet subband and the next adjacent target wavelet subband of the succeeding target wavelet subband as the third feature map of the succeeding target wavelet subband; In response to the two second current subbands including a target wavelet subband located at the last position in the second order, the fourth feature map of the non-last target wavelet subband among the two second current subbands is the second fused feature map of the first two second current subbands of the two second current subbands, and the third feature map of the last target wavelet subband is the initial feature map of the last target wavelet subband.
11. The method according to claim 9, characterized in that The obtaining of the target feature map based on the second fusion features corresponding to the last two second current sub-bands includes: The target wavelet subband at the last position and the second fused features corresponding to the last two second current subbands are fused to obtain the target feature map.
12. The method according to claim 1, characterized in that The denoising process is performed on the image to be denoised based on the target feature map to obtain a target clean image corresponding to the image to be denoised, including: Performing denoising on the image to be denoised based on the target feature map to obtain at least one initial clean image; The initial clean images are fused to obtain the target clean image.
13. The method according to claim 12, characterized in that The step of performing denoising processing on the image to be denoised based on the target feature map to obtain at least one initial clean image is performed by the second network; And / or, the second network includes at least one denoising module group, each of the denoising module groups includes a plurality of denoising sub-modules, and different denoising module groups share at least two denoising sub-modules; The plurality of denoising submodules in each of the denoising module groups include a first denoising submodule for outputting an initial clean image, and at least two second denoising submodules for providing feature representations of different scales for the first denoising submodule.
14. The method according to claim 13, characterized in that Each of the denoising submodules is composed of at least one of a downsampling layer, an upsampling layer and a convolutional layer; And / or, the method further comprises: Obtaining sample objects corresponding to a number of sample clean images respectively; wherein the sample objects include sample noise images corresponding to the sample clean images; Obtaining multiple levels of sample wavelet subbands obtained after wavelet decomposition of each sample object; Processing the sample wavelet subbands at multiple levels using the first network to obtain a sample feature map; Processing the sample feature map using a third network to obtain a sample denoised image corresponding to the sample noise image; Based on the difference between the sample denoised image and the corresponding sample clean image, the network parameters of the first network and the third network are adjusted.
15. The method according to claim 14, characterized in that The third network includes several denoising module groups; The method further comprises: Pruning is performed on several denoising module groups in the third network, and the pruned third network is used as the second network.
16. The method according to claim 1, characterized in that The target object also includes a stripe noise prior and / or at least one adjacent frame image of the image to be denoised; the stripe noise prior and / or the at least one adjacent frame image are used to be combined with the image to be denoised to obtain at least one image group, each of the image groups includes the stripe noise prior and / or the at least one adjacent frame image, and the image to be denoised; Each of the image groups corresponds to a group of target wavelet subbands of different levels. For each of the image groups, the target wavelet subbands of each level corresponding to the image group are obtained by splicing the wavelet subbands of the corresponding levels of each image in the image group. The number of the target feature maps is the same as the number of the image groups, and each of the target feature maps is obtained by processing target wavelet subbands of various levels corresponding to each of the image groups using the first network.
17. An electronic device, characterized in that: comprising a memory and a processor coupled to each other, The memory stores program instructions; The processor is used to execute the program instructions stored in the memory to implement the method according to any one of claims 1 to 16.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions that can be run by a processor, and the program instructions can be executed by the processor to implement the method according to any one of claims 1 to 16.
Citation Information
Cited By
Double-domain collaborative image denoising method, device and equipment and readable storage medium
CN121685315A