Image noise reduction method and device, electronic device, and storage medium
By decomposing the image in the time and space domains and using pyramid matching and motion compensation technology for fine alignment and fusion, the problems of large computational complexity and poor accuracy in existing technologies are solved, and efficient image denoising effects are achieved.
Patent Information
- Application Number
- CN202211174866.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-09-26
AI Technical Summary
The existing image denoising methods are computationally intensive and have poor accuracy, which can easily lead to artifacts.
By decomposing the image in the time domain and spatial domain, performing multi-layer pyramid decomposition respectively, and using pyramid matching and motion compensation technology for fine alignment and fusion, the amount of calculation is reduced and the matching accuracy is improved.
It improves the comprehensiveness and reliability of image denoising, reduces the amount of calculation, avoids artifact problems, and improves image quality.
Smart Images

Figure CN115439369B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of imaging technology, and in particular to an image noise reduction method and device, an electronic device, and a computer-readable storage medium. Background Art
[0002] In the image processing process, image noise reduction is a common way to improve image quality.
[0003] Related techniques involve converting all matched image blocks to the frequency domain, performing frequency domain denoising on each frequency component to reduce high-frequency noise, and then performing an inverse transform to the spatial domain to obtain the denoised image. This approach is computationally intensive and suffers from poor accuracy. Furthermore, it can lead to artifacts in some areas, resulting in poor noise reduction results.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0005] The purpose of the present disclosure is to provide an image noise reduction method and device, an electronic device, and a storage medium, thereby overcoming, at least to a certain extent, the problem of poor noise reduction effect caused by the limitations and defects of related technologies.
[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0007] According to a first aspect of the present disclosure, an image denoising method is provided, comprising: acquiring a time-domain denoised image of a current image and a reference frame image of the current image, and decomposing the time-domain denoised image to obtain multiple layers of time-domain images; performing spatial-domain denoising on the current image to obtain a spatial-domain denoised image corresponding to the current image, and decomposing the spatial-domain denoised image to obtain multiple layers of spatial-domain images; matching each layer of the time-domain image with each layer of the spatial-domain image to determine an aligned image of the time-domain denoised image; and fusing the spatial-domain denoised image and the aligned image to obtain a denoised image corresponding to the current image.
[0008] According to a second aspect of the present disclosure, an image denoising device is provided, comprising: a first image decomposition module, configured to obtain a time-domain denoised image of a current image and a reference frame image of the current image, and decompose the time-domain denoised image to obtain a multi-layer time-domain image; a second image decomposition module, configured to perform spatial-domain denoising on the current image to obtain a spatial-domain denoised image corresponding to the current image, and decompose the spatial-domain denoised image to obtain a multi-layer spatial-domain image; an image matching module, configured to match each layer of the time-domain image with each layer of the spatial-domain image to determine an aligned image of the time-domain denoised image; and an image fusion module, configured to fuse the spatial-domain denoised image with the aligned image to obtain a denoised image corresponding to the current image.
[0009] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the image denoising method of the first aspect and its possible implementation methods by executing the executable instructions.
[0010] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the image denoising method of the first aspect and its possible implementation methods are implemented.
[0011] In the technical solution provided in the embodiments of the present disclosure, on the one hand, the image is first finely aligned using each layer of temporal images and each layer of spatial images, and then temporal noise reduction is achieved on the current image by fusing the aligned images with the spatial noise reduction images. This avoids the limitation of related technologies that only relatively good noise reduction effects can be achieved in certain areas, avoids the problem of artifacts that are easily generated, increases the noise reduction range, and improves the comprehensiveness and reliability of noise reduction. On the other hand, by matching each layer of spatial images and each layer of temporal images, the layered matching can increase the search range during image matching and improve the matching accuracy, avoid the conversion between the temporal domain and the spatial domain, reduce the amount of computation during image matching, enhance the noise reduction effect, and improve the image quality of the noise-reduced image.
[0012] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0014] Figure 1 A schematic diagram shows an application scenario in which the image noise reduction method according to an embodiment of the present disclosure can be applied.
[0015] Figure 2 A schematic diagram schematically illustrates an image noise reduction method according to an embodiment of the present disclosure.
[0016] Figure 3 A schematic diagram schematically illustrates the decomposition of a time-domain denoised image in an embodiment of the present disclosure.
[0017] Figure 4 A schematic diagram schematically illustrates the decomposition of a spatially denoised image in an embodiment of the present disclosure.
[0018] Figure 5 The following schematically illustrates a flow chart of obtaining an aligned image in an embodiment of the present disclosure.
[0019] Figure 6 The flowchart of the first block matching according to the embodiment of the present disclosure is schematically shown.
[0020] Figure 7 The specific flow chart of the first block matching in the embodiment of the present disclosure is schematically shown.
[0021] Figure 8 The specific flow chart of the second block matching in the embodiment of the present disclosure is schematically shown.
[0022] Figure 9 The specific flow chart of the third block matching in the embodiment of the present disclosure is schematically shown.
[0023] Figure 10 The following schematically illustrates the process flow of image fusion in an embodiment of the present disclosure.
[0024] Figure 11 The overall process flow of time domain noise reduction in an embodiment of the present disclosure is schematically shown.
[0025] Figure 12 A block diagram of an image noise reduction device in an embodiment of the present disclosure is schematically shown.
[0026] Figure 13 A block diagram schematically illustrates an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0028] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0029] In related technologies, all matching image blocks can be converted to the frequency domain, and frequency domain denoising is performed on each frequency component in the frequency domain to reduce high-frequency noise, and then the image is inversely transformed to the spatial domain to obtain the denoised image. In the above-mentioned denoising scheme, on the one hand, the amount of calculation is very large. In order to ensure the denoising effect, it is necessary to search for matching blocks within a larger search window, and then all the reference blocks found must be converted to the frequency domain, and then each frequency component must also be converted to the frequency domain. After the frequency domain denoising is completed, it is inversely transformed back to the spatial domain twice to obtain the final denoised image. If the search window is too small, although the amount of calculation is reduced, the denoising effect will also decrease. On the other hand, the denoising effect of the image does not take into account the influence of the time domain. Sometimes, denoising only in the spatial domain will convert high-frequency noise into low-frequency noise, and the low-frequency noise will flicker when the video is played.
[0030] In order to solve the technical problems in the related art, an image denoising method is provided in an embodiment of the present disclosure, which can be applied to the image denoising process to enhance the application scenario of the image denoising effect. Figure 1 A schematic diagram showing a system architecture to which the image noise reduction method and apparatus according to embodiments of the present disclosure can be applied is shown.
[0031] like Figure 1As shown, the terminal 101 may be a smart device with an image processing function, such as a smart phone, a computer, a tablet computer, a smart speaker, a smart watch, an in-vehicle device, a wearable device, a monitoring device, or the like. The current image may be a captured image, each frame of a captured video, or an image stored in the terminal or a frame of a video.
[0032] In the disclosed embodiment, the terminal 101 may include a memory 102 and a processor 103. The memory is used to store images, and the processor is used to process images. The memory 102 may store a current image 104. The terminal 101 obtains the current image 104 and the reference frame image 105 from the memory 102 and sends them to the processor 103. The processor 103 performs spatial denoising and pyramid decomposition on the current image to obtain a multi-layer spatial image; performs temporal denoising and pyramid decomposition on the reference frame image to obtain a multi-layer temporal image; performs hierarchical alignment on the multi-layer spatial image and the multi-layer temporal image to obtain an aligned image; and fuses the aligned image with the spatial denoised image to generate a denoised image 106 after temporal denoising.
[0033] It should be noted that the image noise reduction method provided in the embodiment of the present disclosure may be executed by the terminal 101. The image noise reduction method may also be set in the terminal.
[0034] Figure 2 The image noise reduction method in the embodiment of the present disclosure is schematically shown in FIG, which specifically includes the following steps:
[0035] Step S210, obtaining a time-domain denoised image of a current image and a reference frame image of the current image, and decomposing the time-domain denoised image to obtain a multi-layer time-domain image;
[0036] Step S220, performing spatial denoising on the current image to obtain a spatial denoised image corresponding to the current image, and decomposing the spatial denoised image to obtain a multi-layer spatial image;
[0037] Step S230, matching each layer of the time domain image with each layer of the spatial domain image to determine an aligned image of the time domain denoised image;
[0038] Step S240: Fusing the spatial denoised image and the aligned image to obtain a denoised image corresponding to the current image.
[0039] In the disclosed embodiment, the current image may be a current frame image. Spatial denoising may be performed on the current image to obtain a spatially denoised image, and the spatially denoised image may be decomposed to obtain a multi-layer spatial image. For example, a four-layer pyramid decomposition may be performed on the spatially denoised image to obtain a multi-layer spatial image, and the spatially denoised image A1 may be decomposed into multi-layer spatial images a0 / a1 / a2 / a3.
[0040] At the same time, a reference frame image for the current image can be obtained. The reference frame image can be a previous frame image adjacent to the current frame. Temporal denoising can be performed on the reference frame image to obtain a temporal denoised image, and the temporal denoised image can be decomposed to obtain a multi-layer temporal image. For example, a four-layer pyramid decomposition can be performed on the temporal denoised image B0 to obtain a multi-layer temporal image b0 / b1 / b2 / b3.
[0041] Furthermore, pyramid matching can be performed on each layer of temporal and spatial images, with the pyramid matching performed multiple times at different scales. During each pyramid matching, matching blocks in each spatial image layer can be aligned within each spatial image layer of the same resolution, in order of resolution. This achieves multiple layered, multi-scale matching, resulting in an aligned image of the temporal denoised image.
[0042] On this basis, the aligned image can be fused with the spatial denoised image to obtain the denoised image after temporal denoising of the current image.
[0043] In the disclosed embodiments, the images are first finely aligned, and then temporal noise reduction is achieved by fusing the aligned images with the spatial noise reduction images. This avoids the limitation of related technologies that only achieves good noise reduction in certain areas, avoids the artifacts that are easily generated, and improves the noise reduction effect. Furthermore, by matching each layer of spatial and temporal images, the layered matching method increases the search range and improves matching accuracy, while reducing the computational complexity during matching, enhancing the noise reduction effect, and improving image quality.
[0044] Next, refer to Figure 2 As shown, each step in the image noise reduction method in the embodiment of the present disclosure is described in detail.
[0045] In step S210, a time-domain denoised image of a current image and a reference frame image of the current image is obtained, and the time-domain denoised image is decomposed to obtain a multi-layer time-domain image.
[0046] In the embodiment of the present disclosure, the current image may be a current frame image, may be each frame image in a video obtained by photographing the object to be photographed through the camera module of the terminal, or may be each frame image in a video obtained from the memory of the terminal.
[0047] The reference frame can be one or more of the n frames preceding the current image, such as a frame immediately preceding the current image. The temporal denoised image can be a denoised image obtained through temporal denoising. Based on this, the temporal denoised image of the reference frame can be the temporal denoised image of the previous frame. Temporal denoising is a 3D denoising method that exploits the temporal correlation of multiple frames to achieve denoising. The temporal denoised image of the reference frame can be represented, for example, as B0.
[0048] In some embodiments, to increase the search range, the temporal denoised image of the reference frame image can be decomposed to obtain multiple layers of temporal images. For example, the temporal denoised image of the reference frame image can be subjected to pyramid decomposition. Pyramid decomposition can be either Gaussian or Laplacian, with Gaussian pyramid decomposition used as an example. The bottom layer of the Gaussian pyramid is the image to be decomposed, i.e., the temporal denoised image of the reference frame image. Each layer upward is obtained by Gaussian filtering and downsampling the image in the adjacent layer. The downsampling here can be, for example, 1 / 2 sampling.
[0049] refer to Figure 3 The pyramid structure shown in the figure includes bottom-level images with successively lower resolutions and at least one top-level image. The bottom-level image is a time-domain denoised image used to represent a reference frame image. The at least one top-level image is obtained by Gaussian filtering and downsampling. In some embodiments, the top-level image includes a top-level image and an intermediate-level image. The top-level image is the top-level image with the lowest resolution, and all images in the top-level image except the top-level image are intermediate-level images. For example Figure 3 As shown, the pyramid includes multiple layers of images with different resolutions, and the resolution of any layer in the pyramid is higher than the resolution of the image above it. In other words, the resolution decreases from the bottom of the pyramid to the top, with the bottom image having the highest resolution and the top image having the lowest resolution.
[0050] It should be noted that the number of intermediate layer images can be one or more. When there are multiple intermediate layer images, the closer the intermediate layer images are to the top image, the lower the resolution. The number of intermediate layer images is determined by the number of Gaussian pyramid decompositions. For example, if four Gaussian pyramid downsamplings are performed, a four-layer image is obtained, including the bottom image, two intermediate layer images, and the top image.
[0051] Based on this, when performing four-fold Gaussian pyramid downsampling, the temporal denoised image of the reference frame image can first be used as the base image of the pyramid. The base image is then Gaussian filtered and downsampled to obtain a first intermediate image layer with a lower resolution than the base image. The first intermediate image layer is then Gaussian filtered and downsampled to obtain a second intermediate image layer with a lower resolution than the first intermediate image layer. The second intermediate image layer is further Gaussian filtered and downsampled to obtain a top image layer with a lower resolution than the second intermediate image layer. It should be noted that in some embodiments, the length and width of the upper image layer in the pyramid can be half the length and width of the adjacent lower image layer.
[0052] refer to Figure 3 As shown in , the time-domain denoised image B0 of the reference frame image can be downsampled by a four-layer Gaussian pyramid to obtain multi-layer time-domain images b0, b1, b2, and b3, wherein the resolutions of the multi-layer time-domain images b0, b1, b2, and b3 decrease in sequence from bottom to top, that is, the resolution of b0 is the largest and the resolution of b3 is the smallest.
[0053] In step S220, spatial denoising is performed on the current image to obtain a spatial denoised image corresponding to the current image, and the spatial denoised image is decomposed to obtain a multi-layer spatial image.
[0054] In the disclosed embodiments, spatial denoising can be a 2D denoising method that only addresses noise within a single frame. Spatial denoising can be achieved using various edge-preserving spatial denoising algorithms. These algorithms may include, but are not limited to, NonLocalMean, bilateral filtering, or other deep learning-based spatial denoising algorithms. The bilateral filtering algorithm is used as an example for illustration.
[0055] The bilateral filtering algorithm is a nonlinear filter that preserves edges, reduces noise, and smooths images. It also uses a weighted average method, representing the intensity of a pixel by the weighted average of the brightness values of surrounding pixels. This weighted average is based on a Gaussian distribution. The bilateral filter consists of two parts: one that is related to the spatial distance between pixels, and the other to the pixel difference. The weight of the bilateral filter is equal to the product of the spatial proximity factor and the brightness similarity factor. The spatial proximity factor is the Gaussian filter coefficient, and the brightness similarity factor is related to the spatial pixel difference.
[0056] For the current image, spatial denoising can be performed first to obtain a spatial denoised image to initially reduce noise and improve the accuracy of subsequent matching. For example, spatial denoising can be performed on the current image A0 to obtain a spatial denoised image A1.
[0057] In the disclosed embodiments, a spatially denoised image can be subjected to pyramid decomposition to obtain a multi-layer spatial image. Pyramid decomposition can be either Gaussian or Laplacian, with Gaussian pyramid decomposition used as an example. The bottom layer of the Gaussian pyramid is the image to be decomposed, i.e., the spatially denoised image. Each upward layer is obtained through Gaussian filtering and 1 / 2 downsampling.
[0058] refer to Figure 4 As shown in , the spatially denoised image A1 corresponding to the current image can be downsampled using a four-layer Gaussian pyramid to produce multi-layer spatial images a0, a1, a2, and a3 with varying resolutions, where a0 has the highest resolution and a3 has the lowest. Pyramid decomposition of the spatially denoised image to obtain the corresponding multi-layer spatial images facilitates matching the largest possible motion range and reduces computational effort during subsequent layered matching.
[0059] In step S230 , each layer of the time domain image and each layer of the spatial domain image are matched to determine an aligned image of the time domain denoised image.
[0060] In an embodiment of the present disclosure, after decomposing the multi-layer time domain image and the multi-layer spatial domain image into multi-layer images, the multi-layer time domain image and the multi-layer spatial domain image can be subjected to multiple layered multi-scale matching. Multiple layered multi-scale matching is used to perform multiple layered alignment of images. Multiple layered alignment refers to multiple layered alignments at different scales. Layered alignment can be used to perform image alignment on each layer of time domain image and each layer of spatial domain image. Layered alignment refers to performing block matching on each layer in the order of resolution to determine the motion vector MV (Motion Vector) of each matching block, thereby aligning each layer of the multi-layer time domain image and the multi-layer spatial domain image according to the motion vector. It should be noted that each layer of time domain image is aligned with each layer of spatial domain image of the same resolution. For example, the third layer of spatial domain image is aligned with the third layer of time domain image, the second layer of spatial domain image is aligned with the second layer of time domain image, and so on.
[0061] Based on this, when performing multiple image alignments at different scales on each layer of a multi-layer temporal image and a multi-layer spatial image, the number of these alignments can be set based on actual needs, for example, three, four, or so. The scale can be represented by the size of the matching block, which can be determined based on actual needs, for example, 16×16, 5×5, or so. The scale of each image alignment can be adjusted as long as the image alignment scale decreases with the number of alignments. For example, the temporal denoised image and the spatial denoised image of the reference frame image are decomposed into four layers according to pyramid decomposition, and three image alignments are performed: the first matching block size is 16×16, the second matching block size is 8×8, and the third matching block size is 5×5. By gradually reducing the matching block size and the image alignment scale, a more refined matching process can be achieved, improving matching accuracy and efficiency.
[0062] When performing hierarchical multi-scale image alignment, motion vector estimation can be performed for each spatial image layer in each temporal image layer, sequentially according to the order of resolution. Motion compensation can then be performed based on the motion vectors to achieve image alignment for each temporal image layer, until multiple image alignments at different scales are performed to determine an aligned image. The resolutions can be arranged in ascending order, i.e., from the top to the bottom of a pyramid. Therefore, motion vector estimation can be performed for each spatial image layer in each temporal image layer, sequentially according to the order of resolution, to determine the movement of each spatial image layer relative to each temporal image layer, thereby achieving image alignment. Furthermore, each hierarchical image alignment is performed in ascending order of resolution. For example, motion vector estimation is performed for each layer of temporal images, in the order b3, b2, b1, and b0.
[0063] Motion vectors are used to represent the position changes of corresponding areas between different images. Motion vectors can be determined by block matching. The block matching method can divide a frame of image into multiple non-overlapping image blocks, such as image blocks of size N×N. Each image block then searches for the most matching image block within the search window of the previous frame according to a certain matching criterion. The resulting displacement difference is called a motion vector. For example, each layer of time domain image and each layer of spatial domain image can be divided into multiple non-overlapping image blocks, and it is assumed that all pixels within an image block undergo translational motion at the same speed, and the displacement of all image blocks is the same. Based on this, for each image block in each layer of spatial domain image, a similar image block can be searched within the search window of each layer of time domain image as a similar block, and the position of the similar block in each layer of time domain image is considered to be the position of the image block in each layer of spatial domain image before displacement. The motion vector can be calculated by the difference between the coordinates of the similar block and the image block.
[0064] In the disclosed embodiments, during each hierarchical alignment, a matching block may be used for motion vector estimation. The matching block used in each layer of each hierarchical alignment is the same, but the matching blocks used in different hierarchical alignment processes may be different. For example, the size of the matching block in the hierarchical alignment decreases as the number of hierarchical alignments increases.
[0065] Each spatial image can be used as a reference image for each layer, and each temporal image can be used as an image to be aligned for each layer. Furthermore, during the first block matching, a motion vector is estimated for each first matching block in the current layer reference image in the image to be aligned for the current layer, and a motion vector for the current layer reference image is determined, thereby determining a motion vector for each layer reference image. The current layer can be each layer.
[0066] In some embodiments, during each block matching process, the third-layer spatial image a3 is used as the third-layer reference image, and the third-layer temporal image b3 is used as the third-layer image to be aligned. An MV estimation is performed on the third-layer reference image in the third-layer image to be aligned b3. The specific estimation method may include: for each first matching block blockA in the third-layer reference image a3, searching within a search window in the third-layer image to be aligned b3 for a block most similar to the first matching block as a similar block blockB. The first matching block blockA is a block of size R0×R0, where R0×R0 may be 15×15. The search window may be of size W0×W0. Assuming the coordinates of the similar block blockB found are (b_x, b_y), and the coordinates of the first matching block blockA are (a_x, a_y), the motion vector of the first matching block blockA in the third-layer image to be aligned b3 is the sum of the two coordinate vectors, which can be expressed as (b_x-a_x, b_y-a_y). Based on this, the motion vectors of all first matching blocks in the third spatial domain image in the third layer of the image to be aligned can be obtained. Repeating the above method, the motion vectors of all matching blocks in each spatial domain image layer in each temporal domain image layer can be obtained.
[0067] Motion compensation describes the process of moving each small block from the previous frame to a specific position in the current frame. Specifically, it describes the process of moving each similar block in each layer of the aligned image to a corresponding position in each spatial image layer. Based on the motion vector of each matching block, the image blocks at each motion vector are extracted from the target layer's aligned image and assembled block by block. The resulting image after motion compensation is called the aligned image.
[0068] For example, for example, the motion vector MV of the first matching block in the target layer reference image a0 is 2, and the image block with a distance of 2 from the first matching block in the target layer image to be aligned b0 is a similar block. In order to compensate, the compensation coordinates are obtained according to the coordinates and motion vector of the first matching block, and the image block corresponding to the compensation coordinates is obtained from the target layer image to be aligned b0 to obtain a new image to obtain the aligned image. For example, the coordinates of the first matching block are (100, 100) and the motion vector is (2, 2). The two coordinates are added to obtain the compensation coordinates (102, 102). Therefore, the image block corresponding to the compensation coordinates (102, 102) in b0 can be obtained. Based on these image blocks, a new image is obtained, that is, the aligned image after alignment, and it is used as the first compensated image. The aligned image obtained by the first motion compensation is the first compensated image B01. In the embodiment of the present disclosure, through motion compensation, it is possible to achieve temporal and spatial image alignment for each layer of temporal image and each layer of spatial image, thereby improving the effect of image alignment.
[0069] On this basis, the motion vector of each layer of spatial domain image is estimated in turn by block matching, and the process of motion compensation based on the motion vector can be achieved through three-way block matching, and then the aligned image is determined based on the matching results of the three-way block matching. Figure 5 The flowchart for determining the aligned image is schematically shown in FIG. Figure 5 As shown in , it mainly includes the following steps:
[0070] In step S510, motion vector estimation is performed on the first matching block in each layer of spatial domain image through the first block matching in each layer of temporal domain image in turn, a motion vector diagram of each layer of spatial domain image is determined, and motion compensation is performed based on the motion vector diagram to determine a first differential image.
[0071] In this step, first, the first matching blocks in each layer of spatial domain image can be arranged in order of resolution from small to large, and the first block matching can be performed in each layer of time domain image. The motion vector of the first matching block is estimated through the first block matching to obtain the motion vector of each first matching block, and the motion vector map of each layer of spatial domain image is determined. Motion compensation is performed based on the motion vector map to determine the first differential image.
[0072] Among them, reference Figure 6 As shown in , the process of the first block matching may include the following steps, and Figure 6 The steps in are specific implementations of step S510, wherein:
[0073] Step S610: using the spatial domain image of the current layer as the reference image of the current layer, and using the temporal domain image of the current layer as the image to be aligned of the current layer;
[0074] Step S620, estimating a motion vector for each first matching block in the current layer reference image in the image to be aligned, determining a motion vector of the current layer reference image, and determining a motion vector of each layer reference image;
[0075] Step S630, performing motion compensation according to the motion vector map of the target layer reference image to determine a first compensated image;
[0076] Step S640: Determine the first differential image based on the first compensated image.
[0077] Exemplarily, the current layer reference image can be a reference image for each layer, and the image to be aligned for the current layer can be an image to be aligned for each layer. During the first block matching, a motion vector for each first matching block in the current layer reference image can be determined within the search window of the image to be aligned for the current layer based on similar blocks to each first matching block in the current layer reference image. The search windows for each layer can be the same or different. In some embodiments, different methods can be selected to determine the search window in the image to be aligned for the current layer depending on whether the current layer reference image is the top layer reference image. If the current layer reference image is the top layer reference image, the search window for the current layer is determined directly based on a first preset window. The first preset window can be a W0×W0 window, and the first preset window can be directly configured without any correlation with other parameters. If the current layer reference image is not the top layer reference image, for example, a middle or bottom layer, its motion estimation can be performed based on the motion estimation of the previous layer. In this case, the search window for the current layer can be determined by combining the motion vector map of the previous layer reference image, the coordinates of the first matching block, and the second preset window. The previous layer reference image here refers to a reference image adjacent to the current layer and close to the top layer, and the resolution of the previous layer reference image is lower than that of the current layer reference image. For example, if the current layer reference image is a1, then the previous layer reference image is a2. Specifically, a center point can be determined based on the motion vector and coordinates of the first matching block in the previous layer reference image. Based on this center point, a second preset window size is determined to serve as the search window for the current layer. The second preset window can be a W1×W1 window, and its size can be the same as or different from the first preset window, depending on actual needs.
[0078] For example, in the second layer, if the coordinates of the first matching block blockA are (a_x, a_y), and the motion vector of the first matching block blockA obtained from the reference image of the previous layer is (mv_x, mv_y), then the search window of the second layer can be the range of W1×W1 centered at (a_x+mv_x, a_y+mv_y) in the image to be aligned b2 of the second layer. In a similar way, the search window of the first layer b1 and the bottom layer b0 can be calculated.
[0079] After determining the search window, a search can be performed in each layer of the image to be aligned based on the search window to obtain similar blocks that are similar to the first matching block. Similar blocks can be determined using the minimum SAD (Sum of Absolute Difference), minimum MAD (Mean Absolute Difference), or MSE (Mean Squared Error), which can be selected based on computational requirements. The size and shape of the search window used to search for similar blocks, as well as the size of the offset when searching for each similar block, can also be set according to actual needs and are not specifically limited here.
[0080] Based on the similar blocks, a motion vector of each first image block in each layer of the reference image in the image to be aligned can be determined based on a vector represented by the difference between the coordinates of the similar blocks and the coordinates of each first matching block. Furthermore, a motion vector map for each layer of the reference image can be determined based on a graph composed of the motion vectors of each first matching block in each layer of the reference image.
[0081] On this basis, the motion vector of each first matching block in the target layer reference image can be determined based on the motion vector map of the target layer reference image; the compensation coordinates are determined based on the motion vector of each first matching block and the coordinates of each first matching block, and the image blocks corresponding to the compensation coordinates are obtained from the target layer image to be aligned for motion compensation to obtain the first aligned image. The target layer reference image can be the underlying reference image, and the target layer image to be aligned can be the underlying time domain image. The motion vector of each first matching block and the coordinates of each first matching block can be added to obtain the compensation coordinates. The image blocks corresponding to the compensation coordinates in the underlying image to be aligned are extracted to form a new image to determine the aligned image, and the aligned image is the first compensated image obtained by the first block matching. When performing motion compensation, the image block (pixel block) corresponding to the compensation coordinate in the underlying image to be aligned can be placed at the position in the new image, that is, the position of the current first matching block, for motion compensation.
[0082] Since the first compensated image obtained by motion compensation is relatively rough, there may be incomplete alignment, such as misalignment in the texture area. Therefore, a first differential image can be determined based on the first compensated image to obtain the output result of the first block matching stage. Exemplarily, differential processing can be performed based on the first compensated image and the target layer reference image to obtain the first differential image. The first differential image can be used to reflect the alignment effect of the first compensated image and the target layer reference image. The area in the first differential image where the value is a preset value represents the aligned area, and the area where the value is not the preset value represents the misaligned area. The preset value can be, for example, 0, which is used to indicate that the type of the area is an aligned area.
[0083] It should be noted that, during the first block matching, the size of the first matching block R1 may be relatively large, for example, 16 × 16. Using a larger first matching block for the first block matching can speed up the matching process and improve the matching efficiency.
[0084] In some embodiments, a3 / a2 / a1 / a0 is used as the reference image for each layer, and b3 / b2 / b1 / b0 is used as the image to be aligned for each layer. For example, the size of the matching block is R, and the interval between matching blocks is also R. The process of the first block matching is described in detail. Figure 7 As shown in:
[0085] Step S701: Using the third-layer spatial domain image a3 as a reference image in the third layer, and performing motion vector estimation based on the third-layer image to be aligned b3, the specific process may include: searching for a similar block that is most similar to the first search block within a search window of size W0×W0 in the third-layer image to be aligned for each first matching block blockA of size R1×R1 in the third-layer reference image a3. Assuming that the coordinates of the similar block blockB found are (b_x, b_y), and the coordinates of the first matching block blockA are (a_x, a_y), then the motion vector of the first matching block blockA in the third-layer image to be aligned b3 is (b_x-a_x, b_y-a_y). After all the first matching blocks are traversed, a graph consisting of the motion vectors of each first matching block in the third-layer reference image is obtained, i.e., the motion vector MVmap3 of the third-layer reference image a3.
[0086] In step S702, when estimating the motion vector of the second-layer image, motion estimation is performed using the second-layer spatial domain image a2 as the reference image. However, the motion estimation of the second-layer image is based on the motion estimation of the previous layer (the third layer). For example, if the coordinates of the first matching block blockA are (a_x, a_y), and the motion vector of the first matching block is (mv_x, mv_y) according to the motion vector MVmap3 of the third-layer reference image a3, the second-layer motion estimation can be performed within a search window of size W1×W1 centered at (a_x+mv_x, a_y+mv_y) in the second-layer image to be aligned b2. The calculation method for obtaining the motion vectors of similar blocks and each first matching block can be the same as the calculation method for the motion vector of the third layer, and will not be repeated here. Furthermore, after the motion vector of each first matching block in the second-layer reference image a2 is calculated, the motion vector MVmap2 of the second-layer reference image a2 can be obtained.
[0087] Similarly, in step S703 , the motion vector MVmap1 of the first layer reference image a1 is calculated with reference to the motion vector MVmap2 of the second layer reference image a2 .
[0088] Step S704 , the motion vector Mvmap0 of the zeroth layer reference image (bottom layer reference image) a0 is calculated with reference to the motion vector Mvmap1 of the first layer reference image a1 , thereby calculating the displacement of each first matching block of the bottom layer reference image a0 in the bottom layer image to be aligned b0 .
[0089] Step S705 , according to the motion vector of each first matching block, image blocks at each motion vector are extracted from the underlying image to be aligned b0 , and assembled block by block to perform the first motion compensation. The image obtained after the first motion compensation is the first compensated image B01 .
[0090] In step S706, a first differential image Diffmask1 is obtained by performing a differential process on the first compensated image B01 and the underlying reference image a0. This first differential image Diffmask1 reflects the alignment between the first compensated image B01 and the underlying reference image a0. Specifically, the differential process can be pixel subtraction to reduce similarities and highlight differences. A value of 0 in Diffmask1 indicates that the alignment is successful; conversely, a non-zero value indicates that the alignment failed.
[0091] Through the first block matching from steps S701 to S706, image matching and alignment are performed on each layer of the multi-layer temporal image and the multi-layer spatial image to obtain a first difference image. This layered matching increases the matching range and reduces the computational complexity. Furthermore, fast matching can be achieved using the first matching block.
[0092] In step S520, based on the first differential image, motion vector estimation is performed on the second matching blocks of each layer of spatial domain image in each layer of temporal domain image through a second block matching, a motion vector diagram of each layer of spatial domain image is determined, and motion compensation is performed based on the motion vector diagram to determine the second differential image.
[0093] In the disclosed embodiment, since the first block matching has already aligned some areas, to reduce the amount of computation, only the misaligned areas in the first block matching need to be processed. Based on this, the first difference image obtained from the first block matching can be used as a guide, and fine matching can be performed only on the misaligned areas corresponding to the first difference image. Specifically, the first misaligned area of each layer of the time domain image can be determined based on the first difference image. The first misaligned area refers to the misaligned area determined based on the first difference image. Specifically, the misaligned area can be determined based on the area with a value of 0 in the first difference image.
[0094] During the second block matching process, motion vector estimation can be performed based on the second matching block. The size of the second matching block R2 can be smaller than the first matching block R1. For example, R2×R2 can be 8×8. Hierarchical matching with smaller second matching blocks can achieve fine matching.
[0095] The second block matching process is basically the same as the first block matching process, and mainly includes the following steps:
[0096] Each spatial image layer is used as the reference image layer, and each temporal image layer is used as the image to be aligned layer. A first non-aligned region is determined from each image to be aligned based on the first differential image. A second block matching is performed on the second matching blocks in each reference image layer in the first non-aligned region. A motion vector is determined for each second matching block based on its similarity to a block in each reference image layer. A motion vector map of each reference image layer in each temporal image layer is determined based on the motion vector map of each second matching block.
[0097] Based on the motion vector map of the target layer reference image represented by the underlying reference image, the motion vector of each second matching block is determined, the compensation coordinates are determined based on the sum of the motion vector and the coordinates of the second matching block, and the image blocks are obtained from the target layer image to be aligned based on the compensation coordinates for motion compensation to determine the second compensated image that is not fully aligned.
[0098] A second differential image is obtained by performing differential processing on the second compensated image and the target layer reference image. The second differential image may include an aligned region and a non-aligned region, but the non-aligned region included therein is smaller than the non-aligned region in the first differential image, and the non-aligned region may still be an area represented by a value of 0.
[0099] During the second block matching, matching can be performed based on the first difference image. Since the first difference image is obtained based on the target layer reference image, for other layers, the first difference image can be scaled according to the resolution relationship between each layer and the target layer to obtain the corresponding difference image for each layer. This allows alignment of each layer's spatial and temporal images based on the corresponding difference image. The resolution relationship can be the resolution relationship between the bottom layer reference image and the middle and top layers.
[0100] In some embodiments, during the second block matching alignment, only the areas where the first block matching alignment failed need to be fine-tuned. Figure 8 As shown in , the second block matching is guided by the first differential image Diffmask1 and only matches the non-zero regions in the first differential image Diffmask1. The second block matching process includes steps S801 to S806. The specific execution steps are the same as steps S701 to S706 and are not repeated here.
[0101] It's important to note that during the second block matching, the size of the second matching block, R2 × R2, is smaller than that of the first matching block. This allows for more precise matching of non-aligned areas based on the smaller matching block size. After motion compensation, the resulting image for the second layered alignment is the second compensated image B02. This image is then subtracted from the target layer reference image a0 to produce a new differential image, the second differential image Diffmask2.
[0102] In step S530, based on the second differential image, the third matching block of each layer of reference image is sequentially subjected to a third block matching in each layer of time domain image to determine the motion vector of each layer of reference image, and motion compensation is performed according to the motion vector to determine the aligned image.
[0103] In the embodiment of the present disclosure, the third block matching may be performed with the second differential image as a guide, and the size of the third matching block used in the third block matching is the smallest.
[0104] Because the first and second block matching processes have already aligned some areas, to reduce computational complexity, the third block matching process only needs to process the misaligned areas from the second block matching process. Specifically, the second misaligned area of each layer of the temporal image can be determined based on the second difference image. The second misaligned area refers to the misaligned area determined based on the second difference image. Specifically, it can be determined based on the area with a value of 0 in the second difference image and can be different from the first misaligned area.
[0105] During the third block matching, matching is performed based on the second difference image. Since the second difference image is obtained based on the target layer reference image, for other layers, the second difference image can be scaled according to the resolution relationship to obtain the corresponding difference images of the other layers. This allows alignment of each layer's spatial and temporal images based on the corresponding difference images. The resolution relationship can be the resolution correspondence between the bottom layer reference image and the reference images of the other layers.
[0106] During the third block matching process, motion vector estimation can be performed based on the third matching block. The size of the third matching block R3 can be smaller than the second matching block R2. For example, R3×R3 can be 5×5. Hierarchical matching with smaller third matching blocks can achieve fine matching.
[0107] The third block matching process is basically the same as the second block matching process, and mainly includes the following steps:
[0108] First, a second non-aligned area is determined. A third block matching is performed on the third matching blocks in each layer of the reference image in the second non-aligned area. A motion vector of each third matching block is determined based on a similar block of each third matching block in each layer of the reference image. A motion vector diagram of each layer of the reference image in each layer of the time domain image is determined based on a graph composed of the motion vectors of each third matching block.
[0109] According to the motion vector map of the target layer reference image represented by the underlying reference image, the motion vector of each third matching block is determined, and the compensation coordinates are determined based on the sum of the motion vector and the coordinates of the third matching block. The image blocks are obtained from the target layer image to be aligned according to the compensation coordinates for motion compensation to determine the incompletely aligned third compensated image. It should be noted that when performing the third block matching, since it is the last matching, all areas need to be aligned, and the differential image is no longer obtained. The third compensated image obtained by motion compensation is directly used as the aligned image. The aligned image refers to the image after the time domain denoised image and the spatial domain denoised image of the reference frame image are finally aligned.
[0110] refer to Figure 9As shown in , the third block matching is guided by the second differential image Diffmask2, and only the non-zero areas in the second differential image Diffmask2 are matched. The process of the third block matching includes steps S901 to S905, and the specific execution steps are the same as those of steps S801 to S805, which will not be repeated here. It should be noted that during the third block matching, the size R3×R3 of the third matching block is smaller than the size of the first matching block and smaller than the size of the second matching block, so as to achieve a more precise matching based on the small-sized matching block. After completing the motion compensation, the image of the third hierarchical alignment is obtained, that is, the third compensated image B03, and the third compensated image can be determined as the final aligned image B1.
[0111] It should be noted that when performing motion compensation, the compensation block corresponding to the matching block in each spatial image layer can first be determined. The compensation block refers to the image block used for motion compensation and can include part or all of the matching block. Different block matching processes have different compensation block sizes, so different methods can be used to determine the compensation block based on the number of block matches. For example, during the first and second block matching processes, the entire area of the corresponding matching block is used as the compensation block. Specifically, during the first block matching process, the entire area of the first matching block is used as the compensation block; during the second block matching process, the entire area of the second matching block is used as the compensation block. For example, based on the motion vector of the first or second matching block, the image block corresponding to the compensation coordinates is obtained from the target layer's image to be aligned. The image block (pixel block) corresponding to the compensation coordinates is placed at the corresponding position in the new image, i.e., the position of the current first or second matching block. Motion compensation is performed on the compensation block of the same size as the first or second matching block. For example, if the first matching block is 16×16, compensation is also performed on the 16×16 first matching block.
[0112] In addition, during the third block matching, a portion of the matching block corresponding to the third block matching is used as a compensation block, and the compensation block is updated according to the offset. Specifically, a portion of the third matching block can be used as the compensation block. For example, the third matching block is 5×5, and a fourth matching block at the center of the third matching block can be used as the compensation block. The size of the fourth matching block can be smaller than the third matching block, for example, 2×2. Furthermore, since the compensation block is smaller than the third matching block, the compensation block can be updated according to the offset. The offset represents the step size of the compensation block movement and can be set according to actual needs, for example, 2 pixels. Based on this, the center of a new compensation block can be obtained according to the offset in a preset direction. The center of the new compensation block is then expanded in each direction to obtain a new third matching block, thereby obtaining a new compensation block, i.e., a new 2×2 fourth matching block, until the entire image is traversed. The preset direction can be determined according to actual needs, such as horizontal or vertical. Furthermore, based on the motion vector of the third matching block, the image block corresponding to the compensation coordinates can be obtained from the target layer image to be aligned, and motion compensation can be performed on the 2×2 compensation block to achieve fine compensation.
[0113] During the first and second block matching, compensation blocks of the same size as the matching blocks are used for compensation, which improves the efficiency of motion compensation. During the third block matching, compensation blocks of smaller sizes are used for compensation, which enables fine compensation, improves compensation accuracy, and realizes fast and precise compensation.
[0114] In the disclosed embodiments, pyramid decomposition is performed on the temporally denoised image and the spatially denoised image, followed by multiple layered matching at different scales. The temporally denoised image is aligned with the spatially denoised image. This multi-layer matching approach increases the search range during matching and reduces the computational complexity. Furthermore, multiple layered matching at different scales improves the accuracy of each layer's matching. By performing layered matching at progressively smaller scales, the precision of image matching is enhanced. Layered matching employs a multi-resolution, coarse-to-fine search strategy, using matching blocks of different sizes at different layers of different resolutions. Motion vectors are estimated recursively at different scales. The first layer uses large blocks and a large search window, reliably estimating the basic direction of block motion. The next layer uses the motion vectors from the previous layer for a more accurate estimate. Because the previous layer ensures the overall reliability of the motion vectors, it overcomes mismatches caused by small-block matching and improves matching accuracy. Furthermore, since the motion vectors of the matching blocks are estimated during denoising, noise reduction is achieved across all image regions.
[0115] In step S240, the spatial domain denoised image and the aligned image are fused to obtain a denoised image corresponding to the current image.
[0116] In the embodiment of the present disclosure, the aligned images of the spatial domain denoised image and the frequency domain denoised image can be fused to achieve spatial and temporal domain fusion, and temporal domain denoising can be performed on the motion area and the non-motion area.
[0117] Figure 10 The flowchart for fusion is shown schematically in FIG. Figure 10 As shown in , it mainly includes the following steps:
[0118] In step S1001, motion detection is performed on the spatial denoised image 1010 and the aligned image 1020 to determine a difference map 1030;
[0119] In step S1002, a fusion weight map 1040 is determined according to the difference map;
[0120] In step S1003 , the spatial denoised image and the aligned image are fused according to the fusion weight map to obtain the denoised image 1050 .
[0121] The spatial denoised image and the aligned image can be smoothed separately; the smoothed spatial denoised image and the smoothed aligned image can be subtracted to obtain a difference map, and the difference map can be morphologically processed to obtain the difference map. Specifically, the morphological processing can include removing scattered points and filling holes. Based on this, for example, the spatial denoised image 1010 and the aligned image 1020 can be smoothed separately and then subtracted to obtain a difference map between the two. The difference map can then be morphologically processed to remove isolated scattered points and fill holes to obtain the difference map.
[0122] Next, the Motion Mask difference map is used as a guide to generate the Alpha Map. This Alpha Map represents the degree of fusion between the spatially denoised image and the temporally denoised image. The fusion principle is: smaller values in the Motion Mask difference map indicate greater similarity between the spatially denoised image and the temporally denoised image, resulting in a larger Alpha Map value. Conversely, larger values in the Motion Mask difference map indicate greater dissimilarity between the spatially denoised image and the temporally denoised image, resulting in a smaller Alpha Map value.
[0123] Based on this, fusion can be performed according to the fusion weight map. Specifically, the fusion strategy is denoised image C0 = A1*(1.0-AlphaMap)+B1*AlphaMap, where denoised image C0 is the image after temporal denoising of the current image. On this basis, the denoised image after temporal denoising of the current image will be used as the image to be aligned for the next frame during temporal denoising of the next frame, i.e., the temporal denoised image of the next frame. This process is then repeated, and all frames are subjected to combined spatial and temporal denoising to obtain the temporal denoising result for each frame.
[0124] Figure 11 The flowchart of combining spatial and temporal domain noise reduction is shown schematically in FIG. Figure 11 As shown in , it mainly includes the following steps:
[0125] In step S1101 , the current image A0 is obtained.
[0126] In step S1102 , spatial denoising is performed on the current image to obtain a spatial denoised image A1 .
[0127] In step S1103 , a spatial domain denoised image A1 is acquired.
[0128] In step S1104 , pyramid decomposition is performed on the spatial denoised image A1 to obtain multi-layer spatial images a0 / a1 / a2 / a3 .
[0129] In step S1105 , a time-domain denoised image B0 is acquired.
[0130] In step S1106 , pyramid decomposition is performed on the time-domain denoised image B0 to obtain multi-layer time-domain images b0 / b1 / b2 / b3 .
[0131] In step S1107 , the multi-layer temporal domain images and the multi-layer spatial domain images are hierarchically aligned.
[0132] In step S1108 , the aligned image B1 is acquired.
[0133] In step S1109 , motion detection is performed on the aligned image B1 and the spatially denoised image A1 .
[0134] In step S1110 , a difference map corresponding to motion detection is obtained.
[0135] In step S1111 , a fusion weight map is determined based on the difference map, and the aligned image and the spatial domain denoised image are double-frame fused based on the fusion weight map.
[0136] In step S1112, a denoised image C0 represented by the fusion result is obtained. The denoised image may be an image corresponding to the current image A0 after denoising in the time domain.
[0137] The technical solution in the disclosed embodiment first performs fine alignment on the image and then performs time domain noise reduction, thereby avoiding the problem of artifacts that are easily generated by directly performing time domain fusion without alignment, and also avoiding the problem of poor noise reduction effect in texture areas caused by not aligning or coarse alignment in areas with complex textures. It improves the noise reduction effect in non-motion areas and motion areas, and also improves the quality, reliability and comprehensiveness of noise reduction. Image alignment and matching are performed by using a pyramid layer and multiple scale matching methods. Since each layer has a corresponding matching range, the matching motion range can be expanded and the amount of calculation is reduced compared to the need to convert to the time domain in related technologies. Through multiple matching at different scales, fine block matching is performed based on the last non-aligned area, so that the two images are aligned more accurately, thereby improving the accuracy of matching. The proportion of non-aligned areas during fusion is small, which improves the effect of time domain noise reduction, thereby improving the image quality of the denoised image.
[0138] The present disclosure provides an image noise reduction device, referring to Figure 12 As shown in , the image noise reduction device 1200 may include:
[0139] A first image decomposition module 1201 is configured to obtain a time-domain denoised image of a current image and a reference frame image of the current image, and decompose the time-domain denoised image to obtain a multi-layer time-domain image;
[0140] A second image decomposition module 1202 is configured to perform spatial denoising on the current image to obtain a spatial denoised image corresponding to the current image, and decompose the spatial denoised image to obtain a multi-layer spatial image;
[0141] An image matching module 1203 is configured to match each layer of the time domain image with each layer of the spatial domain image to determine an aligned image of the time domain denoised image;
[0142] The image fusion module 1204 is configured to fuse the spatial denoised image and the aligned image to obtain a denoised image corresponding to the current image.
[0143] In an exemplary embodiment of the present disclosure, the image matching module includes: a multi-scale alignment module, configured to perform multiple image alignments of different scales on each layer of temporal domain images and each layer of spatial domain images to determine the aligned images.
[0144] In an exemplary embodiment of the present disclosure, the first image decomposition module includes: a downsampling module, configured to downsample the time-domain denoised image to obtain multi-layer time-domain images, where the resolutions of the multi-layer time-domain images are different.
[0145] In an exemplary embodiment of the present disclosure, the image fusion module includes: a difference map acquisition module, which is used to perform motion detection on the spatial denoised image and the aligned image to determine a difference map; a weight fusion module, which is used to determine a fusion weight map based on the difference map, and fuse the spatial denoised image and the aligned image according to the fusion weight map to obtain the denoised image.
[0146] In an exemplary embodiment of the present disclosure, the multi-scale alignment module includes: an alignment control module, which is used to estimate the motion vector of each layer of spatial domain image in each layer of temporal domain image in sequence according to the order of resolution, and perform motion compensation based on the motion vector until multiple image alignments are performed to determine the aligned image.
[0147] In an exemplary embodiment of the present disclosure, the alignment control module includes: a block matching module, which is used to use a block matching method to sequentially estimate the motion vector of the matching blocks of each layer of spatial domain image in each layer of temporal domain image, and perform motion compensation based on the motion vector to determine the aligned image.
[0148] In an exemplary embodiment of the present disclosure, the block matching module includes: a first block matching module, which is used to perform motion vector estimation on the first matching block in each layer of spatial domain image in each layer of time domain image through the first block matching in turn, determine the motion vector diagram of each layer of spatial domain image, and perform motion compensation based on the motion vector diagram to determine a first differential image; a second block matching module, which is used to perform motion vector estimation on the second matching block in each layer of spatial domain image in turn, based on the first differential image, determine the motion vector diagram of each layer of spatial domain image, and perform motion compensation based on the motion vector diagram to determine a second differential image; a third block matching module, which is used to perform motion vector estimation on the third matching block in each layer of spatial domain image in turn, based on the second differential image, determine the motion vector diagram of each layer of spatial domain image, and perform motion compensation based on the motion vector diagram to determine the aligned image.
[0149] In an exemplary embodiment of the present disclosure, the first block matching module includes: an image configuration module, which is used to use the current layer spatial domain image as the current layer reference image and the current layer time domain image as the current layer to be aligned image; a motion vector estimation module, which is used to perform motion vector estimation on each first matching block in the current layer reference image in the to-be-aligned image, determine the motion vector diagram of the current layer reference image, and determine the motion vector diagram of each layer reference image; a motion compensation module, which is used to perform motion compensation according to the motion vector diagram of the target layer reference image and determine the first compensated image; and a differential image acquisition module, which is used to determine the first differential image based on the first compensated image.
[0150] In an exemplary embodiment of the present disclosure, the motion vector estimation module includes: a motion vector determination module, which is used to determine the motion vector of each first matching block in the current layer reference image based on the similar block of each first matching block in the image to be aligned; and a motion vector diagram determination module, which is used to determine the motion vector diagram based on a diagram composed of the motion vectors of each first matching block in the current layer reference image.
[0151] In an exemplary embodiment of the present disclosure, the motion vector determination module includes: a similar block search module, used to search for similar blocks of each first matching block in the image to be aligned based on a search window of the current layer; and a motion vector calculation module, used to determine the motion vector of each first matching block in the image to be aligned based on the difference between the coordinates of each first matching block and the coordinates of the similar block.
[0152] In an exemplary embodiment of the present disclosure, the device also includes: a first search window determination module, which is used to determine the search window of the current layer according to the first preset window if the current layer reference image is a top layer reference image; and a second search window determination module, which is used to determine the search window of the current layer in combination with the motion vector map of the previous layer reference image of the current layer reference image, the coordinates of the first matching block and the second preset window if the current layer reference image is not a top layer reference image.
[0153] In an exemplary embodiment of the present disclosure, the second search window determination module includes: a window control module, which is used to determine the center point based on the motion vector of the first matching block in the previous layer reference image and the coordinates of the first matching block, and determine the search window of the current layer based on the center point and the second preset window.
[0154] In an exemplary embodiment of the present disclosure, the motion compensation module includes: a motion vector acquisition module, which is used to determine the motion vector of each first matching block in the target layer reference image based on the motion vector map of the target layer reference image; a compensation control module, which is used to determine the compensation coordinates according to the motion vector of each first matching block and the coordinates of each first matching block, and obtain the image blocks corresponding to the compensation coordinates from the target layer image to be aligned for motion compensation to obtain a first compensated image.
[0155] In an exemplary embodiment of the present disclosure, the first differential image acquisition module includes: a differential processing module, configured to perform differential processing based on the first compensation image and the target layer reference image to acquire the first differential image.
[0156] In an exemplary embodiment of the present disclosure, the second block matching module includes: a first non-aligned area determination module, which is used to determine the first non-aligned area of each layer of the time domain image based on the first differential image; a block matching module, which is used to perform a second block matching on the second matching block in each layer of the spatial domain image in the first non-aligned area in turn, and determine the motion vector of each layer of the spatial domain image in each layer of the time domain image; a motion compensation module, which is used to perform motion compensation based on the motion vector of the target layer spatial domain image to determine the second compensated image; and a second differential image acquisition module, which is used to perform differential processing based on the second compensated image and the target layer spatial domain image to obtain the second differential image.
[0157] In an exemplary embodiment of the present disclosure, the third block matching module includes: a second non-aligned area determination module, which is used to determine the second non-aligned area of each layer of the time domain image based on the second differential image; a block matching module, which is used to perform a third block matching in the second non-aligned area on the third matching block in each layer of the spatial domain image in turn to determine the motion vector graph of each layer of the spatial domain image; a motion compensation module, which is used to perform motion compensation based on the motion vector graph of the target layer spatial domain image to determine a third compensated image; and an aligned image determination module, which is used to determine the third compensated image as the aligned image.
[0158] In an exemplary embodiment of the present disclosure, the difference map acquisition module includes: a smoothing processing module for smoothing the spatial denoised image and the aligned image respectively; a morphological processing module for subtracting the smoothed spatial denoised image and the smoothed aligned image to obtain a difference map, and performing morphological processing on the difference map to obtain the difference map.
[0159] In an exemplary embodiment of the present disclosure, the motion compensation module includes: a compensation block determination module, which is used to determine the compensation block for motion compensation based on the matching block in each layer of the spatial domain image; a compensation control module, which is used to determine the compensation coordinates based on the motion vector of each matching block in the target layer spatial domain image, and perform motion compensation on the compensation block of each layer of the spatial domain image based on the image block corresponding to the compensation coordinates.
[0160] In an exemplary embodiment of the present disclosure, the compensation block determination module includes: a first compensation block determination module, which is used to use the entire area of the matching block as the compensation block during the first block matching and the second block matching; and a second compensation block determination module, which is used to use a partial area of the matching block as the compensation block during the third block matching and update the compensation block according to the offset.
[0161] It should be noted that the specific details of each part of the above-mentioned image noise reduction device have been described in detail in the implementation of the image noise reduction method. The undisclosed details can be found in the implementation of the method part, and will not be repeated here.
[0162] The exemplary embodiments of the present disclosure further provide an electronic device. The electronic device may be the aforementioned terminal 101. Generally, the electronic device may include a processor and a memory, wherein the memory is configured to store executable instructions of the processor, and the processor is configured to execute the aforementioned image noise reduction method by executing the executable instructions.
[0163] Below is Figure 13 The structure of the electronic device is exemplified by taking the mobile terminal 1300 in FIG. 1 as an example. It should be understood by those skilled in the art that, in addition to the components specifically used for mobile purposes, Figure 13 The construction in can also be applied to fixed type equipment.
[0164] like Figure 13 As shown, the mobile terminal 1300 may specifically include: a processor 1301, a memory 1302, a bus 1303, a mobile communication module 1304, an antenna 1, a wireless communication module 1305, an antenna 2, a display screen 1306, a camera module 1307, an audio module 1308, a power module 1309 and a sensor module 1310.
[0165] The processor 1301 may include one or more processing units, for example, an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit). The image denoising method in this exemplary embodiment may be executed by an AP, a GPU, or a DSP. When the method involves neural network-related processing, it may be executed by an NPU. For example, the NPU may load neural network parameters and execute neural network-related algorithm instructions.
[0166] The encoder can encode (i.e., compress) an image or video to reduce the data size for easy storage or transmission. The decoder can decode (i.e., decompress) the encoded data of the image or video to restore the image or video data. The mobile terminal 1300 can support one or more encoders and decoders, such as: image formats such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap), and video formats such as MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0167] The processor 1301 may be connected to the memory 1302 or other components via a bus 1303 .
[0168] Memory 1302 can be used to store computer-executable program code, which includes instructions. Processor 1301 executes various functional applications and data processing of mobile terminal 1300 by running the instructions stored in memory 1302. Memory 1302 can also store application data, such as images, videos, and other files.
[0169] The communication functions of mobile terminal 1300 are implemented through mobile communication module 1304, antenna 1, wireless communication module 1305, antenna 2, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1304 can provide 3G, 4G, and 5G mobile communication solutions for mobile terminal 1300. Wireless communication module 1305 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1300.
[0170] The display screen 1306 is used to implement display functions, such as displaying user interfaces, images, videos, etc. The camera module 1307 is used to implement shooting functions, such as shooting images, videos, etc., and the camera module may include a color temperature sensor array. The audio module 1308 is used to implement audio functions, such as playing audio, collecting voice, etc. The power module 1309 is used to implement power management functions, such as charging the battery, powering the device, monitoring the battery status, etc. The sensor module 1310 may include one or more sensors for implementing corresponding sensing detection functions. For example, the sensor module 1310 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 1300 and output inertial sensing data.
[0171] It should be noted that a computer-readable storage medium is also provided in an embodiment of the present disclosure. The computer-readable storage medium may be included in the electronic device described in the above embodiment; or it may exist independently without being assembled into the electronic device.
[0172] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device.
[0173] Computer-readable storage media can transmit, propagate, or transfer programs for use by or in conjunction with an instruction execution system, apparatus, or device. Program code contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, or any suitable combination thereof.
[0174] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments.
[0175] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0176] Furthermore, the figures above are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the figures above do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0177] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0178] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing what is disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims. It should be understood that the present disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and that various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image denoising method, characterized in that: include: Acquire a time-domain denoised image of a current image and a reference frame image of the current image, and decompose the time-domain denoised image to obtain a multi-layer time-domain image; Performing spatial denoising on the current image to obtain a spatial denoised image corresponding to the current image, and decomposing the spatial denoised image to obtain a multi-layer spatial image; Matching each layer of temporal domain images with each layer of spatial domain images to determine an aligned image of the temporal domain denoised image, including: performing motion detection on the spatial domain denoised image and the aligned image to determine a difference map; determining a fusion weight map based on the difference map, and fusing the spatial domain denoised image and the aligned image according to the fusion weight map to obtain the denoised image; The spatial denoised image and the aligned image are fused to obtain a denoised image corresponding to the current image, including: performing motion detection on the spatial denoised image and the aligned image to determine a difference map; determining a fusion weight map based on the difference map, and fusing the spatial denoised image and the aligned image according to the fusion weight map to obtain the denoised image.
2. The image denoising method according to claim 1, wherein: Decomposing the time-domain denoised image to obtain a multi-layer time-domain image includes: The time-domain denoised image is down-sampled to obtain multi-layer time-domain images, where the resolutions of the multi-layer time-domain images are different.
3. The image denoising method according to claim 1, wherein: Performing multiple image alignments of different scales on each layer of temporal domain images and each layer of spatial domain images to determine the aligned images includes: According to the arrangement order of resolution, motion vector estimation is performed on each layer of spatial domain image in each layer of temporal domain image in turn, and motion compensation is performed according to the motion vector until multiple image alignments are performed to determine the aligned image.
4. The image denoising method according to claim 3, wherein: The step of sequentially estimating the motion vector of each layer of spatial domain image in each layer of temporal domain image and performing motion compensation according to the motion vector includes: The motion vector of the matching block of each layer of spatial domain image is estimated in each layer of temporal domain image in turn by adopting block matching, and motion compensation is performed according to the motion vector to determine the aligned image.
5. The image denoising method according to claim 4, wherein: The method of using block matching to sequentially estimate motion vectors of matching blocks of each layer of spatial domain image in each layer of temporal domain image, and performing motion compensation according to the motion vectors to determine the aligned image, includes: For the first matching block in each layer of spatial domain image, a motion vector is estimated by first block matching in each layer of temporal domain image to determine a motion vector of each layer of spatial domain image, and motion compensation is performed based on the motion vector to determine a first differential image; Based on the first differential image, sequentially performing motion vector estimation on the second matching block of each layer of the spatial domain image in each layer of the temporal domain image by a second block matching, determining a motion vector map of each layer of the spatial domain image, and performing motion compensation based on the motion vector map to determine a second differential image; Based on the second differential image, motion vector estimation is performed on the third matching block of each layer of spatial domain image in each layer of time domain image through a third block matching, a motion vector diagram of each layer of spatial domain image is determined, and motion compensation is performed according to the motion vector diagram to determine the aligned image.
6. The image denoising method according to claim 5, wherein: The method sequentially estimates the motion vector of the first matching block in each layer of the spatial domain image by first block matching in each layer of the temporal domain image, determines the motion vector of each layer of the spatial domain image, and performs motion compensation according to the motion vector to determine the first differential image, including: The spatial domain image of the current layer is used as the reference image of the current layer, and the temporal domain image of the current layer is used as the image to be aligned of the current layer; Estimating a motion vector for each first matching block in the current layer reference image in the image to be aligned, determining a motion vector of the current layer reference image, and determining a motion vector of each layer reference image; Perform motion compensation according to the motion vector diagram of the target layer reference image to determine a first compensated image; The first difference image is determined based on the first compensated image.
7. The image denoising method according to claim 6, wherein: The estimating a motion vector of each first matching block in the current layer reference image in the image to be aligned to determine a motion vector diagram of the current layer reference image includes: Determine a motion vector of each first matching block in the current layer reference image according to a similar block in the image to be aligned to each first matching block; The motion vector map is determined according to a map composed of motion vectors of each first matching block in the current layer reference image.
8. The image denoising method according to claim 7, wherein: The determining of the motion vector of each first matching block in the current layer reference image according to a similar block in the image to be aligned includes: Based on the search window of the current layer, searching for similar blocks of each first matching block in the image to be aligned; A motion vector of each first matching block in the image to be aligned is determined according to a difference between the coordinates of each first matching block and the coordinates of the similar block.
9. The image denoising method according to claim 8, wherein: The method further comprises: If the current layer reference image is a top layer reference image, determining a search window for the current layer according to a first preset window; If the current layer reference image is not a top layer reference image, the search window of the current layer is determined by combining the motion vector of the previous layer reference image of the current layer reference image, the coordinates of the first matching block and the second preset window.
10. The image denoising method according to claim 9, wherein: The determining the search window of the current layer reference image by combining the motion vector map of the previous layer reference image of the current layer reference image, the coordinates of the first matching block, and the second preset window includes: A center point is determined according to the motion vector of the first matching block in the previous layer reference image and the coordinates of the first matching block, and a search window of the current layer is determined based on the center point and a second preset window.
11. The image denoising method according to claim 6, wherein: The performing motion compensation according to the motion vector diagram of the target layer reference image to determine the first compensated image includes: Determining a motion vector of each first matching block in the target layer reference image based on a motion vector diagram of the target layer reference image; The compensation coordinates are determined according to the motion vector and coordinates of each first matching block, and the image blocks corresponding to the compensation coordinates are obtained from the image to be aligned in the target layer for motion compensation to obtain a first compensated image.
12. The image denoising method according to claim 6, wherein: The determining the first differential image based on the first compensated image includes: Perform differential processing on the first compensation image and the target layer reference image to obtain the first differential image.
13. The image denoising method according to claim 5, wherein: The method of estimating a motion vector of the second matching block of each layer of the spatial domain image in each layer of the temporal domain image by a second block matching based on the first differential image, determining a motion vector map of each layer of the spatial domain image, and performing motion compensation according to the motion vector map to determine a second differential image includes: determining a first non-aligned region of each layer of the time domain image according to the first differential image; performing a second block matching on the second matching blocks in each layer of the spatial domain image in the first non-aligned area in turn to determine a motion vector diagram of each layer of the spatial domain image in each layer of the temporal domain image; Perform motion compensation according to the motion vector diagram of the target layer spatial domain image to determine a second compensated image; Perform differential processing based on the second compensation image and the target layer spatial domain image to obtain the second differential image.
14. The image denoising method according to claim 5, wherein: The method of estimating the motion vector of the third matching block of each layer of the spatial domain image in each layer of the temporal domain image by a third block matching based on the second differential image, determining the motion vector of each layer of the spatial domain image, and performing motion compensation according to the motion vector to determine the aligned image includes: determining a second non-aligned region of each layer of the time domain image according to the second differential image; performing a third block matching in the second non-aligned area on the third matching block in each layer of the spatial domain image in turn to determine a motion vector diagram of each layer of the spatial domain image; Perform motion compensation according to the motion vector diagram of the target layer spatial domain image to determine a third compensated image; The third compensated image is determined as the alignment image.
15. The image denoising method according to claim 1, wherein: The performing motion detection on the spatial denoised image and the aligned image to determine a difference map includes: performing smoothing processing on the spatial domain denoised image and the aligned image respectively; The smoothed spatial denoised image and the smoothed aligned image are subtracted to obtain a difference image, and the difference image is subjected to morphological processing to obtain the difference image.
16. The image denoising method according to claim 3, wherein: The performing motion compensation according to the motion vector includes: Determine a compensation block for motion compensation according to a matching block in each layer of spatial domain image; The compensation coordinates are determined according to the motion vector of each matching block in the target layer spatial domain image, and motion compensation is performed on the compensation blocks in each layer of the spatial domain image according to the image blocks corresponding to the compensation coordinates.
17. The image denoising method according to claim 16, wherein: The determining of the compensation block for motion compensation according to the matching block in each layer of the spatial domain image includes: During the first block matching and the second block matching, the entire area of the matching block is used as a compensation block; During the third block matching, a partial area of the matching block is used as a compensation block, and the compensation block is updated according to the offset.
18. An image noise reduction device, characterized in that: include: a first image decomposition module, configured to obtain a time-domain denoised image of a current image and a reference frame image of the current image, and decompose the time-domain denoised image to obtain a multi-layer time-domain image; a second image decomposition module, configured to perform spatial denoising on the current image to obtain a spatial denoised image corresponding to the current image, and decompose the spatial denoised image to obtain a multi-layer spatial image; An image matching module is configured to match each layer of temporal domain images with each layer of spatial domain images to determine an aligned image of the temporal domain denoised image, including: performing motion detection on the spatial domain denoised image and the aligned image to determine a difference map; determining a fusion weight map based on the difference map, and fusing the spatial domain denoised image and the aligned image according to the fusion weight map to obtain the denoised image; An image fusion module is used to fuse the spatial denoised image and the aligned image to obtain a denoised image corresponding to the current image, including: performing motion detection on the spatial denoised image and the aligned image to determine a difference map; determining a fusion weight map based on the difference map, and fusing the spatial denoised image and the aligned image according to the fusion weight map to obtain the denoised image.
19. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the image denoising method according to any one of claims 1 to 17 by executing the executable instructions.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image denoising method according to any one of claims 1 to 17 is implemented.
Citation Information
Patent Citations
Noise reduction processing method and device and electronic equipment
CN113012061A