Time domain noise reduction method based on scene detection, scene detection method and storage medium
Through the time domain noise reduction method based on scene detection, appropriate processing is carried out for static and moving scenes, and the problems of storm and detail loss during noise reduction in motion scenes in the prior art are solved, achieving better noise reduction effect and detail retention.
Patent Information
- Application Number
- CN202311579764.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Existing time-domain noise reduction methods can easily lead to loss of shadowing and texture details when dealing with moving scenes.
The time domain noise reduction method based on scene detection is adopted to determine whether the pixel points of the current frame are in a stationary scene or a motion scene through scene detection, pre-filtering and weighting fusion are performed for pixel points of the stationary scene, and motion compensation and weighting fusion are performed for pixel points of the motion scene.
It effectively reduces the shadow-sweeping phenomenon in the moving scene, while retaining the details of the still scene, improving the noise reduction effect of the image.
Smart Images

Figure CN120070226A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to a time-domain noise reduction method based on scene detection, a scene detection method, and a computer-readable storage medium. Background Art
[0002] There are two traditional time-domain noise reduction methods. One is an adaptive time-domain noise reduction method, and the formula is as follows: fout(t)(x,y) = fin(t)(x,y)*(1 - w) + ref(t - 1)(x',y')*w, where fin(t)(x,y) is the pixel value of the pixel point (x,y) of the input frame at time t (i.e., the current frame), ref(t - 1)(x',y') is the pixel value of the pixel point (x',y') of the reference frame at time t, the pixel point (x,y) of the current frame corresponds to the pixel point (x',y') of the reference frame, (x,y) represents the coordinates of the pixel point of the current frame, (x',y') represents the coordinates of the pixel point of the reference frame. Here, the reference frame is the noise-reduced result of the input frame at time (t - 1), w is the fusion weight, and fout(t)(x,y) is the noise-reduced pixel value of the pixel point (x,y) of the current frame. This noise reduction method has a good noise reduction effect on the pixel points of the current frame in a static scene. However, for the pixel points of the current frame in a moving scene, due to the lack of motion compensation, after using this noise reduction method to perform time-domain noise reduction on the pixel points of the current frame in a moving scene, serious ghosting phenomena will occur in the current frame. The other is a time-domain noise reduction method based on motion estimation and motion compensation, and the formula is as follows: fout(t)(x,y) = fin(t)(x,y)*(1 - w) + ComRef(t - 1)(x',y')*w, where ComRef(t - 1)(x',y') is the pixel value of the pixel point (x',y') of the reference frame at time t after motion compensation. If the motion estimation is accurate, then the motion-compensated results of the pixel point (x,y) of the current frame and the pixel point (x',y') of the reference frame correspond accurately. In this case, this noise reduction method can perform time-domain noise reduction with a good effect on the pixel points of the current frame and can also better retain the texture details of the current frame. However, in reality, due to the influence of noise or the limitation of the search area, the motion estimation is often inaccurate, which also leads to the fact that the motion-compensated result of the pixel point (x',y') of the reference frame does not actually correspond to the pixel point (x,y) of the current frame. In this case, after using this noise reduction method to perform time-domain noise reduction on the pixel point (x,y) of the current frame, serious ghosting phenomena and loss of texture details will occur in the current frame.
[0003] Disclosure
[0004] Embodiments of the present disclosure propose a time-domain noise reduction method based on scene detection, a scene detection method, and a computer-readable storage medium, which can better reduce noise in the static scenes of an image and retain details to the greatest extent, and can also reduce the smear of the moving scenes of the image.
[0005] In a first aspect, embodiments of the present disclosure provide a time-domain noise reduction method based on scene detection, including: obtaining a current frame of an image and a reference frame corresponding to the current frame; performing scene detection to determine whether the pixel points of the current frame are in a static scene; determining that the pixel points of the current frame are in a static scene, then performing pre-filtering related to the current frame on the pixel points of the reference frame corresponding to the pixel points of the current frame to obtain pre-filtered pixel values of the pixel points of the reference frame; performing first weighted fusion on the original pixel values and the pre-filtered pixel values of the pixel points of the reference frame to obtain first weighted fusion pixel values of the pixel points of the reference frame; and performing second weighted fusion on the pixel values of the pixel points of the current frame and the first weighted fusion pixel values of the pixel points of the reference frame corresponding to the pixel points of the current frame to obtain time-domain noise reduction pixel values of the pixel points of the current frame.
[0006] In some embodiments, performing scene detection to determine whether the pixel points of the current frame are in a static scene includes: calculating a fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame according to the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame; and determining the motion state of the pixel points of the current frame according to the fusion frame difference, where the motion estimation neighborhood of the pixel points of the current frame includes a set of adjacent pixel points centered on the pixel points of the current frame.
[0007] In some embodiments, calculating the fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame includes: selecting a neighborhood range m×n centered on the pixel point (x, y) of the current frame as the motion estimation neighborhood, and calculating the sum of the absolute differences SAD between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference ASD between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame as follows,
[0008]
[0009]
[0010] (x, y) represents the coordinates of the pixel point in the current frame. fin(t)(x + j, y + i) represents the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) in the current frame at time t. The pixel point (x, y) in the current frame corresponds to the pixel point (x', y') in the reference frame. (x', y') represents the coordinates of the pixel point in the reference frame. ref(t - 1)(x' + j, y' + i) represents the pixel value of the pixel point (x', y') in the reference frame at time t. The neighborhood range m×n is a pixel array centered on the pixel point (x, y) in the current frame. One of m and n is the number of rows of the pixel array, and the other is the number of columns of the pixel array. Both m and n are odd numbers; and determine the fusion frame difference fusion_diff between the pixel point (x, y) in the current frame and the pixel point (x', y') in the reference frame. fusion_diff = SAD * α + ASD * (1 - α), where α is related to the noise intensity of the area where the pixel point (x, y) in the current frame is located. The greater the noise intensity of the area where the pixel point (x, y) in the current frame is located, the smaller α is. The smaller the noise intensity of the area where the pixel point (x, y) in the current frame is located, the greater α is. 0 ≤ α ≤ 1.
[0011] In some embodiments, determining the motion state of the pixel point in the current frame according to the fusion frame difference includes: calculating the motion intensity MotionStr of the pixel point (x, y) in the current frame, MotionStr = fusion_diff / NoiseStr, where NoiseStr is the noise intensity of the area where the pixel point (x, y) in the current frame is located; and if the motion intensity of the pixel point (x, y) in the current frame is not greater than a predetermined first motion threshold, determine that the pixel point (x, y) in the current frame is in a static scene. If the motion intensity of the pixel point (x, y) in the current frame is greater than the predetermined first motion threshold, determine that the pixel point (x, y) in the current frame is not in a static scene.
[0012] In some embodiments, performing scene detection to determine whether the pixel point in the current frame is in a static scene further includes: calculating the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point in the current frame and the pixel values of the corresponding pixel points in the reference frame as the frame difference between the pixel point in the current frame and the pixel point in the reference frame corresponding to the pixel point in the current frame; and determining the motion state of the pixel point in the current frame according to the frame difference.
[0013] In some embodiments, performing pre-filtering related to the current frame on the pixel points of the reference frame corresponding to the pixel points of the current frame includes: calculating the block differences between the pixel point (x, y) of the current frame and each pixel point in the search neighborhood sw×sh centered on the pixel point (x', y') of the reference frame, where the pixel point (x', y') of the reference frame corresponds to the pixel point (x, y) of the current frame, (x, y) represents the coordinates of the pixel point of the current frame, and (x', y') represents the coordinates of the pixel point of the reference frame; determining, based on the calculated block differences, the weight value w of each pixel point in the search neighborhood of the reference frame relative to the pixel point (x, y) of the current frame, where the weight value of each pixel point is a linear or non-linear function of its corresponding block difference, and the smaller the block difference corresponding to each pixel point, the larger its weight value; and performing the following pre-filtering on the pixel point (x', y') of the reference frame,
[0014]
[0015] ref(t - 1)(x', y') is the original pixel value of the pixel point (x', y') of the reference frame at time t, f1(ref(t - 1)(x', y')) is the pre-filtered pixel value of the pixel point (x', y') of the reference frame at time t, the search neighborhood sw×sh is a pixel array centered on the pixel point (x', y') of the reference frame, one of sw and sh is the number of rows of the pixel array, the other is the number of columns of the pixel array, both sw and sh are odd numbers, Prn is the original pixel value of each pixel point in the search neighborhood, wn is the weight value of each pixel point in the search neighborhood, n is the serial number of each pixel point in the search neighborhood, and n is a natural number, n∈sw*sh indicates that the pixel point with serial number n is within the search neighborhood sw×sh of the reference frame.
[0016] In some embodiments, calculating the block differences between the pixel point (x, y) of the current frame and each pixel point in the search neighborhood sw×sh centered on the pixel point (x', y') of the reference frame includes: selecting a block neighborhood pw×ph centered on the pixel point (x, y) of the current frame and a block neighborhood pw×ph centered on each pixel point in the search neighborhood of the reference frame; and calculating the sum of the absolute differences between the pixel values of each pixel point in the block neighborhood centered on the pixel point (x, y) of the current frame and the pixel values of the corresponding pixel points in each block neighborhood of the reference frame as the block difference between the pixel point (x, y) of the current frame and the pixel point serving as the center of each neighborhood block of the reference frame, where the block neighborhood pw×ph is a pixel array, one of pw and ph is the number of rows of the pixel array, the other is the number of columns of the pixel array, and both pw and ph are odd numbers.
[0017] In some embodiments, the first weighted fusion of the original pixel value and the pre-filtered pixel value of the pixel points of the reference frame includes: calculating the first weighted fusion pixel value of the pixel points of the reference frame according to the following formula,
[0018] ref_filtered(t-1)(x',y') = ref(t-1)(x',y')*(1-β)+f1(ref(t-1)(x',y'))*β, where ref_filtered(t-1)(x',y') is the first weighted fusion pixel value of the pixel point (x',y') of the reference frame at time t, β is a parameter related to the texture details of the image. The less the texture details of the image, the larger β is; the more the texture details of the image, the smaller β is, and 0≤β≤1.
[0019] In some embodiments, the second weighted fusion of the pixel value of the pixel points of the current frame and the first weighted fusion pixel value of the pixel points of the reference frame corresponding to the pixel points of the current frame includes: calculating the temporally denoised pixel value of the pixel points of the current frame according to the following formula,
[0020] fin_out(t)(x,y) = fin(t)(x,y)*(1-ρ)+ref_filtered(t-1)(x',y')*ρ, where fin_out(t)(x,y) is the temporally denoised pixel value of the pixel point (x,y) of the current frame at time t, fin(t)(x,y) is the original pixel value of the pixel point (x,y) of the current frame at time t, ρ is related to the motion intensity of the pixel point (x,y) of the current frame. The larger the motion intensity of the pixel point (x,y) of the current frame, the smaller ρ is; the smaller the motion intensity of the pixel point (x,y) of the current frame, the larger ρ is, and 0≤ρ≤1.
[0021] In some embodiments, the temporally denoising method further includes: determining that the pixel points of the current frame are not in a static scene, then determining that the pixel points of the current frame are in a motion scene; performing motion compensation on the pixel points of the reference frame corresponding to the pixel points of the current frame; and performing a third weighted fusion on the pixel value of the pixel points of the current frame and the pixel value of the pixel points of the reference frame corresponding to the pixel points of the current frame after motion compensation to obtain the temporally denoised pixel value of the pixel points of the current frame.
[0022] In some embodiments, performing motion compensation on the pixel points of the reference frame corresponding to the pixel points of the current frame includes: calculating the block difference between the pixel point (x, y) of the current frame and each pixel point in the search neighborhood sw×sh centered on the pixel point (x', y') of the reference frame, where the pixel point (x', y') of the reference frame corresponds to the pixel point (x, y) of the current frame, (x, y) represents the coordinates of the pixel point of the current frame, and (x', y') represents the coordinates of the pixel point of the reference frame; if the block difference between the pixel point (x, y) of the current frame and the pixel point (x1, y1) in the search neighborhood of the reference frame is the smallest, then determining the Euclidean distance between the pixel point (x, y) of the current frame and the pixel point (x1, y1) of the reference frame as the motion vector mv of the pixel point (x, y) of the current frame, mv = (x1 - x, y1 - y), where (x1, y1) represents the coordinates of the pixel point in the search neighborhood of the reference frame; and performing the following motion compensation on the pixel point (x', y') of the reference frame,
[0023] ComRef(t - 1)(x', y') = ref(t - 1)(x' + mv(0), y' + mv(1)), mv(0) = x1 - x, mv(1) = y1 - y, ComRef(t - 1)(x', y') is the pixel value after motion compensation of the pixel point (x', y') of the reference frame at time t, ref(t - 1)(x', y') is the original pixel value of the pixel point (x', y') of the reference frame at time t, the search neighborhood sw×sh is a pixel array centered on the pixel point (x', y') of the reference frame, one of sw and sh is the number of rows of the pixel array, the other is the number of columns of the pixel array, and both sw and sh are odd numbers.
[0024] In some embodiments, performing third weighted fusion on the pixel value of the pixel point of the current frame and the pixel value after motion compensation of the pixel point of the reference frame corresponding to the pixel point of the current frame includes:
[0025] Calculating fin_out(t)(x, y) = fin(t)(x, y) * (1 - γ) + ComRef(t - 1)(x', y') * γ, where fin_out(t)(x, y) is the pixel value after temporal noise reduction of the pixel point (x, y) of the current frame at time t, fin(t)(x, y) is the original pixel value of the pixel point (x, y) of the current frame at time t, γ is related to the motion intensity of the pixel point (x, y) of the current frame, the greater the motion intensity of the pixel point (x, y) of the current frame, the smaller γ is, the smaller the motion intensity of the pixel point (x, y) of the current frame, the greater γ is, and 0 ≤ γ ≤ 1.
[0026] In some embodiments, the time-domain noise reduction method further includes: after determining that the pixel points of the current frame are not in a static scene, determining whether the pixel points of the current frame are in a violently moving scene; and if it is determined that the pixel points of the current frame are in a violently moving scene, performing a fourth weighted fusion on the pixel values of the pixel points of the current frame and the pixel values of the pixel points of the reference frame corresponding to the pixel points of the current frame to obtain the pixel values of the pixel points of the current frame after time-domain noise reduction.
[0027] In some embodiments, performing a fourth weighted fusion on the pixel values of the pixel points of the current frame and the pixel values of the pixel points of the reference frame corresponding to the pixel points of the current frame includes: calculating fin_out(t)(x,y) = fin(t)(x,y)*(1 - ω)+ref(t - 1)(x',y')*ω, where fin(t)(x,y) is the original pixel value of the pixel point (x,y) of the current frame at time t, ref(t - 1)(x',y') is the original pixel value of the pixel point (x',y') of the reference frame at time t, fin_out(t)(x,y) is the pixel value of the pixel point (x,y) of the current frame after time-domain noise reduction, the pixel point (x',y') of the reference frame corresponds to the pixel point (x,y) of the current frame, (x,y) represents the coordinates of the pixel point of the current frame, (x',y') represents the coordinates of the pixel point of the reference frame, ω is related to the motion intensity of the pixel point (x,y) of the current frame, the greater the motion intensity of the pixel point (x,y) of the current frame, the smaller ω is, the smaller the motion intensity of the pixel point (x,y) of the current frame, the greater ω is, and 0 ≤ ω ≤ 1.
[0028] In a second aspect, an embodiment of the present disclosure provides a scene detection method, including: obtaining a current frame of an image and a reference frame corresponding to the current frame; calculating a fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame according to the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame; and determining whether the pixel points of the current frame are in a static scene or a moving scene according to the fusion frame difference.
[0029] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and the computer program enables a processor to implement the method provided by the above embodiments of the present disclosure.
[0030] In the time-domain noise reduction method based on scene detection provided by the embodiments of the present disclosure, scene detection is first performed to determine whether the pixel points of the current frame serving as the input frame are in a static scene. For the pixel points of the current frame in the static scene, pre-filtering related to the current frame is performed on the pixel points of the reference frame corresponding to the pixel points of the current frame to obtain the pre-filtered pixel values of the pixel points of the reference frame. Then, first weighted fusion is performed on the original pixel values and the pre-filtered pixel values of the pixel points of the reference frame to obtain the first weighted fusion pixel values of the pixel points of the reference frame. Next, second weighted fusion is performed on the pixel values of the pixel points of the current frame and the first weighted fusion pixel values of the pixel points of the reference frame corresponding to the pixel points of the current frame to obtain the time-domain noise-reduced pixel values of the pixel points of the current frame. By performing pre-filtering related to the current frame on the pixel points of the reference frame corresponding to the pixel points of the current frame, the time-domain noise reduction of the pixel points of the current frame can be made more effective. By performing first weighted fusion on the original pixel values and the pre-filtered pixel values of the pixel points of the reference frame, the texture details of the current frame after time-domain noise reduction can be retained as much as possible. By performing second weighted fusion on the pixel values of the pixel points of the current frame and the first weighted fusion pixel values of the pixel points of the reference frame corresponding to the pixel points of the current frame, the noise reduction effect and the retention of texture details of the current frame after time-domain noise reduction can be effectively balanced.
[0031] Further, in the time-domain noise reduction method based on scene detection provided by the embodiments of the present disclosure, by performing different noise reduction processes on the pixel points of the current frame in the static scene and the pixel points in the moving scene, the noise reduction effect on the pixel points of the current frame in each scene can be greatly improved, the noise reduction effect on the pixel points in the static scene and the retention of texture details can be better balanced, and the occurrence of ghosting can be reduced while ensuring the noise reduction effect on the pixel points in the moving scene.
[0032] Further, in the time-domain noise reduction method based on scene detection provided by the embodiments of the present disclosure, for the pixel points of the current frame in the violent motion scene, motion compensation is not performed on the pixel points of the reference frame corresponding to the pixel points of the current frame, but directly weighted fusion is performed on the pixel values of the pixel points of the current frame and the pixel values of the pixel points of the reference frame corresponding to the pixel points of the current frame to obtain the time-domain noise-reduced pixel values of the pixel points of the current frame, avoiding inappropriate motion compensation for the pixel points of the reference frame corresponding to the pixel points of the current frame in the violent motion scene due to inaccurate motion offset estimation of the pixel points of the current frame in the violent motion scene, and also avoiding resource shortage caused by a large amount of calculation for motion offset estimation of the pixel points of the current frame in the violent motion scene.
[0033] In addition, in the scene detection method provided by the embodiments of the present disclosure, the motion state of the pixel points of the current frame can be determined by referring to the scene noise (i.e., the noise intensity of the area where the pixel points of the current frame are located), so as to determine whether the pixel points of the current frame are in a static scene or a moving scene. When the scene noise is small, the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame can be used as the frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame, and the motion state of the pixel points of the current frame can be determined according to the frame difference. As the scene noise increases, the fusion of the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame can be used as the fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame, and the motion state of the pixel points of the current frame can be determined according to the fusion frame difference. When the scene noise is large, the noise distribution may fluctuate. In this case, the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame may be large. However, a large sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame does not necessarily mean that the pixel points of the current frame are in a moving scene. Therefore, when the scene noise is large, it may be inaccurate to determine the motion state of the pixel points of the current frame according to the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame. In the scene detection method provided by the embodiments of the present disclosure, when the scene noise is large, the fusion of the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame can be used as the fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame, and the motion state of the pixel points of the current frame can be determined according to the fusion frame difference. By referring to the absolute difference between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame, the influence of partial noise distribution fluctuations can be removed, so the accuracy of scene detection can be improved. Description of the Drawings
[0034] Figure 1 It is a flowchart of a time-domain noise reduction method based on scene detection according to an embodiment of the present disclosure.
[0035] Figure 2Schematic diagram of the scene detection instance according to an embodiment of the present disclosure.
[0036] Figure 3 Schematic diagram of the block difference calculation according to an embodiment of the present disclosure.
[0037] Figure 4 Flowchart of the time-domain noise reduction method based on scene detection according to an embodiment of the present disclosure.
[0038] Figure 5 Flowchart of the time-domain noise reduction method based on scene detection according to an embodiment of the present disclosure.
[0039] Figure 6 An example of the time-domain noise reduction method based on scene detection according to an embodiment of the present disclosure.
[0040] Figure 7 An example of the time-domain noise reduction method based on scene detection according to an embodiment of the present disclosure.
[0041] Figure 8 Flowchart of the scene detection method according to an embodiment of the present disclosure.
[0042] Figure 9 Flowchart of the scene detection method according to an embodiment of the present disclosure.
[0043] Figure 10 Schematic diagram of the computer-readable storage medium according to an embodiment of the present disclosure. Detailed implementation manners
[0044] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the time-domain noise reduction method based on scene detection, the corresponding scene detection method, and the computer-readable storage medium provided by the present disclosure will be described in detail below with reference to the accompanying drawings.
[0045] In the following, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments may be embodied in different forms and the present disclosure should not be construed as limited to the embodiments set forth herein. The purpose of providing these embodiments is to make the present disclosure more thorough and complete, and to enable those skilled in the art to fully understand the scope of the present disclosure.
[0046] Without conflict, the various embodiments of the present disclosure and the features in the embodiments may be combined with each other. As used herein, terms such as "at least one", "and / or" include any and all combinations of one or more of the associated listed items.
[0047] It should be understood that the terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. Although the terms "first", "second", etc. are used herein to describe elements (e.g., weighted fusion), these elements are not limited to these terms, and these terms do not denote any order, quantity, or importance, but are only used for distinction. For example, without departing from the scope of the present disclosure, the first weighted fusion may also be referred to as the second or third weighted fusion, and similarly, the second weighted fusion may also be referred to as the first or third weighted fusion, and the third weighted fusion may also be referred to as the first or second weighted fusion.
[0048] In addition, as used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. Depending on the context, the terms "if", "in response to...", etc. used herein may be interpreted as "during...", or "when...". It will also be understood that the term "comprising" used in this specification specifies the presence of a particular feature, entity, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, entities, steps, operations, elements, components, and / or combinations thereof.
[0049] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0050] In a first aspect, an embodiment of the present disclosure provides a time-domain noise reduction method based on scene detection, as Figure 1 shown, including steps S1 to S5.
[0051] Step S1, obtain the current frame of the image and the reference frame corresponding to the current frame.
[0052] In an embodiment of the present disclosure, the input frame fin(t) at time t is used as the current frame, and the noise-reduced result ref(t - 1) of the input frame at time (t - 1) is used as the reference frame corresponding to the current frame fin(t).
[0053] Step S2, perform scene detection to determine whether the pixel points of the current frame are in a static scene.
[0054] In the embodiments of the present disclosure, the sum of the absolute differences between the pixel values of each pixel in the motion estimation neighborhood of the pixel of the current frame and the pixel values of the corresponding pixels in the reference frame can be calculated as the frame difference between the pixel of the current frame and the pixel of the reference frame corresponding to the pixel of the current frame. Then, the motion intensity of the pixel of the current frame can be determined according to the frame difference, so as to determine the motion state of the pixel of the current frame according to the motion intensity of the pixel of the current frame. If the motion intensity of the pixel of the current frame is not greater than a predetermined first motion threshold, it can be determined that the pixel of the current frame is in a static scene. If the motion intensity of the pixel of the current frame is greater than the predetermined first motion threshold, it can be determined that the pixel of the current frame is not in a static scene.
[0055] To improve the accuracy of scene detection, the sum of the absolute differences between the pixel values of each pixel in the motion estimation neighborhood of the pixel of the current frame and the pixel values of the corresponding pixels in the reference frame, and the absolute difference between the sum of the pixel values of each pixel in the motion estimation neighborhood of the pixel of the current frame and the sum of the pixel values of the corresponding pixels in the reference frame can also be used to calculate the fused frame difference between the pixel of the current frame and the pixel of the reference frame corresponding to the pixel of the current frame. Then, the motion intensity of the pixel of the current frame can be calculated according to the fused frame difference, so as to determine the motion state of the pixel of the current frame according to the motion intensity of the pixel of the current frame. Similarly, if the motion intensity of the pixel of the current frame is not greater than a predetermined first motion threshold, it can be determined that the pixel of the current frame is in a static scene. If the motion intensity of the pixel of the current frame is greater than the predetermined first motion threshold, it can be determined that the pixel of the current frame is not in a static scene.
[0056] In the embodiments of the present disclosure, the motion estimation neighborhood of the pixel of the current frame includes a set of adjacent pixels centered on the pixel of the current frame.
[0057] In some embodiments, if it is determined that the pixel of the current frame is not in a static scene, that is, the motion intensity of the pixel of the current frame is greater than a predetermined first motion threshold, it can be determined that the pixel of the current frame is in a motion area.
[0058] Of course, a second motion threshold greater than the first motion threshold can also be preset. If the motion intensity of the pixel of the current frame is not only greater than the preset first motion threshold but also greater than the preset second motion threshold, it can be determined that the pixel of the current frame is in a violent motion scene.
[0059] Both the above-mentioned first motion threshold and second motion threshold can be preset and / or adjusted according to actual situations such as practical experience, expectations, resource limitations, etc. For example, they can be preset and / or adjusted according to the desired overall noise reduction effect on the current frame, the desired degree of preservation of texture details of the current frame, the acceptable cost of resource consumption, etc. The embodiments of the present disclosure do not make specific limitations.
[0060] It should be understood that the pixel value of the pixel point in the embodiment of the present disclosure can be a grayscale value.
[0061] Figure 2 A schematic diagram showing a scene detection example of the embodiment of the present disclosure is shown.
[0062] As Figure 2 shown, the sum of absolute differences SAD between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference ASD between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel point of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame can be calculated based on the current frame fin(t) and the reference frame ref(t - 1) corresponding to the current frame.
[0063] If the noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located is small, SAD can better reflect the motion state of the pixel point of the current frame. Therefore, the motion state of the pixel point of the current frame can be determined based on SAD.
[0064] As the noise intensity (i.e., as the scene noise) in the region where the pixel point of the current frame is located increases, the noise distribution may fluctuate. In this case, SAD may be large. However, a large SAD does not necessarily mean that the pixel point of the current frame is in a motion scene. Therefore, if the noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located is large, the motion state of the pixel point of the current frame can be determined based on the fusion of SAD and ASD. By referring to ASD, the influence of partial noise distribution fluctuations can be removed, and the accuracy of scene detection can be improved.
[0065] It should be understood that if the noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located is large to a certain extent, the motion state of the pixel point of the current frame can also be determined only based on ASD. That is, when the noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located is large to a certain extent, for the motion state of the pixel point of the current frame, SAD may have lost its reference value.
[0066] In some embodiments, a neighborhood range m×n centered on the pixel point (x, y) of the current frame can be selected as the motion estimation neighborhood. The sum of absolute differences SAD between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame and the pixel values of the corresponding pixel points in the reference frame is calculated according to the following formula (1). The absolute difference ASD between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame and the sum of the pixel values of the corresponding pixel points in the reference frame is calculated according to the following formula (2).
[0067]
[0068]
[0069] In the above formulas (1) and (2), (x, y) represents the coordinates of the pixel point of the current frame, fin(t)(x + j, y + i) represents the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame at time t. The pixel point (x, y) of the current frame corresponds to the pixel point (x', y') of the reference frame. (x', y') represents the coordinates of the pixel point of the reference frame, and ref(t - 1)(x' + j, y' + i) represents the pixel value of the pixel point (x', y') of the reference frame at time t. The neighborhood range m×n is a pixel array centered on the pixel point (x, y) of the current frame. One of m and n is the number of rows of the pixel array, and the other is the number of columns of the pixel array. Both m and n are odd numbers.
[0070] It should be understood that the size of the above neighborhood range m×n can be selected according to actual situations such as practical experience, expectations, resource limitations, etc., and is not specifically limited in the embodiments of the present disclosure.
[0071] Then, the motion state of the pixel point (x, y) of the current frame can be determined according to the calculated SAD, or according to the fusion of the calculated SAD and ASD.
[0072] In some embodiments, the fusion frame difference fusion_diff between the pixel point (x, y) of the current frame and the pixel point (x', y') of the reference frame can be determined according to the following formula (3), so as to determine the motion state of the pixel point (x, y) of the current frame.
[0073] fusion_diff = SAD*α+ASD*(1-α) (3)
[0074] In the above formula (3), α is related to the noise intensity of the area where the pixel point (x, y) of the current frame is located. The greater the noise intensity of the area where the pixel point (x, y) of the current frame is located, the smaller α is. The smaller the noise intensity of the area where the pixel point (x, y) of the current frame is located, the greater α is. 0≤α≤1.
[0075] That is to say, α can be adaptively adjusted according to the noise intensity in the region where the pixel point (x, y) of the current frame is located, so that based on the noise intensity in the region where the pixel point (x, y) of the current frame is located (i.e., scene noise), through the frame difference or fused frame difference between the pixel point (x, y) of the current frame and the pixel point (x', y') of the reference frame, the motion state of the pixel point (x, y) of the current frame can be more accurately reflected.
[0076] Obviously, when α = 1, SAD is used as the frame difference between the pixel point (x, y) of the current frame and the pixel point (x', y') of the reference frame to reflect the motion state of the pixel point (x, y) of the current frame. When α ≠ 1, the fusion of SAD and ASD is used as the fused frame difference between the pixel point (x, y) of the current frame and the pixel point (x', y') of the reference frame to reflect the motion state of the pixel point (x, y) of the current frame. When α = 0, only ASD is used to reflect the motion state of the pixel point (x, y) of the current frame.
[0077] In practice, α can also be a fixed value. For example, it can be preset with reference to practical experience according to scene noise, etc.
[0078] The embodiments of the present disclosure do not specifically limit how to determine and / or calculate the noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located. The noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located can be determined and / or calculated according to the methods in the relevant technical fields.
[0079] Moreover, the embodiments of the present disclosure do not limit the specific numerical value or numerical range of "larger" or "smaller" of the noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located. In some embodiments, if the noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located is not greater than a predetermined threshold, it is determined that the noise intensity in the region where the pixel point of the current frame is located is small. If the noise intensity (i.e., scene noise) in the region where the pixel point of the current frame is located is greater than the predetermined threshold, it is determined that the noise intensity in the region where the pixel point of the current frame is located is large. Those skilled in the art can preset and / or adjust the predetermined threshold according to actual situations such as practical experience, expectations, resource limitations, etc. For example, the predetermined threshold can be preset and / or adjusted according to the expected overall noise reduction effect on the current frame, the acceptable resource consumption cost, etc.
[0080] In some embodiments, the normalized motion intensity of the pixel point of the current frame relative to the noise intensity in its region can be calculated, and then the motion state of the pixel point of the current frame, that is, being in a static scene or a motion scene, can be determined according to the calculated motion intensity.
[0081] For example, the motion intensity MotionStr of the pixel point (x, y) in the current frame can be calculated as MotionStr = fusion_diff / NoiseStr, where NoiseStr is the noise intensity of the region where the pixel point (x, y) in the current frame is located. If the motion intensity of the pixel point (x, y) in the current frame is not greater than a predetermined first motion threshold LowMotionThreshold, it is determined that the pixel point (x, y) in the current frame is in a static scene. If the motion intensity of the pixel point (x, y) in the current frame is greater than the predetermined first motion threshold LowMotionThreshold, it is determined that the pixel point (x, y) in the current frame is not in a static scene. In addition, if the motion intensity of the pixel point (x, y) in the current frame is not only greater than the predetermined first motion threshold LowMotionThreshold but also greater than a preset second motion threshold HighMotionThreshold, where LowMotionThreshold ≤ HighMotionThreshold, it can be further determined that the pixel point (x, y) in the current frame is in a violent motion scene.
[0082] In step S3, if it is determined that the pixel point in the current frame is in a static scene, pre-filtering related to the current frame is performed on the pixel point of the reference frame corresponding to the pixel point in the current frame to obtain the pre-filtered pixel value of the pixel point of the reference frame.
[0083] Performing pre-filtering related to the current frame on the pixel point of the reference frame corresponding to the pixel point in the current frame means performing weighted filtering based on the pixel point in the current frame on the pixel point of the reference frame corresponding to the pixel point in the current frame. The filtering weight for the pixel point of the reference frame is strongly correlated with the pixel point in the current frame, so that the pixel point of the reference frame is closer to the corresponding pixel point in the current frame, and the noise reduction effect for the pixel point in the current frame can be improved.
[0084] In some embodiments, performing pre-filtering related to the current frame on the pixel point of the reference frame corresponding to the pixel point in the current frame includes: calculating the block difference between the pixel point (x, y) in the current frame and each pixel point in the search neighborhood sw × sh centered on the pixel point (x', y') in the reference frame. The pixel point (x', y') in the reference frame corresponds to the pixel point (x, y) in the current frame, (x, y) represents the coordinates of the pixel point in the current frame, and (x', y') represents the coordinates of the pixel point in the reference frame; determining the weight value w of each pixel point in the search neighborhood of the reference frame relative to the pixel point (x, y) in the current frame according to the calculated block difference. The weight value of each pixel point is a linear or non-linear function of its corresponding block difference, and the smaller the block difference corresponding to each pixel point, the larger its weight value; and pre-filtering the pixel point (x', y') in the reference frame according to the following formula (4).
[0085]
[0086] In the above formula (4), ref(t - 1)(x', y') is the original pixel value of the pixel point (x', y') in the reference frame at time t, f1(ref(t - 1)(x', y')) is the pre-filtered pixel value of the pixel point (x', y') in the reference frame at time t, the search neighborhood sw × sh is a pixel array centered on the pixel point (x', y') in the reference frame, one of sw and sh is the number of rows of the pixel array, the other is the number of columns of the pixel array, both sw and sh are odd numbers, Prn is the original pixel value of each pixel point in the search neighborhood, wn is the weight value of each pixel point in the search area, n is the serial number of each pixel point in the search area, which is a natural number, and n ∈ sw * sh means that the pixel point with serial number n is within the search neighborhood sw × sh of the reference frame.
[0087] In some embodiments, calculating the block difference between the pixel point (x, y) of the current frame and each pixel point in the search neighborhood sw × sh centered on the pixel point (x', y') of the reference frame includes: selecting the block neighborhood pw × ph centered on the pixel point (x, y) of the current frame and the block neighborhood pw × ph centered on each pixel point in the search neighborhood of the reference frame; and calculating the sum of absolute differences SAD between the pixel values of each pixel point in the block neighborhood centered on the pixel point (x, y) of the current frame and the pixel values of the corresponding pixel points in each block neighborhood of the reference frame, as the block difference between the pixel point (x, y) of the current frame and the pixel point serving as the center of each neighborhood block of the reference frame. The block neighborhood pw × ph is a pixel array, one of pw and ph is the number of rows of the pixel array, the other is the number of columns of the pixel array, and both pw and ph are odd numbers.
[0088] In this case, for example, the weight value w of each pixel point in the search neighborhood sw × sh of the reference frame relative to the pixel point (x, y) of the current frame is w = f(SAD), and f can be a linear transformation or a non-linear transformation.
[0089] In the embodiments of the present disclosure, the definition of SAD as the block difference is consistent with the definition of SAD in the above scene detection.
[0090] It should be understood that the sizes of the above search neighborhood sw × sh and block neighborhood pw × ph can be selected according to actual situations such as practical experience, expectations, resource limitations, etc., and are not specifically limited in the embodiments of the present disclosure.
[0091] For example, the size of the search neighborhood sw × sh can be larger than the size of the block neighborhood pw × ph.
[0092] Figure 3Shows a schematic diagram of block difference calculation according to an embodiment of the present disclosure.
[0093] As Figure 3 shown, the block difference (e.g., SAD) between the pixel point (x, y) of the current frame and the pixel point (x1, y1) of the reference frame is small. Therefore, the weight value w1 of the pixel point (x1, y1) of the reference frame relative to the pixel point (x, y) of the current frame is large. The block difference (e.g., SAD) between the pixel point (x, y) of the current frame and the pixel point (x2, y2) of the reference frame is large. Therefore, the weight value w2 of the pixel point (x2, y2) of the reference frame relative to the pixel point (x, y) of the current frame is small.
[0094] Performing pre-filtering related to the current frame on the pixel point of the reference frame corresponding to the pixel point of the current frame is beneficial to improving the noise reduction effect. However, it is not beneficial to retain the texture details of the image.
[0095] Step S4: Perform first weighted fusion on the original pixel value and the pre-filtered pixel value of the pixel point of the reference frame to obtain the first weighted fusion pixel value of the pixel point of the reference frame.
[0096] Performing first weighted fusion on the original pixel value and the pre-filtered pixel value of the pixel point of the reference frame is beneficial to retaining the texture details of the image.
[0097] In some embodiments, performing first weighted fusion on the original pixel value and the pre-filtered pixel value of the pixel point of the reference frame includes: calculating the first weighted fusion pixel value of the pixel point of the reference frame according to the following formula (4).
[0098] ref_filtered(t - 1)(x', y') = ref(t - 1)(x', y') * (1 - β) + f1(ref(t - 1)(x', y')) * β (5)
[0099] In the above formula (5), ref_filtered(t - 1)(x', y') is the first weighted fusion pixel value of the pixel point (x', y') of the reference frame at time t, β is a parameter related to the texture details of the image. The less the texture details of the image, the larger β is, and the first weighted fusion pixel value of the pixel point (x', y') of the reference frame tends to be the pre-filtered pixel value of the pixel point (x', y') of the reference frame. On the contrary, the more the texture details of the image, the smaller β is, and the first weighted fusion pixel value of the pixel point (x', y') of the reference frame tends to be the original pixel value of the pixel point (x', y') of the reference frame, 0 ≤ β ≤ 1. Therefore, performing first weighted fusion on the original pixel value and the pre-filtered pixel value of the pixel point of the reference frame can better balance the noise reduction effect and the retention of texture details.
[0100] It should be understood that β can be adaptively adjusted according to the specific texture details of the image. β = 0 and β = 1 correspond to extreme cases. That is, in the case of β = 0, the pixel points of the reference frame may not be pre-filtered, while in the case of β = 1, the pixel values after pre-filtering of the pixel points of the reference frame will be separately used for time-domain noise reduction of the pixel points of the current frame. In practice, β can also be a fixed value. For example, it can be preset according to the image type, etc., referring to practical experience, and the embodiments of the present disclosure do not make specific limitations.
[0101] Step S5: Perform a second weighted fusion on the pixel value of the pixel point of the current frame and the pixel value of the pixel point of the reference frame corresponding to the pixel point of the current frame after the first weighted fusion, to obtain the pixel value of the pixel point of the current frame after time-domain noise reduction.
[0102] In some embodiments, performing a second weighted fusion on the pixel value of the pixel point of the current frame and the pixel value of the pixel point of the reference frame corresponding to the pixel point of the current frame after the first weighted fusion includes: calculating the pixel value of the pixel point of the current frame after time-domain noise reduction according to the following formula (6).
[0103] fin_out(t)(x,y) = fin(t)(x,y) * (1 - ρ) + ref_filtered(t - 1)(x',y') * ρ (6)
[0104] In the above formula (6), fin_out(t)(x,y) is the pixel value of the pixel point (x,y) of the current frame at time t after time-domain noise reduction, fin(t)(x,y) is the original pixel value of the pixel point (x,y) of the current frame at time t, ρ is related to the motion intensity of the pixel point (x,y) of the current frame. The greater the motion intensity of the pixel point (x,y) of the current frame, the smaller ρ is, and the pixel value of the pixel point (x,y) of the current frame after time-domain noise reduction tends to be closer to the original pixel value of the pixel point (x,y) of the current frame, that is, the degree of noise reduction for the pixel point (x,y) of the current frame is relatively small. On the contrary, the smaller the motion intensity of the pixel point (x,y) of the current frame, the larger ρ is, and the pixel value of the pixel point (x,y) of the current frame after time-domain noise reduction tends to be closer to the pixel value of the pixel point (x',y') of the reference frame after the first weighted fusion, that is, the degree of noise reduction for the pixel point (x,y) of the current frame is relatively large, and 0 ≤ ρ ≤ 1.
[0105] It should be understood that ρ can be adaptively adjusted according to the motion intensity of the pixel point (x, y) of the current frame. ρ = 0 and ρ = 1 correspond to extreme cases. That is, in the case of ρ = 0, time-domain noise reduction is substantially not performed on the pixel points of the current frame, while in the case of ρ = 1, the original pixel value of the pixel point of the reference frame, or the fusion of the original pixel value of the pixel point of the reference frame and the pre-filtered pixel value, is directly used as the pixel value after time-domain noise reduction of the pixel point of the current frame. In practice, ρ can also be a fixed value. For example, it can be preset with reference to practical experience according to the image type, etc. The embodiments of the present disclosure do not make specific limitations.
[0106] The time-domain noise reduction of the pixel points of the current frame in the static scene has been described in detail above. However, it is possible to determine in step S2 above that the pixel points of the current frame are not in the static scene. In this case, as Figure 4 shown, the time-domain noise reduction method of the embodiments of the present disclosure may further include the following steps S3' to S5'.
[0107] Step S3', if it is determined that the pixel points of the current frame are not in the static scene, then it is determined that the pixel points of the current frame are in the motion scene.
[0108] That is to say, in the above scene detection, if the motion intensity of the pixel point (x, y) of the current frame is greater than a predetermined first motion threshold (for example, LowMotionThreshold), it can be determined that the pixel point (x, y) of the current frame is in the motion scene.
[0109] Step S4', perform motion compensation on the pixel points of the reference frame corresponding to the pixel points of the current frame.
[0110] In some embodiments, performing motion compensation on the pixel points of the reference frame corresponding to the pixel points of the current frame includes: calculating the block difference between the pixel point (x, y) of the current frame and each pixel point in the search neighborhood sw×sh centered on the pixel point (x', y') of the reference frame. The pixel point (x', y') of the reference frame corresponds to the pixel point (x, y) of the current frame. (x, y) represents the coordinates of the pixel point of the current frame, and (x', y') represents the coordinates of the pixel point of the reference frame. The minimum block difference between the pixel point (x, y) of the current frame and the pixel point (x1, y1) in the search neighborhood of the reference frame means that the pixel point (x, y) of the current frame is closest to the pixel point (x1, y1) in the search neighborhood of the reference frame. Then, the Euclidean distance between the pixel point (x, y) of the current frame and the pixel point (x1, y1) of the reference frame is determined as the motion vector mv of the pixel point (x, y) of the current frame, mv = (x1 - x, y1 - y), and (x1, y1) represents the coordinates of the pixel point in the search neighborhood of the reference frame. And perform motion compensation on the pixel point (x', y') of the reference frame according to the following formula (7).
[0111] ComRef(t - 1)(x', y') = ref(t - 1)(x' + mv(0), y' + mv(1)) (7)
[0112] In the above formula (7), mv(0) = x1 - x, mv(1) = y1 - y. ComRef(t - 1)(x', y') is the pixel value after motion compensation of the pixel point (x', y') in the reference frame at time t, ref(t - 1)(x', y') is the original pixel value of the pixel point (x', y') in the reference frame at time t. The search neighborhood sw × sh is a pixel array centered on the pixel point (x', y') of the reference frame. One of sw and sh is the number of rows of the pixel array, and the other is the number of columns of the pixel array. Both sw and sh are odd numbers.
[0113] It should be understood that the size of the search neighborhood sw × sh selected for motion compensation here can be the same as or different from the size of the search neighborhood sw × sh selected when pre-filtering the pixel points of the reference frame. The embodiments of the present disclosure do not make specific limitations, and those skilled in the art can select according to needs.
[0114] Step S5', perform third weighted fusion on the pixel value of the pixel point of the current frame and the pixel value after motion compensation of the pixel point of the reference frame corresponding to the pixel point of the current frame, to obtain the pixel value after temporal noise reduction of the pixel point of the current frame.
[0115] In some embodiments, performing third weighted fusion on the pixel value of the pixel point of the current frame and the pixel value after motion compensation of the pixel point of the reference frame corresponding to the pixel point of the current frame includes: calculating the pixel value after temporal noise reduction of the pixel point (x, y) of the current frame at time t according to the following formula (8).
[0116] fin_out(t)(x, y) = fin(t)(x, y) * (1 - γ) + ComRef(t - 1)(x', y') * γ (8)
[0117] In the above formula (8), fin_out(t)(x, y) is the pixel value after temporal noise reduction of the pixel point (x, y) of the current frame at time t, fin(t)(x, y) is the original pixel value of the pixel point (x, y) of the current frame at time t. γ is related to the motion intensity of the pixel point (x, y) of the current frame. The greater the motion intensity of the pixel point (x, y) of the current frame, the smaller γ is; the smaller the motion intensity of the pixel point (x, y) of the current frame, the greater γ is. 0 ≤ γ ≤ 1.
[0118] It should be understood that γ can be adaptively adjusted according to the motion intensity of the pixel point (x, y) of the current frame. γ = 0 and γ = 1 correspond to extreme cases. That is, in the case of γ = 0, time-domain noise reduction is not substantially performed on the pixel points of the current frame, while in the case of γ = 1, the pixel value after motion compensation of the pixel point of the reference frame is directly used as the pixel value after time-domain noise reduction of the pixel point of the current frame. In practice, γ can also be a fixed value. For example, it can be preset with reference to practical experience according to the image type, etc., and the embodiments of the present disclosure do not make specific limitations.
[0119] By performing third weighted fusion on the pixel value of the pixel point of the current frame and the pixel value after motion compensation of the pixel point of the reference frame corresponding to the pixel point of the current frame, the pixel value after time-domain noise reduction of the pixel point of the current frame is obtained, which can reduce the occurrence of smear phenomenon while ensuring the noise reduction effect on the pixel points in the motion scene.
[0120] In addition, if it is determined in step S2 above that the pixel point of the current frame is not in a static scene, that is, the motion intensity of the pixel point of the current frame is greater than a predetermined first motion threshold, and the motion intensity of the pixel point of the current frame is also greater than a preset second motion threshold, and the preset second motion threshold is greater than the preset first motion threshold. In this case, as Figure 5 shown, the time-domain noise reduction method of the embodiments of the present disclosure may further include steps S3” to S4”.
[0121] Step S3”, after determining that the pixel point of the current frame is not in a static scene, determine whether the pixel point of the current frame is in a violent motion scene.
[0122] That is to say, in the above scene detection, if the motion intensity of the pixel point (x, y) of the current frame is not only greater than a predetermined first motion threshold (for example, LowMotionThreshold), but also greater than a preset second motion threshold (for example, HighMotionThreshold), and LowMotionThreshold ≤ HighMotionThreshold, then it can be further determined that the pixel point (x, y) of the current frame is in a violent motion scene.
[0123] In practice, both the first motion threshold (for example, LowMotionThreshold) and the second motion threshold (for example, HighMotionThreshold) can be fixed values. For example, they can be preset with reference to practical experience according to the image type, etc.
[0124] Step S4”, if it is determined that the pixel point of the current frame is in a violent motion scene, then perform fourth weighted fusion on the pixel value of the pixel point of the current frame and the pixel value of the pixel point of the reference frame corresponding to the pixel point of the current frame to obtain the pixel value after time-domain noise reduction of the pixel point of the current frame.
[0125] In some embodiments, performing fourth weighted fusion on the pixel value of a pixel point in the current frame and the pixel value of the pixel point in the reference frame corresponding to the pixel point in the current frame includes: calculating the temporally denoised pixel value of the pixel point (x, y) in the current frame according to the following formula (9).
[0126] fin_out(t)(x,y) = fin(t)(x,y)*(1-ω)+ref(t-1)(x',y')*ω (9)
[0127] In the above formula (9), fin(t)(x,y) is the original pixel value of the pixel point (x, y) in the current frame at time t, ref(t-1)(x',y') is the original pixel value of the pixel point (x', y') in the reference frame at time t, fin_out(t)(x,y) is the temporally denoised pixel value of the pixel point (x, y) in the current frame at time t, the pixel point (x', y') in the reference frame corresponds to the pixel point (x, y) in the current frame, (x, y) represents the coordinates of the pixel point in the current frame, (x', y') represents the coordinates of the pixel point in the reference frame, ω is related to the motion intensity of the pixel point (x, y) in the current frame. The greater the motion intensity of the pixel point (x, y) in the current frame, the smaller ω is; the smaller the motion intensity of the pixel point (x, y) in the current frame, the larger ω is, and 0≤ω≤1.
[0128] It should be understood that ω can be adaptively adjusted according to the motion intensity of the pixel point (x, y) in the current frame. ω = 0 and ω = 1 correspond to extreme cases. That is, in the case of ω = 0, the pixel point in the current frame is substantially not temporally denoised, and in the case of ω = 1, the original pixel value of the pixel point in the reference frame is directly used as the temporally denoised pixel value of the pixel point in the current frame. In practice, ω can also be a fixed value. For example, it can be preset with reference to practical experience according to the image type, etc., and the embodiments of the present disclosure do not make specific limitations.
[0129] Since it is difficult to perform accurate motion estimation and motion compensation on pixel points in a severely moving scene and a large amount of resources are required, motion estimation and motion compensation are not performed on pixel points in the current frame in a severely moving scene, and instead, weighted fusion is directly performed on the pixel value of the pixel point in the current frame and the pixel value of the pixel point in the reference frame corresponding to the pixel point in the current frame to obtain the temporally denoised pixel value of the pixel point in the current frame, which can avoid inappropriate denoising caused by inaccurate motion estimation and motion compensation and can also avoid resource shortages.
[0130] It should be understood that, without conflict, the above-mentioned temporally denoising methods for pixel points in different scenes (including static scenes, moving scenes, and severely moving scenes) of the current frame can be arbitrarily combined. The following is illustrated by examples.
[0131] Example 1
[0132] Figure 6 FIG. shows a schematic diagram of Example 1 of the time-domain noise reduction method according to an embodiment of the present disclosure.
[0133] As Figure 6 shown, in this example, according to the scene detection result, different time-domain noise reduction can be performed on the pixel points in the static scene and the pixel points in the moving scene of the current frame at time t (i.e., the input frame). Finally, the noise reduction results of the pixel points in different scenes of the current frame at time t are combined as the time-domain noise reduction result of the current frame at time t. At the same time, the time-domain noise reduction result of the current frame at time t can be stored, for example, stored in a memory such as a double data rate synchronous dynamic random access memory (DDR), so as to be used as a reference frame at time t+1.
[0134] Specifically, for the time-domain noise reduction method for the pixel points in the static scene of the current frame, reference can be made to the description above in combination with Figure 1 For the time-domain noise reduction method for the pixel points in the moving scene of the current frame, reference can be made to the description above in combination with Figure 4 The description will not be repeated here.
[0135] Example 2
[0136] Figure 7 FIG. shows a schematic diagram of Example 2 of the time-domain noise reduction method according to an embodiment of the present disclosure.
[0137] As Figure 7 shown, in this example, according to the scene detection result, different time-domain noise reduction can be performed on the pixel points in the static scene, the pixel points in the general motion scene, and the pixel points in the violent motion scene of the current frame at time t (i.e., the input frame). Finally, the noise reduction results of the pixel points in different scenes of the current frame at time t are combined as the time-domain noise reduction result of the current frame at time t. At the same time, the time-domain noise reduction result of the current frame at time t can be stored, for example, stored in a memory such as a double data rate synchronous dynamic random access memory (DDR), so as to be used as a reference frame at time t+1.
[0138] Specifically, for the time-domain noise reduction method for the pixel points in the static scene of the current frame, reference can be made to the description above in combination with Figure 1 For the time-domain noise reduction method for the pixel points in the general motion scene of the current frame, reference can be made to the description above in combination with Figure 4 For the time-domain noise reduction method for the pixel points in the violent motion scene of the current frame, reference can be made to the description above in combination with Figure 5 The description will not be repeated here.
[0139] According to the above Examples 1 and 2, obviously, the scene where the pixel points of the current frame are located can be divided more finely. For example, the static scene can be further divided into an absolute static scene and a suspected static scene (which may include small movements), and the moving scene can be further divided into a small movement scene, a medium movement scene, a large movement scene, and a violent movement scene. Moreover, corresponding motion thresholds can be set to determine the specific scene where the pixel points of the current frame are located. Each motion threshold can be preset and / or adjusted according to actual situations such as practical experience, expectations, and resource limitations.
[0140] In this case, the temporal noise reduction for the pixel points in the absolute static scene of the current frame and the temporal noise reduction for the pixel points in the suspected static scene can be distinguished by adjusting ρ in the above formula (6), the temporal noise reduction for the pixel points in the small movement scene of the current frame and the temporal noise reduction for the pixel points in the medium movement scene can be distinguished by adjusting γ in the above formula (8), and the temporal noise reduction for the pixel points in the large movement scene of the current frame and the temporal noise reduction for the pixel points in the violent movement scene can be distinguished by adjusting ω in the above formula (9).
[0141] It should be understood that ω ≤ γ ≤ ρ can be set to avoid sudden changes in the noise reduction degree and cause unstable noise reduction effects.
[0142] In a second aspect, an embodiment of the present disclosure provides a scene detection method, as Figure 8 shown, including steps S11 to S13.
[0143] Step S11, obtaining the current frame of the image and the reference frame corresponding to the current frame.
[0144] In the embodiment of the present disclosure, the input frame fin(t) at time t is used as the current frame, and the noise-reduced result ref(t - 1) of the input frame at time (t - 1) is used as the reference frame corresponding to the current frame fin(t).
[0145] Step S12, calculating the fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame according to the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame.
[0146] It should be understood that when the noise intensity (i.e., scene noise) in the area where the pixel points of the current frame are located is relatively large, the noise distribution may fluctuate. In this case, the sum of absolute differences SAD between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points in the reference frame may be relatively large. However, a large SAD does not necessarily mean that the pixel points of the current frame are in a moving scene. Therefore, the absolute difference ASD between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the sum of the pixel values of the corresponding pixel points in the reference frame can be referred to, and the motion state of the pixel points of the current frame can be determined based on the fusion of SAD and ASD. By referring to ASD, the influence of partial noise distribution fluctuations can be removed, and the accuracy of scene detection can be improved.
[0147] In some embodiments, the neighborhood range m×n centered on the pixel point (x, y) of the current frame can be selected as the motion estimation neighborhood, and the sum of absolute differences SAD between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame and the pixel values of the corresponding pixel points in the reference frame can be calculated according to the foregoing formula (1), and the absolute difference ASD between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame and the sum of the pixel values of the corresponding pixel points in the reference frame can be calculated according to the foregoing formula (2).
[0148] It should be understood that the neighborhood range m×n can be a pixel array centered on the pixel point (x, y) of the current frame. One of m and n is the number of rows of the pixel array, and the other is the number of columns of the pixel array. Both m and n can be odd numbers. The size of the neighborhood range m×n can be selected according to actual situations such as practical experience, expectations, and resource limitations, and the embodiments of the present disclosure do not make specific limitations.
[0149] In some embodiments, the fusion frame difference fusion_diff between the pixel point (x, y) of the current frame and the pixel point (x', y') of the reference frame corresponding to the pixel point (x, y) of the current frame can be determined according to the foregoing formula (3).
[0150] Step S13, determine whether the pixel points of the current frame are in a static scene or a moving scene according to the fusion frame difference.
[0151] In some embodiments, the normalized motion intensity of the pixel points of the current frame relative to the noise intensity in its area can be calculated, and then whether the pixel points of the current frame are in a static scene or a moving scene can be determined according to the calculated motion intensity.
[0152] For example, the motion intensity MotionStr of the pixel point (x, y) in the current frame can be calculated as MotionStr = fusion_diff / NoiseStr, where NoiseStr is the noise intensity of the region where the pixel point (x, y) in the current frame is located.
[0153] In some embodiments, if the motion intensity of the pixel point (x, y) in the current frame is not greater than a predetermined first motion threshold LowMotionThreshold, it is determined that the pixel point (x, y) in the current frame is in a static scene. If the motion intensity of the pixel point (x, y) in the current frame is greater than the predetermined first motion threshold LowMotionThreshold, it is determined that the pixel point (x, y) in the current frame is not in a static scene.
[0154] In addition, if the motion intensity of the pixel point (x, y) in the current frame is not only greater than the predetermined first motion threshold LowMotionThreshold, but also greater than a preset second motion threshold HighMotionThreshold, where LowMotionThreshold ≤ HighMotionThreshold, it can be further determined that the pixel point (x, y) in the current frame is in a violent motion scene.
[0155] In practice, both the first motion threshold (e.g., LowMotionThreshold) and the second motion threshold (e.g., HighMotionThreshold) can be preset and / or adjusted according to the image type, etc., with reference to practical experience. For example, both can be fixed values.
[0156] It should be understood that when the noise intensity (i.e., scene noise) of the region where the pixel point in the current frame is located is small, as Figure 9 shown, the scene detection method of the embodiments of the present disclosure may further include steps S12' to S13'.
[0157] Step S12', calculate the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point in the current frame and the pixel values of the corresponding pixel points in the reference frame, as the frame difference between the pixel point in the current frame and the pixel points in the reference frame corresponding to the pixel point in the current frame.
[0158] When the noise intensity (i.e., scene noise) of the region where the pixel point in the current frame is located is small, the sum of the absolute differences SAD between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point in the current frame and the pixel values of the corresponding pixel points in the reference frame can better reflect the motion state of the pixel point in the current frame. Therefore, SAD can be used as the frame difference between the pixel point in the current frame and the pixel points in the reference frame corresponding to the pixel point in the current frame.
[0159] Step S13’, determine whether the pixel points of the current frame are in a static scene or a moving scene according to the frame difference.
[0160] When the noise intensity (i.e., scene noise) of the area where the pixel points of the current frame are located is small, SAD can better reflect the motion state of the pixel points of the current frame. Therefore, it is possible to determine whether the pixel points of the current frame are in a static scene or a moving scene according to the frame difference represented by SAD.
[0161] It should be understood that in the scene detection method of the embodiments of the present disclosure, it is possible to adaptively select SAD as the frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame, or the fusion of SAD and ASD as the fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame according to the noise intensity (i.e., scene noise) of the area where the pixel points of the current frame are located. Then, according to the frame difference or fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame, determine whether the pixel points of the current frame are in a static scene or a moving scene, which will not be elaborated here.
[0162] Similarly, the embodiments of the present disclosure do not specifically limit how to determine and / or calculate the noise intensity (i.e., scene noise) of the area where the pixel points of the current frame are located. The noise intensity (i.e., scene noise) of the area where the pixel points of the current frame are located can be determined and / or calculated according to the methods in the relevant technical fields.
[0163] Moreover, the embodiments of the present disclosure do not limit the specific numerical value or numerical range of “large” or “small” of the noise intensity (i.e., scene noise) of the area where the pixel points of the current frame are located. In some embodiments, if the noise intensity (i.e., scene noise) of the area where the pixel points of the current frame are located is not greater than a predetermined threshold, it is determined that the noise intensity of the area where the pixel points of the current frame are located is small. If the noise intensity (i.e., scene noise) of the area where the pixel points of the current frame are located is greater than the predetermined threshold, it is determined that the noise intensity of the area where the pixel points of the current frame are located is large. Those skilled in the art can preset and / or adjust the predetermined threshold according to practical situations such as practical experience, expectations, and resource limitations. For example, the predetermined threshold can be preset and / or adjusted according to the expected overall noise reduction effect on the current frame and the acceptable resource consumption cost.
[0164] In a third aspect, the embodiments of the present disclosure provide a computer-readable storage medium, as Figure 10 shown, on which a computer program is stored, and the computer program enables a processor to implement the above-mentioned time-domain noise reduction method based on scene detection or the above-mentioned scene detection method of the embodiments of the present disclosure.
[0165] It should be understood that in the time-domain noise reduction method based on scene detection and the corresponding scene detection method provided by the embodiments of the present disclosure, the execution order of each step is not limited to the order described above in combination with the drawings. Without departing from the scope of the present disclosure, the execution order of each step can be adjusted, and moreover, some steps can be executed in parallel.
[0166] In addition, those of ordinary skill in the art can understand that all or some of the steps and functional modules / units in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components. For example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some physical components or all physical components can be implemented as software executed by a processor (such as a central processing unit, a digital signal processor, or a microprocessor), or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). The computer storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disc (DVD), or other optical disc storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, the communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0167] Example embodiments have been disclosed herein, and although specific terms have been used, they are used for and should be construed only as general illustrative meanings and not for the purpose of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, features, characteristics, and / or elements described in connection with a particular embodiment can be used alone or in combination with features, characteristics, and / or elements described in connection with other embodiments. Accordingly, those skilled in the art will understand that various forms and details changes can be made without departing from the scope of the present disclosure as set forth by the appended claims.
Claims
1. A time-domain noise reduction method based on scene detection, comprising: obtaining a current frame of an image and a reference frame corresponding to the current frame; performing scene detection to determine whether the pixel points of the current frame are in a static scene; if it is determined that the pixel points of the current frame are in a static scene, performing pre-filtering related to the current frame on the pixel points of the reference frame corresponding to the pixel points of the current frame to obtain pre-filtered pixel values of the pixel points of the reference frame; performing first weighted fusion on the original pixel values and the pre-filtered pixel values of the pixel points of the reference frame to obtain first weighted fusion pixel values of the pixel points of the reference frame; and performing second weighted fusion on the pixel values of the pixel points of the current frame and the first weighted fusion pixel values of the pixel points of the reference frame corresponding to the pixel points of the current frame to obtain time-domain noise reduction pixel values of the pixel points of the current frame.
2. The time-domain noise reduction method according to claim 1, wherein, performing scene detection to determine whether the pixel points of the current frame are in a static scene includes: calculating a fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame according to the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel points of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame; and determining the motion state of the pixel points of the current frame according to the fusion frame difference, wherein, the motion estimation neighborhood of the pixel points of the current frame includes a set of adjacent pixel points centered on the pixel points of the current frame.
3. The time-domain noise reduction method according to claim 2, wherein, calculating a fusion frame difference between the pixel points of the current frame and the pixel points of the reference frame corresponding to the pixel points of the current frame includes: selecting a neighborhood range m×n centered on the pixel point (x, y) of the current frame as the motion estimation neighborhood, and calculating the sum of the absolute differences SAD between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference ASD between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame as follows, (x, y) represents the coordinates of the pixel points of the current frame, fin(t)(x + j, y + i) represents the pixel values of each pixel point in the motion estimation neighborhood of the pixel point (x, y) of the current frame at time t, the pixel point (x, y) of the current frame corresponds to the pixel point (x', y') of the reference frame, (x', y') represents the coordinates of the pixel points of the reference frame, ref(t - 1)(x' + j, y' + i) represents the pixel values of the pixel point (x', y') of the reference frame at time t, the neighborhood range m×n is a pixel array centered on the pixel point (x, y) of the current frame, one of m and n is the number of rows of the pixel array, the other is the number of columns of the pixel array, and both m and n are odd numbers; and Determine the fusion frame difference fusion_diff between the pixel point (x, y) of the current frame and the pixel point (x', y') of the reference frame. fusion_diff = SAD * α + ASD * (1 - α), where α is related to the noise intensity of the area where the pixel point (x, y) of the current frame is located. The greater the noise intensity of the area where the pixel point (x, y) of the current frame is located, the smaller α is; the smaller the noise intensity of the area where the pixel point (x, y) of the current frame is located, the greater α is, and 0 ≤ α ≤ 1.
4. The time-domain noise reduction method according to claim 3, wherein, Determining the motion state of the pixel point of the current frame according to the fusion frame difference includes: Calculating the motion intensity MotionStr of the pixel point (x, y) of the current frame. MotionStr = fusion_diff / NoiseStr, where NoiseStr is the noise intensity of the area where the pixel point (x, y) of the current frame is located; and If the motion intensity of the pixel point (x, y) of the current frame is not greater than a predetermined first motion threshold, it is determined that the pixel point (x, y) of the current frame is in a static scene. If the motion intensity of the pixel point (x, y) of the current frame is greater than the predetermined first motion threshold, it is determined that the pixel point (x, y) of the current frame is not in a static scene.
5. The time-domain noise reduction method according to claim 2, wherein, Performing scene detection to determine whether the pixel point of the current frame is in a static scene further includes: Calculating the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point of the current frame and the pixel values of the corresponding pixel points of the reference frame as the frame difference between the pixel point of the current frame and the pixel point of the reference frame corresponding to the pixel point of the current frame; and Determining the motion state of the pixel point of the current frame according to the frame difference.
6. The time-domain noise reduction method according to claim 1, wherein, Performing pre-filtering related to the current frame on the pixel point of the reference frame corresponding to the pixel point of the current frame includes: Calculating the block differences between the pixel point (x, y) of the current frame and each pixel point in the search neighborhood sw × sh centered on the pixel point (x', y') of the reference frame. The pixel point (x', y') of the reference frame corresponds to the pixel point (x, y) of the current frame. (x, y) represents the coordinates of the pixel point of the current frame, and (x', y') represents the coordinates of the pixel point of the reference frame; According to the calculated block differences, determining the weight value w of each pixel point in the search neighborhood of the reference frame relative to the pixel point (x, y) of the current frame. The weight value of each pixel point is a linear or non-linear function of its corresponding block difference. The smaller the block difference corresponding to each pixel point, the greater its weight value; and Performing the following pre-filtering on the pixel point (x', y') of the reference frame Among them, ref(t - 1)(x', y') is the original pixel value of the pixel point (x', y') in the reference frame at time t, f1(ref(t - 1)(x', y')) is the pre-filtered pixel value of the pixel point (x', y') in the reference frame at time t, the search neighborhood sw×sh is a pixel array centered on the pixel point (x', y') in the reference frame, one of sw and sh is the number of rows of the pixel array, the other is the number of columns of the pixel array, both sw and sh are odd numbers, Prn is the original pixel value of each pixel point in the search neighborhood, wn is the weight value of each pixel point in the search neighborhood, n is the serial number of each pixel point in the search neighborhood, n is a natural number, and n∈sw*sh means that the pixel point with serial number n is within the search neighborhood sw×sh of the reference frame.
7. The temporal domain noise reduction method according to claim 6, wherein, calculating the block difference between the pixel point (x, y) of the current frame and each pixel point in the search neighborhood sw×sh centered on the pixel point (x', y') of the reference frame includes: selecting a block neighborhood pw×ph centered on the pixel point (x, y) of the current frame and a block neighborhood pw×ph centered on each pixel point in the search neighborhood of the reference frame; and calculating the sum of the absolute differences between the pixel values of each pixel point in the block neighborhood centered on the pixel point (x, y) of the current frame and the pixel values of the corresponding pixel points in each block neighborhood of the reference frame, as the block difference between the pixel point (x, y) of the current frame and the pixel point centered on each neighborhood block of the reference frame, wherein, the block neighborhood pw×ph is a pixel array, one of pw and ph is the number of rows of the pixel array, the other is the number of columns of the pixel array, and both pw and ph are odd numbers.
8. The temporal domain noise reduction method according to claim 6, wherein, performing the first weighted fusion on the original pixel value and the pre-filtered pixel value of the pixel point of the reference frame includes: calculating the first weighted fusion pixel value of the pixel point of the reference frame according to the following formula, ref_filtered(t - 1)(x', y') = ref(t - 1)(x', y') * (1 - β) + f1(ref(t - 1)(x', y')) * β, where ref_filtered(t - 1)(x', y') is the first weighted fusion pixel value of the pixel point (x', y') in the reference frame at time t, β is a parameter related to the texture details of the image, the less the texture details of the image, the larger β, the more the texture details of the image, the smaller β, and 0≤β≤1.
9. The temporal domain noise reduction method according to claim 8, wherein, performing the second weighted fusion on the pixel value of the pixel point of the current frame and the first weighted fusion pixel value of the pixel point of the reference frame corresponding to the pixel point of the current frame includes: The pixel value after temporal noise reduction of the pixel points in the current frame is calculated according to the following formula: fin_out(t)(x,y) = fin(t)(x,y) * (1 - ρ) + ref_filtered(t - 1)(x',y') * ρ, where fin_out(t)(x,y) is the pixel value after temporal noise reduction of the pixel point (x,y) in the current frame at time t, fin(t)(x,y) is the original pixel value of the pixel point (x,y) in the current frame at time t, ρ is related to the motion intensity of the pixel point (x,y) in the current frame. The greater the motion intensity of the pixel point (x,y) in the current frame, the smaller ρ is; the smaller the motion intensity of the pixel point (x,y) in the current frame, the greater ρ is, and 0 ≤ ρ ≤ 1.
10. The temporal noise reduction method according to claim 1, further comprises: determining that the pixel points in the current frame are not in a static scene, and then determining that the pixel points in the current frame are in a motion scene; performing motion compensation on the pixel points in the reference frame corresponding to the pixel points in the current frame; and performing third weighted fusion on the pixel value of the pixel points in the current frame and the pixel value after motion compensation of the pixel points in the reference frame corresponding to the pixel points in the current frame to obtain the pixel value after temporal noise reduction of the pixel points in the current frame.
11. The temporal noise reduction method according to claim 10, wherein, performing motion compensation on the pixel points in the reference frame corresponding to the pixel points in the current frame comprises: calculating the block difference between the pixel point (x,y) in the current frame and each pixel point in the search neighborhood sw×sh centered on the pixel point (x',y') in the reference frame. The pixel point (x',y') in the reference frame corresponds to the pixel point (x,y) in the current frame. (x,y) represents the coordinates of the pixel point in the current frame, and (x',y') represents the coordinates of the pixel point in the reference frame; if the block difference between the pixel point (x,y) in the current frame and the pixel point (x1,y1) in the search neighborhood of the reference frame is the smallest, then determining the Euclidean distance between the pixel point (x,y) in the current frame and the pixel point (x1,y1) in the reference frame as the motion vector mv of the pixel point (x,y) in the current frame, mv = (x1 - x, y1 - y), where (x1,y1) represents the coordinates of the pixel point in the search neighborhood of the reference frame; and performing the following motion compensation on the pixel point (x',y') in the reference frame, ComRef(t - 1)(x',y') = ref(t - 1)(x' + mv(0), y' + mv(1)), where mv(0) = x1 - x, mv(1) = y1 - y, ComRef(t - 1)(x',y') is the pixel value after motion compensation of the pixel point (x',y') in the reference frame at time t, ref(t - 1)(x',y') is the original pixel value of the pixel point (x',y') in the reference frame at time t, the search neighborhood sw×sh is a pixel array centered on the pixel point (x',y') in the reference frame, one of sw and sh is the number of rows of the pixel array, the other is the number of columns of the pixel array, and both sw and sh are odd numbers.
12. The time-domain noise reduction method according to claim 11, wherein, performing third weighted fusion on the pixel value of the pixel point of the current frame and the pixel value after motion compensation of the pixel point of the reference frame corresponding to the pixel point of the current frame includes: calculating fin_out(t)(x,y) = fin(t)(x,y)*(1 - γ)+ComRef(t - 1)(x',y')*γ, wherein, fin_out(t)(x,y) is the pixel value after time-domain noise reduction of the pixel point (x,y) of the current frame at time t, fin(t)(x,y) is the original pixel value of the pixel point (x,y) of the current frame at time t, γ is related to the motion intensity of the pixel point (x,y) of the current frame, the greater the motion intensity of the pixel point (x,y) of the current frame, the smaller γ is, the smaller the motion intensity of the pixel point (x,y) of the current frame, the greater γ is, and 0 ≤ γ ≤ 1.
13. The time-domain noise reduction method according to claim 1, further including: after determining that the pixel point of the current frame is not in a static scene, determining whether the pixel point of the current frame is in a violently moving scene; and if it is determined that the pixel point of the current frame is in a violently moving scene, performing fourth weighted fusion on the pixel value of the pixel point of the current frame and the pixel value of the pixel point of the reference frame corresponding to the pixel point of the current frame to obtain the pixel value after time-domain noise reduction of the pixel point of the current frame.
14. The time-domain noise reduction method according to claim 13, wherein, performing fourth weighted fusion on the pixel value of the pixel point of the current frame and the pixel value of the pixel point of the reference frame corresponding to the pixel point of the current frame includes: calculating fin_out(t)(x,y) = fin(t)(x,y)*(1 - ω)+ref(t - 1)(x',y')*ω, wherein, fin(t)(x,y) is the original pixel value of the pixel point (x,y) of the current frame at time t, ref(t - 1)(x',y') is the original pixel value of the pixel point (x',y') of the reference frame at time t, fin_out(t)(x,y) is the pixel value after time-domain noise reduction of the pixel point (x,y) of the current frame at time t, the pixel point (x',y') of the reference frame corresponds to the pixel point (x,y) of the current frame, (x,y) represents the coordinates of the pixel point of the current frame, (x',y') represents the coordinates of the pixel point of the reference frame, ω is related to the motion intensity of the pixel point (x,y) of the current frame, the greater the motion intensity of the pixel point (x,y) of the current frame, the smaller ω is, the smaller the motion intensity of the pixel point (x,y) of the current frame, the greater ω is, and 0 ≤ ω ≤ 1.
15. A scene detection method, including: acquiring the current frame of the image and the reference frame corresponding to the current frame; calculating the fusion frame difference between the pixel point of the current frame and the pixel point of the reference frame corresponding to the pixel point of the current frame according to the sum of the absolute differences between the pixel values of each pixel point in the motion estimation neighborhood of the pixel point of the current frame and the pixel values of the corresponding pixel points of the reference frame, and the absolute difference between the sum of the pixel values of each pixel point in the motion estimation neighborhood of the pixel point of the current frame and the sum of the pixel values of the corresponding pixel points of the reference frame; and Determine whether the pixel points of the current frame are in a static scene or a moving scene according to the fused frame difference.
16. A computer-readable storage medium having a computer program stored thereon, the computer program causing a processor to implement the method according to any one of claims 1 to 15.