Temporally consistent video denoising
Patent Information
- Application Number
- PCT/US2025/018221
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-06
- Filing Date
- 2025-03-03
- Publication Date
- 2025-10-02
AI Technical Summary
Conventional methods struggle to efficiently remove film-grain noise from videos while preserving the cinematic quality and texture, leading to high coding costs and visual degradation.
A machine learning-based framework using a noise-estimation and image-denoising module architecture, incorporating neural networks and entropy filters, to estimate and remove film-grain noise while maintaining temporal consistency and texture.
Effectively reduces film-grain noise, enhancing video quality by preserving essential details and reducing coding costs through temporally consistent denoising.
Smart Images

Figure US2025018221_02102025_PF_FP_ABST
Abstract
Description
TEMPORALLY CONSISTENT VIDEO DENOISING 1. Cross-Reference to Related Applications
[0001] This application claims the benefit of priority from European Application No. 24180475.6 filed on June 6, 2024 and U.S. Provisional Application No.63 / 562,505, filed on March 7, 2024, each of which is incorporated by reference herein in its entirety. 2. Field of the Disclosure
[0002] Various example embodiments relate generally to image and video denoising and, more specifically but not exclusively, to the digital film-grain technology. 3. Background
[0003] On physical film, film grain is the random physical texture made from small metallic silver particles found on processed photographic celluloid. In digital photography, the visual and artistic effects of film grain can be simulated by adding a digital grain pattern to a digital image after the digital image is taken. Because film grain may be difficult to encode, e.g., due to its (quasi)random nature, a video encoder may be configured to remove the film grain during encoding. In some cases, during playback, the corresponding video decoder may be configured to synthesize the film grain and add it back in. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS
[0004] Disclosed herein are various embodiments of temporally consistent video denoising directed at removing image noise. According to one example, a denoising algorithm uses a noise estimation block and an image denoising block operatively connected to one another. In various examples, the noise estimation block is implemented using a neural network or a filter arrangement including an entropy filter and is designed to analyze noisy frame sequences and generate a noise strength map. With this noise strength map, the image denoising block operates to perform the image denoising more efficiently, effectively reducing the image noise while preserving the texture of the original footage. In some examples, a joint loss objective is used to find configurations of both blocks that result in nearly optimal performance of the denoising algorithm.
[0005] According to an example embodiment, provided is a method of denoising a video sequence comprising: generating a noise strength map based on a set of consecutive video frames of the video sequence, the noise strength map representing an estimate of image noise for a middle frame of the set; and producing a denoised frame using a first neural network trained to 1 remove image noise from the middle frame of the set in response to the set and the noise strength map being applied thereto as inputs.
[0006] According to another example embodiment, provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above method.
[0007] According to yet another example embodiment, provided is an apparatus for denoising a video sequence comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: generate a noise strength map based on a set of consecutive video frames of the video sequence, the noise strength map representing an estimate of image noise for a middle frame of the set; and produce a denoised frame using a first neural network trained to remove image noise from the middle frame of the set in response to the set and the noise strength map being applied thereto as inputs. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:
[0009] FIG.1 is a block diagram illustrating a noise removal (NR) apparatus according to some examples.
[0010] FIG.2 is a block diagram illustrating a noise estimation block that can be used in the NR apparatus of FIG.1 according to some examples.
[0011] FIG.3 is a block diagram illustrating a neural network that can be used in the noise estimation block of FIG.2 according to one example.
[0012] FIG.4 is a block diagram illustrating a noise estimation block that can be used in the NR apparatus of FIG.1 according to some other examples.
[0013] FIGS.5A-5B pictorially illustrate the performance of a Gaussian filter used in the noise estimation block of FIG.4 according to one example.
[0014] FIG.6 pictorially illustrates the performance of an entropy filter used in the noise estimation block of FIG.4 according to one example.
[0015] FIGS.7A-7B pictorially illustrate the performance of some of entropy-map postprocessing operations used in the noise estimation block of FIG.4 according to one example.
[0016] FIG.8 graphically illustrates a set of different logistic curves for a logistic mapper used in the noise estimation block of FIG.4 according to some examples. 2
[0017] FIG.9 graphically illustrates an optimized logistic curve that can be used in the logistic mapper of the noise estimation block of FIG.4 according to one example.
[0018] FIG.10 pictorially illustrates the performance of the logistic mapper used in the noise estimation block of FIG.4 according to one example.
[0019] FIG.11 pictorially illustrates the performance of an optical flow module and a motion compensation module used in the noise estimation block of FIG.4 according to one example.
[0020] FIGS.12A-12D pictorially illustrate the warping operations performed in the motion compensation module of the noise estimation block of FIG.4 according to one example.
[0021] FIG.13 pictorially illustrates the performance of a self-guided filter used in the noise estimation block of FIG.4 according to one example.
[0022] FIG.14 is a block diagram illustrating an image denoising block that can be used in the NR apparatus of FIG.1 according to some examples.
[0023] FIG.15 is a block diagram illustrating a neural network that can be used in the image denoising block of FIG.14 according to one example.
[0024] FIG.16 is a flowchart of a method of denoising a video sequence that can be implemented with the NR apparatus of FIG.1 according to some examples.
[0025] FIG.17 is a block diagram illustrating a computing device, one or more instances of which can be used to implement the NR apparatus of FIG.1 according to various examples. DETAILED DESCRIPTION
[0026] Video denoising is typically applied to mitigate noise and remove artifacts, thereby significantly enhancing the quality of video for a multitude of applications, including but not limited to surveillance, medical imaging, entertainment, and remote sensing. Noise in videos can arise from various sources and manifest itself in diverse forms. For example, spatial noise, characterized by random fluctuations in pixel values, often occurs during image acquisition and transmission. Temporal noise manifests itself as flickering or irregularities in consecutive frames, e.g., due to such factors as sensor imperfections and / or compression artifacts. The process of video compression typically presents its own unique set of noise and quality degradation issues.
[0027] Various embodiments disclosed herein can beneficially be used for temporally consistent video denoising with respect to at least some of the different noise types mentioned above. For illustration purposes and without any implied limitations, example embodiments are described below with specific reference to film-grain noise. Based on the provided description, a 3 person of ordinary skill in the pertinent art will be able to make and use other embodiments directed to mitigation of other types of video noise without any undue experimentation.
[0028] Film-grain noise, which emerges from the development of analog film via corresponding chemical reactions of silver-halide crystals, is a unique image characteristic that is responsible for a vintage look of the footage. Example features of the film-grain noise include: (i) film-grain noise can typically be modeled as a zero-mean Gaussian noise, and the noise variance varies with the intensity of noise-free signal; (ii) film-grain noise is temporally independent; and (iii) film-grain noise is spatially correlated. For images acquired with digital camera sensors, film-grain noise is not present in the corresponding images and video, thereby offering improved visual quality and robustness compared to the analog film medium. Nevertheless, in some cases, video-content creators prefer that the visual and artistic effects of film grain be present for their aesthetic and / or creative appeal, finding the digital medium to be too sharp or lacking in emotion. To emulate the cinematic quality of analog film, post- processing, such as color palette adjustments, contrast adjustments, and artificial film grain insertion, are applied to digital content, thereby transforming the film grain into an intentional visual effect from previously being an incidental byproduct of inhomogeneous chemical reactions. Due to its temporally independent nature and presence of high-frequency details, film- grain noise cannot typically be accurately predicted by conventional motion compensation, and significant film-grain noise components remain in the prediction residue. The presence of such components typically inflicts a relatively high coding cost in the form of additional bits to be encoded in the discrete cosine transform (DCT) domain. Thus, efficient methods directed at removing film-grain noise are desirable.
[0029] At least some of the above-indicated problems in the state of the art can beneficially be addressed using at least some embodiments disclosed herein. For example, one embodiment provides a machine learning-based framework for removing film-grain noise from videos using noise-estimation and image-denoising modules coupled to one another. Example features of this architecture include: o The noise-estimation module is configured to estimate the amount of film-grain noise which is then used to drive a neural network used in the denoising module. ^ At least two different implementations of the noise-estimation module can be used for this purpose: (i) a convolutional-neural-network-based film- grain noise estimator and (ii) an entropy-filter-based film-grain noise estimator. ^ In some examples, to make the film-grain noise estimate temporally consistent, we perform additional re-mapping of the estimate based on the 4 denoising module’s sensitivity to noise with a custom logistic function, followed by optical flow and motion compensation. Furthermore, to smooth out the remains of temporal flicker in the noise estimates, some examples also employ a guided image filter. o Given the film-grain noise estimate generated by the noise-estimation module, a neural-network-based denoising component of the denoising module is driven to remove the corresponding film-grain noise from the content while preserving texture and temporal stability. Mathematical Formulation
[0030] In this section, we outline a mathematical formulation that can be used in the above-indicated framework for removing film-grain noise from images and videos according to some examples.
[0031] We denote a clean video sequence as ^^ ൌ ^^^^^^்௧ୀ^, where each clean frame ^^^^represents the luminance (Y-channel) in the YUV color space. The relationship between theclean frames ^^^^ and film-grain noisy frames, ^^ ൌ ^^^^^^்௧ୀ^, where each noisy frame is denoted as ^^^^, is modeled as: ^^^^^^^, ^^^ ൌ ^^^^^^^,^^^ ^ ^^^^^^^^^^^^^^^^^^^^^,^^^ (1)where ^^^^^^^^^^^^^^^^^^is the example steps:^^^^^^^^^^^^^^^^^^^^^, ^^^ ൌ ^^^^^^^,^^^ ∗ ^^^^^^^^^,^^^ (2)The notations used in
[0032] For each frame, film-grain noise can be constructed using a two-dimensional (2D) random field in the frequency domain. With ^^^as the block size, a (^^^,^^^) sized film grain noise block ^^^^^^^,^^^ can be rendered in accordance with Eq. ^^^^^^^,^^^ ൌ ^^^^^^^^^ ^^^^^,^^^ ^where ^^^^^,^^^ denotes the DCT coefficients of a block of size); ^^^^^^^^^. ^ is theinverse DCT; ^^^,^^^ is a pixel location in the spatial domain; and ^^^,^^^ are indices in the frequency domain. Some alternating current (AC) coefficients of ^^^^^,^^^ are Gaussian random numbers of mean 0, with standard deviation of 1. The other coefficients, including the direct current (DC) coefficient, are all zeros. When some frequency band of AC coefficients has a non- zero value in the DCT domain, the corresponding ^^^^^^^,^^^ appears as film-grain noise in the pixel domain. Let ^^^and ^^^be the start and end frequencies for a block-based noise pattern for 5 each 2D direction (e.g., horizontal and vertical). With the frequency range ^^^≤ f ≤ ^^^determining the non-zero frequencies, the resulting noise block is [^1,1].
[0033] The above process can be repeated for each block of size (^^^,^^^) until the desired film grain noise of size (ℎ^^^^^^ℎ^^,^^^^^^^^ℎ) ^^^^^^is built block by block. After this procedure, the blocks can be stitched together using a 3-tap filter applied across the block-boundary pixels to smooth out any visible lines or discontinuities at the edges of the block patches.
[0034] The variables at pixel location^^^,^^^and time ^^ used in the above equations include: ^ LUT(l) is a custom lookup table (LUT) designed as a function of pixel intensity. As an example, with ^^ as the luminance value of the pixel in the narrow range representing ^^ െ ^^^^^^^^, ^^^^^ as minimum luminance value, ^^^^௫ as minimumluminance value, ^^ representing ^^ െ ^^^^^^^^ with ^^^^^ and ^^^^௫ as minimum andmaximum values, the LUT can be defined by a triangle function with the its peak at ൫^^^^௩^௧, ^^^^௫൯, which can be mathematically represented as:^^^^௫ ^ ^^^൫^^ െ ^^ ൯, ^^^^ ^^ ^ ^^ ^ ^^^^^^^^^^^^ ൌ ^ ^^௩^௧ ^^^ ^^௩^௧(6) ^^^^^ ^ ^^ ^Note: In some examples, a film grain module’s LUT may involve more complexity than this two-piece linear formulation, potentially mapping each distinct input value to a unique output value through a more complex transformation. In some examples of our training process described in more detail below, we employ the LUT defined by Eq. (6) with ^^^^^= 16, ^^^^௫ = 235, ^^^^^ = 0.0, ^^^^௫ = 1.0, and ൫^^^^௩^௧ , ^^^^௫൯ =we can also employ random pivots during training formore robustness. ^ ^^^^^^is the film-grain noise scaling matrix based on ^^^^′^^ intensity obtained through LUT(x). ^ ^^^^^^^,^^^ is the clean intensity pixel. ^ ^^^^ is the film-grain noise block of size (ωG x ωG ) obtained via a 2D-inverse DCT^⋅^ based function independent of ^^^^. ^ ^^^^^^is the film-grain noise (with the same spatial size as ^^^^) obtained through a film- synthesis process explained in more detail below. 6 ^ γ is the noise strength scaling hyperparameter. In some examples, γ is in the range between 1 and 50. ^ ^^^^is the noise strength map. ^ ^^^^^^^^^^^^^^^^^^is the scaled film grain noise.
[0035] Based on the above, the noisy frame ^^^^^^^,^^^ can be mathematically represented as a function (^^^⋅^) of the corresponding clean frame ^^^^^^^,^^^ as follows: ^^^^^^^, ^^^ ൌ ^^^ ^^^^^^^,^^^^ (7)Video denoising can be then represented using an inverse form of Eq. (7) as follows: ^^^^^^^,^^^ ൌ ^^ି^൫^^^^^^^,^^^൯ (8)
[0036] In practice, ^^ି^^⋅^ be a relatively hardproblem. In some learnable sub-problems ℎ^⋅^and ^^^⋅^. Given the noisy pixel data of the we predict ^^^^^^^^,^^^ representing thenoise strength estimate of film-grain noise at pixel ^^^, ^^^ through the function ℎ^⋅^. Then,denoising processing is performed with the awareness of the predicted noise strength estimate^^^^^^^^,^^^to obtain^^^^^^^^,^^^via the function ^^^⋅^, where^^^^^ is the predicted clean frame at time ^^.Mathematically, this processing sequence can be expressed as: ^^^^^^,^^^ ൌ ℎ൫^^^^^^^, ^^^൯ (9)
[0037] In some as being universalfunction approximators, we introduce a video denoising architecture including a first neural network ^^^^^^^^ே^ and a second neural network ^^^^^^^ோ^, where ^^^ே& ^^ோare the learnable parameters optimized through a learning objective. These neural networks are trained to approximate the above-described functions ℎ^⋅^ and ^^^⋅^ as follows: NE^^^ோ^ ^ ℎ^⋅^ (11)^^^^^^^^ே^ ^ ^^^⋅^ (12)Noise Removal
[0038] FIG.1 is a block diagram illustrating a noise removal (NR) apparatus 100 according to some examples. The NR apparatus 100 is configured to denoise an input sequence 102 of film-grain noisy image frames ^^^^^^^,^^^ using a sliding-window mechanism to generate acorresponding output sequence 122 of denoised image frames^^^^^^^^,^^^. In the example shown,the sliding window 101 includes five frames and is centeredimage frame ^^^^^^^,^^^. In other examples, the sliding window 101 may include a different (from five) number of film-grain noisy image frames ^^^^^^^,^^^. For a given position of the sliding window 101 within the input 7 sequence 102, the NR apparatus 100 outputs a single denoised image frame ^ ^^^^^^^, ^^^corresponding to the image frame ^^^^^^^,^^^. The output sequence 122 is generated by sequentially shifting the sliding window 101 by one frame and generating the correspondingdenoised image frame ^ ^^^^^^^, ^^^ for each shifted position of the sliding window 101.
[0039] The NR apparatus 100 includes a noise estimation block 110 and an image denoising block 120. Each of the blocks 110, 120 receives a copy of the input sequence 102.For each image frame ^^^^^^^, ^^^, the noise estimation block 110 operates to generate a noisestrength map ^^^^^^^,^^^. A sequence 112 of such noise strength maps ^^^^^^^,^^^ corresponding to different positions of the sliding window 101 along the input sequence 102 is directed to the image denoising block 120. For each position of the sliding window 101, the image denoising block 120 operates to denoise the middle frame of the sliding window 101, with guidance obtained from the corresponding noise strength map ^^^^^^^,^^^of the sequence 112. Example embodiments of the noise estimator block 110 and the image denoising block 120 are described in more detail below.
[0040] In different examples, the noise estimation block 110 is configured to generate the noise strength map ^^^^^^^,^^^ in different ways, including but not limited to (i) with an appropriately trained neural network and (ii) with an entropy-based filter arrangement. Depending on the visual content of the input sequence 102, one way may be more advantageous than another way. For example, for image regions including human faces and / or skin, noise estimation tends to be more accurate with the entropy-based filter implementation of the noise estimation block 110. In such cases, this implementation ensures that, after the film-grain noise is removed, essential details of those regions of the image remain intact, thereby beneficially enhancing the overall visually perceived quality of the denoised output 122.
[0041] FIG.2 is a block diagram illustrating the noise estimation block 110 according to some examples. In the example shown, the noise estimation block 110 is implemented using a neural network (also see FIG.3). For each time t, a corresponding set {^^^^}^^of noisy frames located within the width 2^ of the sliding window 101 is applied to the neural network of the noise estimation block 110. The output of the neural network is a predicted noise strength map^^^ t corresponding to the middle frame ^^^^ of the set {^^^^}^^.
[0042] FIG.3 is a block diagram illustrating a neural network 300 used in the noise estimation block 110 of FIG.2 according to one example. In the example shown, the neural network 300 has a depth of five convolutional layers, which are labeled in FIG.3 using the reference numerals 310-350. In other examples, a different (from five) number of convolutional layers can alternatively be used in the neural network 300. 8
[0043] The input convolutional layer 310 has a channel depth equivalent to the number of frames in the sliding window 101. In this example, each frame has the size of (96^96) pixels, and there are five frames in the sliding window 101. The input convolutional layer 310 transforms five channels into 32 channels. Each of the four convolutional layers 320-350 maintains this 32-channel depth, until the final layer 350 transforms the 32 channels into a single-channel output representing the predicted per pixel film-grain noise strength map ^^^^^of the middle frame ^^^^of the sliding window 101. Following each convolution operation of the layers 310-350, the ReLU activation function is used to ensure non-linearity, which beneficially enables the neural network 300 to capture different relatively complex noise characteristics. In other embodiments, other suitable activation functions can also be used.
[0044] FIG.4 is a block diagram illustrating a filter arrangement 400 used in the noise estimation block 110 of the NR apparatus 100 (FIG.1) according to some other examples. The filter arrangement 400 includes a plurality of filter stages described in more detail below. The description of the filter arrangement 400 that follows is given with continued reference to FIG.4 and with further reference to FIGS.5-13.
[0045] For time t, the input to the filter arrangement 400 is the set {^^^^}^^located within the sliding window 101, where ^=2. This set is applied a Gaussian filter 410, and the output of the Gaussian filter 410 is a corresponding smoothened set 412, {^^^^}^^, of image frames. Gaussian filtering performed in the Gaussian filter 410 helps to reduce at least some of the prevalent noise from the frames the set {^^^^}^^. With the noise having been smoothed out by the Gaussian filter 410, a downstream entropy filter 420 can operate more effectively, e.g., due to a better effective focus on extracting the true texture information without being misled by noise artifacts. For ^=2, the contents of the set 412, {^^^^}^2, can be expressed as: ^^^^ା^^^^^,^^^ ൌ ^^^,ఙ^^^௧ା^^,∀ ^^ ∈ ^െ2,െ1,0,1,2^ (13)where ^^^,ఙ^⋅^, is the Gaussian filter with a size ^^ and a standard deviation ^^. In some examples, the parameter values ^^ = 6 and ^^ = 0.5 may provide a nearly optimal performance.
[0046] FIGS.5A-5B pictorially illustrate the performance of the Gaussian filter 410 according to one example. More specifically, FIG.5A shows an example of the input noisy image frame ^^௧located within the sliding window 101. FIG.5B shows a corresponding example of the smoothened image frame ^^^^of the set 412 generated with the Gaussian 410. Note that the Gaussian-filter smoothing manifests itself in the smoothened image frame ^^^^of FIG.5B, e.g., as somewhat blurred boundaries between different objects in the image, such as, for example, the sky and the trees, compared to those in the input image frame ^^௧of FIG.5A. 9
[0047] The Gaussian corrected frames ^^^^ା^^generated with the Gaussian filter 410 are applied to the entropy filter 420. The output of the entropy filter 420 is a set 422 of entropy maps ^^^^ା^^. In a representative example, an entropy map ^^^^ା^^assigns lower values to non- textured areas of the frame and higher values to textured areas. The set 422 of entropy maps ^^^^ା^^is then subjected to one or more postprocessing operations 424. In one example, a plurality of postprocessing operations 424 includes (i) normalization, e.g., using the 95th percentile, to reduce the number of outliers and (ii) clipping the result of the normalization operation to lie between 0 and 1. The result of the normalization and clipping operations of the plurality ofpostprocessing operations 424 includes normalized entropy maps, ^^^^^ା^^. The plurality ofpostprocessing operations 424 also includes subjecting the map ^^^^^ା^^to an inversion operation,thereby transforming the map ^^^^^ା^^ into a corresponding complemented entropy map ^^^^^^ା^^.The inversion operation inverts the pixel values within the 0, 1 interval to produce lower values in the textured areas of the frame and highlight noisier areas of the frame, resulting in lower values in the textured areas and highlighting noisier areas of the frame, thereby aiding the denoising algorithm of the image denoising block 120 in preserving the details where it matters most. Collectively, the postprocessing operations 424 transform the set 422 of entropy maps ^^^^ା^^into a corresponding set 426 of complemented entropy maps ^^^^^^ା^^. In mathematical terms, the above-described example of the plurality of postprocessing operations 424 can be expressed as follows: ^^^^ା^^^^^,^^^ ൌ െ∑^ ^^൫^^ห^^௪^^^,^^^൯ log ^^^൫^^ห^^௪^^^, ^^^൯^ ∀ ^^ ∈ ^െ2,െ1,0,1,2^ (14)^^௧ା^వఱ ൌ ^^^^^^^^^^ ^^^^ 95th percentile of ^^^^ା^^ ∀ ^^ ∈ ^െ2,െ1,0,1,2^ (15)^^^^^ା^^^^^,^^^ ൌ ^^^^^^^^ ൬ ^^^^శ^^^௫,௬^ா^శೖ ,^^^^^^ ൌ 0,^^^^^^ ൌ 1^ ∀ ^^ ∈ ^െ2,െ1,0,1,2^ (16)వఱ^^^^^^ା^^^^^,^^^ ൌ 1.0 െ ^^^^^ା^^^^^,^^^ ∀ ^^ ∈ ^െ2,െ1,0,1,2^ (17)where ^^^^^^ is the probability of occurrence of intensity level ^^ in the local window ^^௪aroundpixel ^^^, ^^^ in frame ^^^^ା^^^^^,^^^ within the window.
[0048] FIG.6 pictorially illustrates the performance of the entropy filter 420 according to one example. More specifically, the entropy map ^^^^shown in FIG.6 is generated by the entropy filter 420 in response to the smoothened image frame ^^^^of FIG.5B.
[0049] FIGS.7A-7B pictorially illustrate some of the postprocessing operations 424 according to one example. More specifically, FIG.7A shows an example of the normalizedtexture information map ^^^^^ generated by the postprocessing operations corresponding to Eqs.(15)-(16) applied to the entropy map ^^^^shown in FIG.6. FIG.7B shows an example of the 10 complemented entropy map ^^^^^^by the postprocessing operations corresponding to Eq. (17)applied to the normalized texture information map ^^^^^ of FIG. 7A.
[0050] The complemented entropy maps ^^^^^^ା^^are applied to a logistic mapper 440. A corresponding output 442 of the logistic mapper 440 includes the logistic-scaled entropy maps ^^^^ା^^. The logistic mapping applied by the logistic mapper 440 aids in scaling and mapping the values from the entropy texture maps into values whose range is more suitable for the given sensitivity of the image denoising block 120 for efficient texture preservation. In some examples, the logistic curve used in the logistic mapper 440 is parameterized as follows: ^^ ^^^,^^^ ൌ ^^ ^ெ ^^ା^^^ା^ష್൫^^^^^^శ^^^^,^^ష^బ൯∀ ^^ ∈ ^െ2,െ1,0,1,2^ (18)where the curve; ^^^isthe x value setting the minimum value.
[0051] FIG.8 graphically illustrates a set 800 of different logistic curves that can be used in the logistic mapper 440 according to various examples. For illustration purposes, the set 800 is shown with x as a linearly spaced array in the range from ^10 to 10, consisting of 1000 points,with the midpoint of the curve being at ^^^ = 0.0, ^^ = 1.0, and the steepness ^^ ∈^0.5,1,1.5,2,3,5^.
[0052] FIG.9 graphically illustrates a logistic curve 902 that can be used in the logistic mapper 440 according to one example. The logistic curve 902 has its hyperparameters optimized on the shown set of entropy map data extracted from representative image frames. The resulting optimized hyperparameters for Eq. (18) are: ^^ = 0.95, ^^ = 0.1, ^^ = 20.7, ^^^^= 0.35. In operation, the logistic curve 902 is used to scale the complemented entropy maps ^^^^^^ା^^to obtain the corresponding entropy maps ^^^^ା^^.
[0053] FIG.10 pictorially illustrates the performance of the logistic mapper 440 according to one example. In the example shown, the logistic mapper 440 is configured to use the logistic curve 902 of FIG.9. The entropy map ^^^^shown in FIG.10 is generated by the logistic mapper 440 in response to the complemented entropy map ^^^^^^of FIG.7B.
[0054] For temporal smoothing and temporal stability of logistic-scaled entropy maps ^^^^ା^^, the filter arrangement 400 is configured to use the Lucas-Kanade method. For this purpose, the filter arrangement 400 employs an optical flow module 450 configured to extract flow information 452 corresponding to the maps ^^^^ା^^. In some examples, the optical flow module 450 operates to calculate the movement of object’s pixels using two consecutive frames ^^^^ା^^and based on the assumption that the pixel intensities of a moving object remain constant over the interframe time. Once the optical flow information 452 is determined, it is provided to a 11 motion compensation module 460 together with the set 442 of the logistic-scaled entropy maps ^^^^ା^^.
[0055] Example operations performed in the motion compensation module 460 include warping operations. The warping adjusts the pixels of subsequent entropy maps ^^^^ା^^based on the flow vectors computed based on the optical flow information 452. This adjustment effectively realigns each of the entropy maps ^^^^ା^^with respect to the target entropy map ^^^^. To further enhance stability and reduce temporal noise, the motion compensation module 460 is also configured to average the middle entropy map ^^^^with its motion-compensated neighbors ^^^^ା^^. This averaging operation combines the information from multiple temporally adjacent frames, thereby providing a smoothed representation 462, MC^^^^^,^^^, that maintains the core visual content while significantly reducing noise, which in turn allows for more accurate denoising inthe image denoising block 120.
[0056] Mathematical expressions for the above-described operations performed in the optical flow module 450 and the motion compensation module 460 are as follows: ^^^^ା^^→^^^^^,^^^ ൌ LucasKanade^^^^^ା^^,^^^^^^^^,^^^ for ^^ ∈ ^െ2,െ1,1,2^ (19)Warp^^ା^^→^^^^^,^^^ ൌ ^^௧ା^൫^^ ^ ^^^^ା^^→^^^^^,^^^௫,^^ ^ ^^^^ା^^→^^^^^,^^^௬൯ for ^^ ∈ ^െ2,െ1,1,2^ (20)^where ^^^^ା^^→^^^^^,^^^ is the optical flow vector from ^^^^ା^^^^^^ ^^^^; Warp^^^ା^^^→^^^^^,^^^ is the warpedimage with respect to the flow vector information; and MC^^^^^, ^^^462 of the motioncompensation module 460.
[0057] FIG.11 pictorially illustrates the performance of the optical flow module 450 and the motion compensation module 460 according to one example. In the example shown, the optical flow module 450 and the motion compensation module 460 operate on the set 442 in which the entropy map ^^^^shown in FIG.10 is the middle map. The motion compensated map MC^^shown in FIG.11 is the corresponding output map 462 of the motion compensation module 460.
[0058] FIGS.12A-12D pictorially illustrate the warping operations performed in the motion compensation module 460 according to one example. In the example shown, the motion compensation module 460 operates on the set 442 in which the entropy map ^^^^shown in FIG.10is the middle map. The resulting warped maps, Warp^^^ା^^^→^^ ∀ ^^ ∈ ^െ2,െ1,1,2^, are shown inFIGS.12A, 12B, 12C, and 12D, respectively.
[0059] The motion-compensated entropy map ^^^^^^462 generated by the motion compensation module 460 is applied to a self-guided filter 470 configured to use the map’s own structure for spatially coherent smoothing, which preserves essential edges while selectively 12 smoothing uniform areas of the map. This spatially coherent smoothing effectively reduces noise artifacts and enhances texture clarity, resulting in a refined noise strength estimate map ^^^^^^472 that captures subtle textural patterns of the input frame sequence 101. In mathematical terms, operations of the self-guided filter 470 can be represented as follows: ^^^^^^^^^,^^^ ൌ ^^^^^^^^^^^^,^^^; ^^, ^^ (22)where ^^^^^^^^^,^^^ is the output of the self-guided filter 470 at pixel ^^^, ^^^ corresponding to themotion compensated texture map ^^^^^^462; r is the radius of the neighborhood; and ^^ is the degree of smoothness. Both r and ^^ are hyperparameters that can be tuned for optimal filtering. For some examples, the optimal values of these hyperparameters are r = 32 and ^^ = 6.
[0060] FIG.13 pictorially illustrates the performance of the self-guided filter 470 according to one example. In the example shown, the self-guided filter 470 operates on the motion compensated map MC^^462 shown in FIG.11. The noise strength estimate map ^^^^^^shown in FIG.13 is the corresponding output map 472 generated by the self-guided filter 470.
[0061] In various cases, the above-described processing sequence implemented with the filter arrangement 400 beneficially ensures that the derived texture map, represented as the noise strength estimate map 472, is not only representative of the inherent texture in the frames of the window 101 but is also resilient to various motion artifacts.
[0062] FIG.14 is a block diagram illustrating the image denoising block 120 according to some examples. In the example shown, the image denoising block 120 includes a neural network 1400 having two stages, labeled 1401 and 1402, respectively. In other examples, a different (from two) number of stages can also be used. Both of the stages 1401 and 1402 are constructed using a plurality of denoising blocks 1410. More specifically, the first stage 1401 has three denoising blocks 1410, which are labeled 14101-14103, respectively. The second stage 1402 has one denoising block 1410, which is labeled 14104.
[0063] In some examples, the denoising blocks 14101-14103 of the first stage 1401 share the same weights, which reduces memory requirements and facilitates network training. The denoising blocks 14101-14103 are configured to processes the set {^^^^}^2 of noisy frames located within the sliding window 101 along with the middle frame’s noise strength map ^^^^received from the noise estimation block 110. The first stage 1401 operates to learn temporal information across the received noisy frames and noise characteristics from the noise strength map and outputs intermediate processing results for the second stage 1402 to process.
[0064] In the example shown, the second stage 1402 has one denoising block 14044 that has the same network architecture as the denoising blocks 14101-14103of the first stage 1401 but 13 its own trainable weights. The output of the denoising block 14044 is the denoised frame^^^^^corresponding to the middle frame of the input set {^^^^}^2.
[0065] FIG.15 is a block diagram illustrating a neural network 1500 used in a denoising block 1410n of FIG.14 according to one example. In the example shown, the neural network 1500 has a depth of sixteen convolutional layers, which are labeled in FIG.15 using the reference numerals 1501-1516. In other examples, a different (from sixteen) number of convolutional layers can alternatively be used in the neural network 1500.
[0066] The input convolutional layer 1501 has a channel depth equivalent to the number of frames in the sliding window 101 plus one (to accommodate the noise strength map ^^^^received from the noise estimation block 110). In this example, each frame has the size of (96^96) pixels, and there are five frames in the sliding window 101, thereby causing the input convolutional layer 1501 has a channel depth of six. The input convolutional layer 1501 transforms six channels into 90 channels. The second convolutional layers 1502 decreases the channel depth to 32. Subsequent convolutional layers 1503-1516 gradually increase the channel depth to 128 and then decrease it back down to one. Following each convolution operation of the layers 1501-1516, the ReLU activation function is used. In other embodiments, other suitable activation functions can also be used.
[0067] The neural network 1500 is designed using a modified U-Net framework with skip connections that keeps the spatial dimensions (width and height) of feature maps constant throughout, while adjusting their depth. At the beginning, the neural network 1500 increases the depth of feature maps similar to the downsampling branch of a U-Net but does not reduce the map’s spatial size. This characteristic enables the neural network 1500 to preserve the spatial context and avoids checkerboard artifacts. Thereafter, the neural network 1500 decreases the depth, much like the upsampling branch of U-Net, to refine the features. This architecture effectively combines the depth variation benefits of a U-Net with consistent spatial dimensions, making the resulting model more suitable for grain-noise removal tasks.
[0068] FIG.16 is a flowchart of a method 1600 of denoising a video sequence that can be implemented with the NR apparatus 100 according to some examples. The method 1600 includes the NR apparatus 100 generating a noise strength map based on a set of consecutive video frames of the video sequence (in a block 1602). The generated noise strength map represents an estimate of image noise (e.g., of film-grain noise) for a middle frame of the frame set located within the window 101. In some examples, operations of the block 1602 are performed using the noise estimation block 110 of the NR apparatus 100. 14
[0069] The method 1600 also includes the NR apparatus 100 producing a denoised frame using the neural network 1400 (in a block 1604). The neural network 1400 is trained, e.g., as described in more detail in the next section, to remove image noise from the middle frame of the frame set located within the window 101. The inputs to the neural network 1400 include the frame set located within the window 101 and the noise strength map generated in the block 1602.In response to these inputs, the neural network 1400 generates and outputs the denoised frame ^^^^^ .
[0070] The method 1600 also includes the NR apparatus 100 determining whether the end of the video sequence is reached (in a decision block 1606). When it is determined that the end of the video sequence is not yet reached (“No” at the decision block 1606), the processing of the method 1600 is directed to a block 1608. When it is determined that the end of the video sequence has been reached (“Yes” at the decision block 1606), the method 1600 is terminated.
[0071] Operations of the block 1608 include the NR apparatus 100 moving the sliding window 101 to a next position along the video sequence by a fixed increment. In some examples, the increment is one frame. In some other examples, other increment values can also be used. Upon advancing the window 101 along the video sequence by the fixed increment, the processing of the method 1600 is looped back to the block 1602 to repeat the operations thereof and then also the operations of the block 1604 for this next position of the sliding window 101. Training
[0072] In one example, the NR apparatus 100 can be trained using the DAVIS (Densely Annotated Video Segmentation) dataset. DAVIS is a collection of high-quality, high-resolution video sequences for a diverse range of challenges in video processing. These challenges include but are not limited to variations in object occlusions, fast motion dynamics, and the intricate details of object shapes and sizes.
[0073] In one example, the training dataset includes approximately 150 video sequences with an average frame resolution of 480 × 854. Each video sequence has a duration of at least 30 seconds, providing rich dataset for training the above-described models. In some examples, we focused on using generated grayscale (Y channel) video sequences from DAVIS. Synthesizing film grain noisy dataset from the clean Y channel data is described above in the Mathematical Formulation section.
[0074] In some examples, the neural network of the NR apparatus 100 is fed with the above-described dataset(s), with assistance of NVIDIA Data Loading Library (DALI), which provides a powerful and efficient tool that facilitates GPU-accelerated data loading and augmentation, ensuring streamlined operations and reduced data preprocessing overhead. A total 15 of ^^=256000 training samples are extracted from the training set of the DAVIS dataset. The spatial size of the patches is 96×96, while the temporal size is 2^ +1=5.
[0075] In some examples, we synthesize film-grain noise for each clean image patch with a spatial dimension of 96×96 as described above. In some examples, the following film-grain synthesis parameters are used: ^^^=2, ^^^=7, ^^^=16. The ^^^^^^^^^^’s parameters are chosen as ^^^^^=16, ^^^^௫ = 235 as ^^ follows the narrow range, ^^^^^ = 0.0, ^^^^௫ = 1.0 and ൫^^^^௩^௧, ^^^^௫൯ =^110, 1.0^ that form a triangle LUT function with its peak at mid-intensity. We selected a γvalue uniformly distributed between 0 and 50 at each training step.
[0076] As described in the Mathematical Formulation section, we compute the noisy image patch ^^^^and the ground truth noise strength map ^^^^with the functions defined and the parameters described above. It should be noted that these parameters can be randomized for on- the-fly film grain noise generation.
[0077] In some examples, for training the neural network(s), we train both the noise estimation block 110 and the image denoising block 120 together as a whole network that learns on the combined loss function described below.
[0078] In some examples, for efficient neural-network training, we define the loss function therefor as a combination of two Mean Squared Error (MSE) losses. The first MSE loss ensures that the image denoising block 120 recovers the original image by minimizing the difference between the mid patch of the noisy patch sequence and its clean counterpart. The second MSE loss aims to train the noise estimation block 110 to estimate the per pixel noise strength accurately by comparing the predicted noise strength to the true noise strength added during the training. In mathematical terms, the combined loss function, ℒ, is expressed as follows: ℒൌ β ∗ ^^^^^^ ൫^^^^^ , ^^^^ ൯ ^ α ∗ ^^^^^^ ൫^^^^^ ,^^௧൯ (23)where: ^ ^^,^^ are hyperparameters that weight the two losses in the combined loss function (^^=10, ^^=1 are found to be nearly optimal for the above-described purposes). ^^^^^^ is the predicted clean image patch obtained from the image denoising block 120.^ ^^^^is the clean image patch of the noisy patch sequence. ^ ^^^^is the true film-grain noise strength map of the added film-grain noise. 16 ^ ^^^^^ is the film-grain noise strength map predicted by the noise estimation block 110.^ ^^^^ , ^^^^^ are ground truth and the predicted values, respectively.^ ^^௧is the number of observations in the current batch. ^ θ represents the collective set of parameters from both the noise estimator network (^^ோ) and the denoising network (^^^ா) ^ ∇ఏdenotes the gradient with respect to the neural network parameters ^^. ^ ^^ is the learning rate that controls the step size in the gradient descent. In at least some examples, the above combined loss function ℒ ensures both the recovery of the original image and an accurate noise strength estimation, thereby enabling our model to provide high-quality, temporally stable denoised images while accurately predicting the noise levels. Example Hardware
[0079] FIG.17 is a block diagram of an example computing device 1700 according to various examples. In some examples, the computing device 1700 is configured to perform at least some operations of the method 1600. In some examples, one or more instances of the computing device 1700 are used in apparatus 100.
[0080] The computing device 1700 of FIG.17 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 1700 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices 1702 and one or more storage devices 1704). Additionally, in various embodiments, the computing device 1700 may not include one or more of the components illustrated in FIG.17, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device 1700 may not include a display device 1710, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device 1710 may be coupled.
[0081] The computing device 1700 includes a processing device 1702 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may 17 be stored in registers and / or memory. In various embodiments, the processing device 1702 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.
[0082] The computing device 1700 also includes a storage device 1704 (e.g., one or more storage devices). In various embodiments, the storage device 1704 may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 1704 may include memory that shares a die with the processing device 1702. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device 1704 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device 1702), cause the computing device 1700 to perform any appropriate ones of the methods disclosed herein below or portions of such methods.
[0083] The computing device 1700 further includes an interface device 1706 (e.g., one or more interface devices 1706). In various embodiments, the interface device 1706 may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device 1700 and other computing devices. For example, the interface device 1706 may include circuitry for managing wireless communications for the transfer of data to and from the computing device 1700. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 1706 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with a Global System for Mobile 18 Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E- HSPA), or LTE network. In some embodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 1706 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.
[0084] In some embodiments, the interface device 1706 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 1706 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 1706 may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 1706 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device 1706 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device 1706 may be dedicated to wireless communications, and a second set of circuitry of the interface device 1706 may be dedicated to wired communications.
[0085] The computing device 1700 also includes battery / power circuitry 1708. In various embodiments, the battery / power circuitry 1708 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 1700 to an energy source separate from the computing device 1700 (e.g., to AC line power).
[0086] The computing device 1700 also includes a display device 1710 (e.g., one or multiple individual display devices). In various embodiments, the display device 1710 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a 19 touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.
[0087] The computing device 1700 also includes additional input / output (I / O) devices 1712. In various embodiments, the I / O devices 1712 may include one or more data / signal transfer interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc.
[0088] Depending on the specific embodiment, various components of the interface devices 1706 and / or I / O devices 1712 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices 1706 and / or I / O devices 1712 include one or more analog-to- digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 1702 and / or the storage device 1704. In some additional examples, the interface devices 1706 and / or I / O devices 1712 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 1702 and / or the storage device 1704 into an analog form suitable for being transmitted through a communication channel.
[0089] According to an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGs.1-17, provided is apparatus for denoising a video sequence, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: generate a noise strength map based on a set of consecutive video frames of the video sequence, the noise strength map representing an estimate of image noise for a middle frame of the set; and produce a denoised frame using a first neural network trained to remove image noise from the middle frame of the set in response to the set and the noise strength map being applied thereto as inputs.
[0090] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGs.1-17, provided is a method of denoising a video sequence, the method comprising: generating a noise strength map based on a set of consecutive video frames of the video sequence, the noise strength map representing an estimate of image noise for a middle frame of the set; and producing a denoised 20 frame using a first neural network trained to remove image noise from the middle frame of the set in response to the set and the noise strength map being applied thereto as inputs.
[0091] In some embodiments of the above method, the image noise includes film-grain noise.
[0092] In some embodiments of any of the above methods, the first neural network has been trained using a loss function configured drive the first neural network toward a configuration in which a difference between a denoised frame of a training set and a corresponding ground-truth noise-free frame is substantially minimized.
[0093] In some embodiments of any of the above methods, the method further comprises: generating the film-grain noise for the training set by applying an inverse discrete cosine transform (DCT) to a set of DCT coefficients corresponding to a selected film-grain block size; and representing a subset of alternating current (AC) coefficients of the set of DCT coefficients with random numbers.
[0094] In some embodiments of any of the above methods, the first neural network includes: a first stage having a plurality of denoising blocks connected in parallel to one another; and a second stage serially connected with and downstream from the first stage, the second stage having at least one denoising block, wherein the plurality of denoising blocks and the at least one denoising block are of a same architecture.
[0095] In some embodiments of any of the above methods, the same architecture is such that a denoising block of said architecture includes a sequence of convolution layers of varying channel depth.
[0096] In some embodiments of any of the above methods, the first neural network includes: a first stage having a first plurality of denoising blocks connected in parallel to one another; and a second stage serially connected with and downstream from the first stage, the second stage having at least one denoising block, wherein the first plurality of denoising blocks and the at least one denoising block are of a same architecture.
[0097] In some embodiments of any of the above methods, the second stage has a second plurality of denoising blocks connected in parallel to one another; wherein the first neural network includes a third stage serially connected with and downstream from the first and second stages, the third stage having at least one denoising block; wherein the second stage has fewer denoising blocks than the first stage; and wherein the third stage has fewer denoising blocks than the second stage.
[0098] In some embodiments of any of the above methods, the sequence of convolution layers includes a first convolution layer, a second convolution layer, and a third convolution layer, the second convolution layer being between the first convolution layer and the third 21 convolution layer in the sequence of convolution layers; wherein the second convolution layer has a larger channel depth than the first convolution layer; and wherein the second convolution layer has a larger channel depth than the third convolution layer.
[0099] In some embodiments of any of the above methods, each of the plurality of denoising blocks is configured to receive a different respective subset of the set and a respective copy of the noise strength map as inputs.
[0100] In some embodiments of any of the above methods, the generating includes using a second neural network trained to generate the noise strength map in response to the set being applied thereto as an input.
[0101] In some embodiments of any of the above methods, the second neural network includes a sequence of convolution layers of varying channel depth.
[0102] In some embodiments of any of the above methods, the first and second neural networks have been trained together using a loss function including a weighted sum of a first mean squared error (MSE) loss and second MSE loss; wherein the first MSE loss is configured to drive the first neural network toward a configuration in which a difference between a denoised frame of a training set and a corresponding ground-truth noise-free frame is substantially minimized; and wherein the second MSE loss is configured to drive the second neural network toward a configuration in which a difference between an estimate of the film-grain noise for a middle frame of the training set and a corresponding ground-truth film-grain noise is substantially minimized.
[0103] In some embodiments of any of the above methods, the first and second neural networks have been trained together using a loss function including a weighted sum of a first loss and second loss; wherein the first loss is configured to drive the first neural network toward a configuration in which a difference between a denoised frame of a training set and a corresponding ground-truth noise-free frame is substantially minimized; and wherein the second loss is configured to drive the second neural network toward a configuration in which a difference between an estimate of the image noise for a middle frame of the training set and a corresponding ground-truth image noise is substantially minimized.
[0104] In some embodiments of any of the above methods, each of the first loss and the second loss is selected from the group consisting of: a mean squared error (MSE) loss; a feature- based loss; a Visual Geometry Group (VGG) loss; and learned perceptual image patch similarity (LPIPS) loss.
[0105] In some embodiments of any of the above methods, the set includes at least five consecutive video frames of the video sequence. 22
[0106] In some embodiments of any of the above methods, the set includes a plurality of consecutive video frames of the video sequence located within a window of a fixed width; and wherein the method further comprises: moving the window to a next position along the video sequence; and repeating the generating and the producing for the next position.
[0107] In some embodiments of any of the above methods, the generating includes using a filter arrangement to generate the noise strength map in response to the set being applied thereto as an input, the filter arrangement including an entropy filter.
[0108] In some embodiments of any of the above methods, the filter arrangement further includes a logistic mapper and a motion compensation module.
[0109] In some embodiments of any of the above methods, the filter arrangement further includes a Gaussian filter and self-guided filter.
[0110] In some embodiments of any of the above methods, the generating includes: subjecting the set of consecutive video frames to Gaussian filtering to generate a Gaussian- filtered set of frames; applying entropy filtering to the Gaussian-filtered set of frames to generate a set of entropy maps; applying one or more postprocessing operations to the set of entropy maps to generate a set of postprocessed entropy maps; applying logistic mapping to the set of postprocessed entropy maps to generate a set of scaled entropy maps; generating a single motion-compensated entropy map using motion compensation applied to the set of scaled entropy maps; and passing the single motion-compensated entropy map through a self-guided filter to generate the noise strength map.
[0111] In some embodiments of any of the above methods, the one or more postprocessing operations include one or more operations selected from the group consisting of normalization, removal of outliers, clipping, and computing complementing values within a [0, 1] interval.
[0112] In some embodiments of any of the above methods, the generating further includes using a Lucas-Kanade method to determine optical flow information corresponding to the set of scaled entropy maps; and wherein said motion compensation includes warping at least some of the scaled entropy maps to a projection corresponding to the middle frame based on the determined optical flow information.
[0113] In some embodiments of any of the above methods, the method further comprises selecting hyperparameters for the logistic mapping based on entropy map data corresponding to a training set of frames.
[0114] In some embodiments of any of the above methods, said generating the single motion-compensated entropy map includes averaging the set of motion-compensated scaled entropy maps. 23
[0115] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any of the above methods.
[0116] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.
[0117] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0118] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.
[0119] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter. 24
[0120] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.
[0121] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.
[0122] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.
[0123] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.
[0124] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.
[0125] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.
[0126] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same 25 embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”
[0127] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.
[0128] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”
[0129] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.
[0130] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.
[0131] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile 26 storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
[0132] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0133] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.
[0134] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. 27
Claims
CLAIMS What is claimed is:
1. A method of denoising a video sequence, the method comprising: generating a noise strength map based on a set of consecutive video frames of the video sequence, the noise strength map representing an estimate of image noise for a middle frame of the set; and producing a denoised frame using a first neural network trained to remove image noise from the middle frame of the set in response to the set and the noise strength map being applied thereto as inputs.
2. The method of claim 1, wherein the image noise includes film-grain noise.
3. The method of claim 2, wherein the first neural network has been trained using a loss function configured to drive the first neural network toward a configuration in which a difference between a denoised frame of a training set and a corresponding ground-truth noise-free frame is substantially minimized.
4. The method of claim 3, further comprising: generating the film-grain noise for the training set by applying an inverse discrete cosine transform (DCT) to a set of DCT coefficients corresponding to a selected film-grain block size; and representing a subset of alternating current (AC) coefficients of the set of DCT coefficients with random numbers.
5. The method of any preceding claim, wherein the first neural network includes: a first stage having a first plurality of denoising blocks connected in parallel to one another; and a second stage serially connected with and downstream from the first stage, the second stage having at least one denoising block, wherein the first plurality of denoising blocks and the at least one denoising block are of a same architecture.
6. The method of claim 5, 28 wherein the second stage has a second plurality of denoising blocks connected in parallel to one another; wherein the first neural network includes a third stage serially connected with and downstream from the first and second stages, the third stage having at least one denoising block; wherein the second stage has fewer denoising blocks than the first stage; and wherein the third stage has fewer denoising blocks than the second stage.
7. The method of claim 5 or 6, wherein the same architecture is such that a denoising block of said architecture includes a sequence of convolution layers of varying channel depth.
8. The method of claim 7, wherein the same architecture is a U-Net architecture.
9. The method of claim 7, wherein the sequence of convolution layers includes a first convolution layer, a second convolution layer, and a third convolution layer, the second convolution layer being between the first convolution layer and the third convolution layer in the sequence of convolution layers; wherein the second convolution layer has a larger channel depth than the first convolution layer; and wherein the second convolution layer has a larger channel depth than the third convolution layer.
10. The method of any one of claims 5 to 9, wherein each of the plurality of denoising blocks is configured to receive a different respective subset of the set and a respective copy of the noise strength map as inputs.
11. The method of any preceding claim, wherein the generating includes using a second neural network trained to generate the noise strength map in response to the set being applied thereto as an input.
12. The method of claim 11, wherein the second neural network includes a sequence of convolution layers of varying channel depth.
13. The method of claim 11 or 12, 29 wherein the first and second neural networks have been trained together using a loss function including a weighted sum of a first loss and a second loss; wherein the first loss is configured to drive the first neural network toward a configuration in which a difference between a denoised frame of a training set and a corresponding ground-truth noise-free frame is substantially minimized; and wherein the second loss is configured to drive the second neural network toward a configuration in which a difference between an estimate of the image noise for a middle frame of the training set and a corresponding ground-truth image noise is substantially minimized.
14. The method of claim 13, wherein each of the first loss and the second loss is selected from the group consisting of: a mean squared error (MSE) loss; a feature-based loss; a Visual Geometry Group (VGG) loss; and learned perceptual image patch similarity (LPIPS) loss.
15. The method of any preceding claim, wherein the set includes at least five consecutive video frames of the video sequence.
16. The method of any preceding claim, wherein the set includes a plurality of consecutive video frames of the video sequence located within a window of a fixed width; and wherein the method further comprises: moving the window to a next position along the video sequence; and repeating the generating and the producing for the next position.
17. The method of any preceding claim, wherein the generating includes using a filter arrangement to generate the noise strength map in response to the set being applied thereto as an input, the filter arrangement including an entropy filter.
18. The method of claim 17, wherein the filter arrangement further includes a logistic mapper and a motion compensation module.
19. The method of claim 18, wherein the filter arrangement further includes a Gaussian filter and self-guided filter. 30 20. The method of any preceding claim, wherein the generating includes: subjecting the set of consecutive video frames to Gaussian filtering to generate a Gaussian-filtered set of frames; applying entropy filtering to the Gaussian-filtered set of frames to generate a set of entropy maps; applying one or more postprocessing operations to the set of entropy maps to generate a set of postprocessed entropy maps; applying logistic mapping to the set of postprocessed entropy maps to generate a set of scaled entropy maps; generating a single motion-compensated entropy map using motion compensation applied to the set of scaled entropy maps; and passing the single motion-compensated entropy map through a self-guided filter to generate the noise strength map.
21. The method of claim 20, wherein the one or more postprocessing operations include one or more operations selected from the group consisting of normalization, removal of outliers, clipping, and computing complementing values within a [0, 1] interval.
22. The method of claim 20 or 21, wherein the generating further includes using a Lucas-Kanade method to determine optical flow information corresponding to the set of scaled entropy maps; and wherein said motion compensation includes warping at least some of the scaled entropy maps to a projection corresponding to the middle frame based on the determined optical flow information.
23. The method of any one of claims 20 to 22, further comprising selecting hyperparameters for the logistic mapping based on entropy map data corresponding to a training set of frames.
24. The method of any one of claims 20 to 23, wherein said generating the single motion- compensated entropy map includes averaging the set of motion-compensated scaled entropy maps. 31 25. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any one of claims 1 to 24.
26. An apparatus for denoising a video sequence, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: generate a noise strength map based on a set of consecutive video frames of the video sequence, the noise strength map representing an estimate of image noise for a middle frame of the set; and produce a denoised frame using a first neural network trained to remove image noise from the middle frame of the set in response to the set and the noise strength map being applied thereto as inputs. 32