Metadata-assisted film grain removal

Metadata-assisted film grain removal techniques in video decoders effectively address the challenge of encoding quasi-random film grain, ensuring accurate removal and improved playback quality through spatial and temporal domain processing.

JP7819354B2Active Publication Date: 2026-02-24DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024558973
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-19
Filing Date
2023-04-18
Publication Date
2026-02-24
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

Digital film grain is difficult to encode due to its quasi-random nature, leading to incomplete removal during video encoding and subsequent synthesis, which results in suboptimal playback quality.

Method used

Metadata-assisted film grain removal methods are employed in video decoders to accurately remove film grain from compressed video signals using spatial and temporal domain processing, leveraging metadata provided by the encoder to estimate and subtract film grain components from decompressed images.

Benefits of technology

The method enables complete removal of film grain from digitally compressed video, improving playback quality by reducing residual noise and enhancing aesthetic appearance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007819354000057
    Figure 0007819354000057
  • Figure 0007819354000058
    Figure 0007819354000058
  • Figure 0007819354000059
    Figure 0007819354000059
Patent Text Reader

Abstract

Metadata-assisted film grain removal method and corresponding apparatus. Certain exemplary embodiments enable a video decoder to substantially completely remove film grain from a digital video signal that has undergone lossy video compression and then video decompression. Different embodiments may rely on only spatial domain grain removal processing, only time domain grain removal processing, or a combination of spatial and time domain grain removal processing. Both spatial and time domain grain removal processing may use metadata provided by a corresponding video encoder, the metadata including one or more parameters corresponding to digital film grain injected into the host video at the encoder. In order for different film grain injection formats to be accepted by the video decoder, signal pre-processing is used that is directed to providing an input to the film grain removal module of the video decoder that is compatible with the film grain removal method implemented therein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] 1. CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority from the following priority applications: U.S. Provisional Patent Application No. 63 / 332,332 (Docket No. D22003USP01), filed April 19, 2022, and European Patent Application No. 22168812.0 (Docket No. D22003EP), filed April 19, 2022, each of which is incorporated herein by reference in its entirety.

[0002] 2. Areas of Disclosure Various exemplary embodiments relate to image and video processing, and more particularly, but not exclusively, to digital film grain techniques. [Background technology]

[0003] 3. background On physical film, film grain is the random physical texture created by tiny metallic silver particles found on processed photographic celluloid. In digital photography, the visual and artistic effect of film grain can be simulated by adding a digital grain pattern to a digital image after the image is captured. Because digital film grain can be difficult to encode, for example, due to its (quasi-)random nature, video encoders typically remove the film grain during encoding; then, during playback, a corresponding video decoder may synthesize the film grain and add it back. Summary of the Invention [Means for solving the problem]

[0004] Various embodiments of metadata-assisted film grain removal methods and corresponding apparatuses are disclosed herein. One exemplary embodiment enables a video decoder to substantially completely remove film grain from a digital video signal that has undergone lossy video compression and subsequent video decompression. Different embodiments may rely solely on spatial-domain grain removal processing, solely on temporal-domain grain removal processing, or on a combination of spatial and temporal-domain grain removal processing. Both spatial-domain and temporal-domain grain removal processing may use metadata provided by a corresponding video encoder, the metadata including one or more parameters corresponding to the digital film grain injected into the host video at the encoder. To enable different film grain injection formats to be accepted by a video decoder, signal preprocessing is used to provide the video decoder's film grain removal module with input compatible with the film grain removal method implemented therein.

[0005] According to an example embodiment, there is provided a video delivery system capable of film grain removal, comprising: an input interface that receives an encoded bitstream comprising a compressed video bitstream and metadata, wherein the compressed video bitstream has encoded therein a sequence of images comprising digital film grain, the metadata including one or more parameters corresponding to the digital film grain; and a processor configured to: decompress the compressed video bitstream to generate a respective decompressed representation of each image, each decompressed representation comprising a respective host image component and a respective film grain component; calculate, for each decompressed representation based on the metadata, a respective estimate of the respective film grain component; and remove the respective estimate from each decompressed representation to generate a corresponding estimate of the respective host image component.

[0006] According to another exemplary embodiment, there is provided a video delivery system capable of film grain removal, comprising: a video decoder comprising: an input interface that receives an encoded bitstream including a compressed video bitstream and metadata, wherein the compressed video bitstream has encoded therein a sequence of images including digital film grain, and the metadata includes one or more parameters corresponding to the digital film grain; and a processor configured to perform the following steps: decompressing the compressed video bitstream to generate a respective decompressed representation of each image, each decompressed representation including a respective host image component and a respective film grain component; and temporally averaging the sequence of the respective decompressed representations using a plurality of temporal sliding windows, each temporal sliding window corresponding to a respective image frame pixel, at least some of the temporal sliding windows having different respective lengths selected based on the metadata.

[0007] According to yet another exemplary embodiment, there is provided a machine-implemented method for removing film grain from video data, the method including: receiving an encoded bitstream including a compressed video bitstream and metadata, wherein the compressed video bitstream has encoded therein a sequence of images including digital film grain, the metadata including one or more parameters corresponding to the digital film grain; decompressing the compressed video bitstream to generate a respective decompressed representation of each image, each decompressed representation including a respective host image component and a respective film grain component; calculating, for each decompressed representation based on the metadata, a respective estimate of the respective film grain component; and removing the respective estimate from each decompressed representation to generate a corresponding estimate of the respective host image component.

[0008] According to yet another exemplary embodiment, a non-transitory machine-readable medium is provided having program code encoded thereon, the program code, when executed by the machine, causing the machine to perform a method comprising: receiving an encoded bitstream comprising a compressed video bitstream and metadata, the compressed video bitstream having encoded therein a sequence of images comprising digital film grain, the metadata comprising one or more parameters corresponding to the digital film grain; decompressing the compressed video bitstream to generate a respective decompressed representation of each image, each decompressed representation comprising a respective host image component and a respective film grain component; calculating, for each decompressed representation based on the metadata, a respective estimate of the respective film grain component; and removing the respective estimate from each decompressed representation to generate a corresponding estimate of the respective host image component. [Brief explanation of the drawings]

[0009] Other aspects, features, and advantages of the various disclosed embodiments will become more fully apparent from the following detailed description and the accompanying drawings, by way of example.

[0010] [Figure 1] 1 illustrates an example process for a video delivery pipeline in which at least some embodiments may be implemented.

[0011] [Figure 2] 2 is a block diagram illustrating a video encoder that can be used in the process of FIG. 1 according to one embodiment.

[0012] [Figure 3] 2 is a block diagram illustrating a video decoder that can be used in the process of FIG. 1 according to one embodiment.

[0013] [Figure 4] 4 is a process flow diagram illustrating a spatial domain grain removal process that can be implemented in the video decoder of FIG. 3 according to one embodiment.

[0014] [Figure 5] 4 graphically illustrates exemplary film grain removal results that may be achievable using a scaled Gaussian filter in the video decoder of FIG. 3 according to first embodiments.

[0015] [Figure 6] 4 graphically illustrates exemplary film grain removal results that may be achievable using a scaled Gaussian filter in the video decoder of FIG. 3, according to second embodiments;

[0016] [Figure 7]4 graphically illustrates exemplary film grain removal results that may be achievable using a scaled Gaussian filter in the video decoder of FIG. 3, according to third embodiments;

[0017] [Figure 8] 10 graphically illustrates exemplary film grain removal results that may be achievable using a scaled Gaussian filter in the video decoder of FIG. 3, according to fourth embodiments.

[0018] [Figure 9] 4 is a process flow diagram illustrating a time-domain grain removal process that can be implemented in the video decoder of FIG. 3 according to an embodiment.

[0019] [Figure 10] 4 illustrates an example of an intensity-dependent scaling factor that can be accommodated in the video decoder of FIG. 3 using signal pre-processing, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0020] The present disclosure and aspects thereof may be embodied in various forms, including hardware, devices or circuits controlled by computer-implemented methods, computer program products, computer systems and networks, user interfaces and application programming interfaces, as well as hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. The foregoing is intended only to give a general idea of ​​various aspects of the present disclosure and is not intended to limit the scope of the disclosure in any way.

[0021] In the following description, numerous details are set forth, such as optical device configurations, timing, operation, etc., to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to those skilled in the art that these specific details are merely examples and are not intended to limit the scope of the present application.

[0022] Additionally, while this disclosure primarily focuses on examples in which various circuits are used in digital projection systems, it will be understood that these are merely examples. It will be further understood that the disclosed systems and methods can be used in any device that needs to project light, such as cinema, consumer, and other commercial projection systems, heads-up displays, virtual reality displays, etc. The disclosed systems and methods may also be implemented in additional display devices, including OLED displays, LCD displays, quantum dot displays, etc.

[0023] As used herein, the term "dynamic range" (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) in an image, e.g., from darkest gray (black) to brightest white (highlight). In this sense, DR relates to "scene-referenced" intensities. DR may also relate to the ability of a display device to adequately and / or approximately render an intensity range of a particular width. In this sense, DR relates to "display-referenced" intensities. Unless a particular meaning is explicitly specified as having particular importance at any point in the description herein, it should be presumed that the terms may be used in either sense, e.g., interchangeably.

[0024] As used herein, the term "high dynamic range" (HDR) refers to a DR width spanning approximately 14 to 15 orders of magnitude of the human visual system (HVS). In practice, the DR at which humans can simultaneously perceive a wide range of intensities may be somewhat truncated relative to HDR. As used herein, the terms "enhanced dynamic range" (EDR) or "visual dynamic range" (VDR), individually or interchangeably, may refer to the DR perceivable in a scene or image by the human visual system, including eye movements, taking into account any light adaptation changes across the scene or image.

[0025] In practice, an image includes one or more color components (e.g., luma Y and chroma Cb and Cr), each represented with n bits of precision per pixel (e.g., n=8). Using linear luminance encoding, images with n≦8 may be considered standard dynamic range images, while images with n>8 (e.g., color 24-bit JPEG images) may be considered enhanced dynamic range images. EDR and HDR images may also be stored and distributed using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR file format developed by Industrial Light and Magic.

[0026] As used herein, the term "metadata" refers to any auxiliary information transmitted as part of an encoded bitstream that assists a decoder in rendering the corresponding image(s). For television broadcasting and video streaming, video metadata may be used to provide side information about specific video and audio streams or files. The metadata is either embedded directly in the video or included as a separate file within a container such as MP4 or MKV. The metadata may include information about the entire video stream or file, or about specific video frames. Metadata created by cameras, encoders, and other video processing elements (e.g., see 115, 120 in Figure 1) may include, but is not limited to, timestamps, video resolution, digital film grain parameters, color space or gamut information, reference display parameters, auxiliary signal parameters, file size, closed captions, audio language, ad insertion points, color spaces, error messages, etc. Additional examples of metadata relevant to the disclosed embodiments are described herein below.

[0027] Many consumer desktop displays have a brightness of 200-300 cd / m 2 Many consumer HDTVs are in the 300-500 nits range, with newer models supporting 1000 nits (cd / m 2 ). Such conventional displays, with respect to HDR or EDR, are representative of low dynamic range (LDR), also referred to as standard dynamic range (SDR). As the availability of HDR content increases due to advances in both image capture devices (e.g., cameras) and HDR displays (e.g., the PRM-4200 Professional Reference Monitor from Dolby Laboratories), HDR content can be color graded and displayed on HDR displays that support a higher dynamic range (e.g., 1,000 nits to 5,000 nits or more).

[0028] Video Encoding According to an Exemplary Embodiment FIG. 1 illustrates an exemplary process of a video delivery pipeline (100) showing various stages from video capture to video content display, according to one embodiment. A sequence of video frames (102) may be captured or generated using an image generation block (105). The video frames (102) may be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) may be captured on film by a film camera. The film may then be converted to a digital format (107) to provide the video data.

[0029] In the production phase (110), the video data (107) may be edited to provide a video production stream (112). The data in the video production stream (112) may then be provided to a processor (or one or more processors, such as a central processing unit (CPU)) in a post-production block (115) for post-production editing. The post-production editing in block (115) may include, for example, adjusting or modifying the color or brightness in specific areas of the image to improve image quality or achieve a particular look for the image according to the video creator's creative intent. This part of post-production editing is sometimes referred to as "color timing" or "color grading." Other editing (e.g., scene selection and ordering, image cropping, adding computer-generated visual special effects, etc.) may be performed in block (115) to generate a "final" version (117) of the production for distribution. During post-production editing (115), the video image may be viewed on a reference display (125).

[0030] Following post-production (115), the final version (117) of the video data may be delivered to an encoding block (120) for delivery to downstream decoding and playback devices, such as television sets, set-top boxes, movie theaters, etc. In some embodiments, the encoding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-Ray, and other delivery formats, to generate an encoded bitstream (122). Some methods described herein below may be performed by corresponding processors in the encoding block (120). At the receiver, the encoded bitstream (122) is decoded by a decoding unit (130) to generate a corresponding decoded signal (132) that represents a copy or approximation of the signal (117). The receiver may be attached to a target display (140), which may have characteristics that are slightly or completely different from the reference display (125). In such cases, a display management block (135) may be used to map the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137). Some methods described herein below may be performed by the decode unit (130) and / or the display management block (135). Depending on the embodiment, the decode unit (130) and the display management block (135) may include individual processors or may be based on a single integrated processing unit.

[0031] As already mentioned above, film grain can provide an effective way to improve the appearance of digital video in terms of aesthetics and sharpness. Film grain can also be used to reduce banding artifacts and / or mask compression artifacts. Film grain may be added in post-production (115). Traditionally, film grain is removed in an encoder in a coding block (120), and the resulting "clean" video signal is compressed, for example, using a lossy video compression code, to generate a coded bitstream (122). In a corresponding decoder, the coded bitstream (122) is decoded, and film grain is added using a decoding unit (130) to generate a decoded signal (132) that includes film grain.

[0032] In some use cases, film grain may not be removed at the encoder in the encoding block (120), or alternatively, basic film grain may be added there to a "clean" video signal before video compression. This type of processing may be performed, for example, to ensure that film grain features are present in the encoded bitstream (122) and allow "simpler" decoders to provide playback including film grain, even if the decoder itself is not capable of synthesizing and overlaying film grain. However, if the decoder has film grain synthesis capabilities, the basic film grain in the encoded bitstream (122) may be removed at the decoder using the decode unit (130), where further video processing may be applied. In some use cases, such further video processing may include adding another (e.g., adaptive) type of film grain according to the viewing environment corresponding to the target display (140).

[0033] In one example of the above use case, the base layer of the bitstream (122) carries an SDR signal with SDR-specific film grain injected therein. If the received bitstream is to be converted to HDR, the SDR-specific film grain may be removed in the decoding unit (130) and replaced by a different (e.g., HDR-specific) film grain.

[0034] In another example of the above use case, the bitstream (122) may have film grain injected therein to reduce the banding artifacts mentioned above. If the decoding unit (130) operates to convert an 8-bit based profile to a 10-bit based profile, the corresponding processing may include removing the existing film grain and then applying a de-banding filter. The resulting filtered signal may be further processed to, among other things, insert different film grain.

[0035] Exemplary embodiments disclosed herein are generally directed to performing film grain removal from a received bitstream (122) in a decode unit (130). Some embodiments may rely solely on spatial domain grain removal processing or solely on temporal domain grain removal processing. Some other embodiments may rely on both spatial and temporal domain grain removal processing. In exemplary embodiments, the grain removal processing is performed using metadata from the received bitstream (122). The metadata may provide useful information about the film grain in the received bitstream, thereby facilitating the grain removal processing in the decode unit (130).

[0036] In one example embodiment, the pipeline blocks (115, 120) can generate film grain for the encoded bitstream (122) using a reproducible film grain model, where a random or pseudo-random seed and a selected intensity modulation function may be used. The seed information and film grain parameters of the model may be transmitted as metadata in the bitstream 122. As a result, the decoder in the decoding unit (130) is provided with certain details of the stream's film grain pattern, which the decoder can beneficially utilize to efficiently generate an accurate approximation of a corresponding "clean" (i.e., grain-free) video signal, for example, as described in more detail below.

[0037] Figure 2 is a block diagram illustrating a video encoder (200) according to one embodiment. The encoder (200) may be used, for example, to implement the encoding block (120) of the pipeline (100) (see also Figure 1). For illustrative purposes, and without implying any limitation, Figure 2 uses reference numerals (117) and (122) to better illustrate the exemplary relationship between the block diagrams of Figures 1 and 2. Those skilled in the art will readily appreciate that alternative configurations for the encoder (200) within the encoding block (120) are possible.

[0038] As shown in Figure 2, the encoder (200) operates to convert an input video signal (117) into an output bitstream (122). In this embodiment, the input video signal (117) does not have film grain encoded therein. The processing performed by the encoder (200) causes the corresponding video signal encoded in the output bitstream (122) to have digital film grain therein. Additionally, the encoder (200) typically causes the output bitstream (122) to carry metadata corresponding to and / or characterizing the digital film grain.

[0039] For each image in the output bitstream (122), the encoder (200) generates a film grain image (220) using a film grain model (210). The film grain model (210) is also typically "known" to a corresponding video decoder (see, e.g., FIG. 3). The specific parameter values ​​of the film grain model (210) used by the encoder (200) in generating the film grain image (220) may typically be communicated to the corresponding video decoder as film grain model metadata (212). Thus, the corresponding video decoder can compute an exact copy of the film grain image (220) (see also FIG. 3).

[0040] The film grain injection block (240) of the encoder (200) operates to modulate the noise intensity of the film grain image (220) on a pixel-by-pixel basis based on the corresponding intensity of the input video signal (117). The resulting modulated film grain image is then added to the host image to generate a corresponding film grain-injected image. Such a sequence of film grain-injected images (242) is compressed in the video compression module (250) of the encoder (200), thereby generating a compressed video bitstream (252). In some embodiments, the video compression module (250) can perform lossy compression in accordance with a video compression standard. The metadata generator (230), with parameter inputs (222, 244, 254) shown in FIG. 2, generates film grain removal metadata (232) that includes one or more parameters that can be used for efficient and effective film grain removal in a corresponding video decoder. Examples of such metadata (232) are described in more detail below. The combiner module (260) operates to appropriately combine the metadata (212, 232) with the compressed video bitstream (252) to generate the output bitstream (122).

[0041] Figure 3 is a block diagram illustrating a video decoder (300) according to one embodiment. The decoder (300) may be used, for example, to implement the decode unit (130) of the pipeline (100) (see also Figure 1). The decoder (300) operates to convert a received bitstream (122) into a video signal (342). In one exemplary embodiment, the video signal (342) has the aforementioned digital film grain substantially completely removed therefrom, thereby representing a relatively accurate approximation of the source video signal (117) (see also Figure 2).

[0042] The decoder (300) includes a separator module (310) configured to appropriately extract the above-mentioned metadata (212, 232) and the compressed video bitstream (252) from the received bitstream (122). As already indicated above, the film grain model (210) is "known" to the decoder (300) (e.g., stored in the decoder's (300's) memory). This knowledge, together with the film grain model metadata (212) extracted by the separator module (310), allows the decoder (300) to recalculate the film grain image (220) as shown in FIG. 3. The film grain image (220) may also be referred to as a simulated film grain image.

[0043] The decoder (300) further includes a video decompression module (320) configured to appropriately decompress the compressed video bitstream (252) received from the separator module (310), thereby generating a decompressed video signal (330). Note that if lossy compression is used in the video compression module (250) of the encoder (200), the decompressed video signal (330) may differ from the sequence (242). Because the film grain injection performed in the film grain injection block (240) of the encoder (200) involves modulating the film grain image (220) with luma, the actual injected film grain does not need to be determined from the decompressed video signal (330) but may instead be estimated. This estimation is performed by the film grain removal module (340) of the decoder (300) and generally relies on the recalculated film grain image (220), the metadata (232) extracted by the separator module (310) from the received bitstream (122), and the decompressed video signal (330). An exemplary embodiment of a film grain removal method implemented in the film grain removal module (340) is described in more detail below.

[0044] Film Grain Injection Process Let the ith pixel of the jth host image frame be s ji and the corresponding film grain pixel is denoted by n ji The luma modulation function is expressed as f(). The modulation function f(s ji ) is typically used in conjunction with the host signal s ji The film grain injected pixel v ji teeth,

number

number

number

[0045] Film Grain Removal Process: Spatial Domain Figure 4 is a process flow diagram (400) illustrating a spatial domain grain removal process that can be implemented in a decoder (300) according to one embodiment. The process (400) includes a processing block (440) where the decompressed video signal (330) is converted to an output video signal (342) (see also Figure 3). The processing block (440) also receives input from the processing blocks (420, 430) that aid in the conversion process. In general, the grain removal process (400) operates on a compressed film grain injection signal (v ji ] (see also equations (1)-(2)) to obtain the signal obtained by applying video decompression to the host signal s ji Depending on the embodiment, the actual host signal recovered in the decoder (300) may be a copy of the host signal used to generate the bitstream (122) in the encoder (200), or may be a relatively accurate estimate of that host signal.

[0046] The input provided by the processing block (420) to the processing block (440) is based on the luma modulation function f() (410) described above. Relevant parameters representing the function f() may be provided to the decoder (300), for example, as part of the film grain model (210) and / or metadata (212) (see also Figures 2 and 3). In one exemplary embodiment, the processing block (420) operates to calculate a local first-order polynomial approximation of the function f(), which may be used, for example, in the grain removal process implemented in the processing block (440), as described in more detail below.

[0047] The input provided by processing block (430) to processing block (440) is an estimate of the distorted film grain image (220), where the distortion is typically caused by the sequential application of lossy compression (250) and corresponding decompression (320) (see also Figures 2 and 3). An example of how such an estimate is calculated is described in more detail below.

[0048] Lossless compression When the video compression module (250) of the encoder (200) applies lossless compression, the function g( ) defined by equations (2)-(3) is effectively a bypass function with the following properties:

number

[0049] The linear function f() can be expressed as follows:

number

number

number

[0050] If the luma modulation function f() is nonlinear, a closed-form solution for film grain removal such as equation (8) may not be available. However, in view of equation (8), an approximate solution can be constructed based on a local linear approximation of the function f(), which can be expressed as follows:

number

number

[0051] Local first-order polynomial parameters (a x ,bx ) into vector m x Presented as:

number

number

number

[0052] Similar to equation (7), equation (1) can now be approximated by equation (16) as follows:

number

number

number

number

[0053] Lossy compression If the video compression module (250) of the encoder (200) applies lossy compression, the above-described film grain removal process adapted for lossless compression may not produce optimal results. For example, note that lossy compression can cause signal distortion, such that both the host video signal and the film grain component lose some of their higher frequency components. Because the above-described "lossless" film grain removal does not take into account the distortion caused by such compression, residual film grain may be very noticeable when such removal is applied in conjunction with lossy compression.

[0054] To address this issue, equation (18) can be further modified to arrive at equation (19).

number

number

[0055] According to one example embodiment, the tilde n ji The value of is estimated using a scaled Gaussian filter G(), expressed as:

number

[0056] Figures 5-8 graphically illustrate exemplary film grain removal results that can be achieved using the scaled Gaussian filter of Equation (21). More specifically, each of Figures 5-7 shows the effect of the parameter σ at different respective bit rates. jThe dependence of the Peak Signal-to-Noise Ratio (PSNR) on σ is shown in the graphs. Figure 5 corresponds to a bit rate of 20 Mbps, Figure 6 corresponds to a bit rate of 10 Mbps, and Figure 7 corresponds to a bit rate of 5 Mbps. j The PSNR value corresponding to σ = 0 represents the result when no filter is applied. j = 0.75 and the bit rate is 20Mbps, the parameter k j 5-8 show graphs of PSNR dependence on . Collectively, the data shown in Figures 5-8 indicate that PSNR gains of up to about 2.5 dB are achievable with these embodiments.

[0057] According to another exemplary embodiment, n with a tilde ji is estimated using an FIR filter implemented according to equation (22).

number

number

[0058] Optimal filter coefficient w jk can be obtained, for example, as follows: In the encoder (200), v with a tilde ji and s ji The value of is known and can be expressed as:

number

number

[0059] An approximate solution to the problem of equation (26) can be obtained using the LMS algorithm, for example, as follows: First, the filter coefficients w jk Organize the sets in vector form.

number

number

number

number

number

[0060] The PSNR gain of FIR filter implementations generally increases with increasing filter size, however, this improvement may come at the cost of additional implementation overhead.

[0061] With respect to the guided filter embodiment, we observe that the most noticeable residual film grain noise may be present in image regions with relatively flat and / or smooth image patterns. For example, pixel values ​​in flat regions may be nearly constant, and local variations may be primarily due to added or incompletely / insufficiently removed film grain. As a preliminary observation, in such flat regions, the estimated compressed film grain

number

number

[0062] Here, the original film grain pattern (n ji ) and reconstructed film grain (n with ^ ji ) is the difference between ji Therefore, the following holds true:

number

number

number

[0063] Then, the average of the local linear model constants in a neighborhood can be calculated, for example as follows:

number

number

number

number

number

[0064] Note that, in general, guided filter embodiments may fall between Gaussian and FIR filter embodiments in terms of PSNR gain. However, because guided filter embodiments perform best for flat, smooth image regions, the subjectively perceptible improvement by a human observer may typically be more noticeable under the guided filter approach.

[0065] Film Grain Removal Process: Time Domain In an exemplary embodiment, film grain may have a zero-mean property, meaning that an image with film grain and the original image have the same DC value. As a result, the averaged pixel value for a given pixel location in a static scene tends to average the film grain back to "zero." Certain embodiments of the time-domain grain removal disclosed herein are designed to take advantage of this property to remove film grain.

[0066] 9 is a process flow diagram (900) illustrating a time-domain particle removal process that can be implemented in the decoder (300) according to one embodiment. For purposes of illustration, and without implying any limitation, the process (900) takes as its input the sequence of images (342) generated by the process (400). j-L ,…,342 j ,…,342 j+L), however, embodiments of the process (900) are not so limited. For example, in another embodiment, the process (900) can operate directly on a sequence of images (330) (see also FIG. 3). The sliding window length L can depend on the maximum amplitude of the film grain pattern. In general, larger values ​​of L can be more beneficial when the maximum amplitude is higher. In operation, the processing block (910) of the process (900) converts an input sequence of images (342) (FIG. 4) or images (330) (FIG. 3) into a corresponding output sequence of images (342′), e.g., as described below. The film grain pattern in the output sequence is less pronounced than in the input sequence.

[0067] According to an exemplary embodiment, the jth frame (342 j ) for the ith spatially filtered pixel in the temporal sliding window [w ji L ,w ji H ], the time domain filtering of the process (900) is carried out to obtain the next signal ^s for the sequence (342'). ji T Give.

number

number

[0068] w ji L The value of is, for example, the sequence (342 j-1 ,342 j-2 ,...) by iteratively checking the condition expressed by inequality (45).

number

number

[0069] w ji H The value of is similarly a sequence (342 j+1 ,342 j+2 ,...) by iteratively checking the condition expressed by inequality (47).

number

number

[0070] Once w ji L and w ji H Once the value of is determined as above, the time averaging in processing block (910) can be performed according to equation (42). Note that the time averaging described above is a pixel-based operation. The buffer size D L and D R may be provided to the decoder (300) as part of the metadata (232).

[0071] Signal Preprocessing to Accommodate Alternative Film Grain Models Some embodiments may be adapted to be compatible with alternative film grain models, i.e., film grain models different from those described above. Such embodiments generally rely on signal preprocessing to provide a preprocessed signal that is compatible with the grain removal process described above. As a non-limiting example of such preprocessing, this specification describes an exemplary preprocessing applicable to the MPEG film grain model defined in the corresponding standard(s).

[0072] An example of an MPEG film grain model is a piecewise model according to pixel intensity values. In each piece, different film grain model parameters may be used, such as different cutoff frequencies, different scaling factors for scaling the film grain, etc. If there are T such pieces in the jth frame, and the division points for the pieces are {p j t}, the film grain image in the t-th piece is expressed as {n ji t}, and the scaling factor for the t-th piece is a j t It is expressed as:

[0073] Figure 10 shows an example of an intensity-dependent scaling factor that can be accommodated using the pre-processing described above according to one embodiment. In Figure 10, it can be seen that the scaling factor can take on several discrete values ​​ranging between 0 and 30. In this case, the film grain injected pixel v ji teeth, v ji =s ji +a j t n ji t (49) where p j t ≦s ji <p j t+1 is.

[0074] Comparing Equation (49) with Equation (1), the piecewise scaling factor (a j t ) is the nonlinear modulation function f(s ji ) can be considered as a special use case of the nonlinear modulation function f(s ji The methodology described above with reference to A(s) is applicable to this particular use case. ji ) and B(s ji ) (see also equations (12) to (15)) can be calculated, for example, as follows: A(s ji )=a j t where p j t ≦s ji <p j t+1 (50) B(s ji )=0 where p j t ≦s ji <p j t+1 (51)

[0075] On the other hand, the different cutoff frequencies of the noise pattern in each piece add an additional dimension that needs to be addressed in pre-processing. ji t Complete film grain image from ji Once a complete film grain image is constructed through pre-processing, subsequent grain removal may rely on one or more of the grain removal processes described above.

[0076] To construct a complete film grain image, v in the decoder (300) ji Consider the i-th pixel of (see equation (1)). If the i-th pixel value is not close to one of the boundaries of different intensity pivot points (see, for example, Figure 10), e.g., the corresponding v ji Values ​​in the range p j t +Δ <v ji <p j t+1 If it is within −Δ, the film grain value for the complete film grain image can follow equation (52). n ji =n ji t (52) where Δ denotes the film grain amplitude, which is typically relatively small.

[0077] However, if the pixel value is relatively close to the piece boundary, e.g., p j t -Δ <v ji <p j t +Δ, the following ambiguity needs to be resolved: v ji =s ji (α) +a j t-1 n ji t-1 (53) v ji =s ji(β) +a j t n ji t (54) where a j t-1 n ji t-1 and a j t n ji t can be determined based on the received bitstream and metadata. ji indicates the decoded pixel value. However, s ji The determination of the host pixel value for s can be done in two different options, namely, s defined by equations (53) and (54): ji (α) or s ji (β) needs to be dealt with.

[0078] In one exemplary embodiment, two possible options s ji (α) or s ji (β) The choice between is made by the local neighborhood / patch (Ω i ) by analyzing the preprocessing results. More specifically, the results corresponding to a correctly selected alternative are expected to be smoother (e.g., exhibit lower residual film grain noise) than the results corresponding to the incorrect one of the two alternatives. Based on this observation, an example embodiment is configured to calculate the standard deviation for each of the two alternatives in such a patch. The standard deviation is expected to be smaller for the correctly selected one of the two alternatives.

[0079] Equations (55) to (56) are the standard deviation ω i (α) and ω i (β) provides a formula for the calculation of

number

number

[0080] According to example embodiments disclosed above, for example, in the Overview section and / or with respect to any one or any combination of some or all of FIGS. 1-10, a video processing system may include: (A) an input interface (e.g., 310 of FIG. 3) for receiving an encoded bitstream (e.g., 122 of FIGS. 1, 3) including a compressed video bitstream (e.g., 252 of FIG. 3) and metadata (e.g., 212, 232 of FIG. 3), wherein the compressed video bitstream has encoded therein a sequence of images including digital film grain, and the metadata includes one or more parameters corresponding to the digital film grain; (B) decompressing the video data (e.g., at 320 of FIG. 3 ) to generate a respective decompressed representation of each image (e.g., 330 of FIG. 3 ), each decompressed representation including a respective host image component and a respective film grain component; (B) calculating, based on the metadata, a respective estimate of the respective film grain component for each decompressed representation; and (C) removing (e.g., at 340 of FIG. 3 ) the respective estimate from each decompressed representation to generate a corresponding estimate of the respective host image component (e.g., 342 of FIG. 4 ).

[0081] In some embodiments of the above apparatus, the processor is further configured to calculate (e.g., at 210 in FIG. 3 ) a film grain image (e.g., 220 in FIG. 3 ) based on the metadata, and to calculate the respective estimates using the film grain image.

[0082] In some embodiments of any of the above devices, the processor is further configured to adjust the film grain image (e.g., at 430 in FIG. 4) to account for distortion caused by at least one of video compression and decompression.

[0083] In some embodiments of any of the above devices, the processor is configured to adjust the film grain image using a Gaussian filter (e.g., as shown in Figures 5-8).

[0084] In some embodiments of any of the above devices, the processor is configured to adjust the film grain image using a finite impulse response filter (eg, according to equations (22)-(31)).

[0085] In some embodiments of any of the above devices, the processor is configured to adjust the film grain image using a guided filter (eg, according to equations (32)-(41)).

[0086] In some embodiments of any of the above devices, the processor is further configured to calculate, for the patch of image pixels, a standard deviation corresponding to the estimated host image pixel values ​​(e.g., according to equations (55)-(56)), and to calculate a film grain image based on the standard deviation (e.g., according to equation (58)).

[0087] In some embodiments of any of the above devices, the processor is further configured to calculate (e.g., at 420 of FIG. 4 ) a first-order polynomial approximation of the nonlinear luma modulation function (e.g., f() in equation (1)) and calculate the respective estimates using said approximation.

[0088] In some embodiments of any of the above apparatus, the processor is further configured to calculate the approximation using metadata.

[0089] In some embodiments of any of the above devices, the processor is further configured to calculate the respective estimates using a pre-computed lookup table (e.g., 420 in FIG. 4) that stores parameters of a first-order polynomial approximation of the nonlinear luma modulation function (e.g., f() in Equation (1)).

[0090] In some embodiments of any of the above apparatus, the processor is further configured to temporally average the sequence of each decompressed representation using a plurality of temporal sliding windows (e.g., at 910 in FIG. 9 ; according to equation (42)), each of the temporal sliding windows corresponding to a respective image frame pixel, at least some of the temporal sliding windows having different respective lengths (e.g., selected according to equations (45), (47)).

[0091] In some embodiments of any of the above devices, the processor may be configured to: A and τ B ) and further configured to select a length of the temporal sliding window based on

[0092] In some embodiments of any of the above devices, the device further comprises a video encoder (e.g., 120 of FIG. 1 ) comprising an output interface (e.g., 260 of FIG. 2 ) for outputting an encoded bitstream for the video encoder, and a video compression module (e.g., 250 of FIG. 2 ) for generating a compressed video bitstream using lossy compression in accordance with a video compression standard.

[0093] According to another exemplary embodiment disclosed above, e.g., in the Overview section and / or with respect to any one or any combination of some or all of FIGS. 1-10, an input interface (e.g., 310 of FIG. 3) for receiving an encoded bitstream (e.g., 122 of FIGS. 1, 3) including a compressed video bitstream (e.g., 252 of FIG. 3) and metadata (e.g., 212, 232 of FIG. 3), wherein the compressed video bitstream has encoded therein a sequence of images including digital film grain, and the metadata includes one or more parameters corresponding to the digital film grain. decompressing the compressed video bitstream (e.g., at 320 in FIG. 3 ) to generate a respective decompressed representation (e.g., at 330 in FIG. 3 ) of each image, each decompressed representation including a respective host image component and a respective film grain component; temporally averaging (e.g., at 910 in FIG. 9 ; according to equation (42)) the sequence of the respective decompressed representations using a plurality of temporal sliding windows, each of the temporal sliding windows corresponding to a respective image frame pixel, at least some of the temporal sliding windows being in correspondence with metadata (e.g., τ in equations (43) and (44)). A and τ B and a processor (e.g., 320, 340 of FIG. 3 ) configured to perform steps (a) to (c) and (d) having different respective lengths selected based on (e.g., according to equations (45), (47)).

[0094] In some embodiments of the above apparatus, the apparatus further comprises a video encoder (e.g., 120 in FIG. 1 ) comprising an output interface (e.g., 260 in FIG. 2 ) for outputting an encoded bitstream for the video encoder, and a video compression module (e.g., 250 in FIG. 2 ) for generating a compressed video bitstream using lossy compression in accordance with a video compression standard.

[0095] According to yet another exemplary embodiment disclosed above, for example, in the Summary section and / or with respect to any one or any combination of some or all of FIGS. 1-10, there is provided a machine-implemented method for removing film grain from video data, the method comprising: (A) receiving an encoded bitstream (e.g., 122 of FIGS. 1 and 3) including a compressed video bitstream (e.g., 252 of FIG. 3) and metadata (e.g., 212, 232 of FIG. 3), wherein the compressed video bitstream has encoded therein a sequence of images including digital film grain, and the metadata includes one or more digital film grains corresponding to the digital film grain; or a plurality of parameters; (B) decompressing the compressed video bitstream (e.g., at 320 of FIG. 3) to generate a respective decompressed representation of each image (e.g., 330 of FIG. 3), each decompressed representation including a respective host image component and a respective film grain component; (C) calculating, based on the metadata, a respective estimate of the respective film grain component for each decompressed representation; and (D) removing the respective estimate from each decompressed representation (e.g., at 340 of FIG. 3) to generate a corresponding estimate of the respective host image component (e.g., 342 of FIG. 4).

[0096] In some embodiments of the above method, the method further includes calculating (e.g., at 210 in FIG. 3 ) a film grain image (e.g., 220 in FIG. 3 ) based on the metadata, and wherein calculating the respective estimates includes using the film grain image.

[0097] In some embodiments of any of the above methods, the method further includes calculating (e.g., at 420 of FIG. 4) a first-order polynomial approximation of a nonlinear luma modulation function (e.g., f() of FIG. 1), and calculating the respective estimates includes using the approximation.

[0098] In some embodiments of any of the above methods, the method further includes temporally averaging the sequence of each decompressed representation using a plurality of temporal sliding windows (e.g., at 910 of FIG. 9 ; according to equation (42)), each of the temporal sliding windows corresponding to a respective image frame pixel, at least some of the temporal sliding windows having different respective lengths (e.g., selected according to equations (45), (47)).

[0099] In some embodiments of any of the above methods, the compressed video bitstream is generated using lossy compression according to a video compression standard.

[0100] According to yet another exemplary embodiment disclosed above, for example, in the Summary section and / or with respect to any one or any combination of some or all of Figures 1-10, there is provided a non-transitory machine-readable medium encoded with program code that, when executed by a machine, causes the machine to perform a method. The method includes: (A) receiving an encoded bitstream (e.g., 122 in FIGS. 1 and 3) including a compressed video bitstream (e.g., 252 in FIG. 3) and metadata (e.g., 212, 232 in FIG. 3), where the compressed video bitstream has encoded therein a sequence of images including digital film grain, and the metadata includes one or more parameters corresponding to the digital film grain; (B) decompressing (e.g., at 320 in FIG. 3) the compressed video bitstream to generate a respective decompressed representation (e.g., 330 in FIG. 3) of each image, where each decompressed representation includes a respective host image component and a respective film grain component; (C) calculating, based on the metadata, a respective estimate of each film grain component for each decompressed representation; and (D) removing (e.g., at 340 in FIG. 3) each estimate from each decompressed representation to generate a corresponding estimate of each host image component (e.g., 342 in FIG. 4).

[0101] With respect to the processes, systems, methods, heuristics, etc. described herein, although steps of such processes, etc. are described as occurring according to a certain ordered sequence, it should be understood that such processes can be implemented with the described steps occurring in an order other than the order described herein. Furthermore, it should be understood that certain steps can be performed simultaneously, other steps can be added, or certain steps described herein can be omitted. In other words, the process descriptions herein are provided for the purpose of illustrating certain embodiments and should not be construed as limiting the scope of the claims in any way.

[0102] Thus, it should be understood that the above description is intended to be illustrative, and not restrictive. Many embodiments and applications other than the examples provided will be apparent from reading the above description. The scope should not be determined with reference to the above description, but instead should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments in the technology described herein will occur, and that the disclosed systems and methods will be incorporated into such future embodiments. In short, it should be understood that this application is capable of modification and variation.

[0103] All terms used in the claims are intended to be given their broadest reasonable interpretation and their ordinary meaning as understood by one skilled in the art described herein, unless expressly indicated to the contrary herein. In particular, the use of singular articles such as "a," "the," "said," etc., should be read as describing one or more of the indicated elements unless the claim describes an express limitation to the contrary.

[0104] The Abstract of the present disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of improving the flow of the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Accordingly, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as separately claimed subject matter.

[0105] While the present disclosure includes reference to exemplary embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the present disclosure that are apparent to those skilled in the art to which the present disclosure pertains, are deemed to be within the principles and scope of the present disclosure, as expressed, for example, in the following claims.

[0106] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.

[0107] Some embodiments may be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments may also be embodied in the form of program code recorded on tangible media, such as magnetic recording media, optical recording media, solid-state memory, floppy disks, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, which, when loaded into and executed by a machine, such as a computer, makes the machine an apparatus for practicing the patented invention. Some embodiments may also be embodied in the form of program code stored on non-transitory machine-readable storage media, including, for example, being loaded into and / or executed by a machine, which, when loaded into and executed by a machine, such as a computer or processor, makes the machine an apparatus for practicing the patented invention. When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.

[0108] Unless expressly stated otherwise, each numerical value and range should be construed as an approximation, as if the word "about" or "approximately" preceded the value or range.

[0109] Use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter to facilitate claim interpretation, and such use should not be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.

[0110] Although elements in the following method claims are described in a particular order with corresponding labeling, the elements are not necessarily intended to be limited to being implemented in that particular order, unless the claim description implies a particular order for implementing some or all of the elements.

[0111] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of the present disclosure. Appearances of the phrase "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, nor are separate or alternative embodiments necessarily exclusive of other embodiments. The same applies to the term "implementation."

[0112] Unless otherwise specified herein, the use of the ordinal adjectives "first," "second," "third," etc. to refer to one object of a plurality of similar objects merely indicates that different instances of such similar objects are being referred to and does not imply that the similar objects so referred to must be in a corresponding order or sequence in time, space, ranking, or in any way whatsoever.

[0113] Unless otherwise specified herein, in addition to its plain meaning, the conjunction "if" may also or alternatively be interpreted to mean "when" or "upon" or "response to determining" or "response to detecting," and the interpretation may depend on the particular context in which it is associated. For example, the phrase "upon determining" or "upon detecting [the stated condition or event]" may be interpreted to mean "upon determining" or "response to determining" or "upon detecting [the stated condition or event]" or "response to detecting [the stated condition or event]."

[0114] Also, for purposes of this document, the terms "couple," "coupled," "coupled," "connect," "connection," or "connected" refer to any manner known or later developed in the art that permits energy to be transferred between two or more elements, where the intervening presence of one or more additional elements is contemplated, but not required. Conversely, terms such as "directly coupled," "directly connected," and the like imply the absence of such additional elements.

[0115] As used herein with respect to elements and standards, the term conforming means that the element communicates with other elements in a manner specified, fully or partially, by the standard and is recognized by other elements as being sufficiently capable of communicating with other elements in the manner specified by the standard. A conforming element need not operate internally in the manner specified by the standard.

[0116] The functions of the various elements illustrated in the figures, including any functional blocks labeled "processor" and / or "controller," may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. If provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by multiple individual processors, some of which may be shared. Furthermore, explicit use of the terms "processor" or "controller" should not be construed as referring only to hardware capable of executing software, but may implicitly include, without limitation, digital signal processor (DSP) hardware, network processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), read-only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Those functions may be performed through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, with the particular technique being selectable by the implementer as more particularly understood from the context.

[0117] As used in this application, the terms "circuit" and "circuitry" may refer to one or more or all of the following: (a) a hardware-only circuit implementation (e.g., an implementation with only analog and / or digital circuitry); (b) a combination of hardware circuitry and software, such as (where applicable) (i) a combination of analog and / or digital hardware circuitry with software / firmware, and (ii) any portion of a hardware processor with software (including a digital signal processor, software, and memory that cooperate to cause a device such as a mobile phone or server to perform various functions); and (c) a hardware circuit and / or processor, such as a microprocessor or portion of a microprocessor, that requires software (e.g., firmware) to operate, but that may be absent when not required for operation. This definition of circuit applies to all uses of the term in this application, including any claims. As a further example, as used in this application, the term circuit also covers simply a hardware circuit or processor (or processors), or a portion of a hardware circuit or processor, and its (or their) accompanying software and / or firmware implementation. The term circuitry also encompasses, for example, baseband or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices, where applicable to particular claim elements.

[0118] It should be understood by those skilled in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the present disclosure. Similarly, any flowcharts, flow diagrams, state transition diagrams, pseudocode, etc., may be substantially represented on a computer-readable medium and, therefore, will be understood to represent various processes that may be performed by such a computer or processor, whether or not a computer or processor is explicitly shown.

[0119] This Summary is intended to introduce some example embodiments; additional embodiments are described in the Detailed Description and / or with reference to one or more of the Figures. This Summary is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0120] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE). [EEE1] 1. A video delivery system capable of film grain removal, the system having a video decoder, the video decoder comprising: an input interface for receiving an encoded bitstream including a compressed video bitstream and metadata, the compressed video bitstream encoding a sequence of images including digital film grain, and the metadata including one or more parameters corresponding to the digital film grain; A processor comprising: decompressing the compressed video bitstream to generate a respective decompressed representation of each image, each decompressed representation including a respective host image component and a respective film grain component; calculating an estimate of each of the respective film grain components for each of the decompressed representations based on the metadata; removing said respective estimates from said respective decompressed representations to generate corresponding estimates of said respective host image components; A video delivery system comprising: [EEE2] The processor: calculating a film grain image based on said metadata; Calculating each of the estimates using the film grain image. The video delivery system of EEE1, further configured as follows: [EEE3] The processor: For a patch of image pixels, calculate the standard deviation corresponding to the estimated host image pixel values; Calculating the film grain image based on the standard deviation 3. The video delivery system of claim 1, further configured to: [EEE4] The video delivery system of EEE2 or 3, wherein the processor is further configured to adjust the film grain image to account for distortion caused by at least one of video compression and decompression. [EEE5] The video delivery system of EEE4, wherein the processor is configured to adjust the film grain image using a Gaussian filter. [EEE6] The video delivery system of any one of EEE4 and EEE5, wherein the processor is configured to adjust the film grain image using a finite impulse response filter. [EEE7] 7. The video delivery system of any one of claims 8 to 10, wherein the processor is configured to adjust the film grain image using a guided filter. [EEE8] The processor: Computes a first-order polynomial approximation of the nonlinear luma modulation function; Calculate each of the estimates using the approximations 8. The video delivery system of any one of EEE1 to EEE7, further configured to: [EEE9] 8. The video delivery system of claim 8, wherein the processor is further configured to calculate the approximation using the metadata. [EEE10] 10. The video delivery system of any one of EEE1 to EEE9, wherein the processor is further configured to calculate the respective estimates using a pre-computed lookup table storing parameters of a first order polynomial approximation of a non-linear luma modulation function. [EEE11] The processor: temporally averaging the sequence of each decompressed representation using a plurality of temporal sliding windows, each of the temporal sliding windows corresponding to a respective image frame pixel, and at least some of the temporal sliding windows having different respective lengths; selecting a length of the temporal sliding window based on the metadata; 11. The video delivery system of any one of EEE1 to EEE10, further configured to: [EEE12] 12. The video delivery system of any one of claims EE1 to 11, further comprising a video encoder, wherein the video encoder: an output interface for outputting the encoded bitstream for the video encoder; a video compression module generating said compressed video bitstream using lossy compression according to a video compression standard; A video delivery system comprising: [EEE13] A video delivery system capable of film grain removal is provided, the system having a video decoder, the video decoder comprising: an input interface for receiving an encoded bitstream including a compressed video bitstream and metadata, the compressed video bitstream encoding a sequence of images including digital film grain, and the metadata including one or more parameters corresponding to the digital film grain; A processor comprising: decompressing the compressed video bitstream to generate a respective decompressed representation of each image, each decompressed representation including a respective host image component and a respective film grain component; a processor configured to perform the step of temporally averaging the sequence of the respective decompressed representations using a plurality of temporal sliding windows, each of the temporal sliding windows corresponding to a respective image frame pixel, at least some of the temporal sliding windows having different respective lengths selected based on the metadata; A video delivery system comprising: [EEE14] 10. The video delivery system according to EEE13, further comprising a video encoder, wherein the video encoder: an output interface for outputting the encoded bitstream for the video encoder; a video compression module generating said compressed video bitstream using lossy compression according to a video compression standard; A video delivery system comprising: [EEE15] 1. A machine-implemented method for removing film grain from video data, the method comprising: receiving an encoded bitstream comprising a compressed video bitstream and metadata, wherein the compressed video bitstream encodes a sequence of images comprising digital film grain, and the metadata comprises one or more parameters corresponding to the digital film grain; decompressing the compressed video bitstream to generate a respective decompressed representation of each image, each decompressed representation including a respective host image component and a respective film grain component; calculating an estimate of each of the respective film grain components for each of the decompressed representations based on the metadata; and removing said respective estimates from said respective decompressed representations to produce corresponding estimates of said respective host image components. method. [EEE16] further comprising calculating a film grain image based on the metadata; calculating the respective estimates includes using the film grain image; Method as described in EEE15. [EEE17] further comprising calculating a first order polynomial approximation of the nonlinear luma modulation function; calculating the respective estimates includes using the approximations; The method according to EEE15 or EEE16. [EEE18] 18. The method of any one of EEE15 to 17, further comprising temporally averaging the sequence of said respective decompressed representations using a plurality of temporal sliding windows, each said temporal sliding window corresponding to a respective image frame pixel, and at least some of said temporal sliding windows having different respective lengths. [EEE19] 8. The method of any one of EEE15 to 18, wherein the compressed video bitstream has been generated using lossy compression according to a video compression standard. [EEE20] A non-transitory machine-readable medium encoded with program code that, when executed by a machine, causes the machine to perform operations that constitute a method according to any one of claims EEE15 to EEE19.

Claims

1. 1. A method for removing digital film grain from video data, performed by a video decoder, the method comprising: receiving an encoded bitstream including a compressed video bitstream and metadata, the compressed video bitstream encoding a sequence of images including digital film grain, the metadata including film grain model metadata (212) and film grain removal metadata (232), the film grain model metadata including one or more parameters corresponding to the digital film grain; pixel [Equation 1] Decompress the respective decompressed representations of each image generating [Equation 2] represents the ith pixel of the jth film grain injected image frame after video compression extracted from the compressed video bitstream, and each of the decompressed representations is a respective host image component s ji and each film grain component n ji Contains s ji represents the ith pixel of the jth host image frame without film grain, and n ji represents the i-th pixel of the j-th film grain image frame, and Film grain injected pixels after video compression [Equation 3] For , the estimated value of the compressed film grain component [Equation 4] and each estimate is [Equation 5] is determined by calculating [Equation 6] represents the estimate of the ith pixel of the jth film grain image frame after video compression, k j represents a scaling factor included in the film grain removal metadata, G() represents a scaled Gaussian filter available in the video decoder; n ji represents the i-th pixel of the j-th film grain image frame obtained using the film grain model metadata from a known film grain model available in the video decoder; σ j represents the spectral width of the scaled Gaussian filter included in the film grain removal metadata; Film grain injected pixels after video compression [Equation 7] for each of said respective host image components, [Equation 8] wherein each estimate is: [Equation 9] is determined by calculating [Equation 10] represents the estimate of the ith pixel of the jth host image frame without film grain, A() and B() represent coefficients of a local linear approximation of a luma modulation function used to modulate the film grain image onto the respective host image, and are obtained from the known film grain model available in the video decoder or from the film grain model metadata; method.

2. 2. The method of claim 1, further comprising temporally averaging the sequence of the respective decompressed representations using a plurality of temporal sliding windows, each of the temporal sliding windows corresponding to a respective image frame pixel, and at least some of the temporal sliding windows having different respective lengths.

3. 10. The method of claim 1, wherein the compressed video bitstream is generated using lossy compression according to a video compression standard.

4. 1. A video delivery system capable of digital film grain removal, the system having a video decoder, the video decoder comprising: an input interface for receiving an encoded bitstream including a compressed video bitstream and metadata, the compressed video bitstream encoding a sequence of images including digital film grain, the metadata including film grain model metadata (212) and film grain removal metadata (232), the film grain model metadata including one or more parameters corresponding to the digital film grain; a processor configured to perform the method of any one of claims 1 to 3; Video delivery system.

5. 5. The video delivery system of claim 4, further comprising a video encoder, the video encoder comprising: an output interface for outputting the encoded bitstream; a video compression module for generating the compressed video bitstream using lossy compression according to a video compression standard; A video delivery system comprising:

6. A non-transitory computer-readable medium on which a computer program for causing a computer to execute a method according to any one of claims 1 to 3 is stored.

Citation Information

Patent Citations

  • Techniques for Simulating Film Grain in Encoded Video

    JP2006524013A

  • Adjusting film grain properties in digital images

    US5641596A