Film grain repetitive pattern scoring

The GraPQA method addresses perceptible repetitive film grain patterns by using mean filtering and local similarity correlation to reconstruct film grain, ensuring efficient compression and maintaining artistic intent in video distribution.

WO2025202987A1PCT designated stage Publication Date: 2025-10-02IMAX CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/053293
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-03-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing video encoders fail to preserve the artistic intent of film grain in video distribution, leading to perceptible repetitive film grain patterns that detract from the viewing experience due to inefficient compression and synthesis techniques.

Method used

A method for film grain pattern quality assessment (GraPQA) is implemented, involving mean filtering, local similarity correlation measurements, and scoring to optimize the detection and penalization of repetitive film grain patterns, using a film grain template to reconstruct film grain in a way that aligns with human visual perception.

Benefits of technology

The GraPQA method effectively reduces perceptible repetitive film grain patterns, enhancing the viewing experience by maintaining the artistic intent of film grain while optimizing video compression and synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025053293_02102025_PF_FP_ABST
    Figure IB2025053293_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Grain pattern quality assessment (GraPQA) is performed. Mean filtering is performed on a luminance channel of an input frame to generate a mean frame. The mean frame is subtracted from the input frame to create a mean-subtracted frame. Local similarity correlation measurements are performed between the mean-subtracted frame and a film grain template. Results of the local similarity correlation measurements are scored to generate a grain pattern quality score for the input frame.
Need to check novelty before this filing date? Find Prior Art

Description

FILM GRAIN REPETITIVE PATTERN SCORINGCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. provisional application Serial No. 63 / 571,623 filed March 29, 2024, the disclosure of which is hereby incorporated in its entirety by reference herein.TECHNICAL FIELD

[0002] Aspects of the disclosure generally relate to the detection and penalization of repetitive film grain patterns.BACKGROUND

[0003] Film grain is essentially specific random noise in a video or still image frame. Film grain may originate from an analog film acquisition process as a random optical texture of processed photographic film due to the presence of small particles of metallic silver. Film grain may also be synthesized purposely and digitally in the content production and post-production pipelines. In either case, the visual feel of the film grain may be considered to be part of the artistic or creative intent of the content producers.SUMMARY

[0004] Preserving such creative intent in video distribution may be challenging when encoders are used in image content streaming process, such as encoders used to put content on networks to be received and decoded for viewing devices that have a display to view images. For example, visual artifacts may become apparent on a viewing device as a result of the processes applied to get the image data to the viewing device, such as processes to get image data onto networks and receiving image data from networks that is further processed within viewing devices or set top units that process image data that is passed on to a viewing device. Video encoders, which treat film grains no differently from other types of noise may cause the encoder to consume a largenumber of bits to encode the film grain. Reducing the number of bits to encode and decode to get image content with film grain displayed on a viewing device requires developing techniques to minimize or eliminate undesirable film grain related artifacts on viewing displays. Disclosed herein is a technique to address film grain related artifacts.

[0005] One film grain related image artifact that may occur is a perceptible repetitive film grain pattern that becomes visible on a display that has received decoded image content intended to be shown as an image with film grain. Perceptible repetitive film grain pattern refers to an image frame where portions of the image exhibit a perceptible repeat pattern of film grain image texture that is visually unnatural or is unacceptable image quality with respect to the source image or original image. To address perceptible repetitive film grain pattern artifacts a repetitive film grain pattern quality metric may be created to allow a perceptible grain pattern quality assessment to be performed to facilitate optimization to reduce or eliminate perceptible repetitive film grain pattern artifacts from becoming visible on a viewing device.

[0006] In one or more illustrative examples, a method for performing grain pattern quality assessment (GraPQA) is performed. Mean filtering is performed on a luminance channel of an input frame to generate a mean frame. The mean frame is subtracted from the input frame to create a mean-subtracted frame. Local similarity correlation measurements are performed between the mean-subtracted frame and a film grain template. Results of the local similarity correlation measurements are scored to generate a grain pattern quality score for the input frame.

[0007] In one or more illustrative examples, a system for performing grain pattern quality assessment (GraPQA) includes one or more hardware computing devices configured to perform mean filtering on a luminance channel of an input frame to generate a mean frame; subtract the mean frame from the input frame to create a mean-subtracted frame; perform local similarity correlation measurements between the mean-subtracted frame and a film grain template; and score results of the local similarity correlation measurements to generate a grain pattern quality score for the input frame.BRIEF DESCRIPTION OF THE DRAWINGS[0008| FIG. 1 illustrates an example of an end-to-end system for grain-aware video coding, transmission, and scoring of repetitive film grain patterns;[0009) FIG. 2A illustrates an example data flow for the performance of a grain pattern quality assessment of one embodiment;[00101 FIG. 2B illustrates an example data flow for the performance of a grain pattern quality assessment of a second embodiment;

[0011] FIG. 3 shows an example the effectiveness of the data flow in detecting film grain patterns in a sample decoded frame of a video asset;

[0012] FIG. 4 illustrates an example graph of the relationship of GraPQA scores to block sizes of the film grain template;

[0013] FIG. 5 illustrates an example of detecting film grain patterns in a sample decoded frame of a video asset for three different block sizes;

[0014] FIG. 6 illustrates an example process for the determination of the GraPQA scores for a video asset; and

[0015] FIG. 7 illustrates an example of a computing device for use in the determination of the GraPQA scores for a video asset.DETAILED DESCRIPTION

[0016] As required, detailed embodiments of the present invention are disclosed herein; however, it is to be understood that the disclosed embodiments are merely exemplary of the invention that may be embodied in various and alternative forms. The figures are not necessarily to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein are not to be interpreted aslimiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present invention.

[0017] During encoding, denoisers may be used to separate authentic film grain from video frames and estimate film grain model parameters, which are then transmitted alongside video bitstreams. In the decoding stage, received film grain parameters are utilized to generate film grain templates. These templates are cropped into grain blocks based on a film grain block size parameter, randomized, and reintegrated into the video to reconstruct film grains. However, the choice of film grain block size significantly influences the quality of reconstructed film grain. Larger block sizes may lead to noticeable repetitive noise patterns, detracting from the viewing experience.

[0018] Aspects of the disclosure generally relate to detection and penalization of repetitive film grain patterns. As discussed in detail herein, a quality assessment approach is described that is tailored to detect and penalize film grain repetitive patterns efficiently. This approach demonstrates high correlation with human judgments of video quality, affirming its effectiveness in evaluating decoded videos containing synthesized film grain.

[0019] FIG. 1 illustrates an example of an end-to-end system 100 for grain-aware video coding, transmission, and scoring of repetitive film grain patterns. In the illustrated example, a video asset 102 is received from a content source 104. The video asset 102 is provided to a denoiser 106 to produce a denoised video asset 108. The video asset 102 and the denoised video asset 108 are then provided to a film grain model parameterization 110, which generates film grain model parameters 112 based on these two inputs. The denoised video asset 108 and the film grain model parameters 112 are then provided to an encoder 114, which generates an encoded video asset 116 including the film grain model parameters 112 as metadata. The encoded video asset 116 is supplied to a network 118 for transmission.

[0020] The network 118 then provides the encoded video asset 116 to a decoder 120, which decodes the encoded video asset 116 into a decoded video asset 122. Additionally, the decoder 120 extracts the film grain model parameters 112 from the metadata of the encoded video asset 116. The film grain model parameters 112 and the decoded video asset 122 are then provided to a film grain synthesizer 124, which uses the inputs to generate a film grain template 126, which inturn is used to create a decoded video with synthesized film grain 128. The decoded video with synthesized film grain 128 may then be played back on a viewer device 130. It should be noted that the end-to-end system 100 may be geographically diverse and that the calculations may occur co-located or in a distributed manner. For example, denoising of the video asset 102 can be processed by a different processor at a different location then the encoding process or the denoising and the encoding could be done by the same processing unit. In another example the denoising and encoding process can be done at a different location then the decoding and grain synthesis processing.

[0021] A video asset 102 can be multiple image frames such as a sequence of image frames and may include, as some examples, live video feeds from current events, pre-recorded shows or movies, and advertisements or other clips to be inserted into other video feeds. The video asset 102 may include just video in some examples, but in many cases the video asset 102 further includes additional content such as audio, subtitles, and metadata information descriptive of the content and / or format of the video. The video asset 102 may comprise a series of one or more frames to be displayed in succession, which may be encoded in any of various resolutions (e.g., 720p, 1080p, WUXGA, 2K, ultra-high definition (UHD), cinema 4K, 8K, etc.), frame rates (e.g., 24fps, 30fps, 60 fps, etc.), dynamic ranges (e.g., standard dynamic range (SDR), high dynamic range (HDR), and color spaces (Y’Cb’Cr’, RGB, etc.).

[0022] The end-to-end system 100 may include one or more content sources 104 of the video assets 102. Content sources 104 may include video networks, live camera feeds, sports events, advertisement feeds, post production sources like Hollywood post production content, etc.

[0023] The denoiser 106 is configured to receive the video asset 102 and remove film grain from the video asset 102 to generate the denoised video asset 108. To do so, the denoiser 106 may use one or more denoising algorithms to filter the input video asset 102 to smooth out random noise from each frame of the video asset 102. Example denoising algorithms may include Digital Media Remastering (DMR), Block-Matching and 3D filtering denoising algorithm (BM3D), and / or High Quality DeNoise 3D (Hqdn3d), as some non-limiting examples. The denoised video asset 108 is the resultant video output from the denoiser 106.

[0024] The film grain model parameterization 110 is a process whereby frames of the video asset 102 are compared to those of the denoised video asset 108, resulting in film grain model parameters 112 that characterize the noise present in the video asset 102 compared to the denoised video asset 108. In an example, the film grain model parameters 112 are determined by performing a difference of the frames of the video asset 102 and denoised video asset 108, and then analyzing the high frequency components of flat or smooth regions of the difference between the noisy and de-noised versions. Using the regions that are relatively flat or smooth is useful because high frequency components from edges and textures can adversely affect estimation of the grain parameters. In an example, the smooth areas can be found by applying an edge detector to the denoised image followed by an optional dilation operation. Based on the comparison of these regions, film grain model parameters 112 may be determined such as grain size, grain contrast intensity, grain density and grain pattern. In some examples, the noise is modeled using a 2D causal autoregressive process with pseudo-Gaussian noise input. As the film grain strength can vary with signal intensity, the film grain strength may be modeled using a non-linear color space Y’Cb’Cr’ representation, as a piece-wise linear function for each of Y, Cb, and Cr color components.

[0025] The encoder 114 may include electronic circuits and / or software configured to compress the denoised video asset 108 into an encoded video asset 116. The encoded video asset 116 may be encoded to a format that conforms with one or more standard video compression specifications. Examples of video encoding formats include MPEG-2 Part 2, MPEG-4 Part 2, H.264 (MPEG-4 Part 10), HEVC, Theora, RealVideo RV40, VP9, and AOMedia Video 1 (AVI). In many cases, the compressed video 116 lacks image information present in the original video asset 102 such as video asset that has film grain, which is referred to as lossy compression. Significantly, because film grain is random and very difficult to compress, the removal of film grain allows for the more efficient compression of the denoised video asset 108 into the encoded video asset 116 as compared to using the video asset 102 directly. In addition to compressing the video content, the encoder 114 may also include the film grain model parameters 112 in the data stream of the encoded video asset 116 to produce an encoded video asset with metadata 116 where the metadata may contain film grain modelling parameters.

[0026] The network 118 may include a geographically-distributed network of servers and data centers configured by a provider to provide the encoded video asset 116 from a source or an originlocation to a destination location. For example, the encoded video asset with metadata 116 may be transported to a designation location for viewing on a viewer device 130

[0027] The decoder 120 may receive the encoded video asset with metadata 116 and may decode the encoded video asset with metadata 116 into a decoded video asset 122 and metadata having film grain parameters. The decoder 120 may use a corresponding decoding algorithm corresponding to the encoding algorithm used by the encoder 114. The determination of what algorithm to use may be identified based on metadata of the encoded video asset 116. Additionally, the decoder 120 may also extract the film grain model parameters 112 from the metadata of the encoded video asset 116.

[0028] The film grain synthesizer 124 may receive the decoded video asset 122 and also the film grain model parameters 112. Based on these inputs, the film grain synthesizer 124 may construct simulated film grain, based on the film grain model parameters 112, in an attempt to recreate the look of the film grain that was originally present but removed by the denoiser 106. The film grain synthesizer 124 may include various operations, such as grain template generation; randomization; local adaptation; deblocking; and blending.

[0029] The grain template generation is an operation that may include use of an auto-regressive model to generate a film grain template 126. The film grain template 126 may represent an image area that is a smaller area than that of the full image frame area, such as 64x64 or 73x82 image pixel area, as some nonlimiting examples. The coefficients of the auto-regressive model used to generate the film grain template 126 can be coefficients estimated by the encoder 114 and transmitted to the decoder 120 in the metadata of the encoded video asset 116. In some examples, more than one film grain template 126 can be generated for a single frame or sequence of frames.

[0030] The randomization is an operation that involves a process that includes steps to extend the film grain template 126 area to the size of the full image frame. The randomization process can include dividing the image frame to be grained into blocks smaller in size than the generated film grain template(s) 126 size. For each block, the randomization process selects a random region of the film grain template 126 to apply film grain to that the smaller block.

[0031] Local adaptation is an operation that involves a process of choosing the relevant film grain template 126 and applying scaling according to the characteristics of the pixel values of a smoothed input frame that is the decoded film frame that is a denoised and is a smoothed frame. Local adaptation can be performed pixel-wise or block-wise, where in the scaling can be calculating a local block or pixel intensity average for each block. The deblocking is performed because the block-based randomization can result in blocking artifacts that can be removed or minimized by implementing deblocking such as a low-pass filtering of block borders. In another example, the deblocking may include overlapping the film grain values at film grain block boundaries to ensure possible artifacts are properly attenuated.

[0032] Blending is an operation that involves a process of blending the synthesized film grain 128 generated by the synthesizer operations described above with the smoothed input frame, the blending operation can be performed either by addition or multiplication. The blending may be performed mathematically to the luma and / or chroma channels, by adding the luma and / or chroma channels of the deblocked film grain onto the corresponding luma and / or chroma channels of the decoded video asset 122. The result of the blending operation is the decoded video with synthesized film grain 128 being outputted.

[0033] The viewer device 130 may be a video player to play back the decoded video with synthesized film grain 128 received by the viewer device 130. The viewer device 130 may include, as some examples, a set-top box connected to a television or other video screen, a tablet computing device, and / or a mobile phone, or a cinema projection system in a theatre. Notably, these varied viewer devices 130 may have different viewing condition (including illumination and viewing distance, etc.), display spatial resolution (e.g., SD, HD, full-HD, UHD, 4K, etc.), frame rate (15, 24, 30, 60, 120 frames per second, etc.), dynamic range (8 bits, 10 bits, and 12 bits per pixel per color, etc.).

[0034] The operations within the grain synthesizer 124 can be optimized using a quality measure determined by a process that does a Grain Pattern Quality Assessment (GraPQA). FIGS. 2A and 2B illustrates two examples of data flow 200 of a process that does GraPQA and produces a quality score that is a GraPQA score. The GraPQA can be determined for different video assets 102 where film grain is synthesized for the video asset 102. For example, the video asset 102 may be a digitalvideo asset 102 received from a source in which film grain like characteristics are to be added. Such a video asset 102 can be received 103 as shown in FIG. 2A in which the input frame 204 or frames from the video asset 102 are applied to the mean filter 206. The film grain model parameters 112 to be used by the film grain template generator 202 may come from a separate source 222 or provided by a user.

[0035] In another example the video asset 102 may be images with film grain removed. The image with removed film grain is encoded before being received 103 by the data flow 200 as shown in FIG. 2B. In this situation, as shown in FIG. 2B the encoded video asset 116 is received 103 and decoded by the decoder 120 that produces a decoded video asset input frame 204 or frames which are provided to the mean filter 206 and the subtractor 210. This scenario is a way an end- to-end system 100 may deliver original video image content that has film grain to a user viewing device where the film grain is removed from the video asset 102 before encoding and delivering the encoded asset to the end user device where film grain is resynthesized and added to the decoded video asset 122. Film grain modelling parameters associated with the video frame are sent with the encoded video and may be sent as metadata.

[0036] To determine a GraPQA score for a video asset 102 from an encoded video asset 116 with metadata 116, a decoded input frame 204 is extracted from the encoded video asset 116 with metadata 116 by the decoder 120. The film grain model parameters 112 are computed or extracted from the encoded video asset 116 with metadata 116 decoded by the decoder 120. A film grain template generator 202 generates the film grain template 126 using the film grain model parameters 112. The decoded input frame 204 is applied to a mean filter 206 to generate a mean frame 208. A frame that is a difference between the mean frame 208 and the decoded input frame 204 is generated by a subtractor 210 and results in a frame that is referred to as the mean- subtracted frame 212. The mean- subtracted frame 212 and the film grain template 126 are applied to a local similarity correlation measurement 214, which determines a local similarity correlation frame 216. The local similarity correlation frame may show the extent of the local similarity correlation for corresponding pixels and blocks between the film grain template 126 with its locally modified and cropped replications over a frame generated by the film grain template generator 202 and the corresponding mean- subtracted frame 212 of the decoded video asset 122. The local similarity correlation measurement may be established such that the greater the local similarity correlationthe lower the degree of mismatch and the lower the local similarity correlation the higher the degree of mismatch. The local similarity correlation can be a ratio of or difference between a vector expression of the frame generated by the film grain template 126 and an expression of the corresponding mean-subtracted frame 212. An example of one representation that may be used for the local similarity correlation measurement as described above may be the Absolute Cosine similarity correlation. This local similarity correlation frame 216 is scored by a scoring algorithm 218, to produce a GraPQA score 220. Significantly, this assessment of the GraPQA score 220 takes advantage of the absolute cosine similarity correlation between the film grain template 126 with its locally modified and cropped replications over a frame generated by the film grain template generator 202 and the corresponding mean-subtracted frame 212 of the decoded video asset 122. Other local similarity correlation expressions may be used in fitting the conditions stated above. For example, other correlation measures may be a dot product correlation measure or a Spearman’s rank correlation measure or a Pearson correlation coefficient measure. To further illustrate local similarity correlation frame more discussion using the absolute cosine similarity correlation expressions is provided further into this specification.

[0037] To better illustrate the advantages of the process to detect and reduce or eliminate a perceptible repetitive film grain pattern the local absolute cosine correlation measurement is used as the local similarity correlation measurement in the description in the examples below. Therefore, certain blocks such as block 214 and 216 described in FIGS. 2A, 2B and block 612 in FIG. 6 use the terminology “Local Similarity Correlation” are referred to in the descriptive examples below using the terminology “Local Absolute Cosine Similarity”.

[0038] The film grain template generator 202 may be configured to generate the film grain template 126, consistent with how the film grain synthesizer 124 generates the film grain template 126 for generation of the decoded video with synthesized film grain 128. A sample film grain template 126 is illustrated in FIG. 3.

[0039] Continuing with FIG. 2A and 2B, the mean filter 206 is applied to the input frame 204 (whether received as encoded and decoded or received in a format where decoding is not required) to generate a mean frame 208. The mean filter 206 process may use the luma channel of a nonlinear color space (e.g. Y’ Cb’ Cr’) or luminance channel if using a linear color space (e.g., Y CbCr) of the extracted input frame 204 in generating the mean frame 208. Other channels, such as chrominance and Chroma may be used, but it should be noted that chrominance channels and chroma channels may have a lower density of structural information in the blocks as compared to the luminance channel and luma channel. The mean filter 206 may be implemented as a smoothing filter that replaces the brightness of a pixel by the mean of the brightness of the pixel itself and its neighbors.

[0040] As noted above, local adaptation is the process of choosing the relevant film grain template 126 and scaling according to the characteristics of the pixel values of the smoothed input frame. The local adaptation may include amplifying / scaling the film grain template 126 according to the pixel values of the smoothed frame. This local adaptation may follow following equation (1): r = T + / (T)GL(1) whereY is the smoothed frame pixel,GLis the luma film grain sample,Y' is the reconstructed frame pixel that contains film grain, and f is a piecewise linear function of Y.The scaling function f might be different for each individual frame. Accordingly, the film grain model parameters 112 may specify pivot points of the scaling function to enable the reconstruction of the scaling function by the decoder 120.

[0041] As shown in FIG. 2A and 2B, the absolute cosine similarity is to be found between GLand Y' in local regions. However, according to the equation (1) above, this cosine similarity (correlation) might be simply perturbed by Y and f(Y). A closer look at Y and f(Y) is taken to explore their effect on the cosine similarity between GLand Y' in details.

[0042] Referring again to the definition of f (T) as a piecewise linear function, it can be seen that this approach benefits performing the GraPQA because of the visual masking effect of Human Visual Systems (HVS). The visual masking effect of HVS states that high-contrast (e.g., textured)regions of an image provide a masking effect for underlying errors and noises, such as noises and distortions are more visible in low-contrast and uniform regions of the image. This is consistent with empirical observations of grain repetitive patterns, which confirms that the repetitive patterns are only visible in flat (low-contrast) regions of the input frames 204.

[0043] Referring to the definition of / (K) as a piecewise linear function, it can be understood that: (1) for flat (low-contrast) regions of the frame, most pixels of the smoothed frame ( / ) are close to each other in their values and end up receiving (approximately) the same scaling factor according to the / (K) function. Therefore, most of the local grains GLare linearly scaled, and the absolute cosine similarity between the grain template and local regions of the input frame 204 is hence mostly preserved; (2) for textured (high-contrast) regions of the frame, most pixels of the smoothed frame ( / ) are highly different from each other in their values and end up receiving hugely different scaling factors according to the / (K) function, which drastically perturbs the absolute cosine similarity between GLand Y' to the benefit of the GraPQA.

[0044] As noted above, the mean filter 206 is applied to the input frame 204. This is done to counteract the effect of Y on the absolute cosine similarity between the film grain template 126 and local regions of the input frame 204 of the video asset 102. The aim of this operation is to obtain an approximation of the underlying smoothed frame Y considering that the film grain template 126, GL, is a zero-mean pseudo-random noise. This approximation is required because there is no direct access to the smoothed frame pixel Y.

[0045] The mean frame 208 is then normalized by the subtractor 210 by subtracting the mean frame 208 from the input frame 204. This is accomplished to obtain an approximation of f(Y)~ GL, herein called the mean- subtracted frame 212.

[0046] The validity of this approximation and effect of this operation on the absolute cosine similarity for both flat and textured local regions of the input frame 204 may be observed. In flat regions, most pixels of the smoothed frame ( / ) are close to each other in their values, and the proposed normalization provides a good approximation of f(Y)~ GL, which results in a high value for the absolute cosine similarity between the film grain template 126 and flat local regions of the input frame 204. However, in textured regions, pixels of the smoothed frame ( / ) are highlyvarying, and the mean-subtracted frame 212 is not a good approximation of f(Y)GL, which in turn attenuates the absolute cosine similarity between the film grain template 126 and textured local regions of the input frame 204. As a result, the mean filter 206 operation benefits the GraPQA as it is consistent with what is observed with the visual masking effect of HVS.

[0047] The mean-subtracted frame 212 and the film grain template 126 may then be applied to a local absolute cosine similarity measurement 214. Given the mean-subtracted frame 212 and the generated film grain template 126, the film grain template 126 may be moved over the mean- subtracted frame 212 in a sliding window manner. For each local window position, the absolute cosine similarity may be found between the film grain template 126 and the underlying local region of the mean-subtracted frame 212. This computation may be performed according to the following equation (2):Here, G_ and X_ represent flattened ID vectors of the 2D grain template and local region of the mean-subtracted frame 212, respectively. The Frobenius norm of the local region of the mean- subtracted frame 212, i.e. may be obtained by sliding an all-ones kernel of the same size ofthe film grain template 126 over the squared mean-subtracted frame 212 and taking the square root of the result. Based on the mean frame 208 output from the mean filter 206 and the definition of the absolute cosine similarity, it should be expected to obtain high similarity values in flat regions of the mean frame 208 that contain (cropped and scaled) repetitive patterns of the film grain template 126. The results of these sliding window operations may be combined into a resultant absolute cosine similarity frame 216 for further processing. An example absolute cosine similarity frame 216 is shown in the example of FIG. 3.

[0048] Continuing with FIG. 2 A and 2B, the absolute cosine similarity frame 216 may be provided to a scoring algorithm 218 for evaluation. The scoring algorithm 218 may utilize the information in the absolute cosine similarity frame 216 to compute a scalar GraPQA score 220 corresponding to the input frame 204 using the film grain model parameters 112.

[0049] In an example, the scoring algorithm 218 may receive the absolute cosine similarity frame 216 and locate the N highest value pixels of the absolute cosine similarity frame 216 (e.g., the 100 greatest magnitude of absolute cosine similarity pixel values). Using those pixels, the scoring algorithm 218 may calculate an average of their pixel values (e.g., of the largest or brightest pixels along a scale from 0 to 1), and subtract this average from a highest value (e.g., from 1). This results in a GraPQA score 220 that is in the of range [0, 1], A lower value of the GraPQA score 220 indicates the detection of more instances of the repetitive film grain pattern and hence a lower perceptual quality. To obtain a grain pattern quality score for the overall video asset 102, the GraPQA score 220 may be calculated for all (or a representative subset) of the input frames 204 and then averaged over the quantity of scored input frames 204 of the encoded video asset 116.

[0050] The data flow 200 in FIG. 2B may be utilized at the end of an end-to-end system 100 to optimize film grain resynthesis. For example, the GraPQA score 220 may be improved in an iterative approach by modifying the film grain template 126 (e.g., block size) created in the film grain synthesizer 124 and regenerate an improved GraPQA score 220. In this approach the processor in the player or set top box or user device may be updated with the process outlined in FIG. 2B. Another implementation of the data flow 200 of FIG. 2B may be to configure the processor at the end of the end-to-end system 100 to provide information related to the GraPQA score 220 or user device type being fed back to the processor at the beginning of the end-to-end process to modify a film grain model parameter 112 such that the change leads to an improved GraPQA score 220 determined by the processor at the end of the end-to-end system 100.

[0051] The data flow 200 of FIG. 2B may be implemented in the processor at the beginning of the end-to-end system 100 to help anticipate or predict a GraPQA score 220 to get a better GraPQA score 220 by the processor at the end of the end-to-end system 100. For example, the processor at the beginning of the end-to-end system 100 is configured to use data flow 200 in FIG. 2B where the encoded video asset 116 determined at the beginning of the end-to-end system 100 is decoded and a film grain template 126 generated to determine a predicted GraPQA score 220 such that film model parameters and encoding parameters at the beginning of the end-to-end system 100 may be optimized. This approach could be applied based on a categorized sequence of frames as opposed to applying on a frame-by-frame basis. A categorized sequence of frames can refer to multipleframes related to a common image theme such as images of faces or a snow scene or a cloud scene or a grass field scene or a low detail scene or a high detail scene are a few examples.

[0052] The data flow 200 explained in FIG. 2A and 2B, may include three major time-consuming operations: (1) 2D convolution between the mean filter 206 and the input frame 204; (2) 2D convolution between the squared mean-subtracted frame 212 and an all-ones kernel when performing the local absolute cosine similarity measurement 214; and (3) 2D convolution between the mean- subtracted frame 212 and the film grain template 126. Given the high resolution of the input frames 204 and the film grain template 126, running the proposed method in the 2D spatial domain is very slow and time consuming, on the order of — 30s per 4K frame with a film grain template 126 of 73x82 image pixel resolution on a GeForce RTX 3060 GPU. Thus, to improve performance, these 2D convolutions may be implemented as 2D fast Fourier transform (FFT) operations and run in the frequency domain, which enhances the efficiency of the scoring. This results in a 2 frames per second (FPS) processing speed for an input frame 204 of 4K resolution and with a film grain template 126 of 73x82 resolution on a GeForce RTX 3060 GPU.

[0053] FIG. 3 shows an example 300 the effectiveness of the data flow 200 in detecting film grain patterns in a sample input frame 204 of a video asset 102. A film grain block size of 64x64 image pixels has been used by the film grain template generator 202 to synthesize the film grain template 126 for the input frame 204. Below the input frame 204, the absolute cosine similarity frame 216 is shown for the same position and size of local window 302. In the local window 302, a set of largest pixels 304 of the local window 302 of the absolute cosine similarity frame 216 are highlighted. Other such largest pixels 304 can be seen in horizontal rows across the absolute cosine similarity frame 216. In this illustration, absolute cosine similarity frame 216 and the local windows 302 are enhanced in terms of their contrast for better visualization of repetitive pattern detections.

[0054] Several observations may be made on the example 300. First, the film grain template 126 has been partially replicated in the local window 302 of the input frame 204, and the approach successfully detects replications of the film grain template 126 in the local window 302 by the cosine similarity measurements. Second, the repetitive film grain patterns in the input frame 204 are easily visible when observed on a HDR display, and the detections obtained by this approachare in line with those observations. Third, it can be seen that the approach only detects film grain patterns in flat regions (e.g., in the sky, on the building’s door) and ignores grain patterns in high- textured regions (e.g., on the train, on the railway) by design. This aligns with the visual masking effect of HVS and empirical observations of the input frame 204.

[0055] FIG. 4 illustrates an example graph 400 of the relationship of GraPQA scores 220 for five different image pixel block sizes (64x64, 32x32, 16x16, 8x8, 4x4) of the film grain template 126 for three input frames 204 of four different encoded video assets 116 labeled Cl, C2, C3, and C4. As shown, decreasing the film grain block size results in less significant film grain repetitive pattern detections. Thus, film grain block sizes have a strong effect on the look and feel of the synthesized film grain 128 based on the film grain template 126 and thus on the decoded video with synthesized film grain 128. Specifically, larger block sizes (increasing along the X-Axis as shown) create more visible and hence more annoying film grain repetitive patterns (resulting in lower GraPQA scores 220 as shown along the Y-Axis). These patterns become less apparent as the film grain block sizes are decreased. It can be observed in this example that the difference between the GraPQA scores 220 for smaller block sizes (e.g., 4x4 vs. 8x8) is negligible. However, the difference in the perceived quality increases with increase in the film grain block size.

[0056] FIG. 5 illustrates an example of detecting film grain patterns in a sample input frame 204 of a video asset 102 for three different block sizes. In the topmost absolute cosine similarity frame 216A, a block size of 64x64 image pixels is used, in the middle absolute cosine similarity frame 216B, a block size of 16x16 image pixels is used, and in the bottom absolute cosine similarity frame 216C, a block size of 4x4 image pixels is used.

[0057] A local window 302 A shows an example closeup of a portion of the absolute cosine similarity frame 216A, a local window 302B shows an example closeup of the same position in the absolute cosine similarity frame 216B, and a local window 302C shows an example closeup of the same position in the absolute cosine similarity frame 216C. These frames 216A, 216B, 216C and local windows 302 A, 302B, 302C are enhanced in terms of their contrast for better visualization of repetitive pattern detections.

[0058] Again, it can be seen that the scores increase as the block size decreases. Using the disclosed approach, the GraPQA score 220 for the absolute cosine similarity frame 216A iscomputed to be 0.6370, the GraPQA score 220 for the absolute cosine similarity frame 216B is computed to be 0.8644, and the GraPQA score 220 for the absolute cosine similarity frame 216C is computed to be 0.9040.

[0059] FIG. 6 illustrates an example process 600 for the determination of the GraPQA scores 220 for a video asset 102. In an example, the process 600 may be performed by one or more computing devices 702 as discussed in detail herein. It should be noted that while the operations are shown as being performed in a specific ordering, one or more operations may be performed concurrently and / or in a different ordering than shown.

[0060] At operation 602, a video asset can be an encoded video asset 116 received. The encoded video asset 116 may be any of a live video feed from current events, a pre-recorded show or movie, an advertisement, or another clip or other video clip. The encoded video asset 116 may include just video in some examples, but in many cases the encoded video asset 116 further includes additional content such as audio, subtitles, and metadata information descriptive of the content and / or format of the video. The encoded video asset 116 may comprise a series of one or more frames to be displayed in succession, which may be encoded in any of various resolutions, frame rates, dynamic ranges, and color spaces.

[0061] At operation 604, based on the received decoded video asset 122 and metadata, the decoder 120 extracts a decoded input frame 204 from the decoded video asset 122 and extracts film grain model parameters 112 from the metadata. Luma channel or luminance channel information of the extracted decoded input frame 204 can be identified for mean filter processing in operation 606. Channel information other than the Luma channel and luminance information can be used such as chroma or chrominance information, however, it should be noted that chroma channels and chrominance channels may have a lower density of structural information in the blocks as compared to the luma channel and luminance channel.

[0062] At operation 606, a mean filtering is performed to the input frame 204 by the mean filter 206 to generate a mean frame 208. The mean filter 206 may be implemented as a smoothing filter that replaces the brightness of a pixel by the mean of the brightness of the pixel itself and its neighbors. In an example, the mean filter 206 is set with the same filtering area size as a pixel block area in the film grain template 126.

[0063] At operation 608, e.g., following operation 604 and 606, the mean frame 208 from block 606 is subtracted from the input frame 204 from block 604. The mean frame 208 is thus normalized by its subtraction by the subtractor 210 from the input frame 204. This is accomplished to obtain an approximation of f(Y)~ GL, herein called the mean-subtracted frame 212.

[0064] At operation 610, a film grain template 126 is generated using the film grain template generator 202. In an example, the film grain template generator 202 generates the film grain template 126, consistent with how the film grain synthesizer 124 generates the film grain template 126 for generation of the decoded video with synthesized film grain 128. The film grain model parameters 112 used in the generated film grain template 126 may have been determined by an encoder 114 and therefore extracted from the metadata of the encoded video asset 116 to inform the film grain template generator 202. In other examples, different film grain model parameters 112 may be supplied by another source such as from a data base or a user and / or by an algorithmic process 618 that searches among different film grain model parameters 112 to find the best scoring film grain model parameters 112. If the image frame does not have associated film modelling parameters, then film modelling parameters from another source would have to be provided as shown in FIG. 2A.

[0065] At operation 612, local similarity correlation measurements are performed between the mean-subtracted frame 212 and a film grain template 126. The local similarity measurement can be an absolute local cosine similarity measurement which may be performed by sliding the film grain template 126 over the mean-subtracted frame 212 to identify a local window 302, and computing the absolute cosine similarity between the film grain template 126 and the local window 302 of the mean- subtracted frame 212. In an example, the sliding window approach is performed using a sliding local window 302 size corresponding to a grain block size of the film grain template 126.

[0066] At operation 614, the scoring algorithm 218 is used to compute the GraPQA score 220 for the input frame 204. In an example, the scoring algorithm 218 may receive the local similarity correlation frame 216 which can be an absolute cosine similarity frame from the output of operation 612 and locate the N highest value pixels 304 of the frame (e.g., the 100 highest value pixels). Using those pixels, the scoring algorithm 218 may calculate an average of their pixelvalues, and subtract this average from a highest value (e.g., from 1). This results in a GraPQA score 220 that is in the of range [0, 1], A lower value of the GraPQA score 220 indicates the detection of more instances of the repetitive film grain pattern and hence a lower perceptual quality. To obtain a grain pattern quality score for the overall decoded video asset 122, the GraPQA score 220 may be calculated for all (or a representative subset) of the input frames 204 and then averaged over the quantity of scored decoded frames 204 of the decoded video asset 122.

[0067] At operation 616, if there are additional frames, the process 600 returns to operation 604 to extract an additional input frame 204 to continue the scoring. If not, the process 600 ends.

[0068] FIG. 7 illustrates an example 700 of a computing device 702 for use in the determination of the GraPQA scores 220 for a decoded video asset 122. Referring to FIG. 7, and with reference to FIGS. 1-6, the devices and modules discussed herein may be examples of such computing devices 702. For instance, the operations performed by the denoiser 106, film grain model parameterization 110, encoder 114, network 118, decoder 120, film grain synthesizer 124, viewer device 130, film grain template generator 202, mean filter 206, subtractor 210, local absolute cosine similarity measurement 214, and scoring algorithm 218, etc., as well as the operations discussed in the data flow 200 and in the process 600 may be performed by such computing devices 702. As shown, the computing device 702 includes a processor 704 that is operatively connected to a storage 706, a network device 708, an output device 710, and an input device 712. It should be noted that this is merely an example, and computing devices 702 with more, fewer, or different components may be used.

[0069] The processor 704 may include one or more integrated circuits that implement the functionality of a central processing unit (CPU) and / or graphics processing unit (GPU). In some examples, the processors 704 are a system on a chip (SoC) that integrates the functionality of the CPU and GPU. The SoC may optionally include other components such as, for example, the storage 706 and the network device 708 into a single integrated device. In other examples, the CPU and GPU are connected to each other via a peripheral connection device such as peripheral component interconnect (PCI) express or another suitable peripheral data connection. In one example, the CPU is a commercially available central processing device that implements aninstruction set such as one of the x86, ARM, Power, or microprocessor without interlocked pipeline stage (MIPS) instruction set families.

[0070] Regardless of the specifics, during operation the processor 704 executes stored program instructions that are retrieved from the storage 706. The stored program instructions, accordingly, include software that controls the operation of the processors 704 to perform the operations described herein. The storage 706 may include both non-volatile memory and volatile memory devices. The non-volatile memory includes solid-state memories, such as not and (NAND) flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the system is deactivated or loses electrical power. The volatile memory includes static and dynamic random-access memory (RAM) that stores program instructions and data during operation of the end-to-end system 100.

[0071] The GPU may include hardware and software for display of at least two-dimensional (2D) and optionally 3D graphics to the output device 710. The output device 710 may include a graphical or visual display device, such as an electronic display screen, projector, printer, or any other suitable device that reproduces a graphical display. As another example, the output device 710 may include an audio device, such as a loudspeaker or headphone. As yet a further example, the output device 710 may include a tactile device, such as a mechanically raisable device that may, in an example, be configured to display braille or another physical output that may be touched to provide information to a user.

[0072] The input device 712 may include any of various devices that enable the computing device 702 to receive control input from users. Examples of suitable devices that receive human interface inputs may include keyboards, mice, trackballs, touchscreens, voice capture devices, graphics tablets, and the like.

[0073] The network devices 708 may each include any of various devices that enable the devices to send and / or receive data from external devices over networks. Examples of suitable network devices 708 include an Ethernet interface, a Wi-Fi transceiver, a cellular transceiver, or a BLUETOOTH or Bluetooth Low Energy (BLE) transceiver, ultra-wideband (UWB) transceiver, or other network adapter or peripheral interconnection device that receives data from anothercomputer or external data storage device, which can be useful for receiving large sets of data in an efficient manner.

[0074] The processes, methods, or algorithms disclosed herein can be deliverable to / implemented by a processing device, controller, or computer, which can include any existing programmable electronic control unit or dedicated electronic control unit. Similarly, the processes, methods, or algorithms can be stored as data and instructions executable by a controller or computer in many forms including, but not limited to, information permanently stored on non- writable storage media such as read-only memory (ROM) devices and information alterably stored on writable storage media such as floppy disks, magnetic tapes, compact discs (CDs), RAM devices, and other magnetic and optical media. The processes, methods, or algorithms can also be implemented in a software executable object. Alternatively, the processes, methods, or algorithms can be embodied in whole or in part using suitable hardware components, such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), state machines, controllers or other hardware components or devices, or a combination of hardware, software and firmware components.

[0075] While exemplary embodiments are described above, it is not intended that these embodiments describe all possible forms encompassed by the claims. The words used in the specification are words of description rather than limitation, and it is understood that various changes can be made without departing from the spirit and scope of the disclosure. As previously described, the features of various embodiments can be combined to form further embodiments of the invention that may not be explicitly described or illustrated. While various embodiments could have been described as providing advantages or being preferred over other embodiments or prior art implementations with respect to one or more desired characteristics, those of ordinary skill in the art recognize that one or more features or characteristics can be compromised to achieve desired overall system attributes, which depend on the specific application and implementation. These attributes can include, but are not limited to strength, durability, life cycle, marketability, appearance, packaging, size, serviceability, weight, manufacturability, ease of assembly, etc. As such, to the extent any embodiments are described as less desirable than other embodiments or prior art implementations with respect to one or more characteristics, these embodiments are not outside the scope of the disclosure and can be desirable for particular applications.

Claims

WHAT IS CLAIMED IS:

1. A method for performing grain pattern quality assessment (GraPQA) of an input frame, comprising: creating a mean-subtracted frame by performing mean filtering of the input frame to generate a mean frame and subtracting the mean frame from the input frame; generating a film grain template based on a film grain model parameter of the input frame; performing local similarity correlation measurements between the mean- subtracted frame and the film grain template to create a local similarity correlation frame; and generating a GraPQA score for the input frame based on the local similarity correlation frame.

2. The method of claim 1, wherein the input frame is a digital image frame where synthesized film grain is to be added and the film grain model parameter is received from a source other than the input frame.

3. The method of claim 1, wherein the input frame is an encoded frame that is decoded, and the film grain model parameter is decoded from the encoded input frame.

4. The method of claim 1, wherein the mean filtering is performed on a luma channel or luminance channel of the input frame.

5. The method of claim 1, further comprising generating the film grain template for a luma channel or luminance channel of the input frame as a predefined fixed grain block size.

6. The method of claim 1, wherein the local similarity correlation measurements are performed by: sliding the grain template over the mean-subtracted frame to identify a local window; andcomputing the similarity correlation between the film grain template and the local window.

7. The method of claim 6, wherein the sliding is performed using a sliding window size corresponding to a grain block size of the film grain template.

8. The method of claim 6, wherein the local similarity correlation is an absolute local cosine similarity measurement that is performed according to the following:where G_ represents a flattened single dimension vector of the film grain template and X represents the local window of the mean- subtracted frame.

9. The method of claim 8, further comprising: locating a predefined quantity of highest value pixels of the local similarity correlation frame; calculating an average value of the highest value pixels; and scaling the average value into a predefined range of values resulting in the GraPQA score. wherein the predefined quantity is 100, and wherein the GraPQA score is scaled in the range [0, 1],10. The method of claim 1 , wherein one or more of a 2D convolution between the mean frame and the input frame, a 2D convolution between the mean-subtracted frame in a squared form and an all-ones kernel, and a 2D convolution between the mean-subtracted frame and the grain template are performed as 2D fast Fourier transformations in the frequency domain.

11. The method of claim 1, further comprising: locating a predefined quantity of highest value pixels of the local similarity correlation frame;calculating an average value of the highest value pixels; and scaling the average value into a predefined range of values resulting in the GraPQA score.

12. The method of claim 11, wherein the predefined quantity is 100, and wherein the GraPQA score is scaled in the range [0, 1],13. The method of claim 1, wherein a lower value of the GraPQA score indicates detection of more instances of repetitive film grain pattern and hence lower perceptual quality, and a higher value of the GraPQA score indicates fewer instances of the repetitive film grain pattern and hence higher perceptual quality.

14. The method of claim 1, wherein the input frame is a frame of a video, and further comprising: repeating the GraPQA score for a plurality of frames of the video; and averaging the GraPQA scores over the plurality of frames of the video to obtain an overall pattern quality score for the video.

15. A system for performing grain pattern quality assessment (GraPQA), comprising: one or more hardware computing devices configured to: perform mean filtering of an input frame to generate a mean frame; subtract the mean frame from the input frame to create a mean- subtracted frame; generate a film grain template based on a film grain model parameter of the input frame; perform local similarity correlation measurements between the mean- subtracted frame and the film grain template to produce a local similarity correlation frame; and score results of the local similarity correlated frame to generate a GraPQA score for the input frame.

16. The system of claim 15, wherein the one or more hardware computing devices are further configured to process the input frame that is a digital image frame where synthesized film grain is to be added and the film grain model parameter is received from a source other than the input frame.

17. The system of claim 15, wherein the one or more hardware computing devices are further configured to decode an encoded input frame wherein the film grain model parameter is decoded from the encoded input frame.

18. The system of claim 15, wherein the one or more hardware computing devices are further configured to perform mean filtering on a luma channel or luminance channel of the input frame.

19. The system of claim 15, wherein the one or more hardware computing devices are further configured to generate the film grain template for a luma channel or luminance channel of the input frame as a predefined fixed grain block size.

20. The system of claim 15, wherein the one or more hardware computing devices are further configured to perform the local similarity correlation measurements by operations including to: slide the grain template over the mean- subtracted frame to identify a local region; and compute the locale similarity correlation between the film grain template and the local region.

21. The system of claim 20, wherein the slide is performed using a sliding window size corresponding to a grain block size of the film grain template.

22. The system of claim 20, wherein the local similarity correlation measurement is an absolute local cosine similarity measurement performed according to the following:where G_ represents a flattened single dimension vector of the film grain template and X_ represents the local region of the mean-subtracted frame.

23. The system of claim 22, wherein the one or more hardware computing devices are further configured to: locate a predefined quantity of highest value pixels of the local similarity correlation frame; calculate an average value of the highest value pixels; and scale the average value into a predefined range of values resulting in the GraPQA score, wherein the predefined quantity is 100, and wherein the GraPQA score is scaled in the range [0, 1],24. The system of claim 15, wherein one or more of a 2D convolution between the mean frame and the input frame, a 2D convolution between the mean-subtracted frame in a squared form and an all-ones kernel, and a 2D convolution between the mean-subtracted frame and the grain template are performed as 2D fast Fourier transformations in the frequency domain.

25. The system of claim 15, wherein the one or more hardware computing devices are further configured to: locate a predefined quantity of highest value pixels of the local similarity correlation frame; calculate an average value of the highest value pixels; and scale the average value into a predefined range of values resulting in the GraPQA score.

26. The system of claim 25, wherein the predefined quantity is 100, and wherein the GraPQA score is scaled in the range [0, 1],27. The system of claim 15, wherein a lower value of the GraPQA score indicates detection of more instances of repetitive film grain pattern and hence lower perceptual quality, and a higher value of the GraPQA indicates fewer instances of the repetitive film grain pattern and hence higher perceptual quality.

28. The system of claim 15, wherein the input frame is a frame of a video, and wherein the one or more hardware computing devices are further configured to: repeat the GraPQA score for a plurality of frames of the video; and average the GraPQA scores over the plurality of frames of the video to obtain an overall pattern quality score for the video.

29. The system of claim 15, wherein the system is integrated as an upgrade into a processor of an end-to-end system for delivering video.

30. The system of claim 29, wherein the processor is located before a network of the end-to-end system configured for transmission of encoded video assets.

31. The system of claim 29, wherein the processor is located after a network of the end-to-end system configured for transmission of encoded video assets.

32. A non- transitory computer-readable medium comprising instructions for performing GraPQA assessment of repetitive film grain patterns that, when executed by one or more hardware computing devices, cause the one or more hardware computing devices to perform operations including to: perform mean filtering of an input frame to generate a mean frame; subtract the mean frame from the input frame to create a mean- subtracted frame; generate a film grain template based on film grain model parameters of the input frame; perform local similarity correlation measurements between the mean- subtracted frame and a film grain template to produce a local similarity correlation frame; andscore results of the local similarity correlated frame to generate a GraPQA score for the input frame.

Citation Information

Patent Citations

  • Film grain simulation technique for use in media playback devices

    US20060133686A1

  • Saliency-weighted video quality assessment

    US20170154415A1

  • Metadata signaling and conversion for film grain encoding

    WO2022212792A1

  • Systems and methods for grain-aware video coding and transmission

    WO2024214077A1