Systems and methods for assessing perceptual quality of synthesized film grain
Patent Information
- Application Number
- PCT/IB2025/052353
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-05
- Filing Date
- 2025-03-04
- Publication Date
- 2025-10-02
AI Technical Summary
Existing video encoding technologies fail to accurately preserve the artistic intent of film grain, leading to inefficient bit usage and compromised fidelity in synthesized film grain due to the lack of a robust objective film grain score.
An objective model using a data-driven approach with a neural network for film grain similarity assessment, providing a film grain synthesis score based on statistical measurements of source and test features to optimize the perceptual quality of synthesized grain.
The model effectively aligns with human perception, enabling improved fidelity and efficient encoding of synthesized film grain, reducing computational burden and enhancing artistic integrity.
Smart Images

Figure IB2025052353_02102025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR ASSESSING PERCEPTUAL QUALITY OF SYNTHESIZED FILM GRAINCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. provisional application Serial No. 63 / 561,530 filed March 5, 2024, the disclosure of which is hereby incorporated in its entirety by reference herein.TECHNICAL FIELD
[0002] Aspects of the disclosure generally relate to perceptual quality assessment in relation to objective film grain synthesis.BACKGROUND
[0003] Film grain is essentially specific random noise in a video or still image frame. Film grain may originate from the analog film acquisition processes. Film grain may also be synthesized purposely and digitally in the content production and post-production pipelines. In either case, the visual feel of the film grain may be considered to be part of the artistic or creative intent of the content producers. Preserving such creative intent in video distribution may be costly for video encoders, which treat film grains no differently from other types of noise. This may cause the encoder to consume a large number of bits to encode the film grain to assist in efficiently preserving the creative intent of the content producers.SUMMARY
[0004] There is a need to assess the perceptual appearance of the synthesized grain as compared to the original film grain or to a reference more indicative of the original film grain. Based on this fidelity assessment, the grain synthesis parameters for the addition of the simulated film grain may be adjusted to improve the perceptual appearance of the synthesized grain. To perform this assessment, an objective model is introduced that can produce a film grain assessment metric as described herein.
[0005] In one or more illustrative examples, a method for perceptual quality assessment of objective film grain synthesis is performed. A source frame with source regions of interest and a corresponding test frame with corresponding test regions of interest, the source frame including native film grain, the test frame including simulated film grain. The source frame with the source regions of interest and the corresponding test frame with the corresponding test regions of interest are applied to a feature extraction using a neural network trained on image data for classification to produce source features and corresponding test features. Statistical measures are applied to the source features and the corresponding test features to determine a distance between statistical measurements of the source features of the source regions of interest and the corresponding test features of the corresponding test regions of interest, wherein the distance between the statistical measurements is used to produce a model score, and a film grain synthesis score is based on the model score.
[0006] In one or more illustrative examples, a system for performing perceptual quality assessment of objective film grain synthesis includes one or more hardware computing devices that provide a decoded and synthesized grain video asset to a viewing device and that are configured to determine a model score for a perceptual quality assessment of objective film grain synthesis. The one or more hardware computing devices are configured to receive a source frame with source regions of interest and a corresponding test frame with corresponding test regions of interest, the source frame including native film grain, the test frame including simulated film grain; apply the source frame with the source regions of interest and the corresponding test frame with the corresponding test regions of interest to a feature extraction using a neural network trained on image data for classification to produce source features and corresponding test features; and apply statistical measures to the source features and the corresponding test features to determine a distance between statistical measurements of the source features of the source regions of interest and the corresponding test features of the corresponding test regions of interest, wherein the distance between the statistical measurements is used to produce a model score, and a film grain synthesis score is based on the model score.
[0007] In one or more illustrative examples, a non-transitory computer-readable medium includes instructions for performing perceptual quality assessment of objective film grain synthesis that, when executed by one or more hardware computing devices, cause the one or more hardware computing devices to perform operations including to receive a source frame with source regions of interest and a corresponding test frame with corresponding test regions of interest, the source frame includingnative film grain, the test frame including simulated film grain; apply the source frame with the source regions of interest and the corresponding test frame with the corresponding test regions of interest to a feature extraction using a neural network trained on image data for classification to produce source features and corresponding test features; and apply statistical measures to the source features and the corresponding test features to determine a distance between statistical measurements of the source features of the source regions of interest and the corresponding test features of the corresponding test regions of interest, wherein the distance between the statistical measurements is used to produce a model score, and a film grain synthesis score is based on the model score.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 illustrates an example of an end-to-end system for grain-aware video coding and transmission;
[0009] FIG. 2 illustrates an example data flow for a process of grain-aware video coding and transmission;
[0010] FIG. 3 illustrates an example data flow for the determination of a film grain synthesis (FGS) score for a test video in comparison to a source video;
[0011] FIG. 4 illustrates further aspects of the operation of the model in determining the FGS score from the source regions of interest and the test regions of interest;
[0012] FIG. 5 illustrates an example method for using the model for performing the image quality analysis to determine the FGS score;
[0013] FIG. 6 illustrates an example of use of the FGS score on a plurality of test regions of interest corresponding to a source region of interest; and
[0014] FIG. 7 illustrates an example of a computing device for use in the end-to-end system for grain- aware video coding transmission and for determining the FGS scores.DETAILED DESCRIPTION
[0015] As required, detailed embodiments of the present invention are disclosed herein; however, it is to be understood that the disclosed embodiments are merely exemplary of the invention that may be embodied in various and alternative forms. The figures are not necessarily to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present invention.
[0016] Though digital cinematography has progressed, many artists favor film rolls for their unique texture and essence. Some artists consider film grain to be one of the integral aspects to their artistic expression. Yet, over-the-top (OTT) providers and streamers have challenges with the high-entropy signal inherent in film grain which is unfriendly to compression.
[0017] Limited bandwidth by such providers and streams prompts various efforts to compress videos. One strategy to preserve film grain involves removing film grain at the source, encoding the video for transport, transporting the video, decoding the video at its destination, and resynthesizing the film grain that was removed after decoding. To allow for the synthesizing at the destination, film grain synthesis parameters may be injected into the bit stream as metadata. Codecs such as AOMedia Video 1 (AVI) and Versatile Video Coding (WC) offer such solutions but may compromise grain fidelity without an accurate objective film grain score. Additionally, existing AVI or WC implementations can fail in replicating the original film grain’s appearance due to the straightforward and simple methods described in the codecs’ standards. Although various approaches may demonstrate technical robustness, without considering the problem through the lens of perceptual video quality, the synthesized film grain may not meet the expectations of filmmakers.
[0018] Improving film grain synthesis models involves subjectively assessing the fidelity of the synthesized grain to the original film grain. Based on this analysis, synthesis parameters for the addition of simulated film grain may be adjusted to improve the fidelity of the synthesized grain.
[0019] In an example, mean opinion score (MOS) may be used to assess the simulated grain fidelity. MOS is a perceptual fidelity score representing the overall quality of a test video in relation to an original video. MOS may be determined as an arithmetic mean of individual values on a predefined scale that a subject assigns to his opinion of the performance of a system quality. As an illustrativeexample, the predefined scale may be from 1 to 5, with 1 being the worst and 5 being the best. In another such example, the predefined scale may be from 1 to 100. The MOS ratings may be gathered from users who take part in a subjective quality evaluation test. However, this subjective evaluation process is notably intricate. Moreover, it may be preferable to utilize a computed measure that does not require human evaluation.
[0020] An alternative method involves developing an objective film grain similarity model that may be computed algorithmically, but that closely aligns with human perception. Crafting and assessing such a metric poses genuine challenges, as conventional full-reference image or video quality measures relying on pixel-based distance are impractical in film grain synthesis scenarios. This may be due to the uncertainty of the spatial alignment between the synthesized and original grain in the source.
[0021] To address this, an objective model is introduced using a data-driven approach capable of gauging film grain similarity, demonstrating a high correlation with subjective studies. This metric may also be used to optimize auto-regression film grain synthesis. Subsequent subjective studies confirm that optimizing film grain synthesis (FGS) parameters based on this objective similarity metric results in a more faithful replication of the original film grain.
[0022] Perceptual video quality assessment may be performed in relation to objective film grain synthesis. An assessment of the perceptual video quality of the film grain synthesis may include receiving an example source frame, with native film grain, and comparing that source frame to various patches with synthesized film grain. The comparing may be performed using a metric that produces a score indicative of user perception of the patch as compared to the source frame’s native grain. Further aspects of the objective model approach are discussed in detail herein.
[0023] FIG. 1 illustrates an example of an end-to-end system 100 for grain-aware video coding transmission. In the illustrated example, a video delivery chain receives a video asset 102. The video asset 102 is provided to a pre-delivery processor 104, which in turn can provide an encoded video asset 114 to a content delivery network 106. The content delivery network 106 then provides a delivered encoded video asset 116 to a post-delivery processor 108, which in turn provides a decoded and synthesized grain video asset 118 to a viewer device 110 for display. A network monitor 112 maybe configured to communicate with a pre-delivery processor 104, content delivery network 106, and post-delivery processor 108. It should be noted that the end-to-end system 100 may be geographically diverse and that the calculations may occur co-located or in a distributed manner.
[0024] A video asset 102 may include, as some examples, live video feeds from current events, prerecorded shows or movies, and advertisements or other clips to be inserted into other video feeds. The video asset 102 may include just video in some examples, but in many cases the video asset 102 further includes additional content such as audio, subtitles, and metadata information descriptive of the content and / or format of the video. The video asset 102 may comprise a series of one or more frames to be displayed in succession, which may be encoded in any of various resolutions (e.g., 720p, 1080p, WUXGA, 2K, ultra-high definition (UHD), cinema 4K, 8K, etc.), frame rates (e.g., 24fps, 30fps, 60 fps, etc.), dynamic ranges (e.g., standard dynamic range (SDR), high dynamic range (HDR), and color spaces (Y’Cb’Cr’, RGB, etc.). The end-to-end system 100 may handle one or more sources of video assets 102. In general, when a video distributor receives source video, the distributor passes the video asset 102 through encoders, transcoders, packagers, origins, network connections, and consumer devices to ultimately present the video content to a user.
[0025] A pre-delivery processor 104 can be a software module or a unit that is a device or an assembly that receives a video asset 102. The pre-delivery processor 104 may be configured to determine aspects of the video asset 102 as well as prepare the video asset 102 for the grain- aware video coding. The video aspects may include, for example, detecting grain in the video asset 102, assess or model the grain in the video asset 102, reduce the grain in the video asset 102, and provide grain parameters with respect to the film grain in the video asset 102. Film grain can refer to randomly distributed noise across frames of the video asset 102 where this noise may be characterized by various parameters, such as grain size, grain density, and / or grain contrast. Grain aware video coding refers to a process of configuring the video asset 102 with film grain more efficiently for encoding of the video asset 102 or delivery of the coded video asset 102 to a content delivery network 106 (e.g., removal of the film grain). The pre-delivery processor 104 can output an encoded video asset 114 for transmission by a content delivery network 106 or the pre-delivery processor 104 can output the coded video asset 102 to be encoded externally for transmission by a content delivery network 106.
[0026] The content delivery network 106 may be configured to transmit the coded or encoded video asset 102 towards the viewer device 110. Configuring the content delivery network 106 may include, for example, receiving the coded or encoded video asset 114 from the pre-delivery processor 104, utilizing one or more encoders or transcoders to compress and / or re-encode the video content for transmission into a format that conforms with one or more standard video compression specifications. The content delivery network 106 may also use one or more packagers and / or origins to create segmented video files to be delivered to clients that then stitch the segments together to form a contiguous video stream. The content delivery network 106 can also be configured to perform the above operations in awareness of the grain in what the pre-delivery process outputs to the content delivery network 106.
[0027] As discussed herein, film grain generally refers to randomly distributed noise across frames of the video asset 102. This noise may be characterized by various parameters, such as grain size, grain density, and / or grain contrast.
[0028] The content delivery network 106 provides a delivered encoded video asset 116 to a postdelivery processor 108. The post-delivery processor 108 can be a software module or a unit that is a device or an assembly. The post-delivery processor 108 may be configured to perform aspects of the characterization and restoration of grain in the delivered encoded video asset 116. The characterization aspects may include, for example, to identify locations and parameters of the grain, parameters used in film grain removal, parameters used in compressing the video asset 102 with the grain removed, parameters associated with the transport of the encoded video asset 114 in compressed form, parameters used in the decoding to decompress the delivered encoded video asset 116, parameters created within synthesizing grain and / or re-graining to produce an output for the viewer device 110. The viewer device 110 may be a cinema display such as a light emitting cinema display in a theatre, cinema display in which an image is projected onto the display by a projector in a theatre, home image projector display, television display, mobile phone display, or other display device onto which the output from the post- delivery processor may be played back for viewing. For example, the postdelivery processor 108 can output a decoded and synthesized grain video asset 118. The post-delivery processor 108 can be integrated with the viewer device 110 or a stand alone unit such as a set-top box that provides a video asset 102 that is a decoded and synthesized grain video asset 118 to the userviewing device. The viewer device 110 can be a remote viewer device 110 located far (miles or thousands of miles) from the pre-delivery processor 104 location.
[0029] The network monitor 112 may be configured to monitor the video asset 102 as it is provided along the end-to-end system 100. In one example, the network monitor 112 may identify perceptual fidelity scores for the video asset 102 at various points along the end-to-end system 100. The components of the end-to-end system 100 may also be configured to perform additional aspects related to the characterization and restoration of grain in the grain-aware video coding and transmission, as discussed in detail herein. The network monitor 112 represents a functional block in hardware that can communicate, exchange or query needed and available information within the pre-delivery processor 104, the post-delivery processor 108 and / or the content delivery network 106. The network monitor 112 can provide further processing that can be used by the pre- delivery processor and the post-delivery processor 108. Processing associated with the network monitor 112 can also be done in the predelivery processor 104 device and the post-delivery processor 108 device. It should be noted that the end-to-end system 100 may be geographically diverse and that the calculations may occur co-located or in a distributed manner.
[0030] The decoded and synthesized grain video asset 118 can be rendered for a display by grain aware rendering operation if the display and the post-delivery are one and the same unit. If the display is a viewer device 110 separate from the post-delivery processor 108 unit, the post-delivery processor 108 unit can output the decoded and synthesized grain video asset 118 to the remote viewing device to be rendered by the remote viewing device or the remote viewing device may expose display rendering parameters to the post-delivery processor 108 unit to perform the grain aware rendering.
[0031] An existing pre-delivery processor 104 unit and an existing post-delivery processor 108 unit in an end-to-end system 100 can be upgraded to become capable of executing enhanced grain- aware processes disclosed. The upgrade can be a software upgrade to the pre-delivery processor 104 unit and the post-delivery processor 108 unit, or the upgrade can be a hardware upgrade. For example, an existing pre-delivery processor 104 unit is replaced by an upgraded pre-delivery processor 104 unit and an existing post-delivery processor 108 unit is replaced by an upgraded post- delivery processor unit.
[0032] Described within are flow charts detailing the flow of video image content through various operations and the interactions of the various operations specified within the pre- delivery processor unit and the post-delivery processor 108 unit of the intended end-to-end systems 100 for delivering video asset 102 with film grain to a viewing device to get an improved or optimized perceptual appearance of the synthesized grain.
[0033] FIG. 2 illustrates an example data flow 200 for a process of grain-aware video coding and transmission. In an example, the overall data flow 200 may be performed using the components of the end-to-end system 100. As shown, a media source provides the video asset 102 that passes through multiple stages of operation along the end-to-end system 100 before the media stream reaches the viewer device 110. The video asset 102 may refer to various types of video content as noted above, while the viewer device 110 may refer to one or more consumer devices. The operations performed along the data flow 200 may include grain detection 202, grain assessment / modeling 204, grain reduction 206, grain-aware encoding / rate control 208, grain-aware streaming / decoding 210, grain synthesis 212 (sometimes referred to as re-grain), and grain-aware rendering / display 214. Each of these is discussed in turn.
[0034] The grain detection 202 may include one or more processes to identify and characterize the amount of grain in the video asset 102. In an example, these processes may be performed by the predelivery processor 104. The grain detection 202 may perform various statistical analyses on the video asset 102, such as computing mean, variance, and distribution of image features to identify the presence of film grain. In another example, a frequency domain analysis may be performed to transform the video asset 102 into the frequency domain to analyze the power spectrum to identify the presence of film grain. In such an analysis, the film grain may present as a high-frequency pattern that appears as a peak in the power spectrum at high frequencies. In some examples, the grain detection 202 may be performed by using an edge detector to identify regions of interest in the video asset 102 having relatively consistent coloration (e.g., flat regions, regions with gradients, regions with low activity in terms of luminance variance, etc.). For instance, a mask of the regions of interest to analyze may be determined by the grain detection 202 as a grain detection result.
[0035] The grain assessment / modeling 204 may include one or more processes to determine the extent of the grain indicated by the grain detection 202. In an example, these processes may be performed bythe pre-delivery processor 104. The grain assessment / modeling 204 may identify the characteristics of the film grain in the video asset 102, which may be for the entire frame or only for regions of interest (such as flat regions, regions with gradients, regions with low activity in terms of luminance variance, etc.). The grain assessment / modeling 204 may be configured to provide an output indicative of the parameters of the grain that are located in the video asset 102. The grain model parameters may include parameters such as grain size, grain density, grain contrast, and / or variation with respect to differences in signal level.
[0036] In one example, the video asset 102 may be denoised, and the regions of the denoised version of the video asset 102 may be compared with the original video asset 102 to determine the parameters of the noise. For instance, these grain model parameters may include one or more of grain size, grain density, and / or grain contrast. In another example, the film grain may be modeled using root-meansquare (RMS) granularity, which is a numerical quantification of density non-uniformity, equal to the RMS fluctuations in optical density. In some examples, an autoregressive model may be used to characterize the film grain. In some examples, the film grain strength may vary with signal intensity, and the grain model parameters may further model these differences in level, e.g., as parameters of a linear function of luma (i.e. luminance).
[0037] The grain reduction 206 may be configured to reduce the grain that is present in the video asset 102. The grain reduction 206 may be performed in various ways by the pre-delivery processor 104. In an example, noise reduction filters may analyze the pixels in the video asset 102 to identify an average color or brightness and / or identify those pixels that are outliers. The filter may then apply a smoothing algorithm that averages the pixel values in the area, effectively reducing the noise or grain. In another example, a machine learning model may be trained on a large dataset of grainy and non-grainy images, to learn to identify and remove film grain.
[0038] The grain-aware encoding / rate control 208 may perform encoding on the video asset 102 after the grain reduction 206. In an example, the grain-aware encoding / rate control 208 may utilize an encoder of the pre-delivery processor 104 to convert the video asset 102 into a format for transmission along the end-to-end system 100. This video encoding may take into account the parameters about the types and properties of the grain, which may be included in the bit stream as metadata. The encoding may also include a rate and / or quality control operation to ensure that the encoding achieves a giventarget bitrate and / or quality level. Examples of the video encoding / rate control parameters may include the video bite rate, spatial resolution, frame rate or temporal resolution, encoding mode selections at video, frame, and local block levels, and the quantization step parameters at video, frame, and local block levels.
[0039] The grain-aware streaming / decoding 210 may perform transfer of the encoded video asset 114 along the end-to-end system 100 as well as decoding of the video asset 102 post transfer. Accordingly, the encoded video asset 114 may be streamed to the receiver side and decoded. In an example, the streaming aspect may be performed using components of the content delivery network 106, while the decoding aspect may be performed by the post-delivery processor 108.
[0040] The grain synthesis 212 may include one or more processes performed by the post-delivery processor 108 to decode the delivered encoded video asset 116 and add grain to produce a decoded and synthesized grain video asset 118. This added grain may be consistent with the grain as assessed by the grain assessment / modeling 204 and / or as reduced by the grain reduction 206. For example, the grain synthesis 212 may be configured to add noise characterized by various parameters, such as the grain size, grain density, and / or grain contrast noted above. In an example, these parameters may be the same as or consistent with the grain model parameters determined by the grain assessment / modeling 204.
[0041] The grain-aware rendering / display 214 may include providing the decoded and synthesized grain video asset 118 to the viewer device 110. By reincorporating the grain after compression, transmission, and decompression, the perceptual quality of the decoded and synthesized grain video asset 118 may be maintained at the viewer device 110, while also allowing the video asset 102 to be more efficiently encoded without the presence of the film grain.
[0042] While this process may aid in transmission of the video asset 102, it may be desirable have a fidelity assessment of the perceptual appearance of the synthesized grain as compared to the original film grain. Based on the fidelity assessment, the grain synthesis parameters for the addition of the simulated film grain may be adjusted to improve the perceptual appearance of the synthesized grain. To perform the fidelity assessment, an objective model is disclosed herein. One approach to optimize the perceptual appearance of the synthesized grain in an end-to-end system 100 can be done by usinga system in which the processes of the pre-delivery processor 104, content delivery network 106 and post-delivery processor 108 are known. For example, an end-to-end process where one has access and control of the elements within the end-to-end system 100. With a known end to end setup, the setup can be optimized for perceptual appearance of the synthesized grain by being able to maximize the effectiveness of integrating and performing the fidelity assessment using the objective model disclosed herein. Further optimization can be done by considering various end user viewer devices 110 when optimizing the end-to-end system 100. Having access to the pre-delivery processor 104 hardware, the content delivery network 106 and post-delivery processor 108 hardware allows the hardware and / or software changes to be made to better facilitate optimizing the end-to-end system 100 to improve perceptual appearance of the synthesized grain. The optimization of the end-to-end system 100 can also be updated to accommodate different video assets 102 such as video assets 102 with different film grain characteristics.
[0043] Another approach to optimize the perceptual appearance of the synthesized grain in an end-to- end system 100 can be done by configuring the video asset 102 with metadata information descriptive of film grain related parameters of the content and / or format of the video that allows for better grain synthesis 212 at the post-delivery processor 108. Such an approach may allow existing algorithms in the pre-delivery processor 104 and the post-delivery processor 108 to achieve an improvement in the optimization of the perceptual appearance of the synthesized grain for an existing end-to-end system 100. This approach may not produce as good a perceptual appearance of the synthesized grain in the end-to-end system 100 compared to the approach in which full access to the elements of the end-to- end system 100 (e.g., the pre-delivery processor 104, the content delivery network 106, and the postdelivery processor 108) that allows upgrades to maximize performance of the fidelity assessment with the objective model disclosed herein.
[0044] FIG. 3 illustrates an example data flow 300 for the determination of an FGS score 320 for a test video 304 in comparison to a source video 302. The source video 302 may include a video asset 102 as discussed herein, including its natural film grain. The source video 302 can be the video asset 102 with film grain received by the pre-delivery processor 104. The test video 304 may include the same video asset 102 with simulated film grain or the test video 304 can be the decoded and synthesized grain video asset 118 produced at the post-delivery processor 108. The source video 302 and test video 304 may be received by a frame splitter 306. The frame splitter 306 may generate oneor more source frames 308 and test frames 310 based on the source video 302 and test video 304. A region of interest (ROI) detection 312 may be performed on the source frames 308 to identify regions of the source video 302 and test video 304 to compare for performing the analysis of the grain synthesis 212. Source regions of interest 314 and test regions of interest 316 identified by the ROI detection 312 are applied to the operation using an objective model, which is referred to herein as an FGS model 318. The FGS model 318 has an operation within that uses a neural network for analysis in generating the FGS score 320.
[0045] In further detail, the source video 302 and the test video 304 may be applied to the frame splitter 306 to split one or more individual frames from each of the source video 302 and the test video 304. In an example, the frame splitter 306 may select a subset of the frames of the source video 302 for analysis. This may include selection of one or more initial frames, periodic selection of frames, random selection of frames, selection of frames that include specific objects or a mix of objects based on a determination of an object detection model, etc. Regardless of approach, the result of the frame splitter 306 is one or more source frames 308 selected from the source video 302, as well as one or more test frames 310 selected from the test video 304 that correspond to the source frames 308 (e.g., via frame count, synchronized frame, etc.). It should be noted that in other examples, the source frame 308 and test frame 310 may be input directly to the ROI detection 312, without using the frame splitter 306.
[0046] The ROI detection 312 may be performed on the source frame 308. It should be noted that film grain is a type of texture. Because of this, the FGS score 320 may be sensitive to texture generally. This sensitivity may lead to other textures in the source frame 308 and test frame 310 potentially causing an incorrect FGS score 320 prediction. This inaccuracy may manifest as a preference by the FGS metric (discussed later with respect to FIG. 4) and the FGS score 320 for test frames 310 having overly smooth simulated film grain. To minimize this effect, the operation with the FGS model 318 may be focused on regions where film grain is more apparent. The ROI detection 312 can be configured to determine which regions to focus on.
[0047] As shown, the ROI detection 312 may be placed before the operation with the FGS model 318.In many examples, the ROI detection 312 may be used to select flat region patches. This may be usefulbecause film grain is more apparent in flat regions to improve the perceptual appearance of the synthesized grain that can be based on visual verification.
[0048] The ROI detection 312 may split the source frames 308 into patches. Note, patches can sometimes be referred to as blocks. In some examples, these patches may be of a predefined size. For instance, the patch size may be NxN (i.e. N image pixels x N image pixels), such as 128x128 of image pixels in one example. It should be noted that this size is configurable. In some examples, the patch size is configured to allow for a whole number of non- overlapping patches to be defined across a first portion of the image area and where a second portion of the image lacks a patch, (e.g., in a grid of non- overlapping patches). In still other examples, the patches may be overlapping and may be defined using a sliding window approach. In yet further examples, the patches may be of varying sizes.
[0049] Regardless of the approach used to construct the patches that are blocks, after splitting the frame into blocks the pixel level variance of each block may be calculated. In an example, the ROI detection 312 may be computed using a pixel-level variance-based algorithm. This variance of the pixels may be computed within each patch and may be computed mathematically using one or more channels of the patch. For instance, the pixel level variance may be computed mathematically over a Y’(luma) channel from the (Y’, Cb’, Cr’) non-linear color space of the patch or the Y (luminance) channel from the (Y, Cb, Cr) linear color space of the patch depending on how the color space of the patch is expressed. The pixel level variance based algorithm may be performed based on the Y’ (luma) channel variance or the Y (luminance) channel variance which can be more representative of the image contrast structure within the patch. Other channels, such as chroma / chrominance channels may be used, but it should be noted that chroma / chrominance channels may have a lower density of image contrast structural information in the blocks as compared to the luma channel that is the Y’ channel or the luminance channel that is the Y channel. The term luminance will be used onwards, however, the term luma can be used in place of luminance in terms of the operations described herein as this only depends on the color space the patch is expressed in.
[0050] Once the variance is computed, the blocks may be ranked by pixel level variance such as the lowest level to the highest level. Using the ranking, a configurable percentile of the blocks with the lowest variance may be selected for use in determining FGS scoring. For example, 50% or another predefined percentage of the lowest ranked pixel level variance may be selected for use in determiningFGS scoring. The source regions of interest 314 are extracted from the source frames 308 and the test regions of interest 316 are extracted from the test frames 310. This percentile may be set based on the observation that a low variance patch is a patch likely to include a flat region suitable for measuring film grain as opposed to including other image textures which may affect the resultant scoring. The source regions of interest 314 and the test regions of interest 316 are accordingly applied to the FGS model 318 for analysis. The analysis may result in the FGS score 320. It should be noted that while flat regions with low activity in terms of luminance variance is given as an example, it should be noted that the regions of interest may be regions satisfying various other criteria, such as being regions with low level of changes in contrast details such as images with low contrast change gradients in pixel level variances (e.g., sky, clouds, etc.).
[0051] FIG. 4 illustrates further details of data flow 400 aspects of the operation with the FGS model 318 in determining the FGS score 320 from the source regions of interest 314 and the test regions of interest 316. As shown in FIG. 4, and with continuing reference to FIG. 3, the source regions of interest 314 and the corresponding test regions of interest 316 are applied to the FGS model 318. The operation of the FGS model 318 can involve feature extraction 404 on the source regions of interest 314 to generate source features 406 and also performs the corresponding feature extraction 404 on the test regions of interest 316 to generate test features 408. The operation of statistical measurement 410 and then an aggregation 412 is performed on the source features 406 and the corresponding test features 408. The result of the aggregation 412 yields a model score 414. The model score 414 may be perceptually non-linear, so the model score 414 is provided to a score mapping 416 to convert the model score 414 into the perceptually linear FGS score 320. In this context, a perceptually linear score refers to a score in which a delta in a one range of the score would lead to the same perceptual difference as the same delta in a different range of the score.
[0052] For sake of explanation, two possible FGS models 318 for determining the model score 414 are discussed. The first example FGS model 318 may use Gram matrices and mean squared error for performing an image quality score (IQ A) for determining the model score 414. The second example FGS model 318 may use global means and global pixel level variance.
[0053] Sometimes referred to as an FGS metric, the model score 414 may be a value indicative of how similar two frames would be perceived. In an example, the FGS model 318 may be used to determinehow similar native film grain would be perceived in the example source frame 308 (e.g., in the source regions of interest 314) as compared to the test frames 310 having synthesized film grain (e.g., in the test regions of interest 316).
[0054] The FGS model 318 may be used for transforming the source regions of interest 314 and the test regions of interest 316 into a new representation. Another example of an FGS model 318 may be a Convolutional Neural Network (CNN) in many examples. The CNN may be a neural network trained on image data to be able to perform various image related tasks. For example, the CNN may have been trained on image data for performing tasks such as image classification, object classification into various object categories, object recognition, and / or texture synthesis.
[0055] In another example the FGS model 318 may be a pretrained Visual Geometry Group (VGG) model, in an example. The VGG model may be constructed as a deep CNN architecture with multiple layers and with small convolutional filters, having a minimal receptive field ( i.e., 3x3 image pixel area), that is small but still able to capture the image pixel level influence of the immediate surrounding image pixels. The VGG- 16 model includes 13 convolutional layers followed by 3 fully connected layers (13 + 3 total), hence the name 16. The final connected layer defines the output channels, one for each class of image that can be detected. As another example, the VGG-19 network may be used, which is similar to VGG- 16 but includes 3 additional convolutional layers bringing the total layer count to 19. It should be noted that using the VGG is only one example, and other models may be used for feature extraction 404. As some other examples, the MobileNet model, the Resnet model, or a custom model may be used instead of VGG.
[0056] Using the CNN, the extraction of feature maps is performed. For instance, five layers, six layers, etc., of the CNN may be used for analysis. Continuing with the example of using VGG, the extraction may include extraction of feature maps from the different stages of the VGG- 16 network or from the VGG-19 network. The extracted feature maps from the reference image x (e.g., Source frame 308) and from the test image y (e.g., test frame 310) are illustrated in the FGS model 318 as the vectors within the central dotted area (e.g., in FIG. 4). The feature maps that are extracted from the source regions of interest 314 are referred to as the source features 406, while the feature maps that are extracted from the test regions of interest 316 are referred to as the test features 408.
[0057] Within this representation, statistical measurements 410 may be performed to capture the appearance of a variety of film grain. In one example, the statistical measurements 410 operation may include flattening the source features 406 and the test features 408 and computing Gram matrices of the flattened feature map vectors. This flattening may be performed to facilitate the mathematical computation. A Gram matrix G may be defined as an inner dot product between the flattened feature map and its transpose. These Gram matrices may be computed for the reference image x, and for the test image y, and for each of the extracted feature maps i. This allows for the calculating of perceptual distance as a mean squared error (MSE) between the corresponding gram matrices for the selected feature maps from each stage. The MSE from each stage may be applied to an aggregation 412 by taking the weighted average using weights wLcorresponding to each of the feature maps i.
[0058] In an example using a five-layer feature extraction 404, this aggregation 412 may be shown mathematically as shown in Equation (1):
[0059] The weights wLmay be defined empirically in an example. In another example, the weights wLmay be learned to fit with MOS ratings that may be gathered from users in a subjective quality evaluation test. In such a learned approach, the automated aggregation 412 may be taught to approximate the results that were manually identified using the MOS ratings.
[0060] Referring to the second example FGS model 318 for determining the model score 414, instead of Gram matrices, the statistical measurements 410 may include calculating global means and global variance of the extracted source feature 406 from the source regions of interest 314 and the test features 408 from the test regions of interest 316.
[0061] Turning to global means and global variance of the extracted feature maps, mathematically as shown in Equations (2):where:rePresent the global means and variances of x, h)and yland the global covariance betweenand y^, respectively, and and c2are small positive constants to avoid numerical instability when the denominators are close to zero.
[0062] Then in the statistical measurements 410 operation, a distance between the statistical measurements 410 between the source frame 308 and the test frame 310 may be computed. This may include to calculate similarity and / or distance between the statistical measurements 410 from the corresponding source feature 406 of the source regions of interest 314 and the test features 408 of the test regions of interest 316. These statistical measurements 410 are accordingly aggregated within the FGS model 318 to produce the model score 414.
[0063] The aggregation 412 may include modeling texture similarity as well as creating a structure score using a weighted average of the source features 406 and test features 408 extracted from the feature maps. The aggregation 412 of the similarity and distance may be performed using a weighted average of the calculated distance across the source features 406 and test features 408 extracted from the feature maps. Mathematically, the distance D may be found as shown in Equation (3):where weights a and [3 are learned such that their sum is equal to 1.
[0064] Additional variations of the FGS model 318 may be implemented in various combinations. For example, variations may be used for the feature extraction 404 transform (e.g., VGG-16, VGG-19, MobileNet, Resnet, or another image classification model). Additionally, or alternatively, variations may be used for the statistical measurements 410 (e.g., Gram matrix (2nd order stats), global means (1st order stats), covariance, etc.) Additionally, or alternately, variations may be used for the perceptual distance measure (e.g., structural similarity (SSIM) luminance component, MSE, cosine similarity, etc.) Additionally, or alternately, the aggregation 412 may involve weighting a straight average, learnable weights, tunable weights, etc.
[0065] Decisions on which variations of the FGS model 318 to use may depend on factors such as available computational resources, required accuracy, etc. For instance, MobileNet is less computational complex to execute than VGG-19. Also, a global mean measure is less computational complex than a 2nd order measure gram matrix. In addition to MSE as the perceptual distance measure, a cosine-like similarity measure such as SSIM that can account for the homogeneity of textural patterns may be beneficial to address potential issues with unboundedness output scoring which may occur with MSE.
[0066] As one non-limiting set of potential variations, 16 potential variations may be constructed using different combinations of VGG-19 or MobileNet, Gram matrix or Global means, SSIM luminance measure or MSE, and average weights or learnable weights.
[0067] A good implementation that produces the FGS score 320 should accurately rank the order of candidate test videos 304 according to their film grain similarity against the source video 302 content. One target use case for the FGS score 320 is to alleviate human’s burden of selecting the most similar FGS test video 304 with respect to a source video 302. As there could be hundreds of possible FGS test video 304 candidates for each source video 302, an accurate FGS score 320 may be desirable to prefdter candidates so that a human-in-the-loop would only need to review the top-k candidates. In another example, the FGS score 320 could be used to rank the FGS test video 304 candidates without human intervention.
[0068] Regardless of application, a relevant evaluation criterion is the accuracy of the FGS score 320 as compared to MOS. The best candidate for content i may be denoted as Ct, the set of all contents asI, the top-k candidates subset as Sk, and total number of source contents as T. With these definitions, the accuracy can be defined as shown in Equation (4):Accuracy =
[0069] As defined, an FGS video selection process does not include cross content comparison in the current setting. Therefore, in some examples per-content Spearman-Rank Correlation (SRCC) may be used as a second criterion for the FGS selection criteria.
[0070] Returning to FIG. 4, the score mapping 416 may be performed on the model score 414 to generate the FGS score 320. Taking the model score 414 outputted from the aggregation 412, and a non-linearly mapping (for example a logistic regression), the model score 414 may be transformed into an easily interpretable range, e.g., of 0 to 100.
[0071] For example, model score 414 outputted from the aggregation 412 may be a value bounded from zero to one. However, simply multiplying this value by 100 may not offer a useful result because the difference in perception from 0.7 to 0.8 may be perceptually different from the difference between 0.8 to 0.9, as one example.
[0072] This mapping of the model score 414 into the FGS score 320 may be performed because the output value of the aggregation 412 may be non-linear with respect to perceptual difference. This means that a delta in one range of the model score 414 may have a different perceptual difference as compared to same delta in a different range of the model score 414. Yet, it may be desirable for the FGS score 320 to be perceptually linear, e.g., that a delta in one range of the FGS score 320 indicates the same perceptual difference as the same delta in a different range of the FGS score 320. Taking the output of the aggregation 412, and a non-linearity mapping (e.g., a logistic regression) performed by the score mapping 416, the model score 414 may be mapped into the FGS score 320.
[0073] It should be noted that the mapping of the model score 414 to the FGS score 320 may be based on the viewer device 110. For instance, different viewer devices 110 may have different viewer device parameters, such as screen size, screen resolution, ambient light level, dynamic range, viewer distance, etc. If these parameters are available, then the score mapping 416 may be performed specific to thosedevice parameters. For example, a smaller screen may utilize a mapping of the model score 414 to the FGS score 320 that outputs a higher FGS score 320 for the same model score 414, due to the perceptual differences between the screen sizes.
[0074] A subjective study such as the MOS discussed above, may be used to train the score mapping 416 how to map from the linear mapping output of the FGS model 318 into a more perceptually meaningful result. For example, different MOS studies may be performed for different viewer device characteristics, which may be used to inform the mapping of the model score 414 to the FGS score 320 for the specific viewer device parameters of the viewer device 110 (e.g., screen size, ambient light level, dynamic range, etc.).
[0075] In some examples, this mapping may account for human visual system (HVS) modeling by assessing the visual media input in terms of human visual contrast sensitivity, luminance and texture masking effects, and / or visual saliency and attention effects, and produce the overall score mapping 416 of the model score 414 to generate the FGS score 320 in accordance with these factors.
[0076] It should be noted that the FGS score 320 is codec-independent. It should also be noted that the FGS score 320 is also independent of other aspects of the source video 302, such as resolution, bit depth, dynamic range, frame rate, color space, etc.
[0077] As noted above, many different possible variations may be used to create the FGS model 318. As some additional variations, the FGS model 318 can be trained with or without using ROI detection 312 and may be used to perform inference with or without using ROI detection 312. In one nonlimiting example, a VGG-19 model may be used with Gram matrices, SSIM luminance measure, and learnable weights. The FGS model 318 may be (i) trained and tested using ROI detection 312, (ii) trained without ROI detection 312 but tested with ROI detection 312, (iii) trained with ROI detection 312 but tested without ROI detection 312, or (iv) trained and tested without ROI detection 312.
[0078] FIG. 5 illustrates an example method 500 for using the FGS model 318 for performing the IQ A to determine the FGS score 320. The FGS model 318 may be any of the different variations as discussed in detail herein. An assessment of perceptual video quality of the film grain synthesis 212 may include receiving an example source video 302, with native film grain, and comparing that sourcevideo 302 to various patches of the same source video 302 with synthesized film grain. For example, the native film grain may be removed and then replaced with synthesized film grain.
[0079] At operation 502, the source video 302 and the test video 304 are received for processing. In an example, the source video 302 including native film grain and the test video 304 including simulated film grain may be received for analysis. While this flow is shown with a single test video 304, in some examples, many test videos 304 may be compared to a single source video 302, e.g., to determine which test video 304 has the best perceptual quality of synthesized film grain.
[0080] At operation 504, source frames 308 and test frames 310 are identified from the source video 302 and test video 304. In some examples, the source frames 308 and the test frames 310 may be received directly for processing (e.g., as single frames received at operation 502), while in other examples, the frame splitter 306 may be applied to the source video 302 and the test video 304 to select the source frames 308 and test frames 310 for analysis.
[0081] At operation 506, regions of interest are detected within the source frames 308 of the source video 302. In an example, the ROI detection 312 component may be used to split a source frame 308 of the source video 302 into blocks. These blocks may be non-overlapping, windowed and overlapping, of equal size, of varied size, etc. Regardless of the specifics of the block splitting, a variance of each block may be determined, and the blocks may be ranked according to the determined variance. A subset of the blocks with the lowest variance may be identified as being the regions of interest selected for extraction.
[0082] At operation 508, the regions of interest are extracted from both the source frames 308 and from the test frames 310. These regions of interest may define the subset of the source frame 308 and the test frame 310 to be compared.
[0083] At operation 510, the source regions of interest 314 and the test regions of interest 316 are applied to a deep neural network for analysis. The analysis may include feature extraction 404, statistical measurement 410, and similarity measurement. These aspects may be performed using the FGS model 318 in any of the variations discussed herein. The output of the aggregation 412 may be a film grain similarity output model score 414 indicative of the similarity of the source regions of interest 314 and the test regions of interest 316. 1
[0084] At operation 512, a map output to FGS score 320 procedure is performed. In an example, the model score 414 output of the aggregation 412 is mapped using the score mapping 416 to a perceptually linear value representative of relative film grain perception to a user. For example, taking the output of the aggregation 412, and a non-linearity mapping (e.g., a logistic regression) performed by the score mapping 416, the perceptually non-linear model score 414 may be mapped into the perceptually linear FGS score 320 and outputted by the FGS model 318. In some examples, viewer device parameters of the viewer device 110 may be available to the score mapping 416, such as screen size, ambient light level, dynamic range, etc. If so, the score mapping 416 may utilize the parameters to perform the mapping based on the device parameters. The FGS score 320 may be an easily interpretable score, such as a value from 0 to 100, where higher values indicate greater perceptual similarity of the test video 304 to the source video 302, and lower values indicate less perceptual similarity of the of the test video 304 to the source video 302.
[0085] Thus, the film grain similarity output model score 414 from the aggregation 412 is mapped to an objective FGS score 320 representative of similarity in perception of the simulated film grain of the test video 304 as compared to perception of the native film grain in the source video 302. After operation 512, the method 500 ends.
[0086] FIG. 6 illustrates an example 600 of use of the FGS score 320 on a plurality of test regions of interest 316 corresponding to a plurality of source regions of interest 314. Each of the plurality of test regions of interest 316 uses different grain synthesis 212 parameters and therefore achieves a different film grain appearance. The FGS score 320 is shown for each of the plurality of test regions of interest 316, as computed via the operations 502-512 of the method 500. This performance evaluation may be completed to determine which of these variations produces the best IQA scoring of film grain replacement. As can be seen, the FGS scores 320 provide a wide range of values indicative of user perception of the various test regions of interests 316. Using these results, the upper left hand test regions of interest 316 may be automatically selected for use by the network monitor 112 of the end- to-end system 100, e.g., by using the grain synthesis 212 parameters having the highest FGS score 320 for transmission of the source video 302 over the content delivery network 106.
[0087] In another example, the operations 502-512 of the method 500 may be performed for a plurality of different video assets 102. In such an example, the objective FGS scores 320 of the plurality ofvideo assets 102 may be ranked. This may allow for the perceptual quality of film grain synthesis 212 algorithms and / or grain model parameters to be compared across video assets 102.
[0088] In another example, the operations 502-512 of the method 500 may be performed for a plurality of different grain synthesis 212 and / or grain model parameters. Then, the objective FGS scores 320 of the plurality of grain synthesis 212 and / or grain model parameters may be ranked. This may allow for the different perceptual quality of film grain synthesis 212 algorithms and / or grain model parameters to be compared. For instance, based on this analysis, the grain synthesis 212 and / or grain model parameters with the highest ranked objective film grain score may be utilized for performing grain reduction 206 of the video asset 102 before streaming and / or for performing grain synthesis 212 to regrain the video asset 102 after the streaming.
[0089] FIG. 7 illustrates an example 700 of a computing device 702 for use in the end-to-end system 100 for grain-aware video coding transmission. Referring to FIG. 7, and with reference to FIGS. 1-6, the devices and modules discussed herein may be examples of such computing devices 702. For instance, the operations performed by the pre-delivery processor 104, content delivery network 106, post-delivery processor 108, viewer device 110, network monitor 112, frame splitter 306, ROI detection 312, FGS model 318, aggregation 412, score mapping 416, etc., as well as the operations discussed in the data flows 200 and 300 and in the method 500 may be performed by such computing devices 702. As shown, the computing device 702 includes a processor 704 that is operatively connected to a storage 706, a network device 708, an output device 710, and an input device 712. It should be noted that this is merely an example, and computing devices 702 with more, fewer, or different components may be used.
[0090] The processor 704 may include one or more integrated circuits that implement the functionality of a central processing unit (CPU) and / or graphics processing unit (GPU). In some examples, the processors 704 are a system on a chip (SoC) that integrates the functionality of the CPU and GPU. The SoC may optionally include other components such as, for example, the storage 706 and the network device 708 into a single integrated device. In other examples, the CPU and GPU are connected to each other via a peripheral connection device such as peripheral component interconnect (PCI) express or another suitable peripheral data connection. In one example, the CPU is a commerciallyavailable central processing device that implements an instruction set such as one of the x86, ARM, Power, or microprocessor without interlocked pipeline stage (MIPS) instruction set families.
[0091] Regardless of the specifics, during operation the processor 704 executes stored program instructions that are retrieved from the storage 706. The stored program instructions, accordingly, include software that controls the operation of the processors 704 to perform the operations described herein. The storage 706 may include both non-volatile memory and volatile memory devices. The nonvolatile memory includes solid-state memories, such as not and (NAND) flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the system is deactivated or loses electrical power. The volatile memory includes static and dynamic random-access memory (RAM) that stores program instructions and data during operation of the end-to-end system 100.
[0092] The GPU may include hardware and software for display of at least two-dimensional (2D) and optionally 3D graphics to the output device 710. The output device 710 may include a graphical or visual display device, such as an electronic display screen, projector, printer, or any other suitable device that reproduces a graphical display. As another example, the output device 710 may include an audio device, such as a loudspeaker or headphone. As yet a further example, the output device 710 may include a tactile device, such as a mechanically raisable device that may, in an example, be configured to display braille or another physical output that may be touched to provide information to a user.
[0093] The input device 712 may include any of various devices that enable the computing device 702 to receive control input from users. Examples of suitable input devices that receive human interface inputs may include keyboards, mice, trackballs, touchscreens, voice input devices, graphics tablets, and the like.
[0094] The network devices 708 may each include any of various devices that enable the devices to send and / or receive data from external devices over networks. Examples of suitable network devices 708 include an Ethernet interface, a Wi-Fi transceiver, a cellular transceiver, or a BLUETOOTH or Bluetooth Low Energy (BLE) transceiver, ultra-wideband (UWB) transceiver, or other networkadapter or peripheral interconnection device that receives data from another computer or external data storage device, which can be useful for receiving large sets of data in an efficient manner.
[0095] The processes, methods, or algorithms disclosed herein can be deliverable to be implemented by a processing device, controller, or computer, which can include any existing programmable electronic control unit or dedicated electronic control unit. Similarly, the processes, methods, or algorithms can be stored as data and instructions executable by a controller or computer in many forms including, but not limited to, information permanently stored on non-writable storage media such as read-only memory (ROM) devices and information alterably stored on writable storage media such as floppy disks, magnetic tapes, compact discs (CDs), RAM devices, and other magnetic and optical media. The processes, methods, or algorithms can also be implemented in a software executable object. Alternatively, the processes, methods, or algorithms can be embodied in whole or in part using suitable hardware components, such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), state machines, controllers or other hardware components or devices, or a combination of hardware, software and firmware components.
[0096] While exemplary embodiments are described above, it is not intended that these embodiments describe all possible forms encompassed by the claims. The words used in the specification are words of description rather than limitation, and it is understood that various changes can be made without departing from the spirit and scope of the disclosure. As previously described, the features of various embodiments can be combined to form further embodiments of the invention that may not be explicitly described or illustrated. While various embodiments could have been described as providing advantages or being preferred over other embodiments or prior art implementations with respect to one or more desired characteristics, those of ordinary skill in the art recognize that one or more features or characteristics can be compromised to achieve desired overall system attributes, which depend on the specific application and implementation. These attributes can include, but are not limited to strength, durability, life cycle, marketability, appearance, packaging, size, serviceability, weight, manufacturability, ease of assembly, etc. As such, to the extent any embodiments are described as less desirable than other embodiments or prior art implementations with respect to one or more characteristics, these embodiments are not outside the scope of the disclosure and can be desirable for particular applications.
[0097] While exemplary embodiments are described above, it is not intended that these embodiments describe all possible forms of the invention. Rather, the words used in the specification are words of description rather than limitation, and it is understood that various changes may be made without departing from the spirit and scope of the invention. Additionally, the features of various implementing embodiments may be combined to form further embodiments of the invention.
Claims
WHAT IS CLAIMED IS:
1. A method for perceptual quality assessment of objective film grain synthesis, comprising: receiving a source frame with source regions of interest and a corresponding test frame with corresponding test regions of interest, the source frame including native film grain, the test frame including simulated film grain; applying the source frame with the source regions of interest and the corresponding test frame with the corresponding test regions of interest to a feature extraction using a neural network trained on image data for classification to produce source features and corresponding test features; and applying statistical measures to the source features and the corresponding test features to determine a distance between statistical measurements of the source features of the source regions of interest and the corresponding test features of the corresponding test regions of interest, wherein the distance between the statistical measurements is used to produce a model score, and a film grain synthesis score is based on the model score.
2. The method of claim 1, wherein the source frame is a frame of a source video including the native film grain, and the test frame is a corresponding frame of a test video including the simulated film grain.
3. The method of claim 1 , wherein applying the source regions of interest and test regions of interest to the neural network includes: extracting feature maps from a plurality of stages of a Convolutional Neural Network (CNN); using the extracted feature maps to compute a set of statistical measurements indicative of appearance of film grain present in the source regions of interest and the test regions of interest; calculating a perceptual distance of the statistical measurements between the source regions of interest and the test regions of interest for each extracted feature map; and aggregating the perceptual distances to determine the model score.
4. The method of claim 3, wherein the CNN is trained on training set image data for performing image classification, object classification, object recognition, and / or texture synthesis.
5. The method of claim 3, further comprising flattening the extracted feature maps from a multidimensional array representation into flattened feature maps of a single dimensional vector to be used as an input to calculating the set of statistical measurements.
6. The method of claim 3, wherein the set of statistical measurements includes calculation of a Gram matrix of the feature maps.
7. The method of claim 6, wherein the perceptual distance is calculated as mean squared error (MSE) between corresponding gram matrices for the extracted feature maps.
8. The method of claim 7, wherein the aggregating the perceptual distances includes computing a weighted average of the MSE of each of extracted feature maps.
9. The method of claim 3, wherein the set of statistical measurements includes luminance measures and structural similarity measures.
10. The method of claim 9, wherein the luminance measures are weighted by a and the structural similarity measures are weighted by 0, such that a + 0 = 1 .
11. The method of claim 1, further comprising mapping the model score to a perceptually linear film grain score representative of similarity in perception of the simulated film grain of the test frame as compared to perception of the native film grain in the source frame.
12. The method of claim 11, wherein the mapping of the model score to the perceptually linear film grain score includes assessing the film grain score in terms of human visual system (HVS) factors, the HVS factors including one or more of contrast sensitivity, luminance and texture masking effects, and / or visual saliency and attention effects.
13. The method of claim 11 , further comprising: receiving viewer device parameters of a viewer device displaying the test frame; and mapping the model score to the perceptually linear film grain score, accounting for the viewer device parameters.
14. The method of claim 13, wherein the viewer device parameters include one or more of screen size of the viewer device, screen resolution of the viewer device, ambient light level surrounding the viewer device, dynamic range of the viewer device, and viewer distance from the viewer device.
15. The method of claim 14, further comprising training the mapping of the model score to the perceptually linear film grain score based on mean opinion score (MOS) studies performed for a plurality of different viewer device parameters.
16. The method of claim 1, further comprising: detecting the source regions of interest within the source frame; and extracting the source regions of interest from the source frame and extracting the corresponding test regions of interest from the corresponding test frame.
17. The method of claim 16, wherein the detecting of the regions of interest includes: splitting the source frame into a plurality of blocks; computing pixel-level variance of each block of the plurality of blocks; ranking the plurality of blocks according to the computed pixel-level variance; and selecting the regions of interest as being those of the plurality of blocks having a lowest ranked variance.
18. The method of claim 17, wherein the plurality of blocks are non-overlapping.
19. The method of claim 17, wherein the plurality of blocks are equally-sized blocks.
20. The method of claim 17, wherein the plurality of blocks are formed using a sliding window.
21. The method of claim 17, wherein the pixel-level variance is computed for each block using the variance of a Y luminance component of the respective block.
22. The method of claim 17, wherein a subset of the plurality of blocks having the lowest ranked variance are selected as being the regions of interest.
23. The method of claim 1, further comprising: determining grain model parameters indicative of aspects of the native film grain that is present in the source frame; performing grain reduction of the source frame according to grain reduction parameters indicative of aspects of removal of the native film grain from the source frame; and performing grain synthesis to re-grain the source frame using grain synthesis parameters indicative of properties of the simulated film grain to be added to create the corresponding test frame.
24. The method of claim 23, further comprising: performing the operations of claim 23 for a plurality of different video assets; and ranking the model scores of the plurality of different video assets.
25. The method of claim 23, further comprising: performing the operations of claim 23 for a plurality of different grain synthesis and / or grain model parameters; and ranking the model scores of the plurality of different grain synthesis and / or grain model parameters.
26. The method of claim 25, further comprising utilizing the grain synthesis and / or grain model parameters with the highest ranked model score for performing grain reduction of a videoasset before streaming and for performing the grain synthesis to re-grain the video asset after the streaming.
27. The method of claim 26, further comprising embedding the grain synthesis and / or grain model parameters into a bit stream of the video asset for informing the grain synthesis after the streaming.
28. A system for performing perceptual quality assessment of objective film grain synthesis, comprising: one or more hardware computing devices that provide a decoded and synthesized grain video asset to a viewing device and that are configured to determine a model score for a perceptual quality assessment of objective film gram synthesis, configured to: receive a source frame with source regions of interest and a corresponding test frame with corresponding test regions of interest, the source frame including native film grain, the test frame including simulated film grain; apply the source frame with the source regions of interest and the corresponding test frame with the corresponding test regions of interest to a feature extraction using a neural network trained on image data for classification to produce source features and corresponding test features; and apply statistical measures to the source features and the corresponding test features to determine a distance between statistical measurements of the source features of the source regions of interest and the corresponding test features of the corresponding test regions of interest, wherein the distance between the statistical measurements is used to produce a model score, and a film grain synthesis score is based on the model score.
29. The system of claim 28, wherein the system is integrated as an upgrade into a pre-delivery processor of an end-to-end system for delivering video.
30. The system of claim 28, wherein the system is integrated as an upgrade into a post-delivery processor of an end-to-end system for delivering video.
31. The system of claim 28, wherein the source frame is a frame of a source video including the native film grain, and the test frame is a corresponding frame of a test video including the simulated film grain.
32. The system of claim 28, wherein the one or more hardware computing devices are further configured to apply the source regions of interest and test regions of interest to the neural network by performing operations including to: extract feature maps from a plurality of stages of a Convolutional Neural Network (CNN); use the extracted feature maps to compute a set of statistical measurements indicative of appearance of film grain present in the source regions of interest and the test regions of interest; calculate a perceptual distance of the statistical measurements between the source regions of interest and the test regions of interest for each extracted feature map; and aggregate the perceptual distances to determine the model score.
33. The system of claim 32, wherein the CNN is trained on training set image data for performing image classification, object classification, object recognition, and / or texture synthesis.
34. The system of claim 32, wherein the one or more hardware computing devices are further configured to flatten the extracted feature maps from a multidimensional array representation into flattened feature maps of a single dimensional vector to be used as an input to calculating the set of statistical measurements.
35. The system of claim 32, wherein the set of statistical measurements includes calculation of a Gram matrix of the feature maps.
36. The system of claim 35, wherein the perceptual distance is calculated as mean squared error (MSE) between corresponding gram matrices for the extracted feature maps.
37. The system of claim 36, wherein to aggregate the perceptual distances includes to compute a weighted average of the MSE of each of extracted feature maps.
38. The system of claim 32, wherein the set of statistical measurements includes luminance measures and structural similarity measures.
39. The system of claim 38, wherein the luminance measures are weighted by a and the structural similarity measures are weighted by 0, such that a + 0 = 1 .
40. The system of claim 28, wherein the one or more hardware computing devices are further configured to map the model score to a perceptually linear film grain score representative of similarity in perception of the simulated film grain of the test frame as compared to perception of the native film grain in the source frame.
41. The system of claim 40, wherein to map the model score to the perceptually linear film grain score includes to assess the film grain score in terms of human visual system (HVS) factors, the HVS factors including one or more of contrast sensitivity, luminance and texture masking effects, and / or visual saliency and attention effects.
42. The system of claim 40, wherein the one or more hardware computing devices are further configured to: receive viewer device parameters of a viewer device displaying the test frame; and map the model score to the perceptually linear film grain score, accounting for the viewer device parameters.
43. The system of claim 42, wherein the viewer device parameters include one or more of screen size of the viewer device, screen resolution of the viewer device, ambient light level surrounding the viewer device, dynamic range of the viewer device, and viewer distance from the viewer device.
44. The system of claim 43, wherein the one or more hardware computing devices are further configured to train the mapping of the model score to the perceptually linear film grainscore based on mean opinion score (MOS) studies performed for a plurality of different viewer device parameters.
45. The system of claim 28, wherein the one or more hardware computing devices are further configured to: detect the source regions of interest within the source frame; and extract the source regions of interest from the source frame and extracting the corresponding test regions of interest from the corresponding test frame.
46. The system of claim 45, wherein to detect the regions of interest includes to: split the source frame into a plurality of blocks; compute pixel-level variance of each block of the plurality of blocks; rank the plurality of blocks according to the computed pixel-level variance; and select the regions of interest as being those of the plurality of blocks having a lowest ranked variance.
47. The system of claim 46, wherein the plurality of blocks are non- overlapping.
48. The system of claim 46, wherein the plurality of blocks are equally-sized blocks.
49. The system of claim 46, wherein the plurality of blocks are formed using a sliding window.
50. The system of claim 46, wherein the pixel-level variance is computed for each block using the variance of a Y luminance component of the respective block.
51. The system of claim 46, wherein a subset of the plurality of blocks having the lowest ranked variance are selected as being the regions of interest.
52. The system of claim 28, wherein the one or more hardware computing devices are further configured to: determine grain model parameters indicative of aspects of the native film grain that is present in the source frame; perform grain reduction of the source frame according to grain reduction parameters indicative of aspects of removal of the native film grain from the source frame; and perform grain synthesis to re-grain the source frame using grain synthesis parameters indicative of properties of the simulated film grain to be added to create the corresponding test frame.
53. The system of claim 52, wherein the one or more hardware computing devices are further configured to: perform the operations of claim 52 for a plurality of different video assets; and rank the model scores of the plurality of different video assets.
54. The system of claim 52, wherein the one or more hardware computing devices are further configured to: perform the operations of claim 52 for a plurality of different grain synthesis and / or grain model parameters; and rank the model scores of the plurality of different grain synthesis and / or grain model parameters.
55. The system of claim 54, wherein the one or more hardware computing devices are further configured to utilize the grain synthesis and / or grain model parameters with the highest ranked model score for performing grain reduction of the video asset before streaming and for performing the grain synthesis to re-grain the video asset after the streaming.
56. The system of claim 55, wherein the one or more hardware computing devices are further configured to embed the grain synthesis and / or grain model parameters into a bit stream of the video asset for informing the grain synthesis after the streaming.
57. A non- transitory computer-readable medium comprising instructions for performing perceptual quality assessment of objective film grain synthesis that, when executed by one or more hardware computing devices, cause the one or more hardware computing devices to perform operations including to: receive a source frame with source regions of interest and a corresponding test frame with corresponding test regions of interest, the source frame including native film grain, the test frame including simulated film grain; apply the source frame with the source regions of interest and the corresponding test frame with the corresponding test regions of interest to a feature extraction using a neural network trained on image data for classification to produce source features and corresponding test features; and apply statistical measures to the source features and the corresponding test features to determine a distance between statistical measurements of the source features of the source regions of interest and the corresponding test features of the corresponding test regions of interest, wherein the distance between the statistical measurements is used to produce a model score, and a film grain synthesis score is based on the model score.