Film grain measurement using adaptive area selection
The sub-band based film grain assessment process addresses the challenges of managing film grain in video streaming by converting images to the frequency domain for accurate similarity evaluation, enhancing efficiency and user experience.
Patent Information
- Application Number
- JP2024173961
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-18
- Filing Date
- 2024-10-03
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2044-10-03
AI Technical Summary
Existing methods for managing film grain in video streaming face challenges such as high bandwidth requirements and difficulty in accurately synthesizing film grain, leading to inefficient compression and potential visual artifacts.
A sub-band based film grain assessment process that evaluates film grain similarity by converting images from the spatial domain to the frequency domain, using techniques like steerable pyramids and fast Fourier transforms to analyze noise power spectra and generate an evaluation score.
This approach provides a more accurate assessment of film grain similarity, enabling efficient bandwidth usage and consistent film grain synthesis across videos, improving user perception and reducing visual artifacts.
Smart Images

Figure 2025109174000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications
[0001] This application claims the benefit of the filing date of U.S. Provisional Application No. 63 / 620,134, entitled "FREQUENCY DOMAIN FILM GRAIN OBJECTIVE METRICS WITH ADAPTIVE REGION SELECTION", filed on January 11, 2024, which is hereby incorporated by reference in its entirety for all purposes.
Background Art
[0002]
[0002] Film grain can be one of the prominent characteristics of video produced by conventional film cameras (e.g., shows or movies produced by the film industry). Film grain can be a perceptually pleasing noise that can be shown for artistic intent. However, including film grain in video streamed to client devices over a network can pose technical challenges, such as the need for a higher bitrate to encode the video containing film grain. This results in high bandwidth requirements that may not be compatible with streaming environments. To conserve bandwidth, film grain can be removed from the video before streaming it to the client device. However, viewers may not be satisfied with the visual quality of video without film grain.
[0003]
[0003] One technique may be to synthesize film grain on a client device to re-add film grain on a decoded video frame to mimic the film grain in the source video. This process may re-introduce film grain into the video. However, film grain synthesis can be a difficult task to restore an exact replica of the original film grain in the source video. Additionally, film grain synthesis may add unwanted effects such as new visual artifacts to the video frame.
Summary of the Invention
[0004]
[0004] The accompanying drawings are for illustrative purposes only and serve only to provide examples of possible structures and operations for the disclosed invention's system, apparatus, method, and computer program product. These drawings in no way limit any changes in form and detail that may be made by those skilled in the art without departing from the spirit and scope of the disclosed embodiments.
Brief Description of the Drawings
[0005]
Figure 1
[0005] A diagram showing a simplified system for evaluating film grain in an image, according to some embodiments.
Figure 2
[0006] A diagram showing a simplified system for synthesizing film grain in a video, according to some embodiments.
Figure 3
[0007] A diagram showing a more detailed example of a film grain synthesis system, according to some embodiments.
Figure 4
[0008] A diagram showing a system that performs a conversion from a spatial domain to a frequency domain and then generates an evaluation score, according to some embodiments.
Figure 5
[0009] A diagram showing an example of a steerable pyramid bandpass filter in four directions within a spatial domain, according to some embodiments.
Figure 6
[0010] A diagram showing an example of the output of a steerable pyramid according to some embodiments.
Figure 7
[0011] A simplified flowchart of a method for performing spatial / frequency conversion according to some embodiments.
Figure 8
[0012] A simplified flowchart of a method for performing metric calculation according to some embodiments.
Figure 9
[0013] A diagram showing an example of a server system according to some embodiments.
Figure 10
[0014] A diagram showing a more detailed example of a server system according to some embodiments.
Figure 11
[0015] A diagram showing an input image, an edge map of the input image, and the output of non-texture region detection according to some embodiments.
Figure 12
[0016] A simplified flowchart of a method for determining an evaluation score using adaptive region selection according to some embodiments.
Figure 13
[0017] A diagram showing an example of a computing device according to some embodiments.
DETAILED DESCRIPTION OF THE INVENTION
[0006]
[0018] In this specification, techniques for a video processing system are described. In the following description, for the purpose of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of some embodiments. Some embodiments defined by the claims may include some or all of the features in these examples alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.
[0007]
[0019] The system performs a subband based film grain assessment (SFGA) process. This process may evaluate the film grain in two images. For example, this evaluation may assess the similarity of the film grain in the two images. In this process, a first image (e.g., a reference image) and a second image (e.g., a test image) can be input into the film grain evaluation system. The film grain of the test image can be analyzed to determine its similarity to the film grain of the reference image. The film grain evaluation system converts the reference image and the test image from the spatial domain to the frequency domain. This results in a frequency domain representation for each of the reference image and the test image. In some embodiments, the film grain evaluation system may select subbands in the frequency domain for each of the reference image and the test image, or the entire frequency domain representation may be used. Next, the film grain evaluation system compares the frequency distribution (e.g., subband noise power spectrum) between the reference image and the test image. The film grain evaluation system may analyze the difference between the distributions to determine an evaluation score. For example, this score may measure the similarity of the film grain between the reference image and the test image. In some embodiments, a higher score may indicate that the film grain in the reference image and the test image is more similar compared to a lower score which may indicate that the film grain in the reference image and the test image is less similar.
[0008]
[0020] The sub-band based film grain evaluation process offers many advantages. For example, film grain may appear as noise in video. In some cases, comparing film grain within a spatial region for each pixel in two images may not provide an accurate indicator of film grain similarity. For example, film grain may have different values for each pixel in two different images, which would indicate that the film grains are not similar. However, the nature of film grain may cause it to appear similar in two images from a human perspective. For example, even if the values are different, the noise representing the film grain may appear similar to a human viewer. Converting from the spatial domain to the frequency domain can capture the presence of similar film grain in two images because the frequency domain depends on the frequencies of the film grain characteristics. The noise that appears as film grain in an image may have slightly different values when analyzed in the spatial domain, but in the frequency domain, the frequency similarity is captured. Frequency domain analysis can better extract the characteristics of the noise and describe the appearance of the noise compared to the spatial domain. When two similar film grain images are analyzed, their frequency domain characteristics can be very similar to each other, but the pixel level / spatial domain characteristics can be very different because the film grain may appear like random noise. When the similarity of two film grain images is calculated by a pixel-by-pixel method in the spatial domain, the similarity can be very low when the films appear very similar to human perception. Therefore, the evaluation score output by the film grain evaluation process can be more accurate than a comparison in the spatial domain when the goal is to detect similar film grain in an image as perceived by a human viewer.
[0009]
[0021] System
[0022] FIG. 1 shows a simplified system 100 for evaluating film grain in an image according to some embodiments. The server system 102 may include one or more computing devices capable of evaluating film grain in images such as a reference image and a test image. The server system 102 includes a film grain evaluation system 104 and a processing system 106.
[0010]
[0023] The film grain evaluation system 104 receives a reference image and a test image and outputs an evaluation score. The evaluation score may evaluate the film grain found in the reference image and the test image. In some embodiments, the evaluation score may measure the similarity of the film grain in the reference image and the test image. Although two images are described as being compared, three or more images may be compared.
[0011]
[0024] The reference image and the test image may be different types of images. For example, the reference image and the test image may be two actual captures of images, such as both the reference image and the test image being captured by a camera. Also, the reference image and the test image may be two synthesized film grain images. For example, an image may have synthesized film grain and no captured film grain. Also, one of the images may be an actual capture, and one of the references may be an image with synthesized film grain. Also, an image may be just film grain. For example, an image containing content and film grain may be separated into a clean image and an image of film grain. The image of film grain may be compared to a synthesized film grain, and neither image contains the content of a clean image. Other combinations may also be understood.
[0012]
[0025] The processing system 106 may receive an evaluation score and use the evaluation score to execute an action. For example, in the film grain synthesis system described in more detail in FIGS. 2 and 3, the evaluation score may be used to evaluate the similarity between the film grain on the encoder side and the synthesized film grain on the decoder side. The action may be to change the parameters for synthesizing the film grain based on the evaluation score. For example, the parameters may be automatically changed to make the synthesized film grain more similar to the original film grain. Also, some images may be flagged as containing dissimilar film grain in a quality control process. Other use cases may also include helping to control the film grain addition process, transferring film grain styling from one video to another, or performing quality control to ensure that the film grain is consistent across multiple videos. For example, performance quality control can ensure that the film grain is consistent over similar regions within the video, especially areas such as walls where there are moving people. The film grain on the wall should be consistent. This system aims to provide general film grain measurement and comparison. Given any pair of film grains, which can be two actual captures, or two synthesized film grains, or one actual capture and one synthesized one, the system can evaluate the film grains within the image to determine the similarity between the film grains within the image.
[0013]
[0026] Next, the film grain evaluation system will be described in more detail. The film grain synthesis system will be described first, but the film grain evaluation system is not limited to being used in that system.
[0014]
[0027] Film Grain Synthesis
[0028] Figure 2 shows a simplified system 200 for synthesizing film grain of a video according to some embodiments. A content provider may operate a video delivery system 206 to provide a content delivery service that enables entities to request and receive media content. The content provider may use the video delivery system 206 to adjust the delivery of media content to the client 204. The media content may be different types of content such as on-demand videos from a library of videos and live videos. In some embodiments, a live video may be when the video is available based on a linear schedule. The video may also be provided on demand. An on-demand video may be content that can be requested at any time and is not limited to viewing on a linear schedule. The video may be a program such as a movie, a show, an advertisement, etc. The server system 102 may receive a source video that may include different types of content such as video, audio, or other types of content information. The source video may be transcoded (e.g., encoded) to create an encoded version of the source video, which may be delivered to the client 204 as an encoded bitstream. Although the delivery of the video from the video delivery system 206 to the client 204 is shown, the video delivery system 206 may use a content delivery network (not shown) to deliver the video to the client 204.
[0015]
[0029] Encoder system 208 can encode source video into an encoded bitstream. Different types of encoders can be used, such as encoders that use different coding specifications. In some embodiments, the source video may have film grain, but the film grain may be removed from the source video and not included in the encoded video frames of the encoded bitstream. In other embodiments, the source video may not include film grain, but it may be desirable to add film grain to the decoded video frames.
[0016]
[0030] Client 204 can include different computing devices such as smartphones, living room devices, televisions, set-top boxes, tablet devices, etc. Client 204 includes an interface 212 and a media player 210 for playing content such as video. In client 204, decoder system 216 receives the encoded bitstream and decodes the encoded bitstream into decoded video frames.
[0017]
[0031] Film grain may exist in some videos such as shows and movies shot with conventional film cameras. Film grain can be a visible texture or pattern that appears in video shooting on film. Film grain may appear as noise in video. As described above, film grain may be removed from the source video before encoding, and the encoded bitstream may not contain film grain from the source video when sent from the server system 102 to the client device 204. Saving film grain from the source video in the encoded bitstream can be difficult for several reasons. For example, when film grain exists in the original source video, the bitrate of the encoded bitstream can increase. Also, the random nature of film grain in the source video can randomly change the bitrate as the bitrate of the frame increases when film grain is encountered, which can affect the delivery of the encoded bitstream to the client 204. The random nature can affect the playback experience as the bitrate changes during playback, which can cause rebuffering. Furthermore, the random nature of film grain in the video makes it difficult to predict when (e.g., which frame) and where (e.g., where within the frame) film grain occurs in the source video using prediction methods in video coding specifications. This can make compression inefficient. As described above, a digital camera may not generate film grain in the video, or the frames of the source video may not contain film grain, but the system may still add film grain to the video.
[0018]
[0032] Based on the above, the encoded bitstream may not contain film grain from the source video, but the film grain synthesis system 214 can synthesize film grain that can be added to the decoded video frame. That is, the film grain is removed on the encoder side and added on the decoder side.
[0019]
[0033] Figure 3 shows a more detailed example of a film grain synthesis system according to some embodiments. The source video is received in the grain removal system 302. The grain removal system 302 can remove film grain from the source video. The output is the de-grained video that can be encoded by the encoder system 208 into the encoded video. Here, a clean image without film grain is encoded.
[0020]
[0034] The difference between the source video and the de-grained video results in a residual video that can represent the film grain found in the source video. The film grain modeling system 306 receives the residual video and determines film grain parameters. The film grain parameters can represent the parameters determined to synthesize the film grain found in the source video.
[0021]
[0035] The encoded video in the film grain parameters can be sent to the client 204 in the video bitstream. The decoder system 216 receives the encoded video and decodes the encoded video into the decoded video. Also, the film grain synthesis system 310 receives the film grain parameters. The film grain synthesis system 310 uses the parameters to synthesize film grain. For example, the film grain may be modeled based on the film grain parameters. Then, the synthesized film grain is combined with the decoded video, and the result is a video with synthesized film grain as the output.
[0022]
[0036] The similarity of the synthesized film grain to the original film grain can vary. In some embodiments, an image from a video having the synthesized film grain can be compared to an image from a source video that includes the captured film grain. For example, a company may wish to evaluate the similarity of the synthesized film grain to the film grain that appears in the original captured video. Although this use case is described, other use cases can be understood. For example, images having only film grain, such as an image having only a residual video and an image having the synthesized film grain, can be compared.
[0023]
[0037] Next, the evaluation of the film grain is described below.
[0024]
[0038] Conversion from the spatial domain to the frequency domain
[0039] FIG. 4 shows a system that performs a conversion from the spatial domain to the frequency domain and then generates an evaluation score, according to some embodiments. A reference image R and a test image T can be compared. A film grain evaluation system 104 converts the reference image R and the test image T from the spatial domain to the frequency domain. As described above, the reference image R and the test image T can be different types of images, such as an image having the original film grain and an image having the synthesized film grain, an image having only the original film grain and the synthesized film grain, or other types of images.
[0025]
[0040] The spatial / frequency conversion method determination system 402 (hereinafter referred to as the determination system 402) receives a reference image R. The determination system 402 can analyze the characteristics of the reference image R and determine the settings of the spatial domain / frequency domain conversion system 404 (hereinafter referred to as the conversion system 404). In some embodiments, settings such as how many directions or which sub-bands should be used in the frequency domain, the range of frequency sub-bands to be used, the band-pass filter used in the fast Fourier transform, the fast Fourier transform to be used, etc., can be used to perform the conversion from the spatial domain to the frequency domain. The frequency band can be determined using different methods, such as using rules for determining the sub-bands to be used based on the characteristics of the image. Also, the prediction network may receive an image as the input and output sub-bands to be used.
[0026]
[0041] The conversion system 404 receives the reference image R and the test image T. The conversion system 404 converts the reference image R and the test image T from the spatial domain to the frequency domain. The reference image R and the test image T can be individually converted into their respective frequency domain representations. The conversion from the spatial domain to the frequency domain can convert the image from a representation of pixel values (e.g., intensity) to a representation of frequency components. The spatial domain can be a space where the pixel values represent the intensity of the image at spatial positions such as x, y coordinates. The frequency domain represents an image regarding its frequency components that can describe how the pixel values change across the image. For example, the frequency components can indicate how quickly the pixel values (e.g., intensity) change across space. In the frequency domain, low frequencies may represent a gradual change in pixel values, such as a non-texture area, and high frequencies may represent a rapid change in pixel values, such as an edge or fine detail (e.g., a texture area).
[0027]
[0042] As will be described in more detail below, the conversion system 404 may use a process to convert an image into sub-bands. In some embodiments, a steerable pyramid may be used to derive several sub-bands. A steerable pyramid is a linear multi-scale, multi-directional image decomposition that provides a useful front-end for image processing and computer vision applications. A steerable pyramid can perform multi-scale, multi-directional image decomposition. Other filters may be, for example, a Gaussian pyramid, a Laplacian pyramid, a wavelet filter, etc. Different frequency scales in a steerable pyramid can result in different sub-bands having multiple directions. The use of the steerable pyramid will be described in more detail below with reference to FIGS. 5, 6, and 7.
[0028]
[0043] The conversion system 404 may select some of the sub-bands based on the settings received from the determination system 402. For example, higher frequency sub-bands may be selected because these sub-bands may contain content that most affects how film grain appears. However, other frequency sub-bands may be selected if it is determined that those frequency sub-bands have a greater impact on film grain. The selected sub-bands are then converted from the spatial domain to the frequency domain. The output of the conversion system 404 may be sub-bands in the frequency domain, which may be represented as F R0 , F R1 ,..., F Rn for the reference image R and as F T0 , F T1 ,..., F Tn for the test image T. The subscripts R0, R1,..., Rn may represent different sub-bands for the reference image R. Similarly, the subscripts T0, T1,..., Tn represent the same sub-bands for the test image T. In some embodiments, the conversion system 404 performs a conversion, such as a fast Fourier transform, to convert the sub-bands from the spatial domain to the frequency domain.
[0029]
[0044] The metric calculation system 406 receives the frequency domain representation of the sub-band and performs metric calculations. The sub-band can be a two-dimensional (2D) representation in the frequency domain, such as a two-dimensional array of the values of the frequency components. The metric calculation system 406 converts the 2D representation into a one-dimensional (1D) representation for the sub-band. Different methods can be used to convert a two-dimensional array representing the sub-band in the frequency domain into a one-dimensional array representing the sub-band in the frequency domain. The one-dimensional representation may be a vector. For example, the values of the one-dimensional array represent the frequency components for each respective sub-band.
[0030]
[0045] The metric calculation system 406 then generates a distribution of frequency values, such as the noise power spectrum, for each sub-band. The noise power spectrum can describe how the noise varies with frequency, such as the distribution of the noise power for the values of the frequency of each sub-band.
[0031]
[0046] The metric calculation system 406 then compares the distributions for each selected sub-band of the reference image and the test image to generate a respective score for each respective sub-band. For example, the output can be scores 0, score 1,..., score N for sub-bands 0, 1,..., N respectively. In some embodiments, the score is based on the difference between the two distributions of the sub-band. In some embodiments, the Jensen-Shannon divergence of the noise power spectra for the sub-bands of the reference image and the test image may be used. The Jensen-Shannon divergence measures the similarity or divergence across the noise power spectra of the two sub-bands for the reference image and the test image. Although the Jensen-Shannon divergence is used, other methods for determining the similarity or divergence of the distributions for each respective sub-band of the reference image and the test image may be used.
[0032]
[0047] The system can better extract the characteristics of noise and describe the appearance of noise. Some widely used spatial domain quality metrics, such as PSNR, SSIM, and MS-SSIM, fail in film grain quality evaluation. When the system looks at two similar film grains, their frequency domain characteristics are very similar to each other, but since film grain may appear like some kind of random noise, the pixel level / spatial domain characteristics are very different. Therefore, when the system calculates the similarity of two film grain images by a per-pixel method, the similarity will be very low when the film grains may appear very similar to a human viewer.
[0033]
[0048] The score adaptive fusion system 408 can receive scores and combine the scores into an evaluation score. In some embodiments, the fusion method determination system 410 may receive a reference image R and determine the fusion method to use. For example, average, weighted average, maximum value of weighted average, minimum value of weighted average, machine learning or deep learning methods, or other methods may be used. The weights can be determined based on the characteristics of the reference image or machine learning methods. For example, based on analyzing the characteristics of the reference image R, weights for different sub-bands that may be more important can be determined. In some embodiments, more important sub-bands may include characteristics such as low texture regions where film grain may be more prominent and may be weighted higher. In some embodiments, the following metric calculations may be used.
[0034]
Number
[0035] Here, f i1 and f i2 are sub-bands of the reference image (e.g., Image 1) and the test image (e.g., Image 2). d(·) is a function that describes the distance / similarity between sub-bands f i1 and f i2 The weight w iis the weight assigned to the sub - band,
[0036]
Number
[0037] where N is the number of sub - bands.
[0038]
[0049] As described above, the steerable pyramid can be used in the conversion from the spatial domain to the frequency domain. Next, the use of the steerable pyramid will be described.
[0039]
[0050] Steerable Pyramid
[0051] As described above, the conversion system 404 can use a steerable pyramid to derive sub - bands. The steerable pyramid can be a multi - scale, multi - direction image decomposition that is translation - invariant and includes direction representation. The basic function is a directional (e.g., steerable) filter that is localized in both space and frequency. Other filters that can be used include Gaussian pyramids, Laplacian pyramids, wavelet filters, etc.
[0040]
[0052] The steerable pyramid may decompose an image into sub - bands, and each sub - band represents spatial frequency content at different directions and scales. FIG. 5 shows an example of a steerable pyramid band - pass filter in four directions within the spatial domain according to some embodiments. The four directions shown may be 0 degrees at 502, 45 degrees at 504, 90 degrees at 506, and 135 degrees at 508. Here, the filters are oriented in various angles. The various directions may capture feature points and edges arranged along the horizontal axis by the 0 - degree filter, and may capture feature points and edges arranged along the 45 - degree direction by the 45 - degree filter, etc., and can capture feature points and edges arranged in the direction. These directions are used, but other directions may be understood, and different numbers of directions may be used.
[0041]
[0053] FIG. 6 shows an example of the output of a steerable pyramid according to some embodiments. Here, there are four sub-bands and four directions. However, the number of sub-bands and directions may be adjusted. In some embodiments, four sub-bands and six directions may be used. At 602, an input of an image such as a reference image or a test image is shown. Different sub-bands are shown at 604-1, 604-2, 604-3, and 604-4. Each sub-band may include several directions. For example, the directions corresponding to four steerable pyramid filters are shown at 606-1, 606-2, 606-3, and 606-4.
[0042]
[0054] Each sub-band may correspond to different scales of frequencies within the image, such as from low frequencies to high frequencies. For example, sub-band 604-1 may include higher frequency components, and sub-band 604-4 may include lower frequency components. Higher frequency components are at the highest level of the steerable pyramid and may correspond to the finest detail levels within the image. The height may define the levels of the steerable pyramid such that the highest level is height 00, the next highest level is height 01, then height 02, and height 03, etc. The highest level may capture the highest spatial frequencies, which means representing the smallest and sharpest feature points within the image, such as fine edges or textures. Lower frequency sub-bands may be coarser scales compared to height 00. For example, height 01 may capture lower spatial frequencies and may represent slightly larger feature points and wider patterns within the image. As the height moves from height 00 to height 01, height 02, and height 03, etc., a decrease in resolution and a concentration on wider and more general feature points within the image are captured.
[0043]
[0055] Higher frequency sub-bands may capture finer details within an image, which may better capture film grain within the image. Wider feature points located within sub-bands having lower frequencies, such as height 02 and height 03, may not capture the details of film grain, similar to the sub-bands at height 00 and height 01. By separating an image within the frequency domain into different sub-bands, the evaluation may focus on sub-bands that can better characterize film grain. By removing sub-bands that do not accurately represent film grain, the evaluation may be improved. In some embodiments, the system uses four sub-bands and six directions and selects a high-frequency band (corresponding to height 00 in FIG. 6) as the input to the fast Fourier transform, although other numbers of bands and directions may be used. For example, all sub-bands may be used and the sub-bands may be weighted based on a determined importance. In some embodiments, the use of sub-bands may improve the analysis. A full frequency band-based metric may not capture the main feature points of film grain. Also, the human eye is sensitive only to some frequency sub-bands when viewing film grain. When taking the entire band as input to the metric, some frequency bands to which the human eye is not sensitive may cause interference in the metric result.
[0044]
[0056] Spatial / Frequency Conversion Method
[0057] FIG. 7 shows a simplified flowchart 700 of a method for performing spatial / frequency conversion according to some embodiments. At 702, the conversion system 404 receives a spatial / frequency conversion method determination. This determination may specify settings for performing spatial / frequency conversion, such as the number of directions to use in the steerable pyramid.
[0045]
[0058] At 704, the conversion system 404 inputs an image into a steerable pyramid. The steerable pyramid may use several filters in different directions. At 706, the conversion system 404 determines sub-bands from the output of the steerable pyramid. For example, the output of the steerable pyramid can be multiple sub-bands in different directions that capture the details of the image at a specific scale and direction. An example of the output is shown in FIG. 6.
[0046]
[0059] At 708, the conversion system 404 selects one or more of the sub-bands based on the determined conversion method. For example, in some videos, it may be more useful to use high-frequency sub-bands, and in some videos, it may be useful to use low-frequency sub-bands for analyzing film grain. As described above, one or more of the higher-frequency sub-bands can be used.
[0047]
[0060] At 710, the conversion system 404 converts the sub-bands into the frequency domain. For example, a fast Fourier transform can be used to convert the sub-bands from the spatial domain to the frequency domain. The result can be multiple frequency-domain representations for the direction of the sub-bands. For example, for sub-band 604-1, each representation for the four directions is converted into four representations in the frequency domain. At 712, the conversion system 404 outputs the frequency-domain representations of one or more sub-bands to the metric calculation system 406. Then, metric calculation can be performed.
[0048]
[0061] In some embodiments, the steerable pyramid can be applied after the conversion to the frequency domain. Here, the image can be converted to the frequency domain using a fast Fourier transform. Then, the steerable pyramid is applied to the frequency representation of the image to generate sub-bands in different directions. The resulting sub-bands can be similar to the sub-bands generated above.
[0049]
[0062] Metric calculation
[0063] FIG. 8 shows a simplified flowchart 800 for performing metric calculation according to some embodiments. At 802, the metric calculation system 406 receives sub-bands in the frequency domain for a reference image and a test image. The representation can be a two-dimensional array of values representing the output of a steerable pyramid in the frequency domain. The following can be performed for each direction in the sub-bands of the reference image and the test image. At 804, the metric calculation system 406 converts the 2D representation of the sub-band into a 1D representation. For example, each representation for a direction is converted into a 1D representation. Different methods can be used to convert a two-dimensional array representing a sub-band in the frequency domain into a one-dimensional array representing the sub-band in the frequency domain. The 1D representation may be a vector.
[0050]
[0064] At 806, the metric calculation system 406 generates a frequency distribution of the 1D representation of the sub-band. The distribution can be generated for each direction of the sub-band.
[0051]
[0065] At 808, the metric calculation system 406 determines the difference between the frequency distributions of the respective sub-bands for the reference image and the test image. Each direction for each band can be compared to each other. Then, the differences for the sub-bands can be combined. The differences for different directions of the sub-band can be combined in various ways. For example, a weighted average of different directions for a sub-band can be used to determine a score for the sub-band. Each sub-band can be associated with a score. In other embodiments, the directions may be combined for the sub-bands of the reference image and the test image, and then the difference is determined. Then, at 810, the metric conversion system 406 has scores 0, score 1,..., score nOutput scores for each sub-band such as this. In 812, the score adaptive fusion system 408 executes the fusion of these scores for the sub-bands in order to generate an evaluation score. As described above, the fusion can combine these scores using various methods. The following formula can be used.
[0052]
Number
[0053]
[0066] The evaluation may enable the processing system 106 to assess the similarity or divergence of film grain in the reference image and the test image. By using the frequency domain to perform the assessment, the evaluation of the similarity between film grains is improved. For example, even if the film grains have different values in the spatial domain, the film grains may appear similar to the user. Using a comparison in the frequency domain can capture film grains that appear similar but may have various values in the spatial domain. Also, decomposing an image into sub-bands using a steerable pyramid may enable the system to focus on sub-bands that can represent the film grain more accurately.
[0054]
[0067] Adaptive Region Detection
[0068] The server system 102 may use adaptive region detection to adaptively select the region to analyze for the comparison of film grain between the reference image and the test image. Adaptive region detection may select a region where the comparison of film grain may result in a more accurate evaluation score. For example, there may be regions where the film grain may be more prominent to a human viewer. In some embodiments, non-texture regions such as flat regions may be regions where the film grain may be more prominent. This is in contrast to high-texture regions where the texture hides the details of the film grain and makes the film grain less perceptible to a human viewer.
[0055]
[0069] The non-texture regions may include characteristics of lack of smoothness and detail, such as relatively uniform values or slight variations in pixel intensity in low-texture regions or non-texture regions. Also, the non-texture regions may have a minimum pattern or repeating structure. Examples of non-texture regions include plain walls, sky, etc. The texture regions include significant changes in pixel intensity and may include various patterns or structures that can be repeated. The visual content within the texture regions can change frequently. Different examples of texture regions may include hair, trees, etc. Slight changes or no changes in non-texture regions may enable film grain noise to be more easily perceptible to a human viewer, while the complex patterns and structures in texture regions may hide the fact that film grain noise is perceptible by a human viewer.
[0056]
[0070] When the entire image is used and in the comparison between a reference image and a test image, some regions may obscure the result of the comparison. For example, using information from high-texture regions in the comparison may distort the accuracy of the evaluation score. If film grain cannot be perceived by a human viewer in high-texture regions, the contribution to the evaluation score for film grain found in texture regions is appropriate compared to the comparison of film grain in low-texture regions where the difference may be significant to a human viewer. Thus, when regions containing film grain that can be more easily perceived by a human viewer are selected, the evaluation score may be a more accurate capture of the difference in film grain.
[0057]
[0071] System
[0072] FIG. 9 shows an example of a server system 102 according to some embodiments. The server system 102 includes an adaptive region detection system 902, a film grain evaluation system 104, and a processing system 106. The output of the adaptive region detection system 902 is regions R1,..., R N and regions T1,..., TN That is. Here, there are a plurality of regions for the reference image R, and the same regions may exist for the test image T. The adaptive region detection system 902 may determine a region by analyzing the reference image R to select a region, and then the corresponding region in the test image T is used. For example, other methods may also be used by analyzing the test image T or by analyzing a combination of the reference image R and the test image T.
[0058]
[0073] The film grain evaluation system 104 may analyze each region from the reference image R and the test image T. For example, region R1 and region T1 may be analyzed, or region R2 and region T2 may be analyzed, and so on. This may result in a plurality of scores from the regions, such as scores 1, 2,..., N for N regions. Then, the film grain evaluation system 104 may combine these scores into an evaluation score.
[0059]
[0074] Then, the processing system 106 may use the evaluation score to perform actions such as the same actions as described above with respect to FIG. 1.
[0060]
[0075] Region detection
[0076] The adaptive region detection system 902 can use various methods to select a region. In some embodiments, edge detection can be used. However, other features of the image can be used to select a region, such as segmentation for selecting a region, a neural network-based method for selecting a region, etc. In some embodiments, the adaptive region detection system 902 generates a frequency domain film grain metric using adaptive region selection. A film grain metric involving region selection for content and film grain can improve performance. This also significantly improves the integrity of the method and expands the application scenarios. If the region is flat, it mainly has a low frequency band and does not affect the high frequency film grain in the texture region. Therefore, the high frequency content is removed by region selection, and the image used can be content and film grain that do not affect the evaluation score due to the presence of film grain in the texture region.
[0061]
[0077] FIG. 10 shows a more detailed example of the server system 102 according to some embodiments. The edge detection system 1002 may use different edge detection methods such as a Sobel detector, a Canny detector, etc. The edge detection system 1002 may perform edge detection using non-overlapping shapes in the image. For example, blocks may be used to analyze a part of the image at a time. An N×N block may be used. In some examples, the size of the block can be 32×32, but other sizes such as 128×128, 256×256, etc. can also be used.
[0062]
[0078] Edge detection system 1002 can identify non-texture regions that may be regions where the texture to be evaluated does not meet the threshold. For each block, edge detection system 1002 may determine whether it is classified as a non-texture region or a texture region. The classification can be based on characteristic values of the block, such as gradient, variance, luma value, etc. In some embodiments, in 10-bit video, when the luma value is less than 130 or greater than 850, film grain becomes almost unobservable to human viewers. Evaluating the quality or similarity of film grain in these areas may be of no use. Therefore, edge detection system 1002 may consider these pixels to be high-texture pixels. Also, when the variance of a block is less than the threshold, this block may also have no film grain. Film grain can be random pixel values. If an area such as a pure white / black block has film grain, the variance of this block should be at least greater than 0. In other areas, if the variance is too small, the film grain is extremely small and becomes something different from what humans perceive. Therefore, this area may not need to be analyzed. Therefore, edge detection system 1002 may classify this block as a texture region. When the above threshold is not met (e.g., higher than the threshold), the block is classified as a non-texture region.
[0063]
[0079] Edge detection system 1002 outputs an edge map that summarizes the edges of an image. Non-texture area detection system 1004 can classify the edge map of each block as a non-texture area or a texture area. The film grain evaluation system may analyze each block as one area. The area can be formed from a single block. Also, multiple blocks may be combined into one area. For example, when multiple adjacent blocks are classified into the same classification, non-texture area detection system 1004 may combine the blocks of the same classification into one area. A single instance of film grain evaluation system 104 may be used.
[0064]
[0080] Multiple examples of film grain evaluation systems 104-1 to 104-N are shown that can analyze regions 1, 2,..., N in parallel. Each film grain evaluation system 104 may output a score for one region, such as scores 1, 2, and N being output for regions 1, 2,..., N. If each region is a block, there are a total of N blocks and N scores. Score adaptive fusion system 1006 may combine these scores and output an evaluation score that assesses the similarity between the film grain and the reference image and the test image. The final score may be calculated based on combining the N scores, such as using a weighted average of the N scores. The final score may be calculated by a weighted average of these N scores, such as using the following formula.
[0065]
Number
[0066] Here, w1, w2,..., w n are weights. Also
[0067]
Number
[0068] It is as follows.
[0069]
[0081] The weights can be obtained based on the ratings of the regions. For example, a region containing less texture can be weighted higher compared to a region containing more texture. The weights can also be determined based on other factors such as the position in the image, the size of some adjacent blocks, etc. A larger region of blocks can have more prominent film grain compared to an isolated single block, so the blocks can be weighted higher in a larger region compared to an isolated single block in a section of the image.
[0070]
[0082] FIG. 11 shows an input image, an edge map of the input image, and the output of non-texture region detection according to some embodiments. At 1100, the input image is shown. The input image can contain different contents. Here, the input image contains different striped bands of different colors. At 1102, the edge map of the input image is shown. The black or dark portions of the edge map may represent the positions where edges are detected within the original image, which are texture regions. These may be points where there are edges in the original input image and there are changes in pixel values such as intensity or color. The white or brighter regions may be regions where few or no edges are detected, which are non-texture regions. These regions may be places where there are few or no changes in pixel values such as intensity or color, which indicates a smoother and more uniform area in the image change.
[0071]
[0083] At 1104, the non-texture region detection is shown. The white or brighter regions may contain non-texture regions, and the black or darker regions may contain texture regions. The sizes of the blocks may be different and may form continuous portions of the image. However, within the continuous portions, there may be some N×N blocks that can be analyzed as special regions.
[0072]
[0084] FIG. 12 shows a simplified flowchart 1200 of a method for determining an evaluation score using adaptive region selection, according to some embodiments. At 1202, edge detection system 1002 processes an input image to determine a representation of the image. For example, an edge map of the image may be determined. The input image here may be a reference image or a test image, and processing may be performed on both images.
[0073]
[0085] At 1204, non-texture region detection system 1004 detects non-texture regions from the representation. For example, blocks of the representation may be analyzed to determine regions that are classified as non-texture regions.
[0074]
[0086] At 1206, the non-texture regions are input into film grain evaluation system 104 through output scores for each non-texture region. For example, each non-texture region may be analyzed and a respective score may be determined. At 1208, score adaptation fusion system 106 combines the scores to generate a fusion score. The combination may use weighted averaging or other methods. At 1210, processing system 106 uses the fusion score to determine an evaluation score.
[0075]
[0087] Accordingly, the evaluation score can be improved by using regions classified as non-texture regions. For example, by removing some regions where film grain may not be easily perceptible by a human viewer, the similarity score can be made more accurate by comparing regions where film grain may be more perceptible by a human viewer. This results in a more accurate evaluation score when determining whether the film grain in two images is similar.
[0076]
[0088] System
[0089] FIG. 13 shows an example of a computing device according to some embodiments. According to various embodiments, a system 1300 suitable for implementing the embodiments described herein includes a processor 1301, a memory module 1303, a storage device 1305, an interface 1311, and a bus 1315 (e.g., a PCI bus or other interconnect fabric). The system 1300 can operate as various devices such as any other device or service described herein. Although a particular configuration is described, various alternative configurations are possible. The processor 1301 can execute operations such as those described herein. Instructions for performing such operations can be embodied on one or more non-transitory computer-readable media in the memory 1303 or on some other storage device. Instead of or in addition to the processor 1301, various specially configured devices can also be used. The memory 1303 can be a random access memory (RAM) or other dynamic storage device. The storage device 1305 can include a non-transitory computer-readable storage medium that holds information, instructions, or some combination thereof, such as instructions that, when executed by the processor 1301, cause the processor 1301 to be configured to or capable of performing one or more operations of the methods described herein. The bus 1315 or other communication component can support the communication of information within the system 1300. The interface 1311 can be connected to the bus 1315 and configured to send and receive data packets over a network.Examples of supported interfaces include, but are not limited to, Ethernet®, Fast Ethernet, Gigabit Ethernet, Frame Relay, Cable, Digital Subscriber Line (DSL), Token Ring, Asynchronous Transfer Mode (ATM), High-Speed Serial Interface (HSSI), and Fiber Distributed Data Interface (FDDI). These interfaces may include ports suitable for communication with an appropriate medium. They may also include an independent processor and / or volatile RAM. A computer system or computing device may include, or communicate with, a monitor, printer, or other suitable display for providing any of the results described herein to a user.
[0077]
[0090] Any of the disclosed implementations may be embodied in various types of hardware, software, firmware, computer-readable media, and combinations thereof. For example, some of the techniques disclosed herein may be at least partially implemented by a non-transitory computer-readable medium that includes program instructions, state information, etc. for configuring a computing system to perform the various services and operations described herein. Examples of program instructions include both machine code generated by a compiler and higher-level code that can be executed via an interpreter. The instructions may be embodied in any suitable language, such as Java®, Python, C++, C, HTML, any other markup language, JavaScript®, ActiveX, VBScript, or Perl. Examples of non-transitory computer-readable media include, but are not limited to, magnetic media such as hard disks and magnetic tapes, optical media such as flash memory, compact discs (CDs), or digital versatile discs (DVDs), magneto-optical media, and other hardware devices such as read-only memory ("ROM") devices and random access memory ("RAM") devices. The non-transitory computer-readable media can be any combination of such storage devices.
[0078]
[0091] In the above specification, various techniques and mechanisms may sometimes be described in the singular for clarity. However, note that some embodiments include multiple repetitions of techniques or multiple instantiations of mechanisms, unless otherwise noted. For example, a system uses processors in various contexts, and multiple processors can be used while remaining within the scope of the present disclosure, unless specifically noted otherwise. Similarly, various techniques and mechanisms may sometimes be described as including a connection between two entities. However, since various other entities (such as bridges, controllers, gateways, etc.) may exist between two entities, a connection does not necessarily mean a direct, unobstructed connection.
[0079]
[0092] Some embodiments may be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to execute the methods described by some embodiments. The computer system may include one or more computing devices. The instructions are configured or operable to execute what is described in some embodiments when executed by one or more computer processors.
[0080]
[0093] As used in the description herein and throughout the following claims, the terms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in the description herein and throughout the following claims, the meaning of "in" includes "in" and "on" unless the context clearly dictates otherwise.
[0081]
[0094] The above description shows various embodiments, along with examples of how aspects of some embodiments may be implemented. The above examples and embodiments should not be regarded as the only embodiments, but are presented to illustrate the flexibility and advantages of some embodiments defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents may be employed without departing from the scope of the invention defined by the claims.
Claims
1. Receiving a first image and a second image for comparison of film grain; Analyzing the first image or the second image to determine a first texture representation of the first image or a second texture representation of the second image; Selecting a set of regions based on the first texture representation or the second texture representation; Converting the set of regions from a spatial domain to a frequency domain to generate a first frequency domain representation for the set of regions in the first image and a second frequency domain representation for the set of regions in the second image; Generating a score for evaluation of the difference in film grain in the first image and the second image based on the first frequency domain representation and the second frequency domain representation A method comprising the steps above.
2. The method according to claim 1, wherein the first texture representation is based on the texture of the content in the first image, or the second texture representation is based on the texture of the content in the second image.
3. Analyzing the first image or the second image comprises detecting a change in the content in the first image for determining the first texture representation or a change in the content in the second image for determining the second texture representation, the method according to claim 1.
4. Analyzing the first image or the second image comprises performing edge detection on the content in the first image for determining the first texture representation or on the content in the second image for determining the second texture representation, the method according to claim 1.
5. Analyzing the first image or the second image comprises comparing the characteristics of a plurality of regions with a threshold value; and adding a region to the set of regions when each characteristic satisfies the threshold value, the method according to claim 1.
6. Comparing the characteristics of the plurality of regions with the threshold value comprises comparing the lum value of a region with a first threshold value and a second threshold value; and adding the region to the set of regions when the lum value is between the first threshold value and the second threshold value, the method according to claim 5.
7. Comparing the characteristics of the plurality of regions with the threshold value includes comparing the variance of a region with the threshold value, and when the variance is greater than the threshold value, adding the region to the set of regions The method according to claim 5, comprising: **Claim 8** Analyzing the first image or the second image includes determining a first portion of a region classified as a non-texture region, determining a second portion of a region classified as a texture region, and adding the first portion of the region to the set of regions The method according to claim 1, comprising: **Claim 9** further comprising not adding the second portion of the region to the set of regions The method according to claim 8, further comprising: **Claim 10** In the method according to claim 9, regions within the second portion of the region include more detected edges than regions within the first portion of the region. **Claim 11** In the method according to claim 1, regions within the set of regions comprise blocks. **Claim 12** generating a set of scores for the set of regions and combining the set of scores to determine the score for the evaluation of the film grain difference The method according to claim 1, further comprising: **Claim 13** In the method according to claim 12, scores within the set of scores are weighted based on respective ratings of regions within the set of regions. **Claim 14** further comprising performing an action based on the score The method according to claim 1, further comprising: **Claim 15** Performing the action includes adjusting parameters of a process used to generate film grain for the second image. The method according to claim 14, comprising: **Claim 16** A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a computing device, cause the computing device to receive a first image and a second image for film grain comparison, analyze the first image or the second image to determine a first texture representation of the first image or a second texture representation of the second image, and select a set of regions based on the first texture representation or the second texture representation To generate a first frequency domain representation for the set of regions in the first image and a second frequency domain representation for the set of regions in the second image, convert the set of regions from the spatial domain to the frequency domain, Generate a score for evaluating the difference in film grain in the first image and the second image based on the first frequency domain representation and the second frequency domain representation A non-transitory computer-readable storage medium that causes the operation to be operable to perform. **Claim 17** The first texture representation is based on the texture of the content in the first image, or The non-transitory computer-readable storage medium according to claim 16, wherein the second texture representation is based on the texture of the content in the second image. **Claim 18** Analyzing the first image or the second image Comparing the characteristics of a plurality of regions with a threshold value, When each characteristic satisfies the threshold value, adding the region to the set of regions The non-transitory computer-readable storage medium according to claim 16, comprising: **Claim 19** Analyzing the first image or the second image Determining a first portion of the region classified as a non-texture region, Determining a second portion of the region classified as a texture region, Adding the first portion of the region to the set of regions The non-transitory computer-readable storage medium according to claim 18, comprising: **Claim 20** An apparatus comprising one or more computer processors and a computer-readable storage medium, wherein the computer-readable storage medium causes the one or more computer processors to Receive a first image and a second image for comparison of film grain, Analyze the first image or the second image to determine a first texture representation of the first image or a second texture representation of the second image, Select a set of regions based on the first texture representation or the second texture representation, Convert the set of regions from the spatial domain to the frequency domain to generate a first frequency domain representation for the set of regions in the first image and a second frequency domain representation for the set of regions in the second image, Generating a score for evaluating the difference in the film grain in the first image and the second image based on the first frequency domain representation and the second frequency domain representation An apparatus comprising instructions for controlling to be operable to perform the above.
Citation Information
Patent Citations
A method or an apparatus for estimating film grain parameters
WO2023275222A1