Scalable system for controlling color management with varying levels of metadata

A scalable color management system using multiple metadata levels addresses display variations to preserve the original intent of image and video content, improving rendering fidelity and reducing distortions.

JP2026032117APending Publication Date: 2026-02-25DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025202765
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2011-05-27
Filing Date
2025-11-25
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Existing image and video rendering technologies fail to accurately preserve the original intent of the content creator due to variations in display capabilities, leading to noticeable artifacts and distortions.

Method used

A scalable color management system that utilizes multiple levels of metadata to adjust image processing based on display characteristics, environmental conditions, and content-specific parameters, ensuring faithful reproduction of the intended image or video.

Benefits of technology

Enhances the fidelity of image and video rendering on diverse displays by maintaining the artistic intent and minimizing distortions, providing a more accurate viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026032117000001_ABST
    Figure 2026032117000001_ABST
Patent Text Reader

Abstract

PROVIDING VIDEO QUALITY WITH IMPROVED IMAGE FIDELITY SOLUTION: Some embodiments of scalable image processing systems and methods are disclosed herein in which the color management processing of source image data to be displayed on a target display (236) is varied according to different levels of metadata (204).SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to image processing, and more particularly to encoding and decoding of image and video signals using metadata, especially several layers of metadata. [Background technology]

[0002] Well-known scalable video encoding and decoding techniques allow for the expansion or compression of video quality depending on the capabilities of the target video display and the quality of the source video data. Summary of the Invention [Problem to be solved by the invention]

[0003] However, the use and application of image metadata, either a single level or several levels of metadata, enhances the rendering of images and / or videos and the viewer's experience. [Means for solving the problem]

[0004] Several aspects of scalable image processing systems and methods are disclosed herein whereby color management processing of source image data displayed on a target display is altered according to various levels of metadata.

[0005] In one aspect, a method is disclosed for processing and rendering image data on a target display with a set of levels of metadata, where the metadata is associated with image content, comprising: inputting image data; ascertaining a set of levels of metadata associated with the image data; performing at least one of a set of image processing steps including switching to default values ​​and adaptively calculating parameter values ​​if no metadata is associated with the image data; and, if metadata is associated with the image data, calculating color management algorithm parameters according to the set of levels of metadata associated with the image data.

[0006] In yet another aspect, a system for decoding and rendering image data on a target display with a series of levels of metadata is disclosed, the system including: a video decoder that receives input image data and outputs intermediate image data; a metadata decoder that receives the input image data, the metadata decoder being capable of detecting a series of levels of metadata associated with the input image data and outputting the intermediate metadata; a color management module that receives the intermediate metadata from the metadata decoder and the intermediate image data from the video decoder and performs image processing on the intermediate image data based on the intermediate metadata; and a target display that receives and displays the image data from the color management module. [Brief explanation of the drawings]

[0007] [Figure 1A] Figure 1A shows one embodiment of a conventional video pipeline, from video signal generation to distribution and processing. [Figure 1B] Figure 1B shows one embodiment of a conventional video pipeline, from video signal generation to distribution and processing. [Figure 1C]Figure 1C shows one embodiment of a conventional video pipeline, from video signal generation to distribution and processing. [Figure 2A] FIG. 2A illustrates one embodiment of a video pipeline including a metadata pipeline in accordance with the teachings of the present application. [Figure 2B] FIG. 2B illustrates one embodiment of a metadata prediction block. [Figure 3] FIG. 3 shows one embodiment of a sigmoid curve using level 1 metadata. [Figure 4] FIG. 4 shows one embodiment of a sigmoid curve using level 2 metadata. [Figure 5] Figure 5 shows one embodiment of a histogram plot based on image / scene analysis that can be used to adjust image / video mapping to a target display. [Figure 6] FIG. 6 shows one embodiment of image / video mapping adjusted based on Level 3 metadata including grading with a second reference display of image / video data. [Figure 7] FIG. 7 shows one embodiment of a linear mapping that can occur when the target display is closely matched to a second reference display that was used to color grade the image / video data. [Figure 8] FIG. 8 illustrates one embodiment of a video / metadata pipeline configured in accordance with the principles of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0008] Other features and advantages of the present system are presented in the following detailed description, which is to be read in conjunction with the drawings presented in this application. Exemplary embodiments are shown in the referenced figures in the drawings: The embodiments and figures disclosed herein are to be considered illustrative and not restrictive.

[0009] Throughout the following description, specific details are set forth to provide a more thorough understanding to those skilled in the art. However, well-known elements may not be shown or described in detail so as not to unnecessarily obscure the disclosure. Accordingly, the description and drawings should be regarded in an illustrative rather than a restrictive sense.

[0010] [overview] One aspect of video quality relates to rendering an image or video on a target display with the same or nearly the same fidelity as intended by the image or video creator. It is desirable to have a color management (CM) scheme that attempts to preserve the original appearance of video content on displays with varying capabilities. To accomplish this goal, it is desirable for such a CM algorithm to be able to predict how the video will appear to a viewer in the post-production environment where the video is finalized.

[0011] To illustrate the challenges inherent in this application and system, one embodiment of a conventional video pipeline 100 that proceeds from the generation, distribution, and processing of a video signal is shown in Figures 1A, 1B, and 1C.

[0012] The video signal generation section 102 may involve color grading 104 of the video signal by a color grader 106 that may grade the signal for various image characteristics, such as brightness, contrast, and color rendering of the input video signal. The color grader 106 may grade the signal to generate an image / video mapping 108, such grading being done to a reference display device 110 that may have, for example, a gamma response curve 112.

[0013] Once the signal is graded, the video signal can be sent out via distribution 114, which in this case is appropriately considered to be widespread. For example, distribution could be via the Internet, DVD, cinema screening, etc. In this example, distribution is shown in FIG. 1A as feeding the signal to a target display 120 with a maximum luminance of 100 nits and a gamma response curve 124. Assuming that the reference display 110 has approximately the same maximum luminance and approximately the same response curve as the target display, then the mapping applied to the video signal can be as simple as a 1:1 mapping 122, e.g., made in accordance with the Rec. 709 standard for color management 118. Keeping all other factors the same (e.g., ambient lighting conditions at the target display), what you see on the reference display is essentially what you will see on the target display.

[0014] This situation may be modified, for example, as shown in FIG. 1B, where the target display 130 differs from the reference display 110 in some respects, such as maximum luminance (500 nits compared to 100 nits for the reference display). In this case, mapping 132 may be a 1:5 mapping for rendering on the target display. In such a case, the mapping is a linear stretching with Rec. 709 CM blocks. The resulting distortion from the reference display's appearance to the target display's appearance may or may not be annoying to the viewer, depending on their individual sensitivity. For example, dark and mid-tone colors may be tolerable even when stretched. Also, MPEG blocking artifacts may become more noticeable.

[0015] FIG. 1C shows an even more extreme example. In this example, the target display 140 may have a more noticeable difference from the reference display. For example, the maximum luminance of the target display 140 is 1000 nits, compared to 100 nits for the reference display. If the same linear stretch mapping 142 were applied to the video signal fed to the target display, a much more noticeable and objectionable distortion may be perceived by the viewer. For example, the video content may be displayed at a significantly higher luminance level (1:10 ratio). Dark and mid-tone colors may be stretched to the point where camera noise from the original capture becomes noticeable, and banding becomes more noticeable in dark areas of the image. Additionally, MPEG blocking artifacts may become more noticeable.

[0016] Rather than exhaustively considering all possible examples of how objectionable artifacts may appear to a viewer, it may be useful to consider a few more. For example, suppose the maximum luminance of a reference display (e.g., 600 nits) is higher than that of a target display (e.g., 100 nits). In this case, the mapping is still a 6:1 linear stretch, which may cause content to be displayed entirely at a lower luminance level, the image to appear dark, and the dark details of the image may contain noticeable crush.

[0017] In yet another example, assume that the maximum luminance of a reference display (e.g., 600 nits) differs from that of a target display (e.g., 1000 nits). When linear scaling is applied, even a slight ratio difference (i.e., close to 1:2) can result in a significant and potentially unpleasant difference in maximum luminance. The difference in luminance can cause the image to appear too bright and painful to view. Midtones can appear unnaturally stretched and washed out. Additionally, both camera noise and compression artifacts can be noticeable and unpleasant. In yet another example, assume that the reference display has a color gamut equivalent to P3 and the target display has a color gamut smaller than Rec709. Assume that content was color-graded on the reference display, but the rendered content has a color gamut equivalent to that of the target display. In this case, mapping the content from the reference display's gamut to the target gamut can unnecessarily compress the content and cause it to appear less saturated.

[0018] Without some intelligent (or at least more accurate) model of image rendering on the target display, slight distortions or unpleasant artifacts will likely be visible to the viewer of the image / video. In fact, what the viewer experiences is likely not what the image / video producer intended. While this discussion focuses on luminance, we see that similar concerns apply to color as well. In fact, if there are differences between the color spaces of the source and target displays, and these differences are not properly taken into account, color distortions will also result in noticeable artifacts. The same considerations apply to ambient differences between the source and target displays.

[0019] [Use Metadata] As these examples show, it is desirable to know the characteristics and capabilities of the reference display, the target display, and the source content in order to produce video that is as faithful to the original intent as possible. There is other data, known as "metadata," that describes and conveys various aspects of the source image data and is useful for such faithful rendering.

[0020] While tone and gamut mappers typically work well for approximately 80–95% of images processed for a particular display, there are problems with using such general-purpose solutions to process images. Such methods typically do not guarantee that the image displayed on the screen will match the director's or original creator's intent. It has also been noted that different tone or gamut mappings may work better with various types of images or better preserve the atmosphere of the image. It has also been noted that various tone and gamut mappings may result in clipping and loss of detail, or color or hue shifts.

[0021] When tone mapping a color-graded image sequence, color grading parameters such as the content's minimum black level and maximum white level may be desirable parameters to drive tone mapping of the color-graded content to a particular display. The color grader has already given the content its desired look (on an image-by-image and even a time basis). It may be desirable to preserve the perceived viewing experience of the image sequence when converting it for a different display. It should be recognized that additional levels of metadata may allow for improved protection of such appearance.

[0022] For example, suppose a sunrise scene is photographed and professionally color-graded on a 1000-nit reference display. In this example, the content is mapped for display on a 200-nit display. Images taken before the sun rise may not use the full range of the reference display (e.g., up to 200 nits). As soon as the sun rises, the image sequence will use the full 1000-nit range, which is the upper limit of the content. Without metadata, many tone mapping techniques use a maximum value (e.g., luminance) as a guideline for how to map the content. In this case, the tone curve applied to the image before sunrise (1:1 mapping) may be different from the tone curve applied to the image after sunrise (5x tone compression). As a result, the image displayed on the target display may have the same peak luminance before and after sunrise, a distortion of the creative intent. The artist intended the image to be darker before sunrise and brighter during sunrise, as created on the reference display. In such scenes, metadata can be defined that fully describes the dynamic range of the scene, and its use can ensure that artistic effects are maintained, and can also be used to minimize temporal variations in luminance from scene to scene.

[0023] As yet another example, consider the reverse of the above. Assume Scene 1 is graded for 350 nits and is shot outdoors in natural light. If Scene 2 is shot in a dark room and displayed in the same range, Scene 2 will appear too dark. In this case, metadata can be used to define the appropriate tone curve and ensure Scene 2 appears appropriate. In yet another example, assume the reference display has a color gamut equivalent to P3 and the target display has a color gamut smaller than Rec709. Assume the content was color graded on the reference display, but the rendered content has a color gamut equivalent to the target display. By using metadata defining the color gamut of the content and the color gamut of the source display, mapping can make intelligent decisions and map the content's color gamut 1:1. This can ensure that the color saturation of the content remains intact.

[0024] In some embodiments of the system, tone and gamut need not be treated as separate entities or terms in a series of images / videos. "Memory colors" are colors in an image that, if adjusted incorrectly, will appear incorrect, even if the viewer may not know the original intent. Skin tones, sky, and grass are good examples of memory colors, and their hues may be altered during tone mapping in a way that appears incorrect. In one embodiment, the gamut mapper has knowledge (as metadata) of protected colors in the image, ensuring that their hues are preserved in the tone mapping process. This metadata can be used to define and highlight protected colors in the image, thereby ensuring accurate processing of memory colors. The ability to define local tone mapping and gamut mapping parameters is an example where metadata is not necessarily derived solely from the parameters of the reference and / or target displays.

[0025] [One embodiment of robust color management] In some embodiments of the present application, systems and methods are disclosed for providing a robust color management scheme, which employs several sources of metadata to provide higher image / video fidelity consistent with the original intent of the content creator. As described in more detail herein, in one embodiment, various sources of metadata can be added to the processing depending on the availability of some metadata.

[0026] FIG. 2A illustrates, by way of example only, a high-level block diagram of an image / video pipeline 200 using metadata. Image generation and post-production can occur in block 202. A video source 208 is input to a video encoder 210. Concurrently with the video source being ingested, metadata 204 is provided to a metadata encoder 206. Examples of metadata 204 have been discussed above and may include items such as the color gamut boundaries and other parameters of the source and / or reference display, the environment of the reference display, and other encoding parameters. In one embodiment, the metadata accompanies the video signal as a subset of the metadata and is temporally and spatially co-located with the video signal being rendered at that time. The metadata encoder 206 and video encoder 210 can be collectively considered a source image encoder.

[0027] The video signals and metadata are then distributed by distribution 212 in any suitable manner, such as multiplexed, serial, parallel, or any other known manner. It should be recognized that distribution 212 should be considered broad for purposes of this application. Suitable distribution methods may include the Internet, DVD, cable, satellite, wireless, wired, etc.

[0028] The video signal and metadata thus delivered are provided to a target display environment 220. A metadata decoder 222 and a video decoder 224 each receive the respective data streams and provide decoding appropriate to the characteristics of the target display, among other factors. The metadata at this point can preferably be sent to either a third-party color management (CM) block 220 and / or one of the embodiments of a CM module 228 of the present application. If the video and metadata are processed by the CM block 228, a CM parameter generator 232 can take the metadata from the metadata decoder 222 as input, along with the metadata from the metadata prediction block 230.

[0029] The metadata prediction block 230 can make some predictions for higher fidelity rendering based on knowledge of previous images or video scenes. The metadata prediction block collects statistics from the input video stream to estimate metadata parameters. One possible embodiment of the metadata prediction block 230 is shown in FIG. 2B. In this embodiment, a histogram 262 of the logarithm of image luminance can be calculated for each frame. An optional low-pass filter 260 can precede the histogram to (a) reduce the histogram's sensitivity to noise and / or (b) partially account for natural blurring in the human visual system (e.g., humans perceive dither patterns as solid-color patches). A minimum point 266 and a maximum point 274 are then obtained. Furthermore, a toe point 268 and a shoulder point 272 can be obtained based on percentile settings (such as 5% and 95%). Furthermore, a geometric mean 270 (logarithmic mean) can be calculated and used as a midpoint. These values ​​may be temporally filtered, for example, to prevent them from changing too quickly. These values ​​may also be reset, if desired, upon a scene change. Scene changes may be detected by the insertion of black frames, or by very abrupt changes in the histogram, or other such techniques. Clearly, scene change detector 264 may detect scene changes from histogram data, as shown, or directly from the video data.

[0030] In yet another embodiment, the system may calculate the average of image intensity values ​​(luminance), which may be scaled by a perceptual weighting such as a logarithm, a power function, or a look-up table (LUT). The system may then estimate highlight and shadow regions (e.g., headroom and footroom in FIG. 5) from predetermined percentile values ​​(e.g., 10% and 90%) of the image histogram. Alternatively, the system may estimate highlight and shadow regions when the slope of the histogram is above or below a certain threshold. Many variations are possible; for example, the system may calculate the maximum and minimum values ​​of the input image, or from predetermined percentile values ​​(e.g., 1% and 99%).

[0031] In other embodiments, the values ​​may be stabilized over time (e.g., from frame to frame), such as by fixed increase and decrease rates. Because sudden changes may indicate a scene change, stabilization over time may not be applied to these values. For example, if the change is below a certain threshold, the system may limit the rate of change; otherwise, it proceeds with the new value. Alternatively, the system may disallow certain values ​​(such as letterbox or zero values) from affecting the shape of the histogram.

[0032] Additionally, the CM parameter generator 232 can employ other (i.e., not necessarily content-creation-based) metadata, such as display parameters, display ambient conditions, and user selection of factors, for color management of the image / video data. As will be apparent, display parameters can be made available to the CM parameter generator 232 through a standard interface, such as Extended Display Identification Data (EDID), via an interface (such as a Display Data Channel (DDC) serial interface, High-Definition Multimedia Interface (HDMI), or Digital Visual Interface (DVI)). Additionally, display ambient conditions data can be provided by an ambient light sensor (not shown) that measures ambient light conditions and their reflection from the target display.

[0033] Once the CM parameter generator 232 receives any appropriate metadata, it can set parameters in downstream CM algorithms 234 that can handle the final mapping of image / video data on target display 236. It should be appreciated that the illustrated branch of functionality between the CM parameter generator 232 and the CM algorithms 234 is not required; indeed, in some embodiments, these functions may be combined into a single block.

[0034] As should be apparent, the various blocks of Figures 2A and 2B are similarly optional from the perspective of this embodiment, and numerous other embodiments can be constructed by incorporating or omitting these enumerated blocks and are within the scope of this application. Additionally, CM processing can be performed at a variety of different points in the image pipeline 200 and not necessarily as shown in Figure 2A. For example, target display CM can be located and included within the target display itself, or such processing can be performed in the set-top box. Alternatively, target display CM can be performed at the point of distribution or post-production, depending on what level of metadata processing is available or deemed appropriate.

[0035] [Scalable color management using various levels of metadata] In some embodiments of the present application, systems and methods are disclosed for providing a scalable color management scheme, whereby several metadata sources can deliver a series of different levels of metadata to provide images / video with greater fidelity to the original intent of the content creator. As described in more detail herein, in one embodiment, depending on the availability of some metadata, different levels of metadata can be added to the processing.

[0036] In many embodiments of the present system, appropriate metadata algorithms can take into account many pieces of information, such as: (1) Encoded video content (2) A method for converting encoded content into linear light. (3) The color gamut boundaries of the source content (both luminance and chrominance) (4) Information about the post-production environment A method for converting to linear light may be desirable so that the actual image appearance (brightness, gamut, etc.) observed by the content creator can be calculated. The gamut boundary is useful for pre-specifying the outermost colors so that they can be mapped to the target display without clipping or too much overhead. Information about the post-production environment is desirable so that external factors that may affect the appearance on the display can be modeled.

[0037] In current video distribution mechanisms, only encoded video content is delivered to the target display. It is assumed that the content was created in a reference studio environment using a reference display compliant with Rec601 / 709 and various SMPTE (Society of Motion Picture and Television Engineers) standards. The target display system is generally assumed to be Rec601 / 709 compliant, and the target display environment is largely ignored. This basic assumption that both the post-production display and the target display are Rec601 / 709 compliant means that neither display may be upgraded without introducing some degree of image distortion. In reality, some distortion may already be introduced because Rec601 and Rec709 use slightly different primary color selections.

[0038] One embodiment of a scalable system of metadata levels is disclosed herein that allows for the use of reference and target displays with a wider and more sophisticated range of capabilities. Several different metadata levels allow the CM algorithm to adjust source content for a given target display with increasing accuracy. The following sections describe several proposed levels of metadata.

[0039] [Level 0] Level 0 metadata is the default case and effectively means zero metadata. Metadata may be absent for several reasons, including:

[0040] (1) The content creator did not include the metadata (or the metadata was lost at some point in the post-production pipeline). (2) The display switches content (i.e., channel surfing or commercial breaks).

[0041] (3) Corruption or loss of data. In one embodiment, it may be desirable for CM processing to address level 0 (i.e., the absence of metadata) by either inferring metadata based on video analysis or by assuming default values.

[0042] In such an embodiment, the color management algorithm may be able to operate in at least two different ways in the absence of metadata: Switch to the default value.

[0043] In this case, the display would operate much like current distribution systems, where the characteristics of a post-production reference display are assumed. The assumed reference display can vary depending on the video encoding format. For example, for 8-bit RGB data, a Rec601 / 709 display can be assumed. For higher bit-depth RGB data or LogYuv-encoded data, the P3 or Rec709 color gamut can be assumed if color-graded on a professional monitor (such as a ProMonitor) in 600 nit mode. This can work well when there is only one standard or de facto standard for higher dynamic range content. On the other hand, if the higher dynamic range content is created under custom conditions, the results may not be significantly improved and may be poor.

[0044] Parameter values ​​are adaptively calculated. In this case, the CM algorithm starts with some default assumptions and can modify those assumptions based on information gained from analyzing the source content. Typically, this involves analyzing the histogram of the video frame to determine how best to adjust the luminance of the input source, possibly by calculating parameter values ​​for the CM algorithm. This can risk creating an "auto-exposed"-type look, where each scene or frame is balanced to the same luminance level. Also, depending on the format, other challenges may be presented; for example, if the source content is in RGB format, there is currently no automated method for determining its color gamut.

[0045] In other embodiments, a combination of the two approaches can be implemented: for example, color gamut and coding parameters (such as gamma) can assume standard default values, and the histogram can be used to adjust the brightness levels.

[0046] [Level 1] In this embodiment, Level 1 metadata provides information describing how the source content was created and packaged. This data allows the CM process to predict how the video content actually appeared to the content creator. Level 1 metadata parameters can be categorized into three areas:

[0047] (1) Video coding parameters (2) Source and display parameters (3) Color gamut parameters of the source content (4) Environmental parameters Video Coding Parameters Because many color management algorithms work, at least in part, in a linear light space, it may be desirable to have a way to convert coded video into a linear (but relative) (X,Y,Z) representation that is either inherent to the coding scheme or provided as metadata itself. For example, coding schemes such as LogYuv, OpenEXR, LogYxy, or LogLuv TIFF inherently contain the information necessary to convert to a linear light format. However, for many RGB or YCbCr formats, additional information such as gamma and primaries may be desired. As an example, the following information may be supplied to process YCbCr or RGB input:

[0048] (1) The coordinates of the primary colors and white point used to encode the source content, which can be used to generate an RGB to XYZ color space transformation matrix, the (x,y) coordinates of red, green, blue, and white, respectively.

[0049] (2) Minimum and maximum code values ​​(e.g., "normal" or "full" range), which can be used to convert code values ​​to normalized input values. (3) Response curves (e.g., "gamma"), either overall or per channel for each primary color, that can be used to linearize intensity values ​​by canceling out any nonlinear responses that may have been applied by the interface or reference display.

[0050] Color gamut parameters of the source display It can be useful for color management algorithms to know the color gamut of the source display. These values ​​correspond to the capabilities of the reference display used to grade the content. The color gamut parameters of the source display, preferably measured in a completely dark environment, can include the following:

[0051] (1) Primary colors given as CIE·xy chromaticity coordinates or XYZ, specified along with maximum luminance. (2) Tristimulus values ​​for white and black, such as CIE XYZ.

[0052] Color gamut parameters of the source content It can be useful for color management algorithms to know the boundaries of the color gamut used to create the source content. Generally, these values ​​correspond to the capabilities of the reference display used to grade the content. However, they may vary depending on software settings or if only a portion of the display's capabilities is used. In some cases, the color gamut of the source content may not match the gamut of the encoded video data. For example, the video data may be encoded in LogYuv (or other encoding) that covers the entire visible spectrum. Source gamut parameters may include:

[0053] (1) Primary colors given as CIE·xy chromaticity coordinates or XYZ, specified along with maximum luminance. (2) Tristimulus values ​​for white and black, such as CIE XYZ.

[0054] Environmental parameters In certain situations, simply knowing the luminance levels of a reference display may not be enough to determine how the source content "appeared" to the viewer in post-production. Information about the light levels produced by the surrounding environment can also be useful. The combined light from both the display and the environment is the signal that the human eye sees, creating a "look." It may be desirable to maintain this look throughout the video pipeline. Environmental parameters, preferably measured in a typical color grading environment, may include:

[0055] (1) The ambient color of a reference monitor given as absolute XYZ values, which can be used to estimate the viewer's level of adaptation to their own environment. (2) The absolute XYZ values ​​of the black level of a reference monitor in a typical color grading environment. These values ​​can be used to determine the effect of ambient lighting on the black level.

[0056] (3) The color temperature of the ambient light, given as absolute XYZ values ​​of a white reflective sample (such as paper) in front of the screen. This value can be used to estimate the viewer's adaptation to the white point. As mentioned above, Level 1 metadata can provide the color gamut, encoding, and environmental parameters of the source content. This can enable CM solutions to predict how the source content will have looked when it was approved. However, it does not provide much guidance on how to best adjust color and brightness for the target display.

[0057] In one embodiment, applying a single sigmoidal curve globally to a video frame in RGB space may be a simple and stable way to map between different source and target dynamic ranges. Furthermore, a single sigmoidal curve can be used to modify each channel (R, G, B) independently. Such a curve can also be sigmoidal in some perceptual space, such as logarithmic or power function. An example curve 300 is shown in FIG. 3. Obviously, other mapping curves are suitable, such as linear mapping (as shown in FIGS. 3, 4, and 6) or other mappings such as gamma.

[0058] In this case, the minimum and maximum points on the curve are known from Level 1 metadata and information about the target display. The exact shape of the curve can be static, or it can be something that has been found to work well on average based on the input and output ranges, or it can be adaptively modified based on the source content.

[0059] [Level 2] Level 2 metadata provides additional information about the characteristics of the source video content. In one embodiment, Level 2 metadata may divide the luminance range of the source content into specific luminance regions. More specifically, in one embodiment, the luminance range of the source content may be divided into five regions, where the regions may be defined by points along the luminance range. Such ranges and regions may be defined in a single image, in a series of images, in a single video scene, or in multiple video scenes.

[0060] For illustrative purposes, one embodiment using Level 2 metadata is shown in Figures 4 and 5. Figure 4 shows a mapping 400 of input luminance to output luminance on a target display. In this case, the mapping 400 is shown as an approximately sigmoidal curve with a series of breakpoints along the curve. The points can correspond to image processing related values, in this figure, min in , foot in , mid in , head in , max in It displays:

[0061] In this embodiment, min in and max in can correspond to the minimum and maximum luminance values ​​in the scene. in can be the median value corresponding to the perceptual "mean" luminance value or "median tone". The last two points, foot in and head in may be footroom and headroom values. The area between the footroom and headroom values ​​may define a significant portion of the dynamic range of the scene. It may be desirable to ensure that content between these points is preserved as much as possible. Content below the footroom may be compressed as needed. Content above the headroom corresponds to the highlights and may be clipped as needed. It should be recognized that these points tend to define the curve itself, and therefore other embodiments may result in a curve that best fits these points. Furthermore, such a curve may assume a linear, gamma, sigmoid, or other appropriate and / or desirable shape.

[0062] Further to this embodiment, FIG. 5 illustrates the minimum, footroom, midpoint, headroom, and maximum points as they appear on a histogram plot 500. Such a histogram may be generated per image, per video scene, or even per series of video scenes, depending on the granularity of the histogram analysis desired to help maintain content fidelity. In one embodiment, these five points may be specified by code values ​​in the same coded representation as the video data. Note that, while min and max typically correspond to the same values ​​in the range of the video signal, this is not always the case.

[0063] Depending on the granularity and frequency of such histogram plots, histogram analysis can be used to dynamically redefine points along the luminance map of FIG. 4, thereby changing the curve over time. This can also be used to improve the fidelity of the content presented to a viewer on a target display. For example, in one embodiment, periodically transmitting the histogram potentially allows the decoder to obtain more information than just min, max, etc. The encoder can also include a new histogram only when there is a significant change, saving the decoder the trouble of having to calculate it on the fly for each frame. In yet another embodiment, the histogram is used to estimate metadata, either to replace missing metadata or to supplement existing metadata.

[0064] [Level 3] In one embodiment, Level 3 metadata can employ parameters from Level 1 and Level 2 metadata for a second reference grading of the source content. For example, the primary grading of the source content can be performed on a reference monitor (e.g., a ProMonitor) using the P3 color gamut at 600 nits of luminance. Level 3 metadata can also provide information about a secondary grading, which can be performed on a CRT reference, for example. In this case, the additional information would indicate Rec601 or Rec709 primaries and a lower luminance, such as 120 nits. The corresponding min, foot, mid, head, and max levels are also fed to the CM algorithm.

[0065] Level 3 metadata allows for the addition of additional data such as color gamut, environment, and color primaries, as well as luminance level information from a second reference grading of the source content. This additional information can then be used to define a sigmoid curve 600 (as shown in Figure 6) that maps the primary input to the range of the reference display. Figure 6 shows an example of how the input and reference display (output) levels can be combined to form an appropriate mapping curve.

[0066] If the capabilities of the target display closely match those of the secondary reference display, this curve can be used directly to map the primary source content. On the other hand, if the capabilities of the target display are intermediate between those of the primary and secondary reference displays, the mapping curve for the secondary display can be used as a lower bound. The curve used for the actual target display can then be an interpolation between no reduction (e.g., linear mapping 700 as shown in Figure 7) and the full range reduction curve generated using the reference levels.

[0067] [Level 4] Level 4 metadata is the same as Level 3 metadata, except that the metadata for the second reference grading is adjusted to match the actual target display.

[0068] Level 4 metadata can also be implemented in over-the-top (OTT) scenarios (i.e., Netflix, mobile streaming, or other video-on-demand (VOD) services) where the actual target display transmits its characteristics to the content provider, and the content is delivered along with the best available curve. In one such embodiment, the target display can communicate with a video streaming service, VOD service, or the like, and provide information, such as its EDID data or other available appropriate metadata, to the data streaming service. Such a communication path is shown in FIG. 2A as dotted path 240 to either a video encoder and / or metadata encoder (210 and 206, respectively), as known in the art for services such as Netflix. Typically, Netflix and other such VOD services monitor the amount of data and data throughput rate to the target device, and the metadata is not necessarily for color management. However, for purposes of this embodiment, it is sufficient if the metadata is transmitted from the target data via delivery 212 or other means (in real time or in advance) to creation or post-production to modify the color, tone, or other characteristics of the image data delivered to the target display.

[0069] The reference luminance levels provided by Level 4 metadata are specific to the target display. In this case, a sigmoid curve can be constructed as shown in Figure 6 and used directly without interpolation or adjustment.

[0070] [Level 5] Level 5 metadata extends Level 3 or Level 4 by identifying distinctive features such as:

[0071] (1) Protected Colors: Colors in an image that are identified as common memory colors that should not be processed, such as skin tones, sky colors, grass colors, etc. Areas of an image with such protected colors can have image data supplied to the target display without modification.

[0072] (2) Prominent highlights: Shows light sources, maximum emissive and specular highlights. (3) Out of Gamut: Features in an image that have been intentionally color graded outside the color gamut of the source content.

[0073] In some embodiments, objects shown this way can be artificially mapped to the display's maximum value if the target display can accommodate higher luminances. If the target display can accommodate lower luminances, the objects can be clipped to the display's maximum value without compensating for detail. In this case, the objects are ignored, and a defined mapping curve is applied to the remaining content, preserving more detail.

[0074] It should be recognized that in some embodiments, such as when attempting to map a VDR down to a lower dynamic range display, knowing the light sources and highlights may be useful because it may be possible to clip them without too much impact. As an example, a brightly lit face that is backlit (i.e., not at all from a light source) is not a feature that you would want to clip. Alternatively, such features may be compressed more gradually. In yet another embodiment, if the target display can accommodate a wider color gamut, those content objects may be stretched to expand to the full capabilities of the display. In yet another embodiment, the system may ignore mapping curves defined to ensure highly saturated colors.

[0075] It should be recognized that in some embodiments of the present application, the levels themselves do not constitute a strict hierarchy of metadata processing. For example, level 5 can be applied to either level 3 or level 4 data. Also, some lower-numbered levels may not exist, and the system can process higher-numbered levels if they exist.

[0076] One embodiment of a system using multiple metadata levels As discussed above, various metadata levels provide increasing information about the source material, allowing the CM algorithm to provide increasingly accurate mappings for the target display. One embodiment employing such scalable levels of metadata is shown in Figure 8.

[0077] The illustrated system 800 shows an overall video / metadata pipeline through five blocks: creation 802, container 808, encoding / distribution 814, decoding 822, and consumption 834. Clearly, numerous variations of several different implementations are possible, some with more blocks and some with fewer blocks. The scope of this application should not be limited to the description of the embodiments herein, and indeed, the scope of this application encompasses these various implementations and embodiments.

[0078] The generator 802 generally acquires image / video content 804 and processes it using a color grading tool 806 as described above. The processed video and metadata are placed in a suitable container 810, e.g., in any suitable format or data structure for subsequent distribution, as known in the art. As an example, the video may be stored and delivered as VDR color-graded video, and the metadata may be delivered as VDR XML formatted metadata. This metadata is separated into several levels, as shown in 812. The container block may embed data in the formatted metadata that encodes up to which level of metadata is available and associated with the image / video data. It should be recognized that not all levels of metadata need be associated with the image / video data; however, downstream decoding and rendering may be able to ascertain and process such available metadata accordingly, regardless of which metadata and levels are associated.

[0079] Encoding can proceed by obtaining the metadata and feeding it to an algorithm parameter determination block 816, while the video can be fed to an AVCVDR encoder 818 which includes a CM block for processing the video prior to distribution 820.

[0080] Once delivered (e.g., via the Internet, DVD, cable, satellite, wireless, wired, etc., which are considered broad), decoding of the video data / metadata can proceed to an AVCVDR decoder 824 (or, optionally, to a legacy decoder 826 if the target display is not VDR-enabled). Both the video data and metadata are restored through decoding (as blocks 830 and 828, respectively, and possibly block 832 if the target display is legacy). The decoder 824 takes the input image / video data and restores the input image data and / or can separate it into an image / video data stream for further processing and rendering, and a metadata stream for calculating parameters for subsequent CM algorithm processing on the rendered image / video data stream. The metadata stream further includes information on whether there is metadata associated with the image / video data stream. If there is no associated metadata, the system can proceed with level 0 processing as described above. Otherwise, the system can proceed with further processing with all metadata associated with the image / video data stream according to the series of different levels of metadata as described above.

[0081] As will be apparent, the presence or absence of metadata associated with the image / video data being rendered can be determined in real time. For example, some sections of a video stream may have no metadata associated with them (either due to data corruption or the content creator's intention that there is no metadata), while other sections may have metadata available, and possibly many metadata at various levels associated with other sections of the video stream. This may be the intention of the content creator, but in at least one embodiment of the present application, it should be possible to make a real-time or substantially dynamic determination as to the presence or absence of metadata, or what level of metadata, associated with such a video stream.

[0082] In the processing blocks, an algorithmic parameter determination block 836 can restore previous parameters, which may have been processed prior to distribution, or recalculate parameters based on metadata from the target display and / or target environment (which may be, for example, from a standard interface such as an EDID or modern VDR interface, as well as input from a viewer or sensors in the target environment, as discussed above in the context of the embodiments of FIGS. 2A and / or 2B). Once the parameters are calculated or restored, they can be sent to one or more of the CM systems (838, 840, and / or 842) for final mapping of source and intermediate image / video data to the target display 844 according to some embodiments disclosed herein.

[0083] In other embodiments, the implementation blocks of Figure 8 need not be subdivided. For example, generally speaking, the processes comprising the algorithm parameter determination and the color management algorithm itself do not necessarily need to be bifurcated as shown in Figure 8, but can contemplate and / or be implemented as a color management module.

[0084] Additionally, while this specification describes a series of various levels of metadata for use in a video / image pipeline, it should be recognized that in practice, the system need not process the image / video data in the order in which the metadata levels are numbered. In fact, some levels of metadata may be available during rendering while others may not. For example, second reference color grading may or may not be performed, and level 3 metadata may or may not be present during rendering. A system configured in accordance with this application will continue to process the metadata optimally at the time, taking into account the presence or absence of various levels of metadata.

[0085] The foregoing provides a detailed description of one or more embodiments of the present invention, to be read in conjunction with the accompanying drawings, illustrating the principles of the invention. It should be recognized that while the invention has been described in connection with such embodiments, the invention is not limited to any particular embodiment. The scope of the present invention is limited only by the claims, and the present invention encompasses numerous alternatives, modifications, and equivalents. Numerous specific details are set forth herein to provide a thorough understanding of the present invention. These details are provided for illustrative purposes, and the present invention may be practiced according to the claims without some or all of these specific details. For purposes of clarity, technical matters known in the art related to the present invention have not been described in detail so as not to unnecessarily obscure the invention.

Claims

1. 1. An apparatus for generating a bitstream containing encoded image data and metadata, comprising: an input for accessing a source video signal to generate encoded image data; and a processor; The processor: generating output metadata associated with the encoded image data; generating a bitstream including the encoded image data and the output metadata; the output metadata includes a first set of metadata and a second set of metadata; the first set of metadata is associated with a portion of the image data and includes one or more parameters that are indicative of characteristics of a reference display device used in creating the image data for color grading the source video signal; The one or more parameters are: a. the white point of the reference display device; b. the three primary colors of the reference display device; c) at least one brightness level of the reference display device; The second set of metadata is associated with the same portion of the image data and includes at least a maximum brightness level of the image data.

2. The apparatus of claim 1 , wherein the first set of metadata and the second set of metadata are separated separately in the bitstream.

3. The apparatus of claim 1 , further comprising: distributing the bitstream.