Multi-step display mapping and metadata reconstruction for HDR video
By using a multi-step display mapping method, the input image is first mapped to a base image in the intermediate dynamic range, and then re-mapped according to the characteristics of the target display device. This solves the limitations of traditional single-step mapping in terms of computing resources and power consumption, and achieves efficient and low-power display management.
Patent Information
- Application Number
- JP2024519016
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-03
- Filing Date
- 2022-09-28
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-09-28
AI Technical Summary
Existing technologies suffer from limitations in computing resources and power consumption when managing the display of high dynamic range (HDR) content, especially on mobile devices. Traditional single-step display mapping processes are complex and energy-intensive, making it difficult to achieve efficient mapping between different display devices.
A multi-step display mapping method is adopted. First, the input image is mapped to the base image of the intermediate dynamic range. Then, it is re-mapped according to the characteristics of the target display device. Combined with the reconstructed metadata, the final mapping curve is generated to match the dynamic range of the target display device.
It simplifies the display mapping process, reduces computational complexity and power consumption, and improves mapping efficiency and image quality consistency across different display devices.
Smart Images

Figure 0007775457000023 
Figure 0007775457000024 
Figure 0007775457000025
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 249,183, filed September 28, 2021, European Patent Application No. 21210178.6, filed November 24, 2021, and U.S. Provisional Patent Application No. 63 / 316,099, filed March 3, 2022, each of which is incorporated herein by reference in its entirety.
[0002] [Technical field] The present invention relates generally to images, and more particularly to dynamic range conversion and display mapping of high dynamic range (HDR) images. [Background technology]
[0003] As used herein, the term "dynamic range (DR)" may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) within an image, e.g., from darkest gray (black) to brightest white (highlight). In this scenario, DR relates to "scene-referred" intensity. DR may also relate to the ability of a display device to properly or approximately render an intensity range of a particular width. In this scenario, DR relates to "display-referred" intensity. At any point in the description herein, unless it is explicitly specified that a particular scene has particular importance, it should be presumed that the terms may be used synonymously for either scene, e.g.,
[0004] As used herein, the term "high dynamic range (HDR)" refers to a DR width that spans approximately 14-15 times or greater than the magnitude of the human visual system (HVS). In fact, DR, where humans can simultaneously perceive a wide range of intensities, may be omitted in some manner in connection with HDR. As used herein, the terms "enhanced dynamic range (EDR)" or "visual dynamic range (VDR)," individually or synonymously, refer to the DR perceivable within a scene or image by the human visual system (HVS), including eye movements, allowing for any light adaptation to change across the scene or image.
[0005] In practice, an image comprises one or more color components (e.g., luma Y and chroma Cb and Cr), each represented with n bits of precision per pixel (e.g., n=8). For example, using gamma luminance coding, an image with n≦8 (e.g., a color 24-bit JPEG image) may be considered a standard dynamic range image, while an image with n≧10 may be considered an extended dynamic range image. EDR and HDR images may be stored and distributed using high-definition (e.g., 16-bit) floating-point formats such as the OpenEXR file format developed by Industrial Light and Magic.
[0006] As used herein, the term "metadata" refers to any auxiliary information transmitted as part of a coded bitstream that assists a decoder in rendering a decoded image. Such metadata may include, but is not limited to, minimum, average, and maximum luminance within an image, color space or gamut information, reference display parameters, and auxiliary signal parameters, as described herein.
[0007] Most consumer desktop displays currently have a brightness of 200-300 cd / m 2 or nits (cd / m). Most consumer HDTVs are in the 300-500 nits range, with newer models reaching 1000 nits (cd / m 2 ). Such conventional displays exhibit lower dynamic range (LDR), also known as standard dynamic range (SDR), as opposed to HDR or EDR. As the availability of HDR content increases due to advances in both capture devices (e.g., cameras) and HDR displays (e.g., Dolby Laboratories' PRM-4200 Professional Reference Monitor), HDR content may be color graded and displayed on HDR displays that support a higher dynamic range (e.g., 1000 nits to 5000 nits, or higher). In general, but not by way of limitation, the methods of this disclosure relate to higher dynamic ranges than SDR.
[0008] As used herein, the term "display management" refers to processing performed at a receiver to render an image for a target display. For example, such processing may include, but is not limited to, tone mapping, gamut mapping, color management, frame rate conversion, etc.
[0009] High Dynamic Range (HDR) content generation and playback is now widespread, as HDR technology provides more realistic and lifelike images than previous formats. However, HDR playback can be constrained by backward compatibility requirements or limited computing power. As recognized by the inventors, improved techniques for display management of images and videos on HDR displays are being developed to improve upon existing display formats.
[0010] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise noted, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, problems identified with one or more approaches should not be assumed to have been recognized in any prior art under this section unless specifically indicated. [Brief explanation of the drawings]
[0011] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the accompanying figures, in which like reference symbols represent similar elements and in which:
[0012] [Figure 1] 1 illustrates an exemplary process for a video delivery pipeline.
[0013] [Figure 2A] 1 illustrates an exemplary process for multi-stage display mapping according to an embodiment of the present invention.
[0014] [Figure 2B] 1 illustrates an exemplary process for generating a bitstream that supports multi-stage display mapping according to an embodiment of the present invention.
[0015] [Figure 3A] 10 illustrates an example of a tone mapping curve for generating reconstruction metadata in multi-stage display mapping according to an embodiment of the present invention. [Figure 3B] 10 illustrates an example of a tone mapping curve for generating reconstruction metadata in multi-stage display mapping according to an embodiment of the present invention. [Figure 3C] 10 illustrates an example of a tone mapping curve for generating reconstruction metadata in multi-stage display mapping according to an embodiment of the present invention. [Figure 3D]10 illustrates an example of a tone mapping curve for generating reconstruction metadata in multi-stage display mapping according to an embodiment of the present invention.
[0016] [Figure 4] 1 illustrates an exemplary process for metadata reconstruction in accordance with an exemplary embodiment of the present invention.
[0017] [Figure 5A] 10 shows an example of tone mapping without "up-mapping" according to one embodiment. [Figure 5B] 10 shows an example of tone mapping after using "upmapping" according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0018] This specification describes a method for multi-step dynamic range conversion and display management for HDR images and videos. Throughout the following detailed description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without some of these specific details. In other instances, well-known structures and devices have not been described in exhaustive detail in order to avoid occluding, obscuring, or obscuring the present invention.
[0019] <Summary> The exemplary embodiments described herein relate to a method for multi-step dynamic range conversion and display management of an image on an HDR display. In one embodiment, a processor includes: receiving input metadata (204) for an input image of a first dynamic range; accessing a base layer image (212) of a second dynamic range, the base layer image being generated based on the input image; accessing base layer parameters (208) that determine the second dynamic range; accessing display parameters (230) for a target display having a target dynamic range; generating reconstruction metadata based on the input metadata, the base layer parameters, and the display parameters; generating an output mapping curve based on the reconstruction metadata and the display parameters; The step of , The output mapping curve is: Mapping the base layer image to the target display It is for , The output mapping curve is used to map the base layer image to the target display in the target dynamic range.
[0020] In a second embodiment, the processor receiving an input image (202) of a first dynamic range; accessing input metadata (204) of the input image; accessing base layer parameters (208) that determine a second dynamic range; generating a base layer image of the second dynamic range based on the input image, the base layer parameters, and the input metadata; accessing display parameters (240) for a target display having a target dynamic range; generating reconstruction metadata based on the input metadata, the base layer parameters, and the display parameters; An output bitstream is generated that includes the base layer image and the reconstruction metadata.
[0021] Multi-step image mapping and display management Video Coding Pipeline 1 illustrates an exemplary process for a conventional video distribution pipeline 100, showing various stages from video capture to video content display. A sequence of video frames (102) is captured or generated using an image generation block (105). The video frames (102) may be captured digitally (e.g., by a digital camera) or computer-generated (e.g., using computer animation) to provide video data (107). Alternatively, the video frames 102 may be captured on film using a film camera. The film is converted to a digital format to provide the video data 107. In a production stage 110, the video data 107 is edited to provide a video production stream 112.
[0022] The video data in the production stream (112) is then provided to a processor for post-production editing in block 115. The post-production editing in block (115) may include adjusting or changing the color or brightness of specific areas of the image to improve image quality or achieve a particular look according to the video creator's production intent. This is sometimes referred to as "color timing" or "color grading." Other edits (e.g., scene selection and sequencing, image cropping, adding computer-generated visual special effects, jitter, etc.) may be performed in block (115) to generate a final version of the production (117) for distribution. During post-production editing (115), the video image is displayed on a reference display (125).
[0023] Following post-production (115), the final production video data (117) may be distributed to a coding block (120) for downstream distribution to decoding and playback devices such as television sets, set-top boxes, movie theaters, etc. In some embodiments, the coding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-Ray, and other distribution formats, to generate a coded bitstream (122). At the receiver, the coded bitstream (122) is decoded by a decoding unit (130) to generate a decoded signal (132) that represents the same as or a close approximation of the signal (117). The receiver may be attached to a target display (140), which may have characteristics quite different from the reference display (125). In that case, a display management block (135) may be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137). Non-limiting examples of display management processes are described in references [1] and [2].
[0024] Single-step and multi-step display mapping In traditional display mapping (DM), the mapping algorithm applies a sigmoid-like function (see, e.g., References [3] and [4]) to map the input dynamic range to the dynamic range of the target display. Such mapping functions can be expressed as piecewise linear or nonlinear polynomials characterized by anchor points, pivots, and other polynomial parameters generated using the characteristics of the input source and the target display. For example, in References [3-4], the mapping function uses anchor points based on the luminance characteristics of the input image and the display (e.g., minimum, mean (average), and maximum luminance). However, other mapping functions can use different statistics, such as luminance variance or luminance standard deviation at the block level or across the entire image. For SDR images, this process can also be aided by additional metadata, either transmitted as part of the transmitted video or calculated by the decoder or display. For example, if a content provider has both SDR and HDR versions of source content, the source can use both versions to generate metadata (such as piecewise linear approximations of forward or reverse reshape functions) to assist the decoder in converting the received SDR image into an HDR image.
[0025] In a typical workflow for HDR data transmission, such as DolbyVision®, display mapping (135) can be considered a single-step process performed at the end of the processing pipeline before the image is displayed on the target display (140). However, it may be necessary or beneficial to perform this mapping in two (or more) processing steps. For example, a DolbyVision (or other HDR format) transmission profile may use a base layer of video coded in HDR10 at 1000 nits to support television sets that do not support DolbyVision but do support the HDR10 format. In that case, a typical workflow process would include the following steps: 1) Map the input image or video from the original HDR master to a base layer (e.g., 1000 nits, ITU-R Rec. 2020) using DolbyVision or other formats. 2) Compute static or dynamic composer metadata that reconstructs the original HDR master image from the mapped base layers. 3) Encode the mapped base layer, embed the original HDR metadata (e.g., minimum, midpoint, and maximum luminance values), and send it downstream to the decoding device along with the composer metadata. 4) At playback time, decode the coded bitstream and a) apply the composer metadata to the base layer to reconstruct the original HDR image from the base layer, and b) use the original HDR metadata to map the reconstructed image to the target display (same as single-step mapping).
[0026] This workflow has the disadvantage that it requires two image processing operations during playback: a) compositing (or prediction) to reconstruct the HDR input, and b) display mapping to map the HDR input to the destination display. In some devices, it may be desirable to perform only a single mapping operation by bypassing the composer, which requires less power consumption and / or simplifies implementation and processing complexity. In an exemplary embodiment, an alternative multi-stage workflow is described that bypasses the composer, allowing for a first mapping to a base layer, followed by a second mapping from the base layer directly to the destination display. This approach can be further extended to include subsequent steps of mapping to additional displays or bitstreams.
[0027] 2A illustrates an exemplary process for multi-stage display mapping. The dotted line and display mapping (DM) unit 205 indicate a conventional single-stage mapping. In this example, without limitation, an input image (202) and its metadata (204) need to be mapped to a target display (225) at 300 nits and P3 color gamut. The characteristics (230) of the target display (e.g., minimum and maximum luminance and color gamut) along with the input (202) and its metadata (e.g., minimum luminance, midpoint luminance, maximum luminance) (204) are provided to the display mapping (DM) process (205), which maps the input to the dynamic range of the target display (225).
[0028] Solid lines and shaded blocks indicate multi-stage mapping. The input image (202), input metadata (204), and parameters related to the base layer (208) are fed to a display mapping unit (210), which generates a mapped base layer (212) (e.g., from the input dynamic range to 1000 nits in Rec. 2020). This step can be performed in the encoder (not shown). During playback, a new processing block, the metadata reconstruction unit (215), uses the target display parameters (230), base layer parameters (208), and input image metadata (204) to adjust the input image metadata to generate reconstructed metadata (217) such that the subsequent mapping (220) of the mapped base layer (212) to the target display (225) is visually identical to the result of a single-stage mapping (205) to the same display.
[0029] For existing (legacy) content containing the base layer and original HDR metadata, a metadata reconstruction block (215) is applied during playback. In some cases, the base layer target information (208) is not available and can be estimated based on other information (e.g., in DolbyVision, using profile information (Profile 8.4, 8.1, etc.)). It is also possible that the mapped base layer (212) is identical to the original HDR master (e.g., 202), in which case metadata reconstruction can be skipped.
[0030] In some embodiments, metadata reconstruction (215) can be applied on the encoder side. For example, due to limited power or computational resources in mobile devices (e.g., phones, tablets, etc.), it may be desirable to pre-compute the reconstructed metadata to save power in the decoder device. This new metadata can be transmitted in addition to the original HDR metadata, in which case the decoder can simply use the reconstructed metadata and skip the reconstruction block. Alternatively, the reconstructed metadata can replace part of the original HDR metadata.
[0031] 2B illustrates an example of a process for reconstructing metadata within an encoder to prepare a bitstream suitable for multi-step display mapping. Given that the encoder is unlikely to know the characteristics of the target display, metadata reconstruction can be applied based on the characteristics of multiple potential displays, such as 100 nits, Rec. 709 (240-1), 400 nits, P3 (240-2), and 600 nits, P3 (240-3). The base layer (212) is constructed as described above, but now the metadata reconstruction process considers multiple target displays to achieve accurate matches for a wide variety of displays. The final output (250) combines the base layer (212), the reconstructed metadata (217), and portions of the original metadata (204) unaffected by the metadata reconstruction process.
[0032] Metadata Restructuring During metadata reconstruction, a portion of the original input metadata (for the input image in the input dynamic range) is combined with information about the characteristics of the base layer (available in the intermediate dynamic range) and the target display (for displaying the image in the target dynamic range) to generate reconstructed metadata for the two-stage (or multi-stage) display mapping. In one embodiment, the metadata reconstruction is performed in four steps:
[0033] Step 1: Single-step mapping As used herein, the term "L1 metadata" refers to the minimum, mean, and maximum luminance values associated with an input frame or image. L1 metadata can be calculated by converting RGB data to luma-chroma format (e.g., YCbCr) and calculating the minimum, mean (average), and maximum values in the Y plane, or it can be calculated directly in RGB space. For example, in one embodiment, L1Min refers to the minimum of the PQ-encoded min(RGB) values of the image while taking into account active areas (e.g., by excluding gray or black bars, letterbox bars, etc.). min(RGB) refers to the minimum of the color component values {R, G, B} of a pixel. The values of L1Mid and L1Max can also be calculated in the same manner by replacing the min() function with the average() and max() functions. For example, L1Mid refers to the average of the PQ-encoded max(RGB) values of the image, and L1Max refers to the maximum of the PQ-encoded max(RGB) values of the image. In some embodiments, L1 metadata can be normalized to [0, 1].
[0034] Consider the L1Min, L1Mid, and L1Max values of the original HDR metadata, as well as the maximum (peak) and minimum (black) luminance of the target display, denoted as Tmax and Tmin. Then, as described in References [3-4], a mapping curve for intensity tone mapping can be generated that maps the intensities of the input image to the dynamic range of the target display. An example of such a curve (305) is shown in Figure 3A. This can be thought of as an ideal single-stage tone mapping curve that is matched using the reconstruction metadata. This direct tone mapping curve is used to map the L1Min, L1Mid, and L1Max values to the corresponding TMin, TMid, and TMax values. In Figures 3A-3D, all input and output values are presented in the PQ domain using SMPTE ST2084. All other calculated metadata values (e.g., BLMin, BLMid, BLMax, TMin, TMid, TMax, TMin', TMid', TMax') are also included in the PQ domain.
[0035] Step 2: Mapping to the base layer Consider as input the L1Min, L1Mid, and L1Max values of the original HDR metadata and the Bmin and Bmax values of the base layer parameters (208), which indicate the black level (minimum luminance) and peak luminance of the base layer stream. Again, a first intensity mapping curve can be derived to map the input data to values in the Bmin and Bmax range. An example of such a curve (310) is shown in Figure 3B. Using this curve, the original L1 values can be mapped to BLMin, BLMid, and BLMax values to be used as the reconstructed L1 metadata in the third step.
[0036] Step 3: Mapping base layers to goals Take BLMin, BLMid, and BLMax from step 2 as updated L1 metadata and map them to the target display (e.g., Tmin and Tmax) using a second display management curve. Using the second curve, the corresponding mapped values of BLMin, BLMid, and BLMax are denoted as TMin', TMid', and TMax'. In Figure 3C, curve (315) shows an example of this mapping. Curve (305) represents a single-stage mapping. The goal is to make the two curves match.
[0037] Step 4: Matching single-step and multi-step mappings As used herein, the term "trim" refers to tone curve adjustments performed by a colorist to improve a tone mapping operation. Trims are typically applied to the SDR range (e.g., 100 nit maximum luminance, 0.005 nit minimum luminance). These values are linearly interpolated to the target luminance range depending only on the maximum luminance. These values modify the default tone curve and are present in all trims.
[0038] Information about the trim may be part of the HDR metadata and may be used to adjust the tone mapping curve generated in steps 1-2 (see references [1-4] and equations (4-8) below). For example, in DolbyVision, the trim can be passed as Level 2 (L2) or Level 8 (L8) metadata, which includes Slope, Offset, and Power variables (collectively called SOP parameters) that represent gain and gamma values for adjusting pixel values. For example, if Slope, Offset, and Power are [-0.5, 0.5], then for a given gain and gamma,
number
[0039] In one embodiment, it may be necessary to use the trim-related reconstruction metadata to match the two mapping curves. To match [TMin', TMid', and TMax'] from step 3 with [TMin, TMid, TMax] from step 1, generate values for Slope, Offset, Power, and TMidContrast, which are used as new (reconstructed) trim metadata (e.g., L8 and / or L2) in the reconstruction metadata.
[0040] Slope, Offset, and Power Calculations: The purpose of the power and TMidContrast calculations is to match [TMin', TMid', and TMax'] from step 2 with [TMin, TMid, TMax] from step 1. These are related to each other by the following formula:
number
number
number
number
[0041] Considering the tone curve y(x) generated according to the input metadata and Tmin and Tmax values (e.g., Ref. [4]), TMidContrast then updates the slope (slopeMid) of the center (see, e.g., point (L1Mid, TMid) (307) in Figure 3A) as follows:
number
number
[0042] In some embodiments, the Slope, Offset, and Power can be applied in normalized space. This has the advantage of reducing the chance of clipping when applying the Power term. In this case, before applying the Slope, Offset, and Power, normalization is performed as follows:
number
number
[0043] TmaxPQ and TminPQ denote PQ-coded luminance values corresponding to the linear luminance values Tmax and Tmin converted to PQ luminance using SMPTE ST2084. In one embodiment, TmaxPQ and TminPQ are in the range of [0, 1], which is expressed as [0 to 4095] / 4095. In this case, normalization of [TMin, TMid, TMax] and [TMin', TMid', TMax'] is performed before step 1, which calculates the slope, offset, and power. Then, in step 3, TMidContrast (see equation (3)) is scaled by (TmaxPQ - TminPQ) as follows:
number
[0044] Figure 4 shows an example process and the steps described above that summarizes the metadata reconstruction process (215) according to one embodiment. As shown in Figure 4, the inputs to the process are input metadata (204), base layer characteristics (208), and target display characteristics (230). Step 405 uses the input metadata and the target display characteristics (e.g., Tmin, Tmax) to generate a direct or single-step mapping tone curve (e.g., 305). This direct mapping curve is used to convert the input luminance metadata (e.g., L1Min, L1Mid, L1Max) to direct-mapped metadata (e.g., TMin, TMid, TMax). Step 410 uses the input metadata and base layer characteristics (e.g., Bmin and Bmax) to generate a first intermediate mapping curve (e.g., 310), which is used to generate a first set of reconstructed luminance metadata (e.g., BLMin, BLMid, BLMax) that correspond to luminance values in the input metadata (e.g., L1Min, L1Mid, L1Max). Step 415 generates a second mapping curve that maps the input with BLMin, BLMid, and BLMax values to the target display (e.g., using Tmin and Tmax). The second tone mapping curve (e.g., 315) can be used to map the first set of reconstruction metadata values (e.g., BLMin, BLMid, BLMax) generated in step 410 to mapped reconstruction metadata values (e.g., TMin', TMid', TMax'). Step 420 generates some additional reconstruction metadata (e.g., SOP parameters Slope, Offset, and Power) that are used to adjust the second tone mapping curve. This step requires solving at least three simultaneous equations with three unknowns: Slope, Offset, and Power, using the direct mapped metadata values (TMin, TMid, TMax) and the corresponding mapped reconstruction metadata values (TMin′, TMid′, and TMax′). Step 425 uses the SOP parameters, the direct mapping curve, and the second mapping curve to generate a slope adjustment parameter (TMidContrast) for further adjusting the second mapping curve. The output reconstruction metadata (212) includes reconstruction luminance metadata (e.g., BLMin, BLMid, BLMax) and reconstruction or new trim path metadata (e.g., TMidContrast, Slope, Power, Offset), which can be used by a decoder to adjust the second mapping curve and generate an output mapping curve for mapping the base layer image to a target display.
[0045] Returning to FIG. 2A, the display mapping process 220: a) generating a tone mapping curve (y(x)) that maps the intensity of the base layer with the reconstructed metadata values BLMin, BLMid, and BLMax to Tmin and Tmax values of the target display (225); b) As described above, use the trim path metadata (e.g., TMidContrast, Slope, Power, Offset) to update this tone mapping curve (e.g., see equation (4-8)).
[0046] In one embodiment, the tone curve can be generated using sampling points different from L1Min, L1Mid, and L1Max. For example, because only a small number of luminance range points are sampled, selecting curve points closer to the center may improve the overall curve match. In another embodiment, the entire curve can be considered during optimization, rather than just three points. Furthermore, if the difference between TMid and TMid' is very small, improvements can be made by allowing a solution with a smaller tolerance for accuracy. For example, rather than resolving the points exactly, allowing a smaller tolerance between points (e.g., 1 / 720) may result in less trimming and an overall better curve match.
[0047] As described in step 1, the tone map intensity curve is the display management tone curve. It is recommended that this curve be as close as possible to the curve used both in generating the base layer and in the target display. Therefore, the version or design of the curve may vary depending on the content or the type of playback device. For example, a curve generated according to [4] may not be supported by older legacy devices that only know how to construct curves according to [3]. Since not all DM curves are supported by all playback devices, the curve used when calculating the tone map intensity should be selected based on the content type and characteristics of the specific playback device. If the exact playback device is unknown (such as when metadata reconstruction is applied in encoding), the closest curve will be selected, but the resulting image may deviate significantly from Single Step Mapping.
[0048] Metadata adjustments to global dimming metadata As used herein, the term "L4 metadata" or "Level 4 metadata" refers to signal metadata that can be used to adjust global dimming parameters. In one embodiment of DolbyVision processing, without limitation, the L4 metadata includes two parameters: FilteredFrameMean and FilteredFramePower, defined as follows:
[0049] FilteredFrameMean (or mean_max for short) is calculated as the temporally filtered mean power of the frame maximum luminance values (e.g., the PQ-encoded maximum RGB values of each frame). In one embodiment, this temporal filtering is reset at scene cuts if such information is available. FilteredFramePower (or std_max for short) is calculated as the temporally filtered standard deviation power of the frame maximum luminance values (e.g., the PQ-encoded maximum RGB values of each frame). Both values can be normalized to 0001. These values represent the mean and standard deviation of the maximum luminance of the image sequence over time and are used to adjust global dimming during display. To improve display output, it is desirable to identify a mapping reconfiguration for the L4 metadata as well.
[0050] In one embodiment, the mapping of std_max values follows a model characterized by the following:
number
[0051] In one embodiment, when Smax=Dmax (e.g., y=1), the standard deviation value should remain the same, so z=x. By substituting these values into equation (9), d=1-b and a=-c are derived, and equation (9) can be rewritten as follows:
number
[0052] In one embodiment, the parameters a and b in equation (10) were derived by applying display mapping to 260 images ranging from a maximum luminance of 4000 nits to 1000, 245, and 100 nits. This mapping provided 780 data points (for Smax, Dmax, std_max) that were curve-fitted to obtain the output model parameters.
number
[0053] Using single-point approximations for a and b, equation (10) can be rewritten as:
number
[0054] Equation (11) expresses a simple relationship regarding how to map L4 metadata, specifically the std_max value. Beyond the mapping described by equations (10) and (11), the properties of equation (11) can be generalized as follows: L4 metadata remapping is linearly proportional, e.g. an image with a higher original std_max value will be remapped to an image with a higher remapped map_std_max value. The ratio of Smax / Dmax reduces the map_std_max value, but at a very slow pace. Thus, an image with a high original std_max value will still be remapped to an image with a relatively high remapped map_std_max value. For example, if Smax / Dmax=1.6, then map_std_max=0.7std_max. If Smax / Dmax=1, no remapping is performed.
[0055] Remapping when Tmax>Smax Let Smax denote the maximum luminance of the reference display. During the direct mapping in step 1, cases where Tmax > Smax are allowed, i.e., the target display can have a higher luminance than the reference display, typically applying a direct one-to-one mapping and no metadata adjustments. Such a one-to-one mapping is shown in FIG. 5A. In one embodiment, a special "up-mapping" step can be used to improve the appearance of the displayed image by allowing mapping of image data up to the Tmax value. This up-mapping step can also be guided by the input trim (L8) metadata.
[0056] In one embodiment, the upmapping is performed as part of step 1 above. For example, consider the case where Smax = 2000 nits and Tmax = 9000 nits. Consider a base layer (Bmax) of 600 nits. Assuming no trim to guide the upmapping, Figure 5B shows an example of upmapping where input (X) PQ values [0.0151, 0.3345, 0.8274] are mapped to output (Y) PQ values [0.0151, 0.3507, 0.9889], with X = Y = 1 corresponding to 10000 nits. Input X = 0.8274 corresponds to Smax = 2000 nits and maps to Y = 0.9889, which corresponds to 9000 nits. Similarly, X = Smid = 0.3345 maps to Tmid = 0.3507, which represents approximately a 5% increase in the original Smid value, and X = 0.0151 maps to Y = 0.0151 using a direct one-to-one mapping. Therefore, in the absence of additional metadata or guide information, when Tmax>Smax, the following anchor points can be used to construct the tone mapping curve: Map Smin (the minimum luminance of the source display) to Tmin. Maps Smid (the estimated average luminance of the source display) to Tmid=Smid+c*Smid, where c is in the range [0, 0.1]. Map Smax to Tmax.
[0057] In another embodiment, if the original metadata includes trims (e.g., L8 metadata) specified for a target display that has a maximum luminance greater than the Smax value, then the upmapping is guided by these trim metadata. For example, consider an Xref[i] luminance point for which a Yref[i] trim is defined. For example,
number
number
[0058] For example, if the trim target is 3000 nits, consider an incoming video source with the following L8 trim: Slope=0.1, Offset=-0.07, Power=0.03 If Smax=2.000 nits, the above trim can be linearly extrapolated to obtain a trim with a target of 9000 nits. Trim extrapolation is done for all trims in L8. The extrapolated trim can be used as part of the direct mapping step in step 1. For example, for the Slope trim value,
number
number
[0059] <Exemplary computer system implementation> Embodiments of the present invention may be implemented by a computer system, a system configured in electronic circuits and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA) or other configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or perform instructions related to image transformations as described herein. The computer and / or IC may calculate any of the various parameters or values related to the multi-step display mapping process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0060] A specific implementation of the present invention includes a computer processor executing software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. may implement the methods associated with the multi-step display mapping described above by executing software instructions in a program memory accessible to the processor. The present invention may be provided in the form of a program product. The program product may include any tangible, non-transitory medium that carries a set of computer-readable signals containing instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. A program product according to the present invention may be in any of a variety of tangible forms. The program product may include physical media such as, for example, magnetic data storage media including floppy disks, hard disk drives, optical data storage media including CD-ROMs, DVDs, ROMs, electronic data storage media including flash RAM, etc. The computer-readable signals on the program product may be optically compressed or encrypted.
[0061] Although components (e.g., software modules, processors, components, devices, circuits, etc.) have been referred to above, unless otherwise indicated, references to those components (including references to "means") should be interpreted to include equivalents of those components, any components that perform the functions of the described components (e.g., are functionally equivalent), and components that are not structurally equivalent to the disclosed structures that perform the functions in the illustrated exemplary embodiments of the present invention.
[0062] <Equivalents, Extensions, Alternatives and Miscellaneous> Exemplary embodiments relating to multi-stage display mapping are described. In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indication of what the invention is, and what the applicant intends to be the invention, is set forth in the claims as issued in particular form hereby, including any subsequent amendments. Any definitions expressly set forth herein for terms contained in such claims shall control the meaning of such terms as used in the claims. Accordingly, any limitation, element, feature, advantage, or attribute not expressly recited in the claims should not in any way limit the scope of the claims. The specification and drawings are, therefore, to be considered in an illustrative, rather than a restrictive, sense.
Claims
1. 1. A method for multi-step display mapping, comprising: accessing input metadata for an input image of a first dynamic range; accessing a base layer image of a second dynamic range, the base layer image having been generated based on the input image; accessing base layer parameters that define the second dynamic range; accessing display parameters for a target display having a target dynamic range; generating reconstruction metadata based on the input metadata, the base layer parameters, and the display parameters; generating an output mapping curve based on the reconstruction metadata and the display parameters, the output mapping curve for mapping the base layer image to the target display; mapping the base layer image to the target display in the target dynamic range using the output mapping curve; A method comprising:
2. 1. A method of multi-step display mapping, comprising: accessing an input image in a first dynamic range; accessing input metadata of the input image; accessing base layer parameters that define a second dynamic range; generating a base layer image of the second dynamic range based on the input image, the base layer parameters, and the input metadata; accessing display parameters for a target display having a target dynamic range; generating reconstruction metadata based on the input metadata, the base layer parameters, and the display parameters; generating an output bitstream including the base layer image and the reconstruction metadata; A method comprising:
3. receiving, at a decoder, the base layer image and the reconstruction metadata; generating an output mapping curve based on the reconstruction metadata and the display parameters, the output mapping curve for mapping the base layer image to the target display; mapping the base layer image to the target display in the target dynamic range using the output mapping curve; The method of claim 1 further comprising:
4. The method of claim 1 , wherein the base layer image has a maximum dynamic range of 1000 nits.
5. The method of claim 1 , wherein the display parameters include a minimum luminance value (Tmin) and a maximum luminance value (Tmax) of the target display.
6. The method of claim 1 , wherein the base layer parameters include a minimum luminance value (Bmin) and a maximum luminance value (Bmax) in the base layer image.
7. The method of claim 1 , wherein the reconstructed metadata includes reconstructed L1 metadata, the reconstructed L1 metadata including a reconstructed minimum (BLMin), a reconstructed mean (BLMid), and a reconstructed maximum (BLMax).
8. The method of claim 7 , wherein the reconstruction metadata further includes Slope, Power, and Offset values.
9. The step of generating reconstruction metadata includes: generating a direct mapping curve that maps the input image to the target dynamic range based on the input metadata and the display parameters; applying the direct mapping curve to luminance values in the input metadata to generate mapped luminance metadata; generating a first mapping curve that maps the input image to the base layer image based on the input metadata and the base layer parameters; mapping luminance values in the input metadata to a first set of reconstructed metadata using the first mapping curve; generating a second mapping curve that maps the base layer image to the target dynamic range based on the first set of reconstruction metadata and the display parameters; mapping the first set of reconstructed metadata to mapped reconstructed metadata using the second mapping curve; generating a second set of reconstruction metadata including Slope, Power, and Offset values based on the mapped luminance metadata and the mapped reconstruction metadata to adjust the second mapping curve; The method of claim 1 , comprising:
10. 10. The method of claim 9, further comprising generating a slope adjustment value for adjusting the second mapping curve based on the direct mapping curve, the second mapping curve, and the Slope, Power, and Offset values.
11. The Slope, Power, and Offset values are as follows: [Equation 1] where N≧3, TM(i) denotes mapped intensity metadata, and TM′(i) denotes mapped reconstruction metadata.
12. 12. The method of claim 11 , wherein the TM(i) values include minimum (TMin), mean (TMid), and maximum (TMax) luminance values corresponding to values mapped using the direct mapping curve of minimum, mean, and maximum luminance values in the input image.
13. generating the direct mapping curve when Tmax is greater than Smax, wherein Tmax represents a maximum luminance value of the target display and Smax represents a maximum luminance value of a reference display; If the input metadata does not include trim metadata, mapping a minimum luminance Smin of a reference display to a minimum luminance Tmin of the target display; mapping Smid, the average luminance of the reference display, to Tmid=Smid+c*Smid, where c is between 0 and 0.2 and Tmid represents the average luminance of the target display; mapping Smax to Tmax; Including, In other cases, Given the Xref[x1, x2] luminance points and the corresponding trim metadata Yref[y1, y2] values, the following equation: [Equation 2] generating an extrapolated trimmed Yout value for luminance point Xin by calculating 10. The method of claim 9, comprising:
14. The input metadata includes global dimming metadata, and given an input global dimming metadata value x, generating a reconstructed dimming metadata value z is performed according to the following formula: [Equation 3] 2. The method of claim 1, further comprising the step of calculating: a = y a b y ... [Request Item 15] [Number 4] 15. The method of claim 14, wherein, for an input video sequence containing the input image, x denotes a time-varying mean or standard deviation of maximum luminance values in the input video sequence.
16. An apparatus including a processor and configured to perform the method of any one of claims 1 to 15.
17. A non-transitory computer-readable storage medium storing computer-executable instructions for carrying out the method by one or more processors according to the method of any one of claims 1 to 15.
Citation Information
Patent Citations
Image encoding device, image decoding device, image processing device, and control method for the same
JP2015035798A
Efficient End-to-End Single-Layer Inverse Display Management Coding
JP2020524446A
Graphics-safe HDR image luminance re-grading
JP2020532805A
Real-time reshaping of single-layer backwards-compatible codec
US20190222866A1