Content-optimized ambient light HDR video adaptation
The method and apparatus adapt HDR video by converting it to SDR or lower dynamic range using luminance mapping functions, addressing the challenge of displaying HDR content on displays with different peak luminance capabilities and ensuring consistent quality across varying ambient lighting.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to optimally display High Dynamic Range (HDR) video on displays with varying peak luminance capabilities, as they do not adequately account for ambient lighting conditions and the need to adapt image brightness to match the display's maximum and minimum luminance levels.
A method and apparatus that adapt HDR video by using luminance mapping functions to convert HDR images to Standard Dynamic Range (SDR) or lower dynamic range images, utilizing metadata to communicate the necessary luminance adjustments, allowing displays to optimally render HDR content on devices with different peak luminance capabilities.
Enables effective display of HDR content on a wide range of displays by ensuring that image brightness is adjusted to match the display's capabilities, providing a consistent viewing experience across varying ambient lighting conditions.
Smart Images

Figure 0007843778000001 
Figure 0007843778000002 
Figure 0007843778000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and apparatus for adapting the image pixel brightness of high dynamic range video to produce a desirable appearance when displaying HDR video under ambient light conditions at a specific viewing location. [Background technology]
[0002] Several years ago, novel methods for high dynamic range (HDR) video coding were introduced, particularly by the present applicant (see, for example, WO2017157977).
[0003] Video coding generally involves either color codes (e.g., lumens and two chromas per pixel) to represent an image, or simply creating, or more precisely, defining, color codes. This is somewhat different from knowing how to optimally display an HDR image (for example, in the simplest way, one could simply use a highly nonlinear optical-optical transfer function OETF to convert the desired luminance to, say, a 10-bit lumen code, and vice versa, using an inverse form of the electrical-optical transfer function EOTF to convert these video pixel lumen codes to display luminances, mapping 10-bit electrical lumen codes to display optical pixel luminances, but more complex systems can deviate in several directions by decoupling image coding, especially from the specific use of coded images).
[0004] The coding and processing of HDR video are in stark contrast to the use of conventional video technologies. According to conventional video technologies, until recently, all video was encoded, which is now called Standard Dynamic Range (SDR) video coding (also known as Low Dynamic Range video coding, LDR). This SDR began in the analog era as PAL or NTSC, and in the digital video era, it transitioned to Rec.709-based coding such as MPEG2 compression.
[0005] While the technology was sufficient for transmitting video in the 20th century, advancements in display technology that surpassed the physical limitations of 20th-century CRT electron beams—namely, global TL-backlit LCDs—made it possible to display images with pixels that were considerably brighter (potentially darker) than conventional displays, which necessitated the ability to encode and create such HDR images.
[0006] In fact, for various reasons, the SDR standard (8-bit Rec.709) couldn't encode much brighter, and in some cases darker, image objects. A method was first invented to technically represent these increased brightness ranges of color, starting with these objects. From there, every rule of video technology had to be rethought, and often reinvented.
[0007] The Rec.709 SDR luma code definition, due to its luma having an OETF function shape that is almost the square root, could only encode a luminance dynamic range of approximately 1000:1 (with 8-bit or 10-bit luma): Y_code=power(2,N) * sqrt(L_norm), where N is the number of bits in the luma channel and L_norm is the normalized version of the physical luminance between 0 and 1.
[0008] Furthermore, in the SDR era, absolute display brightness was not defined, so in practice, maximum relative brightness L_norm_max=100% or 1 was mapped to maximum normalized luma code Yn=1 (e.g., Y_code_max=255) via the square root OETF. This has some technical differences compared to creating an absolute HDR image (i.e., an image pixel coded to be displayed as 200 nits would ideally (i.e., where possible) be displayed as 200 nits on all displays and not as a significantly different display brightness). In the relative paradigm, a coded pixel brightness of 200 nits would be displayed as 300 nits on a brighter display, i.e., a display with a brighter maximum displayable brightness PL_D (aka display maximum brightness), and as 100 nits on a less capable display. Note that absolute encoding can also work with normalized brightness representation and normalized 3D color gamut, in which case 1.0 would uniquely mean, for example, 1000 nits.
[0009] On displays, such relative images have typically been displayed somewhat heuristically by mapping the brightest brightness of the video to the brightest displayable pixel brightness (this is done automatically via the electrical drive of the display panel by maximum luminance Y_code_max, without further brightness mapping). Therefore, if you buy a 200-nit PL_D display, white will appear twice as bright as on a 100-nit PL_D display, but factors such as eye adaptation were considered less important, other than giving a brighter, more enjoyable, and somewhat more beautiful version of the same SDR video image.
[0010] Traditionally, and today, when discussing SDR video images (within an absolute framework), there is usually a video peak luminance of PL_V = 100 nits (e.g., agreed upon according to standards). Therefore, in this application, we assume that the maximum luminance (or SDR grading) of an SDR image is exactly that, or is generalized to approximately that value.
[0011] In this application, "grading" is intended to mean either an activity or a resulting image in which pixels are given brightness as needed, for example, by a human color grader or an automaton. For example, when viewing an image (for example, designing an image), there are multiple image objects, and ideally, it is desirable to give the pixels of those objects brightness that is spread around the average brightness that is optimal for that object and scene, while also considering the overall picture. For example, if an image has the capability for the brightest codeable pixel to be 1000 nits (maximum brightness PL_V of the image or video), one grader might choose to give the pixels an explosion brightness value of 800-1000 nits to make the explosion appear very powerful, while another filmmaker might choose an explosion with a brightness of 500 nits or less, for example, so as not to overshadow the rest of the image at that point (of course, both situations are technically possible).
[0012] The maximum brightness of an HDR image or video can vary considerably and is usually communicated along with the image data as metadata about the HDR video or image (common values are, for example, 1,000 nits, 4,000 nits, 10,000 nits (not limited), and typically, an HDR image can be said to exist when PL_V is at least 600 nits). If a video creator chooses to define their image as PL_V=4,000 nits, they can of course choose to create brighter explosions, but relatively they will not reach 100% of PL_V, for example, up to 50% in the case of such a high PL_V definition for a scene.
[0013] HDR displays may have maximum capabilities, i.e., the highest displayable pixel brightness, such as 600 nits, 1000 nits, or N x 1000 nits (starting with lower-end HDR displays). The maximum brightness or peak brightness PL_D of that display is different from the maximum brightness PL_V of the video, and these two should not be confused. Video creators typically cannot create optimal video for every possible end-user display (i.e., the capabilities of the end-user display are best utilized by the video, and (ideally) the maximum brightness of the video will not exceed the maximum brightness of the display, but also not be lower; i.e., some pixels in the video image should have at least some pixels where the pixel brightness L_p=PL_V=PL_D).
[0014] The creator makes several decisions (e.g., what content to capture, how to capture it) and typically creates a video with a high PL_V so as to accommodate the best PL_D displays of the intended audience today, and possibly the intended audience in the future when higher PL_D displays become available.
[0015] Next, a secondary question arises: how to optimally display an image with a peak brightness PL_V on a display with a (sometimes much) lower peak brightness PL_D, a phenomenon known as display adaptation. In the future, there will be displays that require images with a lower dynamic range than, for example, a 2000 nit PL_V image, which will be created and received over a communication medium. In theory, displays can always be regraded; that is, the brightness of image pixels can be mapped and made displayable by internal heuristics, but if the video creator takes sufficient care in determining the brightness of the pixels, it is also possible to show how to display the image to a lower PL_D value, and ideally, it may be beneficial for displays to largely adhere to these technical desiderators.
[0016] The situation is more complex with respect to the darkest displayable pixel brightness BL_D. While some of this may be fixed physical characteristics of the display, such as LCD cell leak light, even with the best displays, what a viewer can ultimately distinguish as a different darkest black also depends on the lighting in the viewing room, which is not a clearly defined value. This lighting is characterized, for example, as the average illuminance level in lux units, but for video display purposes, it is more elegantly characterized as the minimum pixel brightness. This also involves the human eye, to a stronger degree than the appearance of bright or medium brightness. This is because the human eye may see fewer pixels of high brightness, thus reducing the relevance of darker pixels, especially their absolute brightness. However, one could argue that the eye is not a limiting factor when viewing an image of a scene that is mostly dark but masked by ambient light in front of the display screen. Assuming that humans can only see a noticeable difference of 2%, there is a darkest driving level (or lumens) b, and above that, one can see the next darkest lumens level (i.e., displaying another 2% higher brightness level, or X% higher in displayed brightness).
[0017] In the LDR era, the darkest pixels were hardly a concern. The focus was primarily on the average brightness, which was about 1 / 4 of the maximum PL_V=100nit. When an image was exposed around this value, everything in the scene appeared bright and colorful, except for clipping of the brighter parts of the scene that exceeded the 100% maximum. For the darkest parts of a scene, if they were important enough, an image was created that was captured with a sufficient amount of base lighting in the recording studio or shooting environment. If parts of the scene didn't look good (for example, because they were buried in code Y=0), it was considered normal.
[0018] Therefore, if nothing further is specified, we can assume that the darkest black is zero, or actually something like 0.1 or 0.01 nits. In such a situation, engineers will place more emphasis on pixels that are brighter than the average of the coded and / or displayed HDR image.
[0019] Regarding coding, the difference between HDR and SDR is not only physical (more different pixel luminances are displayed on a display with a greater dynamic range capability), but also technical, potentially involving further technical concepts such as different lumacode assignment capabilities (those using OETF, or inversely, EOTF in an absolute sense), and additional dynamically changing metadata (per image or per set of temporally consecutive images). Metadata specifies how to regrade the pixel luminances of various image objects to obtain images with a secondary dynamic range different from the starting image dynamic range (typically, the two luminance ranges end with peak luminances that differ by at least 1.5 times).
[0020] A simple HDR codec has been introduced to the market. The HDR10 codec is used, for example, to create the recently released Black Jewel Box HDR Blu-ray. This HDR10 video codec uses a function with a logarithmic shape rather than a square root, namely the so-called Perceptual Quantizer (PQ) function standardized in SMPTE2084, as OETF (Inverse OETF). Rather than being limited to 1000:1 like Rec.709 OETF, this PQ OETF can define more (ideally visible) luminance, i.e., 1 / 10,000 nit to 10,000 nit, which is enough for the practical production of HDR video.
[0021] Readers should not confuse HDR with the large number of bits in a Luma codeword. This applies to linear systems such as the number of bits in an analog-to-digital converter, where the number of bits is indeed a logarithmic dynamic range with base 2. However, since the code assignment function can have a highly nonlinear shape, theoretically, one can define an HDR image with only 10 bits of Luma (8 bits per color component HDR image), no matter how desired. This provides the advantage of reusability in already implemented systems (e.g., ICs have a specific bit depth, or video cables have a specific bit depth, etc.).
[0022] After calculating the luma, there is a 10-bit plane of pixel luma Y_code, to which two chrominance components Cb and Cr per pixel are added as a chrominance pixel plane. This image can be handled further down the line as if it were a conventionally mathematically compressed (e.g., MPEG-HEVC compressed) SDR image. The compressor does not need to consider pixel color or luminance. However, a receiving device, e.g., a display (or actually its decoder), usually needs to perform the correct color interpretation of the {Y, Cb, Cr} pixel colors and display an image that looks correct, e.g., not an image with bleach color.
[0023] This is usually processed by communicating together further image definition metadata that defines the image coding, such as the EOTF used (assuming without limitation that the PQ EOTF (or OETF) was used), instructions such as the value of PL_V, etc., along with the three pixelated color component planes.
[0024] More advanced coders may include further image definition metadata, e.g., processing metadata, e.g., a function (explained in more detail in Figure 2) that specifies how to map a normalized version of the luminance of the first image up to PL_V = 1000 nit to the normalized luminance of a secondary reference image, e.g., an SDR reference image with PL_V = 100 nit.
[0025] For clarity, the reference image is not a single pre-fixed image. For each image of a scene (e.g., consecutive images within a video), one or more reference images can be created. These are defined within a particular dynamic range (i.e., for the aforementioned DR, particularly in the case of its maximum luminance PL_V, the various intermediate object luminances have values that depend on the range. Therefore, the reference images represent the manner in which an image is graded for various DRs, and at least two reference images show the manner in which the luminance of one of these two images is re-graded for the DR of the other image). The equivalence of images usually means the equivalence of their pixel colors. Different images with different dynamic ranges can correspond to each other in that the pixel luminance of the first image shifts to different luminances of the corresponding images with different DRs.
[0026] To deepen the understanding of readers with little knowledge about HDR, some interesting aspects in FIG. 1 will be briefly explained. FIG. 1 shows some typical examples of many possible HDR scenes that a future HDR system (e.g., connected to a display with PL_D of 1000 nits) needs to be able to process correctly. The actual technical processing of pixel colors can be performed in various color space definitions in various manners, but the desideratum for re-grading is shown as the absolute luminance mapping between luminance axes spanning different dynamic ranges.
[0027] For example, ImSCN1 is an image of a sunny outdoor scene from a western movie that is mostly bright areas. The first thing that should not be misunderstood is that the pixel luminance of any image is usually not the luminance that can actually be measured in the real world.
[0028] Even if no further human involvement is involved in creating the output HDR image (which acts as a starter image and is referred to as the master HDR grading or image), the camera, for the sake of its iris, always measures relative brightness within the image sensor, no matter how simple it may be by adjusting one parameter. Therefore, there is always some step involved in ensuring that at least the brightest image pixels fall within the available coding brightness range of the master HDR image.
[0029] For example, the specular reflection of the sun on a sheriff's star emblem measures over 100,000 nits in the real world, which cannot be displayed on typical near-future displays and would not be comfortable for viewers watching a movie in a dimly lit room at night. Instead, a video creator can determine that 5000 nits is bright enough for the pixels of the star emblem, and if this is the brightest pixel in the movie, the video creator can decide to create a video with PL_V=5000 nits. The camera is only a relative pixel luminance measuring device for the raw version of the master HDR grading, but it should also have a sufficiently high native dynamic range (full pixels far above the noise floor) to produce a good image. The pixels of a graded 5000 nit image are usually obtained non-linearly from the raw image captured by the camera. For example, a color grader considers such aspects as typical viewing conditions. These viewing conditions are not the same as the actual shooting location, i.e., standing in a hot desert. The best (highest PL_V) image selected to create for this scene ImSCN1, i.e., a 5000 nit image in this example, is the master HDR grading. This is the minimum HDR data required to create and communicate, but it is not the only data communicated by all codecs, and some codecs do not communicate the image at all.
[0030] The availability of such an encodeable high-luminance range DR_1 (e.g., 0.001nit to 5000nit) allows content creators to provide viewers with a better experience of brighter appearances, but of course, provided that viewers have a corresponding high-end PL_D=5000nit display, it also allows for dimming of night scenes (if the entire film is well-graded). A good HDR film balances the luminance of various image objects not only within a single image, but also over time in the story of a film or generally created video material (e.g., a well-composed HDR soccer program).
[0031] The leftmost vertical axis in Figure 1 shows some of the (average) object luminances you would want to see in a 5000-nit PL_V master HDR grading, ideally targeting a 5000-nit PL_D display. For example, in a movie, you might want to show a cowboy bathed in bright sunlight with a pixel luminance of about 500 nits (i.e., 10 times brighter than typical LDR; however, another creator might prefer a slightly less HDR impression, such as 300 nits). This constitutes the best mode for displaying this Western image, according to the creator, giving the end consumer the best possible appearance.
[0032] The need for a higher dynamic range of brightness becomes easier to understand when considering an image that, within the same image, has fairly dark areas, such as the shadowed corners of the cave image ImSCN3, but also relatively large areas of very bright pixels, such as the sunlight-lit outside world seen through the cave entrance. This is why, for example, a streetlamp alone creates a different visual experience from the nighttime image of ImSCN2, which contains high-brightness pixel areas.
[0033] The problem here is that, at present, many consumers still have LDR displays, and even in the future, there is a legitimate reason to create two gradings for a movie instead of a typical single HDR image per coding, so we need to define a PL_V_SDR=100nit SDR image that best corresponds to the master HDR image. This is a technical decideratum separate from the technical choice regarding coding itself, which proves that, for example, if we know how to (reversibly) create this secondary image from one master HDR image and the other, we can choose to code and communicate either one of the pair (effectively communicating two images at the price of one image; i.e., communicating only one pixel color component plane of one image at a time in a video moment).
[0034] In such reduced dynamic range images, it's naturally impossible to define objects with a pixel brightness of 5000 nits, like a truly bright sun. The minimum pixel brightness or deepest black also increases to 0.1 nits instead of the more preferable 0.001 nits.
[0035] Therefore, this corresponding SDR image needs to be created with a reduced luminance dynamic range (DR_2).
[0036] This can be done by an automatic algorithm on the receiving display, which may, for example, use a fixed luminance mapping function, or it may be conditioned by simple metadata such as a PL_V_HDR value or potentially one or more other luminance values.
[0037] However, more complex luminance mapping algorithms can generally be used. However, in this application, without loss of generality, we assume that the mapping is defined by a global luminance mapping function F_L (e.g., one function per image). This defines a method for mapping, for at least one image, all possible luminances in the first image (i.e., 0.0001 to 5000) to the corresponding luminances in the second output image (e.g., 0.1 to 100 nits for an SDR output image). The normalization function is obtained by dividing the luminances along both axes by their respective maximum values. In this context, "global" means that the same function is used for all pixels in an image, regardless of further conditions such as their position within the image (e.g., a more general algorithm that uses several functions for pixels that can be classified according to some criteria).
[0038] Ideally, the video creator should determine how to redistribute all luminance along the available range of the secondary image (SDR image). This is because the video creator is best known for how to quasi-optimize for the reduced dynamic range so that, within the limitations, the SDR image still looks at least as good as the intended master HDR image. The reader will understand that actually defining (positioning) such object luminances corresponds to defining the shape of the luminance mapping function F_L (the details of which are outside the scope of this application).
[0039] Ideally, the shape of the function should change for each different scene, i.e., the cave scene versus the sunny Western scene a little later in the film, or generally for each time image. This is called dynamic metadata (F_L(t), where t represents the moment in the image).
[0040] Ideally, the content creator would create the optimal image for each situation, namely each potential end-user display, for example, a PL_D_MDR=800nit display that requires a corresponding PL_V_MDR=800nit image. However, this is usually too much work for the content creator, even in the case of the most expensive offline video production.
[0041] However, the applicant has demonstrated that it is sufficient to create only two different dynamic range reference gradings for a scene (usually at opposite extremes; for example, 5000 nits is sufficient as the highest required PL_V and 100 nits as the lowest required PL_V). This is because all other gradings can be automatically derived from these two reference gradings (HDR and SDR) via a display adaptive algorithm (usually fixed, e.g., standardized) that is applied to the end-user display receiving the information of the two gradings. Generally, the calculations can be performed on any video receiver, such as a set-top box, TV, computer, or film production equipment. The communication channel for HDR images can also be any communication technology, such as terrestrial or cable broadcast, physical media such as Blu-ray discs, the internet, communication channels to portable devices, or video communication between professional sites.
[0042] This display adaptation typically applies a luminance mapping function to the pixel luminance of a master HDR image, for example. However, the display adaptation algorithm needs to determine a different luminance mapping function from F_L_5000to100 (a reference luminance mapping function that connects the luminances of two reference gradings), namely the display adaptation luminance mapping function FL_DA. This is not necessarily trivially related to the original mapping function F_L between the two reference gradings (there are several variations of the display adaptation algorithm). The luminance mapping function between a master luminance defined with a PL_V dynamic range of 5000 nits and an intermediate dynamic range of 800 nits is described in this text as F_L_5000to800.
[0043] Here, the F_L_5000to100 function does not "simply" map to the expected position when it exceeds the 800nit MDR image luminance range, but rather symbolically illustrates display adaptation with an arrow that maps to a slightly higher position (i.e., in such an image, the cowboy should be slightly brighter, at least according to the selected display adaptation algorithm) (for only one of the average pixel luminances of the object). Therefore, while more complex display adaptation algorithms may place the cowboy in a higher position than indicated, some customers are satisfied with the simpler position where the connection between a 500nit HDR cowboy and an 18nit SDR cowboy exceeds the 800nit PL_V luminance range.
[0044] Typically, a display adaptive algorithm calculates the shape of the display adaptive luminance mapping function FL_DA based on the shape of the original luminance mapping function F_L (or reference luminance mapping function, also known as the reference regrading function).
[0045] The explanation based on Figure 1 constitutes a technical desiderator for HDR video coding and / or processing systems, and Figure 2 shows several exemplary technical systems and their components for realizing the desiderator (not limited to) according to the applicant's codec approach. Those skilled in the art should understand that these components can be embodied in various devices, etc. Those skilled in the art should understand that this example is presented as a pulse prototo (part to be received as a whole) for various HDR codec frameworks, merely to provide a background understanding of some of the principles of operation, and is not intended to particularly limit any of the embodiments of the innovative contributions shown below.
[0046] While possible, the technical communication of two actual, different images in a single moment (HDR and SDR grading, each communicating as three color planes) is expensive, especially considering the amount of data required.
[0047] Furthermore, this is not necessary. Knowing that the luminance of all corresponding secondary image pixels can be calculated based on the luminance of the primary image and the function F_L, it is possible to decide to communicate only the primary image and the function F_L as metadata at a single moment (and it is also possible to choose to communicate either the master HDR image or the SDR image as a representative of both). The receiver, knowing its (usually fixed) display adaptation algorithm, can then determine the FL_DA function at its end based on this data (further metadata may be communicated to control or guide the display adaptation, but this is not currently implemented).
[0048] There are two modes: one for each instantaneous image and another for communicating with the function F_L.
[0049] In the first backward-compatible mode, an SDR image is communicated ("SDR communication mode"). This SDR image can be displayed directly on a conventional SDR display (without requiring further luminance mapping), but an HDR display needs to apply the F_L function or FL_DA function to obtain an HDR image from the SDR image (or a variation of the communicated function, i.e., an upgrading variation or a downgrading variation, and vice versa). Interested readers should refer to the following for all details of the standardized exemplary first mode approach of the applicant: ETSI TS 103 433-1 V1.2.1(2017-08):High-Performance Single Layer High Dynamic Range System for use in Consumer Electronics devices;Part 1:Directly Standard Dynamic Range(SDR) Compatible HDR System(SL-HDR1).
[0050] In another mode, the master HDR image itself is communicated, i.e., a 5000-nit image and a function F_L that allows the calculation of a 100-nit SDR image (or other low dynamic range images via display adaptation) from it ("HDR communication mode"). The master HDR communicated image itself can be encoded, for example, using a PQ EOTF.
[0051] Figure 2 also shows the entire video communication system. On the transmitting side, it starts with image source 201. Depending on whether it is, for example, offline video created from an internet distribution company or actual broadcast, this could be anything from a hard disk to a cable output from a television studio, etc.
[0052] This results in, for example, a master HDR video (MAST_HDR) that has been color-graded by a human color grader, or a shaded version of the camera capture, or a master HDR video created using an automatic brightness redistribution algorithm.
[0053] In addition to grading the master HDR image, a set of often reversible color conversion functions F_ct is defined. Without intending to lose generalization, we assume this includes at least one luminance mapping function F_L (but there may be further functions or data specifying, for example, how pixel saturation changes from HDR to SDR grading).
[0054] This luminance mapping function defines the mapping between HDR and SDR reference grading, as mentioned above (the latter being the SDR image Im_SDR communicated to the receiver in Figure 2, for example, whether the data is compressed via MPEG or other video compression algorithms).
[0055] The color mapping of the color converter 220 should not be confused with that applied to the raw camera feed to obtain the master HDR video, which is assumed to be input here. This is because the color conversion is for obtaining the image being communicated and, at the same time, the desiderator of the regrading, which is technically formulated within the luminance mapping function F_L.
[0056] In the case of an exemplary SDR communication type (i.e., SDR communication mode), the master HDR image is input to the color converter 202. The color converter 202 is configured to apply an F_L luminance mapping to the luminance of the master HDR image (MAST_HDR) to obtain all corresponding luminances written to the output image Im_SDR. For clarity, we assume that the shape of this function is adjusted by a human color grader using color grading software for each shot of images from similar scenes in a film. The applied function F_ct (i.e., at least F_L) is written to (dynamic, processed) metadata communicated with the image, such as MPEG Supplemental Enhancement Information Data SEI(F_ct), or a similar metadata mechanism in other standardized or non-standardized communication methods.
[0057] If the HDR images to be communicated are correctly redefined as the corresponding SDR images Im_SDR, these images are often compressed using existing video compression techniques (e.g., MPEG HEVC, VVC, AV1, etc.) (at least for broadcasting to the end user). This is performed by the video compressor 203, which forms part of the video encoder 221 (the video encoder may be included in various types of video creation devices or systems).
[0058] The compressed image Im_COD is transmitted to at least one receiver via an image communication medium 205 (e.g., satellite, cable, or internet transmission, compliant with ATSC3.0 or DVB, etc.) (however, HDR video signals may also be communicated via cable between two video processing units, for example).
[0059] Typically, further conversion is performed by the transmit formatter 204 before communication. Depending on the system, the transmit formatter 204 applies techniques such as packetization, modulation, and transmission protocol control. This is usually done using an integrated circuit.
[0060] At any receiving site, the corresponding video signal unformatter 206 applies the necessary unformatting method, such as modulation, to reacquire the compressed video as, for example, a set of compressed HEVC images (i.e., HEVC image data).
[0061] The video decompressor 207 performs, for example, HEVC decompression to obtain a pixelated, uncompressed image stream, Im_USDR. In this example, image Im_USDR is an SDR image, but in other modes it would be an HDR image. The video decompressor also unpacks the required luminance mapping function F_L, or more commonly, the color conversion function F_ct, from the SEI message, for example. The image and functions are input to the color converter 208 (of the decoder). The color converter 208 is configured to convert the SDR image to an image with an arbitrary non-SDR dynamic range (i.e., an image with PL_V greater than 100 nits, usually at least several times higher, e.g., 5 times higher).
[0062] For example, a reconfigured 5000nit HDR image Im_RHDR is reconstructed as a close approximation of the master HDR image (MAST_HDR) by applying the inverse color transform IF_ct of the color transform F_ct used on the encoding side to create Im_LDR from MAST_HDR. This image is then sent to, for example, display 210 for further display adaptation. However, the display-adapted image Im_DA_MDR can be created in one step during decoding by using the FL_DA function (determined, for example, in the firmware, in the offline loop) in the color converter instead of the F_L function. Therefore, the color converter may include a display adaptation unit 209 to derive the FL_DA function.
[0063] The optimized, for example, 800nit display-adapted image Im_DA_MDR is sent to, for example, the display 210 if the video decoder 220 is included in, for example, a set-top box or computer. Alternatively, if the decoder is in, for example, a mobile phone, it is sent to the display panel. Or, if the decoder is in, for example, an internet-connected server, it is sent to a cinema projector.
[0064] Figure 3 shows a useful variation of the internal processing of the color converter 300 (i.e., corresponding to 208 in Figure 2) of an HDR decoder (or an encoder that usually has nearly the same topology but uses an inverse function and does not usually include display adaptation). The luminance of a pixel (in this example, a pixel in an SDR image) is input as the corresponding lumens Y'SDR. The chrominance, also known as the chroma components Cb and Cr, is input to the lower processing path of the color converter 300.
[0065] Luma Y'SDR is mapped by the luminance mapping circuit 310 to the required output luminance L'_HDR (e.g., master HDR reconstructed luminance, or other HDR image luminance). The luminance mapping circuit 310 applies an appropriate function for a particular image and maximum display luminance PL_D, e.g., a display-adapted luminance mapping function FL_DA(t), obtained from the display adaptive function calculator 350, which uses a reference luminance mapping function F_L(t) communicated with metadata as input. The display adaptive function calculator 350 can also determine a function suitable for handling chrominance. Here, we simply assume that a set of multiplication coefficients mC[Y] for each possible input image pixel luma Y is stored, for example, in a color LUT 301.
[0066] Indexing the color LUT 301 with the lumens value Y of the currently color-converted (luminance-mapped) pixel yields the required multiplication coefficient mC as the LUT output. This multiplication coefficient mC is then used by the multiplier 302 to multiply the two chrominance values of the current pixel, resulting in the color-converted output chrominance. Cbo=mC * Cb, Cro=mC * Cr
[0067] The chrominance can be converted to normalized, nonlinear R'G'B' coordinates R' / L', G' / L', and B' / L' via a fixed color matrixing processor 303 that applies standard colorimetric calculations.
[0068] The R'G'B' coordinates that give the output image appropriate brightness are obtained by the multiplier 311. The multiplier 311 calculates the following: R'_HDR=(R' / L') * L'_HDR, G'_HDR=(G' / L') * L'_HDR, B'_HDR=(B' / L') * L'_HDR, This can be combined into a color triplet, R'G'B'_HDR.
[0069] Finally, the display mapping circuit 320 may provide further mapping to the format required for the display. This results in a display-driven color D_C, which is not only formulated with the desired color measurement for the display (e.g., HLG OEFT format), but in some variations, the display mapping circuit 320 may also be positioned to perform specific color processing for the display. That is, the display mapping circuit 320 may, for example, further remap a portion of the pixel luminance.
[0070] Several examples of elucidating appropriate display adaptation algorithms for deriving FL_DA functions corresponding to any possible F_L functions determined by the creating grader are taught in WO2016 / 091406 or ETSI TS 103 433-2 V1.1.1(2018-01).
[0071] However, these algorithms do not take into account the smallest black level that can be displayed on the end user's screen.
[0072] In fact, it could be said that the minimum brightness BL_D is so small that it appears to be zero. Therefore, display adaptation primarily considers the differences in the maximum brightness PL_D of various displays compared to the maximum brightness PL_V of the video.
[0073] As can be seen from Figure 18 of the prior application WO2016 / 091406, any input function (in the example, a simple function formed from two linear segments) is typically scaled diagonally based on a metric positioned along an angle of 135 degrees, starting from the horizontal axis of input luminance in a plot normalized to input and output luminances of 1.0. It should be understood that this is merely one example of display adaptation across the entire class of display adaptation algorithms and is not mentioned in any manner intended to limit the applicability of novel display adaptation concepts. For example, specifically, the angle in the metric direction may be other values.
[0074] However, this metric and its action on the reshaped F_L function, i.e., the determined FL_DA function, depend only on the maximum luminance PL_V and the maximum luminance PL_D of the display, which provides an optimally regraded intermediate dynamic range image. For example, the 5000nit position corresponds to the zero-metric point on the diagonal (any position on the diagonal corresponds to the possible pixel luminance of the input image), and the 100nit position (marked PBE) is the point in the original F_L function.
[0075] A useful variation of this method, display adaptation, is summarized in Figure 4 by showing its action on a plot of possible normalized input luminance Ln_in versus normalized output luminance Ln_out (these are converted to actual luminances by multiplying by the maximum luminance of the display associated with the normalized luminance (i.e., the PL_V value)).
[0076] For example, the video creator designed a luminance mapping strategy between two reference gradings, as illustrated in Figure 1. Therefore, for any possible normalized luminance of a pixel in the input image Ln_in, e.g., the master HDR image, this normalized input luminance must be mapped to the normalized output luminance Ln_out of the second reference grading, which is the output image. This regrading of all luminances corresponds to a function F_L, which can have many different shapes determined by a human grader or grading automaton, and the shape of this function is communicated together as dynamic metadata.
[0077] The issue here is the shape that the derived quadratic version of the F_L function should have in this simple display adaptation protocol to map to an MDR image (not a reference SDR image) for an intermediate dynamic range display (again, assuming a mapping starting from an HDR reference grading image as the input image). For example, the metric can be used to calculate that, for example, an 800nit display should have a 50% grading effect. Full 100% is regrading from a master HDR image to a 100nit PL_V SDR image. In general, the metric can be used to determine any point between the point where there is no regrading and the point where there is full regrading to a second reference image for any possible normalized input luminance (Ln_in_pix) of a pixel, indicated as the display-adapted luminance L_P_n. The position of the display-adapted luminance L_P_n depends, of course, on the input normalized luminance, but also on the maximum luminance value (PL_V_out) associated with the output image.
[0078] The corresponding display-adapted luminance mapping FL_DA can be determined as follows (see Figure 4a): Select one of all input luminances (e.g., Ln_in_pix). This corresponds to a starting position on the diagonal (shown as a square) that has equal angles with respect to the input and output axes of the normalized luminance. At each point on the diagonal, place a scaled version of the metric (scaled metric SM) so that it is orthogonal to the diagonal (or 135 degrees counterclockwise from the input axis), starts on the diagonal, and ends at its 100% level at a point on the F_L curve, i.e., the intersection (shown as a pentagon) with the scaled metric SM orthogonal to the F_L curve. Place a point at the 50% level of the metric, i.e., the middle [in this case, the PL_V value of the output image is set to the PL_D value of the display that needs to supply the display-optimized image] (in this example, for this PL_D value of the display that needs to calculate the image). By doing this for all points on the diagonal corresponding to all Ln_in values, the FL_DA curve is obtained, which is shaped similarly to the original curve, i.e., the same regrading is performed, but the maximum luminance is rescaled / adjusted. This function is ready to be applied to calculate the required corresponding optimally regraded / display-adapted 800nit PL_V pixel luminance, given any input HDR luminance value of Ln_in. This function FL_DA is applied by the luminance mapping circuit 310.
[0079] Generally, the properties of this display adaptation are as follows (not intended to be more specific): The orientation of the metric is fixed beforehand as technically desirable. Figure 4b shows another scaled metric, namely the vertically oriented scaled metric SMV (i.e., orthogonal to the axis of the normalized input luminance Ln_in). Here again, 0% and 100% (or 1.0) correspond to no regrading (i.e., identity transformation to input image luminance) and regrading to the second of two reference grading images (associated in this example by a luminance mapping function F_L2 of different shapes), respectively.
[0080] The position of the measurement points on the metric, i.e., the location of values such as 10% and 20%, does, strictly speaking, fluctuate, but is usually nonlinear.
[0081] This is pre-designed in technologies such as television displays. For example, the function described in WO2015007505 can be used. Also, a * The logarithmic function can also be designed such that (log(PL_V)+b) is equal to 1.0 for a PL_V_HDR value (e.g., 5000 nits), the point 0.0 corresponds to a PL_V_SDR reference level of 100 nits, and vice versa. Any PL_V_MDR position for which image brightness needs to be calculated can be obtained from the designed mathematical processing of the metric.
[0082] Figure 5 summarizes the effects of these metrics.
[0083] For example, a display adaptive circuit 510 in a television or set-top box includes a configuration processor 511. The configuration processor 511 sets values for image processing and then executes the pixel color of the incoming image to be processed. For example, the maximum brightness value of the output image PL_V_out optimized for the display is set once in the set-top box by polling from the connected display (i.e., the display communicates its maximum displayable brightness PL_D to the set-top box). Alternatively, if the circuit is in a television, it may be set by the manufacturer, etc.
[0084] In some embodiments, the luminance mapping function F_L is different for each input image (in other variations, it is fixed for a large number of images) and is input from a metadata source 512 (e.g., broadcast as an SEI message or read from a sector of memory such as a Blu-ray disc). This data establishes the normalized height of normalized metrics (Sm1, Sm2, etc.), and then the desired position of the PL_D value can be found from the mathematical equation of the metric.
[0085] When image 513 is input, the consecutive pixel luminances (e.g., Ln_in_pix_33, Ln_in_pix_34, etc.) pass through a color processing pipeline to apply display adaptation, resulting in the corresponding output luminances such as Ln_out_pix_33.
[0086] Furthermore, this approach does not specifically address the minimum black luminance.
[0087] This is because the usual approach is as follows: Black levels depend heavily on actual viewing conditions and vary even more than the display characteristics (i.e., the highest PL_D). A wide range of effects arise, from the physical lighting aspects to the optimal configuration of photosensitive molecules in the human eye.
[0088] Therefore, the goal is to create a superior image "for display," and that's it (i.e., how much better the intended HDR display is in terms of brightness than a regular SDR display). It can then be adjusted later, as needed, to suit the viewing conditions. This is an (undefined) ad-hoc task left to the display.
[0089] Therefore, we typically assume that a display can display all the necessary pixel brightness encoded within the image (here, we assume an MDR image already optimized for the PL_D value, i.e., an image that typically has at least some pixel regions within the image that rise up to PL_D) up to its variable brightness capability, i.e., PL_D. This is because we generally do not want to suffer the harsh consequence of white clipping, although, as mentioned earlier, the blacks in an image are often of less interest.
[0090] In any case, black is "almost" visible, so it's not important if some of it is slightly less visible. At the very least, the potentially very bright pixel brightness of the master HDR grading can be optimally placed within the limited upper limit of the display, e.g., above 200 nits, e.g., 200~PL_D=600 nits (for example, for master HDR brightness up to 5000 nits).
[0091] This is similar to assuming that black is always zero nits for all images and (at least almost) all displays. White clipping is a far more visually unpleasant property than losing some of the black (in which case something is still visible, if not very pleasant).
[0092] However, this approach can be insufficient. In a viewing room with considerable ambient light (for example, a living room with large windows during the daytime where consumer television viewers are present), a significant subrange of the darkest luminances may become invisible or at least undetectable. This differs from the ambient light conditions in a video editing room where video is created, which can be dim or even dark.
[0093] Therefore, for example, you may need to increase the brightness of these pixels using the control buttons typically found on the display (the so-called brightness buttons).
[0094] Taking the electronic behavior model of a television, as described in Rec.ITU-R BT.814-4(07 / 2018), as an example, a television in an HDR scenario acquires lumens and chroma pixels (which actually drive the display) and converts these into nonlinear R', G', and B' nonlinear drive values (according to standard colorimetric calculations) to drive its panel. The display then processes these R', G', and B' nonlinear drive values in a PQ EOTF to determine the brightness of the front screen pixels to be displayed (i.e., how to drive, for example, PLED panel pixels or LCD pixels; there are usually internal processes that still describe the electro-optical physical behavior of LCD materials, but those aspects are irrelevant to this discussion).
[0095] For example, a control knob on the front of the display can provide a luma offset value b (where the smallest black patch is visible when it is 2% above the PLUGE or other test pattern, while -2% black is not visible).
[0096] The original, uncorrected display behavior is as follows, for example: LR_D=EOTF[max(0,R')]=PQ[max(0,R')] LG_D=EOTF[max(0,G')]=PQ[max(0,G')] LB_D=EOTF[max(0,B')]=PQ[max(0,B')] [Formula 1]
[0097] In this equation, LR_D is a (linear) amount of red contribution that is displayed to create a particular pixel color with a specific luminance ((partial) nit), and R' is a nonlinear Lumacode value (e.g., 419 out of 1023 in 10-bit coding).
[0098] The same applies to the blue and green components. For example, if we need to make a particular achromatic color 1 nit (the total brightness of that color to the eye), we need, for example, 1 unit of blue, and the same applies to red and green. If we need to make 100 nits of that same color, we can say LR_D = 100 nits (the visual weight of red is excluded from the equation).
[0099] Next, when this display drive model is controlled via the Luma offset knob, the general equation becomes: LR_D_c=EOTF[max(0,a * R'+b), where a = 1-b / inv_EOTF[PL_D], etc. [Equation 2] inv_EOTF[ ] is the mathematical inverse equation of EOTF[ ].
[0100] This approach doesn't display zero black somewhere hidden within the invisible black of the display, but rather raises the black to a level where it is sufficiently distinguishable. (Note that consumer displays can use mechanisms other than PLUGE. For example, viewer preferences can be used, potentially resulting in another, and possibly suboptimal, luma offset value b that the viewer might prefer.)
[0101] This is a post-processing step for the display after creating an optimally regraded image. Specifically, first, the decoder calculates an optimally theoretically regraded image (e.g., the initial mapping to the reconstructed master HDR image), and then, for example, a luminance remapping to an MDR image with a PL_V of 550 nits, i.e., a luminance remapping that takes into account the display's brightness capability PL_D.
[0102] Then, after this optimal image is determined according to the filmmaker's ideal vision, it is further mapped by the display, taking into account the expected visibility of black in the image.
[0103] According to the inventors, the problem is that this is a rather crude operation of grading that may have been carefully crafted; therefore, there may be an opportunity to invent a better alternative that gives a better-looking image.
[0104] U.S. Patent Application Publication 2019 / 0304379 describes another technical approach to creating brightness regrading optimized for displays, which also takes into account the amount of ambient light.
[0105] The system is based on determining an appropriate amount of average backlight setting (which may vary with local dimming) in combination with the native and achievable dynamic range of the LCD display panel (for example, if closed pixels are transmitting an amount of dim leak light equivalent to 1 / 1000th the amount of light from open pixels).
[0106] Display mapping is based on the design of an optimal S-curve, which is defined by one or more, typically three, metadata values that describe the input image. Along with the metadata values, the optimal S-curve is determined by the image receiver. This S-curve approach worked well in the wet photography era when creating LDR prints from negative captures with a larger dynamic range, and has proven to work well in the digital age as well.
[0107] For example, metadata specifying the average luminance (or luma, if the method applies to the luma region), called the midpoint luminance value, can determine the shape and position of the linear slope of the S-curve. If the input image is fairly dark (e.g., nighttime with poorly lit buildings), readers will understand that mapping the lower luminances of the HDR input range to the middle of the output range (e.g., 20-60 nits in LDR) will make it look good on displays with a lower dynamic range. Similarly, the maximum luminance metadata of the input image helps determine where the S-curve should begin to stabilize at its top. If the maximum is much higher than the average (ultimately needing to display up to 60 nits), and only 40 nits remain available for highlights, then soft clipping of the top of the S-curve should begin early in the luminance input range (this works particularly well if there aren't too many very bright pixels, e.g., a few shiny patches reflecting scene lighting on a metallic object).
[0108] A similar type of tapering is performed for lower available ranges, i.e., output luminances below 20 nits, by determining stepped clipping at the lower end of the S-shaped curve of the luminance mapping (sometimes called the foot or toe of the S-shaped curve).
[0109] Therefore, this display adaptation method (or, as referred to in U.S. Patent No. 379, the display management method) works well with all kinds of input images, but it assumes that the image is downgraded for a display with a smaller dynamic range in the same viewing environment in which the input image was created (for example, in the ambient / environmental light of a grading booth, i.e., darkened with controlled low lighting and graded on a 4000-nit grading display).
[0110] The idea of generalizing to any viewing environment involves designing a strategy based on modifying a graded master input HDR image so that it would have been graded if the grading booth were illuminated in a different manner (for example, under the amount of light inside a train during the day). This redefinition of the graded master image is called a virtual image.
[0111] Here too, since this is just an image, we can use the exact same display adaptation.
[0112] The master / input image is brightened for a brighter environment (all luminances are increased, from darkest to maximum). Therefore, even after standard S-curve display management, if the ambient light level remains at the reference dim light level, it will appear too bright on end displays with a smaller dynamic range. However, both will look good at the new light level. It is important not to forget to calculate the new metadata (minimum, midpoint, maximum) of the brightened (virtual) input image, which is a redefined approach.
[0113] This approach differs from systems that use a single control value, which allows positioning along a metric that can be placed along various points of a reference luminance mapping function that can be specified in the metadata of almost any shape as needed, so that new (adjusted) curve points can be found at these points of the metric, not a control value based on maximum luminance, while redefining it with the black value.
[0114] Document 6 / 115-E of March 27, 2017, of the ITU Radiocommunication Research Committee, titled "High Dynamic Range Television for Production and International Program Exchange" (Question ITU-R142-2 / 6), is a revision of the report BT.2390-1 by Working Group 6C and states the following:
[0115] The image of the reference viewing environment displayed on the reference display is created by modifying the reference OOTF with artistic adjustments (creative intent). The receiving image is display-adjusted for a non-reference display in a non-reference viewing environment. The reference OOTF directly optimizes camera capture. This camera capture is merely a relative brightness measurement of the scene and is not optimal as a good image appearance for a typical display. This approach can be applied to perceptual quantizer HDR coding systems. In other words, the artistically optimized pixel brightness can be encoded as PQ luminance before HDR image ( / video) transmission.
[0116] In addition to the camera-side OOTF, which performs a rough average grading of the camera image, an EETF is defined, which is applied to adapt images created in high dynamic range to displays with a smaller dynamic range (as a lumen-to-lumen mapping for calculating luminance mapping). The codeable PQ range allows coding of images with pixel luminances ranging from as dark as 1 / 10000 nits and potentially as bright as 10,000 nits, and since typical HDR displays currently display from 0.01 to 1000 nits, a fixed remapping curve is presented. However, while 0.01 nits characterizes black on such displays (e.g., commercially available consumer televisions), under various ambient light levels and in front of screen reflections, the minimum achievable black can be, for example, 0.1 nits. Thus, the same S-shaped curve can be calculated for scenarios starting from various minimum black levels. Mapping a specific curve to the available display range is, strictly speaking, quite different from changing the control parameters of a metric-based function shape following a display adaptation algorithm. [Overview of the project]
[0117] Images that look better visually under various ambient lighting levels can be obtained by calculating this through a method that processes the input image to obtain the output image. The input image has pixels with input brightness within a first luminance dynamic range (DR_1), and the first luminance dynamic range has a first maximum brightness (PL_V_HDR). The reference luminance mapping function (F_L) is received as metadata associated with the input image. The reference luminance mapping function specifies the relationship between the luminance of the first reference image and the luminance of the second reference image. The first reference image has a first reference maximum brightness, and the second reference image has a second reference maximum brightness. The input image is equal to one of the first reference image and the second reference image. The output image is not equal to either the first reference image or the second reference image. The above process involves applying the adapted luminance mapping function (FL_DA) to the input pixel luminance to obtain the output luminance. The adapted luminance mapping function (FL_DA) is calculated based on the reference luminance mapping function (F_L) and the maximum luminance value (PLA), which is a control parameter specifying the degree to which the adapted luminance mapping function deviates from the reference luminance mapping function. The above calculation of the adapted luminance mapping function involves finding the metric position corresponding to the maximum luminance value (PLA). The first endpoint of the metric corresponds to the first maximum brightness (PL_V_HDR), and the second endpoint of the metric corresponds to the maximum brightness of one of the first and second reference images that is not equal to the input image. The maximum brightness value (PLA) is calculated based on the maximum brightness (PL_D) of the display from which the output image is supplied, and the black level value (b) of the display. The above calculation is characterized by applying the inverse function of the electric-optical transfer function to the maximum luminance (PL_D), subtracting the black level value (b) from the resulting value, and applying the electric-optical transfer function to the subtraction to obtain the maximum luminance value (PLA) used in the above calculation of the applied luminance function.
[0118] The primary characteristic of the dynamic range of pixel brightness is the maximum brightness. That is, every pixel in an image created (encoded) according to such a dynamic range must not be brighter than this maximum brightness. A commonly used agreement is that, for simplification, the brightness of the components is prescaled with visual weights, so that 1 nit of red, green, and blue gives a pixel color of 1 nit of gray (in this case, the same OETF or EOTF equation can be used). If an image is part of a series of images (also known as a video), not every image in that series must necessarily contain a pixel with a brightness equal to the maximum brightness, but the series of images is defined by this maximum value.
[0119] In prior art display adaptation, the maximum brightness (PL_V) is the maximum brightness (PL_D) obtained from the connected or to be connected display (e.g., a 500-nit display). The content itself is optimized to match the display's typically lower maximum display brightness compared to the content represented by the brighter pixel brightness of the input image. This approach has proven highly applicable, giving video creators a means to control how their created images (typically the highest quality master images suitable for any display, but specified to be optimized for a specific associated target display; the images are specifically color-graded for this display) ultimately look on any available endpoint (e.g., consumer audience) display.
[0120] Generally, we would like to use the PL_V control value associated with the end display on which the image is viewed, but this concept can be redefined.
[0121] This determines the adjusted value of the maximum brightness, which specifically takes into account the display's black level (i.e., what is visible on the display). This allows for a different level of control over the regrading of displays that typically have a smaller dynamic range than the target display of the received image. This is suitable for viewing ambient brightness adaptive regrading, as the inventors have found.
[0122] Any HDR electro-optical transfer function can be selected for this method (usually fixed in advance by the technology's designers). The electro-optical transfer function (EOTF) is a function that defines the display pixel brightness corresponding to the so-called lumacode. These lumacodes are, for example, 8-bit words, 10-bit words, or 12-bit words, where each word (luma level) specifies a particular brightness.
[0123] While not a limiting factor of this innovation, a well-functioning EOTF is a known perceptual quantizer defined in the SMPTE ST.2084 standard. This works perceptually well, much like the HDR EOTFs, which are mostly perceptually uniform.
[0124] The black level, as mentioned above, can be determined by a standard test pattern using a dark black test patch such as PLUGE, but is generally an arbitrary selectable value b that, given any viewing room constraints (e.g., ambient lights), strictly or approximately represents the darkest black still discernible on the display. That is, it is a black that appears slightly brighter than true black, which would obscure all darker pixels. This can be represented in the domain of perceptual quantizer lumas, without any intention of limitation in the elucidation teachings of this patent application (as those skilled in the art will understand, is a set of lumas that can represent the corresponding luminance, which can also be represented in the luminance domains themselves, and by switching between domains where technically advantageous).
[0125] Typically, the reference image differs from both the input and output images, but we assume the input image is identical to one of the reference images (e.g., the HDR image from an HDR / SDR pair). The output image is not identical to one of the reference images and is typically an image with an intermediate dynamic range of a particular display (e.g., lower than the received HDR image with 2000 nits, e.g., with an intermediate maximum brightness of 900 nits). The received image is optimized for this particular display.
[0126] Therefore, the method (or apparatus) is typically as follows: A method for processing an input image and obtaining an output image, The input image has pixels with input brightness within a first luminance dynamic range (DR_1), and the first luminance dynamic range has a first maximum brightness (PL_V_HDR). The reference luminance mapping function (F_L) is received as metadata associated with the input image. The reference luminance mapping function specifies how to map the luminance of an input image to the luminance of a second image having a different dynamic range from the input image. The second image has a reference maximum brightness, The output image is not equal to either the input image or the reference image, and has a different maximum output brightness (for example, between these values, below the lower of the two values, or above the higher of the two values). The above process involves applying the adapted luminance mapping function (FL_DA) to the input pixel luminance to obtain the output luminance. The adapted luminance mapping function (FL_DA) is calculated based on the reference luminance mapping function (F_L) and the controlled maximum luminance value (PLA), which is a control parameter specifying the degree to which the adapted luminance mapping function deviates from the reference luminance mapping function. The above calculation of the adapted luminance mapping function involves finding the metric position corresponding to the maximum luminance value (PLA). The first endpoint of the metric corresponds to the first maximum brightness (PL_V_HDR), and the second endpoint of the metric corresponds to the reference maximum brightness. The maximum brightness value (PLA) is characterized by being calculated based on the maximum brightness (PL_D) of the display to which the output image is supplied and the black level value (b) of the display. The above calculation involves applying the inverse function of the electric-optical transfer function to the maximum luminance (PL_D), subtracting the black level value (b) from the resulting value, and applying the electric-optical transfer function to the subtraction to obtain the maximum luminance value (PLA) used in calculating the applied luminance mapping function.
[0127] Controlled Maximum Brightness should not be confused with the Output Maximum Brightness of the output image. Controlled Maximum Brightness functions to have optimal regrading behavior, including taking black levels into consideration. Output Maximum Brightness, on the other hand, is typically the maximum pixel brightness that occurs in the generated output image (or, in the case of a video of temporally continuous images, typically the maximum brightness of at least one or more pixels of at least one of the video images). PLA determines the position of the metric between the diagonal representing the input brightness and the trajectory of the curve of the reference function, and all positions on the positioned metric define the adjusted function required to obtain the output brightness.
[0128] In many cases, the input image is a high dynamic range image (e.g., a master image color-graded according to the best predictable dynamic range quality, such as a 5000 nit PL_V image). Often, the reference image for which the reference luminance mapping function is defined is an LDR image, i.e., an image with PL_V_LDR = 100 nits (however, other variations are possible, such as the input image being an LDR image or an intermediate dynamic range image). The target display has the capability to range from conventional LDR displays (approximately 100 nits, i.e., PL_D which can be assumed to be approximately 100 nits / 100 nits for all practical purposes) to low-end HDR displays such as 700 nit displays, to high-end HDR displays with PL_D of 1000 nits or more, to even higher-end HDR displays exceeding 3000 nits, and displays exceeding 5000 nits. Viewing conditions vary from a very dark living room (e.g., no lighting other than the TV) to dimly lit indoor and outdoor environments during the day.
[0129] When it says "image with luminance," the reader should understand that pixel luminance is typically encoded as luma (e.g., 10 bits). However, luma can be uniquely converted to corresponding luminance via an EOTF that defines absolute luminance (e.g., known to the receiver by prior agreement in a single EOTF coding system, or notified via metadata if variable). Similarly, luminance processing can indeed be performed as luma processing in the luma region, but is ultimately luminance processing, at least since the end display will show a certain luminance (e.g., a maximum display capability of up to 750 nits).
[0130] The single-control-parameter approach of this regrading format should not be confused with methods that simply perform the transformation based on minimum black (i.e., not via a redefined maximum luminance).
[0131] The key is not just to reuse the display adaptive approach (algorithm), but to ensure that it largely adheres to the shape of a reference luminance mapping function that has been carefully designed by the video creator and indicates the specific luminance regrading needs of various objects in any given HDR scene image.
[0132] Furthermore, this approach does not rely on a fixed display adaptation strategy and can be used with any non-decreasing shape of the reference luminance mapping function. A key feature of this display adaptation approach is that it automatically adapts to any display through mathematical processing that optimally considers the output display's dynamic range, including its PL_D.
[0133] While this mathematical processing of display adaptation can be used to optimize for lower display maximum brightness, just like existing display adaptations, in this approach, the maximum brightness used in this display adaptation—that is, the maximum brightness that determines the metric position of the points in the display-adapted brightness mapping function (FL_DA)—is defined quite differently from the maximum brightness PL_D of the display from which the display-optimized image of the input image is supplied.
[0134] Preferably, display adaptation is performed with a tuned maximum luminance PLA that is not equal to PL_D, after which a black level value (b) is added to the output luminance obtained by the display adaptation (preferably these output luminances are expressed as luma in the luma region of the perceptual quantizer).
[0135] Advantageously, the present invention can be embodied as a color conversion circuit (600), and this color conversion circuit (600) is An input connector (601) that receives an input image, An output connector (699) outputs an output image calculated by a color conversion circuit based on the input image, A metadata input unit (602) receives a reference luminance mapping function (F_L), which is metadata associated with the input image, Equipped with, The input image has pixels with input brightness within a first luminance dynamic range (DR_1), and the first luminance dynamic range has a first maximum brightness (PL_V_HDR). The reference luminance mapping function specifies the relationship between the luminance of the first reference image and the luminance of the second reference image. The first reference image has a first reference maximum brightness, and the second reference image has a second reference maximum brightness. The input image is equal to one of the first reference image and the second reference image. The output image is not equal to either the first reference image or the second reference image. The color conversion circuit (600) includes a luminance mapping circuit (310) that applies an adapted luminance mapping function (FL_DA) to the input pixel luminance to obtain the output luminance. The color conversion circuit (600) includes a luminance mapping function calculation circuit (611) that calculates an adapted luminance mapping function (FL_DA) based on a reference luminance mapping function (F_L) and a maximum luminance value (PLA) as a control parameter that specifies the degree to which the adapted luminance mapping function deviates from the reference luminance mapping function. The above calculation involves finding the metric position corresponding to the maximum luminance value (PLA). The first endpoint of the metric corresponds to the first maximum brightness (PL_V_HDR), and the second endpoint of the metric corresponds to the maximum brightness of one of the first and second reference images that is not equal to the input image. The luminance mapping function calculation circuit (611) is characterized by receiving the maximum luminance (PL_D) of the display to which the output image is supplied and the black level value (b) of the display. The luminance mapping function calculation circuit (611) calculates the maximum luminance value (PLA) by applying the inverse function of the electric-optical transfer function to the maximum luminance (PL_D), subtracting the black level value (b) from the resulting value, and applying the electric-optical transfer function to subtraction to obtain the maximum luminance value (PLA).
[0136] Advantageously, the color conversion circuit (600) has a luminance mapping function calculation circuit (611) that applies a perceptual quantizer function as the electro-optical transfer function.
[0137] Advantageously, the color conversion circuit (600) includes a display mapping circuit (320) that adds a black level value (b) to the luminance value of the output brightness represented in the luminance region of the perceptual quantizer. [Brief explanation of the drawing]
[0138] These and other aspects of the methods and apparatus according to the present invention will be made clear and clarified with reference to the implementations and embodiments described below and to the accompanying drawings. The accompanying drawings serve only as non-limiting specific examples illustrating more general concepts, and dashed lines are used in the drawings to indicate that components are optional. Components not shown by dashed lines are not necessarily required. Dashed lines are also used to indicate that an element is described as required but is hidden inside an object, or for intangible things such as the selection of an object / region.
[0139] [Figure 1] This schematic outlines some typical color transformations that occur when a high dynamic range image is optimally color-graded and mapped to a corresponding image with a lower dynamic range (e.g., a standard dynamic range image with a maximum brightness of 100 nits) that looks similar (given the difference between the first dynamic range DR_1 and the second dynamic range DR_2, it is possible to achieve the desired similarity). This also corresponds, in the case of losslessness, to mapping a received SDR image that is actually encoding an HDR scene to a reconstructed HDR image of that scene. Brightness is shown as a position on the vertical axis from the darkest black to the maximum brightness PL_V. The brightness mapping function is symbolically shown by arrows mapping the average object brightness from the brightness of the first dynamic range to the second dynamic range (those skilled in the art will understand how to plot this equivalently as a conventional function (e.g., an axis normalized to 1, normalized by dividing by each maximum brightness)). [Figure 2]A schematic example of a high-level diagram of a technique recently developed by the applicant for encoding high dynamic range images (i.e., images that can typically have a brightness of at least 600 nits or more (usually 1000 nits or more)). This allows for the communication of HDR images either by themselves or by adding metadata to a correspondingly luminance-regraded SDR image, which includes at least a color conversion function (F_L) appropriately determined to the pixel color, used by a decoder to convert the received SDR image to an HDR image. [Figure 3] As a (non-limiting) preferred embodiment, details of the internal workings of the image decoder, particularly the pixel color processing engine, are shown. [Figure 4] The partial images, Figures 4a and 4b, illustrate two possible variations of display adaptation to obtain the final display-adapted luminance mapping function FL_DA, which is used to compute the optimally display-adapted version of the input image for a given display capability (PL_D), from a reference luminance mapping function F_L that systematizes the luminance regrading needs between two reference images. [Figure 5] This is a more general summary of the display adaptation principle in order to make the formulation of this embodiment and its components in the claims easier to understand. [Figure 6] A receiver (particularly a color processing circuit) according to one embodiment of the present invention is shown. [Figure 7] In particular, we elucidate the distribution of processing between the linear region (which may be advantageous and, if necessary, can be further mapped to a visually and psychologically uniform region within the linear processing) and the perceptual quantizer region, and further embodiments that particularly elucidate advantageous modes for adding a black level b to the display. [Figure 8]Figures 8a, 8b, and 8c are sub-figures that illustrate some of the effects resulting from applying this processing. Figure 8a shows an exemplary image with dark pixel areas, such as the tires under the police car and the interior of a house seen through a window, such as a circular attic window; mid-range pixel areas with average brightness, such as the body of the police car and people; and bright HDR pixel areas, such as the flashing lights on top of the police car. Figure 8b shows these objects and their average brightness, which is spread around the exemplary HDR image brightness range, and also shows the MDR image brightness range obtained by brightness mapping using two different methods, which is twice the same maximum brightness (maximum brightness of the output image ML_V_out). Figure 8c shows the same effect of the two methods using brightness histograms. [Modes for carrying out the invention]
[0140] Figure 6 shows the ambient light level adaptive processing embodied in the color conversion circuit 600. Such circuits are typically found in HDR video decoders of television signal receivers (e.g., broadcast, video distribution over the Internet, etc.). Physically, they may be found in television displays, mobile phones, receiving devices (set-top boxes, etc.). We assume that image pixel colors appear one by one and are encoded in the PQ region with lumens Y' and chromens Cb and Cr (i.e., lumens are calculated from nonlinear red, green, and blue components by applying the inverse EOTF of the perceptual quantizer, and chromens are calculated according to the usual colorimetric definition of video color).
[0141] The luma component Y' (assuming it is the luma of the HDR image, as shown on the leftmost axis in Figure 1) is processed by the luminance mapping circuit (310).
[0142] The metadata input is obtained as the reference luminance mapping function F_L. This is the function to apply when a secondary reference image needs to be generated and when there is no correction for ambient light levels, i.e., the smallest discriminable black on the display.
[0143] The purpose of the luminance mapping function calculation circuit (611) is to determine an appropriate display and environment-adaptive mapping function (FL_DA(t)). Here, time t can usually change over time, not because the display or viewing environment changes, but because the content of the video image changes with respect to its luminance distribution (which usually requires a different luminance regrading function for subsequent images).
[0144] This function FL_DA(t), if specified in the correct region (for example, as an equivalent function on the PQ luma), can be directly applied by the luminance mapping circuit (310) to the input luma Y' to generate the required output luma, i.e., luminance L'_HDR. That is, it is communicated from the luminance mapping function calculation circuit (611) to the luminance mapping circuit (310), and stored in the latter's memory, normally at the beginning of a new image that is color-processed, with the new function stored before that.
[0145] The luminance mapping function calculation circuit (611) requires the following data to perform its calculation: namely, the reference luminance mapping function F_L, the maximum luminance (PL_D) of the connected display, and the black level value (b). The black level value (b) is communicated either as a PQ Luma code value or as luminance in nits convertible to PQ Luma. The latter two come from metadata memory 602, for example, directly from the connected display. The rest of the color processing is as described above in this already explained, merely clarified embodiment (for example, there may be different color processing tracks as this innovation primarily concerns the luminance portion).
[0146] The luminance mapping function calculation circuit (611) needs to calculate a new (adjusted) maximum luminance value (PLA) before executing the display mapping algorithm configured to derive the FL_DA function based on the F_L function. To clarify, this is not the actual maximum value of the display (i.e., PL_D), but merely a value that functions as if it were the PL_D value in a display adaptive algorithm that is otherwise unchanged (as explained in Figures 4 and 5).
[0147] The luminance mapping function calculation circuit (611) calculates the following: PLA=EOTF(inverse_EOTF(PL_D)-b) [Formula 3]
[0148] As mentioned above, preferably, the EOTF to be applied is the perceptual quantizer EOTF (EOTF_PQ). However, other EOTFs also function similarly, especially if they are reasonably uniform in terms of visual psychology, i.e., the k steps of the luma code show approximately as some degree of visual difference, regardless of the starting luma level l to which this step is applied.
[0149] Subsequently, in a preferred embodiment, further processing is performed, namely, repositioning to the black level value b. To achieve this, the display mapping circuit 320 also receives the display black level value b, i.e., for this viewing environment condition. The display mapping circuit 320 may have several processing functions to prepare the pixel color output signal (i.e., display driving color D_C) required for a particular display. A useful embodiment for implementation is to function in the PQ region. That is, the display desires, for example, to obtain R'G'B' values (or Y'CbCr values derived therefrom) of type PQ. Therefore, this component also functions internally in the PQ region. The R'G'B'_HDR values may be in a different region (for example, they may be linear), in which case the display mapping circuit 320 performs the internal conversion to PQ (as described above, using a standardized method specified by SMPTE).
[0150] Perform addition for the following three PQ region color components: R’_out = a * R’ + b; G’_out = a * G’ + b; B’_out = a * B’ + b; [Equation 4]
[0151] Here too, the constant a can be a scaling factor to make the white of the signal displayable as white, but other approaches such as clipping or non-linear equations can also be used (note that since Y’ is a linear combination of the R’G’B’ color components and three weights that sum to 1 in the standard agreed-upon definition, Y’_out = Y’ + b is also available).
[0152] Next, these output signals constitute the D_C output. For example: D_C = {R’_out, G’_out, B’_out}.
[0153] This approach of adding b has the advantage that the darkest color starts from the visibility threshold, but this method is also useful when applying only Equation 3 (i.e., only different white setting adjustments). This is already specifically for enhancing the contrast of an image.
[0154] Figure 7 is even more useful for clarifying different color regions in the process.
[0155] Typically, when color is represented in the PQ domain, the initial part of the processing is in the PQ domain (Note: Color processing, i.e., appropriate color type (e.g., saturated yellow) rather than luminance, is an independently configurable processing path, but it can also be performed in the PQ domain). In Figure 7, a color space conversion circuit 711 is explicitly included. This converts PQ Luma to linear luminance Ylin. This can also be done directly by the luminance mapping circuit 310 function, as other color space conversions can usually be performed internally anyway (the applicant has found it useful to specify the regrading function definition in a display-optimal, visually and psychologically uniform color space, to which all other processing follows), but this aspect is not essential to the present innovation, and therefore, we will not delve into it further to avoid confusion. Assume that processing begins in PQ in subcircuit (or processing unit) 701, that end-to-end is linear between Ylin and L'_HDR for R'G'B'_HDR in subcircuit 702, and that it is converted back to the PQ region in subcircuit 703 for the display mapping circuit 320. Here the required output format is generated, and according to this invention, additional black level correction is performed in the PQ region (those skilled in the art will understand that this principle can be applied to regions other than this preferred embodiment).
[0156] Finally, there is an output connector 699 for communicating the optimized image and its pixel colors. Similar to the image input connector 601, the output connectors physically extend to various technologies available for communicating the image, such as connections to an HDMI® cable or a wireless channel antenna. The metadata input may be a separate input, but advantageously, this metadata is communicated via an internet connection or via a separate data package, such as an HDMI® cable, along with the image color data, i.e., SEI messages embedded in the MPEG video, for example.
[0157] Figure 8 illustrates how this luminance mapping method works and highlights its advantages over existing methods.
[0158] The luminance axis on the left edge of Figure 8b shows the average luminance of various image objects when represented within a 5000-nit HDR image. The geometric structure of the image is shown in Figure 8a. This is the master grading image, i.e., the highest quality image of the scene, which can usually be twice as high as the first reference image, and can be assumed to be the input image (the input image is transmitted, for example, by a broadcaster and received by a television display that performs a color conversion to a 600-nit MDR image optimized for a 600-nit PL_D display. The output image is output to the 600-nit PL_D display).
[0159] The creator of the image (or video / film) may decide to make the flashlight of a police car very bright, such as 4000 nits. The unlit interior of a house, visible from a window (for example, in a night scene), may be made quite dim, illuminated only by streetlights outside. When creating a master image, it is made for the best viewing environment, such as a dimly lit living room of a consumer, where only a small lamp is available. In such circumstances, the viewer can see many different dark colors, so ideally, the master image should contain information for all of these indoor pixels. Thus, the structure of indoor objects is encoded within the master image at an average value of, for example, 2 nits (which may be reduced to as low as 0.01 nits in some cases). The problem then becomes, what if a neighbor watches the same movie, i.e., the same communicated 5000 master HDR image, containing the same dark pixels, in a very bright room, possibly in daylight? This neighbor may not be able to see the darkest pixels, but in some cases, they may not see anything darker than 7 nits.
[0160] The right side of Figure 8b shows how the received reference image is adapted to the MDR image. This MDR image not only has a lower maximum brightness, but more importantly, in this teaching, the black level value b is higher, for example, 7 nits, even though the MDR image has the same maximum value of 5000 nits.
[0161] Method METH1 is the method described above using Equation 2, and is a method by which a person skilled in the art can usually solve the problem. That is, it is determined that there are some invisible dark pixels, and therefore all of them are within the visible range above black level b.
[0162] How this works is illustrated by the 20-nit window and, furthermore, by the first histogram 810 in Figure 8c. The shift to 20 nits is somewhat large, resulting in suboptimal visual contrast for the resulting MDR image. Furthermore, this correction method extends the offset correction for dark levels considerably further, actually across the entire range, and therefore, for example, human mid-luminances are also shifted significantly to brighter values (which may also have a suboptimal effect on these areas (i.e., the overall contrast of the image)). This can be seen from the shift of the histogram lobes for dark and mid-luminance objects in the histogram of the conventional method relative to the second histogram 820. The second histogram 820, along with the luminance axis labeled METH2, represents the novel approach.
[0163] Not only do images generally look better with more realistic contrast, but applying this method to the display adaptation process allows for a more realistic adherence to the filmmaker's desiderata, specifically as represented by the reference regrading luminance mapping function F_L. For example, a grader could specify this function, and thus the darkest luminance subrange, e.g., indoor pixels, would be significantly boosted when mapped to a lower maximum luminance. Meanwhile, the local slope of the mapping in the HDR luminance subrange where the car's luminance falls could remain low. The grader could then indicate that the interior of the house is important in this story (compared to a scenario where the interior is not important and is simply buried in invisible black). Thus, each time luminance mapping is performed on a maximum luminance image that gradually decreases, a normalized subrange of gradually increasing available luminance is used for the darkest indoor pixels, resulting in a relatively good view even on dimly lit displays (in this case, also relatively good on brighter displays, but with lower visibility of dark pixels). This approach can adhere to these technical desiderata as accurately as possible. The ideal mapping to available conditions is to avoid redistributing 0 nits to PL_Dnit, but to redistribute the actually available luminance b nits to PL_Dnit, and then to use this range communicated in the F_L function as best as possible, as intended by the content creator. This is usually satisfactory for the viewer. This is at least an elegant and not difficult practical method.
[0164] Numerical examples illustrating the nonlinear aspects of EOTF_PQ are shown. The lowest luma corresponds to very dim luminances up to fairly high values. Therefore, relatively bright dim levels usually correspond to relatively high luma values. Consequently, the PLA value is significantly lower than the PL_D value, but it is still optimally used in this method. For example, let PL_D = 1000 nits and EOTF_PQ(b) = 0.3 nits. In this case, b as PQ luma is 0.097. *1023 = 99. Inverse_EOTF(1000) is 0.75. * 1023 = 767. The difference between these two is 668.
[0165] The question here is which white level, adjusted to control display adaptation, corresponds to this luma level.
[0166] EOTF_PQ(668 / 1023) = EOTF_PQ(0.65) = 408 nits. Therefore, the PLA value (quantified as minimum visibility on the display screen) of the brightness of such a display and viewing environment is 408 nits. This value is input to one of the embodiment of display adaptation used by the luminance mapping function calculation circuit (611) to generate the correct display-adapted luminance mapping function FL_DA. This function is optimized not only for the display but also for the viewing environment.
[0167] The algorithmic components disclosed in this text may actually be implemented (in whole or in part) as hardware (e.g., as part of an application-specific IC) or as software running on a specialized digital signal processor or general-purpose processor.
[0168] Those skilled in the art should be able to understand from this description which components are optional improvements and can be realized in combination with other components, and how the (optional) steps of the method correspond to each means of the apparatus, and vice versa. The word "apparatus" in this application is used in its broadest sense, that is, a group of means that enable the achievement of a particular purpose, and therefore may be, for example, an IC or a dedicated device (e.g., a device with a display) (a small circuit part of it), or part of a networked system. "Arrangement" is also intended to be used in its broadest sense, and may include, in particular, a single apparatus, a part of an apparatus, a collection of cooperating apparatuses (or parts of them).
[0169] The explicit meaning of a computer program product should be understood as encompassing any physical embodiment of a set of commands that, after a series of loading steps (which may include intermediate translation steps such as translation into an intermediate language and a final processor language), enable a general-purpose or special-purpose processor to input commands to the processor and execute any of the characteristic functions of the present invention. In particular, a computer program product can be realized, for example, as data on a carrier such as a disk or tape, data residing in memory, data via a network connection (wired or wireless), or program code written on paper. Apart from the program code, characteristic data required for a program can also be realized as a computer program product.
[0170] Some of the steps required to operate a method, such as data input and data output steps, may already exist in the processor's functionality rather than being described in the computer program product.
[0171] The embodiments described above are illustrative and not limiting to the present invention. For the sake of brevity, if a person skilled in the art can easily map the presented embodiments to other areas of the claims, all of these options have not been explored in detail. Other combinations of elements are possible besides those combined in the claims. Any combination of elements can be realized in a single dedicated element.
[0172] Any reference symbols between parentheses in the claims are not intended to limit the claims. The word “contains” does not preclude the existence of elements or aspects not enumerated in the claims. A singular element does not preclude the existence of multiple elements.
Claims
1. A method for processing an input image and obtaining an output image, The input image has pixels having input brightness within a first brightness dynamic range, and the first brightness dynamic range has a first maximum brightness. The reference luminance mapping function is received as metadata associated with the input image. The aforementioned reference luminance mapping function specifies the relationship between the luminance of the first reference image and the luminance of the second reference image, The first reference image has a first reference maximum brightness, and the second reference image has a second reference maximum brightness. The input image is equal to one of the first reference image and the second reference image, The output image is not equal to either the first reference image or the second reference image. The process includes applying the adapted luminance mapping function to the input pixel luminance to obtain the output luminance, The adapted luminance mapping function is calculated based on the reference luminance mapping function and a maximum luminance value as a control parameter specifying the degree to which the adapted luminance mapping function deviates from the reference luminance mapping function. The calculation of the adapted luminance mapping function includes finding the position on the metric corresponding to the maximum luminance value, The first endpoint of the metric corresponds to the first maximum brightness, and the second endpoint of the metric corresponds to the maximum brightness of one of the first reference image and the second reference image that is not equal to the input image. The maximum brightness value is calculated based on the maximum brightness of the display to which the output image is supplied and the black level value of the display. The calculation method includes applying the inverse function of the electric-optical transfer function to the maximum luminance, subtracting the black level value from the resulting value, and applying the electric-optical transfer function to the subtraction to obtain the maximum luminance value used in the calculation of the adapted luminance mapping function.
2. The method according to claim 1, wherein a perceptual quantizer function is used as the electric-optical transfer function.
3. The method according to claim 1 or 2, wherein the black level value is added to the output luminance, which is represented as luma in the luma region of the perceptual quantizer.
4. An input connector that receives an input image, An output connector that outputs an output image calculated by a color conversion circuit based on the input image, A metadata input unit that receives a reference luminance mapping function, which is metadata associated with the input image, The color conversion circuit comprises, The input image has pixels having input brightness within a first brightness dynamic range, and the first brightness dynamic range has a first maximum brightness. The aforementioned reference luminance mapping function specifies the relationship between the luminance of the first reference image and the luminance of the second reference image, The first reference image has a first reference maximum brightness, and the second reference image has a second reference maximum brightness. The input image is equal to one of the first reference image and the second reference image, The output image is not equal to either the first reference image or the second reference image. In a color conversion circuit that includes a luminance mapping circuit that applies an adapted luminance mapping function to the input pixel luminance to obtain the output luminance, The color conversion circuit includes a luminance mapping function calculation circuit that calculates the adapted luminance mapping function based on the reference luminance mapping function and a maximum luminance value as a control parameter that specifies the degree to which the adapted luminance mapping function deviates from the reference luminance mapping function. The calculation includes finding the position on the metric corresponding to the maximum brightness value, The first endpoint of the metric corresponds to the first maximum brightness, and the second endpoint of the metric corresponds to the maximum brightness of one of the first reference image and the second reference image that is not equal to the input image. The luminance mapping function calculation circuit receives the maximum luminance of the display to which the output image is supplied and the black level value of the display, The luminance mapping function calculation circuit is a color conversion circuit characterized by calculating the maximum luminance value by applying the inverse function of the electric-optical transfer function to the maximum luminance, subtracting the black level value from the resulting value, and applying the electric-optical transfer function to subtraction to obtain the maximum luminance value.
5. The color conversion circuit according to claim 4, wherein the luminance mapping function calculation circuit applies a perceptual quantizer function as the electro-optical transfer function.
6. The color conversion circuit according to claim 4 or 5, further comprising a display mapping circuit that adds the black level value to the lumens value of the output luminance represented in the lumens region of the perceptual quantizer in the lumens region of the perceptual quantizer.
Citation Information
Patent Citations
Apparatus and method for converting the dynamic range of an image
JP2014531821A
Display mapping of high dynamic range images on power-limited displays
JP2022500969A
Apparatus and method for dynamic range transforming of images
US20140225941A1
Ambient light-adaptive display management
US20190304379A1
Display mapping for high dynamic range images on power-limiting displays
US20220044615A1