Display-optimized ambient HDR video adaptation

The method of creating reference gradings for HDR and SDR images, combined with display adaptation algorithms, addresses the challenge of adapting HDR images to varying display conditions, ensuring optimal image quality on both high and low dynamic range displays.

JP7809138B2Active Publication Date: 2026-01-30KONINKLIJKE PHILIPS NV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023568027
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-07
Filing Date
2022-04-29
Publication Date
2026-01-30
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

Existing video technologies struggle to adapt high dynamic range (HDR) images for display conditions that vary in ambient lighting, particularly when lower dynamic range displays are used, as they lack effective methods to optimize pixel brightness and maintain image quality across different display capabilities.

Method used

A method and apparatus for adapting HDR video by creating two reference gradings, one for HDR and one for SDR, allowing for display adaptation algorithms to transform pixel brightness based on metadata and luminance mapping functions, ensuring optimal image quality on both high and low dynamic range displays.

Benefits of technology

Enables efficient conversion of HDR images to SDR images, maintaining image quality and adapting to varying display capabilities, providing a better viewing experience across different ambient lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007809138000001
    Figure 0007809138000001
  • Figure 0007809138000002
    Figure 0007809138000002
  • Figure 0007809138000003
    Figure 0007809138000003
Patent Text Reader

Abstract

In order to obtain better viewable images for a variety of potentially significantly different viewing environment light levels in a practical manner, the inventors have proposed a method of processing an input image to obtain an output image, the method comprising the steps of obtaining a starting luma for a pixel of the input image by applying a photoelectric transfer function OETF_psy to an input luminance L_in of the input image; obtaining a target display minimum luminance mL_VD, the target display corresponding to an end-user display 210, to which the output image may be provided for displaying the output image; obtaining an end-user display minimum luminance mL_De in a viewing room, the end-user display minimum luminance mL_De depending on the amount of illumination in the viewing room; calculating a difference dif by subtracting the end-user display luminance mL_De from the target display minimum luminance mL_VD; and calculating the difference dif as a luma starting luma by applying the difference to the photoelectric transfer function OETF_psy as an input to the photoelectric transfer function OETF_psy. a transforming step of transforming the starting luma into a difference Ydif, thereby resulting in a luma difference as output; a mapping step of mapping the starting luma by applying a linear function that applies the luma difference multiplied by -1.0 as an additive constant and multiplies the starting luma using the luma difference incremented by a value of 1.0 as a multiplier, thereby resulting in a mapped luma Yim; and a transforming step of transforming the mapped luma Yim by applying the inverse of the photoelectric transfer function to obtain a normalized mid-luminance Lim. The method includes the steps of subtracting the second minimum luminance mL_De2 of the end user display divided by the maximum luminance PL_O of the output image from the intermediate luminance PL_O to obtain a final normalized luminance Ln_f, and scaling the subtraction by the result of subtracting the second minimum luminance mL_De2 of the end user display divided by the maximum luminance PL_O of the output image from 1.0 to obtain an output luminance, multiplying the final normalized luminance Ln_f by the maximum luminance PL_O of the output image to obtain an output luminance, and outputting the output luminance in a color representation of a pixel of the output image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for adapting image pixel brightness of high dynamic range video to provide a desired appearance for HDR video display conditions under the ambient lighting conditions of a particular viewing site. [Background technology]

[0002] A few years ago, a novel technique for high dynamic range (HDR) video coding was introduced, inter alia, by the applicant (see, for example, WO2017157977).

[0003] Video coding is generally primarily or solely concerned with creating or more precisely defining color codes (e.g., luma and two chromas per pixel) to represent an image, which is distinct from knowing how to optimally display an HDR image (e.g., the simplest method simply utilizes a highly nonlinear electro-optical transfer function (OETF) to convert the desired luminance into, for example, a 10-bit luma code, and vice versa, converting those video pixel luma codes into the luminance to be displayed by using an inversely formed electro-optical transfer function (EOTF) to map the 10-bit electrical luma code to the optical pixel luminance to be displayed, but more complex systems deviate in several directions, particularly by decoupling the coding of an image from the specific use of the coded image).

[0004] Encoding and handling HDR video stands in stark contrast to the way older video technologies were used, by which all video was coded until recently, and which today is called standard dynamic range (SDR) video coding (also known as low dynamic range video coding, or LDR). SDR began in the analog era as PAL or NTSC and transitioned to Rec. 709-based coding, e.g., MPEG-2 compression, in the digital video era.

[0005] While a satisfactory technology for communicating moving images in the 20th century, advances in display technology beyond the physical limitations of the electron beam of 20th century CRTs or the world-famous TL backlit LCDs have made it possible to show images with significantly brighter (and potentially even darker) pixels than older displays, which has necessitated the need to be able to encode and create such HDR images.

[0006] In fact, starting from the inability to encode very bright and sometimes even darker image objects in the SDR standard (8-bit Rec. 709) for various reasons, ways of technically representing colors with such a wide luminance range had to be invented first, and then, one by one, all the rules of video technology had to be rethought and, in many cases, reinvented.

[0007] The Rec. 709 SDR luma code definition could only encode (at 8 or 10-bit luma) an approximately 1000:1 luminance dynamic range due to the approximately square-root OETF function shape luma Y_code=power(2,N)*sqrt(L_norm), where N is the number of bits in the luma channel and L_norm is the normalized version of physical luminance between 0 and 1.

[0008] Furthermore, in the SDR era, the absolute luminance to be displayed was not defined, so in practice, the maximum relative luminance L_norm_max=100%, or 1, was mapped via the square-root OETF to the maximum normalized luma code Yn=1, corresponding to, for example, Y_code_max=255. This has some technical differences compared to creating absolute HDR images: an image pixel coded to display as 200 nits would ideally (i.e., where possible) be displayed as 200 nits on all displays, not at a completely different display luminance. In the relative paradigm, a 200-nit coded pixel luminance would be displayed as 300 nits on a brighter display, i.e., a display with a brighter maximum displayable luminance PL_D (a.k.a., display maximum luminance), and would be displayed as 100 nits on a display with less capability, for example. Furthermore, note that absolute coding works with normalized luminance representations or normalized 3D color gamuts, but 1.0, for example, uniquely means 1000 nits.

[0009] On displays, such relative images were typically displayed somewhat heuristically by mapping the brightest luminance in the video to the brightest displayable pixel luminance (which was done automatically via the electronic driving of the display panel by the maximum luma Y_code_max, without the need for further luminance mapping). So if you bought a 200 nit PL_D display, white would appear twice as bright as on a 100 nit PL_D display, but given factors such as eye adaptation, that wasn't considered to be of much importance other than making the same SDR video image a brighter, more viewable, and somewhat more beautiful version.

[0010] Conventionally, when we talk about an SDR video image today (in an absolute framework), it generally has a video peak luminance of PL_V=100 nits (e.g., as agreed upon in accordance with standards), and so in this application we consider the maximum luminance of an SDR image (or SDR grading) to be exactly that value or generalized around that value.

[0011] Grading in this application is intended to mean either an activity or a resulting image in which desired pixels are given brightness, for example, by a human color grader or an automaton. When viewing an image, for example, when designing an image, there are several image objects, and ideally, one wants to give the pixels of those objects brightnesses that are spread around an average brightness that is optimal for that object, also taking into account the overall image and scene. For example, if the image capacity available is such that the brightest encodable pixel of the image is 1000 nits (image or video maximum brightness PL_V), one grader might choose a brightness value between 800 nits and 1000 nits for the explosion pixels to make the explosion look punchy, while another filmmaker might choose an explosion that is no brighter than 500 nits, for example, so that it does not interfere too much with the rest of the image at that moment (of course, technology can handle both situations).

[0012] The maximum brightness of an HDR image or video varies considerably and is generally communicated with the image data as metadata for the HDR video or image (typical values ​​are, for example, 1000 nits, or 4000 nits, or 10,000 nits, but are not limiting; generally, one is said to have an HDR image when the PL_V is at least 600 nits). If a video creator chooses to define an image as PL_V=4000 nits, the video creator can of course choose to create a brighter burst, but relatively, it will not reach the 100% level of PL_V, but rather, for example, only reach 50% for such a high PL_V definition of the scene.

[0013] An HDR display has a maximum capability, i.e., a highest displayable pixel brightness that is, for example, 600 nits, or 1000 nits, or N times 1000 nits (starting with the lowest HDR display). The maximum (or peak) brightness of the display, PL_D, is separate from the maximum brightness of the video, PL_V, and the two should not be confused. Video creators generally cannot create optimal video for each possible end-user display (i.e., the capabilities of the end-user display are optimally used by the video, and the maximum brightness of the video never exceeds the maximum brightness of the display (ideally), but is not lower either; i.e., some of the video images should have at least some pixels with pixel brightness L_p=PL_V, which, with successive optimizations to specific displays, will further include PL_V=PL_D).

[0014] Creators make some of their own decisions (e.g., what kind of content to capture and how) and generally make videos with a PL_V very high to serve the highest PL_D displays of their intended audience, at least now, and perhaps in the future as higher PL_D displays emerge.

[0015] A secondary problem then arises: how best to display an image with a peak luminance PL_V on a display with a lower (often much lower) display peak luminance PL_D, which is called display adaptation. Even in the future, there will still be displays that require a lower dynamic range image received over a communications medium than the created, say, 2000 nit PL_V image. In theory, a display will always re-grade, or map, the luminance of image pixels so that they are displayable according to its own internal heuristics, but if the video creator is careful enough in determining pixel luminance, it is beneficial for the video creator to also indicate how the image should be display-adapted to the lower PL_D value, and ideally, for the display to follow to a large extent what is technically required.

[0016] The situation is more complicated when it comes to the darkest displayable pixel luminance, BL_D. While some of this is due to fixed physical characteristics of the display, such as LCD cell leakage, even with the best displays, what a viewer ultimately perceives as distinct darkest blacks also depends on the viewing room illumination, which is not a well-defined value. This illumination can be characterized as an average illuminance level, for example, in lux, but for video display purposes, it is more conveniently characterized as minimum pixel luminance. This is also generally more relevant to the human eye than the perception of bright or medium luminance. This is because when the human eye sees many high-brightness pixels, darker pixels, especially their absolute luminance, become less relevant. However, for example, when viewing a generally dark scene image that is still masked by ambient light in front of the display screen, we can assume that the eye is not a limiting factor. If we assume that humans can see a noticeable difference of just 2%, there is a darkest drive level (or luma), b, above which the next darker luma level (i.e., displaying a luminance level that is X% higher, e.g., 2% more) can still be seen.

[0017] In the LDR era, the darkest pixels were of no concern at all. They were primarily concerned with an average brightness of about 1 / 4 of the maximum PL_V=100nit. If an image was exposed near this value, everything in the scene looked nice and bright and colorful, except for clipping in the bright parts of the scene that exceeded the 100% maximum. If the darkest parts of a scene were important enough, the captured image was created with a sufficient amount of base lighting in a recording studio or filming environment. If part of the scene was not visible well, for example, it was buried at code Y=0, so it was considered normal.

[0018] Therefore, if nothing further is specified, we assume that the blackest black is zero, or in practice something like 0.1 nit or 0.01 nit. In such situations, engineers are more interested in pixels that are brighter than average within the encoded and / or displayed HDR image.

[0019] In terms of coding, the differences between HDR and SDR are not only physical differences (more different pixel brightnesses can be displayed on a display with a larger dynamic range capability), but also technical differences, including different luma code allocation functions (which use the OETF, or in absolute terms the inverse of the EOTF), and potentially even further technical HDR concepts, such as additional dynamically changing (for each image, or for each set of temporally consecutive images) metadata that specifies how to re-grade the pixel brightnesses of various image objects to obtain images with a secondary dynamic range that differs from the starting image dynamic range (the two brightness ranges generally end with peak brightnesses that differ by at least a factor of 1.5).

[0020] A simple HDR codec, the HDR10 codec, has been introduced to the market and is used, for example, to create the recently released Black Jewel Box HDR Blu-ray. This HDR10 video codec uses a logarithmic function rather than a square root as the OETF (inverse OETF), i.e., the so-called perceptual quantizer (PQ) function standardized in SMPTE 2084. Instead of being limited to 1000:1 as in the Rec. 709 OETF, this PQ OETF allows for the definition of luma over a much larger range (ideally, the desired display brightness) that is sufficient for practical HDR video production: between 1 / 10,000 nit and 10,000 nit.

[0021] The reader should note that HDR should not be confused simply with a large number of bits in the luma codeword. That applies to linear systems, such as the amount of bits in an analog-to-digital converter, and in fact the amount of bits follows the base-2 logarithm of the dynamic range. However, since the code allocation function has a fairly nonlinear shape, theoretically one could define HDR images with only 10-bit luma (and even 8-bit HDR images per color component) however one wanted, which would offer the advantage of reusing systems already deployed (e.g., ICs may have a specific bit depth, or video cables, etc.).

[0022] After luma calculation, we have 10 bit planes of pixel luma Y_code, to which two chrominance components Cb and Cr per pixel are added as chrominance pixel planes. This image is further processed mathematically classically "as if" it were an SDR image, e.g., MPEG-HEVC compressed. The compressor does not actually need to care about pixel color or luminance.

[0023] However, the receiving device, e.g., a display (or indeed a decoder thereof), generally needs to perform a correct color interpretation of the {Y,Cb,Cr} pixel colors to display a correct looking image, rather than, e.g., an image with washed out colors.

[0024] This is typically handled by communicating the three pixelated color component planes along with additional image definition metadata that defines the image encoding, such as an indication of which EOTF is used (for which, without limitation, we assume that the PQ EOTF (or OETF) is used), and the PL_V value.

[0025] More sophisticated codecs include further image defining metadata, e.g., handling metadata, e.g., functions that specify how to map a normalized version of the luminance of a first image up to PL_V=1000 nit to the normalized luminance of a secondary reference image, e.g., an SDR reference image with PL_V=100 nit (as explained in more detail in Figure 2).

[0026] For the convenience of readers less knowledgeable about HDR, some interesting aspects are briefly described in Figure 1. Figure 1 shows some typical illustrative examples of the many possible HDR scenes that a future HDR system (e.g., connected to a 1000 nit PL_D display) will need to be able to process correctly. While the actual technical processing of pixel colors is done in different ways in different color space definitions, what is required for re-grading is shown as an absolute luminance mapping between luminance axes spanning different dynamic ranges.

[0027] For example, ImSCN1 is a sunny outdoor image from a Western movie with mostly bright areas. First, it should be noted that the pixel brightness of any image is generally not a brightness that can actually be measured in the real world.

[0028] Even without further human intervention during the creation of the output HDR image (which serves as a starter image and is called the master HDR grading or image), by tweaking one parameter, however simple, the camera will always at least measure the relative luminance at the image sensor because of the aperture, so there is always some step involved where at least the brightest image pixels end up in the available encoded luminance range of the master HDR image.

[0029] For example, the specular reflection of the sun on a sheriff's star badge might be measured at over 100,000 nits in the real world, which would be neither visible on a typical near-future display nor comfortable for a viewer watching a movie image in, say, a dimly lit room in the evening. Alternatively, if a video creator determines that 5000 nits is bright enough for a pixel on the badge and therefore intends this pixel to be the brightest pixel in the movie, the video creator decides to create a video with PL_V=5000 nits. While a relative pixel brightness measurement device is only for the RAW version of the master HDR grade, the camera should also have a sufficiently high native dynamic range (all pixels well above the noise floor) to produce a good image. The pixels of the graded 5000-nit image are generally derived nonlinearly from the RAW image captured by the camera; for example, the color grader takes into account aspects of the actual shooting location, such as typical viewing conditions that are not the same as standing in a hot desert. The best (highest PL_V) image selected to produce this scene, ImSCN1, i.e., the 5000 nit image in this example, is the master HDR grading. This is the minimum required HDR data to be created and communicated, but in all codecs it is not the only data communicated, or in some codecs it is not even an image communicated at all.

[0030] Making such a codable high luminance range DR_1 available, for example between 0.001 nit and 5000 nit, allows content makers to provide viewers with a better experience of brighter looking scenes, but also of course a better experience of dimmer night scenes (when well graded throughout the movie), assuming the viewer also has a corresponding high-end PL_D=5000 nit display. A good HDR movie balances the luminance of various image objects, not just in a single image, but also over time in the movie's story or in the video material generally produced (e.g., a well-designed HDR soccer program).

[0031] The leftmost vertical axis in Figure 1 shows some (average) object luminances that one would like to see in a 5000 nit PL_V master HDR grading, ideally intended for a 5000 nit PL_D display. For example, in a movie, one creator might want to show cowboys in bright sunshine with a pixel luminance of about 500 nits (i.e., generally 10 times brighter than LDR, while another creator might want slightly less HDR punch, say 300 nits), thereby configuring the best way to display this Western image by the creator, which will give the best possible look to the end consumer.

[0032] The need for a higher dynamic range of luminance is more easily understood by considering an image in which there are fairly dark regions, such as the dark corners of the cave image ImSCN3, but also relatively large areas of very bright pixels, such as the sunlit outside world seen from the cave entrance, which creates a different visual experience than, for example, the nighttime image ImSCN2, where only street lights contain areas of high luminance pixels.

[0033] Now, the problem is that at this time, many consumers still have LDR displays, and even in the future there will be good reasons to make two grades of a movie instead of the typical encoding of only the HDR image itself, so we need to be able to define an SDR image with PL_V_SDR=100 nits that best corresponds to the master HDR image. This is a technical desire, and it is separate from the technical choice regarding the encoding itself; it states, for example, that if we know how to create (invert) the master HDR image and one of these secondary images from the other, we can choose to encode and communicate only one of the pair (effectively communicating two images for the price of one, i.e., communicating only one image of the pixel color component planes per video time instant).

[0034] Of course, with such a reduced dynamic range image, it is not possible to define a 5000 nit pixel brightness object like the bright sun in reality. The minimum pixel brightness or deepest black is also as high as 0.1 nit, rather than the more preferred 0.001 nit.

[0035] So anyway, it should be possible to create this corresponding SDR image with a reduced luminance dynamic range DR_2.

[0036] This is done by some automatic algorithm in the receiving display, for example using a fixed luminance mapping function, or perhaps one conditioned by simple metadata like the PL_V_HDR value and potentially one or more other luminance values.

[0037] However, while more complex luminance mapping algorithms may generally be used, in this application, without loss of generality, we assume that the mapping is defined by a global luminance mapping function F_L (e.g., one function per image) that defines, for at least one image, how all possible luminances (i.e., for example, 0.0001 to 5000) occurring in the first image should be mapped to corresponding luminances in the second output image (e.g., 0.1 to 100 nits for an SDR output image). The normalization function is obtained by dividing the luminance along both axes by their respective maximum values. Global in this context means that the same function is used for all pixels of the image, regardless of further conditions, such as their position within the image (more general algorithms, for example, use several functions for pixels that can be classified according to some criteria).

[0038] Ideally, how all luminance should be redistributed along the available range of the SDR image of the secondary images should be decided by the video creator, since, when limited, the video creator knows well how to sub-optimize for the reduced dynamic range so that the SDR image still looks at least as good as possible like the intended master HDR image. The reader can understand that actually defining (locating) such object luminance corresponds to defining the shape of the luminance mapping function F_L, the details of which are beyond the scope of this application.

[0039] Ideally, the shape of the function should also change for different scenes, i.e., a cave scene in a movie versus a sunny western scene a little later, or generally for different images in time: this is called dynamic metadata (F_L(t), where t denotes image time).

[0040] Now, ideally, the content creator would create an optimal image for each situation, i.e., for each potentially served end-user display, e.g., a PL_D_MDR=800 nit display requiring a corresponding PL_V_MDR=800 nit image, but that is generally too much effort for the content creator, even in the most expensive offline video production.

[0041] However, it has previously been demonstrated by the applicant that it is sufficient to create (only) two different dynamic range reference gradings of a scene (generally at the extreme ends, e.g., 5000 nits being the highest required PL_V and 100 nits being generally sufficient as the lowest required PL_V). This is because all other gradings can then be derived automatically from those two reference gradings (HDR and SDR), e.g., via a (generally fixed, e.g., standardized) display adaptation algorithm applied to the end-user's display receiving the information of the two gradings. Generally, the calculations are performed in any video receiver, e.g., a set-top box, a television, a computer, a cinema device, etc. The communication channel for the HDR images can also be any communication technology, e.g., terrestrial or cable broadcasting, a physical medium such as a Blu-ray disc, the Internet, a communication channel to a portable device, professional inter-site video communication, etc.

[0042] This display adaptation generally also applies a luminance mapping function to, for example, the pixel luminance of the master HDR image. However, the display adaptation algorithm needs to determine a luminance mapping function, i.e., a display adaptation luminance mapping function FL_DA, that is different from F_L_5000to100 (which is the reference luminance mapping function connecting the luminances of the two reference gradings), which is not necessarily trivially related to the original mapping function between the two reference gradings F_L (there are several variants of the display adaptation algorithm). The luminance mapping function between the master luminance defined by the 5000 nit PL_V dynamic range and the 800 nit intermediate dynamic range is written as F_L_5000to800 in this text.

[0043] Rather than mapping to where the F_L_5000to100 function would "naively" expect across the 800 nit MDR image luminance range, for example, we've indicated the display adaptation symbolically (for only one of the average object pixel luminances) by an arrow, mapping to a slightly higher location (i.e., in such an image, the cowboy should be slightly brighter, at least according to the chosen display adaptation algorithm). So, while some more complex display adaptation algorithms may place the cowboy at the indicated higher location, some customers are satisfied with the simpler location where the connection between the 500 nit HDR cowboy and the 18 nit SDR cowboy crosses the 800 nit PL_V luminance range.

[0044] In general, a display adaptation algorithm calculates the shape of the display adaptation luminance mapping function FL_DA based on the shape of the original luminance mapping function F_L (or reference luminance mapping function, also known as reference regrading function).

[0045] While this description based on Figure 1 constitutes the technical requirements of any HDR video encoding and / or processing system, Figure 2 illustrates some exemplary technical systems and their components for realizing the requirements (non-limiting) in accordance with Applicant's codec approach. Those skilled in the art will understand that these components may be embodied in a variety of devices, etc. Those skilled in the art will understand that this example is presented merely as a representative portion of various HDR codec frameworks to provide a background understanding of some principles of operation, and is not intended to specifically limit any of the embodiments of the innovative contributions presented below.

[0046] Although possible, the technical communication of two actual different images at each time (HDR and SDR gradings each communicated as three respective color planes) is expensive, especially in terms of the amount of data required.

[0047] Also, it is not necessary, since if it is known that all corresponding secondary image pixel intensities are calculated based on the intensities of the primary image and function F_L, it can decide to communicate only the primary image and function F_L at each time instant as metadata (and can choose to communicate either the master HDR or SDR image as representative of both).The receiver knows the (generally fixed) display adaptation algorithm, so it determines the FL_DA function at its end based on this data (additional metadata that controls or guides the display adaptation may be communicated, but is not currently in place).

[0048] There are two modes for communicating a unique image and function F_L at each time.

[0049] In a first backwards compatible mode, an SDR image is communicated ("SDR communication mode"). The SDR image can be displayed directly (without further luminance mapping) on ​​a legacy SDR display, but an HDR display must apply the F_L or FL_DA function to obtain an HDR image from the SDR image (or vice versa, depending on which variant of the function is communicated, i.e., upgrading or downgrading). The interested reader can find full details of Applicant's exemplary standardized first mode approach below.

[0050] ETSI TS 103 433-1 V1.2.1 (2017-08): High-Performance Single Layer High Dynamic Range System for use in Consumer Electronics devices; Part 1: Directly Standard Dynamic Range (SDR) Compatible HDR System (SL-HDR1).

[0051] Another mode communicates the master HDR image itself ("HDR communication mode"), i.e., for example, a 5000 nit image, and a function F_L that allows computing from it a 100 nit SDR image (or any other lower dynamic range image via display adaptation). The master HDR communication image itself is encoded, for example, by using the PQ EOTF.

[0052] Figure 2 further illustrates the overall video communication system. On the transmitting side, it starts with the source of the image 201. This can be anything from a hard disk, to a cable output from, for example, a television studio, etc., depending on whether it is an offline produced video from, for example, an internet distribution company, or a real broadcast.

[0053] This results in a master HDR video (MAST_HDR) that has been color graded, for example, by a human color grader, a shaded version of the camera capture, or by an automatic brightness redistribution algorithm, etc.

[0054] In addition to the grading of the master HDR image, a set of reversible color transformation functions F_ct is often defined. Without loss of generality, we assume that this includes at least one luminance mapping function F_L (however, there may be further functions and data that specify, for example, how the saturation of pixels should change from HDR to SDR grading).

[0055] This luminance mapping function, as described above, defines the mapping between the HDR reference grading and the SDR reference grading (the latter in FIG. 2 is the SDR image Im_SDR to be communicated to the receiver; it may or may not have been data compressed, e.g., via MPEG or other image compression algorithms).

[0056] The color mapping of color converter 220 should not be confused with that applied to the raw camera feed to obtain the master HDR video, which is assumed here to be already input, since this color conversion is to obtain the image to be communicated, and at the same time, to obtain what is needed for re-grading as technically formulated in the luminance mapping function F_L.

[0057] In an exemplary SDR communication type (i.e., SDR communication mode), the master HDR image is input to a color converter 202 configured to apply an F_L luminance mapping to the luminance of the master HDR image (MAST_HDR) to obtain all corresponding luminances that are written to the output image Im_SDR. For purposes of illustration, let's assume that the shape of this function is fine-tuned for each shot of images of similar scenes in a movie by a human color grader using color grading software. The applied function F_ct (i.e., at least F_L) is written into (dynamic, processing) metadata to be co-communicated with the image, such as the exemplary MPEG Supplemental Enhancement Information data SEI(F_ct), or into a similar metadata mechanism in other standardized or non-standardized communication methods.

[0058] After properly redefining the HDR images to be communicated as corresponding SDR images Im_SDR, they are often compressed (at least, e.g., for broadcast to end users) using existing image compression techniques (e.g., MPEG HEVC, VVC, or AV1, etc.). This is performed in a video compressor 203 that forms part of the video encoder 221 (and which may also be included in various forms of video production devices or systems).

[0059] The compressed image Im_COD is transmitted to at least one receiver by some image communication medium 205 (e.g., satellite, cable, or internet transmission according to, for example, ATSC3.0, or DVB, etc.; however, the HDR video signal may also be communicated by cable, for example, between two video processing devices).

[0060] Typically, prior to communication, further conversion is performed by a transmit formatter 204, which applies techniques such as packetization, modulation, transmission protocol control, etc. depending on the system, which typically applies integrated circuits.

[0061] At the receiving site, a corresponding video signal unformatter 206 applies the necessary unformatting methods, eg, demodulation, etc., to recapture, for example, a set of compressed HEVC images (ie, HEVC image data).

[0062] The video decompressor 207 performs, for example, HEVC decompression to obtain a stream of pixelated decompressed images Im_USDR, which are SDR images in this example but HDR images in other modes. The video decompressor also unpacks the required luminance mapping function F_L, or in general the color transformation function F_ct, for example from the SEI message.

[0063] The image and function are input to a (decoder) color converter 208, which is configured to convert the SDR image into an image with a non-SDR dynamic range (i.e., a PL_V higher than 100 nits, typically at least several times higher, e.g., 5 times higher).

[0064] For example, a 5000 nit reconstructed HDR image Im_RHDR is reconstructed as very close to the master HDR image (MAST_HDR) by applying the inverse color transform IF_ct of the color transform F_ct used on the encoding side to create Im_LDR from MAST_HDR. This image is then sent to the display 210, for example, for further display adaptation, but creating the display-adapted image Im_DA_MDR is also done in one go during decoding by using the FL_DA function (determined in an offline loop, for example, in firmware) instead of the F_L function in the color converter. Therefore, the color converter further includes a display adaptation unit 209 to derive the FL_DA function.

[0065] The optimized, e.g., 800 nit, display-adaptive image Im_DA_MDR is sent to, e.g., the display 210 if the video decoder 220 is included in, e.g., a set-top box or a computer, or is sent to a display panel if the decoder is in, e.g., a mobile phone, or is communicated to a movie theater projector if the decoder is in, e.g., an internet-connected server, etc.

[0066] FIG. 3 shows a useful variation of the internal processing of a color converter 300 of an HDR decoder (or encoder, which generally has the same topology but uses an inverse function and generally does not include display adaptation), i.e., corresponding to 208 in FIG. 2.

[0067] The luminance of a pixel, in this example an SDR image pixel, is input as the corresponding luma Y'SDR. The chrominance, also known as chroma components Cb and Cr, are input to the downstream processing path of the color converter 300.

[0068] The luma Y'SDR is mapped to the required output luminance L'_HDR (e.g., the master HDR reconstruction luminance, or some other HDR image luminance) by the luminance mapping circuit 310. It applies an appropriate function, e.g., the display adaptation luminance mapping function FL_DA(t), for the particular image and maximum display luminance PL_D as obtained from the display adaptation function calculator 350, which uses as input the reference luminance mapping function F_L(t) associated with the metadata. The display adaptation function calculator 350 also determines the appropriate function for processing the chrominance. For the moment, we simply assume that a set of multiplication coefficients mC[Y] for each possible input image pixel luma Y is stored, for example, in the color LUT 301. The exact nature of the color processing can vary. For example, one might want to keep pixel saturation constant by first normalizing the chrominance by the input luma (the corresponding hyperbola in the color LUT) and then correcting the output luma, although differential saturation processing may be used as well. Because both chrominances are multiplied by the same multiplier, hue is generally maintained. Indexing color LUT 301 with the luma value of the currently color-transformed (luminance-mapped) pixel Y yields the required multiplication coefficient mC as the LUT output, which is used by multiplier 302 to multiply it by the two chrominance values ​​of the current pixel, i.e., to yield the color-transformed output chrominance. Cbo=mC*Cb Cro=mC*Cr

[0069] Via a fixed color matrixing processor 303 that applies standard colorimetric calculations, the chrominance is converted to lightness-deficient normalized nonlinear R'G'B coordinates R' / L', G' / L', and B' / L'.

[0070] The R'G'B' coordinates that give the output image the appropriate brightness are obtained by multiplier 311, which calculates: R'_HDR=(R' / L')*L'_HDR, G'_HDR=(G' / L')*L'_HDR, B'_HDR=(B' / L')*L'_HDR which are then grouped into a color triplet R'G'B'_HDR.

[0071] Finally, further mapping to the format required by the display is performed by a display mapping circuit 320. This results in the display driving colors D_C, which are not only formulated into the colorimetry desired by the display (e.g., even HLG OEFT format), but also this display mapping circuit 320 is configured in some variants to perform some specific color processing for the display, i.e., for example, further remapping some of the pixel luminances.

[0072] Some examples illustrating some suitable display adaptation algorithms for deriving corresponding FL_DA functions for possible F_L functions determined by the producing grader are taught in WO2016 / 091406 or ETSI TS 103 433-2 V1.1.1(2018-01).

[0073] However, these algorithms do not give much consideration to the minimum displayable black on the end user's display.

[0074] In fact, the algorithms pretend that the minimum luminance BL_D is small enough to be zero. Therefore, display adaptation primarily addresses the difference in the maximum luminance PL_D of various displays compared to the maximum luminance PL_V of the video.

[0075] As can be seen in drawing 18 of prior application WO2016 / 091406, any input function (in the illustrated example, a simple function formed from two linear segments) is typically scaled diagonally based on a metric positioned along an angle of 135 degrees from the horizontal axis of input luminance in a plot normalized to an input / output luminance of 1.0. This is only one example of a full range of display adaptation algorithms, and it is not intended to limit the applicability of the inventors' novel display adaptation concepts; for example, it should be understood that the angle of the metric direction may have other values, among others.

[0076] However, this metric and its effect on the reshaped F_L function, i.e., the determined FL_DA function, depends only on the maximum display luminances PL_V and PL_D to be delivered with the optimally regraded mid-dynamic range image. For example, the 5000 nit location corresponds to the zero metric point located on the diagonal (for any location along the diagonal corresponding to a possible pixel luminance in the input image), and the 100 nit location (marked PBE) is a point in the original F_L function.

[0077] Display adaptation as a useful variant of this method is summarized in Figure 4 by showing its effect on a plot of possible normalized input luminance Ln_in versus normalized output luminance Ln_out (which is converted to actual luminance, i.e., PL_V value, by multiplying it by the maximum luminance of the display relative to the normalized luminance).

[0078] For example, a video creator designs a luminance mapping strategy between two reference gradings, as described in Figure 1. Therefore, for each possible normalized luminance Ln_in of a pixel in an input image, e.g., a master HDR image, this normalized input luminance must be mapped to a normalized output luminance Ln_out of a second reference grading, which is the output image. This re-grading of all luminances corresponds to a function F_L, which has many different shapes determined by a human grader or a grading automaton, and the shape of this function is communicated as dynamic metadata.

[0079] The question now becomes, in this simple display adaptation protocol, what shape should the derived quadratic version of the F_L function have to map to an MDR image for a mid-dynamic range display (instead of the reference SDR image) (assuming the mapping again starts with the HDR reference-graded image as the input image)? For example, based on the metric, it can be calculated as follows: an 800-nit display should have 50% of the grading effect, while a full 100% is a regrading of the master HDR image to a 100-nit PL_V SDR image. In general, via the metric, for a possible normalized input luminance (Ln_in_pix) of a pixel, represented as display adaptation luminance L_P_n, one determines any point between no regrading to the second reference image and full regrading, the location of which naturally depends on the input normalized luminance, but also on the value of the maximum luminance (PL_V_out) associated with the output image. Those skilled in the art will understand that while the function can be expressed in normalized luminance terms, it can equally be expressed in any normalized luma terms defined according to any OETF.

[0080] The corresponding display-adaptive luminance mapping FL_DA is determined as follows (see Figure 4a): Take any one of all input luminances, e.g., Ln_in_pix. This corresponds to a starting position (depicted as a square) on a diagonal that has an equal angle with the normalized luminance input / output axis. For each point on the diagonal, place a scaled version of the metric (the scaled metric SM) perpendicular to the diagonal (or 135 degrees counterclockwise from the input axis), starting at the diagonal and ending at a point on the F_L curve (at the 100% level), i.e., the intersection (depicted as a pentagon) of the F_L curve with the vertically scaled metric SM. In this example, place the point at the 50% level, i.e., the middle, of the metric (for this PL_D value of the display for which the image must be calculated). [Note that in this case, the PL_V value of the output image is set equal to the PL_D value of the display to which the display-optimized image must be delivered]. By doing this for all points on the diagonal corresponding to all Ln_in values, an FL_DA curve is obtained that is shaped similarly to the original, i.e., with the same regrading, but with maximum luminance rescaling / adjustment. This function is now ready to be applied to calculate the corresponding optimally regraded and / or display-adapted 800 nit PL_V pixel luminance required, given any input HDR luminance value of Ln_in. This function FL_DA is applied by the luminance mapping circuit 310.

[0081] In general, the characteristics of this display adaptation are as follows (not intended to be particularly limiting): The direction of the metric may be fixed in advance as technically desired. Figure 4b shows another scaling metric, namely the vertical scaling metric SMV (i.e., perpendicular to the axis of the normalized input luminance Ln_in). Again, 0% and 100% (or 1.0) correspond, respectively, to no regrading (i.e., an identity transformation on the input image luminance) and regrading to the second of the two reference grading images (related in this example by a luminance mapping function F_L2 of a different shape).

[0082] The location of the measurement points on the metric, i.e., where the 10%, 20%, etc. values ​​are located, is also subject to engineering variation but is generally non-linear.

[0083] It is technically pre-designed, for example, in a television display. For example, a function such as that described in WO2015007505 is used. The logarithmic function can also be designed so that a*(log(PL_V)+b) is equal to the PL_V_HDR value of 1.0 (for example, 5000 nit) and the 0.0 point corresponds to the PL_V_SDR reference level of 100 nit, or vice versa. The position of the PL_V_MDR from which the image brightness needs to be calculated is then obtained from the pre-designed mathematics of the metric.

[0084] The behavior of such a metric is summarized in FIG.

[0085] The display adaptation circuitry 510 includes a configuration processor 511, for example in a television or a set-top box, etc., which sets values ​​for processing of an image before the actual pixel colors of that image are processed. For example, the maximum luminance value of the display-optimized output image PL_V_out may be set once in the set-top box by polling it from the connected display (i.e., the display communicates its maximum displayable luminance PL_D to the set-top box), or if the circuitry is present in the television, it may be set by the manufacturer, etc.

[0086] The luminance mapping function F_L, in some embodiments, varies for each input image (in other variations it is fixed for many images) and is input from some source of metadata information 512 (e.g., it is broadcast as an SEI message, read from a sector of memory such as a Blu-ray disc, etc.). This data establishes the normalized height of the normalized metric (Sm1, Sm2, etc.) on which the desired location of the PL_D value is found from the mathematical formula of the metric.

[0087] When an input image 513 is input, successive pixel intensities (eg, Ln_in_pix_33 and Ln_in_pix_34, or luma) are passed through a color processing pipeline that applies display adaptation, resulting in corresponding output intensities such as Ln_out_pix_33.

[0088] Note that none of this techniques specifically provides a minimum black luminance.

[0089] The reason is that the usual approach is that the black level depends heavily on the actual viewing situation, which is even more variable than the display characteristics (i.e. the very first PL_D). All sorts of influences arise, ranging from physical lighting aspects to the optimal configuration of the light-sensitive molecules in the human eye.

[0090] So you make a good image "for the display," and that's it (i.e., in terms of how much more capable (brightness-wise) the intended HDR display is than a typical SDR display). Then, if necessary, you can post-correct a little later for the viewing situation, which is a special (undefined) task left to the display.

[0091] Therefore, we generally assume that the display has a variable high brightness capability, i.e., is capable of displaying all the necessary pixel intensities coded in the image up to PL_D (for the moment we assume that the image is already MDR optimized for the PL_D value, i.e., that there are generally at least some pixel regions in the image that reach up to PL_D), since we generally do not want to suffer the severe consequences of white clipping, but as mentioned above, the black of the image is often not of interest.

[0092] Black is "almost" visible anyway, so if some of it is less visible, it's not the most important thing. At the very least, the potentially very bright pixel luminance of the master HDR grading can be optimally constrained to the limited upper range of the display, say above 200 nit, e.g. from 200 nit to PL_D=600 nit (for master HDR luminance up to e.g. 5000 nit).

[0093] This is similar to assuming that black is always (at least approximately) zero nits for all images and all displays. White clipping is a much more visually annoying property than losing some of the black, which often still allows you to see something, but is less pleasant.

[0094] However, sometimes that approach is not sufficient because a significant subrange of the darkest luminances becomes invisible, or at least not fully visible, under the significant ambient light of a viewing room (e.g., a consumer television viewer's living room with large windows during the day), which may be different from, and dim or even darker than, the ambient light conditions in a video editing room where the video is created.

[0095] It is therefore necessary, for example, to increase the brightness of those pixels, typically using a control button on the display (a so-called brightness button).

[0096] When using a television electronic behavior model such as in Rec. ITU-R BT.814-4 (07 / 2018), a television in an HDR scenario takes the luma+chroma pixel colors (which actually drive the display) and converts them to nonlinear R', G', B' drive values ​​(according to standard colorimetric calculations) to drive the panel. The display then processes these R', G', B' nonlinear drive values ​​with a PQ EOTF to know what front screen pixel luminance to display (i.e., generally, how to drive an OLED panel pixel or an LCD pixel, if there is still internal processing that accounts for the electro-optical physical behavior of the LCD material, but that aspect is irrelevant to this discussion).

[0097] A control knob, for example on the front of the display, then gives the luma offset value b (the moment at which a black patch 2% above the minimum in the PLUGE or other test pattern becomes visible, while -2% black becomes invisible).

[0098] The original uncorrected display behavior is LR_D=EOTF[max(0,R')]=PQ[max(0,R')] LG_D=EOTF[max(0,G')]=PQ[max(0,G')] LB_D=EOTF[max(0,B')]=PQ[max(0,B')] [Formula 1] If so.

[0099] In this formula, LR_D is the amount of red contribution (linear) that should be displayed to create a particular pixel color with a particular luminance (in (fractional) nits), and R' is the non-linear luma code value, e.g., 419 out of 1023 values ​​in a 10-bit encoding.

[0100] The same happens for the blue and green components. For example, if you need to make a particular color 1 nit (the total brightness of that color to the eye), you need, say, 0.33 units of blue, and the same for red and green. If you need to make 100 nits of that same color, you can say that LR_D=100*0.33 nits.

[0101] Now, if this display driving model is controlled via the luma offset knob, the general equation becomes: LR_D_c=EOTF[max(0,a*R'+b)], where a=1-b / OETF[PL_D], etc. [Equation 2]

[0102] Instead of displaying the zero black of the image hidden somewhere in the invisible display black, this technique raises the zero black to the exact level where black becomes sufficiently discernible. (Note that in consumer displays, mechanisms other than PLUGE may be used, e.g., despite the viewer's preferred value of the available luma offset b, viewer preference may potentially lead to another possible suboptimal.)

[0103] This is a post-processing step for the display after creating an optimally regraded image: first, the optimal theoretically regraded image is calculated by the decoder and, for example, first mapped to the reconstructed master HDR image, and then luminance-remapped to, for example, a 550 nit PL_V MDR image, i.e., taking into account the display's brightness capabilities PL_D. Then, after this optimal image has been determined according to the filmmaker's ideal vision, it is further mapped by the display, taking into account the expected visibility of black in the image.

[0104] According to the inventors, the problem is that this is a rather crude way of adapting the viewing room ambient light level of the image to be displayed; an alternative approach has been developed. Summary of the Invention

[0105] Visually better looking images for various ambient lighting levels are obtained by a method of processing an input image to obtain an output image, the method comprising: obtaining a starting luma (Yn_CC) for a pixel of an input image; obtaining a minimum luminance (mL_VD) of a target display, the target display corresponding to an end-user display (210) to which the output image may be supplied for displaying the output image; obtaining a minimum luminance (mL_De) of the end user display in the viewing room, the minimum luminance (mL_De) of the end user display being dependent on the amount of illumination in the viewing room; Calculating a difference (dif) by subtracting the luminance of the end user display (mL_De) from the minimum luminance of the target display (mL_VD); converting the difference (dif) to a luma difference (Ydif) by applying the difference as an input to an opto-electrical transfer function (OETF_psy), thereby providing the luma difference as an output; Mapping the starting luma by applying a linear function, the linear function being: a) applying the luma difference multiplied by -1.0 as an additive factor; b) applying a linear multiplicative factor with the luma difference increased by a value of 1.0 and multiplying the starting luma by that multiplicative factor; a mapping step, the sum of the two resulting in the mapped luma (Yim); Transforming the mapped luma (Yim) by applying the inverse of the photoelectric transfer function to obtain a normalized middle luminance (Ln_im); subtracting the second minimum luminance of the end-user display (mL_De2) divided by the maximum luminance of the output image (PL_O) from the intermediate luminance (Ln_im) and scaling the subtraction by 1.0 minus the second minimum luminance of the end-user display (mL_De2) divided by the maximum luminance of the output image (PL_O) to obtain a final normalized luminance (Ln_f); multiplying the final normalized luminance (Ln_f) by the maximum luminance of the output image (PL_O) to obtain the output luminance; outputting the output luminance in a color representation of the pixel of the output image; It has.

[0106] The starting luma is calculated, for example, by applying an opto-electrical transfer function (OETF_psy) to the input luminance (L_in) of the input image, where L_in is input. Instead of the starting luma being defined directly from the input luminance by applying the OETF, further luma-to-luma mapping may be included to obtain the starting luma, particularly luma mapping that optimizes the distribution of the image's lumas for a particular display capability. In some methods or devices, the input image pixel color has a luma component.

[0107] Although the target display, specifically the minimum luminance (mL_VD) of the target display, corresponds to the (actual) end-user display (210), in particular a consumer's display having a specific limited maximum luminance, the target display is still characterized in terms of the ideal display on which the content should be displayed, i.e., the technical characteristics of the image (defined by a metadata-based joint association of the technical details of the ideal display) over and above the specific actual display.

[0108] This method uses two different characteristics for the display minimum: one is mL_De, which includes the black behavior of the display itself (e.g., LCD leakage); the other is mL_De2, which depends only on the ambient lighting in the viewing room (e.g., screen reflections) but does not depend on the physical display black, so the factor is zero (i.e., assumed to be zeroed) at full darkness. The reason is that the psychovisual correction takes both black offsets into account, but the final compensation only takes ambient light into account (which will preserve the image black level).

[0109] The photoelectric transfer function (OETF_psy) used is preferably a psychovisually uniform photoelectric transfer function defined by determining (typically experimentally in a laboratory) the shape of a mapping function from luminance to luma so that a second luma a fixed integer number N lumas higher than a first luma chosen somewhere in the luma range corresponds approximately to a similar difference in perceived lightness to a human observer.

[0110] That is, human vision is non-linear, and therefore does not perceive the difference between 10 nit and (1.05)*10 nit in the same way as, for example, the difference between 2000 nit and (1.05)*2000 nit. The equalized perception curve also depends on what the human is looking at, i.e., in particular the dynamic range (in a particular surrounding) of the display being viewed, i.e., the maximum luminance PL_D and the minimum luminance.

[0111] So, ideally, for processing purposes we define an OETF (or its inverse, EOTF) with the following properties:

[0112] If we take the first luma, for example, in 10 bits, luma_1=10, and then we move the N=5 luma code higher, for example, and get the second luma, luma_2=15. This corresponds to a change in the lightness sensation (i.e., the value that characterizes what a human viewer experiences as the lightness of a displayed patch at a particular physically displayed luminance; it is determined by applying the EOTF).

[0113] So, let's assume that in the brightness range of 1 to 200, a luma of 10 gives an impression of brightness of 5, and a luma of 15 gives an impression of 7, i.e., 2 units brighter.

[0114] Next, take two lumas that encode brighter luminance, for example, 800 and 800+5. Then, if the viewer experiences the same difference in lightness due to the luma difference, the luma scale is approximately visually uniform. For example, luma 800 gives a lightness perception of 160, while luma 805 appears to be lightness 162, i.e., again, there is a 2-unit difference in lightness. The function does not need to be exactly the lightness-determining function, since a reasonably perceptually uniform luma definition already works well.

[0115] The effects of changes due to luminance processing would be less visually unpleasant (because they are more easily suppressed by the brain) if they were performed in such a psychovisually uniform system.

[0116] It should be understood that the minimum luminance of the target display (mL_VD) is similar to the maximum luminance of a virtual target or target display associated with an image, but it says something about the image capabilities rather than a specific actual display. Because the image representation is associated with metadata that characterizes the image on a corresponding ideal target display, it is again an ideal value associated with the image. It is not the minimum achievable (displayable, i.e., generally visible or discernible) luminance, also known as black luminance, of the actual end-user display, because they are compared in a specific manner in the algorithm. However, it is a variable minimum value associated with the actual end-user display, and in particular, as shown below, it is the maximum luminance of the end-user display (PL_D), but also the minimum luminance selected by the image creator for reference grading (master HDR image and corresponding typical SDR image). All of them are selected for different reasons, but must technically correspond as the present innovation ensures. This correspondence, then, involves a one-to-one relationship between the selectable end-user displays, which are combined or can be combined, and the value of the minimum luminance of the target display (mL_VD) calculated as shown in the embodiments. What is important is that there exists such a minimum luminance (mL_VD) for the target display.

[0117] The minimum luminance (mL_De) of an end-user display is also variable but depends on the actual viewing situation, i.e., the end-user display present in the viewing environment, such as a living room for watching television at night. It has display-dependent properties but varies primarily with the viewing environment conditions (e.g., how many lamps are present, where they are located in the room, which lamps are respectively on or off, etc.).

[0118] The maximum luminance of the output image (PL_O) should be understood as the maximum possible luminance of the output image, not necessarily the luminance of a pixel present in a particular image of a temporally consecutive set of images (video). For an image, it is defined as the maximum value that occurs (e.g., represented by the maximum luma code in an N-bit system), commonly annotated as image describing metadata, and when that image is optimized for a particular display, it is the maximum displayable luminance, for example by setting the maximum drive signal to the driver, which consequently generally corresponds to the maximum energy output that the display can achieve, and in particular the light luminance for this drive (i.e., generally the brightest pixel that the actual display can display).

[0119] The brightness of the output image naturally depends on the conditions under which the output image is derived. The present invention exists, for example, in a system that obtains an image already optimized for the maximum brightness capability of the display. For example, the display indicates to the image source that it is a 1000 nit display, i.e., that it requires an image optimized for PL_V_HDR=1000 nit as input. The embodiment then performs a corresponding optimization for the darkest brightness for any (still variable) viewing environment lighting conditions. In that case, PL_O simply remains PL_V_HDR.

[0120] In another embodiment, it is also taken into account that the display has a PL_D lower than the PL_V_HDR of the input image, and therefore the method performs display adaptation (also known as display optimization), the method comprising the steps of: calculating a display optimization luminance mapping function (FL_DA) based on a reference luminance mapping function (F_L) and a maximum luminance (PL_D) of the end user display; and using the display optimization luminance mapping function (FL_DA) to map input lumas (Yn_CC0) of pixels of the input image to output lumas, thereby forming the starting lumas (Yn_CC) for the remainder of the processing; the display-optimized luminance mapping function (FL_DA) has a less steep slope for the darkest input luminance compared to the shape of the reference luminance mapping function (F_L); A reference luminance mapping function specifies a relationship between the luminance of a first reference image and the luminance of a second reference image having respective first and second reference maximum luminances, such that the maximum luminance of the end-user display (PL_D) falls between the first and second reference maximum luminances.

[0121] That is, any of the possible embodiments to make lesser variations of the specially formulated F_L function (corresponding to the specific required luminance remapping, a.k.a. re-grading needs for the particular HDR image or scene at hand, to obtain a corresponding lower maximum luminance secondary grading) as described in Figures 4 and 5, or similar techniques, are used. The luminance mapping function is represented as a luma mapping function, i.e., by converting both normalized luminance axes to the luma axis via the appropriate OETF (normalization using the appropriate maximum luminance value of the input and output images).

[0122] In particular, the secondary grading for which the F_L function is defined is advantageously a 100 nit grading, which is suitable to meet the requirements of most future display situations.

[0123] Advantageously, the step of calculating the display optimized luminance mapping function (FL_DA) comprises, for example, finding a position (pos) on a metric (SM) that determines the position of the maximum luminance value, which position corresponds to the maximum luminance (PL_D) of the end user display; a first endpoint of the metric corresponds to a first maximum luminance (PL_V_HDR) of the input image, and a second endpoint of the metric corresponds to a maximum luminance of a second reference image; a first endpoint of the metric is located at a diagonal point having horizontal and vertical coordinates equal to the input luma for any normalized input luma (Yn_CC0); The second endpoint is concatenated with the output value of a reference luminance mapping function (F_L) determined by the direction of the metric, where the first endpoint is on the diagonal and the second endpoint is on where the locus of the F_L function lies, for example, for a vertical metric direction line, i.e., at the point Yn_out=F_L(Yn_CC0) having a particular (all possible) input luma as horizontal coordinate and the output luma obtained by applying the F_L function as output, where the luminance mapping function is generally not represented in the luminance domain per se, but generally by an appropriate OETF-ization of the coordinate axes in the psychovisually equalized luma domain.

[0124] Some advantageous device embodiments result from innovative modifications, for example: An apparatus for processing an input image to obtain an output image, comprising: an input unit for obtaining an image start luma (Yn_CC), such as a circuit including a first photoelectric conversion circuit (601) configured to apply a photoelectric transfer function (OETF_psy) to an input luminance (L_in) of the input image to obtain an input image start luma (Yn_CC); a first minimum metadata input (693) for receiving a minimum luminance (mL_VD) of a target display, the target display corresponding to an end-user display (210) on which the output image may be provided for displaying the output image; a second minimum metadata input (694) for receiving a minimum brightness of the end user display (mL_De) in the viewing room, the second minimum metadata input (694) being dependent on the amount of lighting in the viewing room; a brightness difference calculator (610) configured to calculate a difference (dif) by subtracting the brightness of the end user display (mL_De) from the minimum brightness of the target display (mL_VD); a second photoelectric conversion circuit (611) configured to convert the difference (dif) into a luma difference (Ydif) by applying the difference to a photoelectric transfer function (OETF_psy) as an input to the photoelectric transfer function (OETF_psy), thereby providing the luma difference as an output; a linear scaling circuit (603) configured to map the starting luma by applying a linear function that applies the luma difference multiplied by -1.0 as an additive constant and multiplies the starting luma using the luma difference incremented by a value of 1.0 as a multiplier, thereby resulting in a mapped luma (Yim); an electro-optical conversion circuit (604) configured to convert the mapped luma (Yim) by applying an inverse of an electro-optical transfer function to obtain a normalized middle luminance (Lim); a third minimum metadata input (695) for receiving a second minimum brightness (mL_De2) of the end user display; a final surround adjustment circuit (605) configured to subtract a second minimum luminance of the end-user display (mL_De2) divided by a maximum luminance of the output image (PL_O) from the intermediate luminance (Ln_im) and scale the subtraction by 1.0 minus a result of dividing the second minimum luminance of the end-user display (mL_De2) by the maximum luminance of the output image (PL_O) to obtain a final normalized luminance (Ln_f); a multiplier configured to multiply the final normalized luminance (Ln_f) by a maximum luminance (PL_O) of the output image to obtain an output luminance; a pixel color output unit 699 for outputting an output luminance in a color representation of a pixel of the output image; Includes.

[0125] The technical realization of the input unit or circuit will depend on the configuration and which variant of the input color definition is input to the device, but the input that is useful to the device at the output of the input stage is the starting luma.

[0126] The apparatus for processing an input image (of any of the combinations of technical elements of the embodiments) further includes a maximum determination circuit (671) configured to determine the maximum luminance of the output image (PL_O) as either the maximum luminance of an end-user display (PL_D) that can display the output image or the maximum luminance of the input image (PL_V_HDR).

[0127] An apparatus for processing an input image (in any combination of embodiments) comprising: a display optimization circuit (670) configured to calculate a display optimized luminance mapping function (FL_DA) based on the reference luminance mapping function (F_L) and a maximum luminance of the end user display (PL_D); a luminance mapping circuit (602) configured to apply a display-optimized luminance mapping function (FL_DA) to an input luminance, specifically expressed as luma, to obtain an output luma (and hence luminance); further comprising the display-optimized luminance mapping function (FL_DA) has a less steep slope for the darkest input luminance compared to the shape of the reference luminance mapping function (F_L); A reference luminance mapping function specifies a relationship between the luminance of a first reference image and the luminance of a second reference image having respective first and second reference maximum luminances, such that the maximum luminance of the end-user display (PL_D) falls between the first and second reference maximum luminances.

[0128] The second reference image is a standard dynamic range image with a second reference maximum luminance equal to 100 nits, or a psychovisually uniform OETF and EOTF are applied by the circuitry that applies them or the device for processing the input image has display-adaptive specific modifications, such as in the display optimization circuit 670.

[0129] In particular, those skilled in the art will understand that these technical elements may be embodied in a variety of processing elements, such as ASICs, FPGAs, processors, etc., and may be present in a variety of consumer or non-consumer devices, whether including displays or non-display devices externally connected to a display; that images and metadata may be communicated to and from over a variety of image communication technologies, such as over-the-air broadcast, cable-based communication, etc.; and that devices may be used in a variety of image communication and / or usage ecosystems, such as, for example, television broadcast, on-demand via the Internet, etc.

[0130] These and other aspects of the method and apparatus according to the invention will become apparent from and will be explained with reference to the implementations and embodiments described below, and with reference to the accompanying drawings, which merely serve as non-limiting specific illustrations illustrating more general concepts, in which dashes are used to indicate that a component is optional, and that an un-dashed component is not necessarily essential. Dashes are also used to indicate elements that are described as essential but are hidden inside an object, or something intangible, such as a selection of an object / region, for example. [Brief explanation of the drawings]

[0131] [Figure 1]1A and 1B are schematic diagrams illustrating some typical color transformations. Color transformation occurs when optimally mapping a high dynamic range image to a corresponding, optimally color-graded and similar-looking (as similar as desired and feasible, given the differences in the first and second dynamic ranges DR_1 and DR_2, respectively) lower dynamic range image, e.g., a standard dynamic range image with a 100 nit maximum luminance, which, in the lossless case, also corresponds to mapping a received SDR image that actually encodes an HDR scene to a reconstructed HDR image of that scene. Luminance is shown as a location on the vertical axis, from darkest black to maximum luminance PL_V. The luminance mapping function is symbolically represented by an arrow that maps the average object luminance from the luminance on the first dynamic range to the second dynamic range (those skilled in the art will know how to equivalently depict this as a classical function on axes normalized by dividing by the respective maximum luminance, e.g., normalized to 1). [Figure 2] 1 shows a schematic example of a high-level view of applicant's recently developed technique for encoding high dynamic range images, i.e., images that can typically have a brightness of at least 600 nits or more (typically 1000 nits or more), which in effect communicates an HDR image by itself or as a corresponding brightness-regraded SDR image plus metadata encoding a color transformation function that includes at least an appropriate determined brightness mapping function (F_L) for pixel colors used by a decoder to convert the received SDR image to an HDR image. [Figure 3] FIG. 1 shows the internal details of an image decoder, in particular a pixel color processing engine, in a (non-limiting) preferred embodiment. [Figure 4]Consisting of sub-images Figures 4a and 4b, they illustrate two possible variations of display adaptation to obtain a final display-adapted luminance mapping function FL_DA that is used to calculate the optimal display-adapted version of an input image for a specific display capability (PL_D) from a reference luminance mapping function F_L that codifies the luminance re-grading needs between two reference images. [Figure 5] 1 is a diagram summarizing the principles of display adaptation more generally, for easier understanding of the principles of display adaptation as an element in the formulation of the present embodiments and claims. [Figure 6] FIG. 1 illustrates an exemplary device for explaining technical elements related to ambient lighting compensation processing. [Figure 7] FIG. 1 illustrates the concept of virtual target display correlation, in particular the target display's minimum luminance (mL_VD) and its relationship or relevance to a real display, such as a viewer's end-user display, on which luminance remapping is typically performed. [Figure 8] FIG. 1 illustrates some of the metadata categories that are available for long-standing professional HDR encoding frameworks (e.g., typically associated with pixel color images), so that any receiver has all the data available that it needs, or at least elegantly utilizes, to optimize the received image specifically for the particular end-viewing situation (display and environment). [Figure 9] FIG. 1 illustrates a method for optimizing image contrast, particularly to compensate for different ambient viewing light levels or conditions, particularly when designing based on existing display adaptation techniques. [Figure 10] FIG. 7 shows an example of what happens when applying the viewing environment lighting adaptation technique (described in FIG. 6) for several image luminances using a particular reference mapping function F_L (chosen to be a simple linear curve for ease of understanding). [Figure 11]This figure shows an example of what happens when applying the contrast improvement technique described in Figure 9 (without ambient lighting, although the two processes can also be combined, which introduces an additional offset to the blackest black). DETAILED DESCRIPTION OF THE INVENTION

[0132] FIG. 6 illustrates an apparatus for generally describing how to implement elements of the innovations, particularly color processing circuit 600. This will be generally described, followed by some details of embodiment variations. We will assume that the ambient light adaptive color processing is performed within an application-specific integrated circuit (ASIC), but those skilled in the art will understand how to similarly implement color processing in other related devices. Some aspects vary depending on whether this ASIC resides, for example, in a television display (typically an end-user display) or in another device connected to the display, such as a set-top box, to perform processing for the display and provide it with an already brightness-optimized image. (Those skilled in the art will also be able to map this apparatus block diagram to a method flow diagram of a corresponding method.)

[0133] The input luminance (L_in) of the input image to be processed (without loss of generality, we assume it to be the master HDR image MAST_HDR) is input via the image pixel data input unit 690 and then first converted to the corresponding luma (e.g., 10-bit coded luma) by using an optical-to-electronic transfer function, which is generally fixed in the device by the manufacturer (although it is configurable).

[0134] It is useful to have a perceptually uniform OETF (OETF_psy).

[0135] Assume the following OETF is used (defined as the output of luma Yn_CC0): Yn_CC0=v(L_in;PL_V_in)=log[1+(RHO-1)*power(L_in;p)] / log[RHO] [Formula 3] where RHO is a constant that depends on the input maximum luminance PL_V_in according to the formula RHO(PL_V_in)=1+32*power((PL_V_in / 10,000);p), where p is preferably a power of a power function equal to 1 / (2.4).

[0136] Input Maximum Luminance is the maximum luminance associated with the input image (it is not necessarily the luminance of the brightest pixel in each image of the video sequence, but is metadata that characterizes the image or video as an absolute upper limit).

[0137] It is generally configurable and input via first maximum data input 691 as PL_V_HDR, for example 2000 nits, or in other variations, a fixed value for the image communication and / or processing system, for example 5000 nits, and therefore a fixed value in the processing of photoelectric conversion circuit 601 (thus the vertical arrow representing the PL_V_HDR data input is shown as a dotted line, as it is not present in all embodiments; note that the open circle symbolizes the branching of this data supply, so as not to be confused with a non-intermixed overlapping data bus).

[0138] Then, in some embodiments, there is further luminance mapping by luminance mapping circuit 602 in order for the method or color processing device to obtain a starting luma Yn_CC that is optimally adapted to the particular viewing environment. Such further luminance mapping is typically display adaptation to pre-adapt the image to the particular display maximum luminance PL_D of the connected display.

[0139] In simpler embodiments, this optional luminance mapping does not exist, and so the 2000 nit environment-optimized image is calculated for the input 2000 nit master HDR image (or a fixed 5000 nit HDR image situation), i.e., we will first describe color processing assuming that the maximum luminance value remains the same between the input and output images. In this case, the starting luma Yn_CC is simply equal to the initial starting luma Yn_CC0 output by the photoelectric conversion circuit 601.

[0140] The linear scaling circuit 603 Yim=(Ydif+1)*Yn_CC-1.0*Ydif [Formula 4] Calculate the intermediate luma Yim by applying a function of the type

[0141] The luma difference Ydif is obtained from a (second) photoelectric conversion circuit 611, which converts the luminance difference dif into a corresponding luma difference Ydif. This circuit uses the same OETF formula as circuit 601, i.e., it also uses the same PL_V_in values ​​(and exponents). The luminance difference dif is calculated by a luminance difference calculator 610, which receives the two darkest luminance values, namely the minimum luminance of the target display (mL_VD) and the minimum luminance of the end user display (mL_De) given the specific lighting characteristics of the viewing room, and dif=mL_VD-mL_De [Equation 5] Calculate mL_De is generally a function (typically additive) of the fixed display minimum black (mB_fD) of the end-user display, on the one hand, and luminance (mB_sur), which is a function of the amount of ambient light. A typical example of display minimum black (mB_fD) is the leakage light of an LCD display. When such a display is driven with a code indicating complete black (i.e., ideally, zero photons are output), due to the physical properties of LCD materials, such a display will still always output, for example, 0.05 nits of light, the so-called leakage light. This is regardless of the ambient light in the viewing environment, and therefore applies to a completely dark viewing room.

[0142] The mL_VD and mL_De values ​​are typically obtained via a first minimum metadata input 693 and a second minimum metadata input 694, e.g., from an external memory in the same or a different device, or via circuitry from a light measurement device, etc. The maximum luminance value required in a particular embodiment, e.g., the maximum luminance of a display capable of displaying the color-processed output image (PL_D), is typically input via an input connector such as a second maximum metadata input 692. The output image color is fully written in a pixel color output 699 (those skilled in the art will understand how to realize this in various technical variants, e.g., pins on an IC, a standard video cable connection such as HDMI, wireless channel-based communication of the image's pixel color data, etc.).

[0143] The precise determination of the luminance as a function of the amount of ambient light (mB_sur) (referred to as ambient black luminance for short) is not a typical aspect of the present innovation, as it may be determined in several alternative ways. For example, a viewer would typically use a test signal such as PLUGE, or a variant more convenient for consumers, to determine a value for luminance mB_sur that they consider to represent the masking black resulting from reflections on the front screen of the display. It is further assumed that the viewer simply sets the value, whether for a particular evening or, for example, from a situation where consumers typically keep the room lighting configuration fixed when purchasing a display. Or even a value baked in by a television manufacturer as a value that works well on average for at least one typical consumer viewing situation. When this method is used to adapt luminance for viewing on a mobile device, it is common for the viewing environment to be relatively unstable (e.g., watching a video on a train and the lighting level changing from outdoors to indoors as the train enters a tunnel).

[0144] In such cases, for example, one can use the time-filtered measurements of the built-in illuminance meter (measured less frequently to avoid adapting the treatment over time, e.g., when sitting on a bench in the sun and then walking indoors).

[0145] Such instruments generally measure the average amount of light (in lux) that falls on it, and therefore on the display.

[0146] Although a different photometric quantity, the lux value is converted to ambient black luminance by the well-known photometric formula: mB_sur=R*Ev / pi [Equation 6]

[0147] In this formula, Ev is the ambient illuminance in lux, Pi is the constant 3.1415, and R is the reflectance of the display screen. Mobile display manufacturers typically include their values ​​in the formula.

[0148] For a normal reflective surface in the surroundings, e.g., a house wall, with a color somewhere between average gray and white, an R-value of about 0.3 can be assumed, i.e., as a rule of thumb, the luminance value is about 1 / 10 of the illuminance value stated in lux. The front of a display reflects much less light. Depending on whether special anti-reflection technology is used, the R-value may be, for example, about 1% (although it may be higher, towards the 8% reflectivity of glass, which may be problematic, especially in brighter viewing environments). Therefore, mL_De=mB_fD+mB_sur [Equation 7]

[0149] The technical meaning of the value of the minimum luminance of the target display (mL_VD) is further explained with reference to FIG.

[0150] The first three luminance ranges, starting from the left, are actually "virtual" display ranges, i.e., ranges that correspond to the image, not necessarily to an actual display (that is, to a target display, i.e., a display on which the image could ideally be shown, but which the consumer potentially does not own). These target displays are co-defined because the image is created specifically for the target display (i.e., the luminance is graded toward a specific desired object luminance). For example, an explosion cannot be very bright on a 550-nit display, so a grader may want to reduce the luminance of other image objects so that the explosion appears at least somewhat contrasty. However, it is quite possible that no one owns such a display, and the image still needs to be optimized by display adaptation to the actual display owned by a particular viewer. The range of luminance physically displayable on this end-user display is shown as the luminance range on the far right (EU_DISP).

[0151] This information for the target display(s) constitutes metadata, some of which is typically communicated along with the image itself, i.e., along with the image pixel intensities.

[0152] This approach to characterizing (encoding) HDR images is a significant departure from traditional SDR image coding, and because these aspects are only recently invented, and because important technical elements should not be misunderstood, the concepts involved are summarized for the reader using FIG. 8.

[0153] HDR video is well complemented by metadata, because many aspects may differ (e.g., the maximum luminance of SDR displays has always been in the range of roughly 100 nits, but now people have displays with significantly different display capabilities, e.g., PL_D equal to 50 nits, 500 nits, 1000 nits, 2500 nits, and in the future maybe even 10,000 nits; content characteristics such as the maximum codable luminance PL_V of the video also vary considerably; therefore the luminance distribution between darks and lights that a grader makes for a typical scene will also be very different between a typical SDR image and any HDR image, etc.), and should not suffer from insufficient control over these various unstable ranges.

[0154] As explained above, one must obtain at least one pixelation matrix of pixel colors, including at least pixel luma, otherwise the image shape will not even be visible (even if colorimetrically misrepresented).As mentioned above, by actually communicating only one image per time, it is possible to communicate two different dynamic range images (which can serve as two reference gradings to indicate the need for luminance re-grading of specific video content when different dynamic range images such as MDR images need to be created).

[0155] Assume that the master HDR image itself is being communicated, so that the first data set 801 (image color) contains the color component triplets of the pixels of the communicated image and is the input to a color processing circuit present, for example, in a receiving television display. Typically, such images are digitized as a (e.g., 10-bit) luma and two chroma components Cr and Cb (although non-linear R'G'B' components can also be communicated). However, it is necessary to know which luminance the luma 1023 or, for example, 229 represents.

[0156] Therefore, it communicates container metadata 802. Assume that luma is defined, for example, according to the perceptual quantizer EOTF (or its inverse, OETF), as standardized in SMPTE ST.2084. This is a large container that can specify luma up to 10,000 nits. Even if images are not currently produced with such high pixel luminance, it can be said to be a "theoretical container" that contains actually used luminance up to, for example, 2500 nits (note that it also communicates primary color chromaticity, but those details would simply unnecessarily hinder this description).

[0157] Of interest is the maximum pixel luminance that can actually be (or will be) coded for the video, which is coded into another video characteristic metadata, which is typically the master display color volume metadata 803.

[0158] This is an important luminance value for the receiver to know, because even if the display doesn't care about the specific details of how it remaps all luminances along the range (at least according to the content creator's desired display adaptation), knowing the maximum still guides it roughly what is best to do at all luminances, since it will at least know what luminance the video will not exceed for an image pixel.

[0159] In the 2000 nit example, this Master Display Color Volume (MDCV) metadata 803 includes the master HDR image maximum luminance, i.e., the PL_V_HDR value in FIG. 7, i.e., characterizing the master HDR video (in this example, it is actually also communicated as the PQ pixel luma defined in SMPTE 2084, but that aspect can be ignored for now, since those skilled in the art know how to convert between the two and can understand the principles of innovative color processing as if the (linear) pixel luminance itself comes in).

[0160] This MDCV is the "virtual" display, or target display, for the HDR master image. By indicating this in the metadata, the video creator is indicating where in the movie there are 2000 nits of brightness pixels in the video signal that is communicated to the receiving actual end-user display, so that the end-user display will take that fully into account when processing the current image's brightness.

[0161] These (actual) brightnesses of a set of images are in fact yet another aspect, and therefore there is a further video-related metadata set 804, which gives further information about the actual video, rather than the characteristics of the associated display (i.e., for example, the maximum possible for the video). To easily understand this, suppose two videos are annotated with the same MDCV PL_V_HDR (and EOTF). The first video is a night video, and therefore in fact the images do not reach a pixel brightness higher than, say, 80 nits (although it is still specified with an MDCV of 2000 nits; furthermore, if it is another video, it may have flashlights in at least one image, which has a few pixels that reach the 2000 nit level, or nearly so), while the second video, specified / created according to the exact same encoding technique (i.e., annotated with the same data in 802 and 803), consists only of explosions, i.e., has pixels mostly above 1000 nits.

[0162] On the one hand, there may be something additional to say about this video, but on the other hand, the experienced reader will understand that if we wanted to re-grade both videos from a 2000 nit representation to, say, a 200 nit output representation, we could do it differently (we could scale the explosion by simply dividing the luminance by 10, while keeping the luminance of the night scene the same in the master HDR and the 200 nit output image).

[0163] A possible useful metadata in set 804 annotating the communicated HDR images (optional in the present invention, but mentioned nevertheless for completeness) is the average pixel luminance of all pixels of all time-series images, with MaxFall being, for example, 350 nits. The receiving color processing can then understand from this value that if it is dark, the luminance should be displayed as is, i.e., without mapping, and if it is bright, dimming is required.

[0164] Even if one does not actually communicate SDR video, i.e., only sends metadata (metadata 814), one can also annotate the SDR video (i.e., a second reference-graded image, which indicates how an SDR image should look as similar as possible to the master HDR image by the content creator in the case of reduced dynamic range capabilities).

[0165] Thus, although some HDR codes may also send pixel color triplets, i.e., SDR pixel color data 811, containing pixel lumas of an SDR image (Rec. 709 OETF definition), as explained above, the illustrative codec, i.e., the SLHDR codec, does not actually communicate this SDR image, i.e., pixel colors, or anything that depends on that color, and is therefore elided (if pixel colors are not communicated, there is also no need to communicate SDR container metadata 812, which indicates the container format of how pixel codes are defined and should be decoded into linear RGB pixel color components).

[0166] What should ideally be communicated (although some systems implicitly assume it) is the corresponding SDR target display metadata 813. In such situations, one would typically fill in the value of PL_V_SDR as equal to 100 nits.

[0167] Importantly to the present innovation, this is also typically where the video creator enters the assumed minimum black value of the theoretical target SDR display to which the film was graded, i.e., SDR reference minimum black mB_SDR.

[0168] For example, if an author assumes that they are making a video for a display that cannot go deeper than 0.1 nit (e.g., due to LCD spill), they do not want very important image object pixel luminances to approach this value, so they may start with, say, 0.2 nit SDR image pixels, and maybe go a little above that, with most pixels above the 1 nit level. There is a similar value that characterizes the master HDR image, or more precisely the target display associated with it, namely HDR minimum black mB_HDR (in metadata 803).

[0169] The need for regrading is image dependent and is therefore advantageously encoded into the SDR video metadata 814 according to the codec of the selected description as explained above. As explained above, typically one communicates one (or multiple, even a single time) image-optimized luminance mapping functions (i.e., F_L reference luminance mapping function shapes) for mapping HDR luminance (or luma) normalized to up to 1 to SDR luminance or luma normalized to up to 1 (the exact manner in which these functions are coded is irrelevant to this patent application; examples can be found in the above-mentioned ETSI SLHDR standard). Now that this function shape (combined with the target display's maximum luminance, HDR, and SDR metadata) is used to define the required remapping of all possible image pixel luminances, further metadata about the image, such as maxFall, is not actually required, although it can also be communicated.

[0170] This already constitutes a fairly specialized set of HDR video coding data to which current ambient adaptive luminance remapping techniques can be applied.

[0171] However, two further sets of metadata (particularly useful for contrast optimization embodiments described in more detail below) may be added (in various ways), but are not yet fully standardized: Content creators, while using a perceptual quantizer, may also be working under the implicit assumption that the video will be created in a viewing environment with a particular illumination level, e.g., 10 lux (so that it will ideally also be displayed as closely as possible at the receiving end for optimal appearance) (whether a particular content creator also strictly follows this lighting suggestion is another matter).

[0172] For increased certainty, one can associate a typical (intended) viewing environment for which the master HDR image was specifically graded (HDR Viewing Metadata 805; optional / dotted line). As mentioned above, a 2000 nit maximum master HDR image could be created. However, if this image were intended to be viewed in a 1000 lux viewing environment, the grader would probably not grade too many subtle dark object luminances (such as slightly different dimly lit objects in a dark room seen through an open door in the background of a scene containing a lit first room in front) because the viewer's brain would likely simply see all of this as "flat black." However, the situation is different if the image is created for typical, say, dim evening viewing in a 50 lux or 10 lux room. For typical brighter viewing conditions, e.g., 200 lux, the regrading to the SDR image, particularly the F_L function, can be made more specific and annotated to the SDR ambient metadata 815, and the receiving device can also take advantage of that information if it so desires (or communicate regrading for different SDR images for different intended viewing).

[0173] Returning to FIG. 7, we show the settings (by display adaptation) we wish to create an output image for a 600 nit MDR display, i.e. the output image should be a PL_V_MDR=600 nit output image.

[0174] If one were to optimize an output image with the same maximum luminance as the input image according to the present innovations (the simple situation described above in FIG. 6), the value of the target display's minimum luminance mL_VD would simply be the mB_HDR value. However, now one needs to set the minimum luminance of the MDR display dynamic range (as an appropriate value for the target display's minimum luminance mL_VD), which is interpolated from the luminance range information of the two reference images (i.e., in the standardized embodiment description of FIG. 8, the MDCV and presentation display color volume PDCV are typically jointly communicated or at least obtainable in metadata 803 and 813, respectively).

[0175] The formula is as follows: mL_VD=mB_SDR+(mB_HDR-mB_SDR)*(PL_V_MDR-PL_V_SDR) / (PL_V_HDR-PL_V_SDR) [Formula 8]

[0176] In this equation, the PL_V_MDR value is selected to be equal to the PL_D representation of the display to which the display-optimized image is to be supplied.

[0177] Returning to Figure 6, the dif value is converted to a psychovisually uniform luma difference (Ydif) in the photoelectric conversion circuit 611 by applying the v-function of Equation 3 and substituting the value dif (i.e., normalized luminance difference) divided by PL_V_in for L_in, where the value PL_V_HDR is used to create an ambient adjusted image with the same maximum luminance as the input master HDR image, respectively, as the value of PL_V_in, and in a display scenario adapted to an MDR display, the value of PL_V_in is the PL_D value of the display, e.g., 600 nit.

[0178] The electro-optical conversion circuit 604 calculates a normalized linear version of the intermediate luma Yim, the intermediate luminance Ln_im. It applies the inverse of Equation 3 (the same definition of RHO). For the RHO value, in fact, in the simplest situation where the output image has the same maximum luminance as the input image and only adjusts the luminance of darker pixels, the same PL_V_HDR value is used. However, in the case of display adaptation to MDR maximum luminance, the PL_D value is used to calculate the appropriate RHO value that characterizes the specific shape (steepness) of the v-function. To achieve this, an output maximum determination circuit 671 exists in the color processing circuit 600. It is generally a logic processor that determines whether to use PL_V_HDR or PL_D, respectively, as the value for determining the maximum luminance of the output image PL_O, i.e., the RHO of the EOTF applied by the electro-optical conversion circuit 604, for the configured situation (those skilled in the art will understand that the situation may be configured with a fixed formula in some specific fixed variants).

[0179] The final surround adjustment circuit 605 performs a linear additive offset in the luminance domain by calculating the final normalized luminance Ln_f using the following equation: Ln_f=(Ln_im-(mL_De2 / PL_O)) / (1-(mL_De2 / PL_O)) [Formula 9] mL_De2 is the second minimum luminance of the end-user display (a.k.a., second end-user display minimum luminance), and is typically input via a third minimum metadata input 695 (connected to a light meter via intermediate processing circuitry). It differs from the first minimum luminance of the end-user display mL_De in that mL_De further includes a characteristic value of the display's physical black (mB_fD), while mL_De does not, and characterizes only the amount of ambient light that degrades the displayed image (e.g., by reflection), i.e., only the typically equal mB_sur.

[0180] Inside the final ambient adjustment circuit 605, the mL_De2 value is normalized by the applicable PL_O value, ie, for example, PL_D.

[0181] Finally, in most variants, it is advantageous if the usual (ie, unnormalized) output luminance L_o appears, which is realized by a multiplier 606 that calculates: L_o=Ln_f*PL_O [Equation 10] That is, it is normalized by the maximum applicable luminance of the output image.

[0182] Advantageously, some embodiments do this in color processing that not only adjusts for darker brightness for ambient lighting conditions, but also optimizes for reduced maximum display brightness. In this scenario, brightness mapping circuit 602 applies an appropriate calculated display-optimized brightness mapping function FL_DA(t), which is typically loaded into brightness mapping circuit 602 by display optimization circuit 670. Those skilled in the art will understand that the particular manner of display optimization is merely a variable part of such an embodiment and is not typical for ambient adaptive elements, but some examples are shown using Figures 4 and 5 (the input of configurable PL_O values ​​to 670 is not depicted to avoid overcomplicating Figure 6, as those skilled in the art will understand this). Generally, display adaptation has the property that the smaller the difference between the input maximum luminance and the output maximum luminance, the closer the function approaches the diagonal. That is, the closer the desired maximum luminance of the output image (i.e., PL_V_MDR = PL_D) is to the input maximum luminance (generally assuming PL_V_HDR), and therefore the farther it is from the maximum luminance of the second reference grading (generally PL_V_SDR), the flatter the shape of the function (i.e., a "lighter" regrading function version FL_DA between F_L and the diagonal). Note that downgrading generally has a convex function, which means that darker luminances (below some midpoint) are relatively boosted at the expense of compressing brighter luminances. Therefore, in downgrading, FL_DA generally has a less steep slope to boost the darkest input luminances below the F_L function (which performs full regrading to the second reference grading at the other end).

[0183] Below, a second innovation for optimizing image contrast is taught, which is useful in brighter ambient conditions. These elements may be used in conjunction with the ambient adjustments described above in various embodiments, but each innovation may also be applied independently of the others.

[0184] This method of processing an input image, particularly to improve contrast, generally consists of obtaining a reference luminance mapping function (F_L) associated with the input image, which defines the need for regrading by a luminance (or equivalently luma) mapping to the luminance of a corresponding secondary reference image.

[0185] The input image generally also serves as the first grading reference image.

[0186] The output image generally corresponds to a maximum luminance situation midway between the maximum luminances of the two reference grading images (a.k.a. reference gradings), and is generally calculated for a corresponding MDR display (a medium dynamic range HDR display compared to the master HDR input image), and the optimized output image maximum luminance relationship is PL_V_MDRPL_D_MDR.

[0187] Display Adaptation (of any embodiment) applies as usual to Display Adaptation for lower maximum brightness displays, but here it is applied specifically differently (i.e. most of the technical elements of Display Adaptation remain the same, but some are changed).

[0188] The display adaptation process determines an adaptive luminance mapping function (FL_DA), which is based on a reference luminance mapping function (F_L). This function F_L may be static, i.e., the same for several images (in which case, for example, FL_DA will still change if the ambient lighting changes significantly or upon user control actions), but it may also change over time (F_L(t)). The reference luminance mapping function (F_L) typically comes from the content creator, but may also come from an optimal re-grading function computing automaton in the device receiving the video images (just like the offset determination embodiment described in Figures 6 to 8).

[0189] An adapted luminance mapping function (FL_DA) is applied to the input image pixel luma to obtain the output luminance.

[0190] An important difference with existing display adaptations (while the definition of the metric, i.e., the mathematics for finding the various maximum luminance values, and the direction of the metric are the same; e.g., the metric is scaled by having one point at any position on the diagonal corresponding to the Yn_CC0 luma normalized to 1, and another point somewhere on the locus of the F_L function, e.g., vertically upward, that corresponds to the output luma or luminance of F_L when the other coordinate of the metric positioning end point is the input Yn_CC0 luma to the function F_L) is now that the position on the metric (more precisely, all its scaled versions due to the shape of the F_L function) to obtain the adaptive luminance mapping function (FL_DA) is calculated based on the adjusted maximum luminance value (PL_V_CO) rather than the required maximum value of the output image (generally, PL_V_MDR).

[0191] This adjusted maximum brightness value (PL_V_CO) is Obtaining an ambient illumination value (Lx_sur); Obtaining a relative ambient light level (SurMult) by dividing the reference illumination value (GenVwLx) by the ambient illumination value (Lx_sur); Multiplying the output maximum brightness (PL_V_MDR) by the relative ambient light level (SurMult) results in the maximum brightness value (PL_V_CO). is determined by.

[0192] The ambient illumination value (Lx_sur) can be obtained in various ways, for example a viewer can determine it empirically by checking the visibility of a test pattern, but it typically comes from measurements with a luminance meter 902, which is typically positioned appropriately relative to the display (e.g., at the bezel edge, facing approximately in the same direction as the screen front plate, or integrated into the side of the mobile phone, etc.).

[0193] The reference ambient value GenVwLx may also be determined in a variety of ways, but is generally fixed as it relates to what is expected to be reasonable ("average") ambient lighting for a typical target viewing situation.

[0194] For television display viewing, this is typically the viewing room.

[0195] The actual lighting in a living room can vary considerably, for example, depending on whether the viewer is watching during the day or at night, and in what room configuration (e.g., whether there is a small or large window and how the display is positioned relative to the window, or whether mood lamps are used at night, or whether another member of the family is doing precision work that requires a sufficient amount of lighting).

[0196] For example, even during the day, if the sky suddenly darkens significantly due to an approaching hailstorm, outdoor lighting may be as low as 200 lux (lx), and indoors, light levels from natural lighting are typically 100 times lower, so indoors it's only 2 lx. This begins to have a nocturnal appearance (which is especially strange during the day), so this is something that many users typically turn on at least one lamp for comfort, effectively raising the level again. Normal outdoor levels are 10,000 lx in winter to 100,000 lx in summer, so more than 50 times brighter.

[0197] However, other viewers may find it advantageous to watch video (especially HDR video) in the dark, for example to enjoy a horror movie more frighteningly and / or to see dark scenes better.

[0198] Although common in the Middle Ages, lighting from a single candle is now a lower limit. The reason is simply that, for city dwellers, at such levels, more light would leak from outdoor lamps like city lights. The candela used is defined as the brightness of a typical candle, so if you place a surface one meter away from a candle, you get 1 lx, which still allows you to see things, but makes text on paper, for example, very hard to read. (For reference, 1 lx is also typical of an outdoor moonlit scene.) So, if you light a five-meter-wide room with several candles, that level of illumination is achieved. Even a single 40-watt incandescent bulb already produces about 40 times more brightness than a candle, so for most viewers, one or a few such lamps would result in a more typical ambient light level. So, for viewing in moderate (atmospheric) light, you can expect something like k*10 lx. However, the video is defined so that it can be viewed properly during the day, where the illumination is n*50lx (e.g. a 200W light bulb set placed about 2 meters away gets about 3000 / 50lux; looking at a recipe display in the kitchen requires about three times higher light levels to safely perform cooking tasks like cutting).

[0199] In mobile / outdoor situations, the light level is higher, for example, when sitting near a train window or in the shade under a tree, etc. Then the light level is, for example, 1000 lx. Without intending to be limiting, let us assume that a good value of GenVwLx for video viewing of television programs is 100 lux.

[0200] Assume that the light sensor measures Lx_sur=550 lux.

[0201] The relative ambient light level is then calculated as follows: SurMult=GenVwLx / Lx_sur=100 / 550=0.18. [Formula 11] Then the adjusted maximum brightness value PL_V_CO for a 1000 nit MDR display is: PL_V_CO=PL_D*SurMult=1000*0.18=180nit [Formula 12]

[0202] In embodiments, this is further controlled by the user, who may feel that the automatic contrast correction is too strong or, conversely, too weak (of course, the automatic setting may be preferable in some systems and / or situations to avoid overly bothering the viewer).

[0203] In this case, the display adaptation uses the user-adjusted maximum luminance value (PL_V_CO_U) in the display adaptation process in place of (i.e., instead of) the adjusted maximum luminance value PL_V_CO, which is generally PL_V_CO_U=PL_V_CO*UCBval [Formula 13] is defined by

[0204] The user control value UCBVal is controlled by appropriately scaled user input, e.g., the slider setting (UCBSliVal, e.g., decreasing symmetrically near zero offset or starting at zero offset, etc.) is scaled in such a way that when the slider is at its maximum, the user does not dramatically change the contrast, e.g., so that all dark image areas appear almost like bright HDR white.

[0205] For this purpose, the device (e.g., display) manufacturer pre-designs the appropriate intensity value (EXCS), and then the formula is: UCBVal=PL_D*UCBSliVal*EXCS [Formula 14] is.

[0206] For example, if we want 100% to correspond to an additional maximum luminance-based contrast change of 10%, we get: PL_D*1*EXCS=0.1*PL_D, therefore EXCS=0.1 etc. (a value of 0.75 has been found to work well in certain embodiments).

[0207] FIG. 9 shows possible implementation elements in a typical device configuration.

[0208] The apparatus for processing input images to obtain an output image (also known as an ambient-optimized display optimization apparatus) 900 has a data input (920) for receiving a reference luminance mapping function (F_L), which is metadata associated with the input image. This function again specifies the relationship between the luminance of a first reference image and the luminance of a second reference image. These two images are again, in some embodiments, graded by the video creator and co-communicated with the video itself as metadata, e.g., via satellite television broadcast. However, the appropriate re-grading luminance mapping function F_L is also determined by a re-grading automaton within the receiving device, e.g., a television display. The function changes over time (F_L(t)).

[0209] The input image generally serves as a first reference grading based on which the display-optimized image and the environment-optimized image are determined, which are generally HDR images. The maximum luminance of the display-adapted output image (i.e., output maximum luminance PL_V_MDR) generally falls between the maximum luminances of the two reference images.

[0210] The device typically includes, or is equivalently connected to, an illuminance meter (902) configured to determine the amount of ambient light falling on the display, i.e., the display on which an optimized image for viewing is being provided. This amount of illumination from the viewing environment is expressed as an ambient illumination value (Lx_sur) in units of lux, or, as used in divisions, as an equivalent luminance in units of nits.

[0211] The display adaptation circuit 510 is configured to determine an adaptive luminance mapping function (FL_DA), which is based on the reference luminance mapping function (F_L). It also actually performs pixel color processing and therefore includes a luminance mapper (915) similar to the color converter 202 described above. Color processing is also involved. The configuration processor 511 makes the actual decision of the (ambient-optimized) luminance mapping function to be used before performing pixel-by-pixel processing of the current image. Such an input image (513) is received via an image input (921), e.g., an IC pin, which itself is connected to an image source to the device, e.g., an HDMI® cable, etc.

[0212] An adaptive luminance mapping function (FL_DA) is determined based on the reference luminance mapping function F_L and the value of the maximum luminance, according to a variation of the display adaptation algorithm described above, but where the maximum luminance is not the typical maximum luminance of the connected display (PL_D), but a specially adjusted maximum luminance value (PL_V_CO) adjusted for the lighting conditions of the viewing environment (and potentially further user corrections). The luminance mapper applies the adaptive luminance mapping function (FL_DA) to the input pixel luminance to obtain the output luminance.

[0213] To calculate the adjusted maximum brightness value (PL_V_CO), the device includes a maximum brightness determination unit (901), which is connected to the display adaptation circuit 510 and supplies this adjusted maximum brightness value (PL_V_CO) to the display adaptation circuit (510).

[0214] This maximum luminance determination unit (901) retrieves a reference illumination value (GenVwLx) from memory 905 (e.g., this value may be pre-stored by the device manufacturer, or may be selectable based on what type of image is coming in, or may be loaded with typical intended ambient metadata to be associated with the image, etc.). It also retrieves a maximum luminance (PL_V_MDR), which may be, for example, a fixed value stored in the display or may be configurable in a device (e.g., a set-top box or other image pre-processing device) that can provide images to various displays.

[0215] The adjusted maximum brightness value (PL_V_CO) is determined by the maximum brightness determination unit (901) as follows: first, the relative ambient light level (SurMult) is obtained by dividing the reference illumination value (GenVwLx) by the ambient illumination value (Lx_sur), and then the adjusted maximum brightness value (PL_V_CO) is determined as the result of multiplying the output maximum brightness (PL_V_MDR) by the relative ambient light level (SurMult).

[0216] In some embodiments, the user (viewer) further controls the automatic ambient optimization of the display adaptation according to their preferences.

[0217] Coupled thereto is a user interface control component 904, e.g., a slider, which allows a user to set a higher or lower value, e.g., slider setting UCBSliVal, which is input to a user value circuit 903, which communicates the user control value UCBVal, calculated in Equation 14, to the maximum brightness determination unit 901 (PL_D=PLV_MDR).

[0218] FIG. 10 shows an example of processing with a particular luma mapping function F_L. Assume that the reference mapping function (1000) is a simple luma identity transformation, i.e., clipping above the 600 nit maximum display capability. The luma mapping here assumes that the input normalized luma Yn_i, represented by the equivalent psychovisually equalized luma plot, which can be calculated according to Equation 3, corresponds to an HDR input luma, which is a 1000 nit maximum luma HDR image (i.e., an RHO of PL_V_HDR=1000 nit is used in the equation). For the output normalized luma Yn_o, assume an exemplary display of 600 nits; therefore, the output normalized luma Yn_o is converted back to luma by using the RHO value corresponding to 600 nit. For convenience, the luma corresponding to the luma position is shown on the right and above. Therefore, this selected F_L reference mapping function 1000 performs equal compression in the visually equalized luma domain. The (ambient) adaptive luminance mapping function FL_DA is shown as curve 1001. On the one hand, an offset Bko is seen, which depends, among other things, on the leak black of the display. On the other hand, a curvature that mainly boosts the darkest black is seen, which is due to residual effects of non-linear psychovisually equalized processing in the luma domain.

[0219] 11 shows an example of a contrast boosting embodiment (ambient lighting offset is set to zero, but as before, both processes are combined). When modifying the luminance mapping in the relatively normalized psychovisually equalized luma domain, this relative domain starts from a value of zero. The master HDR image starts with a black value, but at the lower output, a black offset is used to map this to zero (or to the display's minimum black; note: displays may have different behavior for darker inputs, e.g., clipping, so a true zero can be placed there).

[0220] When using modified display adaptation, the black zero input generally maps to black zero, whatever the value of the adjusted maximum luminance value (PL_V_CO). Generally, the zero point of the output luma starts at the virtual black level mL_VD (ideally) or the minimum black mB_fD of the actual end-user display, but in either case, there is a brightness difference in the darkest colors, which results in better visible contrast in dark pictures such as night scenes (the input luminance histogram 1010 is stretched on the normalized luma axis of the output luma as the output luminance histogram 1011, which gives sufficient image contrast despite the lower PL_V_MDR = PL_D = 600 nit). When the two methods are combined, the entire process to obtain the appropriate black level offset Bko is optimally transferred to the process described in, for example, FIG. 6 (i.e., the adjusted contrast process can be performed using an image starting at zero and a function starting at zero). In practice, we simply always define the F_L curve starting from zero (map zero HDR luminance or luma to zero output luminance or luma) because the display-adaptive luminance (luma) mapping simply occurs anyway, regardless of whether zero actually occurs in the input image.

[0221] Note that in another embodiment, there is no illuminance meter, only user controls (904, 903). This embodiment then varies the contrast based on user input as if the automatic ambient correction of PL_V_CO in Equation 13 were not present. The maximum luminance PL_D of the end user display then serves as the neutral point according to Equation 15. PL_V_CO_U=PL_D*UCBval [Equation 15]

[0222] The algorithmic components disclosed herein may actually be implemented (wholly or partially) in hardware (e.g., as part of an application-specific IC) or in software running on a dedicated digital signal processor or a general-purpose processor, etc.

[0223] From the inventors' presentation, it should be possible for a person skilled in the art to understand which components are optional improvements and can be realized in combination with other components, and how the (optional) steps of a method correspond to the respective means of an apparatus, and vice versa. The term "apparatus" in this application is used in the broadest sense, i.e., to refer to a group of means that enable the realization of a specific purpose, and thus includes, for example, (a small circuit portion of) an IC, or a dedicated device (such as a device with a display), or a part of a networked system, etc. The term "arrangement" is also intended to be used in the broadest sense, and thus includes, inter alia, a single device, a part of a device, a collection of (parts of) cooperating devices, etc.

[0224] The meaning of computer program product should be understood to encompass any physical realization of a set of commands that enables a general-purpose or dedicated processor, after a series of loading steps (including intermediate conversion steps, for example, conversion into an intermediate language and a final processor language), to input the commands to the processor and to execute any of the characteristic functions of the invention. In particular, a computer program product is realized as data on a carrier, for example a disk or tape, data in a memory, data traveling via a network connection (wired or wireless), or program code on paper. Apart from the program code, characteristic data required by the program may also be embodied in the computer program product.

[0225] Some of the steps required in the operation of the method may already be present in the functionality of the processor instead of those described in the computer program product, such as data input and output steps.

[0226] It should be noted that the above embodiments illustrate, rather than limit, the present invention. For the sake of brevity, not all these options are detailed, as those skilled in the art can easily realize that the presented examples map to other areas of the claims. Apart from the combinations of elements of the present invention as combined in the claims, other combinations of elements are possible. Any combination of elements may be realized in a single dedicated element.

[0227] Any reference signs placed between parentheses in the claims are not intended to limit the claim. The words "comprises," "comprises," "having," etc. do not exclude the presence of elements or aspects not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements.

Claims

1. 1. A method of processing an input image to obtain an output image, the method comprising: obtaining a starting luma for a pixel of the input image, the starting luma being related to an input luminance via an opto-electrical transfer function; obtaining a minimum luminance of a target display, the target display corresponding to an end user display onto which the output image may be provided for displaying the output image; obtaining a first minimum luminance of the end user display in a viewing room, wherein the first minimum luminance of the end user display is dependent on an amount of illumination in the viewing room; calculating a difference by subtracting the first minimum luminance of the end user display from the minimum luminance of the target display; converting the difference into a luma difference by applying the difference as an input to the photoelectric transfer function, thereby providing the luma difference as an output; mapping the starting luma by applying a linear function that applies the luma difference multiplied by −1.0 as an additive constant and multiplies the starting luma using the luma difference incremented by a value of 1.0 as a multiplier, thereby resulting in a mapped luma; transforming the mapped luma by applying the inverse of the photoelectric transfer function to obtain a normalized intermediate luminance; subtracting a second minimum luminance of an end-user display divided by a maximum luminance of the output image from the intermediate luminance and scaling the subtraction by 1.0 minus a result of the second minimum luminance of an end-user display divided by the maximum luminance of the output image to obtain a final normalized luminance; multiplying the final normalized luminance by the maximum luminance of the output image to obtain an output luminance; outputting the output luminance in a color representation of a pixel of the output image; A method comprising:

2. 2. The method of claim 1, comprising determining the maximum luminance of the output image as either a maximum luminance of an end user display capable of displaying the output image or a maximum luminance of the input image.

3. a display optimization step for calculating a display optimized luminance mapping function based on the reference luminance mapping function and the maximum luminance of the end user display; using the display-optimized luminance mapping function to map input lumas of pixels of the input image to output lumas, the output lumas being the starting lumas; and the display-optimized luminance mapping function has a slope that is less steep for the darkest input luma compared to a slope of the reference luminance mapping function; 3. The method of claim 1, wherein the reference luminance mapping function specifies a relationship between the luminance of a first reference image and a luminance of a second reference image having respective first and second reference maximum luminances, and wherein the maximum luminance of an end-user display falls between the first and second reference maximum luminances.

4. 4. The method of processing an input image of claim 3, wherein the second reference image is a standard dynamic range image whose second reference maximum luminance is equal to 100 nits.

5. and wherein the step of calculating the display optimized luminance mapping function comprises finding a position on a metric that determines a location of a maximum luminance value, the position corresponding to the maximum luminance of the end user display; a first endpoint of the metric corresponds to a first maximum luminance of the input image, and a second endpoint of the metric corresponds to the maximum luminance of the second reference image; the first endpoint of the metric is located, for any normalized input luma, at a diagonal point having horizontal and vertical coordinates equal to the normalized input luma; The method of claim 3 , wherein the second endpoint is associated with an output value of the reference luminance mapping function determined by the direction of the metric.

6. 1. An apparatus for processing an input image to obtain an output image, said apparatus comprising: an input circuit for obtaining a starting luma; a first minimum metadata input for receiving a minimum luminance of a target display, the target display corresponding to an end user display on which the output image may be provided for displaying the output image; a second minimum metadata input for receiving a first minimum luminance of the end user display in a viewing room, wherein the first minimum luminance of the end user display is dependent on an amount of lighting in the viewing room; and a brightness difference calculator for calculating a difference by subtracting the first minimum brightness of the end user display from the minimum brightness of the target display; a second photoelectric conversion circuit for converting the difference into a luma difference by applying the difference as an input to a photoelectric transfer function, thereby providing the luma difference as an output; and a linear scaling circuit for mapping the starting luma by applying a linear function that applies the luma difference multiplied by −1.0 as an additive constant and multiplies the starting luma using the luma difference incremented by a value of 1.0 as a multiplier, thereby resulting in a mapped luma; and an electrical-to-optical conversion circuit for converting the mapped luma by applying an inverse of the electrical-to-optical transfer function to obtain a normalized intermediate luminance; a third minimum metadata input for receiving a second minimum luminance of said end user display; a final surround adjustment circuit for subtracting a second minimum luminance of an end user display divided by a maximum luminance of the output image from the intermediate luminance and scaling the subtraction by 1.0 minus a result of the second minimum luminance of an end user display divided by the maximum luminance of the output image to obtain a final normalized luminance; a multiplier for multiplying the final normalized luminance by the maximum luminance of the output image to obtain an output luminance; a pixel color output for outputting the output luminance in a color representation of a pixel of the output image; 1. An apparatus comprising:

7. 7. The apparatus of claim 6, further comprising a maximum determination circuit for determining the maximum luminance of the output image as either a maximum luminance of an end user display capable of displaying the output image or a maximum luminance of the input image.

8. a display optimization circuit for calculating a display-optimized luminance mapping function based on the reference luminance mapping function and a maximum luminance of the end-user display; a luminance mapping circuit for applying the display-optimized luminance mapping function to an input luma to obtain an output luma, the output luma being the starting luma; Including, the display-optimized luminance mapping function has a slope that is less steep for the darkest input luminance compared to the slope of the reference luminance mapping function; 8. The apparatus of claim 6 or 7, wherein the reference luminance mapping function specifies a relationship between the luminance of a first reference image and a luminance of a second reference image having respective first and second reference maximum luminances, and wherein the maximum luminance of an end user display falls between the first and second reference maximum luminances.

9. 9. The apparatus of claim 8, wherein the second reference image is a standard dynamic range image whose second reference maximum luminance is equal to 100 nits.

Citation Information

Patent Citations

  • Ambient light-adaptive display management

    US20190304379A1

  • Ambient Headroom Adaptation

    US20210096023A1

  • Adjustment of display optimization behaviour for HDR images

    WO2021004839A1