Display-optimized HDR video contrast adaptation

JP2024519606A5Active Publication Date: 2025-05-07KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023568223
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-07
Filing Date
2022-04-28
Publication Date
2025-05-07
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

Existing technologies fail to optimally adapt high dynamic range (HDR) video images for displays with varying ambient light levels, leading to suboptimal viewing experiences due to inadequate consideration of minimum and maximum brightness adjustments.

Method used

A method involving a reference luminance mapping function (F_L) and an adaptive brightness mapping function (FL_DA) is applied to HDR video images, allowing for dynamic adjustment of brightness levels based on the display's maximum brightness (PL_V) and ambient light conditions, ensuring optimal image contrast and visibility across different viewing environments.

Benefits of technology

The method enhances image contrast and visibility in varying ambient light conditions, providing a visually better-looking HDR video experience by effectively mapping brightness levels to match the display's capabilities and environmental lighting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In order to obtain in a practical manner an image that can be viewed better for a variety of potentially quite different viewing environment light levels, the inventor proposes an apparatus 900 for processing an input image having pixels with input luminance that is within a first luminance dynamic range DR_1 having a first maximum luminance PL_V_HDR. The apparatus includes an image input 921 configured to obtain an input image 513, and a data input 920 for receiving a reference luminance mapping function F_L that is metadata associated with the input image specifying how the luminance should be re-graded between two reference images, whereby the output image has an output maximum luminance PL_V_MDR that is different from the first maximum luminance, the apparatus further includes a user value circuit 903 configured to determine and output a user correction value UCBVal as set by a human user of the apparatus, and a maximum luminance determination unit 901 configured to obtain the user correction value UCBVal from the user value circuit 903 and to output an adjusted maximum luminance value PL_V_CO as a result of subtracting the user correction value UCBVal from the output maximum luminance PL_V_MDR, the apparatus further includes a user correction circuit 903 configured to determine and output a user correction value UCBVal from the user value circuit 903 and to output an adjusted maximum luminance value PL_V_CO as a result of subtracting the user correction value UCBVal from the output maximum luminance PL_V_MDR, The display adaptation unit is further configured to: determine an adaptive luminance mapping function FL_DA based on L; and the calculation of the adaptive luminance mapping function FL_DA includes finding a position pos on the metric SM corresponding to the adjusted maximum luminance value PL_V_CO, where a first end point of the metric corresponds to a first maximum luminance PL_V_HDR and a second end point of the metric corresponds to a maximum luminance of a second reference image, where for any normalized input luma Yn_CC0, the first end point of the metric is located at a diagonal point having horizontal and vertical coordinates equal to the normalized input luma, and the second end point is located on the locus of the reference luminance mapping function F_L; and the display adaptation unit is configured to apply the adaptive luminance mapping function FL_DA to the input luminance, or an input luma encoding the input luminance, to obtain an output luminance or output luma, and to output these output lumas or output luminances as pixel colors of the output image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method and apparatus for adapting image pixel brightness of high dynamic range video to provide a desired look for the conditions under which the HDR video is displayed under the ambient lighting conditions of a particular viewing site. [Background technology]

[0002] A few years ago, novel techniques for high dynamic range (HDR) video coding were introduced, inter alia, by the applicant (see, for example, WO2017157977).

[0003] Video coding is generally primarily or solely concerned with creating or more precisely defining color codes (e.g., luma and two chromas per pixel) to represent an image, which is distinct from knowing how to optimally display an HDR image (e.g., the simplest method simply utilizes a highly nonlinear electro-optical transfer function OETF to convert the desired luminance to, say, a 10-bit luma code, and vice versa, converting those video pixel luma codes to the luminance to be displayed by mapping the 10-bit electrical luma code to the optical pixel luminance to be displayed using an inversely formed electro-optical transfer function EOTF, but more complex systems deviate in several directions, particularly by decoupling the coding of the image from the specific use of the coded image).

[0004] Encoding and handling HDR video stands in stark contrast to the way older video technologies were used, by which all video was coded until recently, and which today is called Standard Dynamic Range (SDR) video coding (also known as Low Dynamic Range video coding, or LDR), which began in the analog era as PAL or NTSC and transitioned to Rec. 709-based coding, e.g., MPEG2 compression, in the digital video era.

[0005] While a satisfactory technology for communicating moving images in the 20th century, advances in display technology beyond the physical limitations of the electron beam of the 20th century CRT or the world-famous TL backlit LCDs have made it possible to show images with pixels that are significantly brighter (and potentially even darker) than older displays, which has created the need to be able to encode and create such HDR images.

[0006] In fact, starting from the inability to encode very bright and sometimes even darker image objects in the SDR standard (8-bit Rec.709) for various reasons, ways were first invented that could technically represent colors with such a wide luminance range, and then, one by one, all the rules of video technology had to be rethought and, in many cases, reinvented.

[0007] The Rec.709 SDR luma code definition could only encode (at 8 or 10-bit luma) approximately 1000:1 luminance dynamic range due to the approximately square root OETF function shape luma Y_code=power(2,N)*sqrt(L_norm), where N is the number of bits in the luma channel and L_norm is the normalized version of the physical luminance between 0 and 1.

[0008] Furthermore, in the SDR era, the absolute luminance to be displayed is not defined, so in practice the maximum relative luminance L_norm_max=100%, or 1, was mapped via square root OETF to the maximum normalized luma code Yn=1, which corresponds to, for example, Y_code_max=255. This has some technical differences compared to making absolute HDR images, namely, an image pixel coded to be displayed as 200 nits will ideally (i.e., when possible) be displayed as 200 nits on all displays, and not as a totally different display luminance. In the relative paradigm, a 200 nit coded pixel luminance will be displayed as 300 nits on a brighter display, i.e., a display with a brighter maximum displayable luminance PL_D (aka display maximum luminance), and will be displayed as 100 nits on a display with less capabilities, for example. Furthermore, it should be noted that absolute coding works with a normalized luminance representation or normalized 3D color gamut, but 1.0, for example, uniquely means 1000 nits.

[0009] On displays, such relative images were usually displayed somewhat heuristically by mapping the brightest luminance of the video to the brightest displayable pixel luminance (which was done automatically via the electronic driving of the display panel with the maximum luma Y_code_max, without the need for further luminance mapping). So if you bought a 200nit PL_D display, white would look twice as bright as it would on a 100nit PL_D display, but given factors such as eye adaptation, that wasn't considered to matter much, except to make the same SDR video image a brighter, more viewable, and somewhat prettier version.

[0010] Conventionally, when we speak of an SDR video image today (in an absolute framework), it generally has a video peak luminance of PL_V=100 nits (e.g., agreed upon in accordance with standards), and so in this application we consider the maximum luminance of an SDR image (or SDR grading) to be exactly that value, or generalized around that value.

[0011] Grading in this application is intended to mean either an activity or a resulting image in which pixels are given brightness as desired, for example, by a human color grader or an automaton. When viewing an image, for example, when designing an image, there are several image objects, and ideally, one wants to give the pixels of those objects a brightness that is spread around an average brightness that is optimal for that object, also taking into account the overall image and scene. For example, if one has an image capacity available such that the brightest encodable pixel of the image is 1000 nit (image or video maximum brightness PL_V), one grader may choose a brightness value between 800 nit and 1000 nit for the pixels of the explosion to make the explosion look punchy, while another filmmaker may choose an explosion that is not brighter than 500 nit, for example, so that it does not interfere too much with the rest of the image at that moment (of course, the technology can handle both situations).

[0012] The maximum brightness of an HDR image or video can vary considerably and is generally communicated with the image data as metadata for the HDR video or image (typical values ​​are, for example, 1000 nits, or 4000 nits, or 10,000 nits, but are not limiting; generally, one says to have an HDR image when the PL_V is at least 600 nits). If a video creator chooses to define an image as PL_V=4000 nits, the video creator can of course choose to create a brighter burst, but relatively, it does not reach the 100% level of PL_V, but only reaches, for example, 50% for such a high PL_V definition of the scene.

[0013] An HDR display has a maximum capability, i.e., a maximum displayable pixel luminance of, for example, 600 nits, or 1000 nits, or N times 1000 nits (starting with the lowest HDR display). The maximum (or peak) luminance of the display, PL_D, is separate from the maximum luminance of the video, PL_V, and the two should not be confused. Video creators generally cannot make an optimal video for each possible end-user display (i.e., the capabilities of the end-user display are optimally used by the video, and the maximum luminance of the video never exceeds (ideally) the maximum luminance of the display, but is not lower either, i.e., some of the video images should have at least some pixels with pixel luminance L_p=PL_V, which in successive optimizations to a particular display would further include PL_V=PL_D).

[0014] Creators make some of their own decisions (e.g., what kind of content to capture and how) and generally make videos with PL_V very high to serve the highest PL_D displays of their intended audience, at least now, and perhaps in the future, as higher PL_D displays emerge.

[0015] A secondary problem then emerges of how to best display an image with a peak luminance PL_V on a display with a lower (often very low) display peak luminance PL_D, which is called display adaptation. Even in the future, there will still be displays that require a lower dynamic range image than the created, e.g., 2000 nit PL_V image received over a communication medium. In theory, a display will always re-grade, i.e., map, the luminance of the image pixels such that they are displayable by its own internal heuristics, but if the video creator is careful enough in determining pixel luminance, it is beneficial for the video creator to further indicate how the image should be display-adapted to the lower PL_D value, and ideally the display will follow to a large extent what is technically required.

[0016] With regard to the darkest displayable pixel luminance BL_D, the situation is more complicated. Some of it is the fixed physical properties of the display, e.g. LCD cell leakage light, but even with the best displays, what the viewer can ultimately discern as a distinct darkest black also depends on the illumination of the viewing room, which is not a well-defined value. This illumination is characterized as an average illuminance level, e.g. in lux, but for video display purposes is more conveniently characterized as a minimum pixel luminance. It is also generally more relevant to the human eye than the appearance of bright or medium luminance. The reason is that if the human eye is looking at many high brightness pixels, the darker pixels, especially their absolute luminance, become less relevant. However, it can be assumed that the eye is not a limiting factor, for example, when looking at a generally dark scene image, but still masked by ambient light in front of the display screen. If we assume that humans can see a noticeable difference of just 2%, there is a darkest drive level (or luma) b above which the next darker luma level (i.e., displaying a luminance level that is X% higher, e.g., 2% more) can still be seen.

[0017] In the LDR era, there was no concern at all about the darkest pixels. There was mainly concern about average brightness around 1 / 4 of the maximum PL_V=100nit. If the image was exposed near this value, everything in the scene looked nice and bright and colorful, except for clipping of the bright parts of the scene that exceeded the maximum 100%. For the darkest parts of the scene, if it was important enough, the captured image was made with a sufficient amount of base lighting in the recording studio or filming environment. If some part of the scene was not visible well, for example because it was buried in the code Y=0, it was considered normal.

[0018] Therefore, if nothing further is specified, it is assumed that the blackest black is zero, or in practice something like 0.1 nit or 0.01 nit, etc. In such situations, engineers are more interested in pixels brighter than average in the encoded and / or displayed HDR image.

[0019] In terms of coding, the differences between HDR and SDR are not only physical differences (more different pixel brightnesses that can be displayed on a display with a larger dynamic range capability), but also technical differences including different luma code allocation functions (for which the OETF is used, or in absolute terms the inverse of the EOTF), and potentially even further technical HDR concepts such as additional dynamically (for each image, or for each set of temporally consecutive images) changing metadata that specifies how to re-grade the pixel brightnesses of various image objects to obtain images in a secondary dynamic range that differs from the starting image dynamic range (the two brightness ranges typically end with peak brightnesses that differ by at least a factor of 1.5).

[0020] A simple HDR codec, the HDR10 codec, has been introduced to the market, which is used, for example, to create the recently released Black Jewel Box HDR Blu-ray. This HDR10 video codec uses a logarithmic rather than square root shaped function as the OETF (inverse EOTF), i.e. the so-called Perceptual Quantizer (PQ) function, which is standardized in SMPTE 2084. Instead of being limited to 1000:1 as in the Rec. 709 OETF, this PQ OETF allows to define the luma at many more luminances (ideally to be displayed) that are sufficient for practical HDR video production, i.e. between 1 / 10,000 nits and 10,000 nits.

[0021] Please note that the reader should not confuse HDR simply with a large number of bits in the luma codeword. That applies to a linear system, like the amount of bits in an analog-to-digital converter, and in fact the amount of bits follows the base 2 logarithm of the dynamic range. However, the code allocation function has a fairly non-linear shape, so in theory we could define HDR images with only 10-bit luma (and even 8-bit HDR images per color component) however we wanted, which would bring the advantage of reusing systems already deployed (e.g. an IC may have a certain bit depth, or a video cable, etc.).

[0022] After luma calculation, we can have 10 bit planes of pixel luma Y_code, to which two chrominance components Cb and Cr per pixel are added as chrominance pixel planes. This image is further processed mathematically classically "as if" it were an SDR image, e.g. MPEG-HEVC compressed. The compressor does not actually need to care about pixel color or luminance.

[0023] However, the receiving device, e.g., a display (or indeed a decoder thereof), generally needs to perform a correct color interpretation of the {Y,Cb,Cr} pixel colors to display a correct looking image and not, e.g., an image with washed out colors.

[0024] This is typically handled by communicating additional image definition metadata along with the three pixelated color component planes, which defines the image encoding, such as an indication of which EOTF is used (for which we assume, without limitation, that the PQ EOTF (or OETF) is used), and the PL_V value, etc.

[0025] More sophisticated codecs include further image defining metadata, e.g., handling metadata, e.g., functions that specify how to map a normalized version of the luminance of the first image up to PL_V=1000 nit to the normalized luminance of a secondary reference image, e.g., an SDR reference image with PL_V=100 nit (as explained in more detail in FIG. 2).

[0026] For the convenience of the less knowledgeable reader, some interesting aspects are briefly described in Fig. 1. Fig. 1 shows some typical illustrative examples of many possible HDR scenes that a future HDR system (e.g., connected to a 1000 nit PL_D display) needs to be able to process correctly. While the actual technical processing of pixel colors is done in different ways in different color space definitions, what is needed for regrading is shown as an absolute luminance mapping between luminance axes spanning different dynamic ranges.

[0027] For example, ImSCN1 is a sunny outdoor image from a Western movie with mostly bright areas. The first thing to be clear about is that the pixel luminance of any image is generally not a luminance that can actually be measured in the real world.

[0028] Even without further human intervention during the creation of the output HDR image (which acts as a starter image and is called the master HDR grading or image), by tweaking one parameter no matter how simple, the camera at least always measures the relative brightness at the image sensor since there is an aperture, so there is always some step involved where at least the brightest image pixels end up in the available encoded brightness range of the master HDR image.

[0029] For example, the specular reflection of the sun on a sheriff's star badge may be measured at over 100,000 nits in the real world, which cannot be displayed on a typical near-future display, nor would it be comfortable for a viewer watching a movie image in, for example, a dimly lit room in the evening. Instead, the video creator decides that 5000 nits is bright enough for a pixel on the badge, and therefore, if this pixel is to be the brightest pixel in the movie, the video creator decides to make a video with PL_V=5000 nits. Although a relative pixel brightness measurement device only for the RAW version of the master HDR grade, the camera should also have a high enough native dynamic range (all pixels well above the noise floor) to make a good image. The pixels of the graded 5000 nit image are generally derived nonlinearly from the RAW image captured by the camera, e.g., the color grader takes into account aspects such as the typical viewing conditions that are not the same as the actual shooting location, i.e., standing in a hot desert. The best (highest PL_V) image selected to produce this scene ImSCN1, i.e. the 5000 nit image in this example, is the master HDR grading. This is the minimum required HDR data to be created and communicated, but in all codecs it is not the only data communicated, or in some codecs it is not even an image communicated at all.

[0030] Making such a codable high luminance range DR_1 available, for example between 0.001 nit and 5000 nit, allows content makers to provide viewers with a better experience of bright looking scenes, but also of course dimmer night scenes (when well graded throughout the movie), assuming the viewer also has a corresponding high-end PL_D=5000 nit display. A good HDR movie balances the luminance of various image objects, not only in one image, but also over time in the movie story or in the video material generally produced (e.g. a well-designed HDR soccer program).

[0031] The leftmost vertical axis of Figure 1 shows some (average) object luminances that one would like to see in a PL_V master HDR grading of 5000 nits, ideally intended for a PL_D display of 5000 nits. For example, in a movie, one creator might want to show cowboys in bright sunshine with a pixel luminance of about 500 nits (i.e., typically 10 times brighter than LDR, while another creator might want slightly less HDR punch, say 300 nits), thereby configuring the best way to display this Western imagery that will give the end consumer the best possible look.

[0032] The need for a higher dynamic range of luminance is more easily understood by considering an image in which, within the same image, there are fairly dark regions, such as the dark corners of the cave image ImSCN3, but which also have relatively large areas of very bright pixels, such as the sunlit exterior seen from the cave entrance, creating a different visual experience than, for example, the nighttime image of ImSCN2, where only street lights contain areas of pixels of high luminance.

[0033] Now, the problem is that at present, many consumers still have LDR displays, and even in the future, there is good reason to make two gradings of a movie, instead of the typical only encoding of the HDR image itself, so we need to be able to define an SDR image with PL_V_SDR=100nit that best corresponds to the master HDR image. This is a technical desire, which is separate from the technical choice about the encoding itself, and it states that, for example, if we know how to create (invert) the master HDR image and one of these secondary images from the other, we can choose to encode and communicate only one of the pair (effectively communicating two images for the price of one, i.e., only one image of pixel color component planes per video time instant).

[0034] In such a reduced dynamic range image, it is of course impossible to define a 5000 nit pixel brightness object such as the truly bright sun. The minimum pixel brightness or deepest black is also as high as 0.1 nit, rather than the more preferred 0.001 nit.

[0035] So anyway it should be possible to create this corresponding SDR image with a reduced luminance dynamic range DR_2.

[0036] This is done by some automatic algorithm in the receiving display, for example using a fixed luminance mapping function, or perhaps one conditioned by simple metadata like the PL_V_HDR value and potentially one or more other luminance values.

[0037] In general, however, more complex luminance mapping algorithms may be used, but in this application, without loss of generality, it is assumed that the mapping is defined by a global luminance mapping function F_L (e.g., one function per image) that defines, for at least one image, how all possible luminances occurring in the first image (i.e., for example, from 0.0001 to 5000) should be mapped to the corresponding luminances in the second output image (e.g., from 0.1 to 100 nit for an SDR output image). The normalization function is obtained by dividing the luminances along both axes by their respective maximum values. Global in this context means that the same function is used for all pixels of the image, regardless of further conditions, such as their location in the image (more general algorithms use, for example, several functions for pixels that can be classified according to some criteria).

[0038] Ideally, how all luminance should be redistributed along the available range of the SDR image of the secondary image should be decided by the video creator, because in the case of limitations, the video creator knows well how to sub-optimize for the reduced dynamic range so that the SDR image still looks at least as good as possible like the intended master HDR image. The reader can understand that actually defining (locating) such object luminance corresponds to defining the shape of the luminance mapping function F_L, the details of which are beyond the scope of this application.

[0039] Ideally, the shape of the function should also change for different scenes, i.e. a cave scene in a movie versus a sunny western scene a little later, or generally for different images in time - this is called dynamic metadata (F_L(t), where t denotes image time).

[0040] Now, ideally the content creator would create an optimal image for each situation, i.e. for each potentially served end-user display, e.g. a PL_D_MDR=800 nit display needing a corresponding PL_V_MDR=800 nit image, but that is generally too much effort for the content creator even in the most expensive offline video production.

[0041] However, it has been previously demonstrated by the applicant that it is sufficient to make (only) two different dynamic range reference gradings of the scene (generally at the extreme ends, e.g. 5000 nits being the highest required PL_V and 100 nits being generally sufficient as the lowest required PL_V), because then all other gradings can be derived automatically from those two reference gradings (HDR and SDR), e.g. via a (generally fixed, e.g. standardized) display adaptation algorithm applied to the end user's display receiving the information of the two gradings. In general, the calculations are performed in any video receiver, e.g. set-top box, television, computer, cinema equipment, etc. The communication channel of the HDR image can also be any communication technology, e.g. terrestrial or cable broadcasting, physical media such as Blu-ray discs, the Internet, communication channels to portable devices, professional inter-site video communication, etc.

[0042] This display adaptation generally also applies a luminance mapping function to, for example, pixel luminances of the master HDR image. However, the display adaptation algorithm needs to determine a luminance mapping function different from F_L_5000to100 (which is the reference luminance mapping function that connects the luminances of the two reference gradings), i.e., the display adaptation luminance mapping function FL_DA, which is not necessarily trivially related to the original mapping function between the two reference gradings F_L (there are several variants of the display adaptation algorithm). The luminance mapping function between the master luminance defined with a PL_V dynamic range of 5000 nits and the 800 nit intermediate dynamic range is written as F_L_5000to800 in this text.

[0043] Display adaptation is symbolically indicated (only for one of the average object pixel luminances) by an arrow, which maps, for example, to a slightly higher location (i.e., in such an image, the cowboy should be at least slightly brighter according to the chosen display adaptation algorithm) rather than to a location where the F_L_5000to100 function would "naively" expect to cross the 800 nit MDR image luminance range. So, while some more complex display adaptation algorithms may place the cowboy at a higher location as indicated, some customers are happy with the simpler location where the connection between the 500 nit HDR cowboy and the 18 nit SDR cowboy crosses the 800 nit PL_V luminance range.

[0044] In general, a display adaptation algorithm calculates the shape of the display adaptation luminance mapping function FL_DA based on the shape of the original luminance mapping function F_L (or the reference luminance mapping function, also known as the reference regrading function).

[0045] This description based on Fig. 1 constitutes the technical requirements of any HDR video encoding and / or processing system, and Fig. 2 illustrates some exemplary technical systems and their components for realizing the requirements (non-limiting) according to the applicant's codec approach. It will be understood by those skilled in the art that these components may be embodied in various devices, etc. Those skilled in the art will understand that this example is presented merely as a representative part of various HDR codec frameworks to provide background understanding of some principles of operation, and is not intended to specifically limit any of the embodiments of the inventive contributions presented below.

[0046] Although possible, the technical communication of two actual different images at each instant in time (HDR and SDR grading, each communicated as three respective color planes) is expensive, especially in terms of the amount of data required.

[0047] Also, it is not necessary since if it is known that all corresponding secondary image pixel luminances are calculated based on the luminances of the primary image and the function F_L, one can decide to communicate only the primary image and function F_L for each time instant as metadata (and can choose to communicate either the master HDR or SDR image as representative of both).The receiver knows the (generally fixed) display adaptation algorithm, so it determines the FL_DA function at the receiver end based on this data (additional metadata to control or guide the display adaptation may be communicated, but is not currently in place).

[0048] There are two modes for communicating a unique image and function F_L at each time instant.

[0049] In a first backwards compatible mode, an SDR image is communicated ("SDR communication mode"). The SDR image is displayed directly (without needing further luminance mapping) on ​​a traditional SDR display, but an HDR display needs to apply the F_L or FL_DA function to obtain an HDR image from the SDR image (or vice versa, depending on which variant of the function is communicated, i.e., upgrading or downgrading). The interested reader can find all the details of the applicant's exemplary first mode approach as standardized below.

[0050] ETSI TS 103 433-1 V1.2.1 (2017-08): High-Performance Single Layer High Dynamic Range System for use in Consumer Electronics devices; Part 1: Directly Standard Dynamic Range (SDR) Compatible HDR System (SL-HDR1).

[0051] Another mode communicates the master HDR image itself ("HDR communication mode"), i.e., for example, a 5000 nit image, and a function F_L that allows to calculate from it a 100 nit SDR image (or any other lower dynamic range image via display adaptation). The master HDR communication image itself is encoded, for example, by using the PQ EOTF.

[0052] Figure 2 further illustrates the overall video communication system: On the transmit side, we start with the source of the images 201. This can be anything ranging from a hard disk, to a cable output from, for example, a television studio, etc., depending on whether it is an offline produced video from, for example, an internet distribution company, or a real broadcast.

[0053] This results in a master HDR video (MAST_HDR) that has been color graded, for example, by a human color grader, a shaded version of the camera capture, or by an automatic brightness redistribution algorithm, etc.

[0054] In addition to the grading of the master HDR image, a set of often reversible color transformation functions F_ct is defined. Without loss of generality, we assume that this includes at least one luminance mapping function F_L (however, there may be further functions and data that specify, for example, how the saturation of a pixel should change from HDR to SDR grading).

[0055] This luminance mapping function defines the mapping between the HDR reference grading and the SDR reference grading, as described above (the latter in FIG. 2 is the SDR image Im_SDR to be communicated to the receiver; it may or may not have been data compressed, e.g. via MPEG or other image compression algorithms).

[0056] The color mapping of the color converter 220 should not be confused with that applied to the raw camera feed to obtain the master HDR video, which is assumed here to be already input, since this color conversion is to obtain the image to be communicated and at the same time what is needed for regrading as technically formulated in the luminance mapping function F_L.

[0057] In an exemplary SDR communication type (i.e., SDR communication mode), the master HDR image is input to a color converter 202 configured to apply F_L luminance mapping to the luminance of the master HDR image (MAST_HDR) to obtain all corresponding luminances that are written to the output image Im_SDR. For illustration purposes, let us assume that the shape of this function is fine-tuned for each shot of images of similar scenes in a movie by a human color grader using color grading software. The applied function F_ct (i.e., at least F_L) is written into the (dynamic, processing) metadata to be co-communicated with the image, into the exemplary MPEG supplemental enhancement information data SEI (F_ct), or into a similar metadata mechanism in other standardized or non-standardized communication methods.

[0058] After properly redefining the HDR images to be communicated as corresponding SDR images Im_SDR, they are often compressed (at least for example, for broadcast to end users) using existing image compression techniques (e.g., MPEG HEVC, VVC, or AV1, etc.). This is performed in a video compressor 203 forming part of the video encoder 221 (or even included in various forms of video production devices or systems).

[0059] The compressed image Im_COD is transmitted to at least one receiver by some image communication medium 205 (e.g., satellite, cable, or Internet transmission, e.g. according to ATSC3.0, or DVB, etc.; however, the HDR video signal may also be communicated, e.g., by a cable between two video processing devices).

[0060] Typically, prior to communication, further conversion is performed by a transmit formatter 204, which applies techniques such as packetization, modulation, transmission protocol control, etc., depending on the system, which typically applies integrated circuits.

[0061] At the receiving site, a corresponding video signal unformatter 206 applies the necessary unformatting methods, eg, demodulation, etc., to recapture, for example, a set of compressed HEVC images (ie, HEVC image data).

[0062] The video decompressor 207 performs, for example, HEVC decompression to obtain a stream of pixelated decompressed images Im_USDR, which in this example are SDR images, but in other modes are HDR images. The video decompressor also unpacks, for example from the SEI message, the required luminance mapping function F_L, or in general the color transformation function F_ct.

[0063] The image and function are input to a (decoder) color converter 208, which is configured to convert the SDR image to an image with a non-SDR dynamic range (i.e. a PL_V higher than 100 nits, typically at least several times higher, e.g. 5 times higher).

[0064] For example, a 5000nit reconstructed HDR image Im_RHDR is reconstructed as very close to the master HDR image (MAST_HDR) by applying the inverse color transform IF_ct of the color transform F_ct used at the encoding side to make Im_LDR from MAST_HDR. This image is then sent to the display 210 for further display adaptation, for example, but making the display adapted image Im_DA_MDR is also done in one go during decoding by using the FL_DA function (determined in an offline loop, for example in firmware) instead of the F_L function in the color converter. Therefore, the color converter further includes a display adaptation unit 209 to derive the FL_DA function.

[0065] The optimized, e.g., 800 nit, display adaptive image Im_DA_MDR is sent, e.g., to the display 210 if the video decoder 220 is included in, e.g., a set-top box or a computer, or is sent to a display panel if the decoder is in, e.g., a mobile phone, or is communicated to a movie theater projector if the decoder is in, e.g., an Internet-connected server, etc.

[0066] FIG. 3 shows a useful variant of the internal processing of a color converter 300 of an HDR decoder (or encoder, which generally has the same topology but uses an inverse function and generally does not include display adaptation), i.e., one corresponding to 208 in FIG.

[0067] The luminance of a pixel, in this example an SDR image pixel, is input as the corresponding luma Y'SDR. The chrominance, also known as chroma components Cb and Cr, are input to the downstream processing path of the color converter 300.

[0068] The luma Y'SDR is mapped to the required output luminance L'_HDR (e.g., master HDR reconstruction luminance, or some other HDR image luminance) by the luminance mapping circuit 310. It applies an appropriate function, e.g., display adaptation luminance mapping function FL_DA(t), for the particular image and maximum display luminance PL_D as obtained from the display adaptation function calculator 350 that uses as input a reference luminance mapping function F_L(t) co-communicated with the metadata. The display adaptation function calculator 350 also determines the appropriate function for processing the chrominance. For the moment, we simply assume that a set of multiplication coefficients mC[Y] for each possible input image pixel luma Y is stored, e.g., in the color LUT 301. The exact nature of the color processing may vary. For example, one may want to keep pixel saturation constant by first normalizing the chrominance by the input luma (corresponding hyperbola in the color LUT) and then correcting the output luma, although differential saturation processing may be used as well. Since both chrominances are multiplied by the same multiplier, the hue is generally maintained. Indexing the color LUT 301 with the luma value of the currently color transformed (luminance mapped) pixel Y results in the required multiplication coefficient mC as the LUT output, which is used by multiplier 302 to multiply it with the two chrominance values ​​of the current pixel, i.e., to yield the color transformed output chrominance. Cbo=mC*Cb Cro=mC*Cr

[0069] Via a fixed color matrixing processor 303 that applies standard colorimetric calculations, the chrominance is converted to lightness-deficient normalized non-linear R'G'B coordinates R' / L', G' / L', and B' / L'.

[0070] The R'G'B' coordinates that give the output image the appropriate brightness are obtained by multiplier 311, which: R'_HDR=(R' / L')*L'_HDR, G'_HDR=(G' / L')*L'_HDR, B'_HDR=(B' / L')*L'_HDR which are then collapsed into a color triplet R'G'B'_HDR.

[0071] Finally, further mapping to the format required by the display is performed by the display mapping circuit 320. This results in the display driving colors D_C, which are not only formulated into the colorimetry desired by the display (e.g. even HLG OEFT format), but furthermore, this display mapping circuit 320 is configured in some variants to perform some specific color processing for the display, i.e., for example further remapping some of the pixel luminances.

[0072] Some examples illustrating some suitable display adaptation algorithms for deriving a corresponding FL_DA function for a possible F_L function determined by the producing grader are taught in WO2016 / 091406 or ETSI's TS 103 433-2 V1.1.1 (2018-01).

[0073] However, these algorithms do not give much consideration to the minimum displayable black on the end user's display.

[0074] In fact, one could say that these algorithms pretend that the minimum luminance BL_D is small enough to be said to be zero. Hence, display adaptation primarily deals with the difference in the maximum luminance PL_D of various displays compared to the maximum luminance PL_V of the video.

[0075] As seen in drawing 18 of prior application WO2016 / 091406, any input function (in the illustrated example, a simple function formed from two linear segments) is typically scaled diagonally based on a metric positioned along an angle of 135 degrees from the horizontal axis of input luminance in a plot normalized to an input / output luminance of 1.0. It should be understood that this is only one example of a full range of display adaptations of a display adaptation algorithm, and it is not stated with the intention of limiting the applicability of the inventors' novel display adaptation concepts, e.g., in particular, the angle of the metric direction may have other values.

[0076] However, this metric and its effect on the reshaped F_L function, i.e., the determined FL_DA function, depends only on the maximum luminances PL_V and PL_D of the display to be delivered with the optimally regraded mid-dynamic range image. For example, the 5000 nit position corresponds to a zero metric point located on the diagonal (for any location located along the diagonal corresponding to a possible pixel luminance in the input image), and the 100 nit position (marked PBE) is a point in the original F_L function.

[0077] Display adaptation as a useful variant of this method is summarized in Figure 4 by showing its effect on a plot of possible normalized input luminance Ln_in versus normalized output luminance Ln_out (which is converted to actual luminance, i.e., PL_V value, by multiplying it by the maximum luminance of the display related to the normalized luminance).

[0078] For example, a video creator designs a luminance mapping strategy between two reference gradings as described in Fig. 1. Therefore, for a possible normalized luminance Ln_in of a pixel in an input image, e.g. a master HDR image, this normalized input luminance must be mapped to a normalized output luminance Ln_out of a second reference grading, which is the output image. This re-grading of all luminances corresponds to a function F_L, which has many different shapes determined by a human grader or a grading automaton, and the shape of this function is jointly communicated as dynamic metadata.

[0079] The question is now what shape should the derived quadratic version of the F_L function have in this simple display adaptation protocol to map (instead of the reference SDR image) to an MDR image for a mid-dynamic range display (assuming that the mapping starts again with the HDR reference graded image as the input image). For example, it can be calculated based on the metric as follows, where for example an 800 nit display should have 50% of the grading effect, and a full 100% is a regrading of the master HDR image to a 100 nit PL_V SDR image. In general, via the metric, for a possible normalized input luminance of a pixel (Ln_in_pix) represented as display adaptation luminance L_P_n, determine any point between no regrading to the second reference image and full regrading, the location of which of course depends on the input normalized luminance, but also on the value of the maximum luminance (PL_V_out) associated with the output image. Those skilled in the art will understand that although the function can be expressed in a normalized luminance representation, it can equally be expressed in any normalized luma representation defined according to any OETF.

[0080] The corresponding display-adaptive luminance mapping FL_DA is determined as follows (see Fig. 4a): Take any one of all input luminances, for example Ln_in_pix. This corresponds to a starting position (shown as a square) on a diagonal with equal angles to the input / output axis of normalized luminance. For each point on the diagonal, place a scaled version of the metric (scaled metric SM) that is perpendicular to the diagonal (or 135 degrees counterclockwise from the input axis), starting at the diagonal and ending at a point on the F_L curve (at the 100% level), i.e. the intersection of the F_L curve with the vertically scaled metric SM (shown as a pentagon). Place a point at the 50% level, i.e. in the middle, of the metric (in this example, for this PL_D value of the display for which the image has to be calculated) [note that in this case the PL_V value of the output image is set equal to the PL_D value of the display for which the display-optimized image has to be fed]. By doing this for all points on the diagonal corresponding to all Ln_in values, an FL_DA curve is obtained that is shaped similarly to the original, i.e., with the same regrading, but with maximum luminance rescaling / adjustment. This function is now ready to be applied to calculate the corresponding optimally regraded and / or display adapted 800 nit PL_V pixel luminance required, given any input HDR luminance value of Ln_in. This function FL_DA is applied by the luminance mapping circuit 310.

[0081] In general, the characteristics of this display adaptation are as follows (not intended to be particularly limiting). The direction of the metric may be fixed in advance as technically desired. Figure 4b shows another scaling metric, namely the vertical scaling metric SMV (i.e., perpendicular to the axis of normalized input luminance Ln_in). Again, 0% and 100% (or 1.0) correspond to no regrading (i.e., identity transformation on the input image luminance) and regrading to the second of the two reference grading images (related in this example by a luminance mapping function F_L2 of a different shape), respectively.

[0082] The location of the measurement points on the metric, ie, where the 10%, 20%, etc. values ​​are located, is also subject to engineering variation, but is generally non-linear.

[0083] It is technically pre-designed, for example in a television display. For example, a function as described in WO2015007505 is used. A logarithmic function can also be designed so that a*(log(PL_V)+b) is equal to 1.0 (for example, 5000nit) of the PL_V_HDR value, and the 0.0 point corresponds to the PL_V_SDR reference level of 100nit, or vice versa. The position of the PL_V_MDR where the image brightness needs to be calculated is then obtained from the designed mathematics of the metric.

[0084] The behavior of such a metric is summarized in FIG.

[0085] The display adaptation circuitry 510 includes a configuration processor 511, for example in a television, or a set-top box, etc., which sets values ​​for the processing of an image before the actual pixel colors of that image are processed. For example, the maximum luminance value of the display-optimized output image PL_V_out may be set once in the set-top box by polling it from the connected display (i.e., the display communicates the maximum displayable luminance PL_D to the set-top box), or if the circuitry is present in the television, it may be set by the manufacturer, etc.

[0086] The luminance mapping function F_L, in some embodiments, varies for each input image (in other variations it is fixed for many images) and is input from some source of metadata information 512 (e.g. it is broadcast as an SEI message, read from a sector of memory such as a Blu-ray disc, etc.). This data establishes the normalized height of the normalized metric (Sm1, Sm2, etc.) on which the desired location of the PL_D value is found from the mathematical formula of the metric.

[0087] When an input image 513 is input, successive pixel intensities (eg, Ln_in_pix_33 and Ln_in_pix_34, or luma) are passed through a color processing pipeline that applies display adaptation, resulting in corresponding output intensities such as Ln_out_pix_33.

[0088] Note that none of this techniques specifically provides a minimum black luminance.

[0089] The reason is that the usual approach is that the black level depends heavily on the actual viewing situation, which is even more variable than the display characteristics (i.e. the very first PL_D). All sorts of influences occur, ranging from physical lighting aspects to the optimal configuration of the light-sensitive molecules in the human eye.

[0090] So you make a good image "for the display", that's it (i.e. in terms of how much more capable (luminosity-wise) the intended HDR display is than a typical SDR display). Then you can later post-correct somewhat for the viewing situation if necessary, which is an (undefined) special task left to the display.

[0091] Therefore, in general we assume that the display has a variable high brightness capability, i.e., is capable of displaying all the required pixel brightnesses coded in the image up to PL_D (for the moment we assume that it is already an MDR image optimized for the PL_D value, i.e., that there are generally at least some pixel regions in the image that reach up to PL_D), since in general we do not want to suffer the severe consequences of white clipping, but as mentioned above, the black of the image is often not of interest.

[0092] Black is "mostly" visible anyway, so if some of it is a little less visible, it's not the most important thing. At least the potentially very bright pixel luminances of the master HDR grading can be optimally constrained to a limited upper range of the display, say above 200nit, e.g. from 200nit to PL_D=600nit (for master HDR luminance up to e.g. 5000nit).

[0093] This is similar to assuming that black is always zero nits (at least approximately) for all images and all displays. White clipping is a much more visually bothersome property than losing some of the black, which often still allows you to see something but is less pleasant.

[0094] However, sometimes that approach is not sufficient because a significant subrange of the darkest luminances becomes invisible, or at least not fully visible, under significant ambient lighting in a viewing room (e.g., a consumer television viewer's living room with large windows during the day), which may be different and dimmer or even darker than the ambient lighting conditions in a video editing room where the video is created.

[0095] It is therefore necessary, for example, to increase the brightness of those pixels, generally by means of a control button of the display (a so-called brightness button).

[0096] When using a television electronic behavior model such as in Rec. ITU-R BT.814-4 (07 / 2018), a television in an HDR scenario takes the luma+chroma pixel colors (which actually drive the display) and converts them to non-linear R', G', B' non-linear drive values ​​to drive the panel (according to standard colorimetric calculations). The display then processes these R', G', B' non-linear drive values ​​with the PQ EOTF to know what front screen pixel luminance to display (i.e., generally, how to drive e.g. an OLED panel pixel or an LCD pixel, if there is still internal processing that accounts for the electro-optical physical behavior of the LCD material, but that aspect is irrelevant to this discussion).

[0097] A control knob, for example on the front of the display, then gives the luma offset value b (the moment at which a black patch 2% above the minimum in the PLUGE or other test pattern becomes visible, while -2% black becomes invisible).

[0098] The original uncorrected display behavior is LR_D=EOTF[max(0,R')]=PQ[max(0,R')] LG_D=EOTF[max(0,G')]=PQ[max(0,G')] LB_D=EOTF[max(0,B')]=PQ[max(0,B')] [Formula 1] If so.

[0099] In this formula, LR_D is the amount of red contribution (linear) that should be displayed to create a particular pixel color with a particular luminance (in (fractional) nits), and R' is the non-linear luma code value, e.g., 419 out of 1023 values ​​in a 10-bit encoding.

[0100] The same happens for the blue and green components. For example, if you need to make a particular color 1 nit (the total brightness of that color to the eye), you need, say, 0.33 units of blue, and the same for red and green. If you need to make 100 nits of that same color, you can say that LR_D=100*0.33 nit.

[0101] Now, if this display driving model is controlled via the luma offset knob, the general equation becomes: LR_D_c=EOTF[max(0,a*R'+b)], where a=1-b / OETF[PL_D], etc. [Equation 2]

[0102] Instead of displaying the zero black of the image hidden somewhere in the invisible display black, this technique raises the zero black to just the level where black becomes sufficiently discernible. (Note that in consumer displays, mechanisms other than PLUGE may be used, e.g., despite the viewer's preferred value of the available luma offset b, viewer preferences may potentially lead to another possible suboptimal.)

[0103] This is a display post-processing step after creating an optimally regraded image: first an optimal theoretically regraded image is calculated by the decoder, e.g., first mapped to the reconstructed master HDR image, then luminance remapped to e.g. a 550 nit PL_V MDR image, i.e. taking into account the display's brightness capability PL_D.

[0104] This optimal image is then determined according to the filmmaker's ideal vision and is then further mapped by the display taking into account the expected visibility of black in the image.

[0105] US Patent Application Publication No. 2017025603 teaches how to arrive at an optimal luminance mapping that depends on the amount of ambient light by analyzing the luminance histogram of the input image. The minimum luminance, maximum luminance, and average luminance are used to make a corresponding tripartite mapping, after which an appropriate tone mapping optimizes the luminance, i.e., a function that boosts the brightness more strongly for the darkest image luminances is selected for mostly dark scenes, and a less steep function is selected for bright scenes. The user fine-tunes by using a function that is roughly halfway between these two functions.

[0106] US Patent Application Publication No. 2019304379 teaches how to pre-calculate a virtual image, essentially a 5000 nit master HDR image for example, but with primarily darker image objects brightened to compensate for viewing environments brighter than the 5 nit to which the master HDR image was graded, by using a predefined ambient brightness correction luminance mapping function. A standard display adaptation algorithm, typically determined by minimum, average, and maximum luminance, then uses a sigmoid mapping of the pre-corrected virtual image rather than the original master HDR image. There is also a specific teaching of how a bright (virtual) image can be determined by using human visual contrast characteristics in PQ space.

[0107] U.S. Patent Application Publication No. 20170186141 generally teaches that the mapping from the luminance of an HDR image to the luminance of an SDR image uses a tone mapping function (typically piecewise linear), which is determined by the user of a television or mobile phone. Summary of the Invention [Problem to be solved by the invention]

[0108] The problem, according to the inventors, is that this is a rather crude way of adapting the viewing room ambient light level of the image to be displayed. Alternative ways have been developed, in particular ways that allow the viewer to see a more contrasty image in at least some parts of the image that are important. In general, the prior art also does not address the high re-grading needs of video or image creators. [Means for solving the problem]

[0109] A visually better looking image for various ambient lighting levels is obtained by processing the input image to obtain an output image; the input image has pixels having input luminance within a first luminance dynamic range (DR_1) having a first maximum luminance (PL_V_HDR); A reference luminance mapping function (F_L) is received as metadata associated with the input image; a reference luminance mapping function specifying a relationship between luminances of located pixels in the two images, the two images being graded differently in that pixel luminances of the same image object have different pixel luminances in the two images; a reference luminance mapping function specifying a relationship between the luminance of the input image and the luminance of a secondary reference image having a second reference maximum luminance (PL_V_SDR); the output image has an output maximum luminance (PL_V_MDR) different from the first maximum luminance and the second reference maximum luminance (PL_V_SDR); The process includes determining an adaptive luminance mapping function (FL_DA) based on a reference luminance mapping function (F_L) and an adjusted maximum luminance value (PL_V_CO), where the adjusted maximum luminance value (PL_V_CO) is different from an output maximum luminance (PL_V_MDR); and applying the adaptive luminance mapping function (FL_DA) to input pixel luminances to obtain an output luminance of an output image; Calculating an adaptive luminance mapping function (FL_DA) comprises finding a position (pos) on a metric (SM) that specifies the location of maximum luminance, this position corresponding to an adjusted maximum luminance value (PL_V_CO); a first endpoint of the metric corresponds to a first maximum luminance (PL_V_HDR) and a second endpoint of the metric corresponds to a second reference maximum luminance (PL_V_SDR); a first endpoint of the metric is located, for any normalized input luma (Yn_CC0), at a diagonal point having horizontal and vertical coordinates equal to that normalized input luma; A second end point is located on the locus of a reference luminance mapping function (F_L) determined by the direction of the metric; The adjusted maximum brightness value (PL_V_CO) is Get the user correction value (UCBVal), Determine the adjusted maximum brightness value (PL_V_CO) as the result of subtracting the user correction value (UCBVal) from the output maximum brightness (PL_V_MDR). and Write the output luminance as a pixel color to the output image, and output the output image. It is characterized by:

[0110] The reference luminance mapping function typically indicates how the video should be regraded, for example, when making an image of lower maximum (possible) luminance that generally corresponds visually to the master HDR image (in the video, the maximum luminance is generally an upper limit that the creator can use to maximally brighten some of the pixels corresponding to very bright objects in some of the images according to the technical HDR encoding variant selected). That is, the need for regrading each image of any HDR scene (e.g., the first night image of a movie versus the later daytime image) is specified by a function (e.g., by the content creator's color grader, or an automaton) that maps the luminance between two specifically graded reference images. Generally, the first of those reference images is the input image itself (e.g., the master HDR image with maximum luminance, e.g., 5000 nit, or an SDR image). The second reference image is also called the secondary reference image. The function is generally defined in the luma domain (according to the selected OETF function flavor), and furthermore, the color processing occurs in the luma domain (although the metrics, even when superimposed on such a luma domain plot of the luminance regrading function, generally indicate the location of maximum luminance rather than luma). The luma is generally normalized. The second reference image is typically (without limitation) an SDR, i.e., 100 nit maximum luminance image, generally regraded from a master HDR image (or vice versa, where the SDR image becomes the starting image for regrading the corresponding HDR image, but that aspect does not change the contrast improvement process).

[0111] When proposing his method, the inventors wanted to have a way to take into account this guide of the need for luma or luminance regrading, specified by the shape of the function F_L. The inventors wanted a user control that would work in an easy and practical way under such technical conditions, in order to efficiently improve the contrast of the image, especially when viewed in an environment with a high level of illumination, which is generally much higher than the light level at which the image is ideally viewed. The inventors wanted an easy way to intervene in the regrading curve (i.e., F_L) that would work well for any possible curve that video creators communicate (it could be a simple concave function that relatively boosts the darker pixel lumas for regrading to reduce the output image maximum luminance, while gradually decreasing towards the brightest lumas, but also a fairly complex curve, for example, where the contrast of an important subrange of lumas around the middle of the total range remains stretched by the curve F_L, resulting in, for example, a double-step appearance).

[0112] Therefore, a display adaptation method as invented by the inventor is used, which is now adjusted differently (shape adapted), i.e. the squeezing towards the diagonal occurs with a different control variable, i.e. the adjusted maximum luminance value (PL_V_CO).

[0113] Various metrics are defined, but in general, the first end point (corresponding to the location of maximum luminance of the master HDR grading / image) is always located on a diagonal, i.e., a line in the plot of normalized luma between zero and one, with an angle of 45 degrees with respect to both the input luma axis and the output luma axis. Specifically, for each possible input luma Y_in, this point has a position on the diagonal with a horizontal coordinate (and also a value of the vertical coordinate) equal to Y_in for any Y_in. The location of the metric end point, i.e., the position that typically defines the value of PL_V_SDR, is somewhere on the locus of the shape of the F_L function, depending on which angle has been chosen to perform the display adaptation algorithm (e.g., for a vertical angle, the metric end point is h=Y_in;v=F_L(Y_in)).

[0114] Such direct user control devices do not require measuring an ambient light level indication quantified as luminance, i.e., ambient illumination value (Lx_sur), although such illuminance measurements may optionally be present to further guide the user, for example by specifically suggesting a good working or initial value, i.e., the amount of light for daytime or nighttime television viewing, for example in a living room.

[0115] The reference lighting value (GenVwLx) is a typical value, e.g. an expected value or a value that works well. Ideally, it is determined corresponding to the received image, e.g. communicated as metadata of the image indicating that the image was made for such an environmental lighting or is best viewed under such an environmental lighting. It is, e.g., a value for viewing the master HDR image (e.g., VIEW_MET), or a value derived therefrom, e.g., by applying a formula to obtain a final typical value, taking into account also the value of the SDR grading. This formula is calculated at the receiving side, or the result is calculated by the video or image creator and communicated as metadata. For example, the GenVwLx value is overwritten in memory, which contains a fixed manufacturer value or a commonly communicated value (e.g., from a previous program), etc., if no metadata of the current image (set) is received.

[0116] Luma is calculated by applying an optoelectronic conversion function (OETF_psy) to the input luminance (L_in) of the input image, where the input luminance (L_in) is present as an input. In some methods or devices, the input image pixel color has a luma component.

[0117] The photoelectric transfer function (OETF_psy) used is preferably a psychovisually uniform photoelectric transfer function. This function is defined by determining (typically experimentally in a laboratory) a luminance-to-luma mapping function shape, so that a second luma a fixed integer number N lumas higher than a first luma selected somewhere in the luma range corresponds approximately to a similar difference in perceived lightness compared to the perceived lightness of the first luma for a human observer. A further regrading mapping is then defined in this visually uniformed luma domain.

[0118] That is, human vision is non-linear and therefore does not perceive the difference between 10 nit and (1.05)*10 nit in the same way as, for example, the difference between 2000 nit and (1.05)*2000 nit. The equalized perception curve also depends on what the human is looking at, i.e., in particular the dynamic range of the display being viewed (in a particular surrounding), i.e., the maximum luminance PL_D and the minimum luminance.

[0119] So, ideally, for processing purposes we define an OETF (or its inverse, EOTF) that has the following properties:

[0120] If we take the first luma, for example in 10 bits, luma_1=10, and then we move, for example, the N=5 luma code higher, and get the second luma luma_2=15. This corresponds to a change in the lightness sensation (i.e. the value that characterizes what a human viewer experiences as the lightness of a displayed patch at a particular physically displayed luminance; it is determined by applying the EOTF).

[0121] So, assume that in a lightness range of 1 to 200, luma 10 gives an impression of a lightness of 5 and luma 15 gives an impression of 7, i.e. 2 units lighter.

[0122] Then take two lumas that encode brighter luminance, say 800 and 800+5. Then, if the viewer experiences the same lightness difference due to the luma difference, the luma scale is approximately visually uniform. For example, luma 800 gives a lightness perception of 160, while luma 805 looks like lightness 162, i.e., again, there is a 2 unit difference in lightness. The function does not need to be exactly the lightness determination function, since a reasonably perceptually uniform luma definition already works well.

[0123] The effects of changes due to luminance processing would be less visually unpleasant (because they are more easily suppressed by the brain) if they were performed in such a psychovisually uniform system.

[0124] Display adaptation generally reduces the tilt of darker object luma compared to reference regrading to an SDR image, but now the aim is to keep the tilt at the value required to provide a visually good looking image.

[0125] That is, any of the possible embodiments to make lesser variations of the specially formulated F_L function (corresponding to the specific required luminance remapping, a.k.a. regrading need for the current specific HDR image or scene to obtain a corresponding lower maximum luminance secondary grading) as described in Figures 4 and 5 or similar techniques are used. The luminance mapping function is generally represented as a luma mapping function, i.e. by transforming both normalized luminance axes to the luma axis via an appropriate OETF (normalization using the appropriate maximum luminance value of the input and output images).

[0126] Specifically, the secondary grading for which the F_L function is defined is advantageously a 100nit grading, which is suitable to meet the requirements of most future display situations (note that there will also be display scenarios with low dynamic range quality in the future).

[0127] In a practical embodiment, there would be a scaler and / or clipper to prevent the adjusted maximum brightness value PL_V_CO from becoming too low, e.g. set equal to PL_V_SDR if it becomes lower, or in general there would be a damping function to reduce the slope for higher input values.

[0128] For example, this can be done by the manufacturer presetting an intensity value (EXCS) in the user value circuit (903) which determines how sensitive the device is to user control (e.g., repeated pressing of a button or turning of a knob, etc.); linear user control is generally satisfactory (although more sophisticated non-linear controllers may be used as well).

[0129] Advantageously, the method for processing the input image uses as a reference luminance mapping function (F_L) created by the creator of the input image and received as metadata via the image communication channel.

[0130] The system also determines a unique version of a well-performing regrading function to be used by the device (i.e., determines a good shape for F_L heuristically, for example, based on an analysis of the luma occurrences of several previous images and / or the current image (in the case of delayed output)), but is particularly useful when working with an author-determined regrading function.

[0131] The metrics in the algorithm are expressed in the computations in the electronic circuit as being mathematically presented in a plot of input versus output luma. The locus of the reference luminance mapping function (F_L) in such a plot is the position having as its vertical coordinate the normalized output luma obtained when using the normalized input luma as input to the reference luminance mapping function (F_L).

[0132] The various display adaptation algorithms are selected for the method, in particular by predefining the direction that the algorithm must use. The first point of the metric (which corresponds to the input image and is the identity transformation of the luma when the luma maps the input image to itself) is always on the diagonal. The second end point is the point on the locus of F_L directly above the first end point for any of the possible normalized input lumas (0 to 1.0). At 45 degrees from the axis of the input luma (horizontal axis) or 135 degrees counterclockwise, a line in this direction intersects with the F_L locus at a horizontally offset position, and the second end point is located there. That is, exactly where the second end point of the metric is located on the locus of the function is determined by the direction of the line segments starting from each position on the diagonal. These calculations are accelerated, for example, by calculating the adaptive luminance mapping function (FL_DA) as a 1D LUT with N output entries for N normalized input lumas before receiving the set of images to be processed for which the function F_L is applicable.

[0133] The user control value UCBVal in these embodiments is generally expressed as a quantity having the same dimension as the maximum luminance value, i.e., in nits (SI units Cd / m 2 The text is presented in a format that is easier to typographically translate into English than the original.

[0134] Advantageously, the method for processing the input image sets the output maximum luminance (PL_V_MDR) equal to the maximum displayable pixel luminance of the display to which the output image may be supplied, so that all processing works with the maximum luminance value (PL_V_CO) adjusted based on this value of the connected display, typically the display on which the viewer is watching or intends to watch the video, e.g. a broadcasted television program (however, there may be systems in which another value is used, e.g. the display it communicates with wants a slightly higher PL_V_MDR and still does an internal luminance optimization process, but then the method still works in the same way).

[0135] A simple but sufficient embodiment of a method for processing an input image is: obtaining an input value (UCBSliVal) from a user of the display; calculating a user control value (UCBVal) by multiplying the input value by the intensity value (EXCS) and the output maximum brightness (PL_V_MDR); has.

[0136] Such a method scales correctly and is intuitive to use (note that the effects are still in a psychovisually uniform domain, so the user has visually precise control that is consistent with the need to regrade, i.e., minimally disruptive to the creator's artistic vision).

[0137] Advantageously, the method of processing the input image not only works by modifying the contrast of the darkest pixels in particular (not knowing about any particular black offset setting process), but also performs a process to set the darkest black in the input image to a black offset value as a function of the ambient illumination value (Lx_sur). For example, it starts from an image with the smallest black luma. In principle, any such methods may be combined, but in the following some particularly advantageous methods are taught that cooperate particularly well with contrast optimization.

[0138] Advantageously, the method of processing the input image has the metric direction preset as vertical, which means that the second endpoint for the normalized input luma (Yn_CC0) is at the location having the result of applying a reference luminance mapping function with the normalized input luma as input as the horizontal coordinate and that normalized input luma (F_L(Yn_CC0)) as the vertical coordinate.

[0139] The method is further embodied as an apparatus (900) for processing an input image to obtain an output image, The input image has pixels having an input luminance within a first luminance dynamic range (DR_1) having a first maximum luminance (PL_V_HDR), and the apparatus comprises: an image input unit (921) for acquiring an input image (513); a data input (920) for receiving a reference luminance mapping function (F_L) which is metadata related to the input image; a reference luminance mapping function specifying a relationship between the luminance of the first reference image and the luminance of the second reference image; a reference luminance mapping function specifying a relationship between the luminance of the input image and the luminance of the second reference image; the second reference image has a second reference maximum luminance; a data input (920) in which an output image has an output maximum luminance (PL_V_MDR) different from a first maximum luminance and a second reference maximum luminance; Including, This device, a user value circuit (903) configured to determine and output a user correction value (UCBVal) as set by a human user of the device; Obtain a user correction value (UCBVal) from the user value circuit (903); Outputs the adjusted maximum brightness value (PL_V_CO) as the result of subtracting the user correction value (UCBVal) from the output maximum brightness (PL_V_MDR). A maximum brightness determination unit (901) configured as follows: Further comprising: This device, a display adaptation circuit (510) configured to determine an adaptive luminance mapping function (FL_DA) based on the adjusted maximum luminance value (PL_V_CO) and the reference luminance mapping function (F_L); Calculating an adaptive luminance mapping function (FL_DA) includes finding a position (pos) on the metric (SM) that corresponds to an adjusted maximum luminance value (PL_V_CO); a first endpoint of the metric corresponds to a first maximum luminance (PL_V_HDR) and a second endpoint of the metric corresponds to a maximum luminance of a second reference image; a first endpoint of the metric is located, for any normalized input luma (Yn_CC0), at a diagonal point having horizontal and vertical coordinates equal to the normalized input luma; A second end point is located on the locus of the reference luminance mapping function (F_L); a display adaptation circuit (510) configured to apply an adaptive luminance mapping function (FL_DA) to an input luminance or an input luma encoding the input luminance to obtain an output luminance or output luma, and to output the output luma or output luminance as pixel colors of an output image; Further includes:

[0140] A further useful embodiment of this apparatus for processing an input image comprises a user value circuit (903), which Obtain input value (UCBSliVal) from the user of the display, Calculate the user control value (UCBVal) by multiplying the input value by the intensity value (EXCS) and the output maximum brightness (PL_V_MDR); Output the user control value (UCBVal) to the maximum brightness determination unit (901) It is configured as follows.

[0141] A further useful embodiment of this apparatus for processing an input image includes a black adaptation circuit configured to set the darkest black luma in the input image to the black offset luma as a function of the ambient illumination value (Lx_sur).

[0142] In particular, those skilled in the art will appreciate that these technological elements are embodied in various processing elements such as ASICs (Application Specific Integrated Circuits, i.e., typically an IC designer will have (part of) an IC perform the method), FPGAs, processors, etc., and are present in a variety of consumer or non-consumer devices, whether including displays or non-display devices externally connected to a display, that images and metadata are communicated to and from various image communication technologies such as over-the-air broadcast, cable-based communication, etc., and that the devices are used in a variety of image communication and / or usage ecosystems such as, for example, television broadcast, on-demand via the Internet, etc.

[0143] These and other aspects of the method and apparatus according to the invention will become apparent and will be explained with reference to the implementations and embodiments described below and with reference to the accompanying drawings, which merely serve as non-limiting specific illustrations illustrating the more general concepts, in which dashes are used to indicate that a component is optional and that a non-dashed component is not necessarily essential. Dashes are also used to indicate elements that are described as essential but are hidden inside an object, or intangible things such as, for example, selection of an object / area. [Brief description of the drawings]

[0144] [Figure 1]1 is a schematic diagram of some typical color transformations. Color transformation occurs when optimally mapping a high dynamic range image to a corresponding optimally color-graded and similar looking (as similar as desired and feasible given the differences in the first and second dynamic ranges DR_1 and DR_2, respectively) lower dynamic range image, e.g., a standard dynamic range image with 100 nit maximum luminance, which in the lossless case also corresponds to mapping a received SDR image that actually encodes an HDR scene to a reconstructed HDR image of that scene. Luminance is shown as a location on the vertical axis from darkest black to maximum luminance PL_V. The luminance mapping function is symbolically shown by an arrow that maps the average object luminance from the luminance on the first dynamic range to the second dynamic range (those skilled in the art know how to equivalently draw this as a classical function on axes normalized by dividing by the respective maximum luminance, e.g., normalized to 1). [Diagram 2] FIG. 1 shows a schematic example of a high-level view of Applicant's recently developed technique for encoding high dynamic range images, i.e., images that can have a brightness of at least 600 nits or more (typically 1000 nits or more), which in effect communicates an HDR image by itself or as a corresponding brightness regraded SDR image plus metadata that encodes a color transformation function including at least an appropriate determined brightness mapping function (F_L) for pixel colors used by a decoder to convert the received SDR image to an HDR image. [Diagram 3] FIG. 2 shows the internal details of an image decoder, in particular a pixel color processing engine, in a (non-limiting) preferred embodiment. [Figure 4]Consisting of the sub-images Fig. 4a and Fig. 4b, it illustrates two possible variants of display adaptation to obtain a final display-adapted luminance mapping function FL_DA that is used to calculate the optimal display-adapted version of an input image for a specific display capability (PL_D) from a reference luminance mapping function F_L that codifies the luminance re-grading needs between two reference images. [Diagram 5] FIG. 1 is a diagram summarizing the principles of display adaptation more generally, for easier understanding of the principles of display adaptation as an element in the formulation of the present embodiments and claims. [Figure 6] FIG. 2 illustrates an exemplary device for explaining technical elements related to ambient lighting compensation processing. [Figure 7] FIG. 1 illustrates the correlation concept of a virtual target display, in particular the minimum luminance (mL_VD) of the target display and its relationship or relevance to a real display, such as the viewer's end-user display, on which luminance remapping is typically performed. [Figure 8] FIG. 1 illustrates some of the metadata categories that are available for long-standing professional HDR encoding frameworks (e.g., typically in conjunction with pixel color images) so that any receiver has all the data available that it needs, or at least elegantly utilizes, to optimize the received image for a particular end viewing situation (display and environment). [Figure 9] FIG. 1 illustrates a method for optimizing image contrast, particularly when designing based on existing display adaptation techniques, and in particular for compensating for various viewing environment light levels or conditions. [Figure 10] FIG. 7 shows an example for several image luminances using a particular reference mapping function F_L (chosen to be a simple linear curve for ease of understanding) of what happens when applying the viewing environment lighting adaptation technique (as described in FIG. 6). [Figure 11]A diagram showing an example of what happens when applying the contrast improvement technique described in FIG. 9 (without viewing environment lighting, although the two processes can also be combined, which introduces an additional offset to the blackest black). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0145] FIG. 6 shows an apparatus for generally describing how to implement the elements of the present invention, in particular the color processing circuit 600. It will be generally described, and then some details of the variations of the embodiment will be described. We assume that the ambient light adaptive color processing is performed inside an application specific integrated circuit (ASIC), but those skilled in the art will understand how to similarly implement color processing in other related devices. Some aspects differ depending on whether this ASIC is present, for example, in a television display (typically an end user display), or in another device connected to the display, such as a set-top box, to perform processing on the display and provide it with an image that is already brightness optimized. (Those skilled in the art can also map this apparatus block diagram to a method flow diagram of the corresponding method).

[0146] The input luminance (L_in) of the input image to be processed (without loss of generality, let us assume it is the master HDR image MAST_HDR) is first converted to the corresponding luma (e.g. 10-bit coded luma) after being input via the image pixel data input 690. For that purpose, an optical-electronic transfer function is used, which is generally fixed in the device by the manufacturer (although it is configurable).

[0147] It is useful to have a perceptually uniform OETF (OETF_psy).

[0148] Assume the following OETF is used (defined as the output of luma Yn_CC0): Yn_CC0=v(L_in;PL_V_in)=log[1+(RHO-1)*power(L_in;p)] / log[RHO] [Formula 3] Here, RHO is a constant that depends on the input maximum luminance PL_V_in according to the formula RHO(PL_V_in)=1+32*power((PL_V_in / 10,000);p), where p is preferably a power of a power function equal to 1 / (2.4).

[0149] The input max luminance is the maximum luminance associated with the input image (it is not necessarily the luminance of the brightest pixel in each image of a video sequence, but is metadata that characterizes the image or video as an absolute upper limit value).

[0150] It is generally configurable and input via a first maximum data input 691 as PL_V_HDR, for example 2000 nits, or in other variants, a fixed value for the image communication and / or processing system, for example 5000 nits, and therefore a fixed value in the processing of the photoelectric conversion circuit 601 (therefore, the vertical arrow representing the data input of PL_V_HDR is shown as a dotted line, since it is not present in all embodiments; note that the open circle symbolizes the branching of this data supply, not to be confused with a non-mixed overlapping data bus).

[0151] Then, in some embodiments, there is further luminance mapping by a luminance mapping circuit 602 in order for the method or color processing device to obtain a starting luma Yn_CC that is optimally adapted to the particular viewing environment. Such further luminance mapping is typically display adaptation to pre-adapt the image to the particular display maximum luminance PL_D of the connected display.

[0152] In a simpler embodiment, this optional luminance mapping does not exist, so the 2000nit environment optimized image will be calculated for the input 2000nit master HDR image (or a fixed 5000nit HDR image situation), i.e., we will first describe the color processing assuming that the maximum luminance value remains the same between the input and output images. In this case, the starting luma Yn_CC is simply equal to the initial starting luma Yn_CC0 output by the photoelectric conversion circuit 601.

[0153] The linear scaling circuit 603 is Yim=(Ydif+1)*Yn_CC-1.0*Ydif [Formula 4] Calculate the intermediate luma Yim by applying a function of the type:

[0154] The luma difference Ydif is obtained from a (second) photoelectric conversion circuit 611. This circuit converts the luminance difference dif into a corresponding luma difference Ydif. This circuit uses the same OETF formula as circuit 601, i.e. it also uses the same PL_V_in values ​​(and powers).

[0155] The luminance difference dif is calculated by a luminance difference calculator 610, which receives the two darkest luminance values, namely the minimum luminance of the target display given the particular lighting characteristics of the viewing room (mL_VD) and the minimum luminance of the end user display (mL_De); dif = mL_VD - mL_De [Formula 5] Calculate mL_De is generally a function (usually additive) of the fixed display minimum black (mB_fD) of the end user display on the one hand, and the luminance (mB_sur), which is a function of the amount of ambient light. A typical example of display minimum black (mB_fD) is the leakage light of an LCD display. If such a display is driven with a code showing complete black (i.e. ideally zero photons are output), due to the physical properties of the LCD material, such a display will still always output, for example, 0.05 nits of light, the so-called leakage light. This is true regardless of the ambient light in the viewing environment, and therefore applies to a completely dark viewing room.

[0156] The values ​​of mL_VD and mL_De are typically obtained via a first minimum metadata input 693 and a second minimum metadata input 694, for example from an external memory in the same or a different device, or via a circuit from a light measurement device, etc. The maximum luminance value required in a particular embodiment, for example the maximum luminance (PL_D) of a display capable of displaying the color processed output image, is typically input via an input connector such as the second maximum metadata input 692. The output image color is fully written in the pixel color output 699 (a person skilled in the art will understand how to realize this as various technical variants, for example pins of an IC, a standard video cable connection like HDMI, wireless channel based communication of the pixel color data of the image, etc.).

[0157] The exact determination of the luminance (short called ambient black luminance) as a function of the amount of ambient light (mB_sur) is not a typical aspect of the present invention, since it may be determined in several alternative ways. For example, the viewer determines the value of the luminance mB_sur that he considers to represent the masking black resulting from reflections on the front screen of the display, using a test signal such as PLUGE, or a variant more convenient for the consumer. We further assume that the viewer simply sets the value, whether for a particular evening or from a situation where, for example, the consumer usually keeps the room lighting configuration fixed when purchasing the display. Or even a value that the television manufacturer has baked in as a value that works well on average for at least one of the typical consumer viewing situations. When this method is used to adapt the luminance for viewing on a mobile device, it is common that the viewing environment is relatively unstable (for example, watching a video in a train and the lighting level changes from outdoors to indoors when the train enters a tunnel).

[0158] In such cases, for example, a time-filtered measurement of the built-in illuminance meter can be utilized (measured less frequently in order not to unnecessarily adapt the treatment over time, e.g. when sitting on a bench in the sun and then walking indoors).

[0159] Such instruments generally measure the average amount of light (in lux) that falls on it, and therefore on the display.

[0160] Although a different photometric quantity, the lux value is converted to ambient black luminance by the well-known photometric formula: mB_sur=R*Ev / pi [Equation 6]

[0161] In this formula, Ev is the ambient illuminance in lux, Pi is the constant 3.1415, and R is the reflectance of the display screen. Typically, mobile display manufacturers put their values ​​into the formula.

[0162] For normal reflective surfaces in the surroundings, e.g. the walls of a house, with a color somewhere between average gray and white, an R-value of about 0.3 can be assumed, i.e. as a rule of thumb it can be said that the luminance value is about 1 / 10 of the illuminance value stated in lux. The front of a display reflects much less light. Depending on whether special anti-reflection techniques are utilized, the R-value is for example about 1% (but may be higher, towards the 8% reflectance of glass, which may be problematic especially in brighter viewing surroundings). Therefore, mL_De=mB_fD+mB_sur [Equation 7]

[0163] The technical meaning of the value of the minimum luminance of the target display (mL_VD) is further explained with the aid of FIG.

[0164] The first three luminance ranges starting from the left are actually "virtual" display ranges, i.e. ranges that correspond to the image, not necessarily to an actual display (that is, to a target display, i.e. a display on which the image could ideally be shown, but which the consumer does not potentially own). These target displays are co-defined because the image is created specifically for the target display (i.e. the luminance is graded towards a specific desired object luminance). For example, an explosion cannot be made very bright on a 550nit display, so the grader may want to reduce the luminance of other image objects so that the explosion appears at least somewhat contrasty. However, it may well be that no one owns such a display, and the image still needs to be optimized by display adaptation to the actual display owned by the particular viewer. The range of luminance that is physically displayable on this end user display is shown as the rightmost luminance range (EU_DISP).

[0165] This information for the target display(s) constitutes metadata, some of which is typically communicated along with the image itself, i.e., along with the image pixel intensities.

[0166] This approach to characterizing (encoding) HDR images is a significant departure from traditional SDR image coding, and because these aspects are only recently invented, and because important technical aspects should not be misunderstood, the concepts involved are summarized for the reader using FIG. 8.

[0167] HDR video is well complemented by metadata, since many aspects may differ (e.g. the maximum luminance of SDR displays has always been in the range of approximately 100 nits, but now people have displays with significantly different display capabilities, e.g. PL_D equal to 50 nits, 500 nits, 1000 nits, 2500 nits, maybe even 10,000 nits in the future, content characteristics like the maximum codable luminance PL_V of the video also change considerably, and therefore the distribution of luminance between darks and lights that a grader makes for a typical scene also differs a lot between a typical SDR image and any HDR image, etc.), and should not run into difficulties due to insufficient control over those various unstable ranges.

[0168] As explained above, one should obtain at least one pixelation matrix of pixel colors including at least pixel luma, otherwise the image shape will not even be visible (even if colorimetrically misrepresented).As mentioned above, by actually communicating only one image per time, it is possible to communicate two different dynamic range images (which can serve as two reference gradings to indicate the need for luminance re-grading of a particular video content in case an image with a different dynamic range needs to be created, such as an MDR image).

[0169] It is assumed that the master HDR image itself is communicated, so that the first dataset 801 (image color) contains the color component triplets of the pixels of the image being communicated and will be the input to a color processing circuit present, for example, in a receiving television display.

[0170] Typically such images are digitized as a (for example 10-bit) luma and two chroma components Cr and Cb (although non-linear R'G'B' components can also be communicated). However, it is necessary to know what luminance the luma 1023 or for example 229 represents.

[0171] Therefore, they communicate container metadata 802 together. Assume that luma is defined, for example, according to the perceptual quantizer EOTF (or its inverse OETF), as standardized in SMPTE ST.2084. This is a large container that can specify luma up to 10,000 nits. It can be said to be a "theoretical container" that contains the actually used luminance up to, for example, 2500 nits, even if images are not currently produced with such high pixel luminance. (Note that primary chromaticity is also communicated, but those details simply unnecessarily hinder this description.)

[0172] Of interest is the maximum pixel luminance that can (or will) actually be encoded for the video, which is encoded in another video characteristic metadata, typically the master display color volume metadata 803.

[0173] This is an important luminance value for the receiver to know, because even if it doesn't care about the specific details of how the display remaps all luminances along the range (at least according to the content creator's desired display adaptation), knowing the maximum still guides it roughly what is best to do at all luminances, since it at least knows what luminance the video will not exceed for an image pixel.

[0174] In the 2000 nit example, this Master Display Color Volume (MDCV) metadata 803 includes the master HDR image maximum luminance, i.e., the PL_V_HDR value in FIG. 7, i.e., characterizing the master HDR video (in this example, it is actually also communicated as the PQ pixel luma defined in SMPTE 2084, but that aspect can be ignored for now, since those skilled in the art know how to convert between the two and can understand the inventive color processing principles as if the (linear) pixel luminance itself were coming in).

[0175] This MDCV is the "virtual" display, or target display, for the HDR master image. By indicating this in the metadata, the video creator is indicating where in the movie there are pixels with 2000 nits of brightness in the video signal that is communicated to the receiving actual end-user display, so that the end-user display will take that fully into account when processing the luminance of the current image.

[0176] These (actual) brightnesses of a set of images are in fact yet another aspect, and therefore there is a further video-related metadata set 804, which gives further information of the actual video, rather than the characteristics of the relevant display (i.e., the maximum possible for the video, for example). To easily understand this, assume that two videos are annotated with the same MDCV PL_V_HDR (and EOTF). The first video is a night video, and therefore in fact the images do not reach a pixel brightness higher than, for example, 80 nits (although it is still specified with an MDCV of 2000 nits; furthermore, if it is another video, it may have flashlights in at least one image, which has a few pixels that reach the 2000 nit level, or almost to it), while the second video, specified / created according to exactly the same encoding technique (i.e., annotated with the same data in 802 and 803), consists only of explosions, i.e., has mostly pixels above 1000 nits.

[0177] On the one hand, we may want to say something additional about this video, but on the other hand, those skilled in the art can understand that if we want to regrade both videos from a 2000 nit representation to, say, a 200 nit output representation, we would do it differently (we could scale the explosion by simply dividing the luminance by 10, while keeping the luminance of the night scene the same in the master HDR and the 200 nit output image).

[0178] A possible useful metadata in set 804 annotating the communicated HDR image (optional in the present invention, but described nevertheless for completeness) is the average pixel luminance of all pixels of all time series images, with MaxFall being, for example, 350 nits. The receiving color processing can then understand from this value that if it is dark, the luminance is displayed as is, i.e., without mapping, and if it is bright, dimming is required.

[0179] Even if one does not actually communicate SDR video, i.e., only sends metadata (metadata 814), one can also annotate the SDR video (i.e., a second reference-graded image, which indicates how an SDR image should look as similar as possible to the master HDR image by the content creator in the case of reduced dynamic range capabilities).

[0180] Thus, some HDR codes may also send pixel color triplets, i.e., SDR pixel color data 811, that include the pixel lumas of an SDR image (Rec. 709 OETF definition), but as explained above, the illustrative codec, i.e., the SLHDR codec, does not actually communicate this SDR image, i.e., the pixel colors, or anything that depends on that color, and is therefore a deletion (if pixel colors are not communicated, there is also no need to communicate SDR container metadata 812, which indicates the container format of how the pixel codes are defined and should be decoded into linear RGB pixel color components).

[0181] What should ideally be communicated (although some systems implicitly assume it) is the corresponding SDR target display metadata 813. In such a situation, one would typically fill in the value of PL_V_SDR as equal to 100 nits.

[0182] Importantly for the present invention, this is also where the video creator typically writes in the value of the assumed minimum black of the theoretical target SDR display for which the film was graded, ie, the SDR reference minimum black, mB_SDR.

[0183] For example, if an author assumes that he is making a video for a display that cannot go deeper than 0.1 nit (e.g. due to LCD leakage), he does not want very important image object pixel luminances to be close to this value, so he will start with, say, 0.2 nit SDR image pixels, maybe go a little above that, with most pixels above the 1 nit level. There is a similar value that characterizes the master HDR image, or more precisely the target display associated with it, namely HDR minimum black mB_HDR (in metadata 803).

[0184] The need for regrading is image dependent and is therefore advantageously encoded into the SDR video metadata 814 according to the codec of the selected description as explained above. As explained above, in general, one communicates one (or several, even a single time) image optimized luminance mapping functions (i.e., F_L reference luminance mapping function shapes) for mapping HDR luminance (or luma) normalized to up to 1 to SDR luminance or luma normalized to up to 1 (the exact manner of encoding these functions is irrelevant to this patent application, examples can be found in the above-mentioned ETSI SLHDR standard). Since now, with this function shape (combined with the target display's maximum luminance, HDR, and SDR metadata) the required remapping of all possible image pixel luminances is specified, further metadata about the image, such as maxFall, can also be communicated but is not actually required.

[0185] This already constitutes a fairly specialized set of HDR video coding data, to which current ambient adaptive luminance remapping techniques can be applied.

[0186] However, two further sets of metadata (particularly useful for contrast optimization embodiments, described in more detail below) may be added (in various ways), which are not yet fully standardized: Content creators, using a perceptual quantizer, may work under an implicit assumption that the video will be created in a viewing environment with a particular illuminance level, e.g., 10 lux (and so will ideally also be displayed as best as possible at the receiving end for optimal viewing) (whether a particular content creator also strictly follows this lighting suggestion is another matter).

[0187] If one wants to improve certainty, one can associate the typical (intended) viewing environment for which the master HDR image was specifically graded (HDR Viewing Metadata 805; optional / dotted line). As before, a 2000 nit maximum master HDR image can be made. However, if this image is intended to be viewed in a 1000 lux viewing environment, the grader will probably not make too many subtle graded dark object luminances (like slightly different dimly lit objects in a dark room seen through an open door behind a scene containing a front lit first room) since the viewer's brain will likely simply see all of this as "flat black", whereas the situation is different if the image is made for typical, say, dim evening viewing in a 50 lux or 10 lux room. Regrading to the SDR image, particularly the F_L function, can be done more specifically for typical brighter viewing conditions, e.g., 200 lux, and annotated into the SDR ambient metadata 815, and the receiving device can also take advantage of that information if it so desires (or communicate regrading for different SDR images for different intended viewing).

[0188] Returning to FIG. 7, we show the settings (by display adaptation) we wish to create an output image for a 600 nit MDR display, ie the output image should be a PL_V_MDR=600 nit output image.

[0189] If according to the present invention we were to optimize an output image with the same maximum luminance as the input image (the simple situation described above in FIG. 6), the value of the minimum luminance of the target display (mL_VD) would simply be the mB_HDR value. However, now we need to set the minimum luminance of the MDR display dynamic range (as an appropriate value of the minimum luminance of the target display mL_VD). This is interpolated from the luminance range information of the two reference images (i.e., in the description of the standardized embodiment in FIG. 8, the MDCV and the presentation display color volume PDCV are typically jointly communicated or at least obtainable in the metadata 803 and 813, respectively).

[0190] The formula is as follows: mL_VD=mB_SDR+(mB_HDR-mB_SDR)*(PL_V_MDR-PL_V_SDR) / (PL_V_HDR-PL_V_SDR) [Formula 8]

[0191] In this formula, the PL_V_MDR value is selected to be equal to the PL_D representation of the display to which the display-optimized image is to be provided.

[0192] Returning to FIG. 6, the dif value is converted to a psychovisually uniform luma difference (Ydif) in the photoelectric conversion circuit 611 by applying the v-function of Equation 3 and substituting the value dif (i.e., the normalized luminance difference) divided by PL_V_in for L_in, where as the value of PL_V_in, the value PL_V_HDR is used when producing an ambient adjusted image with the same maximum luminance as the input master HDR image, and in a display scenario adapted to an MDR display, the value of PL_V_in is the PL_D value of the display, e.g. 600 nit.

[0193] The electro-optical conversion circuit 604 calculates the normalized linear version of the intermediate luma Yim, the intermediate luminance Ln_im. It applies the inverse formula of formula 3 (the same definition of RHO). For the RHO value, in fact, in the simplest situation where the output image has the same maximum luminance as the input image and only adjusts the darker pixel luminance, the same PL_V_HDR value is used. However, in case of display adaptation to MDR maximum luminance, the PL_D value is used to calculate the appropriate RHO value that characterizes the specific shape (steepness) of the v-function. To achieve this, an output maximum determination circuit 671 is present in the color processing circuit 600. It is generally a logic processor that determines whether to use, for the configured situation, the maximum luminance of the output image PL_O, i.e., the PL_V_HDR or the PL_D, respectively, as the value for determining the RHO of the EOTF applied by the electro-optical conversion circuit 604 (those skilled in the art will understand that the situation may be configured with a fixed formula in some specific fixed variants).

[0194] The final surround adjustment circuit 605 performs a linear additive offset in the luminance domain by calculating the final normalized luminance Ln_f using the following equation: Ln_f=(Ln_im-(mL_De2 / PL_O)) / (1-(mL_De2 / PL_O)) [Formula 9] mL_De2 is the second minimum luminance of the end-user display (aka second end-user display minimum luminance) and is typically input via a third minimum metadata input 695 (connected to a light meter via an intermediate processing circuit). It differs from the first minimum luminance of the end-user display mL_De in that mL-De further includes a characteristic value of the physical black of the display (mB_fD) while mL_De does not, characterizing only the amount of ambient light that degrades the displayed image (e.g. by reflection), i.e. only the typically equal mB_sur.

[0195] Inside the final perimeter adjustment circuit 605, the mL_De2 value is normalized by the applicable PL_O value, ie, for example, PL_D.

[0196] Finally, in most variants, it is advantageous if a regular (ie, unnormalized) output luminance L_o appears, which is realized by a multiplier 606 that calculates: L_o=Ln_f*PL_O [Equation 10] That is, it is normalized to the maximum applicable luminance of the output image.

[0197] Advantageously, some embodiments do this in color processing that not only adjusts the darker brightness for ambient light conditions, but also optimizes for the reduction in maximum brightness of the display. In this scenario, the brightness mapping circuit 602 applies an appropriate calculated display-optimized brightness mapping function FL_DA(t), which is typically loaded into the brightness mapping circuit 602 by a display optimization circuit 670. Those skilled in the art will understand that the particular manner of display optimization is merely a variable part of such an embodiment and is not typical for the ambient adaptation element, but some examples are shown with Figures 4 and 5 (the input of configurable PL_O values ​​to 670 is not depicted to avoid overcomplicating Figure 6, as those skilled in the art will understand this). In general, display adaptation has the property that the smaller the difference between the input maximum luminance and the output maximum luminance, the closer the function is to the diagonal, i.e., the closer the desired maximum luminance of the output image (i.e., PL_V_MDR=PL_D) is to the input maximum luminance (generally assuming PL_V_HDR), and therefore the further away from the maximum luminance of the second reference grading (generally, PL_V_SDR), the flatter the shape of the function (i.e., the "lighter" regrading function version FL_DA between F_L and the diagonal). Note that in downgrading, we generally have a convex function, which means that darker luminances (below some midpoint) are relatively boosted at the expense of compressing brighter luminances. So, in downgrading, FL_DA generally has a less steep slope to boost the darkest input luminances below the F_L function (with full regrading to the second reference grading at the other extreme end).

[0198] Below, a second invention is taught for optimizing image contrast, which is useful in brighter surroundings. These elements may be used in conjunction with the above-mentioned ambient adjustments in various embodiments, but each invention may also be applied in isolation from the others.

[0199] This method of processing an input image, in particular to improve contrast, generally consists of obtaining a reference luminance mapping function (F_L) associated with the input image, which defines the need for regrading by a luminance (or equivalently luma) mapping to the luminance of a corresponding secondary reference image.

[0200] The input image generally also serves as the first grading reference image.

[0201] The output image generally corresponds to a maximum luminance situation midway between the maximum luminance of the two reference graded images (aka reference grades) and is generally calculated for a corresponding MDR display (a medium dynamic range HDR display compared to the master HDR input image), and the optimized output image maximum luminance relationship is PL_V_MDRPL_D_MDR.

[0202] Display Adaptation (of any embodiment) applies as usual to Display Adaptation for lower maximum brightness displays, but here it is applied specifically differently (i.e. most of the technical elements of Display Adaptation remain the same, but some are changed).

[0203] The display adaptation process determines an adaptive luminance mapping function (FL_DA), which is based on a reference luminance mapping function (F_L). This function F_L is static, i.e. the same for several images (in which circumstances, e.g., if the ambient lighting changes significantly or upon user control actions, FL_DA still changes), but it may also change over time (F_L(t)). The reference luminance mapping function (F_L) typically comes from the content creator, but may also come from an optimal regrading function computing automaton in the receiving device of the video image (just like the offset determination embodiment described in Figures 6 to 8).

[0204] An adapted luminance mapping function (FL_DA) is applied to the input image pixel luma to obtain the output luminance.

[0205] An important difference with existing display adaptations (while the definition of the metric, i.e. the math for finding the various maximum luminance, and the direction of the metric are the same; e.g. the metric is scaled by having one point at any position on the diagonal corresponding to the Yn_CC0 luma normalized to 1, and the other point is somewhere on the locus of the F_L function, e.g. vertically upwards, that corresponds to the output luma or luminance of F_L when the other coordinate of the metric positioning end point is the input Yn_CC0 luma to the function F_L) is now that the position on the metric (more precisely, all its scaled versions due to the shape of the F_L function) to obtain the adaptive luminance mapping function (FL_DA) is calculated based on the adjusted maximum luminance value (PL_V_CO) rather than the required maximum value of the output image (generally PL_V_MDR).

[0206] This adjusted maximum brightness value (PL_V_CO) is determined by the user of the device, typically the viewer of the connected display (or the device may be within a display). The user determines the optimal control value, i.e. the User Correction Value (UCBVal), which typically serves to reduce (downgrade) the maximum brightness used in the display adaptation algorithm, so that this algorithm no longer works with the actual (physical) maximum brightness achievable by the connected display (e.g. in a backlit LCD the maximum light a pixel can output is determined by setting the backlight to maximum, settings by the manufacturer to ensure correct operation, e.g. cooling and longevity, and controlling the liquid crystal pixels to transmit as much light as possible), but with a virtual user control value. Typically the user controls by offsetting in one direction from a starting set point. This set point is optionally influenced by a measurement of the ambient light level, e.g. the ambient illumination value (Lx_sur).

[0207] The ambient illumination value (Lx_sur) can be obtained in a variety of ways, for example a viewer can determine it empirically by checking the visibility of a test pattern, but it typically comes from measurements with a luminance meter 902, which is typically appropriately positioned relative to the display (e.g., at the bezel edge, facing approximately in the same direction as the screen front plate, or integrated into the side of the mobile phone, etc.).

[0208] The reference ambient value GenVwLx may further be determined in various ways, but is generally fixed as it relates to what is expected to be reasonable ("average") ambient lighting for typical target viewing situations.

[0209] This method may be used without a GenVwLx value, however, which serves as an intermediate point for user control.

[0210] For television display viewing, this is typically the viewing room.

[0211] The actual lighting in a living room can vary considerably depending, for example, on whether the viewer is watching during the day or at night, and also on what room configuration they are watching in (e.g., whether there is a small or large window and how the display is positioned relative to the window, or whether mood lamps are used in the evening, or whether other members of the family are doing precision work that requires a sufficient amount of lighting).

[0212] For example, even during the day, when the sky suddenly darkens significantly because of an approaching hailstorm, outdoor lighting can be as low as 200 lux (lx), and indoors, the light level from natural lighting is typically 100 times lower, so indoors it is only 2 lx. This starts to have a nocturnal appearance (particularly strange during the day), so this is something that many users generally turn on at least one lamp for comfort, effectively raising the level again. Normal outdoor levels are 10,000 lx in winter to 100,000 lx in summer, so more than 50 times brighter.

[0213] However, other viewers may find it advantageous to watch videos (especially HDR videos) in the dark, for example to enjoy a horror movie as more frightening and / or to see dark scenes better.

[0214] Although common in the Middle Ages, illumination from a single candle is nowadays a lower limit. The reason is that for city dwellers, at such a level, there is simply more light leaking out of outdoor lamps such as city lights. The candela used is defined as the brightness of a typical candle, so if you place a surface one meter away from a candle, you get 1lx, which still allows you to see things, but makes text on paper, for example, very hard to read (for reference, 1lx is also typical of an outdoor moonlit scene). So, if you light a 5-meter wide room with several candles, that level of illumination is achieved. Even a single 40W incandescent bulb already produces about 40 times more brightness than a candle, so for most viewers, one or a few such lamps will be at a more typical ambient light level. So you can expect something like k*10lx for viewing in not too much (atmospheric) light. However, video is defined to be suitable for viewing during the day, where the illumination is n*50lx (e.g. a 200W light bulb set approximately 2 meters away gets approximately 3000 / 50lux; looking at a recipe display in the kitchen requires approximately 3 times higher light levels to safely perform cooking tasks like cutting).

[0215] In mobile / outdoor situations, the light level is higher, e.g. when sitting near a train window or in the shadow under a tree, etc. Then the light level is e.g. 1000lx.

[0216] Without intending to be limiting, assume that a good value for GenVwLx for video viewing of television programs is 100 lux.

[0217] Assume that the light sensor measures Lx_sur=550 lux.

[0218] The user then controls by setting such a value of UCBVal that the image looks better under levels that are effectively 5 times higher than the ideal 550 lux level. This is done by the user adjusting the adjusted maximum luminance value (PL_V_CO) to such an extent that the darkest luminance subrange for a particular F_L function is stretched to a visually acceptable level. This is done in the maximum luminance determination unit (901) by calculating the value to be output of the adjusted maximum luminance value PL_V_CO as follows: PL_V_CO=PL_V_MDR-UCBVal [Formula 11] PL_V_MDR is generally (but not necessarily) equal to the maximum displayable pixel luminance of the connected display.

[0219] The user control value UCBVal is controlled by appropriately scaled user input, e.g., the slider setting (UCBSliVal, e.g., falls off symmetrically near zero offset or starts at zero offset, etc.) is scaled in such a way that when the slider is at its maximum, the user does not ridiculously change the contrast, e.g., all dark image areas appear almost like bright HDR white.

[0220] For this purpose, the device (e.g., display) manufacturer pre-designs the appropriate intensity value (EXCS), and then the formula is: UCBVal=PL_D*UCBSliVal*EXCS [Formula 12] It is.

[0221] For example, if we want 100% to correspond to an additional maximum luminance-based contrast change of 10%, we get: PL_D*1*EXCS=0.1*PL_D, therefore EXCS=0.1 etc. (a value of 0.75 has been found to work well in particular embodiments).

[0222] FIG. 9 shows possible implementation elements in a typical device configuration.

[0223] The device for processing an input image to obtain an output image (aka ambient optimization display optimization device) 900 has a data input (920) for receiving a reference luminance mapping function (F_L), which is metadata associated with the input image. This function again specifies the relationship between the luminance of a first reference image and the luminance of a second reference image. These two images are again graded in some embodiments by the creator of the video and are co-communicated with the video itself as metadata, for example via satellite television broadcast. However, the appropriate regrading luminance mapping function F_L is also determined by a regrading automaton in the receiving device, for example a television display. The function changes over time (F_L(t)).

[0224] The input image generally serves as the first reference grading, based on which the display-optimized image and the environment-optimized image are determined, which is generally an HDR image. The maximum luminance of the display-adapted output image (i.e., the output maximum luminance PL_V_MDR) generally falls between the maximum luminance of the two reference images.

[0225] The device optionally includes, or is equivalently connected to, a light meter (902) configured to determine the amount of ambient light that falls on the display, i.e., the display on which the image optimized for viewing is being provided, although this is generally not required for current user control.

[0226] The display adaptation circuit 510 is configured to determine an adaptive luminance mapping function (FL_DA), which is based on the reference luminance mapping function (F_L). It also actually performs pixel color processing and therefore includes a luminance mapper (915) similar to the color converter 202 described above. Color processing is also involved. The configuration processor 511 performs the actual decision of the (ambient optimized) luminance mapping function to be used before performing the pixel-by-pixel processing of the current image. Such an input image (513) is received via an image input (921), e.g. an IC pin, which itself is connected to an image source to the device, e.g. an HDMI cable, etc.

[0227] An adaptive luminance mapping function (FL_DA) is determined based on the reference luminance mapping function F_L and the value of the maximum luminance according to a variation of the display adaptation algorithm described above, but where the maximum luminance is not the typical maximum luminance of the connected display (PL_D) but a specially adjusted maximum luminance value (PL_V_CO) adjusted for the lighting conditions of the viewing environment (and potentially further user corrections). The luminance mapper applies the adaptive luminance mapping function (FL_DA) to the input pixel luminance to obtain the output luminance.

[0228] To calculate the adjusted maximum luminance value (PL_V_CO), the device includes a maximum luminance determination unit (901), which is connected to the display adaptation circuit 510 and supplies this adjusted maximum luminance value (PL_V_CO) to the display adaptation circuit (510).

[0229] This maximum brightness determination unit (901) gets a reference illumination value (GenVwLx) from memory 905 (e.g., this value is pre-stored by the device manufacturer, or selectable based on what type of image is coming in, or loaded with typical intended ambient metadata to be associated with the image, etc.). It further gets a maximum brightness (PL_V_MDR), which is, for example, a fixed value stored in the display or configurable in a device (e.g., a set-top box or other image pre-processing device) that can feed images to various displays.

[0230] In some embodiments, the user (viewer) controls the automatic ambient optimization of the display adaptation according to the user's preferences, for example using a user interface control component 904, for example a slider (or rotary knob etc., which does not have to be a physically present button, but is a finger-controllable element on the screen of a mobile phone controlling the device for example) coupled thereto, thereby allowing the user to set a higher or lower value, for example a slider setting UCBSliVal. This value is input to a user value circuit 903, which communicates the user control value UCBVal to the maximum brightness determination unit 901.

[0231] FIG. 10 shows an example of processing with a particular luma mapping function F_L. We assume that the reference mapping function (1000) is a simple luma identity transformation, i.e., clipping above the maximum display capability of 600 nit. The luma mapping is now represented by a plot of the equivalent psychovisually equalized luma, which can be calculated according to Equation 3. We assume that the input normalized luma Yn_i corresponds to the HDR input luma, which is a 1000 nit maximum luma HDR image (i.e., an RHO of PL_V_HDR=1000 nit is used in the formula). For the output normalized luma Yn_o, we assume an exemplary display of 600 nit, and therefore the output normalized luma Yn_o is converted back to luma by using the RHO value corresponding to 600 nit. For convenience, on the right and above, there are luma corresponding to the luma positions. Thus, this selected F_L reference mapping function 1000 performs equal compression in the visually equalized luma domain. The (ambient) adaptive luminance mapping function FL_DA is shown as curve 1001. On the one hand, an offset Bko is seen, which depends, among other things, on the leak black of the display. On the other hand, a curvature is seen that mainly boosts the darkest black, which is due to residual effects of the non-linear psychovisually equalized processing in the luma domain.

[0232] 11 shows an example of an embodiment of contrast boosting (ambient lighting offset is set to zero, but as before both processes are combined). When modifying the luminance mapping in the relatively normalized psychovisually uniformed luma domain, this relative domain starts from a value of zero. The master HDR image starts with some black value, but at the lesser output, a black offset is used to map this to zero (or to the minimum black of the display, note: displays may have different behaviors of darker inputs, e.g. clipping, so a true zero can be placed there).

[0233] When using modified display adaptation, the black zero input generally maps to black zero, whatever the value of the adjusted maximum luminance value (PL_V_CO). Generally, the zero point of the output luma starts at the virtual black level mL_VD (ideally) or at the minimum black mB_fD of the actual end-user display, but in any case there is a brightness difference in the darkest colors, which results in a better visible contrast for dark pictures such as night scenes (the input luminance histogram 1010 is stretched as the output luminance histogram 1011 on the normalized luma axis of the output luma, which gives a sufficient image contrast despite the lower PL_V_MDR=PL_D=600nit). When the two methods are combined, all the processing to obtain the appropriate black level offset Bko is optimally transferred to the processing described in FIG. 6 and the like (i.e., the adjusted contrast processing can be performed with the image starting at zero and the function starting at zero). In practice, we simply always define the F_L curve starting from zero (maps zero HDR luminance or luma to zero output luminance or luma) because the display-adaptive luminance (luma) mapping simply occurs anyway, regardless of whether zero actually occurs in the input image.

[0234] The algorithmic components disclosed herein may in fact be implemented (wholly or partially) as hardware (e.g., as part of an application specific IC) or as software running on a dedicated digital signal processor or a general purpose processor, etc.

[0235] It should be possible for a person skilled in the art to understand from the inventors' presentation which components are optional improvements and are implemented in combination with other components, and how the (optional) steps of the method correspond to the respective means of the apparatus and vice versa. The term "apparatus" in this application is used in the broadest sense, i.e., a group of means enabling the realization of a specific purpose, and thus, for example, an IC (a small circuit part of an IC), or a dedicated device (such as a device with a display), or a part of a networked system, etc. The term "arrangement" is further intended to be used in the broadest sense, and thus includes, inter alia, a single device, a part of a device, a collection of (parts of) cooperating devices, etc.

[0236] The meaning of a computer program product should be understood to encompass any physical realization of a set of commands that enables a general-purpose or dedicated processor, after a series of loading steps (including intermediate conversion steps, e.g. conversion into an intermediate language and into a final processor language), to input the commands to the processor and to execute any of the characteristic functions of the invention. In particular, a computer program product is realized as data on a carrier, e.g. a disk or tape, data residing in a memory, data moving via a network connection (wired or wireless), or as a program code on paper. Apart from the program code, characteristic data required by the program may also be embodied as a computer program product.

[0237] Some of the steps required for the operation of the method may already be present in the functionality of the processor instead of those described in the computer program product, such as data input and output steps.

[0238] It should be noted that the above embodiments are illustrative rather than limiting of the present invention. For the sake of brevity, not all these options are detailed, when a person skilled in the art can easily realize that the presented examples are mapped to other areas of the claims. Apart from the combination of the elements of the present invention as combined in the claims, other combinations of elements are possible. Any combination of elements may be realized in a single dedicated element.

[0239] Any reference signs placed between parentheses in the claims are not intended to limit the claim. The words "comprises", "comprises", "having" do not exclude the presence of elements or aspects not listed in the claims. The word "a" or "an" in the "singular" does not exclude the presence of a plurality of such elements.

Claims

1. 1. A method for processing an input image to obtain an output image, comprising the steps of: the input image having pixels having input luminance within a first luminance dynamic range having a first maximum luminance; A reference luminance mapping function is received as metadata associated with the input image; the reference luminance mapping function specifies a relationship between luminances of located pixels in two images, the two images being graded differently in that pixel luminances of the same image object have different pixel luminances in the two images; the reference luminance mapping function specifies a relationship between the luminance of the input image and the luminance of a secondary reference image having a second reference maximum luminance; the output image has an output maximum luminance different from the first maximum luminance and the second reference maximum luminance; The process comprises: determining an adaptive luminance mapping function based on the reference luminance mapping function and an adjusted maximum luminance value, the adjusted maximum luminance value being different from the output maximum luminance; applying the adaptive luminance mapping function to input pixel luminances to obtain an output luminance of the output image; calculating the adaptive luminance mapping function includes finding a position on a metric that specifies a location of maximum luminance, the position corresponding to the adjusted maximum luminance value; a first endpoint of the metric corresponds to the first maximum luminance and a second endpoint of the metric corresponds to a second reference maximum luminance; the first endpoint of the metric is located, for any normalized input luma, at a diagonal point having horizontal and vertical coordinates equal to that normalized input luma; The method of claim 1, wherein the second end point is located on a locus of the reference luminance mapping function determined by an orientation of the metric, The adjusted maximum brightness value is by obtaining a user correction value and determining the adjusted maximum luminance value as a result of subtracting the user correction value from the output maximum luminance; writing said output luminance as a pixel color into said output image and outputting said output image; 1. A method for processing an input image, comprising:

2. 2. The method of claim 1, wherein the output maximum luminance is set equal to a maximum displayable pixel luminance of a display to which the output image may be supplied.

3. obtaining input from a user of the display; calculating a user control value by multiplying the input value by an intensity value and the output maximum brightness; 3. A method for processing an input image according to claim 1, comprising:

4. 4. A method for processing an input image according to claim 1, wherein said processing is performed in addition to setting the darkest black in the input image as a function of an ambient lighting value to a black offset value.

5. 5. A method for processing an input image according to claim 1, wherein the direction of the metric is preset as vertical and the metric places the second end point relative to a normalized input luma at a location having a result of applying the reference luminance mapping function with the normalized input luma as input as the horizontal coordinate and the normalized input luma as the vertical coordinate.

6. 1. An apparatus for processing an input image to obtain an output image, comprising: The input image has pixels having an input luminance within a first luminance dynamic range, the first luminance dynamic range having a first maximum luminance, and the apparatus: an image input unit for acquiring the input image; a data input for receiving a reference luminance mapping function, the reference luminance mapping function being metadata associated with the input image; the reference luminance mapping function specifies a relationship between a luminance of a first reference image and a luminance of a second reference image; the reference luminance mapping function specifies a relationship between the luminance of the input image and the luminance of the second reference image; the second reference image having a second reference maximum luminance; the output image has an output maximum luminance different from the first maximum luminance and the second reference maximum luminance; The apparatus, a user value circuit for determining and outputting a user correction value as set by a human user of said device; obtaining the user correction value from the user value circuit; Outputting an adjusted maximum luminance value as a result of subtracting the user correction value from the output maximum luminance. Maximum brightness determination unit for Further comprising: The apparatus, a display adaptation circuit for determining an adaptive luminance mapping function based on the adjusted maximum luminance value and the reference luminance mapping function; calculating the adaptive luminance mapping function includes finding a position on a metric that corresponds to the adjusted maximum luminance value; a first endpoint of the metric corresponds to the first maximum luminance and a second endpoint of the metric corresponds to the maximum luminance of the second reference image; the first endpoint of the metric is located, for any normalized input luma, at a diagonal point having horizontal and vertical coordinates equal to the normalized input luma; the second end point is located on the locus of the reference luminance mapping function; An apparatus for processing an input image, wherein a display adaptation unit applies the adaptive luminance mapping function to an input luminance, or an input luma that encodes the input luminance, to obtain an output luminance or output luma, and outputs the output luma or output luminance as pixel colors of the output image.

7. the user value circuit comprising: Get input from the user on the display, calculating a user control value by multiplying said input value by an intensity value and said output maximum brightness; 7. The apparatus for processing an input image according to claim 6, further comprising: a processor for outputting said user control value to said maximum brightness determination unit.

8. 8. An apparatus for processing an input image according to claim 6 or 7, comprising a black adaptation circuit for setting the darkest black luma in the input image to a black offset luma as a function of ambient illumination value.