Optimization of HDR Image Display

The method processes HDR video images by adapting luminance mapping functions to match different display dynamic ranges, addressing the challenge of inconsistent image quality across various displays and achieving visually enhanced HDR images.

JP7695482B2Active Publication Date: 2025-06-18KONINKLIJKE PHILIPS NV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024530478
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-26
Filing Date
2022-11-14
Publication Date
2025-06-18
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively adapt high dynamic range (HDR) video images to various display devices with different dynamic ranges, leading to inconsistent and suboptimal image quality.

Method used

A method and apparatus for processing HDR video images by determining a luminance mapping function adapted based on a reference luminance mapping function, using a pre-fixed display adaptation algorithm to normalize metrics along a line segment, and applying an adjusted luminance mapping function to achieve optimal display adaptation across different dynamic ranges.

Benefits of technology

The solution enables the creation of visually better HDR images by ensuring that the luminance of image subjects is optimally displayed across a wide range of dynamic ranges, maintaining artistic intent while accommodating varying display capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695482000001
    Figure 0007695482000001
  • Figure 0007695482000002
    Figure 0007695482000002
  • Figure 0007695482000003
    Figure 0007695482000003
Patent Text Reader

Abstract

1. In order to obtain a better and more viewable HDR image in a practical manner, several variants of an image processing device have been proposed for processing an input image Im_Comm of an input video to obtain an output image of an output video, the input image having pixels with an input luminance Ln_in_pix falling within a first luminance dynamic range DR_1, the first luminance dynamic range having a first maximum luminance PL_V_HDR, the output image having pixels with an output luminance that can be calculated from the input luminance and falling within a second luminance dynamic range, the second luminance dynamic range having a second maximum luminance PL_V_MDR, the image processing device comprises a video data input 227 configured to receive the input image and a reference luminance mapping function that is encoded as metadata associated with the input image, the reference luminance mapping function specifying a relationship between the luminance of the first reference image and the luminance of a second reference image, the first reference image being the input image and the second reference image having a second reference maximum luminance PL_V_SDR, the image processing device further comprising: a display adaptation unit 209 configured to determine an adapted luminance mapping function FL_DA based on F_L, the display adaptation unit using a pre-fixed display adaptation algorithm, the algorithm identifying, for each point on the diagonal line, a respective metric along a line segment starting on the diagonal line oriented in a pre-fixed direction in a coordinate system of an input luminance normalized to a maximum of 1 and an output luminance normalized to a maximum of 1, the respective metric at each position being normalized by giving a value of 1 to an intersection point between the line segment and a locus of the luminance function F_L in the coordinate system; determining identifying a position on each respective metric corresponding to a second maximum luminance PL_V_MDR, the set of positions on the respective metric being output as the adapted luminance mapping function FL_DA; the image processing device includes a boost decision circuit 700 configured to calculate a booster strength value BO, the boost decision circuit 700 comprising:a histogram calculation unit 721 configured to determine a histogram of intermediate luminances obtained by applying the adapted luminance mapping function FL_DA to the input luminance; an average calculation circuit 705 configured to calculate an average luminance magnitude AB based on the histogram; a first converter unit 706 configured to calculate a first intensity value PosB from the average luminance magnitude AB; a second measurement unit 711 configured to calculate a weighted sum BDE of pixel counts in at least two configurable upper bins of the histogram; and a second conversion unit 712 configured to calculate a second intensity value NegB from the weighted sum BDE; determines a booster intensity value BO based on a value DV equal to the first intensity value PosB minus the second intensity value NegB, and the image processing device is further configured to use a display adaptation unit to calculate an adjusted adapted luminance mapping function F_ALT_B, the adjusted adapted luminance mapping function being calculated as the adjusted adapted luminance mapping function F_ALT_B by using a display adaptation algorithm configured to output a locus of positions equal to the booster intensity value on a metric, and the image processing device includes a color converter 208 configured to apply the adjusted adapted luminance mapping function F_ALT_B to the input luminance to obtain an output luminance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and apparatus for adapting the image pixel luminance of high dynamic range video to obtain a clear image for a specific display.

Background Art

[0002] Over the past few decades, the standard modes of video capturing, coding, and displaying (now called low dynamic range (LDR) or standard dynamic range (SDR)) have been quite sufficient to provide end customers with all kinds of images (in television broadcasts, video conferencing, gaming, etc.).

[0003] The system was based on defining relative brightness (with 0% being the darkest black and 100% being the brightest white) in a state where the luminance dynamic range of the display output was typically between 100:1 and 1000:1. This system can function in the case of a well-lit (almost uniformly lit) scene, so it essentially captures the nature of the subjects, in other words their own brightness, that is, the amount of light they reflect (+ about 4% to + about 90%). In such a scene, such capture is similar to what the brain should expect and what the eyes should see.

[0004] From the creative side, simply when a value for the camera iris is set based on the average amount of light in the scene and the illuminance of the subject is captured based on the filling rate of the photoelectrons collected by the pixels of each camera sensor, a sufficient image is provided. For historical reasons, the signal itself (when normalized, from 0 to 1) is not transmitted, but its square root is transmitted. Analog standards such as PAL represent the signal as a voltage between 0 mV and 700 mV, and today's digital standards quantize it to 255 levels, that is, transmit the relative brightness of the pixels as values between 0 and 255.

[0005] However, this approach had several problems. Real-world scenes are not illuminated in such a tightly controlled manner. Although there is still generally a limit to the ratio between the darkest and the brightest pixels (because bright entities will necessarily illuminate darker areas to some extent), there can be a fairly large dynamic range of luminance in a scene. For example, indoor lighting levels are typically less than 1 / 100th of outdoor levels even with relatively large windows, so the dynamic range is already at a lighting ratio of 100×subject reflectance of 100, i.e., 10,000:1, and small cracks in cave walls will result in an even larger dynamic range.

[0006] Until recently, displays simply did not have the ability to display a dynamic range far exceeding 100:1 or a peak brightness, also known as maximum luminance, far exceeding 100 nits. So when an image (i.e., that which becomes the displayed luminance) is created, the luminance of all scene subjects that cannot be clearly displayed is typically clipped to the highest white (or the darkest black).

[0007] However, several years ago, much better displays emerged. For example, by placing separately controllable LEDs behind the positions of LCD pixels, at least theoretically, the display area can be darkened or brightened as desired, even if the LCD pixels limit how much light from the LEDs they can transmit. Thus, today, it is possible to create displays with a maximum luminance of over 2000 nits and a dynamic range of over 10,000:1.

[0008] Thereby, in dealing with creating those images, it becomes possible to have a completely new style of image display (e.g., having the visual impression of a real environment illuminated differently).

[0009] However, even the conventional coder Rec.709 cannot perform coding over a large range of luminance that becomes the coded luminance in the created image at all. Therefore, a new format for coding, representing, and handling images (so-called high dynamic range (HDR) images) for starters is also needed.

[0010] Several years ago, novel techniques for coding high dynamic range video were introduced, especially by the applicant (see, for example, WO2017157977).

[0011] Video coding involves creating pixel color codes (e.g., one luma and two chroma per pixel) to depict an image. This is sometimes different from knowing how to optimally display an HDR image. For example, the simplest method is to simply utilize a highly non-linear opto-electronic transfer function OETF to convert the desired luminance into, for example, a 10-bit luma code, and vice versa, using an inversely shaped electro-optical transfer function EOTF to map, for example, a 10-bit electrical luma code to the optical pixel luminance to be displayed, thereby converting those video pixel luma codes to the luminance to be displayed. However, more complex systems can deviate in several directions, especially by decoupling the coding of an image from a particular use of the coded image.

[0012] One new approach for defining HDR images is to create them on the dynamic range provided by a so-called object display (which should not be confused with any actual display of the end user). The maximum luminance PL_V of that object display is transmitted simultaneously with the image in metadata, which indicates that the image was created within that range. Instead of directly loading a relative image from a camera, it is now possible to precisely define the luminance required in some master HDR image.

[0013] Therefore, now, for example, it is possible to accurately answer how bright a fireball explosion should be compared to the surrounding pixel luminance of the room (in SDR, typically, the iris exposed to the fireball is changed, i.e., as a result, many of those pixels will fall within the range 0 - 100%).

[0014] Therefore, the first video creator can create, for example, a 4000 nit PL_V video of this scene. In contrast to, for example, the case of a 2000 nit PL_V video, creating a 4000 nit video typically involves creating pixels with a luma code that codes pixels that are approximately 4000 nit (i.e., pixels that should be ideally displayed as long as they can be accepted as a 4000 nit display) somewhere in the video, i.e., for a moment. Given that the available range goes up to 4000 nit, and considering the impact of the fireball in the context of, for example, the size of the amount of pixels or the ratio of the image area included as the target, it may be determined that the brightest pixel of the fireball is 3000 nit and reserve 4000 nit for the subsequent laser beam in the movie.

[0015] The maker of an equivalent 2000 nit version of the same movie, when balancing all aspects (not only the visual impact of the fireball but also the relationship to the luminance of the surrounding pixels and the capabilities of its 2000 nit target display), places, for example, the brightest point of the fireball at 1950 nit. This defines various characteristics of the input image, and then, for example, any particular display with an (actual) display maximum luminance PL_D of 600 nit has to figure out how to map all the luminance values of the image subject pixels to its available dynamic range.

[0016] However, this new way of defining images opens the way to enabling a more predictable way of displaying those images. Displays without sufficient dynamic range actually have to perform some kind of brightness downscaling, but can do so in a way that respects the needs of the images, or more precisely, how the creator has defined their requantization needs. In the SDR era, images were simply displayed using a simple heuristic, that is, mapping the brightest pixel in the image (white, i.e., luma code 255) to the brightest displayable pixel. At that time, on a 200 nit PL_D display, the image appeared twice as bright as on a 100 nit display. When the ratio of the maximum brightness of the displays is so small, this is acceptable to both the viewer and the creator (who, because the eye and brain recalibrate and remove some of the difference, simply purchased a less bright and less expensive display with a somewhat nicer image of the same scene).

[0017] For example, when having a brighter HDR display with a PL_D of 1000 nit or more, not only does showing a 1000 nit white pixel become extreme, but more importantly, by using new techniques for coding and handling HDR images and thus showing realistic and impressive HDR images, that range of the display can be used more fully.

[0018] Using FIG. 1, it is clarified not only what the exact difference is between HDR images (despite the higher value of PL_V), but also how such a difference can be handled (reduced) when creating an equivalent lower dynamic range image, for example, a 100 nit PL_V SDR image corresponding to the HDR master image.

[0019] The grading in this application example is intended to mean, for example, either the human skin tone part, or the activity or the resulting image when the desired luminance is brought to the pixels by automation (for example, in the master image where grading is performed, the fireball is 5 times brighter rather than 4 or 6 times brighter than the surrounding pixels). When a person focuses on an image, for example, when designing an image, there will be several image subjects, and the person ideally wants to give the pixels of those subjects, and further the entire image and scene as given, a luminance spread of approximately the average luminance that is optimal for that subject. For example, if a person has available image capabilities such that the brightest encodable pixel of the image is 1000 nits (the maximum luminance PL_V of the image or video), the human skin tone part selects to give pixels with an explosion luminance value between 800 nits and 1000 nits to make the explosion look quite imposing, while another filmmaker selects, for example, an explosion brightness of 500 nits or less so as not to overly interfere with the rest of the timely image at that moment (of course, the technology should be able to handle both situations, and then for both grading parts, the dynamic range of the target display should be defined in a stable manner).

[0020] (That is, the) maximum luminance of the HDR image or video of the associated target display varies quite a lot. For example, typical values are, for example, 1000 nits or 4000 nits, or even 10,000 nits, and there is no limit (a person generally says that they have an HDR image when PL_V is at least 600 nits).

[0021] The HDR display has the maximum ability, that is, the most highly displayable pixel luminance of, for example, 600 nits or 1000 nits, or a number N × 1000 nits (starting from a lower-end HDR display).

[0022] Video creators generally cannot create videos that are optimal for each possible end-user display (i.e., the capabilities of the end-user display are optimally used by the video in such a way that the maximum brightness of the video never exceeds the maximum brightness of the end-consumer display and is not lowered to fully utilize the available dynamic range). That is, there is something that implements an attribute similar to the "display maximum image brightness as the maximum displayable brightness" in the SDR era, but this is no longer accidentally implemented electronically; rather, it is an attribute that can be precisely defined mathematically (moreover, it is possible to select to give a luminance of 4000 nits to only one pixel and perform grading on all other scene subjects with any desired lower luminance value).

[0023] Second, a second question arises as to how best to display an image including peak brightness PL_V on a display having (often much lower) display peak brightness PL_D, which in this text is referred to as display adaptation (or sometimes as display optimization). Specifically, this term will be used for non-arbitrary adaptation, but in display optimization, as shown by the following guidance regarding the required re-grading of pixel luma, the display optimization described below is done in the simplest technical manner by the video creator creating and communicating at least one luma mapping function. In the future, there will still be displays that require created dynamic range images that are lower than, for example, a 2000 nit PL_V image. In theory, a display can always perform re-grading so that the brightness of the image pixels becomes displayable by its own internal discovery method, i.e., it can map the brightness of the image pixels, but if the video creator has adequately addressed determining the pixel brightness, it is beneficial that he can further show how his image should be displayed when adapted to lower the PL_D value, and ideally the display will greatly comply with these technical requirements. If the display manufacturer performs re-grading of the fireball according to any of its internal display algorithms, it will not look as impressive as originally intended by the creator.

[0024] However, there is some uncertainty as to who will manage the display, i.e., who has the final say on how bright all the image subjects should be displayed. Since the video creator sets the scene, it could be argued that it should be the video creator. However, the end user may also decide, for example, that they do not prefer videos that are as bright as what they do. Finally, according to what the display manufacturer also wants to say about this issue, for example, without being a slave to the creator, by creating a video display that is slightly better than that of a competitor, make his display stand out. By defining such possibilities in the framework of the newly deployed HDR coding and handling techniques, at least, rather than the display manufacturer performing any processing that completely ignores what the image was intended for and as a result no one has a good idea anymore about how the image will finally look on various displays, it is guaranteed that the video creator will have more technical possibilities to have the final say on how the image will look.

[0025] The steps involved in creating and encoding an HDR image are summarized as follows. a) One must define the desired pixel absolute luminance of (the master HDR image, i.e., the first image that grades the HDR scene), and for this person, define the dynamic range of the target display that ends with PL_V (this can be called the grading of the image) b) Since typical video coding uses YCbCr pixel color coding, one needs to encode these luminances, for example, as 10-bit luma, which requires selecting one of various possible HDR electro-optical transfer functions (EOTFs), and this selection is typically also transmitted simultaneously as metadata of the target display.

[0026] This constitutes the basic creation of the HDR image itself. It can be reconstructed into the decrypted HDR image on the receiving side, for example, on a TV display.

[0027] However, in the expert future-proofing system, the video creator will add additional metadata.

[0028] Typically, this requires at least defining a luma mapping function (component c in addition to a and b above), and applying a secondary reference grading, which is typically defined as a luma mapping function and the two are mathematically related to whether a person knows the EOTF that defines the master HDR and the EOTF that defines the luma of the secondary image, to specify, for example, how it can be obtained from a master HDR image of PL_V of 500 nits.

[0029] Advantageously, a person desires to make the second image an SDR image. Then, not only does the person ensure knowing which relative position among the relative positions between which the luminance of all image subjects should have (with respect to the most extreme grading that a person typically encounters, i.e., actually needs) to be made into other intermediate dynamic range gradings, but also, when this SDR image is used as the image to be transmitted as a representative of a pair of gradings, it can be used as a reverse compatibility system where an old display can directly display this received SDR image without the need for knowledge of HDR video coding or HDR video handling. This SDR image can be defined to have PL_V_SDR = 100 nits, but traditionally, it will be interpreted as a dimensionless image where 100% is the white of SDR.

[0030] To explain the components of the method, the reader can focus on a few image subjects of some typical images in Figure 1.

[0031] The left luminance axis is the luminance of the master HDR image, i.e., the "best image" that the creator has created and ideally wants to be displayed.

[0032] For example, in the outdoor western scene ImSCN1, the creator wants to fully utilize the increased dynamic range to render subjects that are burning under the sun. He takes great care to ensure that it is sufficient for those subjects at 500 nits (e.g., the white hat of the cowboy). (Subsequently, the viewer will encounter a scene burned by sunlight as it will appear five times brighter than a normal LDR rendering that could be the previous scene of that movie if it were done indoors, for example). Of course, on the right SDR luminance range with a secondary luminance dynamic range DR_2 that is much smaller than the primary dynamic range DR_1, the luminance of the cowboy hat has to be mapped at a convenient position, for example 18 nits, between 0.1 nit and 100 nits. The set of all these projection arrows defines a luminance mapping function for defining the SDR image based on the HDR image (or, i.e., without including the clip of the same SDR luminance of the HDR luminance and if the function is invertible, SDR to HDR). The same can be done for the night scene ImSCN2 or for the cave with an opening through which a person can see the outside world ImSCN3 illuminated by sunlight.

[0033] Now, if a person wants to create an intermediate dynamic range image (MDR) while, for example, being ready to drive an 800 nit PL_D display with an 800 nit PL_V_MDR, the problem is where various subject luminances should end up. This can be a complex problem as it depends on which subject is in which type of scene, and a person may want to place any particular subject like a cowboy at various possible positions according to a video creator (see arrow FL_DA), because on such a dynamic range, a person not only has to consider the rearrangement of various subject luminances to vary, especially according to the extent of the available third-order dynamic range DR_3, and how much it deviates from the first-order DR_1, but also has to consider regarding the type of image content (there are different needs for a plain type of subject like a cowboy compared to, for example, an explosion or clouds in the sky).

[0034] However, it can be assumed that a person can actually use a theoretically imperfect system and a pre-designed method symbolically shown by simply connecting three positions by a continuous line that links the three positions by some fixed algorithms. However, the display application from the first-order HDR image (also called the master image) to the second-order grading only requires the application of a luma mapping function, and the luma mapping function is specified for this re-grading by a color grading section or an automatic grading algorithm (which will define an average excellent mapping based on the analysis of various characteristics of the master HDR image, such as its luma histogram), and use the display application to obtain a third-order (MDR) image based on this luma mapping function and some pre-deployed display application algorithms, and the display application algorithm will convert the initial, for example, HDR-to-SDR luma mapping function to an HDR-to-MDR luma mapping function (see the following explanation together with Figures 4 and 5).

[0035] Regarding component b, simple HDR encoders and HDR10 encoders have been introduced to the market. This HDR10 video encoder uses the so-called perceptual quantizer (PQ) function standardized in SMPTE2084 as the OETF (inverse EOTF). Instead of being limited to 1000:1 like the Rec.709 OETF, this PQ OETF enables the definition of (as much as possible to be displayed) much more luminance, that is, it is possible to define luma (typically 10 bits) for values between 1 / 10,000 nit and 10,000 nit, which is sufficient for any actual HDR video generation.

[0036] Readers should be cautioned not to confuse HDR too simplistically with a large number of bits in the luma coding word. This is true for linear systems such as the amount of bits in an analog-to-digital converter. In fact, the amount of bits follows the dynamic range as the base-2 logarithm. However, since the code assignment function can have a rather non-linear shape theoretically even if desired, an HDR image can be defined using only 10 bits of luma (and even 8 bits of HDR image per color component), which has led to the advantage of reusability of already deployed systems (for example, an IC having a certain bit depth or video cable, etc.).

[0037] After the calculation of luma, one has a 10-bit pixel luma plane Y_code, to which two reference color components Cb and Cr are added per pixel as reference color pixel planes. This image can be handled "as if" it were, traditionally and further on, technically an SDR image, for example, a compressed MPEG-HEVC, etc. The video compressor actually does not need to consider the color or luminance of the pixels. However, since the colorimetry of the HDR PQ-defined image will not result in the appropriate colors as defined under the Rec.709 colorimetry, additional luminance or color tone mapping may be in place.

[0038] Thus, to accept a device, e.g., a display (or actually its decoder), it is typically necessary to perform an appropriate color interpretation of the {Y, Cb, Cr} pixel colors in order to display an image that looks appropriate, rather than an image with, e.g., bleached colors.

[0039] This is typically handled by simultaneously transmitting additional image - defining metadata in addition to the pixel color component matrix of the video image. The image - defining metadata defines image coding such as an indication of which EOTF is used (assuming for now without limitation that the PQ EOTF (or OETF) is used), and the value of the maximum luminance of the target display associated with, e.g., the PL_V of the HDR image.

[0040] More sophisticated future - proof HDR encoders include functions that specify how to map (or, in other words, re - grade) additional image - defining metadata, e.g., handling metadata such as the normalized luminance of the first image up to PL_V = 1000 nit, to the normalized luminance of a secondary reference image, e.g., an SDR reference image with PL_V = 100 nit (as described in more detail in Figure 2). These mapping functions vary from image to image because, for example, a cave scene should require a different luminance remapping to SDR luminance than a sun - bleached desert scene. That is, one associates, in the relevant metadata, a separate global luminance mapping function F_L (or something equivalently embodied as a luma mapping function via the EOTFs of the input and output luma) with each successive image of the video.

[0041] Figure 2 shows a full HDR video communication system. It uses, for example, without intending to be limiting, the HDR video image coding described in the standard ETSI TS103433 (High Performance Single - Layer High Dynamic Range System for Use in Consumer Electronic Equipment (SL - HDR)). On the transmitting side, there is an image source 201. Depending on whether it has, for example, a video created offline from an Internet delivery company or a real - life broadcast, this can range from a hard disk to a cable output from, for example, a television studio.

[0042] According to some video generation rules, for example, a master HDR video (MAST_HDR) is provided, which is a version where color grading is done by a human color grader, or shading is done by camera capture, or by an automatic brightness redistribution algorithm, etc.

[0043] This can be, for example, a 5000 - nit PL_V video shown on the left side of Figure 1, for example, a western movie with 5000 - nit sunlight brightness.

[0044] From this master HDR image, a communication image for actually transmitting the HDR image is derived. This can be done in various ways. For example, the 5000 - nit HDR image itself can be transmitted by calculating a 10 - bit luma code using the PQ EOTF. Another way to transmit the additional grading for the master HDR's grading as a proxy will be described. For example, an SDR communication image Im_SDR, which is an SDR re - grading version of the master HDR image, is transmitted. When a strictly increasing luminance or a luma mapping function F_L is used, since the luma mapping function F_L is invertible, at any receiver, the HDR master image can be reconstructed by applying the inverse function to the Im_SDR pixel luma. Generally, considering that a person may further desire, for example, to perform chroma changes, there will be a color mapping function F_ct that can be optimized for each image. The color processing of the mapping is done by the color converter 202.

[0045] The communication image essentially looks just like any other 10-bit image. In particular, since the SDR communication image Im_SDR is a true SDR image (in the sense that when it is directly displayed on an old-fashioned LDR display without further color optimization, it will give a picture that looks good enough), this image can be compressed by the video compressor 203. The video compressor 203 applies, depending on which compression format is deployed on any communication medium (whether it is an optical disc, Internet access, terrestrial digital TV, or some professional video communication system for, e.g., surveillance or security), for example, an MPEG algorithm such as HEVC or VVC, or some other video coder such as AVI, to obtain the compressed HDR image Im_COD. The formatter 204 applies formatting typical for communication, such as packetization, frequency modulation, (handshaking), etc. The mapping function, which can be defined in various forms such as parameters that set the form of a function or the exact specification of a function such as an LUT, etc., can be easily transmitted because there are existing mechanisms for transmitting various forms of proprietary data. In the existing mechanisms, one simply has to agree to some video communication standard that includes a luma mapping function defined in a particular way by a specific, e.g., supplementary enhancement information (SEI) message (SEI(F_ct)), and each compliant receiver knows how to understand that SEI message.

[0046] Thus, the reader sees something similar to the explanation using FIG. 1 again. The two components of future-proof HDR coding are: the first technical circuit or process includes the colorimetric specification of the desired image (the master image, more precisely the communication image), and the second circuit or process includes, in particular, the coding scheme itself to generate an image that at least officially (not necessarily colorimetrically) conforms to existing communication technologies. As a result, one can send an HDR image over an existing communication channel with specific defined coding techniques, conditions such as available bandwidth, etc.

[0047] At any receiving site, the corresponding video signal formatter 206 applies the necessary formatting method to re-acquire the compressed video as, for example, a set such as compressed HEVC images (i.e., HEVC image data), such as modulation.

[0048] The video decompressor 207 performs, for example, HEVC decompression to obtain a stream of pixelated uncompressed images Im_USDR, which in this example is an SDR image but could be an HDR image in another mode. The video data input 227 is any medium suitable for transmitting video pixel data and metadata. For example, it could be an external cable such as HDMI (registered trademark), an internal data bus within the device, a connection to a network, etc. The video decompressor will further decompress, for example, the necessary luminance mapping function F_L or general color conversion function F_ct from the SEI message.

[0049] The image and function are input to a (decoder) color converter 208 configured to convert the SDR image to an image of any non-SDR dynamic range (i.e., higher than 100 nits, typically at least several times, e.g., 5 times higher than PL_V).

[0050] For example, the receiver decoder can perform a simple reconstruction of the master's 5000 nit image by applying the inverse color conversion IF_ct of the color conversion F_ct used on the encoding side to calculate the reconstructed HDR image Im_RHDR from Im_USDR in order to create Im_LDR from MAST_HDR.

[0051] Next, this image is sent, for example, to display 210 for further display application, and a 700 nit PL_V image can be obtained to optimally drive the connected 700 nit PL_D end-user display. However, creating the display adaptation image Im_DA_MDR can also be done during decoding at once in the color converter by using the calculated (e.g., identified in the firmware in an offline loop) FL_DA function instead of the F_L function. In such an embodiment, the color converter further has a display adaptation unit 209 to derive the FL_DA function based on the received F_L function using a pre-baked (or configured) display adaptation algorithm.

[0052] The optimized, for example, 700 nit display adaptation image Im_DA_MDR is sent, for example, to display 210 if the video decoder 220 is included in, for example, a set-top box or a computer, to the display panel if the decoder is in, for example, a mobile phone, or transmitted to a cinema projector if the decoder is in, for example, some Internet-connected server.

[0053] FIG. 3 shows a more detailed and useful variant of the internal processing of a HDR decoder (or, mostly typically having the same topology but using inverse functions and not including display adaptation as circuit 350 typically shown by a dotted line in FIG. 3), i.e., the color converter 300 corresponding to FIGS. 2, 208, for further explanation.

[0054] The luminance of the pixel, in this example the SDR image pixel, is input as the corresponding luma Y’SDR (further figures will be used to explain alternative inverse re-gradings for reducing the dynamic range). The chroma components Cb and Cr, also called reference colors, are input into the lower processing path of the color converter 300.

[0055] The luma Y’SDR is mapped by the luminance mapping circuit 310 to the required output luminance L’HDR, for example, to the master HDR reconstructed luminance or some other HDR image luminance. It is obtained from the display adaptation function calculator 350 that uses the metadata simultaneous communication reference luminance mapping function F_L(t) as an input, and for a specific image and, for example, the maximum display luminance PL_D, a suitable function, for example, the display adaptation luminance mapping function FL_DA(t) is applied. In an embodiment, by itself, it consists of several consecutive sub-processes of an upper processing track (circuit 310. However, for the purpose of explaining the current technical contribution to the art, circuit 310 can be assumed to apply a single luma mapping function to all pixels of the currently converted video image.

[0056] The display adaptation function calculator 350 also determines a function suitable for processing the reference colors. For now, it will simply be assumed that a set of multiplication factors mC[Y] for each possible input image pixel luma Y (i.e., for example, 0 to 1023) is stored, for example, in the color LUT 301. The treatment of the exact nature of the color is diverse. For example, one desires to keep the pixel chroma constant by first normalizing the reference color by the input luma (the corresponding hyperbola in the color LUT) and then correcting it for the output luma, although any differential chroma processing can also be used. The color tone will typically be maintained since both reference colors are multiplied by the same multiplier.

[0057] When indexing the color LUT 301 with the luma value Y of the currently color-converted (luminance-mapped) pixel, the required multiplication factor mC is provided as the LUT output. The multiplication factor mC is used by the multiplier 302, which multiplies it by two reference color values of the current pixel, i.e., a color-converted output reference color is provided. Cbo = mC * Cb Cro = mC * Cr [Equation 1] Via a fixed color matrix processor 303, standard colorimetry calculations can be applied, and the reference colors can be converted to lightness-loss normalized non-linear R’G’B’ coordinates R’ / L’, G’ / L’, and B’ / L’.

[0058] The R’G’B’ coordinates that give appropriate luminance to the output image are obtained by a multiplier 311, which R’_HDR = (R’ / L’) * L’_HDR, G’_HDR = (G’ / L’) * L’_HDR, B’_HDR = (B’ / L’) * L’_HDR, [Equation 2] are calculated, and they can be aggregated into the color triplet R’G’B’_HDR.

[0059] Finally, there is further mapping by a display mapping circuit 320 to the format required for the display. This results in the display drive color D_C, whereby not only the colorimetry formulation desired by the display (e.g., also in the HLG OEFT format) is performed, but also this display mapping circuit 320 is configured in some variations to perform some specific color processing for that display. That is, it remaps, for example, some of the pixel luminances further.

[0060] Some examples that explain some suitable display adaptation algorithms for deriving the corresponding FL_DA function for any possible F_L function determined by the authoring side grading unit are taught in WO2016 / 091406 or ETSI TS103433-2V1.1.1 (2018-01).

[0061] The display adaptation method is summarized by showing its movement on a plot of the possible normalized input luminance Ln_in versus the normalized output luminance Ln_out in Figure 4 (these will be converted to the actual luminance by multiplication by the maximum luminance of the display associated with the normalized luminance, i.e., the PL_V value). Note that such a plot can be created as a luma-luma plot, and it is advantageous for the luma to be perceptually uniform. For example, the following equations in this application example are used to derive such a perceptually uniform luma. Y_P = v(L_in; PL_V_in) = log[1 + (RHO - 1) * power(L_in; p)] / log[RHO], where RHO is a constant that depends on the input maximum luminance PL_V_in according to the following equation. RHO(PL_V_in) = 1 + 32 * power((PL_V_in / 10,000); p), and p is a constant typically set to be equal to 2.4. [Equation 3]

[0062] When converting the horizontal axis to perceptual luma, the normalized luminance in Figure 4 is put into the input luminance L_in, and PL_V_in becomes the PL_V for those luminances when converting a 5000 nit PL_V master HDR image to a secondary reference grading. When the secondary reference grading of the perceptually uniform luma between 0 and 1 required on the vertical axis is a 100 nit SDR image, the value PL_V_in = 100 will be used when using Equation 3 to calculate the normalized output luma. The advantage of such an equation is that both axes are already properly normalized so that the eyes want to see different shade values for each respective range appropriately, thereby making the mapping function visually more appropriately defined.

[0063] For example, such a display adaptation algorithm to be applied to the luma of an input image to obtain the final display - compliant luma of a 700 nit MDR image, i.e., an algorithm for calculating the final luma mapping typically based on an input luma mapping from a video creator, typically has the following technical characteristics.

[0064] The input luma mapping function \(F_L\) can be compared to an identification information transformation. This identification information transformation corresponds to mapping the input image onto itself, i.e., on the horizontal input axis, for example, it will have luma that is perceptually uniform for a 5000 nit span of luminance (i.e., \(PL_V_{in}=5000\)), and on the vertical axis, it will have a set of perceptually scaled and divided luma that is exactly the same. Moreover, the identification information transformation is intended such that neither brightening nor darkening occurs for the pixels. Thus, for a particular pixel, an input value of, for example, 0.6 on the horizontal axis will perform an exact mapping to the same value 0.6 on the vertical axis. Thus, this identification information transformation will geometrically come to the position of a 45 - degree slant line in the perceptually uniform luma plot. It stands to reason that the function is scaled by mapping to some intermediate gray - grading of intermediate \(PL_V\). Thus, the 5000 - to - 700 nit display - compliant luma mapping function (\(FL_{DA}\)) typically has the same shape as the 5000 - to - 100 nit input function (\(F_L\)), i.e., the same humps, but is placed closer to the slant line.

[0065] This "opening" of the "fan of functions" is formally formulated as follows. a) One defines a direction to locate the location of an intermediate point of any display - compliant luma mapping function to be calculated, i.e., a secondary gray - grading that forms the end - point of a line segment. In this example, for the SDR gray - grading, any point (\(pos\)) on the slant line corresponding to the input image luma on the function \(F_L\) will come on the line segment along that direction. b) A person defines a metric to place any PL_V value of an intermediate image to be calculated (i.e., a display - compliant image for a 700 nit display, having, for example, a PL_V equal to 700 nit) on that line segment, in the middle between the end - points corresponding to the PL_V values of two start - reference gradings, one of which is the transmitted and received image, which in this example is a master HDR image with a PL_V of 5000 nit corresponding to the hatched identification information mapping, and the other is a secondary grading that can be calculated by applying a luma mapping function received as metadata to the luma of a primary / master HDR image having a value of PL_V = 100 nit. According to some embodiments, the direction is 135 degrees clockwise from the horizontal axis of the input luma, and in other embodiments, it is 90 degrees counter - clockwise. There is a display - compliant algorithm that uses the first direction to calculate the first partial luma mapping function and the second direction for the second partial function.

[0066] For example, a video creator designs a luminance mapping strategy between two reference gradings as described with reference to FIG. 1. Thus, for any possible, normalized luminance Ln_in of a pixel in an input image, e.g., a master HDR image, this normalized input luminance must be mapped to the normalized output luminance Ln_out of a second reference grading that is the output image. This re - grading of all luminances corresponds to some function F_L, which can have many different shapes determined by a human grader or grading automation, and the shape of this function is transmitted simultaneously with dynamic metadata.

[0067] The problem here is, for example, in terms of a metric that can calculate that a display of, for example, 800 nits should have 50% of the grading effect (a full 100% is the re-grading of the 100 nit PL_V SDR image of the master HDR image), which shape should the second-order version of the derived F_L function, which is the display adaptation mapping function FL_DA, have in this simple display adaptation protocol for mapping the input image luma to the MDR image luma. Generally, through the metric, any point between no re-grading at all and full re-grading to a second reference image can be determined, which is for any possible normalized input luminance (Ln_in_pix) of a pixel. The resulting MDR luminance (or luma) is shown as the display adaptation luminance L_P_n, and its location depends, of course, not only on the input normalized luminance but also on the value of the maximum luminance associated with the display adaptation output image (PL_V_out) to be calculated. One skilled in the art understands that one can represent a function of the normalized luminance representation, and one can equally represent a function of any normalized luma representation defined by any OETF.

[0068] The metric is preferably, typically, essentially logarithmic, which is intended such that values between, for example, 100 nit and 1000 nit are first transformed by a non-linear function before assigning them to equally spaced positions, rather than spreading evenly on a line segment every 100 nit. For example, a 950 nit display is close enough to the PL_D = 1000 nit value so that it can be determined that the PL_V = 1000 nit input image requires little to no luminance mapping, while a 150 nit image requires relatively more changes in the mapping compared to a 1000 to 100 nit mapping. Typically, the creator of the coding ecosystem, or at least the creator of a receiving device such as a television display, will choose a fixed value for all options. For example, he will use an orientation of 90 degrees according to the metric pos = 1 - {log(1 + [(PL_V - 100)]) / log(1 + [(1000 - 100)])} [Equation 4].

[0069] The corresponding display - compliant luminance mapping FL_DA can be determined as follows (see Figure 4). Take any one of all the input luminances, for example Ln_in_pix. This corresponds to the starting position on the slanted line that has equal angles with respect to the input axis and the output axis of the normalized luminance (shown like a square). For each position along the slanted line, place the scaled version of the metric (scaled metric SM) at each point on the slanted line. As a result, it starts with a slanted line that is orthogonal to the slanted line (or 135 degrees counterclockwise from the input axis) and ends at its 100% level, or the point on the F_L curve at the normalized pos = 1, that is, at the intersection with the orthogonal scaling metric SM of the F_L curve (shown as a pentagon). At the 50% level of the metric (for this PL_D value of the display for which the image has to be calculated in this example), that is, in the middle, or generally, by obtaining the numerical value of Equation 4 or any similar non - linear equation for defining the severity required for re - grading for any deviation between PL_D and PL_V_in, place a point at any normalized position obtained. By doing this for all points on the slanted line corresponding to all Ln_in values, the FL_DA curve is obtained, which is shaped similarly to the original one, that is, it is re - graded in the same way, that is, the maximum luminance is re - scaled / adjusted in a preferably lower - order manner. Now, this function is ready to be applied to calculate the required corresponding optimally re - graded / display - compliant 700 nit PL_V pixel luminance given any input HDR luminance value Ln_in. This function FL_DA will be applied by the luminance mapping circuit 310 immediately before starting the colorimetric quantitative conversion on all the active pixels of the currently processed video image after receiving this display - compliant luma mapping function.

[0070] Figure 5 generally and formally shows the technical components of display compliance.

[0071] For example, the display adaptation circuit 510 in a television or a set-top box includes a setting processor 511. The setting processor 511 sets values for image processing before the pixel colors of the image during its operation are to be processed. For example, the maximum luminance value PL_V_out of the display optimization output image is once set in the set-top box by polling it from the connected display (i.e., the display transmits its maximum displayable luminance PL_D to the set-top box), or when the circuit is within the television, this is set by the manufacturer or the like.

[0072] The luminance mapping function F_L varies for each received image in some embodiments (in other variations, it is fixed for a number of images) and is input from some source 512 of metadata information (for example, this is broadcast as an SEI message and read from a sector of memory such as a Blu-ray disc). This data establishes the normalized height of the normalized metrics (Sm1, Sm2, etc.), and the desired position of the PL_D value thereon can be found from the mathematical equation of the metric.

[0073] When the input image 513 is input, the consecutive pixel luminances (for example, Ln_in_pix_33 and Ln_in_pix_34, or luma) operate through a color processing pipeline to which display adaptation is applied, resulting in corresponding output luminances such as Ln_out_pix_33. Summary of the Invention Problems to be Solved by the Invention

[0074] There are advantages in terms of technical simplification in establishing such a well-functioning fixed function shift algorithm in advance as the above display adaptation algorithm (i.e., for example, when there are no control parameters for changing the metric equation for different positions along the slant line, only a small amount of human grading input is required). However, by establishing stable re-grading, the grading unit can predict what will be displayed on various end-user displays owned by various consumers as well. In fact, since the display adaptation algorithm functions based on the input F_L function in such a way as to maintain its specific shape including the needs of brightness re-grading of various subjects in the image, it is guaranteed that what exactly prevents all the luminances of the master HDR image from being fully displayed on the display will be clearly revealed on the display due to the graded state of his master image. Of course, the disadvantage of the fixed algorithm is that, due to any specific preference, the algorithm does not always produce a 100% perfect image, so people wish to improve it.

[0075] Therefore, with this deployed technology, a video creator can define videos of HDR images, and further specify how these video images should be automatically re-graded for various displays with different dynamic ranges. And he can adapt the shape of the luma mapping function to the needs of any specific image, that is, the input distribution outputs luma along the dynamic range of luminance for all image subjects respectively. However, according to the inventor, the problem with this approach is that it is still somewhat crude, so potentially, even if it deviates slightly from the graded state of the creator, especially in or on this framework, it is advantageous for display manufacturers and / or end viewers to perform some additional amount of re-grading.

[0076] US20140368531 teaches an algorithm for enhancing the transmissivity of an LCD panel when a relatively dark image is to be shown, so that, otherwise, the backlighting can be dimmed to save energy that could be consumed by closed pixels. Instead of doing this via voltage control, according to what US20140368531 teaches, this can be done when an image processing operation changes the pixel luma code to a higher value, and the gamma 2.2 of the standard LDR display power function is used. The enhanced transmissivity function can be defined as desired as to which mainly enhances the darkest colors to be brighter, and the brighter colors remain largely unchanged. For example, for the darkest colors, a boost of up to 4 times can be applied (with the slope of the curve being zero). Such a strong boost is used for dark images, but not so much for bright images. The boost can be quickly realized by interpolating between the maximum boost and (for bright images where there would be no enhancement, i.e., the image remains as it is) the minimum boost. That is, thus, quickly, the (intermediate) output required for each pixel luma code, for example 166, is interpolated as the interpolation of two corresponding output values. The amount of interpolation, for example 25%, depends on the analysis of the input image histogram (e.g., how many pixels are within the darkest bin out of 8 bins).

Means for Solving the Problem

[0077] An image that looks visually better can be obtained by a method of processing an input image (Im_Comm) of an input video to obtain an output image of an output video, The input image has pixels having an input luminance (Ln_in_pix) that falls within a first luminance dynamic range (DR_1), and the first luminance dynamic range has a first maximum luminance (PL_V_HDR), The output image can be calculated from the input luminance and has pixels having an output luminance that falls within a second luminance dynamic range, and the second luminance dynamic range has a second maximum luminance (PL_V_MDR) The reference luminance mapping function (F_L) is received as metadata associated with the input image, The reference luminance mapping function identifies the relationship between the luminance of the first reference image and the luminance of the second reference image, The first reference image is the input image, The second reference image has a second reference maximum luminance (PL_V_SDR), Processing includes determining a luminance mapping function (FL_DA) adapted based on the reference luminance mapping function (F_L), Determining uses a pre-fixed display adaptation algorithm that, in a coordinate system of input luminance normalized to a maximum of 1 and output luminance normalized to a maximum of 1, for each point on the slant line, identifies each metric along a line segment starting on the slant line and oriented in a pre-fixed direction, Each respective metric at each position is normalized by giving the value 1 to the intersection point between the line segment and the locus of the luminance function (F_L) in the coordinate system, Determining identifies the position on each respective metric corresponding to the second maximum luminance (PL_V_MDR), The set of positions on each respective metric is output as the adapted luminance mapping function (FL_DA), An image that looks visually better includes processing that includes calculating a booster intensity value (BO), and calculating the booster intensity value (BO) includes - determining a histogram (hist) of intermediate luminance obtained by applying the adapted luminance mapping function (FL_DA) to the input luminance, - calculating an average lightness magnitude (AB) based on the histogram, - calculating a first intensity value (PosB) from the average lightness magnitude (AB), - calculating a weighted sum (BDE) of the pixel counts within at least two configurable upper bins of the histogram, - calculating a second intensity value (NegB) from the weighted sum (BDE); - determining a booster intensity value (BO) based on a value (DV) equal to a first intensity value (PosB) minus the second intensity value (NegB); further processing includes calculating an adjusted and adapted luminance mapping function (F_ALT_B) by using a display adaptation algorithm configured to output a locus of positions equal to the booster intensity value on a metric, as the adjusted and adapted luminance mapping function (F_ALT_B); characterized by applying the adjusted and adapted luminance mapping function (F_ALT_B) to an input luminance to obtain an output luminance.

[0078] Note that when we consider luminance, we intend to represent the technical concept of luminance. This is not intended to be read by the reader as being limited to any particular representation, especially normalized physical luminance, but rather typically intends that luminance will be represented in a luma representation. For example, embodiments that function with a luma system that is psychovisually uniform according to Equation 3 will function adequately, but other representations will also function if various elements (e.g., metric scale) are appropriately set.

[0079] When the metric equation is defined, any PL_V_MDR value will, by calculating that equation, be located at some location of the metric between 0 and 1, which depends on which input image comes in and whether upscaling or downscaling is involved, typically values PL_V_HDR (e.g., 5000 nit) and PL_V_SDR (100 nit). While one function F_L is used for illustration, several functions can be concatenated and used to transform one image (the interested reader can find useful embodiments in ETSI TS103433 incorporated by reference). The display adaptation algorithm will typically be once fixed for each video communication ecosystem. For example, in terrestrial ATSC broadcasting, the algorithm is selected by the standard. A movie provider from an Internet server decides whether to use the same display adaptation algorithm or a similar but different one. For any PL_V mastering of an HDR image, the second reference image is usually a 100 nit SDR image because a 100 nit SDR image has an excellent redistribution of the image subject luma for typically producing a low-quality display system. The PosB value and the PosN value are an easy way to indicate the needs and risks of increasing the intended re-grading in a competing manner. The reuse of a pre-fixed display adaptation algorithm ensures predictability and a large retention of the artistic intent for the original content-rated part's image. Endpoint entities, such as display manufacturers, can tailor the system to what they desire (e.g., change the intended grading more aggressively, but, for example, enhance some desirable, pictorial aspects from the intensity of the compressive downscaling), for example, several parameters are available for shaping the conversion function.

[0080] In FIG. 6, what can be done for at least some of the images of several videos can be seen schematically (in an exaggerated manner). Assume that an input HDR image with PL_V = 1000 nits is mapped to a standard 100 nit (or perhaps 200 nit) output. Then, the slanted line represents what is introduced as darkening by the inherent linear scaling (e.g., mapping the brightest pixel encoded in the image to the brightest displayable color (displaying white) according to the traditional SDR display paradigm) in the normalized luminance coordinates. All colors are weakened to one-tenth of what they are assumed to be. This is not a problem (or rather, it becomes a rather handy automatic color mapping) for already fairly bright subjects such as clouds illuminated by sunlight or our strongly illuminated cowboy. For example, a pixel originally assumed to be 500 nits will be displayed at 50 nits. This is, on the one hand, a reasonable value in the more limited dynamic range of an SDR display, but on the other hand, it is bright enough even in SDR to be worth seeing. More problematic are relatively darker subjects. In the input range R_i, the grading of the master 1000 nit HDR image has carefully done the grading of darker subjects while also dealing with the presence of contrasting brighter HDR subjects. They are, for example, indoor subjects, and of particular importance are indoor subjects within areas of a scene that are not very well illuminated. In the real world (i.e., when starting measurements using a luminance meter), indoor subjects are typically 100 times darker than outdoor subjects, for example, when seen through a window in image viewing. In the master HDR grading that ends at PL_V = 1000 nits, a ratio of typically 100:1 will not be used. This is because, on the one hand, it is not the best use of the dynamic range, and on the other hand, it is not what a typical viewer is expected to see (not seeing a real scene but a small rectangle during evening viewing in his living room).Thus, for example, in a master grading that gives values between 10 nit and 1000 nit to external pixels seen through a window and values between 0 nit and 100 nit to indoor pixels, a lighting ratio of 10:1 is used.

[0081] It looks quite accurate when actually shown on a 1000 nit HDR display. However, when scaled by 10, pixels that are supposed to be displayed at 10 nit will be displayed at 1 nit, which is deep black. Considering further effects of ambient lighting at the viewing location, such as reflections on the front glass of the display, these darker areas will rarely be worth seeing.

[0082] This is why the grading unit, in combination with a display adaptation algorithm, creates a luminance (or luma) mapping function FL_DA, which will give a fairly good remapping of all pixel luminances, especially for those indoor pixels of the mentioned indoor scene. It is actually found that the mapping curve stretches the input range Ri relative to the output range Ro, and the output range Ro must be further converted to absolute luminance by multiplying by 100. However, in any case, in the example, approximately 50% of the available range is used, thereby ensuring better visibility.

[0083] However, depending on the situation, one may desire to apply an even more powerful brightening re-grading and, in this example, stretch the darkest luminance over an even larger second output range Ro2.

[0084] It is important to build on this possibility on top of existing re-grading frameworks. Otherwise, a TV manufacturer, which is one of the first entities that desires to apply this further brightening, can perform any luminance mapping to make the manufacturer's display brand look better. And this will not only disrupt the attention of the video creator (who creates a particularly impressive HDR scene image by placing all subject luminances at the desired values in the master grading image through all difficulties, and further specifies how to slightly change this distribution when going to a smaller dynamic range), but also revert one of the creation, coding, handling, and display of HDR in a simple and unpredictable manner. In fact, it can be expected that at least some end users will apply a rather powerful luminance remapping function and, according to the creator, potentially make their images pop out impressively in a rather overdone and dull manner. However, this has even worse effects. For the sake of simplicity, what happens to the darkest luminance of the image has been described, but it turns out that all luminances along the range need to be properly placed to have a reliable and impressive HDR effect within the image (see, for example, the steeper slope between two humps). If one strongly deviates from the desired display adaptation curve, one strongly degrades the image and, instead of an impressive HDR image, there is no good HDR effect at all (not even a bit of scaling down to fit the smaller dynamic range capabilities, but simply gone or at least severely distorted), and it gets closer to an SDR image. The output image looks flashy and varies according to preference, but may vary even more appropriately for some customers, and does not look complete, i.e., does not look like it was created by the grading section of the master HDR video.

[0085] Therefore, an alternative display adaptation function F_ALT_B suitable for the different preferences of non-creators, such as display manufacturers, needs to be created in a carefully designed technical manner.

[0086] A method of processing an input image is advantageous when it utilizes the fact that one of two configurable super bins contains 1% of the brightest pixels within a histogram (hist) of intermediate luminance. This super bin can further advantageously be ignored as non-critical by setting its ci multiplier to zero (see below).

[0087] A method of processing an input image is advantageous when it has a boost intensity value (BO) determined as a customizable multiplier value (mf) multiplied by a relative boost value (RB) normalized between zero and one. This is an easy way for an implementer to set the typical intensity of the original grading shift while balancing a general improvement in the beauty of the initial appearance of the ultimately displayed image.

[0088] Processing the input image to retain the changes associated with the determined boost intensity value as limited or reduced compared to the unfiltered direct change is advantageous when it has a calculation of the boost intensity value performed after temporary filtering.

[0089] When an asymmetric temporary filtering is used that filters the negative differences differently from the positive differences, it is advantageous for the image to change rapidly to lower boost intensity values (since the image cannot withstand large boosting, creates errors, and immediately requires more careful treatment) and change more slowly to higher boost intensity values.

[0090] The method can be further embodied in an image processing apparatus for processing an input image (Im_Comm) of an input video to obtain an output image of an output video. The input image has pixels having an input luminance (Ln_in_pix) that falls within a first luminance dynamic range (DR_1), and the first luminance dynamic range has a first maximum luminance (PL_V_HDR). The output image can be calculated from the input luminance. has pixels with an output luminance that falls within a second luminance dynamic range, the second luminance dynamic range having a second maximum luminance (PL_V_MDR), The image processing apparatus includes a video data input (227) configured to receive an input image and a reference luminance mapping function (F_L) encoded as metadata associated with the input image, The reference luminance mapping function specifies a relationship between the luminance of a first reference image and the luminance of a second reference image, The first reference image is the input image, The second reference image has a second reference maximum luminance (PL_V_SDR), The image processing apparatus includes a display adaptation unit (209) configured to determine a luminance mapping function (FL_DA) adapted based on the reference luminance mapping function (F_L), The display adaptation unit uses a pre-fixed display adaptation algorithm that, in a coordinate system of input luminance normalized to a maximum of 1 and output luminance normalized to a maximum of 1, for each point on the diagonal line, specifies each metric along a line segment starting on the diagonal line and oriented in a pre-fixed direction, Each metric at each position is normalized by giving a value of 1 to the intersection point between the line segment and the locus of the luminance function (F_L) in the coordinate system, Determining includes identifying the positions on each metric corresponding to the second maximum luminance (PL_V_MDR), The set of positions on each metric is output as the adapted luminance mapping function (FL_DA), The method includes the image processing apparatus performing processing that includes a boost determination circuit (700) configured to calculate a boost intensity value (BO), the boost determination circuit (700) being, - A histogram calculation unit (721) configured to determine a histogram (hist) of intermediate luminance obtained by applying a suitable luminance mapping function (FL_DA) to the input luminance; - An average calculation circuit (705) configured to calculate an average lightness magnitude (AB) based on the histogram; - A first converter unit (706) configured to calculate a first intensity value (PosB) from the average lightness magnitude (AB); - A second measurement unit (711) configured to calculate a weighted sum (BDE) of pixel counts within at least two configurable upper bins of the histogram; - A second conversion unit (712) configured to calculate a second intensity value (NegB) from the weighted sum (BDE); and - The boost determination circuit determines a boost intensity value (BO) based on a value (DV) equal to the first intensity value (PosB) minus the second intensity value (NegB). The image processing apparatus further is configured to use a display adaptation unit to calculate an adjusted suitable luminance mapping function (F_ALT_B), and the adjusted suitable luminance mapping function is calculated as the adjusted suitable luminance mapping function (F_ALT_B) by using a display adaptation algorithm that outputs a locus of positions equal to the boost intensity value on a metric, and the image processing apparatus includes a color converter (208) configured to apply the adjusted suitable luminance mapping function (F_ALT_B) to the input luminance in order to obtain the output luminance.

[0091] The image processing apparatus includes a temporary filter (740) configured to calculate a less variable boost intensity value within the boost determination circuit (700).

[0092] These and other aspects of the methods and apparatuses according to the present invention will become apparent from the implementations and embodiments described hereinafter with reference to the accompanying drawings, and will be described hereinafter with reference to the implementations and embodiments described hereinafter. The accompanying drawings serve merely as non-limiting specific illustrations that exemplify more general concepts. The dashed lines in the accompanying drawings are used to indicate that a component is optional, and components that are not dashed lines are not necessarily essential. Dashed lines may also be used to indicate elements that are described as being essential but hidden inside an object, or for non-entities such as, for example, the selection of an object / region.

Brief Description of the Drawings

[0093]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0094] FIG. 7 illustrates aspects of a fully functional embodiment of the current approach and the booster decision circuit 700. Complex image analysis is expensive in simpler ICs when performed, for example, in less expensive mobile phones. Further, often the image is very complex, so the initial insights are relatively easy to obtain, but more detailed insights about the image require exponentially more analytical calculations before the knowledge of the image becomes deeper. Note that it is desirable that the calculations be performed on the current image being processed immediately before the final colorimetric quantitative conversion is performed. This is because only a single image time delay (i.e., 1 / 50 of a second, for example) is required at that time.

[0095] Typically, what an IC can do is to determine the histogram of an image. For example, if a histogram of 128 bins is calculated at 1024 luma levels, one has a count of the amount of pixels in the image with an accuracy of approximately 10 luma units. From such a histogram, already much interesting image information can be obtained to create a fully functional algorithm. The received image (Im_Comm) transmitted is assumed in the current description without limitation to be the master HDR itself by the downscaling function F_L created on the video creation side.

[0096] The IC is assumed to perform a first histogram calculation in a first histogram calculation unit 720 for, for example, the unique coding of the input image, i.e., the coding in which the input image enters the decoder together (those skilled in the art can understand how to obtain the histogram in other ways, so this explanation is taught without intentional limitation). Let's assume this is the PQ luma domain histogram. In fact, we need to have a histogram of the display - compliant luma (reference display compliance by FL_DA without current improvements). Because, even if it is not yet completely as desired, it is an image optimized for a specific display, so it will be measured to understand what (semi -)automatic improvement is needed. This histogram (hist) will be determined by a second histogram calculation unit 721. This unit actually does not need to apply display compliance to the entire image to calculate the histogram. It can simply apply the function FL_DA to the histogram (hi) obtained from the first histogram calculation unit 720. Those skilled in the art understand how the histogram in the second domain can be derived from the histogram in the first domain, for example, in the luminance domain or any luma domain defined according to some EOTF. For example, first, the PQ luma can be converted to normalized luminance by applying the perceptual quantizer EOTF EOTF_PQ of SMPTE2084, and then the function FL_DA is applied. Then, the counts will move to different histogram bins. In a preferred embodiment, our luminance representation and thus the histogram his is expressed in a psychovisual - ly uniform luma domain according to the above - mentioned equation 3, using the maximum of the image to be determined as the value of PL_V_in, i.e., as the value of PL_D_MDR of the MDR display for which the image should be display - compliant, for example, 700 nit or 1500 nit, etc.

[0097] The brightening promoter unit 701 calculates the magnitude of the image brightness (AB) (within the average calculation circuit 705) based on some average value of the histogram. There can be various equivalent calculations, but for now, it is assumed to be the average of the histogram in the perceptually uniform luma domain that works well in practice.

[0098] This value is converted to an intensity value, which is ideally normalized between 0 and 1.

[0099] The reader can refer to Figure 6 to see how this functions. As taught, it fades in to the desired state starting from the HDR image input and can be downscaled metrically to any degree required between 0 and 1 (the reader can understand that, for example, an upscaling starting from the received SDR image functions similarly, even though it is in the other direction, such as a concave curve below the diagonal line), and 1 is the highest regrading to obtain the most extreme other reference grading (e.g., SDR grading). Here, further, this metric can be continued to fade in to something greater than what is required. This is not theoretically the optimal grading for a video creator, but for example, the display manufacturer applies a little more of the useful grading desired, and at least, it follows this defined grading. Advanced embodiments can further modify the creator's reference regrading only over a part of the input luminance range. For example, the lower half brightens only itself and harmonizes with the upper half that remains generally the same or has less variation, but we stick to a simpler embodiment for the sake of explanation.

[0100] Therefore, this additional technical option can be mathematically formulated by extending the metric somewhat (not overly, otherwise the grading will be overly distorted. However, the maximum value can be realized by the developer of the coding framework). Let's assume we can progress to the maximum metric position, which is a boosting intensity value (BO) of 2.0.

[0101] It is advantageous to stamp this value, for example, as a customizable setting by the display manufacturer. The appropriate equation should use a multiplier fixed for a value that is preferably between zero and one, a relative boost value (RB) that will be generated from the boost determination circuit 700. Therefore, BO = mf * RB, where 1 < mf <= VAL_Ph, [Equation 5] VAL_PH is, for example, a maximum value determined by the applicant, for example 2.0.

[0102] Therefore, the reader should understand that the metric position needs to be determined here using a metric position of 1.0 on the outside in order to apply a preselected display adaptation algorithm, but this value should be determined appropriately automatically (or semi-automatically in the sense that an adoptee, for example a mobile phone manufacturer, can set some control parameters such as mf) based on the characteristics of the current image.

[0103] Thereby, no unwanted images or artifacts should be introduced for each image or for all of the video.

[0104] To convert the magnitude AB of the average brightness into a positive intensity value, the first converter unit 706 applies a suitable conversion function.

[0105] This function is designed by the developer of the entire HDR handling framework, i.e., for example, the applicant of this patent application, who leaves it possible for a specific display manufacturer to fine-tune according to their preferences with a few control parameters.

[0106] This function aggregates the requirements regarding when a person desires to enhance re - grading or leave it as it is (i.e., F_ALT_B remains at FL_DA), and for which intensity a specific image should be enhanced according to the automatic algorithm.

[0107] Figure 8 shows a useful embodiment of such a first conversion function. MxPU is a typical value that, over a large image set, is predefined and beyond which boosting of display - compliant re - grading (i.e., boosting where FL_DA becomes F_ALT_B) is not performed. For example, if all pixels are within the top 10 histogram bins, the person already has such a bright image, so boosting has little meaning. In fact, in the practical example of Figure 8, the person stops promoting boosting above control point 2 (Cp2). These control points (and also the first control point Cp1) can be dragged by the display manufacturer to adjust the behavior of the fixed booster algorithm. The designer can preset them, for example, to 25% of MxPU for Cp1 and 75% of MxPU for Cp2. Thus, if the average value of the histogram hist is actually low (lower than the first value ABV1), the person will desire to perform maximum boosting. That is, if the display manufacturer sets mf equal to 2, a value RB of 1 is output from the boost decision circuit 700, and a value BO = 2 is used in the display - compliant algorithm to calculate F_ALT_B. If the image is already somewhat brighter, the person will tend to use a smaller down - grading, i.e., the boosting intensity will decrease from above ABV1 to the second value ABV2, or become the exact value of the (positive / promoting) first intensity value (PosB).

[0108] This is why the algorithm (which over - promotes boosting) is not optimal in this way.

[0109] Therefore, the boost decision circuit 700 includes a boost reduction decision circuit 710. This unit protects the rendering of brighter regions.

[0110] It first further calculates the size of the histogram and then calculates it in the second measurement unit 711.

[0111] This unit creates several (configured) super bins, that is, sets of bins of the histogram hist.

[0112] It creates, for example, two bins, the top 1% and top 5% of the pixels.

[0113] Or it groups the top two bins of the calculated histogram hist into the top super bin and, for example, groups the four bins below it into the second super bin, etc.

[0114] For example, the use of four super bins works well.

[0115] Next, a counting equation including appropriately set coefficients is calculated.

[0116] The result DBE is equal to the sum {ci * SB_i} [Equation 6]. SB_i is the count of pixels entering super-bin i among at least two super-bins, and c_i is the corresponding coefficient. The coefficient can be set, for example, to a value such that the brightest pixels only give some sparkle, such as a metallic reflection, to the image, or, for example, when a person desires a more powerful but less colorimetrically accurate rendering for sports content, for example, when a very dark image is expected, to somewhat exactly evaluate as the brightest pixels (thus, although generally the parameter will be set once to give an average clearly visible result, the parameter can be varied for further aspects such as the type of image, the TV channel being viewed, expert vs. consumer content, etc.). Typically, the c_i value for the highest super-bin will be set to zero, measuring the image content within one or more lower super-bins, and then the size of the highest super-bin is set to include the brightest 1% of the pixels. Since there is clipping in the equation, extreme values are clipped to 0 with no reduction effect, or clipped to 1, or more precisely -1, resulting in a strong reduction. Further non-linearity can exist, for example, when DBE = 1, and there will be no boost even in the case of a high promoting PosB, but a more simple subtractive counting is explained.

[0117] This value is converted to a second intensity value (NegB) (minus / reducing) via a conversion function, such as shown in FIG. 9, within the second conversion unit 712. There may be additional control points, a third control point Cp3, a fourth control point Cp4 to adjust different shapes, but the example shows how an increasing DBE value converts to an increasing normalized NegB value. Finally, NegB is subtracted from PosB by the subtractor 730 to obtain the final relative boost value RB to be used. Typically, the algorithm decides to clip negative values to zero, resulting in no boost or a metric position of 1.0.

[0118] Since we are dealing with video processing, it is advantageous to add a temporary filter 740 to filter the relative boost value based on the value of the previous boost value or the difference from one or more previous boost values. At that time, the deviation between the boosts of temporally consecutive video images, which can lead to visual artifacts, will not be so large. This filter will typically have an asymmetric nature, taking into account the required increase or decrease in relative boost, as shown in FIG. 10.

[0119] The following exemplary types of filtering are used for illustration (it is understood by those skilled in the art that equivalent filters may be used).

[0120] The difference from the previous relative boost value RB_tp (or a combination of several weighted previous relative boost values RB_tp) is calculated. DEL_RB = DV - RB_tp; The weight (Blnd) being held is calculated. Blnd = FF(DEL_PB), where FF is the filter function shape. The output relative boost value is RB = (1 - Blnd)*DV + Blnd*RB_tp [Equation 7] is calculated as.

[0121] The components of the algorithms disclosed in this text are realized (fully or partially) in fact as hardware (e.g., part of an application-specific IC), or as software operating on a dedicated digital signal processor or a general-purpose processor, etc.

[0122] It should be understandable from our presentation to those skilled in the art which components are optional improvements, which can be implemented in combination with other components, and how the (optional) steps of the method correspond to the respective device means, and vice versa. The term "device" in this application example is used in its broadest sense, i.e., a group of means enabling the realization of a specific purpose, and thus can be, for example, an IC (a small circuit part thereof), or a dedicated instrument (such as an instrument including a display), or a part of a networked system, etc. "Constituent" is also intended to be used in the broadest sense, and thus it includes, inter alia, a single device, a part of a device, a collection of cooperating devices (parts thereof).

[0123] The explicit meaning of a computer program product should be understood to encompass any physical realization of a set of commands that enables a general-purpose or dedicated processor, after a series of loading steps (including intermediate conversion steps such as translation into an intermediate language and a final processor language), to place commands into the processor and execute any of the characteristic functions of the invention. In particular, a computer program product can be realized, for example, as data on a carrier such as a disk or tape, data present in a memory, data moving via a wired or wireless network connection, or program code on paper. Apart from the program code, the characteristic data required for the program is also embodied as a computer program product.

[0124] Some of the steps required for the operation of the method, such as data input steps and data output steps, already exist within the functionality of the processor instead of being described in a computer program product.

[0125] It should be noted that the above embodiments do not limit the present invention, but rather exemplify it. A person skilled in the art can easily realize the mapping of the presented examples to other areas of the claims, and for the sake of brevity, all those options have not been deeply referred to. Apart from the combinations of the elements of the present invention such as combinations within the scope of the claims, other combinations of elements are possible. Any combination of elements can be realized in a single dedicated element.

[0126] Any reference signs between parentheses in the claims are not intended to limit the scope of the claims. The word "comprising" does not exclude the existence of elements or aspects not listed in the claims. An element in the singular form does not exclude the existence of a plurality of such elements.

Claims

1. A method for processing an input image of an input video to obtain an output image of an output video, comprising: The input image has pixels having an input luminance that falls within a first luminance dynamic range, and the first luminance dynamic range has a first maximum luminance. The output image has pixels having an output luminance that can be calculated from the input luminance and that falls within a second luminance dynamic range, and the second luminance dynamic range has a second maximum luminance. A reference luminance mapping function is received as metadata associated with the input image. The reference luminance mapping function specifies a relationship between the luminance of a first reference image and the luminance of a second reference image. The first reference image is the input image. The second reference image has a second reference maximum luminance. Processing the input image includes determining a luminance mapping function adapted based on the reference luminance mapping function. Determining the luminance mapping function uses a pre-fixed display adaptation algorithm that, in a coordinate system of the input luminance normalized to a maximum of 1 and the output luminance normalized to a maximum of 1, for each point on the diagonal line, specifies a respective metric along a line segment starting on the diagonal line and oriented in a pre-fixed direction. Each respective metric at each position is normalized in the coordinate system by giving a value of 1 to the intersection point between the line segment and the locus of the luminance mapping function. Determining the luminance mapping function includes identifying the positions on each respective metric corresponding to the second maximum luminance. The set of positions on each respective metric is output as the adapted luminance mapping function, in the method. The method includes processing the input image. A step of calculating a booster strength value (BO), wherein the step of calculating the booster strength value comprises: Determining a histogram of intermediate luminance obtained by applying the adapted luminance mapping function to the input luminance; Calculating the magnitude of the average lightness based on the histogram; Calculating a first strength value from the magnitude of the average lightness; Calculating a weighted sum of pixel counts within at least two configurable upper bins of the histogram; Calculating a second strength value from the weighted sum; Determining the booster strength value based on a value equal to the first strength value minus the second strength value; The step of calculating the booster strength value (BO); A step of calculating an adjusted adapted luminance mapping function, wherein The adjusted adapted luminance mapping function is calculated as the adjusted adapted luminance mapping function by using a display adaptation algorithm that outputs a locus at a position equal to the booster strength value on a metric; Applying the adjusted adapted luminance mapping function to the input luminance to obtain the output luminance. A method characterized by including this.

2. The method according to claim 1, wherein one of the two configurable upper bins contains the 1% brightest pixels within the histogram of intermediate luminance.

3. The method according to claim 1 or 2, wherein the booster strength value is determined as a customizable multiplier value multiplied by a relative boost value normalized between zero and one.

4. The method according to claim 1 or 2, wherein in order to retain as it is the change accompanying the limited booster strength value determined so far, the calculation of the booster strength value requires temporary filtering.

5. The method according to claim 4, wherein the temporary filtering is asymmetric, changing rapidly to a smaller booster strength value and changing more slowly to a higher booster strength value.

6. An image processing apparatus for processing an input image of an input video in order to obtain an output image of an output video, The input image has pixels having an input luminance falling within a first luminance dynamic range, and the first luminance dynamic range has a first maximum luminance, The output image has pixels having an output luminance that can be calculated from the input luminance and falling within a second luminance dynamic range, and the second luminance dynamic range has a second maximum luminance, The image processing apparatus includes a video data input that receives the input image and a reference luminance mapping function encoded as metadata associated with the input image, The reference luminance mapping function specifies the relationship between the luminance of a first reference image and the luminance of a second reference image, The first reference image is the input image, The second reference image has a second reference maximum luminance, The image processing apparatus includes a display adaptation unit that determines a luminance mapping function adapted based on the reference luminance mapping function, The display adaptation unit uses a pre-fixed display adaptation algorithm, and the pre-fixed display adaptation algorithm, in a coordinate system of the input luminance normalized to a maximum of 1 and the output luminance normalized to a maximum of 1, for each point on the slant line, specifies each metric along a line segment starting on the slant line and oriented in a pre-fixed direction. Each of the respective metrics at each position is normalized in the coordinate system by assigning a value of 1 to the intersection point between the line segment and the locus of the reference luminance mapping function, Determining the luminance mapping function identifies the position on each respective metric corresponding to the second reference maximum luminance, In the image processing apparatus, the set of positions on each respective metric is output as the adapted luminance mapping function, The image processing apparatus has a processor including a boost determination circuit that calculates a boost intensity value, The boost determination circuit, A histogram calculation unit that determines a histogram of intermediate luminance obtained by applying the adapted luminance mapping function to the input luminance, An average calculation circuit that calculates the magnitude of average lightness based on the histogram, A first converter unit that calculates a first intensity value from the magnitude of the average lightness, A second measurement unit that calculates a weighted sum of pixel counts in at least two configurable upper bins of the histogram, And a second conversion unit configured to calculate a second intensity value from the weighted sum, The boost determination circuit determines the boost intensity value based on a value equal to the first intensity value minus the second intensity value, The image processing apparatus further, Using the display adaptation unit, calculates an adjusted adapted luminance mapping function, and the adjusted adapted luminance mapping function is calculated by using the display adaptation algorithm of the setting that outputs a locus of positions equal to the boost intensity value on the metric as the adjusted adapted luminance mapping function, The image processing apparatus, characterized in that it includes a color converter that applies the adjusted adapted luminance mapping function to the input luminance to obtain the output luminance. Claim 7 The image processing apparatus according to claim 6, wherein the boost determination circuit includes a temporary filter configured to calculate a booster strength value that varies less.

Citation Information

Patent Citations

  • Dynamic contrast enhancement using dithered gamma remapping

    US20140368531A1

  • Methods and apparatuses for processing or defining luminance / color regimes

    US20160307602A1

  • Optimizing high dynamic range images for particular displays

    US20190311694A1

  • Improved HDR image encoding and decoding methods and devices

    WO2014128586A1

  • Methods and apparatuses for encoding an HDR images, and methods and apparatuses for use of such encoded images

    WO2015180854A1