HDR video contrast adaptive optimized for displays

JP7927013B2Active Publication Date: 2026-09-30KONINKLIJKE PHILIPS NV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023568223
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-07
Filing Date
2022-04-28
Publication Date
2026-09-30
Estimated Expiration
2042-04-28

AI Technical Summary

Benefits of technology

【0143】 本発明による方法及び装置のこれら及び他の態様は、以下に記載される実施態様及び実施形態を参照して、及び添付の図面を参照して、明らかになり、説明され、添付の図面は、より一般的な概念を例示する非限定的な特定の例証として単に役立ち、添付の図面において、ダッシュは、構成要素がオプションであることを示すために使用され、ダッシュされていない構成要素は必ずしも不可欠でない。ダッシュはまた、不可欠であると説明されているがオブジェクトの内部に隠されている要素、又は例えばオブジェクト/領域の選択などの無形のものを示すために使用される。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927013000001
    Figure 0007927013000001
  • Figure 0007927013000002
    Figure 0007927013000002
  • Figure 0007927013000003
    Figure 0007927013000003
Patent Text Reader

Abstract

In order to obtain in a practical manner an image that can be viewed better for a variety of potentially quite different viewing environment light levels, the inventor proposes an apparatus 900 for processing an input image having pixels with input luminance that is within a first luminance dynamic range DR_1 having a first maximum luminance PL_V_HDR. The apparatus includes an image input 921 configured to obtain an input image 513, and a data input 920 for receiving a reference luminance mapping function F_L that is metadata associated with the input image specifying how the luminance should be re-graded between two reference images, whereby the output image has an output maximum luminance PL_V_MDR that is different from the first maximum luminance, the apparatus further includes a user value circuit 903 configured to determine and output a user correction value UCBVal as set by a human user of the apparatus, and a maximum luminance determination unit 901 configured to obtain the user correction value UCBVal from the user value circuit 903 and to output an adjusted maximum luminance value PL_V_CO as a result of subtracting the user correction value UCBVal from the output maximum luminance PL_V_MDR, the apparatus further includes a user correction circuit 903 configured to determine and output a user correction value UCBVal from the user value circuit 903 and to output an adjusted maximum luminance value PL_V_CO as a result of subtracting the user correction value UCBVal from the output maximum luminance PL_V_MDR, The display adaptation unit is further configured to: determine an adaptive luminance mapping function FL_DA based on L; and the calculation of the adaptive luminance mapping function FL_DA includes finding a position pos on the metric SM corresponding to the adjusted maximum luminance value PL_V_CO, where a first end point of the metric corresponds to a first maximum luminance PL_V_HDR and a second end point of the metric corresponds to a maximum luminance of a second reference image, where for any normalized input luma Yn_CC0, the first end point of the metric is located at a diagonal point having horizontal and vertical coordinates equal to the normalized input luma, and the second end point is located on the locus of the reference luminance mapping function F_L; and the display adaptation unit is configured to apply the adaptive luminance mapping function FL_DA to the input luminance, or an input luma encoding the input luminance, to obtain an output luminance or output luma, and to output these output lumas or output luminances as pixel colors of the output image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and an apparatus for adapting image pixel luminance of high dynamic range video to provide a desired appearance when displaying HDR video under ambient light at a specific viewing site.

Background Art

[0002] Several years ago, a novel technique for high dynamic range (HDR) video encoding was introduced, particularly by the present applicant (see, for example, WO2017157977).

[0003] Encoding of video generally primarily or exclusively relates to generating or more precisely defining color codes (e.g., luma and two chroma components per pixel) to represent an image. This is different from knowing how to optimally display an HDR image (for example, the simplest method simply uses a highly non-linear opto-electrical transfer function OETF to convert a desired luminance into, for example, a 10-bit luma code, and vice versa, and maps a 10-bit electrical luma code to an optical pixel luminance to be displayed by using an inversely formed electro-optical transfer function EOTF, thereby converting those video pixel luma codes into the luminance to be displayed, but more complex systems deviate in several directions, in particular by decoupling the encoding of an image from the specific use of the encoded image).

[0004] The encoding and handling of HDR video is in complete contrast to how conventional video technology has been used. According to conventional video technology, all video was encoded until recently, which is nowadays called standard dynamic range (SDR) video encoding (also known as low dynamic range video encoding; LDR). This SDR started as PAL or NTSC in the analog era and transitioned to Rec.709-based encoding, such as MPEG2 compression, in the digital video era.

[0005] While it was a satisfactory technology for transmitting moving images in the 20th century, advancements in display technology that surpassed the physical limitations of 20th-century CRT electron beams or global TL backlight LCDs have made it possible to display images with pixels that are significantly brighter (and potentially even darker) than those of older displays, which has led to the need to encode and create such HDR images.

[0006] In fact, it began with the fact that very bright, and sometimes even darker, image objects could not be encoded in the SDR standard (8-bit Rec.709) for various reasons. First, a way was invented to technically represent those wide-ranging colors, and then, one by one, all the rules of video technology had to be rethought, and in many cases reinvented.

[0007] The Lumacode definition for SDR in Rec.709 could only encode a luminance dynamic range of approximately 1000:1 (with 8 or 10-bit Luma) due to the nearly square root OETF function shape of Luma Y_code = power(2,N)*sqrt(L_norm), where N is the number of bits in the Luma channel and L_norm is the normalized version of the physical luminance between 0 and 1.

[0008] Furthermore, in the SDR era, the absolute brightness to be displayed was not defined; therefore, in practice, the maximum relative brightness L_norm_max=100%, or 1, was mapped via the square root OETF to the maximum normalized lumina code Yn=1, for example, corresponding to Y_code_max=255. This has some technical differences compared to creating absolute HDR images, namely, an image pixel encoded to be displayed as 200 nits will ideally (i.e., where possible) be displayed as 200 nits on all displays and not as a completely different display brightness. In the relative paradigm, a 200 nit encoded pixel brightness will be displayed as 300 nits on a brighter display, i.e., a display with a brighter maximum displayable brightness PL_D (aka display maximum brightness), and as 100 nits on a less capable display, for example. Furthermore, it should be noted that absolute encoding works with normalized brightness representation or normalized 3D color gamut, where 1.0 uniquely means, for example, 1000 nits.

[0009] In displays, such relative images are typically displayed somewhat heuristically by mapping the brightest luminance of the video to the brightest displayable pixel luminance (this was done automatically via the electrical drive of the display panel by the maximum luminance Y_code_max without requiring further luminance mapping). Therefore, if you buy a 200-nit PL_D display, white will appear twice as bright as on a 100-nit PL_D display, but considering factors such as eye adaptation, this was considered not very significant, other than making the same SDR video image a slightly more beautiful, brighter, and better-viewable version.

[0010] Conventionally, when we speak of an SDR video image today (in the absolute framework), it generally has a video peak luminance of PL_V = 100 nits (e.g., agreed upon in accordance with the standard), and therefore, in this application, we assume that the highest luminance of an SDR image (or SDR grading) is exactly that value or generalized around that value.

[0011] In this application, grading is intended to mean either an activity or a resulting image in which pixels are given luminance as desired, for example, by a human color grader or an automaton. When viewing an image, for example, when designing an image, there are several image objects, and ideally, we want to give the pixels of those objects luminances that are spread around an average luminance that is optimal for those objects, also taking into account the overall picture and scene. For example, if the available image capability is such that the brightest encodeable pixel of the image is 1000 nits (highest image or video luminance PL_V), one grader might choose luminance values ​​between 800 and 1000 nits for the pixels of an explosion to make the explosion look punchy, while another filmmaker might choose an explosion that is no brighter than 500 nits so as not to interfere too much with the rest of the image at that moment (of course, the technology can handle both situations).

[0012] The maximum brightness of an HDR image or video can vary considerably and is generally communicated with the image data as metadata related to the HDR video or image (typical values ​​are, for example, 1,000 nits, 4,000 nits, or 10,000 nits, and are not limiting; generally, an image is said to have HDR when PL_V is at least 600 nits). If a video creator chooses to define an image as PL_V = 4,000 nits, the video creator can, of course, choose to create a brighter explosion, but relatively speaking, it will not reach 100% of PL_V, but rather only 50% for such a high PL_V definition in the scene.

[0013] An HDR display has a maximum capability, i.e., a maximum displayable pixel brightness of, for example, 600 nits, or 1000 nits, or N times 1000 nits (starting from the lowest HDR display). The maximum (or peak) brightness PL_D of that display is different from the maximum brightness PL_V of the video, and the two should not be confused. Video creators generally cannot make the video optimal for every possible end-user display (i.e., the capabilities of the end-user display are best utilized by the video, and the maximum brightness of the video never exceeds the maximum brightness of the display (ideally), but never falls below it; i.e., some of the video images should have at least some pixels with pixel brightness L_p=PL_V, which in continuous optimization for a particular display will further include PL_V=PL_D).

[0014] The creator makes some of their own decisions (for example, what kind of content to capture and how), and generally sets the PL_V of the video very high to serve the highest PL_D display of the intended audience, at least now, and perhaps in the future, when higher PL_D displays appear.

[0015] Next, a secondary problem arises: how to best display an image with a peak brightness PL_V on a display with a lower (often very low) display peak brightness PL_D, which is called display adaptation. Even in the future, there will still be displays that require dynamic range images with lower dynamic ranges than the created, for example, 2000 nit PL_V image, which will be received via a communication medium. Theoretically, a display will always regrade, or map, the brightness of image pixels so that the brightness of the image pixels is displayable by its own internal heuristics, but if the video creator is careful enough in determining the pixel brightness, the video creator can further indicate how the image should be display adapted to a lower PL_D value, and ideally, it is beneficial for the display to follow to a considerable extent what is technically required.

[0016] The situation is more complex with respect to the darkest displayable pixel brightness BL_D. Part of this is fixed physical characteristics of the display, such as LCD cell light leakage, but even with the best displays, what a viewer can ultimately distinguish as a different darkest black also depends on the lighting of the viewing room, which is not a well-defined value. This lighting can be characterized, for example, as the average illuminance level in lux units, but for video display purposes, it is more precisely characterized as the minimum pixel brightness. This is also generally more strongly related to the human eye than the appearance of bright or medium brightness. This is because when the human eye is seeing many high-brightness pixels, darker pixels, especially their absolute brightness, become less relevant. However, if, for example, we are looking at a generally dark scene image, but it is still masked by ambient light in front of the display screen, we can assume that the eye is not a limiting factor. Assuming that a human can see a noticeable difference of just 2%, there is a darkest driving level (or lumens) b, and above that, the next darkest lumens level (i.e., displaying a brightness level that is X% higher, e.g., 2% higher) can still be seen.

[0017] In the LDR era, there was no concern whatsoever with the darkest pixels. The primary concern was the average brightness, approximately 1 / 4 of the maximum PL_V = 100 nits. When an image was exposed near this value, everything in the scene appeared clean, bright, and colorful, except for clipping of the brighter parts of the scene exceeding 100%. For the darkest parts of a scene, if they were important enough, a captured image was created using a sufficient amount of base lighting in the recording studio or shooting environment. If a part of the scene was not clearly visible, for example, if it was buried in code Y=0, it was considered normal.

[0018] Therefore, unless otherwise specified, we assume that the darkest black is zero, or actually something like 0.1 nit or 0.01 nit. In such a situation, the engineer is more interested in pixels that are brighter than average in the encoded and / or displayed HDR image.

[0019] Regarding encoding, the difference between HDR and SDR is not only a physical difference (more different pixel luminances that can be displayed on displays with greater dynamic range capabilities), but also a technical difference that includes different lumacode assignment functions (using OETF for it, or the reciprocal of EOTF in the absolute method), potentially, further, a further technical HDR concept such as additional dynamically changing metadata (per image or per set of temporally consecutive images) that specifies how to regrade the pixel luminances of various image objects to obtain an image with a secondary dynamic range different from the starting image dynamic range (two luminance ranges generally end with peak luminances that differ by at least 1.5 times).

[0020] The simple HDR codec HDR10 has been introduced to the market and is used, for example, to create the recently released Black Jewel Box HDR Blu-ray. This HDR10 video codec uses a function with a logarithmic shape rather than a square root, namely the so-called Perceptual Quantizer (PQ) function standardized in SMPTE2084, as OETF (Inverse EOTF). Instead of being limited to 1000:1 like Rec.709OETF, this PQ OETF allows for defining luma at a much larger (ideally displayed) brightness range, namely between 1 / 10,000 nit and 10,000 nit, which is sufficient for practical HDR video production.

[0021] Readers should note that HDR should not be simply confused with the large number of bits in the Lumacode language. This applies to linear systems such as the number of bits in an analog-to-digital converter, where, in practice, the number of bits follows the base-2 logarithm of the dynamic range. However, since the code assignment function has a fairly nonlinear shape, theoretically, one could define an HDR image (and even an 8-bit HDR image per color component) with only 10 bits of Luma, as desired, which would offer the advantage of reusing already deployed systems (e.g., an IC, or a video cable, etc., with a specific bit depth).

[0022] After calculating the lumens, we can have a 10-bit plane of pixel lumens Y_code, to which two chrominance components Cb and Cr per pixel are added as a chrominance pixel plane. This image is then processed classically and even more completely, mathematically "as if" it were an SDR image compressed, for example, MPEG-HEVC. The compressor doesn't actually need to care about pixel color or luminance.

[0023] However, the receiving device, for example, a display (or actually its decoder), generally needs to perform a correct color interpretation of {Y,Cb,Cr} pixel colors to display an image that looks correct, rather than an image with faded colors, for example.

[0024] This is typically handled by co-communicating further image-defining metadata along with three pixelation color component planes, which define image coding such as instructions on which EOTF to use (assuming, without limitation, that a PQ EOTF (or OETF) is used), and the PL_V value.

[0025] More sophisticated codecs include additional image definition metadata, such as handling metadata, including functions that specify how to map a normalized version of the luminance of a first image up to PL_V=1000nit to the normalized luminance of a secondary reference image, such as an SDR reference image with PL_V=100nit (as illustrated in more detail in Figure 2).

[0026] For the convenience of readers with less extensive knowledge of HDR, Figure 1 provides a brief overview of some interesting aspects. Figure 1 shows several typical illustrative examples of the many possible HDR scenes that a future HDR system (e.g., connected to a 1000-nit PL_D display) needs to be able to handle correctly. While the actual technical handling of pixel color is done in various ways across different color space definitions, what is needed for regrading is shown as an absolute luminance mapping between luminance axes spanning different dynamic ranges.

[0027] For example, ImSCN1 is a sunny outdoor image from a Western movie, where most of the area is bright. First, it's important to understand that the pixel brightness of any image is generally not the same as the brightness that can actually be measured in the real world.

[0028] Even if there is no further human involvement during the creation of an output HDR image (which functions as a starter image, referred to as master HDR grading or an image), by fine-tuning one parameter, no matter how simple this process is, since a camera has an aperture, it at least always measures relative luminance with an image sensor. Therefore, in the available encoded luminance range of a master HDR image, there is always some step involved at least where the brightest image pixels end up.

[0029] For example, specular reflection of the sun on a sheriff's star badge may be measured to exceed 100,000 nits in the real world, but this cannot be displayed on typical near-future displays, nor is it comfortable for a viewer watching a movie image, for example, in a dimly lit room in the evening. Instead, if a video creator determines that 5000 nits is sufficiently bright for the pixels of the badge, and thus designates this pixel as the brightest pixel in the movie, the video creator will decide to produce a video with PL_V=5000 nits. Although the camera is only a relative pixel luminance measurement device for the raw version of master HDR grading, the camera should also have a sufficiently high native dynamic range (all pixels sufficiently exceed the noise floor) to produce a good image. Pixels in the graded 5000 nit image are generally derived non-linearly from the RAW image captured by the camera; for example, a color grader takes into account aspects such as the typical viewing situation, which is not the same as when one is standing at the actual shooting location, i.e., a hot desert. The best (highest PL_V) image selected to produce this scene ImSCN1, that is, the 5000 nit image in this example, is the master HDR grading. This is the minimum required HDR data to be created and communicated, but it is not the only data communicated in all codecs, or even not the communicated image at all in some codecs.

[0030] When such an encodable high luminance range DR_1, for example the range between 0.001 nit and 5000 nit, is made available, on the premise that viewers similarly have a corresponding high-end PL_D=5000 nit display, content makers can provide viewers with a better experience of bright appearances, and furthermore, naturally also enable provision of a better experience of darker night views when the movie is sufficiently graded throughout, which is an advantage. An excellent HDR movie balances the luminance of various image objects not only in a single image but also over time in the context of the movie story or generally created video material (for example, a well-designed HDR soccer program).

[0031] On the vertical axis at the left end of FIG. 1, several (average) object luminances viewed in 5000 nit PL_V master HDR grading, which are ideally intended for a 5000 nit PL_D display, are shown. For example, in a movie, one creator may wish to show a cowboy lit on a bright day having a pixel luminance of about 500 nit (that is, generally 10 times brighter than LDR, while another creator may desire slightly less HDR punch, for example 300 nit), whereby the creator constructs the best way to display this western movie image, which gives the best possible appearance to end consumers.

[0032] The need for a higher dynamic range of luminance is more easily understood by considering an image that has considerably dark regions such as the dark corner of a cave image ImSCN3 within the same image, and furthermore also has a relatively large area of very bright pixels such as the sunlit outside world visible from the entrance of the cave. This creates a different visual experience from, for example, a night image ImSCN2 in which only street lamps include high-luminance pixel regions.

[0033] The problem is that, at present, many consumers still have LDR displays, and even in the future, there is good reason to create two gradings for a movie instead of encoding a typical single HDR image itself. Therefore, it is necessary to be able to define an SDR image with PL_V_SDR=100nit that best corresponds to a master HDR image. This is a technical requirement, separate from the technical choice regarding encoding itself, and it explicitly states that, for example, if one of the master HDR image and this secondary image is known how to create one from the other (by inverting it), then one can choose to encode and communicate either one of the pair (effectively communicating two images at one price, i.e., communicating only one image of the pixel color component plane per video time).

[0034] In images with such a reduced dynamic range, it is naturally impossible to define a pixel brightness object like a bright sun with a brightness of 5000 nits. The minimum pixel brightness or deepest black is also 0.1 nits high, rather than the more preferable 0.001 nits.

[0035] Therefore, it should be possible to create this corresponding SDR image with a reduced luminance dynamic range DR_2.

[0036] This is done by some automated algorithm on the receiving display, for example, using a fixed luminance mapping function, or perhaps using simple metadata such as a PL_V_HDR value and potentially one or more other luminance values ​​as conditional.

[0037] However, while more complex luminance mapping algorithms may generally be used, this application assumes, without loss of generality, that the mapping is defined by a global luminance mapping function F_L (e.g., one function per image) that defines how, for at least one image, all possible luminances in the first image (i.e., 0.0001 to 5000) should be mapped to the corresponding luminances in the second output image (e.g., 0.1 to 100 nits in an SDR output image). The normalization function is obtained by dividing the luminances along both axes by their respective maximum values. In this context, "global" means that the same function is used for all pixels in an image, regardless of further conditions such as, for example, their position within the image (more general algorithms might use several functions for pixels that can be classified according to some criterion).

[0038] Ideally, how all luminances should be redistributed along the available range of the secondary SDR image should be determined by the video creator. This is because, given the limitations, video creators are well aware of how to sub-optimize for a reduced dynamic range so that the SDR image still looks at least as good as possible, like the intended master HDR image. Readers will understand that actually defining (positioning) such object luminances corresponds to defining the shape of the luminance mapping function F_L, details of which are beyond the scope of this application.

[0039] Ideally, the shape of the function should also change from scene to scene, i.e., between a cave scene in a film and a slightly later sunny Western scene, or generally from image to image. This is called dynamic metadata (F_L(t), where t represents the image time).

[0040] Ideally, content creators would create the optimal image for each situation, i.e., for each potentially served end-user display, such as a PL_D_MDR=800nit display requiring a corresponding PL_V_MDR=800nit image. However, this is generally too much effort for content creators, even in the most expensive offline video production.

[0041] However, the applicant has previously demonstrated that it is sufficient to create only two different dynamic range reference gradings for a scene (generally, at the extreme edges, e.g., 5000 nits is sufficient as the highest required PL_V and 100 nits is generally sufficient as the lowest required PL_V). The reason is that all other gradings can then be automatically derived from those two reference gradings (HDR and SDR) via a display adaptive algorithm (generally fixed, e.g., standardized) applied to the end-user's display receiving the information of the two gradings. Generally, the calculation is performed in any video receiver, e.g., set-top boxes, televisions, computers, movie equipment, etc. Communication channels for HDR images are also any communication technology, e.g., terrestrial or cable broadcasting, physical media such as Blu-ray discs, the internet, communication channels to portable devices, professional inter-site video communication, etc.

[0042] This display adaptation generally applies the luminance mapping function to the pixel luminance of, for example, the master HDR image. However, the display adaptation algorithm needs to determine a different luminance mapping function from F_L_5000to100 (the reference luminance mapping function that connects the luminances of two reference gradings), namely the display adaptation luminance mapping function FL_DA, which is not necessarily trivially related to the original mapping function between the two reference gradings F_L (there are several variations of the display adaptation algorithm). The luminance mapping function between the master luminance defined with a PL_V dynamic range of 5000 nits and the intermediate dynamic range of 800 nits is written in this text as F_L_5000to800.

[0043] The F_L_5000to100 function maps to a location that one would "naively" expect when crossing the 800nit MDR image luminance range, but rather to a slightly higher location (i.e., in such an image, the cowboy must be at least slightly brighter according to the chosen display adaptation algorithm), symbolically indicated by an arrow (for only one of the average object pixel luminances). Thus, while some more complex display adaptation algorithms could place the cowboy at a higher indicated position, one customer would be satisfied with a simpler position that crosses the 800nit PL_V luminance range, where the connection between a 500nit HDR cowboy and an 18nit SDR cowboy is.

[0044] Generally, a display adaptive algorithm calculates the shape of the display adaptive luminance mapping function FL_DA based on the shape of the original luminance mapping function F_L (or reference luminance mapping function, also known as the reference regrading function).

[0045] The explanation based on Figure 1 constitutes the technical requirements of any HDR video encoding and / or processing system, and Figure 2 shows several exemplary technical systems and their components for realizing the requirements (non-limiting) according to the applicant's codec method. It will be understood by those skilled in the art that these components can be embodied in various devices, etc. This example is presented merely as a representative part of various HDR codec frameworks as a whole to understand the background of some of the principles of operation, and will be understood by those skilled in the art to be not intended to particularly limit any of the embodiments of the inventive contributions presented below.

[0046] While possible, the technical communication of two different actual images at different points in time (HDR and SDR grading, each communicated as three color planes) is expensive, particularly in terms of the amount of data required.

[0047] Furthermore, this is not necessary, as it can be decided to communicate only the primary image and function F_L per time as metadata (and one can choose to communicate either the master HDR or SDR image as a representative of both), since it is known that the luminance of all corresponding secondary image pixels is calculated based on the luminance of the primary image and function F_L. The receiver knows the (generally fixed) display adaptation algorithm and, based on this data, determines the FL_DA function at the receiver end (further metadata that controls or guides the display adaptation may be communicated, but is not currently implemented).

[0048] There are two modes for communicating a single image and function F_L for each time point.

[0049] In the first backward-compatible mode, an SDR image is communicated ("SDR communication mode"). The SDR image is displayed directly on a conventional SDR display (without requiring further luminance mapping), but an HDR display needs to apply the F_L or FL_DA function to obtain an HDR image from the SDR image (or vice versa, depending on which variation of the function is communicated, i.e., upgrading or downgrading). Interested readers can find all the details of the standardized exemplary first mode technique of the applicant below.

[0050] ETSI TS 103 433-1 V1.2.1 (2017-08): High-Performance Single Layer High Dynamic Range System for use in Consumer Electronics devices; Part 1: Directly Standard Dynamic Range (SDR) Compatible HDR System (SL-HDR1).

[0051] Another mode communicates the master HDR image itself ("HDR communication mode"), i.e., a 5000-nit image, and a function F_L that allows the calculation of a 100-nit SDR image (or any other lower dynamic range image via display adaptation) from it. The master HDR communication image itself is encoded, for example, by using PQ EOTF.

[0052] Figure 2 further illustrates the entire video communication system. On the transmitting side, it begins with image source 201. Depending on whether it is offline-created video from, for example, an internet distribution company or live broadcast, this can be anything from a hard disk to, for example, a cable output from a television studio.

[0053] This results in a master HDR video (MAST_HDR) that is, for example, a shaded version of the camera capture, color-graded by a human color grader, or produced by an automatic brightness redistribution algorithm.

[0054] In addition to grading the master HDR image, a set of often reversible color conversion functions F_ct is defined. Without losing generalization, we assume this includes at least one luminance mapping function F_L (however, there may be further functions and data specifying, for example, how pixel saturation should change from HDR to SDR grading).

[0055] This luminance mapping function defines the mapping between HDR-based grading and SDR-based grading, as described above (the latter in Figure 2 is the SDR image Im_SDR to be communicated to the receiver; it may or may not be data-compressed via, for example, MPEG or other image compression algorithms).

[0056] The color mapping of the color converter 220 should not be confused with that applied to the raw camera feed to obtain the master HDR video, as it is assumed here that the master HDR video has already been input. This is because this color conversion is for obtaining the image to be communicated, and at the same time, for obtaining what is needed for regrading, as technically formulated in the luminance mapping function F_L.

[0057] In an exemplary SDR communication type (i.e., SDR communication mode), the master HDR image is input to a color converter 202 configured to apply an F_L luminance mapping to the luminances of the master HDR image (MAST_HDR) to obtain all corresponding luminances to be written to the output image Im_SDR. For illustrative purposes, let us assume that the shape of this function is fine-tuned by a human color grader using color grading software for each shot of an image in a similar scene of a film. The applied function F_ct (i.e., at least F_L) is written to metadata that should be communicated (dynamically processed) with the image, to the exemplary MPEG supplemental enhancement information data SEI(F_ct), or to a similar metadata mechanism in other standardized or non-standardized communication methods.

[0058] After correctly redefining the HDR images to be communicated as corresponding SDR images Im_SDR, these images are often compressed using existing image compression techniques (e.g., MPEG HEVC, VVC, or AV1, etc.) (at least for broadcasting to the end user, for example). This is done by a video compressor 203 that forms part of the video encoder 221 (and is also included in various forms of video creation devices or systems).

[0059] The compressed image Im_COD is transmitted to at least one receiver by some image communication medium 205 (e.g., satellite, cable, or internet transmission according to ATSC3.0, DVB, etc.; however, the HDR video signal may also be communicated, for example, by cable between two video processing units).

[0060] Generally, prior to communication, further transformation is performed by the transmit formatter 204, which applies techniques such as packetization, modulation, and transmission protocol control, depending on the system. This generally involves the application of an integrated circuit.

[0061] At the receiving site, the corresponding video signal unformatter 206 applies the necessary unformatting method, such as demodulation, to reacquire the signal as a set of compressed HEVC images (i.e., HEVC image data).

[0062] The video decompressor 207 performs, for example, HEVC decompression to obtain a stream of pixelated decompressed image Im_USDR. The pixelated decompressed image Im_USDR is an SDR image in this example, but an HDR image in other modes. The video decompressor also unpacks the required luminance mapping function F_L, or generally a color conversion function F_ct, from, for example, SEI messages.

[0063] The image and function are input to the (decoder)color converter 208, which is configured to convert the SDR image to an image with a non-SDR dynamic range (i.e., a PL_V higher than 100 nits, generally at least several times higher, e.g., 5 times higher).

[0064] For example, a reconstructed 5000nit HDR image Im_RHDR is reconstructed to be very close to the master HDR image (MAST_HDR) by applying the inverse color conversion IF_ct of the color conversion F_ct used on the encoding side to create Im_LDR from MAST_HDR. This image is then sent to the display 210 for further display adaptation, but the creation of the display-adapted image Im_DA_MDR is also done in one step during decoding by using the FL_DA function (determined in the offline loop, e.g., firmware) instead of the F_L function in the color converter. Therefore, the color converter further includes a display adaptation unit 209 to derive the FL_DA function.

[0065] The optimized, for example, 800nit display-adapted image Im_DA_MDR is sent to, for example, the display 210 if the video decoder 220 is included in, for example, a set-top box or computer, or sent to the display panel if the decoder is located in, for example, a mobile phone, or communicated to a movie theater projector if the decoder is located in, for example, an internet connection server.

[0066] Figure 3 shows a useful variation of the internal processing of the color converter 300 of an HDR decoder (or encoder, which generally has the same topology but uses an inverse function and generally does not involve display adaptation), corresponding to 208 in Figure 2.

[0067] Pixels, in this example, the luminance of SDR image pixels, are input as the corresponding lumens Y'SDR. Chrominance, also known as the chroma components Cb and Cr, are input to the lower processing path of the color converter 300.

[0068] Luma Y'SDR is mapped by the luminance mapping circuit 310 to the required output luminance L'_HDR (e.g., master HDR reconstructed luminance, or some other HDR image luminance). It then applies an appropriate function for a particular image and the maximum display luminance PL_D, such as the display adaptive luminance mapping function FL_DA(t), which is obtained from the display adaptive function calculator 350, which uses a reference luminance mapping function F_L(t) that communicates with metadata as input. The display adaptive function calculator 350 also determines an appropriate function for handling chrominance. For the time being, we simply assume that a set of multiplication coefficients mC[Y] for each possible input image pixel lumen Y is stored, for example, in the color LUT 301. The exact nature of the color processing may vary. For example, one might want to keep the pixel saturation constant by first normalizing the chrominance by the input lumen (the corresponding hyperbola in the color LUT) and then correcting the output lumen, although differential saturation processing may be used similarly. Since both chrominances are multiplied by the same multiplier, hue is generally preserved. When the color LUT 301 is indexed with the lumens value of the pixel Y that is currently being color-converted (luminance-mapped), the required multiplicative coefficient mC is produced as the LUT output. This multiplicative coefficient mC is used by the multiplier 302, which multiplies it by the two chrominance values ​​of the current pixel, resulting in the color-converted output chrominance. Cbo = mC * Cb Cro = mC * Cr

[0069] Through a fixed-color matrix processor 303 that applies standard colorimetric calculations, the chrominance is converted to lightness-deleting normalized nonlinear R'G'B coordinates R' / L', G' / L', and B' / L'.

[0070] The R'G'B' coordinates that give the output image the appropriate brightness are obtained by the multiplier 311, and the multiplier 311 is, R'_HDR=(R' / L')*L'_HDR, G'_HDR=(G' / L')*L'_HDR, B'_HDR=(B' / L')*L'_HDR These are calculated and combined into a color triplet R'G'B'_HDR.

[0071] Finally, the display mapping circuit 320 performs further mapping to the format required by the display. This results in a display driving color D_C, which is formulated to the colorimetric desired by the display (e.g., even the HLG OEFT format). Furthermore, in some variations, the display mapping circuit 320 is configured to perform some specific color processing for the display, namely, further remapping a portion of the pixel luminance.

[0072] Several examples illustrating some appropriate display adaptation algorithms for deriving the corresponding FL_DA function for possible F_L functions determined by the creating grader are taught in WO2016 / 091406 or ETSI TS 103 433-2 V1.1.1(2018-01).

[0073] However, these algorithms do not give much consideration to the smallest displayable black on the end user's display.

[0074] In fact, these algorithms pretend that the minimum brightness BL_D is small enough to be considered zero. Therefore, display adaptation primarily deals with the difference in maximum brightness PL_D across various displays compared to the maximum brightness PL_V of video.

[0075] As seen in the 18th drawing of prior application WO2016 / 091406, an arbitrary input function (in the illustrative example, a simple function formed from two linear segments) is scaled diagonally based on a metric positioned along an angle of 135 degrees from the horizontal axis of input luminance in a plot typically normalized to input / output luminance of 1.0. This is merely one example of all types of display adaptation in the display adaptation algorithm, and it is not stated to limit the applicability of the inventors' novel concept of display adaptation, and it should be understood that, for example, the angle in the metric direction may have other values.

[0076] However, this metric, and its effect on the reshaped F_L function, i.e., the determined FL_DA function, depends only on the maximum brightness PL_V and PL_D of the display that should be supplied with an optimally regraded intermediate dynamic range image. For example, the 5000nit position corresponds to the zero-metric point located diagonally (for any location along the diagonal corresponding to the possible pixel brightness in the input image), and the 100nit position (marked PBE) is the point in the original F_L function.

[0077] A useful variation of this method, display adaptation, is summarized in Figure 4 by demonstrating its effect on a plot of possible normalized input luminance Ln_in versus normalized output luminance Ln_out (which is converted to actual luminance, i.e., PL_V value, by multiplying it by the maximum luminance of the display associated with the normalized luminance).

[0078] For example, a video creator designs a luminance mapping strategy between two reference gradings, as illustrated in Figure 1. Therefore, for the possible normalized luminance Ln_in of pixels in an input image, e.g., a master HDR image, this normalized input luminance must be mapped to the normalized output luminance Ln_out of a second reference grading, which is the output image. This regrading of all luminances corresponds to a function F_L, which has many different shapes determined by a human grader or grading automaton, and the shape of this function is communicated as dynamic metadata.

[0079] The question now is, in this simple display adaptive protocol, what shape should a derived secondary version of the F_L function take to map to an MDR image for an intermediate dynamic range display (instead of a reference SDR image) (assuming the mapping starts again from an HDR reference grading image as the input image)? For example, it could be calculated based on a metric as follows: for example, an 800nit display should have 50% of the grading effect, and a full 100% is regrading the master HDR image to a 100nit PL_V SDR image. In general, via the metric, for a possible normalized input luminance (Ln_in_pix) of a pixel, represented as display adaptive luminance L_P_n, an arbitrary point between no regrading to a second reference image and full regrading is determined, where this location naturally depends on the input normalized luminance, but also on the value of the maximum luminance (PL_V_out) related to the output image. Those skilled in the art will understand that while the function can be expressed in normalized luminance representation, the function can be equivalently expressed in any normalized luminance representation defined according to any OETF.

[0080] The corresponding display adaptive luminance mapping FL_DA is determined as follows (see Figure 4a): Take any one of all input luminances, e.g., Ln_in_pix. This corresponds to the starting position on the diagonal (shown as a square) that has equal angles with respect to the input and output axes of the normalized luminance. Position a scaled version of the metric (scaled metric SM) for each point on the diagonal, perpendicular to the diagonal (or 135 degrees counterclockwise from the input axis), starting on the diagonal and ending at a point on the F_L curve (at the 100% level), i.e., the intersection of the F_L curve and the vertically scaled metric SM (shown as a pentagon). (In this example, for this PL_D value of the display from which the image must be calculated) position a point at the 50% level of the metric, i.e., in the middle [note that in this case the PL_V value of the output image is set to be equal to the PL_D value of the display from which the display-optimized image must be supplied]. By doing this for all points on the diagonal corresponding to all Ln_in values, the FL_DA curve is obtained, which is formed similarly to the original, i.e., performs the same regrading but with maximum luminance rescaling / adjustment. This function is now ready to be applied to calculate the required corresponding optimally regraded and / or display-adapted 800nit PL_V pixel luminance, given any input HDR luminance value of Ln_in. This function FL_DA is applied by the luminance mapping circuit 310.

[0081] In general, the characteristics of this display adaptation are as follows (not intended to be further limited): The direction of the metric may be fixed in advance as technically desired. Figure 4b shows another scaling metric, namely the vertical scaling metric SMV (i.e., perpendicular to the axis of the normalized input luminance Ln_in). Again, 0% and 100% (or 1.0) correspond to no regrading (i.e., an identity transformation to the input image luminance) and regrading to the second of two reference grading images (associated in this example by a luminance mapping function F_L2 of a different shape), respectively.

[0082] The location of measurement points on a metric, i.e., where values ​​such as 10% and 20% are located, is generally nonlinear, although it can be further modified technically.

[0083] This is technically pre-designed, for example, for television displays. For example, a function like the one described in WO2015007505 is used. A logarithmic function can also be designed such that a*(log(PL_V)+b) is equal to a PL_V_HDR value of 1.0 (e.g., 5000 nits), and the point 0.0 corresponds to a PL_V_SDR reference level of 100 nits, or vice versa. The PL_V_MDR position for which image brightness needs to be calculated is then obtained from the pre-designed mathematics of the metric.

[0084] The effects of such metrics are summarized in Figure 5.

[0085] The display adaptive circuit 510 includes a configuration processor 511, for example, in a television or set-top box. It sets values ​​for processing an image before the running pixel colors of the image are processed. For example, the maximum luminance value of the display optimized output image PL_V_out may be set once to the set-top box by polling it from the connected display (i.e., the display communicates the maximum displayable luminance PL_D to the set-top box), or, if the circuit is present in a television, this may be set by the manufacturer, etc.

[0086] The luminance mapping function F_L varies for each input image in some embodiments (and is fixed for many images in other variations) and is input from several sources of metadata information 512 (for example, this is broadcast as an SEI message and read from a memory sector such as a Blu-ray disc). This data establishes a normalized height for a normalized metric (Sm1, Sm2, etc.), on which the desired position of the PL_D value is found from the mathematical formula of the metric.

[0087] When the input image 513 is input, consecutive pixel luminances (e.g., Ln_in_pix_33 and Ln_in_pix_34, or Luma) pass through a color processing pipeline to which display adaptation is applied, resulting in corresponding output luminances such as Ln_out_pix_33.

[0088] Please note that this method does not provide any minimum black brightness.

[0089] The reason is that the usual method is as follows: Black levels depend heavily on actual viewing conditions, and are even more variable than display characteristics (i.e., the initial PL_D). All sorts of influences arise, from the physical lighting conditions to the optimal configuration of photosensitive molecules in the human eye.

[0090] Therefore, it's simply about creating a good image "for the display" (i.e., how much better (in terms of brightness) the intended HDR display is than a typical SDR display). Then, if necessary, some post-correction can be made later for the viewing conditions, which is the (undefined) special task left to the display.

[0091] Therefore, we generally assume that a display has variable high-brightness capability, i.e., it can display all the necessary pixel brightness encoded in the image up to PL_D (for the time being, we assume that it is an MDR image already optimized for the PL_D value, i.e., that there are generally at least some pixel regions in the image that reach up to PL_D). This is because we generally do not want to suffer the harsh consequences of white clipping, but as mentioned above, the black in the image is often of no interest.

[0092] Since black is "almost" visible anyway, it's not the most important thing if some parts are slightly less visible. At the very least, the potentially very bright pixel brightness of the master HDR grading can be optimally suppressed within the limited upper range of the display, for example, above 200 nits, for example, from 200 nits to PL_D=600 nits (for example, for master HDR brightness up to 5000 nits).

[0093] This is similar to assuming that black is always zero nits for all images and all displays. White clipping is a far more visually jarring characteristic than losing a portion of black, and it is often still possible to see something, but it is not as comfortable.

[0094] However, sometimes this method is insufficient. This is because, under considerable ambient light in a viewing room (for example, a consumer television viewer's living room with large windows during the day), a significant subrange of the darkest luminances becomes invisible, or at least not clearly visible. This differs from the ambient light conditions in a video editing room where video is created, which can be dim or even darker.

[0095] Therefore, for example, it is generally necessary to increase the brightness of those pixels using the display's control buttons (so-called brightness buttons).

[0096] When using a television electronic behavior model such as Rec.ITU-R BT.814-4(07 / 2018), a television in an HDR scenario obtains lumens and chroma pixel colors (which actually drive the display) and converts these into nonlinear R', G', and B' nonlinear drive values ​​to drive the panel (according to standard colorimetric calculations). The display then processes these R', G', and B' nonlinear drive values ​​in the PQ EOTF to determine which front screen pixel brightness to display (i.e., generally, there is still internal processing that causes the electro-optical physical behavior of the LCD material, for example, the way in which OLED panel pixels or LCD pixels are driven, but the manner in which this is irrelevant to this discussion).

[0097] Next, for example, a control knob on the front of the display provides a luma offset value b (the moment when black patches 2% above the minimum value in PLUGE or other test patterns become visible, while -2% black patches disappear).

[0098] The original, uncorrected display behavior is LR_D=EOTF[max(0,R')]=PQ[max(0,R')] LG_D=EOTF[max(0,G')]=PQ[max(0,G')] LB_D=EOTF[max(0,B')]=PQ[max(0,B')] [Formula 1] In that case.

[0099] In this equation, LR_D is the amount of red contribution (linear) that should be displayed to create a specific pixel color with a specific luminance (in (partial) nits), and R' is the nonlinear Luma code value, e.g., 419 out of 1023 in 10-bit coding.

[0100] The same applies to the blue and green components. For example, if you need to create a specific color with 1 nit (the total brightness of that color to the eye), you need, for example, 0.33 units of blue, and the same applies to red and green. If you need to create 100 nits of that same color, you can say that LR_D = 100 * 0.33 nits.

[0101] Here, if this display drive model is controlled via a luma offset knob, the general formula is as follows: LR_D_c = EOTF[max(0, a*R'+b)], where a = 1 - b / OETF[PL_D], etc. [Equation 2]

[0102] Instead of displaying a zero black in the image hidden somewhere within the invisible black of the display, this technique raises the zero black to the level where black is clearly distinguishable (note that consumer displays may use mechanisms other than PLUGE, for example, that despite the viewer's preferred available luma offset b value, viewer preference may potentially lead to another possible suboptimal).

[0103] This is a post-processing step for the display after creating an optimally regraded image. Specifically, first, the optimally theoretically regraded image is calculated by the decoder and, for example, first mapped to a reconstructed master HDR image, and then luminance remapped to, for example, a 550nit PL_V MDR image, i.e., the brightness capability PL_D of the display is taken into account.

[0104] Then, after this optimal image is determined according to the filmmaker's ideal vision, it is further mapped by the display, taking into account the expected visibility of black in the image.

[0105] U.S. Patent Application Publication 2017025603 teaches a method for achieving optimal luminance mapping that depends on the amount of ambient light by analyzing the luminance histogram of the input image. Minimum, maximum, and average luminance are used to perform the corresponding ternary curve mapping, after which an appropriate tone mapping optimizes the luminance; that is, a function that boosts brightness more strongly for the darkest image luminances is selected for mostly dark scenes, and a less steep function is selected for bright scenes. The user fine-tunes by using a function that is roughly in between these two.

[0106] U.S. Patent Application Publication 2019304379 teaches how to pre-calculate a virtual image, essentially a 5000-nit master HDR image, but with darker image objects brightened primarily to compensate for viewing environments brighter than the 5-nit master HDR image to which the master HDR image was graded, by using a predetermined ambient brightness correction luminance mapping function. A standard display adaptive algorithm, generally determined by minimum, average, and maximum luminance, then uses a sigmoid mapping of the pre-corrected virtual image rather than the original master HDR image. There is also specific teaching on how a brighter (virtual) image can be determined by using human visual contrast characteristics in PQ space.

[0107] U.S. Patent Application Publication 20170186141 teaches that, generally, the mapping from luminance in an HDR image to luminance in an SDR image uses a tone mapping function (generally piecewise linear) which is determined by the television or mobile phone user. [Overview of the Initiative] [Problems that the invention aims to solve]

[0108] According to the inventors, the problem is that this is, to some extent, a crude way of adapting the ambient light level of the viewing room for the image to be displayed. Alternative methods have been developed, in particular, that allow viewers to see images with higher contrast in at least some important parts of the image. In general, the prior art also does not address the high regrading requirements of video or image creators. [Means for solving the problem]

[0109] Images that appear visually better to various ambient lighting levels are obtained by processing the input image to acquire the output image. The input image has pixels having input brightness within a first brightness dynamic range (DR_1) having a first maximum brightness (PL_V_HDR), The reference luminance mapping function (F_L) is received as metadata associated with the input image. The reference luminance mapping function specifies the relationship between the luminances of pixels placed in two images, and the two images are graded differently in such a way that the pixel luminances of the same image objects are different in the two images. The reference luminance mapping function specifies the relationship between the luminance of the input image and the luminance of a secondary reference image having a second reference maximum luminance (PL_V_SDR). The output image has an output maximum brightness (PL_V_MDR) that is different from the first maximum brightness and the second reference maximum brightness (PL_V_SDR). The process includes the steps of determining an adaptive luminance mapping function (FL_DA) based on a reference luminance mapping function (F_L) and a regulated maximum luminance value (PL_V_CO), determining that the regulated maximum luminance value (PL_V_CO) is different from the output maximum luminance (PL_V_MDR), and applying the adaptive luminance mapping function (FL_DA) to the input pixel luminance to obtain the output luminance of the output image. The calculation of the adaptive luminance mapping function (FL_DA) involves a step of finding a position (pos) on the metric (SM) that specifies the position of maximum luminance, and a step of finding that this position corresponds to the adjusted maximum luminance value (PL_V_CO). The first endpoint of the metric corresponds to the first maximum brightness (PL_V_HDR), and the second endpoint of the metric corresponds to the second reference maximum brightness (PL_V_SDR). The first endpoint of the metric is located at a point on the diagonal that has horizontal and vertical coordinates equal to the normalized input luma (Yn_CC0) for any normalized input luma. The second endpoint is located on the trajectory of the reference luminance mapping function (F_L) determined by the direction of the metric, The adjusted maximum brightness value (PL_V_CO) is, Obtain the user-corrected value (UCBVal), The adjusted maximum brightness value (PL_V_CO) is determined by subtracting the user correction value (UCBVal) from the maximum output brightness (PL_V_MDR). It is determined by, The output brightness is written to the output image as the pixel color, and the output image is output. It is characterized by the following.

[0110] A reference luminance mapping function typically indicates how a video should be regraded when creating an image with a lower maximum (possible) luminance that visually corresponds to a master HDR image (in video, maximum luminance is generally an upper limit that the creator can use to make some of the pixels corresponding to extremely bright objects in part of the image as bright as possible, according to the chosen variation of the technical HDR encoding). That is, the need to regrade each image in any HDR scene (e.g., the first night image in a movie versus the later day image) is specified by a function (e.g., by the content creator's color grader or automaton) that maps the luminance between two specifically graded reference images. Generally, the first of these reference images is the input image itself (e.g., a master HDR image or SDR image with a maximum luminance, e.g., 5000 nits). The second reference image is also called the secondary reference image. The function is generally defined in the luma domain (according to the selected OETF function flavor), and furthermore, color processing occurs in the luma domain (however, the metric generally indicates the position of maximum luminance rather than luma, even when superimposed on such a luma domain plot of the luminance regrading function). Luma is generally normalized. The second reference image is typically (non-limited), generally an SDR regraded from a master HDR image, i.e., a 100 nit maximum luminance image (or vice versa, where the SDR image becomes the starting image for regrading the corresponding HDR image, but its aspect does not alter the contrast improvement process).

[0111] When proposing his method, the inventors desired to have a way to take into account this guide of the need for luma or luminance regrading, specified by the shape of the function F_L. In particular, the inventors desired user control that works under such technical conditions in an easy and practical way to efficiently improve the contrast of an image when viewed in a high-level lighting environment that is considerably higher than the light level at which the image is ideally viewed. The inventors desired an easy way to interfere with a regrading curve (i.e., F_L) that would work well for any possible curve co-communicated by the video creator (it is a simple concave function that relatively boosts the luma of darker pixels for regrading to reduce the maximum luminance of the output image, while gradually decreasing toward the brightest luma, but is also a fairly complex curve where, for example, the contrast of a significant subrange of luma around the center of the overall range remains stretched by the curve F_L, at which point it takes on the appearance of, for example, a double step).

[0112] Therefore, a display adaptation method such as the one invented by the present inventor is used, which is now adjusted (shape adapted), i.e., diagonal squeezing is produced by a different control variable, i.e., a regulated maximum luminance value (PL_V_CO).

[0113] While various metrics are defined, generally, the first endpoint (corresponding to the master HDR grading / the position of maximum brightness in the image) is always positioned on a line in the plot of normalized luma between zero and one, with a diagonal, i.e., at a 45-degree angle with respect to both the input luma axis and the output luma axis. Specifically, for each possible input luma Y_in, this point has a position on the diagonal (and furthermore, a vertical coordinate value) with a horizontal coordinate equal to Y_in for any Y_in. The position of the endpoint of the metric, i.e., the position that determines the value of PL_V_SDR in a typical example, lies somewhere on the trajectory of the shape of the F_L function, and this position depends on which angle was chosen to perform the display adaptation algorithm (for example, at a vertical angle, the metric endpoint is h=Y_in;v=F_L(Y_in)).

[0114] Such direct user control devices do not require the measurement of ambient light level indicators, i.e., ambient illumination values ​​(Lx_sur), which are quantified as luminance. However, such illuminance measurements may be optionally present to further guide the user, for example, by specifically indicating a good working or initial value, i.e., the amount of light for daytime or nighttime television viewing in a living room.

[0115] The reference illumination value (GenVwLx) is a typical value, e.g., an expected value or a value that performs well. Ideally, it is determined in response to the received image and is co-communicated as image metadata indicating, for example, that the image was made for such ambient lighting or that it looks best under such ambient lighting. It is, for example, a value for viewing a master HDR image (e.g., VIEW_MET), or a value derived from it by applying a formula to obtain a final typical value, taking into account, for example, SDR grading values ​​as well. This formula is calculated at the receiving end, or the result is calculated by the video or image creator and co-communicated as metadata. For example, the GenVwLx value is overwritten in memory and includes a fixed manufacturer value or a generally communicated value (e.g., from a previous program) if metadata for the current image (set) is not received, etc.

[0116] Luma is calculated by applying the photoelectric conversion function (OETF_psy) to the input luminance (L_in) of the input image, provided that input luminance (L_in) is present as input. In some methods or devices, the input image pixel color has a lumen component.

[0117] The photoelectric transfer function (OETF_psy) used is preferably a psychovisually uniform photoelectric transfer function. This function is defined by determining the shape of the luminance-to-luma mapping function (generally experimentally in the laboratory), so that a second luma that is a fixed integer N lumas higher than a first luma selected somewhere within the luma range roughly corresponds to a similar difference in perceived brightness compared to the perceived brightness of the first luma to a human observer. Further regrading mapping is then defined within this visually uniform luma domain.

[0118] In other words, human vision is nonlinear, and therefore, the difference between 10 nits and (1.05)*10 nits is not perceived in the same way as, for example, the difference between 2000 nits and (1.05)*2000 nits. The uniform perceptual curve also depends on what the person is seeing, that is, in particular, the dynamic range of the display being viewed (in a particular environment), i.e., the maximum brightness PL_D and the minimum brightness.

[0119] Therefore, ideally, for processing purposes, we define an OETF (or vice versa EOTF) with the following characteristics.

[0120] When taking the first luma, for example in 10 bits, luma_1 = 10. Then, for example, by moving the N=5 luma code higher, we obtain the second luma, luma_2 = 15. This corresponds to a change in brightness perception (i.e., a value that characterizes what a human viewer experiences as brightness of a displayed patch at a particular physically displayed brightness; it is determined by applying the EOTF).

[0121] Therefore, we assume that, within the range of brightness from 1 to 200, Luma 10 gives the impression of a brightness of 5, and Luma 15 gives the impression of a brightness of 7, meaning that the brightness is 2 units brighter.

[0122] Next, we take two lumens that encode brighter luminances, for example, 800 and 800+5. Then, the lumens scale is visually nearly uniform if the viewer experiences the same brightness difference due to the lumens difference. For example, lumens 800 gives a perceived brightness of 160, and lumens 805 appears as if it were brightness 162, i.e., there is again a difference of 2 units in brightness. Since a perceptually moderately uniform lumens definition already works well, the function does not need to be exactly a brightness determination function.

[0123] The effects of changes caused by brightness processing are not as visually unpleasant when performed in such a psychovisually uniform system (because the brain suppresses them more easily).

[0124] Display adaptation generally reduces the lumens of darker objects compared to standard regrading to SDR images, but now the goal is to maintain the lumens to the values ​​required to provide a visually pleasing image.

[0125] That is, any possible embodiment is used to create a fewer variation of the specially formed F_L function (to correspond to the need for specific required luminance remapping, also known as regrading, for a particular HDR image or scene, in order to obtain a corresponding lower maximum luminance secondary grading) as described in Figures 4 and 5, or similar techniques. The luminance mapping function is generally represented as a lumens mapping function, i.e., by converting both normalized luminance axes to lumens axes via a suitable OETF (normalization using appropriate maximum luminance values ​​of the input and output images).

[0126] Specifically, the quadratic grading for which the F_L function is defined is, favorably, 100-nit grading, which is suitable to meet the requirements of most future display situations (it should be noted that there will also be display scenarios with lower dynamic range quality in the future).

[0127] In actual embodiments, a scaler and / or clipper is present to prevent the adjusted maximum luminance value PL_V_CO from becoming too low, for example, by setting it to be equal to PL_V_SDR if it would be lower than PL_V_SDR, or generally, a damping function is present to reduce the slope for higher input values.

[0128] For example, this is done in a user-value circuit (903) by the manufacturer pre-setting an intensity value (EXCS), which determines how sensitive the device is to user control (e.g., repeatedly pressing a button or turning a knob), and linear user control is generally satisfactory (however, more advanced nonlinear controllers may also be used).

[0129] Advantageously, the method for processing the input image uses a reference luminance mapping function (F_L) created by the creator of the input image and received as metadata via the image communication channel.

[0130] The system is also used in devices that determine a unique version of a well-functioning regrading function (i.e., heuristically determine a good shape of F_L based on an analysis of luma generation in some previous and / or current images (in the case of delayed output)), but is particularly useful when operating with a regrading function determined by the creator.

[0131] In the algorithm, the metric is represented mathematically in the calculations within the electronic circuit as a plot of input versus output luma. The locus of the reference luminance mapping function (F_L) in such a plot is the position whose vertical coordinate is the normalized output luma obtained when the normalized input luma is used as the input to the reference luminance mapping function (F_L).

[0132] Various display adaptive algorithms are selected for this method by predefining the direction that the algorithm must use. The first point of the metric (which corresponds to the input image and is the identity transform of Luma when Luma maps the input image to itself) is always on the diagonal. The second endpoint is a point on the locus of F_L directly above the first endpoint for any of the possible normalized input Luma (0 to 1.0). At a 45-degree angle from the axis (horizontal axis) of the input Luma, or 135 degrees counterclockwise, the line in this direction intersects the F_L locus at a horizontally offset position, and the second endpoint is located there. That is, exactly where the second endpoint of the metric is located on the function's locus is determined by the direction of the line segments starting from each position on the diagonal. These calculations are accelerated, for example, by computing the adaptive luminance mapping function (FL_DA) as a 1D LUT with N output entries for N normalized input Luma, before receiving the set of images to be processed to which the function F_L is applicable.

[0133] In these embodiments, the user-controlled value UCBVal is generally expressed as a quantity with the same dimensions as the maximum luminance value, i.e., in nit units (SI units Cd / m²). 2 It is presented as (a more easily quantifiable expression for printing purposes).

[0134] Advantageously, the method for processing the input image sets the output maximum brightness (PL_V_MDR) to be equal to the maximum displayable pixel brightness of the display from which the output image may be supplied. Therefore, all processing operates with an adjusted maximum brightness value (PL_V_CO) based on this value of the connected display, generally the display on which the viewer is watching or intending to watch video, e.g., a broadcast television program (however, there may be systems where a different value is used, for example, if the display communicating it desires a slightly higher PL_V_MDR and still wants internal brightness optimization processing, then this method still works similarly).

[0135] A simple but sufficient embodiment of a method for processing an input image is: Steps include obtaining input values ​​(UCBSliVal) from the user of the display, The steps include: calculating the user-controlled value (UCBVal) by multiplying the input value by the intensity value (EXCS) and the maximum output brightness (PL_V_MDR); and It holds.

[0136] Such methods should be properly scaled and intuitive to use (it should be noted that the impact remains within a psychovisually uniform domain, so that users have visually precise control, which harmonizes with the need for regrading, i.e., minimizes disruption to the creator's artistic vision).

[0137] Advantageously, the method of processing the input image works not only by correcting the contrast of the darkest pixels in particular (we are not aware of any specific black offsetting process), but also by performing a process to set the darkest black in the input image as a function of the ambient illumination value (Lx_sur) to the black offset value. For example, it starts with the image with the smallest black luma. In principle, any such method may be combined, but below we teach some particularly advantageous methods that work especially well with contrast optimization.

[0138] Advantageously, the method for processing the input image is pre-set with the metric direction as vertical, which means that the second endpoint for the normalized input luma (Yn_CC0) is located where the result of applying a reference luminance mapping function with the normalized input luma as the horizontal coordinate and that normalized input luma (F_L(Yn_CC0)) as the vertical coordinate.

[0139] This method is further embodied as a device (900) for processing the input image in order to acquire the output image. The input image has pixels having an input brightness within a first brightness dynamic range (DR_1) having a first maximum brightness (PL_V_HDR), and this device, An image input unit (921) for acquiring an input image (513), This is a data input unit (920) for receiving a reference luminance mapping function (F_L), which is metadata related to the input image. The reference luminance mapping function specifies the relationship between the luminance of the first reference image and the luminance of the second reference image. The reference luminance mapping function specifies the relationship between the luminance of the input image and the luminance of the second reference image. The second reference image has the second reference maximum brightness, The output image has an output maximum brightness (PL_V_MDR) that is different from the first maximum brightness and the second reference maximum brightness, and the data input section (920) Includes, This device, A user value circuit (903) configured to determine and output a user correction value (UCBVal) as set by the human user of the device, The user correction value (UCBVal) is obtained from the user value circuit (903). The adjusted maximum brightness value (PL_V_CO) is output as a result of subtracting the user correction value (UCBVal) from the maximum output brightness (PL_V_MDR). A maximum brightness determination unit (901) configured as follows and It further includes, This device, The display adaptive circuit (510) is configured to determine an adaptive luminance mapping function (FL_DA) based on the adjusted maximum luminance value (PL_V_CO) and the reference luminance mapping function (F_L). The calculation of the adaptive luminance mapping function (FL_DA) involves finding the position (pos) on the metric (SM) corresponding to the adjusted maximum luminance value (PL_V_CO). The first endpoint of the metric corresponds to the first maximum brightness (PL_V_HDR), and the second endpoint of the metric corresponds to the maximum brightness of the second reference image. The first endpoint of the metric is located at a point on the diagonal that has horizontal and vertical coordinates equal to the normalized input luma (Yn_CC0) for any normalized input luma. The second endpoint is placed on the locus of the reference luminance mapping function (F_L), Display adaptive circuit (510) is configured such that the display adaptive unit applies an adaptive luminance mapping function (FL_DA) to the input luminance, or the input luma encoded from the input luminance, to obtain the output luminance or output luma, and outputs these output luma or output luminance as the pixel color of the output image. It also includes.

[0140] A useful further embodiment of this apparatus for processing input images has a user-value circuit (903), and the user-value circuit (903) is Obtain the input value (UCBSliVal) from the user of the display, The user-controlled value (UCBVal) is calculated by multiplying the input value by the intensity value (EXCS) and the maximum output brightness (PL_V_MDR). The user-controlled value (UCBVal) is output to the maximum brightness determination unit (901). It is configured in this way.

[0141] A useful further embodiment of this apparatus for processing an input image includes a black adaptive circuit configured to set the darkest black luma in the input image as a function of the ambient illumination value (Lx_sur) to a black offset luma.

[0142] In particular, those skilled in the art will understand that these technical elements are embodied in various processing elements such as ASICs (Application-Specific Integrated Circuits, i.e., generally ICs designed by IC designers to have an IC (or part of an IC) perform this method), FPGAs, processors, etc., and are present in various consumer or non-consumer devices, whether or not they include displays or non-display devices externally connected to displays; that images and metadata are transmitted in and out by various image communication technologies such as wireless broadcasting and cable-based communications; and that devices are used in various image communication and / or usage ecosystems, such as television broadcasting and on-demand internet services.

[0143] These and other aspects of the methods and apparatus according to the present invention will be made apparent and described with reference to the embodiments and models described below and with reference to the accompanying drawings, which serve only as non-limiting specific illustrations illustrating more general concepts, and in the accompanying drawings, dashes are used to indicate that a component is optional, and components that are not dashed are not necessarily essential. Dashes are also used to indicate elements that are described as essential but are hidden inside an object, or intangible things such as the selection of an object / region. [Brief explanation of the drawing]

[0144] [Figure 1]This diagram schematically illustrates some typical color conversions. Color conversion is performed when an optimally high dynamic range image is optimally color-graded and looks similar to a corresponding lower dynamic range image, e.g., a standard dynamic range image with a maximum brightness of 100 nits, given the differences in the first dynamic range DR_1 and the second dynamic range DR_2, respectively. It also corresponds, in the case of losslessness, to mapping a received SDR image that actually encodes an HDR scene to a reconstructed HDR image of that scene. Brightness is shown as a location on the vertical axis from the darkest black to the maximum brightness PL_V. The brightness mapping function is symbolically shown by arrows mapping the average object brightness from brightness on the first dynamic range to the second dynamic range (those skilled in the art will know how to plot this equivalently on an axis normalized to, for example, 1, by dividing by the respective maximum brightness, as a classical function). [Figure 2] This figure schematically illustrates an example of a high-level view of a technique recently developed by the applicant for encoding high dynamic range images, i.e., images that can typically have a brightness of at least 600 nits (typically 1000 nits or more). It actually communicates an HDR image either by itself or as a corresponding luminance-regraded SDR image, plus metadata encoding a color conversion function that includes at least a well-determined luminance mapping function (F_L) for pixel colors used by a decoder to convert the received SDR image to an HDR image. [Figure 3] This figure shows the internal details of an image decoder, particularly a pixel color processing engine, as a (non-limiting) preferred embodiment. [Figure 4]The sub-images Figures 4a and 4b show two possible variations of display adaptation to obtain the final display-adapted luminance mapping function FL_DA, which is used to calculate the optimal display-adapted version of the input image for a given display capability (PL_D), from a reference luminance mapping function F_L that systematizes the need for luminance regrading between two reference images. [Figure 5] This diagram provides a more general summary of the display adaptation principle, which is a component in the formulation of this embodiment and the claims, in order to make the principle of display adaptation easier to understand. [Figure 6] This figure shows an exemplary apparatus for illustrating the technical elements related to environmental lighting compensation processing. [Figure 7] This diagram illustrates the correlation concept of a virtual target display, specifically the relationship or association between the minimum brightness (mL_VD) of the target display and the actual display, such as the viewer's end-user display, where brightness remapping is typically performed. [Figure 8] This diagram illustrates some categories of metadata, some of which are available for long-standing professional HDR coding frameworks (e.g., typically co-communicating with pixel color images), and therefore any receiver has all the available data that is necessary, or at least elegantly utilized, to optimize the received image, particularly for specific end-viewing conditions (display and environment). [Figure 9] This diagram illustrates methods for optimizing image contrast, particularly when designing based on existing display adaptation techniques, and especially for compensating for various viewing ambient light levels or conditions. [Figure 10] This figure shows some examples of image luminance using a specific reference mapping function F_L (chosen to be a simple linear curve for ease of understanding) to illustrate what happens when viewing environment lighting adaptation techniques (explained in Figure 6) are applied. [Figure 11]Figure 9 shows an example of what happens when the contrast enhancement techniques described are applied (without ambient lighting, but the two processes can also be combined, which results in an additional offset to the deepest blacks). [Modes for carrying out the invention]

[0145] Figure 6 shows an apparatus for generally illustrating the elements of the present invention, in particular, a method for performing the color processing circuit 600. This is described in general terms, and then some details of variations of the embodiment are described. While it is assumed that ambient light adaptive color processing is performed within an application-specific integrated circuit (ASIC), those skilled in the art will understand how color processing is similarly performed in other related devices. Some embodiments differ depending on whether this ASIC resides, for example, in a television display (typically an end-user display) or in another device connected to the display, such as a set-top box, to perform processing on a display and provide it along with an already brightness-optimized image. (Those skilled in the art may also map this apparatus configuration diagram to a method flowchart of the corresponding method).

[0146] The input luminance (L_in) of the input image to be processed (assuming it is the master HDR image MAST_HDR for generality) is input via the image pixel data input 690 and is first converted to the corresponding lumens (e.g., 10-bit encoded lumens). For this purpose, the optical-electron transfer function is used, which is generally fixed to the device by the manufacturer (although it is configurable).

[0147] Using perceptually uniform OETF (OETF_psy) is useful.

[0148] Assume the following OETF is used (defined as the output of LumaYn_CC0). Yn_CC0=v(L_in;PL_V_in)=log[1+(RHO-1)*power(L_in;p)] / log[RHO] [Formula 3] Here, RHO is a constant that depends on the input maximum brightness PL_V_in, given by the formula RHO(PL_V_in)=1+32*power((PL_V_in / 10,000);p), where p is preferably a power of a power function equal to 1 / (2.4).

[0149] The input maximum brightness is the maximum brightness associated with the input image (it is not necessarily the brightness of the brightest pixel in each image of the video sequence, but rather metadata that characterizes the image or video as an absolute upper limit).

[0150] It is generally configurable and input via the first maximum data input section 691 as PL_V_HDR, for example 2000nit. In other variations, it is further a fixed value for the image communication and / or processing system, for example 5000nit, and therefore a fixed value in the processing of the photoelectric conversion circuit 601 (thus the vertical arrow representing the PL_V_HDR data input is shown as a dotted line, as it is not present in all embodiments; note that the white circle symbolizes this branch of data supply so as not to be confused with the non-mixed duplicate data bus).

[0151] Next, depending on the embodiment, there may be further luminance mapping by a luminance mapping circuit 602 in order to obtain an initial luminance Yn_CC that is optimally adapted to a particular viewing environment. Such further luminance mapping is generally display adaptation to pre-fit the image to a specific maximum display luminance PL_D of the connected display.

[0152] In a simpler embodiment, this optional luminance mapping does not exist, and the 2000nit environment-optimized image is calculated against the input 2000nit master HDR image (or a fixed 5000nit HDR image situation), i.e., assuming that the maximum luminance value remains the same between the input and output images, the color processing will be described first. In this case, the start lumen Yn_CC is simply equal to the initial start lumen Yn_CC0 output by the photoelectric conversion circuit 601.

[0153] The linear scaling circuit 603 is, Yim=(Ydif+1)*Yn_CC-1.0*Ydif [Formula 4] The intermediate luma Yim is calculated by applying a function of type .

[0154] The lumen difference Ydif is obtained from the (second) photoelectric conversion circuit 611. This circuit converts the luminance difference dif to the corresponding lumen difference Ydif. This circuit uses the same OETF formula as circuit 601, i.e., it similarly uses the same PL_V_in value (and power).

[0155] The luminance difference (dif) is calculated by the luminance difference calculator 610, which receives two darkest luminance values, namely the minimum luminance of the target display (mL_VD) and the minimum luminance of the end-user display (mL_De), given specific lighting characteristics of the viewing room. dif = mL_VD - mL_De [Equation 5] Calculate, mL_De is generally a function (generally added) of the fixed minimum display black (mB_fD) of an end-user display and the luminance (mB_sur), which is a function of the amount of ambient light. A typical example of minimum display black (mB_fD) is the stray light of an LCD display. Even if such a display is driven by a code that indicates perfect black (i.e., ideally zero photons output), due to the physical properties of the LCD material, such a display will still always emit light, for example, 0.05 nits, so-called stray light. This is regardless of the ambient light in the viewing environment and therefore applies to a completely dark viewing room.

[0156] The values ​​of mL_VD and mL_De are generally obtained via the first minimum metadata input 693 and the second minimum metadata input 694, for example, from external memory in the same or different devices, or via circuitry from an optical measuring device. The maximum luminance value required in a particular embodiment, for example, the maximum luminance (PL_D) of a display capable of displaying a color-processed output image, is generally input via an input connector such as the second maximum metadata input 692. The output image color is fully written in the pixel color output 699 (those skilled in the art will understand how this can be implemented in various technical variations, e.g., as IC pins, standard video cable connections such as HDMI®, wireless channel-based communication of the image's pixel color data, etc.).

[0157] The precise determination of luminance [referred to as ambient black luminance] as a function of ambient light (mB_sur) is not typical of the present invention, as it may be determined in several alternative ways. For example, a viewer may use a test signal such as PLUGE, or a more consumer-friendly variation, to determine a value of luminance mB_sur that they generally consider to represent the masking black resulting from reflections on the display's front screen. We further assume that the viewer simply sets the value regardless of whether it is a particular evening or, for example, from a situation where the consumer typically keeps the room lighting configuration fixed when purchasing the display. Or it may even be a value that television manufacturers have burned in as a value that works well on average for at least one of a typical consumer viewing situation. When this method is used to adapt luminance for viewing on a mobile device, the viewing environment is generally not very stable (e.g., watching a video on a train, and the lighting level changes from outdoors to indoors as the train enters a tunnel).

[0158] In such cases, for example, time-filtered measurements from the built-in light meter can be used (measurements should not be taken too frequently to avoid unnecessarily adapting the processing over time, for example, when sitting on a bench in the sun and then walking indoors).

[0159] Such instruments generally measure the average amount of light (in lux) that hits the display, and therefore the display itself.

[0160] Although they are different photometric quantities, the lux value is converted to ambient black luminance using a well-known photometric formula. mB_sur = R*Ev / pi [Equation 6]

[0161] In this equation, Ev is the ambient illuminance in lux units, Pi is a constant of 3.1415, and R is the reflectance of the display screen. Generally, mobile display manufacturers include these values ​​in the equation.

[0162] For a typical reflective surface in the surroundings, such as a house wall, which has a color somewhere between average gray and white, we can assume an R value of approximately 0.3. That is, as a rule of thumb, we can say that the luminance value is about 1 / 10 of the illuminance value mentioned in lux units. The front of a display reflects far less light. Depending on whether special anti-reflective technology is used, the R value is, for example, about 1% (although it can be higher, towards the 8% reflectivity of glass, which can be a problem, especially in brighter viewing environments). Therefore, mL_De = mB_fD + mB_sur [Equation 7]

[0163] The technical significance of the minimum brightness (mL_VD) value of the target display is further explained using Figure 7.

[0164] The first three luminance ranges, starting from the left, are actually "virtual" display ranges, i.e., ranges corresponding to images, not necessarily actual displays (i.e., target displays, i.e., displays on which the image could ideally be displayed, but which consumers potentially do not own). These target displays are co-defined because the image is created specifically for these target displays (i.e., luminance is graded toward a specific desired object luminance). For example, an explosion may not be bright enough on a 550-nit display, and therefore the grader might want to lower the luminance of other image objects so that the explosion appears to have at least some contrast. However, it is entirely possible that no one owns such a display, and the image still needs to be optimized by display adaptation to the actual display owned by the particular viewer. The range of luminance that is physically displayable on this end-user display is shown as the luminance range on the far right (EU_DISP).

[0165] This information for one or more target displays constitutes metadata, some of which is generally communicated together with the image itself, i.e., with the image pixel brightness.

[0166] This method for characterizing (encoding) HDR images deviates significantly from conventional SDR image encoding, and since these embodiments are only recently invented, and important technical elements should not be misunderstood, the necessary concepts are summarized for the reader using Figure 8.

[0167] HDR video is well supplemented by metadata, for this reason, because many aspects can differ (for example, the maximum brightness of an SDR display has always been in the range of approximately 100 nits, but now people have displays with considerably different display capabilities, e.g., 50 nits, 500 nits, 1000 nits, 2500 nits, and perhaps even 10,000 nits in the future, PL_D displays, and content characteristics such as the maximum codeable brightness of video PL_V have also changed considerably, and therefore the distribution of brightness between dark and light that a grader makes for a typical scene will also differ greatly between a typical SDR image and any HDR image, etc.), and one should not run into trouble due to insufficient control over these various unstable ranges.

[0168] As explained above, we should obtain at least one pixelation matrix of pixel colors, including at least pixel luma, otherwise the image shape will not even be visible (even if it is colorimetrically incorrect). As mentioned above, by actually communicating only one image per time, it is possible to communicate two different dynamic range images (which can serve as two reference gradings to indicate the need for luminance regrading of a particular video content when it is necessary to create images with different dynamic ranges, such as MDR images).

[0169] Assume that the master HDR image itself is communicated, and therefore the first dataset 801 (image color) contains color component triplets of the pixels of the communicated image, and becomes the input to a color processing circuit present in, for example, a receiving television display.

[0170] Generally, such images are digitized as (e.g., 10-bit) lumens and two chroma components, Cr and Cb (although nonlinear R'G'B' components can also be communicated). However, it is necessary to know which luminance 1023 or, for example, 229 lumens represents.

[0171] Therefore, container metadata 802 is communicated jointly. Luma is assumed to be defined according to, for example, the perceptual quantizer EOTF (or its inverse, OETF), as standardized in SMPTE ST.2084. This is a large container that can specify lumens up to 10,000 nits. Even if images are not currently generated using such high pixel luminances, it can be said that it is a "theoretical container" that includes luminances actually used up to, for example, 2500 nits. (Primary chromaticity is also communicated, but note that the details of these would simply unnecessarily hinder this explanation).

[0172] What's interesting is the maximum pixel brightness that can actually be encoded (or is encoded) in the video, which is encoded into another video feature metadata, which is generally the master display color volume metadata 803.

[0173] This is an important brightness value that the receiver should know. The reason is that even if the display doesn't care about the specific details of how it remaps all brightness along the range (at least according to the content creator's desired display adaptation), knowing the maximum still roughly guides what is best to do at all brightness levels, as it at least tells you what brightness the video will not exceed for each image pixel.

[0174] In the 2000nit example, this Master Display Color Volume (MDCV) metadata 803 includes the Master HDR Image Maximum Brightness, i.e., the PL_V_HDR value in Figure 7, i.e., what characterizes the Master HDR video (in this example, it is also communicated as the PQ Pixel Luma as defined in SMPTE2084, but that aspect can now be ignored, as those skilled in the art know how to convert between the two and can understand the inventive color processing principle as if the (linear) pixel brightness itself were coming in).

[0175] This MDCV is a "virtual" display, or target display, for the HDR master image. By indicating this in the metadata, the video creator indicates in the video signal communicated to the actual end-user display at the receiving end where there are pixels with a brightness of 2000 nits in the film, so that the end-user display can take this into full consideration when processing the brightness of the current image.

[0176] These (actual) luminances of the image set are, in fact, yet another aspect, and therefore there is a further video-related metadata set 804. This gives further information about the actual video, rather than the characteristics of the associated display (i.e., the maximum possible in the video). To easily understand this, let us assume that two videos are annotated with the same MDCV PL_V_HDR (and EOTF). The first video is a night video, and therefore, in reality, the image does not reach pixel luminances higher than, for example, 80 nits (although it is still specified in the MDCV for 2000 nits; furthermore, if it is another video, it may have flashlights in at least one image, which have a small number of pixels that reach the 2000 nit level, or nearly that level), and the second video, specified / created according to exactly the same encoding technique (i.e., annotated with the same data in 802 and 803), consists only of explosions, i.e., most of which have pixels above 1000 nits.

[0177] On the one hand, there may be something to add regarding this video, but on the other hand, those skilled in the art will understand that if you want to regrade both videos from a 2000nit representation to, for example, a 200nit output representation, you can do so in a different way (you can scale the explosion by simply dividing the luminance by 10, while keeping the luminance of the night scene the same in the master HDR and the 200nit output image).

[0178] A possible (optional in this invention, but described nevertheless for completeness) useful metadata in set 804 for annotating the transmitted HDR image is the average pixel brightness of all pixels in all time-series images, where MaxFall is, for example, 350 nits. The receiving color processing can then understand from this value that if it is dark, the brightness is displayed as is, i.e., without mapping, and if it is bright, dimming is required.

[0179] Even if the SDR video is not actually transmitted, i.e., even if only metadata (metadata 814) is sent, it is still possible to annotate the SDR video (i.e., a second reference grading image, which indicates how the SDR image should look to the content creator in the case of reduced dynamic range capability, making it as similar as possible to the master HDR image).

[0180] Therefore, some HDR codes may also send a pixel color triplet containing the pixel luma of an SDR image (as defined in Rec.709 OETF), i.e., SDR pixel color data 811. However, as explained above, the example codec, i.e., the SLHDR codec, does not actually communicate this SDR image, i.e., pixel color, or anything dependent on that color, and is therefore omitted (if pixel color is not communicated, there is also no need to communicate SDR container metadata 812, which indicates the container format in which the pixel code is defined and should be decoded into linear RGB pixel color components).

[0181] Ideally, what should be communicated (although some systems implicitly assume this) is the corresponding SDR target display metadata 813. In such situations, the value of PL_V_SDR is generally entered as equal to 100 nits.

[0182] Important to the present invention, this is also where the video creator typically fills in the assumed minimum black value of the theoretical target SDR display on which the film was graded, i.e., the SDR reference minimum black mB_SDR.

[0183] For example, if we assume that the creator is making video for a display that cannot go deeper than 0.1 nit (e.g., due to LCD light leakage), the creator does not want the brightness of very important image object pixels to approach this value, and starts with, for example, 0.2 nit SDR image pixels, maybe slightly above that, with most pixels exceeding the 1 nit level. There exists a similar value that characterizes the master HDR image or, more precisely, the target display associated with it, namely, the minimum black in mB_HDR (in metadata 803).

[0184] The need for regrading depends on the image, and is therefore advantageously encoded in SDR video metadata 814 according to the selected codec as described above. As described above, generally one (or more, even at a single time) image-optimized luminance mapping function (i.e., F_L reference luminance mapping function shape) is communicated for mapping HDR luminance (or luma) normalized to 1 up to SDR luminance or luma normalized to 1 up to (the exact way these functions are encoded is irrelevant to this patent application, and examples can be found in the ETSI SLHDR standard mentioned above). Now, using this function shape (combined with the target display's maximum luminance, HDR, and SDR metadata), all possible remappings of image pixel luminances are defined, so further metadata about the image, such as maxFall, can be communicated but is not actually required.

[0185] This already constitutes a fairly specialized set of HDR video encoded data, to which current ambient adaptive luminance remapping techniques can be applied.

[0186] However, two additional metadata sets (particularly useful for the contrast optimization embodiments described in more detail below) may be added (in various ways), but they are not yet fully standardized. Content creators may also work under the implicit assumption that video is created in a viewing environment with a specific illumination level, e.g., 10 lux, while using a perceptual quantizer (whether a particular content creator strictly adheres to this lighting suggestion is another matter).

[0187] For greater certainty, a typical (intended) viewing environment can be associated with the master HDR image, which has been specially graded (HDR Viewing Metadata 805; Optional / Dotted Line). As mentioned above, a maximum master HDR image of 2000 nits can be created. However, if this image is intended to be viewed in a 1000 lux viewing environment, the grader will probably not create too many subtle graded dark object luminances (such as slightly different dimly lit objects in a dark room seen through an open door in the background of a scene that includes a first room that is lit in the front), as the viewer's brain is likely to simply perceive all of this as "flat black." The situation is different if the image is made for viewing in a dim evening in a room of, for example, 50 lux or 10 lux. For typical brighter viewing conditions, e.g., 200 lux, the SDR image is regraded more specifically to the F_L function and annotated in the SDR ambient metadata 815, which the receiving device can then use to its advantage (or communicate regrading to different SDR images for different intended viewing).

[0188] Returning to Figure 7, we show the settings (by display adaptation) that we want to create an output image for a 600nit MDR display, i.e., the output image needs to be PL_V_MDR=600nit output image.

[0189] If we were to optimize an output image having the same maximum brightness as the input image according to the present invention (the simple situation described above in Figure 6), the minimum brightness (mL_VD) value of the target display would simply be the mB_HDR value. However, here we need to set the minimum brightness of the MDR display dynamic range (as an appropriate value for the minimum brightness mL_VD of the target display). This is interpolated from the brightness range information of two reference images (i.e., in the description of the standardized embodiment in Figure 8, the MDCV and the presented display color volume PDCV are generally communicated together or at least obtainable in metadata 803 and 813, respectively).

[0190] The formula is as follows: mL_VD=mB_SDR+(mB_HDR-mB_SDR)*(PL_V_MDR-PL_V_SDR) / (PL_V_HDR-PL_V_SDR) [Formula 8]

[0191] In this formula, the PL_V_MDR value is selected to be equal to the PL_D representation of the display to which the display-optimized image is supplied.

[0192] Returning to Figure 6, the dif value is converted to a psychovisually uniform lumen difference (Ydif) by applying the v-function of Equation 3 in the photoelectric conversion circuit 611 and dividing by PL_V_in to obtain the value dif (i.e., normalized luminance difference), which is then substituted into L_in. Here, the value of PL_V_in is PL_V_HDR when creating an ambient adjustment image with the same maximum luminance as the input master HDR image, and in the display scenario adapted to an MDR display, the value of PL_V_in is the PL_D value of the display, for example, 600 nits.

[0193] The photovoltaic conversion circuit 604 calculates the intermediate luminance Ln_im, a normalized linear version of the intermediate luminance Yim. It applies the inverse of equation 3 (the same definition of RHO). For the RHO value, in fact, the same PL_V_HDR value is used in the simplest situation where the output image has the same maximum luminance as the input image and only the luminance of darker pixels is adjusted. However, in the case of display adaptation to MDR maximum luminance, the PL_D value is used to calculate an appropriate RHO value that characterizes a particular shape (sloppiness) of the v-function. To accomplish this, an output maximum determination circuit 671 exists in the color processing circuit 600. It is a logic processor that, generally, determines whether PL_V_HDR or PL_D should be used, respectively, as the value for determining the maximum luminance of the output image PL_O, i.e., the RHO of the EOTF applied by the photovoltaic conversion circuit 604, for a given situation (those skilled in the art will understand that in some specific fixed variations, the situation may be composed of a fixed equation).

[0194] The final ambient adjustment circuit 605 performs a linear additive offset in the luminance domain by calculating the final normalized luminance Ln_f using the following formula. Ln_f=(Ln_im-(mL_De2 / PL_O)) / (1-(mL_De2 / PL_O)) [Formula 9] mL_De2 is the second minimum brightness of the end-user display (also known as the second end-user display minimum brightness) and is generally input via a third minimum metadata input 695 (connected to an illuminometer via an intermediate processing circuit). It differs from the first minimum brightness of the end-user display mL_De in that mL-De further includes the physical black characteristic value of the display (mB_fD), but mL_De does not, and only characterizes the amount of ambient light that degrades the displayed image (e.g., by reflection), i.e., typically equals only mB_sur.

[0195] Inside the final ambient control circuit 605, the mL_De2 value is normalized by the applicable PL_O value, i.e., PL_D.

[0196] Ultimately, in most variations, it is favorable when the normal (i.e., unnormalized) output luminance L_o appears, which is realized by the multiplier 606, which is calculated as follows: L_o = Ln_f * PL_O [Equation 10] In other words, the output image is normalized to the maximum applicable brightness.

[0197] Advantageously, some embodiments do this in color processing, optimizing not only for adjusting to dimmer luminances relative to ambient light conditions but also for reducing the maximum luminance of the display. In this scenario, the luminance mapping circuit 602 applies a suitable calculated display-optimized luminance mapping function FL_DA(t), which is generally loaded into the luminance mapping circuit 602 by the display optimization circuit 670. Those skilled in the art will understand that specific ways of display optimization are merely variable parts of such embodiments and are not typical for ambient adaptive elements, but several examples are shown with reference to Figures 4 and 5 (the input of configurable PL_O values ​​to 670 is not depicted in order to avoid overcomplicating Figure 6, as those skilled in the art can understand this). In general, display adaptation has the characteristic of bringing the function closer to the diagonal as the difference between the maximum input brightness and the maximum output brightness is small; that is, the closer the desired maximum brightness of the output image (i.e., PL_V_MDR=PL_D) is to the maximum input brightness (generally assuming PL_V_HDR), and therefore farther from the maximum brightness of the second reference grading (generally PL_V_SDR), the flatter the shape of the function becomes (i.e., a "lighter" regrading function version FL_DA between F_L and the diagonal). In downgrading, the function generally has a convex shape, but it should be noted that this means that darker brightness is relatively boosted (than some midpoints) at the expense of compression of brighter brightness. Thus, in downgrading, FL_DA generally has a less steep slope to boost the darkest input brightness below the F_L function (which performs a full regrading to the second reference grading at the other end of the diagonal).

[0198] A second invention for optimizing image contrast is taught below, which is useful in brighter ambient conditions. These elements may be used in conjunction with the ambient adjustments described above in various embodiments, but each invention may also be applied independently of the others.

[0199] This method of processing an input image, particularly to improve contrast, generally consists of the step of obtaining a reference luminance mapping function (F_L) associated with the input image, which defines the need for regrading by mapping luminance (or equivalent luma) to the luminance of a corresponding secondary reference image.

[0200] The input image generally also serves as the first grading reference image.

[0201] The output image generally corresponds to the midpoint of the maximum brightness of two reference grading images (also known as reference grading), and the relationship of the maximum brightness of the output image, calculated and optimized for the corresponding MDR display (intermediate dynamic range HDR display compared to the master HDR input image), is generally PL_V_MDRPL_D_MDR.

[0202] Display adaptation (of any embodiment) is applied as usual to display adaptation of lower maximum brightness displays, but here it is applied in a particularly different way (i.e., most of the technical elements of display adaptation remain the same, but some are modified).

[0203] The display adaptive processing determines an adaptive luminance mapping function (FL_DA), which is based on a reference luminance mapping function (F_L). This function F_L is static, i.e., the same for several images (in which case, for example, FL_DA still changes if ambient lighting changes significantly or during user-controlled operations), but it can also change over time (F_L(t)). The reference luminance mapping function (F_L) generally originates from the content creator, but may also originate from the optimal regrading function calculated by the automaton at the receiving device of the video image (just as in the offset determination embodiments described in Figures 6 to 8).

[0204] The adapted luminance mapping function (FL_DA) is applied to the input image pixel luminance to obtain the output luminance.

[0205] A key difference from existing display adaptations is that (on the one hand, the definition of the metric, i.e., the mathematics for finding various maximum luminances, and the direction of the metric are the same; for example, the metric is scaled by having one point at any position on the diagonal corresponding to the Yn_CC0 lumens normalized to 1, and the other point somewhere on the trajectory of the F_L function, e.g., vertically upward, so that it corresponds to the output lumens or luminance of F_L when the coordinates of the other end of the metric positioning endpoint are the input Yn_CC0 lumens to the function F_L), now the positions on the metric (more precisely, all scaled versions of its shape according to the shape of the F_L function) for obtaining the adaptive luminance mapping function (FL_DA) are calculated based on the adjusted maximum luminance value (PL_V_CO) rather than the maximum value of the required output image (generally PL_V_MDR).

[0206] This adjusted maximum brightness value (PL_V_CO) is determined by the user of the device, generally the viewer of the connected display (or the device may be inside the display). The user determines the optimal control value, i.e., the user correction value (UCBVal). The user correction value (UCBVal) generally serves to downgrade the maximum brightness used in the display adaptive algorithm, so that the algorithm no longer operates with the actual (physical) maximum brightness achievable by the connected display (for example, in a backlit LCD, the maximum light that a pixel can output is determined by setting the backlight to maximum, setting it to ensure correct operation, e.g., cooling and lifespan, as configured by the manufacturer, and controlling the liquid crystal pixels to transmit as much light as possible), but rather with the virtual user control value. Generally, the user controls this by offsetting in one direction from a starting setpoint. This setpoint is optionally influenced by a measurement of the ambient light level, e.g., the ambient illumination value (Lx_sur).

[0207] Ambient illumination values ​​(Lx_sur) can be obtained in various ways, for example, by the viewer determining it empirically by checking the visibility of a test pattern, but it is generally derived from measurements by an illuminometer 902, which is generally positioned appropriately relative to the display (e.g., on the bezel edge and facing approximately the same direction as the screen front plate, or integrated into the side of the mobile phone, etc.).

[0208] The reference ambient value GenVwLx is determined in various ways, but is generally fixed because it relates to what is expected to be reasonable ("average") ambient lighting for a typical target viewing situation.

[0209] This method may be used without a GenVwLx value, but it serves as an intermediate point for user control.

[0210] In the case of watching television on a display, this is generally referred to as a viewing room.

[0211] The actual lighting in a living room can vary considerably depending on factors such as whether the viewer is watching during the day or at night, and the room configuration in which they are watching (for example, whether there are small or large windows and how the display is positioned relative to the windows, or whether mood lighting is used at night, or whether another member of the family is doing precise work that requires a sufficient amount of light).

[0212] For example, even during the day, if a hailstorm suddenly darkens the sky considerably, outdoor lighting can be as low as 200 lux (lx), and indoors, the light level from natural lighting is generally 1 / 100th of that, so indoors it's only about 2 lx. This is where the nighttime appearance (which is particularly strange during the day) begins to emerge, and this is why many users generally turn on at least one lamp for comfort, effectively raising the level again. Normal outdoor levels range from 10,000 lx in winter to 100,000 lx in summer, so it's more than 50 times brighter.

[0213] However, other viewers may find it convenient to watch videos (especially HDR videos) in the dark, for example, to enjoy horror movies as scarier and / or to see dark scenes better.

[0214] While common in the Middle Ages, single candlelight is now considered the minimum level of illumination. This is simply because, for urban dwellers, such a level is insufficient as more light spills in from outdoor lamps like city lights. The candela used is defined as the brightness of a typical candle; therefore, placing a surface 1 meter away from the candle results in 1 lx, which is still visible, but makes reading text on paper extremely difficult (for reference, 1 lx is also typical of a moonlit scene outdoors). Thus, illuminating a 5-meter wide room with several candles achieves that level of illumination. Even a single 40W incandescent bulb already produces about 40 times the brightness of a candle, and therefore, for most viewers, one or two or three such lamps provide a more typical ambient light level. Therefore, to view with not much (atmospheric) light, one can expect something like k*10 lx. However, the video is defined to be viewable properly even during the day, in which case the lighting should be n*50 lux (for example, a set of 200W light bulbs placed about 2 meters apart will yield about 3000 / 50 lux; if you are viewing a recipe display in the kitchen, you will need a light level about three times higher to safely perform cooking tasks such as chopping).

[0215] In mobile / outdoor situations, for example, when sitting near a train window or in the shade under a tree, the light level is higher. In such cases, the light level is, for example, 1000 lux.

[0216] I don't mean to be restrictive, but I will assume that a good GenVwLx value for watching television programs on video is 100 lux.

[0217] Let's assume the light sensor measures Lx_sur = 550 lux.

[0218] Next, the user controls this by setting such a UCBVal value so that the image looks better at levels effectively five times higher than the ideal 550 lux level. This is done by the user adjusting the adjusted maximum luminance value (PL_V_CO) to such an extent that the darkest luminance subrange for a particular F_L function is extended to a visually acceptable level. This is done by calculating the value that should be output by the adjusted maximum luminance value PL_V_CO in the maximum luminance determination unit (901) as follows: PL_V_CO=PL_V_MDR-UCBVal [Formula 11] PL_V_MDR is generally (though not always) equal to the maximum displayable pixel brightness of the connected display.

[0219] The user-controlled value UCBVal is controlled by appropriately scaled user input, for example, a slider setting (UCBSliVal, e.g., symmetrically decreasing near zero correction or starting at zero correction) is scaled in such a way that when the slider is at its maximum, the user does not drastically alter the contrast, for example, so that all dark image areas appear almost like bright HDR white.

[0220] Therefore, the manufacturer of the device (e.g., a display) pre-designs an appropriate intensity value (EXCS), and at that time, the formula is: UCBVal=PL_D*UCBSliVal*EXCS [Formula 12] That is the case.

[0221] For example, if you want to make 100% correspond to an additional 10% change in maximum brightness base contrast, you get the following: PL_D*1*EXCS=0.1*PL_D, therefore EXCS=0.1 etc (a value of 0.75 was found to be good in certain embodiments).

[0222] Figure 9 shows possible embodiment elements in a typical device configuration.

[0223] A device 900 for processing an input image to obtain an output image (also known as an ambient display optimization device) has a data input unit (920) for receiving a reference luminance mapping function (F_L), which is metadata associated with the input image. This function again specifies the relationship between the luminance of a first reference image and the luminance of a second reference image. These two images are again, in some embodiments, graded by the video creator and communicated together with the video itself as metadata, for example, via satellite television broadcasting. However, the appropriate regrading luminance mapping function F_L is also determined by the receiving device, for example, a regrading automaton in a television display. The function changes over time (F_L(t)).

[0224] The input image generally serves as the first reference grading, and based on it, the display-optimized image and the environment-optimized image are determined, which are generally HDR images. The maximum brightness of the display-adapted output image (i.e., output maximum brightness PL_V_MDR) generally falls between the maximum brightness of the two reference images.

[0225] The device optionally includes, or is equivalent to, an illuminometer (902) configured to determine the amount of ambient light hitting a display, i.e., a display on which images optimized for viewing are supplied. However, this is generally not required for current user control.

[0226] The display adaptive circuit 510 is configured to determine an adaptive luminance mapping function (FL_DA), which is based on a reference luminance mapping function (F_L). It also actually performs pixel color processing and therefore includes a luminance mapper (915) similar to the color converter 202 described above. Color processing is also involved. The configuration processor 511 makes the actual determination of the (ambient optimized) luminance mapping function to be used before performing pixel-by-pixel processing of the current image. Such an input image (513) is received via an image input unit (921), for example, an IC pin, and the image input unit (921) itself is connected to an image source to the device, for example, an HDMI® cable, etc.

[0227] The adaptive luminance mapping function (FL_DA) is determined based on the reference luminance mapping function F_L and the maximum luminance value, following a variation of the display adaptive algorithm described above. Here, the maximum luminance is not the typical maximum luminance (PL_D) of the connected display, but a specially adjusted maximum luminance value (PL_V_CO) that is adjusted to the light intensity of the viewing environment (and potentially further user adjustments). The luminance mapper applies the adaptive luminance mapping function (FL_DA) to the input pixel luminance to obtain the output luminance.

[0228] To calculate the adjusted maximum luminance value (PL_V_CO), the device includes a maximum luminance determination unit (901), which is connected to a display adaptation circuit 510 and supplies this adjusted maximum luminance value (PL_V_CO) to the display adaptation circuit (510).

[0229] This maximum brightness determination unit (901) obtains a reference illumination value (GenVwLx) from memory 905 (for example, this value may be pre-stored by the device manufacturer, or be selectable based on the type of incoming image, or be loaded along with typical intended ambient metadata that communicates with the image, etc.). It further obtains the maximum brightness (PL_V_MDR), which may be a fixed value stored in the display, for example, or be configurable in a device (e.g., a set-top box or other image preprocessing device) that can supply images to various displays.

[0230] In some embodiments, the user (viewer) controls the automatic ambient optimization of the display adaptation according to the user's preference by using, for example, a user interface control component 904, such as a slider (or rotary knob, which does not have to be a physically present button, but rather a finger-controllable element on the screen of a mobile phone that controls a device), which allows the user to set a higher or lower value, for example, the slider setting UCBLiVal. This value is input to a user value circuit 903, which communicates the user control value UCBVal to the maximum brightness determination unit 901.

[0231] Figure 10 shows an example of processing with a specific luminance mapping function F_L. We assume the reference mapping function (1000) is a simple luminance identity transformation, i.e., clipping above a maximum display capability of 600 nits. The luminance mapping is represented here by a plot of equivalent psychovisually homogenized luma, which can be calculated according to Equation 3. The input normalized luma Yn_i corresponds to the HDR input luminance, which we assume is a 1000-nit maximum luminance HDR image (i.e., the RHO of PL_V_HDR=1000nit is used in the equation). For the output normalized luma Yn_o, we assume an exemplary display of 600 nits, and therefore the output normalized luma Yn_o is re-converted to luminance by using the RHO value corresponding to 600 nits. For convenience, the luminances corresponding to the luma positions are shown on the right and top. Thus, this selected F_L reference mapping function 1000 performs equicompression in the visually homogenized luma domain. The (ambient) adaptive luminance mapping function FL_DA is shown as curve 1001. On the one hand, an offset Bko is observed, which depends particularly on the display's leak black. On the other hand, a curvature is observed that primarily boosts the darkest blacks, which is due to the remaining effects of processing in the non-linear psychovisually uniform luma domain.

[0232] Figure 11 shows an example of an embodiment of contrast boosting (the ambient light offset is set to zero, but as described above, both processes are combined). When modifying the luminance mapping in a relatively normalized psychovisually uniform luma domain, this relative domain starts at a value of zero. The master HDR image starts at a certain black value, but at low output, a black offset is used to map this to zero (or to the minimum black of the display; note: displays may have various behaviors for darker inputs, e.g., clipping, and therefore a true zero can be placed there).

[0233] When using modified display adaptation, a black zero input generally maps to black zero, regardless of the value of the adjusted maximum luminance value (PL_V_CO). Generally, the zero point of the output luma starts at the virtual black level mL_VD (ideally) or the minimum black mB_fD of the actual end-user display, but in either case, there is a difference in brightness at the darkest colors, which results in better visible contrast for dark pictures such as night scenes (the input luminance histogram 1010 is stretched as the output luminance histogram 1011 on the normalized luma axis of the output luma, which gives sufficient image contrast despite the lower PL_V_MDR=PL_D=600nit). When the two methods are combined, all processing to obtain a suitable black level offset Bko is best moved to the processing described in Figure 6, etc. (i.e., adjusted contrast processing can be performed using a zero-starting image and a zero-starting function). In practice, since display adaptive luminance (luma) mapping simply occurs regardless of whether zero actually occurs in the input image, we always simply define an F_L curve starting from zero (mapping zero HDR luminance or luma to zero output luminance or luma).

[0234] The algorithmic components disclosed herein are actually implemented (in whole or in part) as hardware (e.g., as part of an application-specific IC) or as software running on a dedicated digital signal processor or general-purpose processor.

[0235] It should be clear from the inventors' presentation to those skilled in the art which components are optional improvements and how they are realized in combination with other components, and how the (optional) steps of the method correspond to each means of the apparatus, and vice versa. In this application, the term “apparatus” is used in its broadest sense, namely a set of means that enable the achievement of a particular purpose, and therefore includes, for example, an IC (or a small circuit portion thereof), or a dedicated device (such as a device with a display), or part of a networked system. “Arrangement” is also intended to be used in its broadest sense, and therefore includes, among other things, a single device, a part of an apparatus, a collection of cooperating devices (or parts thereof).

[0236] The meaning of a computer program product should be understood as encompassing any physical realization of a set of commands that enables a general-purpose or dedicated processor to input commands into the processor after a series of loading steps (including intermediate translation steps, such as translation to an intermediate language and a final processor language) and to execute any of the characteristic functions of the invention. In particular, a computer program product can be realized as data on a carrier such as a disk or tape, data in memory, data traveling over a network connection (wired or wireless), or program code on paper. Apart from the program code, characteristic data required by the program may also be embodied as a computer program product.

[0237] Some of the steps required for the operation of the method may already exist in the processor's functions, instead of being described in a computer program product such as data input and output steps.

[0238] It should be noted that the embodiments described above are illustrative, not limiting, of the invention. For the sake of brevity, not all of these options are described in detail where those skilled in the art can easily map the presented examples to other areas of the claims. Apart from the combinations of elements of the invention as combined in the claims, other combinations of elements are possible. Any combination of elements may be realized by a single dedicated element.

[0239] Any reference numerals in parentheses in a claim are not intended to limit the claim. The words “includes,” “equipment,” and “possess” do not preclude the existence of elements or aspects not enumerated in the claim. A singular element does not preclude the existence of multiple such elements.

Claims

1. A method for processing an input image in order to obtain an output image, The input image has pixels having an input brightness within a first brightness dynamic range having a first maximum brightness, The reference luminance mapping function is received as metadata associated with the input image. The aforementioned reference luminance mapping function specifies the relationship between the luminances of pixels placed in two images, and the two images are graded differently in such a way that the pixel luminances of the same image objects have different pixel luminances in the two images. The aforementioned reference luminance mapping function specifies the relationship between the luminance of the input image and the luminance of a secondary reference image having a second reference maximum luminance. The output image has an output maximum brightness that is different from the first maximum brightness and the second reference maximum brightness. The aforementioned process is, A step of determining an adaptive luminance mapping function based on the reference luminance mapping function and the adjusted maximum luminance value, wherein the adjusted maximum luminance value is different from the output maximum luminance. The process includes the step of applying the adaptive brightness mapping function to the input pixel brightness in order to obtain the output brightness of the output image, The calculation of the adaptive luminance mapping function includes finding a position on a metric that specifies the position of maximum luminance, where the position corresponds to the adjusted maximum luminance value. The first endpoint of the metric corresponds to the first maximum brightness, and the second endpoint of the metric corresponds to the second reference maximum brightness. The first endpoint of the metric is located at a point on the diagonal having horizontal and vertical coordinates equal to the normalized input luma for any normalized input luma. In a method in which the second endpoint is positioned on the trajectory of the reference luminance mapping function determined by the direction of the metric, The adjusted maximum brightness value is The adjustment of the maximum brightness value is determined by obtaining a user-controlled value and subtracting the user-controlled value from the maximum output brightness. The output brightness is written to the output image as the pixel color, and the output image is output. A method for processing input images, characterized by the following:

2. The method for processing an input image according to claim 1, wherein the maximum output brightness is set to be equal to the maximum displayable pixel brightness of a display on which the output image can be supplied.

3. Steps include obtaining input values ​​from the user of the display, The steps include: calculating the user control value by multiplying the input value by the intensity value and the maximum output brightness; A method for processing an input image according to claim 1 or 2, comprising:

4. A method for processing an input image according to claim 1, wherein, in addition to the step of setting the darkest black in the input image as a function of ambient illumination values ​​as the black offset value, the processing described above is performed.

5. A method for processing an input image according to claim 1, wherein the direction of the metric is pre-set as vertical, and the metric places the second endpoint with respect to the normalized input luma at a location having the result of applying the reference luminance mapping function with the normalized input luma as the horizontal coordinate and the normalized input luma as the vertical coordinate.

6. A device for processing an input image in order to obtain an output image, The input image has pixels having input brightness within a first brightness dynamic range, the first brightness dynamic range has a first maximum brightness, and the device is An image input unit for acquiring the aforementioned input image, It includes a data input unit for receiving a reference luminance mapping function, which is metadata related to the input image, The aforementioned reference luminance mapping function specifies the relationship between the luminance of the first reference image and the luminance of the second reference image, The reference luminance mapping function specifies the relationship between the luminance of the input image and the luminance of the second reference image, The second reference image has a second reference maximum brightness, The output image has an output maximum brightness that is different from the first maximum brightness and the second reference maximum brightness. The aforementioned device A user value circuit for determining and outputting user-controlled values ​​as set by a human user of the device, The user control value is obtained from the user value circuit, The adjusted maximum brightness value is output as a result of subtracting the user-controlled value from the maximum output brightness. Maximum brightness determination unit for It further includes, The aforementioned device The display further includes a display adaptive circuit for determining an adaptive luminance mapping function based on the adjusted maximum luminance value and the reference luminance mapping function, The calculation of the adaptive luminance mapping function includes finding the position on the metric corresponding to the adjusted maximum luminance value, The first endpoint of the metric corresponds to the first maximum brightness, and the second endpoint of the metric corresponds to the maximum brightness of the second reference image. The first endpoint of the metric is located at a point on the diagonal having horizontal and vertical coordinates equal to the normalized input luma for any normalized input luma. The second endpoint is positioned on the trajectory of the reference luminance mapping function, A device for processing an input image, wherein a display adaptive unit applies the adaptive luminance mapping function to the input luminance, or the input luma encoded from the input luminance, to obtain output luminance or output luma, and outputs these output luma or output luminance as the pixel color of the output image.

7. The user value circuit is Obtain input values ​​from the user on the display, The user control value is calculated by multiplying the input value by the intensity value and the maximum output brightness. The apparatus for processing an input image according to claim 6, wherein the user-controlled value is output to the maximum brightness determination unit.

8. An apparatus for processing an input image according to claim 6 or 7, comprising a black adaptive circuit for setting the darkest black luma in the input image as a function of ambient illumination values ​​to a black offset luma.

Citation Information

Patent Citations

  • Conversion method and conversion device

    JP2017050840A

  • Video processing device, display device, video processing method, control program, and recording medium

    JP2019041269A

  • Systems and methods for adjusting video processing curves for high dynamic range images

    JP2020502707A

  • Adjustment of display optimization behaviour for HDR images

    WO2021004839A1