Tone mapping for out-of-HDR gamut adjustment
By defining the maximum brightness of the target display in the high dynamic range image and optimizing the brightness of the image objects, the problem of image quality degradation in the existing technology is solved, and reasonable approximation and visual effects are achieved when displayed on different displays.
Patent Information
- Application Number
- CN202480011425.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-08
- Filing Date
- 2024-01-29
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies have difficulty in effectively processing brightness variations in high dynamic range images, resulting in a decrease in image quality when downgrading to lower maximum brightness output images.
The absolute method maintains image quality during encoding and display by defining the maximum luminance of the target display and optimizing the luminance of image objects to meet the luminance range of the target display.
This method achieves the goal of maintaining reasonable image approximation and visual effects when displaying high dynamic range images on different displays, avoiding the problem of being too bright or too dark.
Smart Images

Figure CN120660341A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to brightness-varying color mapping for high dynamic range images, in particular for downgrading to lower maximum brightness output images. The technique can be used in HDR image or video encoders. Background Art
[0002] Figure 1 The creation, encoding, and display of images according to the so-called Low Dynamic Range (LDR) (also known as the Standard Dynamic Range (SDR) video specification) are shown. This technology has been used in electronic television broadcasting (and similar technologies such as Blu-ray Discs) since about the 1940s. It was technically developed based on the technical possibilities of the time. Although it has its limitations, the technology is considered sufficient to transmit everything from royal marriages to the moon landing. Some aspects related to this innovation are described, and many other aspects (such as how to digitize analog video signals and compress them according to the MPEG standard) are left to the domain of video technicians.
[0003] The original analog SDR video representation systems (NTSC or PAL) were developed around cathode ray tube displays (CRTs) and "reverse" CRT cameras (other camera technologies such as CCD and CMOS, and other display technologies such as plasma displays and liquid crystal displays (LCDs) were later developed, but these were technically configured to behave like CRTs to match the SDR video encoding standards, which remained unchanged).
[0004] The camera 100 transmits a voltage signal (V) to a television display 150 in a consumer's home, for example, via a broadcast tower. This voltage signal drives a cathode 151, which emits a certain number of electrons in an electron beam 152, depending on the control voltage applied to the cathode. The time-varying voltage is deflected by a coil, causing it to scan sequential positions on the display's screen in a zigzag pattern until the entire image has been generated and it is time to fly back to the screen's origin position to display the next image in the video sequence. For any scan position at a certain electron beam deflection angle, the corresponding red, green, or blue phosphor on the screen is excited by the electrons and emits a number of photons that is roughly linearly proportional to the number of impacting electrons (which is controlled by the signal voltage V, which defines the image content). Because the cathode is a quantum mechanical object, the number of electrons emitted by it is not linearly proportional to the drive voltage. Its inherent physical behavior can be modeled by a power function that approximates a square power. This means that if our voltage at time t2 is to encode an object that should appear twice as bright on the screen as another object corresponding to the position ("pixel") or time t1, we should not have V2=2*V1, but rather V2=sqrt(2)*V1, where sqrt is the inverse of the square power, i.e. the square root.
[0005] Since all CRT displays for all consumers behave like this, the broadcaster's camera will be responsible for square root pre-compensation.
[0006] For example, when lens 101 images scene light onto a pixel of a CMOS-type sensor 102, such a CMOS sensor is typically substantially linear. If twice the number of scene photons falls on a pixel, the pixel voltage will be twice as high (or the digital number output from an analog-to-digital converter operating at that voltage will also be twice as high). Therefore, to generate the (digital or analog) output signal V, the camera uses a signal converter 110 that applies a square root function.
[0007] This creates a luminance that encodes the brightness of objects in the scene. relatively (Scene-referenced) systems, which are technically required to display a good-looking representative image to the human eye and brain. For example, in the PAL standard, brightness (or more correctly, the square root of the relative measure of the amount of light from a scene object) is a voltage value between essentially 0 (for black) and 700mV (for the brightest possible brightness, which is perceived and understood as a clean white object). The value 700mV encoding a white object, or the luma (Y) code 255 in the digital era, is known as 100% brightness (or 1.0).
[0008] SDR systems have a definition that allows for a luminance dynamic range of about 1000:1 (i.e., the brightest displayed pixel to the darkest). In analog systems, this is primarily due to the noise on the voltage, but in a typical 8-bit digital representation, such as in the standardized Rec.709 codec, there are only codes for encoding luminance or grayscale values between 1 and 1 / 1000. This corresponds to and is sufficient for typical displays, which—depending on the ambient lighting such as daytime viewing—can be lower than 50:1, but On average As you can see, with SDR technology, viewers can see display pixels between approximately 1 nit black and 100 nit white.
[0009] Figure 1A shows what usually happens from the perspective of the signal data. Scene objects are illuminated by a usually uncontrollable (and unknown at display time) amount of scene illumination, which should usually only comply with the requirement that there should be enough light. In a studio one usually hangs a light grid from the ceiling, which creates a reasonably uniform illumination (base illumination) at all important points in the scene. So, for example, a dark black shoe standing at the back of the scene being captured will be illuminated by roughly the same amount of illumination (IL) as a white sheet of paper on a table in the middle of the scene. Beyond that, one is able to create some dedicated light effects to put, for example, a shiny object in the spotlight, but it may quickly happen that parts of the object become so bright that any further brightness increase is clipped to the same value in the captured signal (for bright scene pixels, just as for even brighter scene pixels, the pixel wells fill to their maximum capacity, the ADC outputs its maximum code, e.g. power (2; 14), and the rec The Y of the 709YCbCr pixel color digital code will be 255. In any case, apart from a little highlight and a little dark shadow, the SDR signal will mainly encode the object's own color (under uniform lighting). This own color, that is, the amount of red, green and blue, or the total amount of light reflected by any object (for example, 90% or 0.9 for white paper), will result in the scene luminance Lsc of the object (for example, the scene luminance of white paper LSc_pw or black shoes Lsc_bs, which typically also reflect 10% or less of the incident light towards the camera lens), and ultimately, in a linear manner, a certain amount of photons falling into the pixel imaging that area of the scene, and therefore ultimately the luminance Y, which is the square root of the relative brightness measurement of the pixel (and sqrt also happens to be a reasonable first approximation of how humans perceive the brightness of objects).
[0010] Another characteristic of SDR imaging technology is that scene brightness is simply irrelevant and does not play a role. Not only is a billion nits of the sun's disk not a reasonable pixel brightness for what we would want to see on screen when watching, for example, a movie or sports program, but all the settings of the camera (such as the iris, shutter, amplifier, neutral density filter, and potentially non-linear modifiers such as the camera function knee) make the relative brightness output signal, i.e., luminance Y in digital SDR video, not easily convertible to scene brightness.
[0011] But importantly (apart from minor gamma corrections due to differences in viewing surroundings between the viewing room and the original scene), the relative scale of the luminances (Ld) of the displayed pixels will be essentially equal to the scale in the scene (apart from clipping etc.) because the sqrt nonlinearity is removed by its inverse square power.
[0012] This means that the adapted human eye and brain will see a reasonable approximation of the original scene, at least seeing objects appear, for example, as a middle grey between black and white objects, as it has about 1 / 4 the relative brightness or luminance of white (or 100%).
[0013] The SDR imaging chain (i.e., capture until display) is relative on the other hand. The brightness code is used as a drive or control signal for the display, so the maximum controllable value (Y=255) drives the display to the maximum value it can (will) display, whatever this is. On average, Y=255 can be displayed as Ld=100nit. However, if people buy a brighter display, it can easily be shown as 200nit (or even up to 500nit for the brightest displays). In addition, there is usually a brightness control knob 170. Even if some consumer buys a TV with 300nit capability, he can dim it with the knob so that the brightest object (defined as Y=255 in the image) will not be displayed brighter than 100nit anyway. Therefore, it is not defined what display brightness Ld the image color codes (especially the brightness Y) will produce, and they do not directly correspond to the scene brightness. However, in relative view, things work like this: the so-called white-on-white principle maps the "white" in the scene (e.g., a bridal gown) to a white code (Y=255), and then at the brightest pixel any consumer-side display will display that (as mentioned before, this may be 100 nits for one consumer and 80 for another, but the adaptation properties of human vision will make the two images look essentially the same, at least when viewed by themselves as they would normally, rather than side by side like in a store).
[0014] However, as newly invented displays (for starters LED matrix backlit LCD displays) had the potential to become much brighter, it became clear that displaying SDR 100% white at 700nit maximum display brightness was no longer a reasonable display of an SDR image (and under certain conditions some customers might consider this exaggeratedly bright), or better stated, a simple SDR signal would no longer be the best image that could be displayed on such high performance displays.
[0015] One might aim to better try to display not only the native colors of objects under essentially uniform lighting, but also natural scenes where there are fairly dark corners and fairly bright areas (e.g., a sunlight scene outside a window, which is typically clipped to an indistinguishable white in SDR, or bright objects such as light bulbs, more realistic specular reflections on metal, etc.).
[0016] The simplest definition of a high dynamic range (HDR) image is any image capable of encoding something 100% brighter, such as 400% or even 1000% (where the system can make some pixels appear 10 times brighter than paper white). A TV doesn't yet know how to handle such new signals in a unique or standardized way, but it can display, for example, 1000% as its maximum displayable brightness (ML_D) of 700 nits, and then other brightness levels (while still following the linear white-on-white approach described above) will fall off to a similar or reasonable approximation. For example, since white paper will appear at 70 nits, which is close to the approximately 100 nits it will appear on an average SDR display. Thus, normal diffuse objects in the scene will appear similar to how they would on an SDR display (which can't display 10 times brighter pixels), while also fitting in reasonably well with brighter display pixels, such as a 700 nit light bulb hanging from the ceiling (as mentioned, the original scene brightness, Lsc, of such a bulb is less important than achieving a natural, brighter-than-average appearance in the displayed HDR image).
[0017] This therefore appears to be a fairly sound technical approach for designing a novel HDR coding framework (although there happen to be some caveats).
[0018] This is exactly how one of the current HDR video codecs on the market, the BBC's Hybrid Loggamma (HLG) codec, defines its HDR approach. They extend the square root function beyond 1.0 (while going from 8 to 10 bits to have enough code points along the extended photoelectric transfer curve that defines how relative brightness from the camera is encoded as corresponding luminance), although not with a continuous square root definition, but instead becoming logarithmic in shape above the halfway point (Y=0.5*1023). This has the advantage that the darkest colors in the image behave exactly as expected, i.e., as under SDR imaging, because the display applies a square power to those luminances, so relative display luminances, and therefore the brightness or grayscale of objects, still behave conservatively, as seen by Figure 1 As explained by A. For brighter scenes or image objects, the relationship between object brightness is distorted logarithmically, but this is needed to compress the larger HDR dynamic range into a smaller range, suitable for direct (i.e., without brightness optimization TV internal processing) display on a lower dynamic range display (such as a traditional LDR display). In addition to being a natural method for HDR, it is a very simple way to create a slightly more impressive dynamic range extended image, but it has its limitations.
[0019] One of the problems lies in the fact that such a completely relative system will still display the brightest coded pixel (which, as mentioned, represents no actual brightness, although it can be pretended that it corresponds to, for example, an approximately desired estimated 1000 nit) as anything that any display will be able to display. So, it may be fine for a 700 nit ML_D capable display, or a 1000 nit display, or even a 2000 nit display, but on, for example, a future 6000 nit display, some images with a Y=1023 maximum brightness code may look too bright. The solution for the HLG framework is to apply some corrective gamma to brighter displays, which at least reduces the mid-level brightness, although it still displays the brightest pixel quite arbitrarily.
[0020] In contrast, scene reference technology proponents argue that any absolute system will have at least two problems. On the one hand, images can be generated that contain much brighter expected pixels (for example, the video creator defines that some pixels in the video should ideally be displayed as bright as 10,000 nit), and many actual consumer displays can display these pixels (for example, ML_D = 500 or 1500 nit displays). On the other hand, the viewing environment can be highly variable (although you may decide to watch a movie under average night lighting, rather than with sun beams directly hitting the TV screen), and an absolute system will not take into account visual adaptation and will also need to post-process the input HDR video for this aspect. But HLG will also need to compensate (still only has one hardware capability HLG-understanding displays and variable viewing environments), and at least by making everything clear and absolute, it is considered in a professional rather than a fast way, thereby implementing any needs.
[0021] Therefore, all other HDR video codec developers, starting with Philips-Technicolor and Dolby, then Samsung's HDR10+ and others, Absolute System The invention works in a similar way to the previous one (working based on defining pixel brightness rather than relative luminance encoding, which is a major paradigm shift in video technology), which will now be described. Although the following innovations may also be useful for relative systems, we will illustrate them in terms of their absolute variants, for which they are particularly useful.
[0022] A modern, future-proof HDR codec framework requires many innovative technical insights, where one change necessitates another, and will now be achieved with the help of Figure 2 Summarize some core aspects.
[0023] The absolute approach is not just about defining "just any brightness" for various pixels, as this still cannot be stably defined. The importance of a stable definition can be understood, as there can be many components in the system that contribute to pixel brightness, and this set of sliding scales makes it impossible to guarantee that the end consumer will see even a vague approximation of the beautiful image that the content creator made and intended everyone to see. For example, the creator can make a flashlight beam of a certain brightness in order to search for a monster in shadows with a certain darkness, but if for example the consumer's TV changes the brightness of the monster and the beam, then anything can come out and the movie may become different. As Figure 2 As shown, several intermediate conversions can be required in at least some market deployments, it is important to define things well, starting with a good first line definition.
[0024] The first line of the stable definition is to set the target display (which has Target The maximum brightness of the display (ML_T) is associated with the video, and a video is created that is tuned to the target display. Choosing a well-functioning ML_T can be compared to selecting an aspect ratio for optimizing a geometric composition. Suppose one wants to depict a still image of life, such as oranges arranged around a vase. An aspect ratio of approximately 1:1 would be chosen. If the oranges had to be positioned around the vase on a 3:1 canvas, they would have to be spread out linearly, which wouldn't show the same composition, or half the canvas would be left blank, which is also suboptimal. On the other hand, when depicting a beautiful landscape, a 1:1 ratio would represent only a disappointingly narrow cutaway of the landscape. Importantly, after the basic selection of the canvas, the exact placement of the oranges around the canvas would still be optimized—slightly to the right or forward, or how the landscape precisely covers the 3:1 canvas.
[0025] The same happens in HDR image generation (in cases where we don't want to restrict, we will call host An HDR image, which is the first image gradation created in cases where at a certain time there may be several different gradations of any image of a scene (SCN); grading means determining pixel colors, and in particular their satisfactory brightness), but then there is the image pixel brightness. The creator will choose for example a 4000nit ML_T target display, because the video he is about to make contains an expressive set of exploding fireballs somewhere in the story. According to for example his artistic preferences, he judges that any pixel brightness below about 90%*4000nit (a reduction in brightness for making yellow) will look too weak. This is still a blowup, and on a less capable display one would have to do so, but this does not mean that one has to define the created image itself with a lower HDR quality. 6000nit ML_D may be too high for such a scene, or unnecessary. After this, having chosen the target display, the actual video image will be optimization To correspond to and specifically for this target display (the name target points to the actual end-user display that ideally the end consumer should buy that has such a display maximum brightness ML_D, although of course the consumer will not buy a different dedicated TV for each different ML_T video, although one can buy a very high ML_D display that can handle most videos; a 7000nit ML_D display could then for example not use the 4000-7000nit range and display 4000nit ML_D video at such brightness, or it could slightly boost e.g. 3500-4000nit incoming main HDR video image pixels so that they appear e.g. in the range of up to 5000nit). When producing video for a 4000nit ML_T target display, one would expect that no pixel in any image would be brighter than 4000nit. However, many or all pixels in many images can be below 4000nit. But it is also expected that if the video is required to be a strong 4000nit HDR video, there will be at least some objects with pixel brightness close to 4000nit in some images (unless it is a video of only dark scenes to be merged with other 4000nit videos).
[0026] Not only will the maximum brightness be determined, but all pixel brightnesses will be optimized, i.e. determined to be at least satisfactory. That is, for each image object in the image of each scene, a brightness value between 0 and 4000 nit must be selected for the object pixel (or for a dark-understood HDR system, one will also think of a minimum black, i.e., e.g., 0.01 nit is practical for home TV viewing, but we will keep the explanation simple by focusing on the bright end). Visually, the impact of the fireball depends on the brightness of the surrounding pixels, and for example, the visibility of dark shadow areas depends on the fireball, its average brightness, spatial extent, and duration, etc. (without light there is no dark, and without dark there is no light). Thus, for example, the brightness of a monster hidden in the dark in a scene after an explosion may and will usually depend on the fireball, and therefore be specified consistently.
[0027] Figure 2A shows a typical image processing result, technically on an axis of possible brightness of some key objects in an exemplary HDR scene image (a dark shoe, a piece of white paper and an ultra-bright object above white as a light bulb). What happens in detail depends on the specific technical embodiment of the video technology (e.g. pre-recorded material unicasted over a dedicated internet connection, etc.), but we have generally illustrated the principles of the extensions required to understand the innovation. Typically, a (typically HDR) camera 201 captures real-world objects (204) in a scene (we ignore for the moment the inclusion of computer-generated image material). The real-world brightness of the scene being interpreted is irrelevant, except for the fact that relative brightness for a master HDR image (Mstr) will be derived for them (this can happen by a simple function similar to broadcast, but in offline film production the relationship may be looser by local adjustments of the image pixels of some objects, to the extent that the original captured image may only be used for the geometric positioning of the image, which results in a completely new colorimetry). Therefore, although complex gradation processing may be involved, for the sake of simplicity of explanation we will assume that any original encoding of the captured image values is transformed by one or more functions Mgrad. For example, R_mgrad=F1(R_cap); G_mgrad=F2(G_cap); B_mgrad=F3(B_cap), and the YCbCr value corresponding to the RGB pixel color value can be calculated according to the classic equation.
[0028] Therefore, what is important is the brightness of the master binning created for all image objects (we will initially focus our explanation on this primary aspect of HDR, namely brightness, expressed as luminance for example). This master binning video has an associated target display maximum brightness, in this example only, the master binning maximum brightness ML_m, taken as 5000 nits. The exemplary brightness of the light bulbs in those scene images is not shown to keep the drawing manageable, but can be set to, for example, Lm_bu = 4500 nits. The paper in the master binning can have an average pixel brightness of Lm_pw = 150 nits, etc.
[0029] The next step is to choose a primary electro-optical transfer function for encoding those luminances (for clarity of technical explanation, we start with luminance) into actual, practical pixel color codes. Digital images compressed with, for example, MPEG-VVC typically expect a YCbCr color encoding, where Y is the luminance code, for example 10 bits, i.e., ranging between 0 and 1023 as the maximum possible code, and Cb and Cr are the chrominance components to give a non-color pixel of a certain brightness a certain color (e.g., pale green, also known as "mint").
[0030] Typically, primary pixel luma may be mapped to corresponding primary luminance (YM) using the perceptual quantizer (PQ) EOTF standardized in SMPTE 2084. The same is true for chroma via the chroma equation.
[0031] For the simplest HDR video communication system, this is about all that is needed: an image of an object with various HDR brightnesses is specified, and it can be encoded as a 10-bit YCbCr color code.
[0032] More advanced, future-proof HDR video coding will use further principles, in particular the actual transmitted image (Im_c) need not be identical to the master image, in particular its transmitted image maximum luminance (ML_com) need not be identical, and therefore all object luminances will also derive different values (transmitted luminance L_com) from the master luminance (L_m). For further details on the applicant's coding (which can work specifically with a 100nit ML_com video function, which serves as a proxy for the master graded image to be transmitted), we can (without wishing to limit the following embodiments) point the reader to WO2017157977 (incorporated by reference). An example of a system for transmitting a further image using a typical minimum dynamic range version that is not equal to the original master grade can be found in WO2016020189 (incorporated by reference).
[0033] Regarding hardware technology (devices), typically the camera is connected to a grading device 202, which includes circuits for performing color processing for grading the main luminance and chrominance components (i.e., the main HDR image object colors) as needed to produce a beautiful HDR program. For an offline Hollywood movie, the color grader may spend several days to produce the master graded movie and store it in some memory for later distribution. For example, he may be watching one or more displays (2020, 2021), such as an HDR display for viewing a master HDR video image with an SDR degraded version side by side on a traditional SDR monitor. The user input module (222) that enables the grader to accurately change the color (i.e., determine the shape of the grading function Mgrad or multiple functions) may include, for example, a dedicated grading keyboard with a trackball or the like. Note that, in general, even if only a single function is determined for mapping all possible input luminances (or corresponding luminances, if the function is defined in the luminance domain) to corresponding output luminances (or luminances), there will be the possibility of shaping the function differently for at least some of each temporally consecutive video image, as the optimal mapping will depend on the type of scene (for example, a cave can be filmed in a reasonably large amount of light so as not to capture too much noise, but then graded well in the primary grading for the visual impact of the film). Similarly, the luminance mapping function FL_cod used for any secondary grading (such as the luminance mapping function used in this example as a proxy image to convey the color of the original primary grading) can also change for several images in the video. If we illustrate such a general system for real-life broadcasts of, for example, sporting events, a much less detailed grading will typically be involved. However, the shader will still verify and accept at least a simple mapping function from the original capture so that the HDR image does not look too bad. Typically, several dials may be involved to set mid-tone contrast, black level, soft clipping behavior of highlights (knee point), etc. In this case, the staging apparatus 202 may reside, for example, in a production truck at the filming location, or in a remote studio that receives contributed raw filming material, for example, via the Internet.
[0034] When the main images (if those images are transmitted themselves) or the corresponding images to be transmitted (Im_c) have been created to satisfaction, a distribution encoder 203 typically encodes and compresses them into any desired format and formats them for output over some video communication medium. In a simple explanation, we have shown the creation of an internet version of, for example, a video signal S1_inet (Im_c; MET), which can be, for example, HDR10+ or SLHDR-1 encoded. Assume that it is a SLHDR_1 encoding. The so-called image essence, that is, a pixelated image of, for example, 4000x 2000 YCbCr color triplets, for example HEVC compressed, will be a 100nit SDR image. The metadata will firstly be a different brightness mapping function for each video image (or in fact SLHDR-1 applies various coarse and fine mapping functions to any image at any moment in the video, but this is too deep a detail for this explanation), and on the other hand a specification of the color mapping required for that image. The distribution encoder can also produce a secondary distribution version, which can be an identical SLHDR-1 signal, or, in this example, an HLG image (a second transmitted image Im2_c, which in this example will have luminance defined by the HLG OETF). In this case, the second metadata MET2 is much less complex, namely, a code that simply specifies the YCbCr color code that the HLG OETF has used to create the second transmitted image. This constitutes the bird's-eye view and generally constitutes the encoding or transmission side system 200. The receiving side device 240 can also have various practical embodiments, depending on whether it is, for example, in a movie theater or in a consumer's room. Although some technical circuitry may reside in, for example, a set-top box or computer, for simplicity, we assume that the receiving side or end-user device is an HDR television display. It typically has a receiver 250 capable of performing one or more HDR functions, typically data management, such as selecting the appropriate image essence and metadata from the signal and sending it to the color processing circuit 251. It may ignore other metadata. Although in a real device some technical functions may be combined in a single dedicated ASIC or run on a general-purpose or dedicated processor (e.g. inside a mobile phone), we illustrate the typical processing steps in a receiving-side device, which decodes the received HDR signal (e.g. the broadcast signal S2_bro (Im2_c; MET2)), i.e. the decoding of the image essence which, after being parsed by the receiver, is not only usually compressed but also has erroneous absolute values or relative positions of the encoded luminance (or color) on an axis that usually does not end at the same maximum value as the maximum luminance (ML_ed) of the actual end-user display. Although three-dimensional processing may occur, the received transmitted luminance Yc and transmitted chrominance components Ci are usually processed in separate dedicated circuits, i.e. the luminance processing circuit 252 and the color processing circuit 253.For example, the luminance processing circuit will apply (directly or after inversion depending on the HDR video definition and the type of communication) the mapping function of the image to be processed received in the metadata (MET) to transform the transmitted luminance into the reconstructed luminance: YR=INV_FL_cod(Yc) for any Yc of any scanned pixel of the image reconstructed in the decoding. This is in. Figure 2 A is shown to produce a reconstructed image (Recs). If the TV is only capable of displaying pixels as bright as up to 750 nit (i.e. the end user display maximum brightness ML_ed=750 nit), then of course a reconstructed image with pixels as bright as 5000 nit (reconstructed image maximum brightness ML_rec) still cannot be displayed directly. Therefore, the display adaptation circuit 254 will apply the display adaptation function FL_DAP, which can typically be a scaled-down version of the function FL_cod to take into account the lower maximum brightness of the user's TV, and which may also include viewing environment adaptation. This process ensures that an image is obtained (a directly displayable display image Im_dis, with all luminances at their correct final values and no further brightness mapping required) that still reasonably maintains the creator's intention, i.e. how much brighter or darker any second image object should ideally be compared to the first image object, but taking into account the unavoidable limitations of the more limited maximum brightness of the display compared to the target display maximum brightness of the input image. Finally, the display formatting circuitry 255 may perform the conversion of the display adaptation luminance to some format required by the display panel 260 (e.g., an 8-bit RGBd code as determined by the final encoding function RQ), and may also involve some post-adjustments to give the colors some flavor according to the TV manufacturer (such as some user control settings for increasing the brightness of the darkest image pixels may be implemented by the display adaptation circuitry 254, but some user control settings may control the internal processing of the display formatting circuitry 255.
[0035] Figure 3This section illustrates some aspects of color processing involved in HDR image processing. If the image is non-color (also known as black and white, i.e., consisting only of colorless gray pixels), HDR image processing (e.g., for encoding) already involves complex new insights. However, color processing makes HDR the most challenging. Human color vision is highly complex due to numerous chemical and neural processes in the eye's cones (e.g., the color of a banana also depends on a memory of what a banana should look like). Therefore, any color space designed with regular mathematical properties (e.g., colors of equal perceived difference are ideally equidistant in the color space) will be highly nonlinear. However, as a general first approximation, most colors can be generated by additive combinations of the red, green, and blue primaries, and the visual appearance of a color can be roughly described as hue, saturation, and brightness. However, even typical technical color spaces exhibit non-linear behavior, making color processing non-trivial. One problem is that the modified color (which can be represented as an arrow from the input color to the resulting output color) can fall outside the gamut of representable colors represented (in the coded white to maximum displayable white frame, this would also correspond to an undisplayable color). For example, if a brightness code ends in a number, such as the power (2;10)-1, we could call it 100% or 1, and the code R=1050 could not exist. As such, you can't create more red than setting the backlight to maximum and turning the red channel on at maximum. Furthermore, the mathematical definition of brightness, which can be highly nonlinear, introduces complexity (and potentially color errors as a result) even within the representable color gamut.
[0036] HLG (standardized in ARIB STD-B67) addresses at least the first issue. Since the same transformation function is applied to the three RGB color channels individually and identically, there is no out-of-gamut issue.
[0037] Figure 3 A to Figure 3 C shows a color error that does occur. The color gamut is the RGB color cube, which is at least geometrically easy. The maximum color component (e.g., the red component R, or the green component G, or the blue component B) can be taken to be 1.0. If we now define a strictly increasing function F_mp that maps the maximum input 1.0 to an output 1.0, this guarantees that there will be no output above 1.0. Since this will be true for the three color components, the output color is guaranteed to always fall within the color gamut cube. However, we show what happens for a typical brightening function. This brightening function will come into play when mapping a higher dynamic range image to a lower dynamic range output image.
[0038] We can represent both images as normalized color triplets, i.e. R<=1, G<=1, B<=1. However, for a 1000nit image, the maximum normalized color (R=G=B) will correspond to 1000nit, while for a 100nit image, the same normalized color will represent only 100nit of white. Suppose we want to define some dark gray in the HDR image. In a linear RGB representation, we can, for example, say R=G=B=1 / 100. This will define a 10nit gray pixel in the HDR representation. If the representation is to be used as an SDR representation (e.g., displaying the encoded color directly without appropriate brightness processing), then on a 100nit display, a 100 darker color will appear as 1nit, i.e., dark black rather than dark gray. Therefore, before display, a suitable SDR image is made from the HDR image, which at least enhances the darkest color (e.g., if not 10 times, then for example 5 times). The brightest colors cannot be boosted because they must end up at the SDR maximum (1.0 representing 100 nits) that can be displayed (or represented in an SDR image). Therefore, a convex brightening curve is usually applied, such as Figure 3 The F_mp in B is shown (the so-called "r curve"). Figure 3 In A, the action of the curve is shown as squeezing toward 1.0 of the original equally spaced colors. For example, arrow 301 maps the original darker input color (or more precisely, its red component) to the output value shown by the circle at the end of the arrow. The r-curve makes darker colors relatively brighter by having a slope greater than 1.0 at the dark end.
[0039] Achromatic (also called achromatic) colors are defined by the property of always having their three color components equal (R=G=B), regardless of their brightness (e.g., R=0.2). Therefore, applying the same curve to the color components individually will actually produce a still-achromatic color that is significantly brighter. A color with a specific hue and saturation (e.g., a pale red, also called pink) will be defined by the ratio of the color components (e.g., if there is much more red than blue and green in the mix, the color is a strong red). This additive combination is like Figure 3C. Assume that the input color (the resulting color RC_1) consists of a small amount of red (normalized value R_1) and a large amount of green (G_1). Perceptually (i.e., when the system is technically used in a typical scenario, such as the expected ambient lighting), this will give a yellowish green color. If we now apply the function F_mp, then the predominantly red component will grow (e.g., double in amplitude) because the green component approaches 1.0 and can no longer change in percentage. Therefore, in the output additive mix, there will be a dominant red component. In fact, the two components are almost equal, which will give a yellow color (the second resulting color RC_2) instead of a certain green. In general, for this independent channel processing, the saturation and hue (and even the expected brightness if not processed correctly) may be incorrect for essentially all colors in the color gamut, and this may be objectionable. It can be a simple way to create and process HDR, but a future-proof approach that may not be considered by everyone as a high-quality HDR system.
[0040] exist Figure 3 In D, we show how alternative systems can handle colorimetry. Future-oriented systems that handle HDR more accurately (e.g., SL-HDR) may also decide to use this color definition and processing.
[0041] Here we show a more realistic representation (and more perception-oriented) of the additive color gamut. It is shown in a partially cylindrical hue / saturation / luminance view (but in practice could also be processed in a diamond version, e.g. by using chrominance CbCr, which for a given object saturation increases with increasing luminance, as Figure 3 E, where we have shown a slice through the color gamut spanned by the blue chromaticity and the luminance Y_n (e.g., luminance defined by PQ to represent normalized luminance).
[0042] The vertical axis represents the brightness of any pixel, specified as normalized luminance (L_n). We can again represent any actual color gamut (which is just a multiplicatively scaled version of the same shape defined given the same RGB primaries), ending up in the normalized version with a maximum target display luminance of, for example, 5000 nit as shown. So, if we overlap the color gamuts of, for example, a 5000 target display input video and a 100 nit output video, we can show the HDR to SDR luminance processing as arrows that move the relatively dark colors of the HDR image to the corresponding relatively bright colors of the SDR image. In the horizontal direction is saturation S, with zero saturation occurring in the middle for the non-color colors that span the luminance axis. So, we can again take the dark gray (C_Ha) of the HDR image and brighten it (with the same function F_mp or another convex function) to obtain a relatively brighter corresponding SDR color C_Sa. If we again employ a function that maps 1.0 to 1.0, no color on the non-color axis will exceed the maximum white (L_n=1.0) of the color representation. However, this color space is not cylindrical, all maximum colors (for any color with a saturation of, say, S=0.5, and some hue angle defined by differently oriented slices through the cylinder) fall at the same high normalized lightness, i.e. L_n equal to 1.0. White is the brightest color, defined by setting all color components to their maximum value. In order to create any other color, some energy of at least one component must be removed (i.e. its value is lowered). For example, maximum yellow (Ye) can appear by removing all blue. Due to the chemistry of the eye, blue appears darker than, e.g., green, i.e. it has a low intrinsic lightness. This means that not much lightness is removed to create maximum yellow (the exact amount depends on the chromaticity chosen for the primaries, e.g. Rec.709 versus Rec.2020 primaries), e.g. 10%. But this also means that the brightest saturated blue ( Figure 3 D, B) will only be as bright as 10% of white. Hence, we get a very asymmetric top of the gamut. Therefore, apply the same mapping curve (same offset or at least the same multiplicative offset) to blue (i.e. Figure 2 In the shown luma-chroma separation process based only on the luminance of blue, a high-luminance blue that is already close to the upper gamut boundary) will cause the HDR blue (C_Hb) to be mapped to an SDR blue (C_Sb) that falls outside the gamut. So what happens in practice is that the blue component will be clipped to its maximum value and therefore become less dominant in the mix, i.e. we can also get a resulting color Cr with the wrong hue, saturation and brightness. But this will only happen for some "difficult" colors, because we can control most of the good in-gamut colors to produce the output exactly as needed, usually a color that is only relatively brightened, but keeps the same hue and saturation as the input color (unlike HLG mapping).
[0043] EP3537697 teaches how to bring a Hybrid-Logggamma defined HDR image into an absolute framework by focusing on where the midpoint between the square root and logarithmic parts of the HLG luminance range will be mapped in the luminance range, involving a variable luminance mapping function shape that can be optimized for the scene.
[0044] US 2021 / 0272497 teaches that color processing for display-adaptive high dynamic range image optimization may require several controlled luminance values for brightness-dependent saturation changes.
[0045] US2019 / 0206360 addresses the non-constant brightness problem of highly nonlinear EOTFs when the chrominance component pixel images are spatially subsampled, for example by deriving more accurate values for brightness through binary search.
[0046] US2021 / 0241055 teaches how HDR images can be printed. The dynamic range of even the best quality prints is much lower than that of video images of even moderate maximum brightness (e.g., 1000 nit), in particular because prints cannot produce luminous areas because they cannot get brighter than paper white, and they cannot produce deep blacks because ink (especially when spread too thickly) is not a perfect absorber. Therefore, significant dynamic range down-conversion is required to make an SDR image for printing that corresponds to the HDR image. An embodiment can focus on the exact content in the image, for example, if there are mainly super bright pixels, these pixels can be mapped to a large sub-range of the printable range by making the mapping function essentially equal to zero for the dark and medium brightness pixels that are almost absent in such an image.
[0047] Therefore, in a similar spirit to the computational simplicity of HLG, it is desirable to have a relatively simple mechanism to excellently handle out-of-gamut issues with HDR image luminance transforms of the luminance / chrominance processing type. Summary of the Invention
[0048] Desirable advantages are achieved by a method of transforming an input pixel color (C_Hb) of a high dynamic range input image into an output color (C_Sb) of an output image, wherein the input color is represented as a gamma-logarithmically defined triplet according to a photoelectric transfer function having a gamma-logarithmic shape, the triplet comprising an initial first color component, an initial second color component, and an initial third color component, wherein the initial first color component, the initial second color component, and the initial third color component are transformed into an intermediate red component (R"_im), an intermediate green component (G"_im), and an intermediate blue component (B"_im) by mapping the initial color components using a function that is the same for each color component, wherein the function has a slope greater than 1 for a sub-range of color component values that includes the darkest of the initial color components, and wherein the intermediate color components are transformed into corresponding red output color components (R"_out), green output color components (G"_out), and blue output color components (B"_out) by applying:
[0049] - calculating a first value (Mx) which, for each input color, is the maximum of the three intermediate color components;
[0050] - subtracting a second value from the first value, the second value being equal to a value obtained from applying the optoelectronic transfer function to a three-level value, the three-level value being equal to a target display maximum luminance (ML_To) of the output image divided by a maximum luminance (ML_PQ) associated with the optoelectronic transfer function, the subtraction producing a fourth value;
[0051] - calculating a fifth value by determining the highest of said fourth value and zero;
[0052] - calculating three output color components by subtracting said fifth value from the corresponding intermediate color component and setting said output color component equal to zero in case said subtraction results in a negative value.
[0053] Problems arise with so-called "low-brightness colors" (i.e., colors that generally appear dark to human vision). These would be colors like blue, front-most, red, and often magenta (in contrast to colors like green and yellow, which appear bright because they have local gamut maximum brightness close to the tip of the white point and are generally not colors that create out-of-gamut projections). An alternative to the proposed per-pixel correction process (at least at the boundaries of the gamut to bring output pixels back into gamut) would be to examine the brightness-mapped image for the problem areas and then process it again with an r-curve that has a shallower slope for the sub-range of the darkest colors (i.e., "black," "dark gray," and "dark colors"). While this would not create, or at least create better out-of-gamut colors, the overall output image would appear less brightened, i.e., its dark areas might now appear undesirably dark again. Therefore, we want some processing that can work based on the brightness mapping as needed, essentially processing only the problematic out-of-gamut pixels while leaving the other calculated good output image colors largely unchanged.
[0054] The inventors have also found that such low-brightness colors often occur for graphic materials, whether computer-generated and inserted, or already in a scene captured by a camera, such as a printed cloth with a red or purple company logo on it. If such an image area would consist of a single color, this may not be a problem for the consumer. However, there will usually be a color shift (hue and / or saturation), and the incorrect color of the logo may be problematic for the company paying for the commercial (they would rather have a slightly darker version, but at least the correct one, such as Ferrari red). If the object color is illuminated with different amounts of lighting and shadow, i.e., the same rose petal red but with different brightness, the output result can show very different hues of the object, such as some purple pixels in addition to the red pixels, because the blue component has been increased due to the transformation, while the red component has been clipped. Even for non-critical consumers, this may be undesirable.
[0055] The present invention utilizes the advantageous properties of the color defined by the gamma logarithm OETF. A skilled person understands what this means. For example, a brightness code is defined by applying the OETF to brightness. For example:
[0056] Luminance = 1023 * OETF_gammalog (normalized luminance) [Equation 1]
[0057] The normalized luminance is obtained by dividing the pixel luminance by a reference maximum luminance (e.g., 10,000 nits). By using the same OETF, e.g., R”_norm = OETF(R_norm), the non-linear additive color components are calculated based on the linear R, G, and B components in additive mixing, where both R”_norm and R_norm are normalized to a maximum value of 1.0 and can be converted to absolute values by using a scaling factor. For example, an 8-bit non-linear component can be calculated as: R”_abs = 255 * R”_norm. The normalized linear component is obtained by dividing by the reference maximum value. For example, a contribution of 3500 for red will result in R_norm = 3500 / 10,000 = 0.35.
[0058] The gamma logarithmic function is a function that starts as a power law for the darkest colors, i.e., if L_norm < k, then Luma_norm = power(L_norm; p), where k << 1. And it becomes logarithmic above the threshold k. For example, the popular gamma logarithmic OETF for which the input color is already defined is the perceptual quantizer OETF, which becomes logarithmic above approximately 20 nits and is thus suitable for implementing the out-of-gamut processing embodiments below. Another exemplary gamma logarithmic OETF is the applicant's version described in WO2015007505, which is incorporated herein by reference. Logarithmic subtraction appears as scaling in the corresponding linear world, which corresponds to dimming without usually any hue or saturation change.
[0059] The input color triple (first color component, second color component, and third color component) can already be in the R”, G”, and B” defined by gamma logarithm (we use the double-prime symbol to indicate that the color components are not linear, similar to the single-prime R’ which is used to define the approximate square-root non-linearity components standardized in Rec. 709), or in YCbCr format. Those skilled in the art know the standard matrixing equations with matrix M to convert between the two systems when necessary.
[0060] Y = a * R” + b * G” + c * B”;
[0061] Cb = d * (B” - Y);
[0062] Cr = d * (R” - Y); [Equation 2] <b
[0063] The constants a, b, c, d and e that define the YCbCr model depend on the chromaticity of the chosen primary, i.e. Rec.2020, or DCI-P3, and are standard colorimetry, which also does not involve the novel insights of the present innovation (which must apply to any primary, luminance mapping function, etc.), so we will not discuss the details further. We will assume that we are able to transform the input color to some intermediate colors (R"_im, G"_im, B"_im), which will be good for the output image (e.g., an SDR image derived for the input master HDR image) except where it projects out of gamut. The skilled person understands that within the circuitry or calculations, it is possible to temporarily make the color representation larger (e.g., the output of the color transform circuit 501) so that it can run up to, for example, 5.0, as long as the final output color is in gamut, i.e., R"_out<=1; G"_out<=1; B"_out<=1. Furthermore, we want to have positive or at least zero color components. We illustrate an example where processing occurs on the R, G, and B components without limitation, as similar processing can be performed on the YCbCr color components of the corresponding gamma logarithms. Brightening of the transform occurs, for example, by applying an r-curve with a slope greater than 1.0 or 45 degrees for the darkest colors (e.g., 20% darkest color, Y < 0.2).
[0064] An example of a useful embodiment of such a curve is taught in WO2017102606, which is incorporated by reference. It is a curve that starts with a configurable linear segment below a first input value threshold Y_in_1 in gamma-log luminance space (i.e., luminance is processed in its corresponding luminance domain), the configuration being able to consist of a threshold position and a slope, followed by a parabolic portion up to a second input value threshold Y_in_2, which typically drops below 1.0, and another configurable linear segment for the brightest colors up to maximum white. Note that these are just exemplary luminance mapping functions for ease of illustration, and that in general any increasing function shape can be used (even in certain embodiments with equal output values for different input values, such as luminance clipping). In practical embodiments, the function can actually be applied to the three color components by multiplying by a common multiplier (mLum). In principle, any such mapping can benefit from the improved method of the present innovation.
[0065] The maximum value Mx of the pixel color minus the OETF (ML_T_o) value calculates the height of the projection above the upper color gamut boundary. The first maximizer avoids negative values by making the result at least as large as zero. This amount is subtracted from the intermediate color components produced by the (uncorrected) desired brightness processing performed by the color transform circuit 501. Typically, all components are high enough so that the colors only shift in hue above the upper color gamut boundary. So, for example, all the reds of a glowing image of a rose will all remain red (there may be some loss of detail inside the rose because all differential brightness is lost due to the simple mapping of all those colors to the same chromaticity (hue and saturation) position at the upper color gamut boundary, but at least the problem of different hues, such as purple appearing inside a red object, is avoided remarkably). Occasionally, due to the subtraction, one of the color components will drop below zero. The third maximizer circuit (557), the fourth maximizer circuit (558) and the fifth maximizer circuit (559) will ensure that all color components of the output color are positive ( Figure 5 and Figure 6 The x in _im indicates any component value of the input, e.g. R"_im minus VF). The value of the target display maximum brightness may typically be the value of a display connected to an IC or device (e.g. a set-top box or computer etc.) that performs the processing, or the value of the television or display itself in case the processing takes place in such a device. It may also be a prepared value for a video that is prepared and stored for later viewing. This can be used in scenarios where (after pre-optimization according to the present method) several videos of different maximum brightness are stored for e.g. video on demand. For the sake of completeness, the word "target" does not refer to the situation as is (i.e. the maximum brightness of the input image), but to the situation of the optimized output image to be created.
[0066] Figure 4 Helps understand because it shows the view Figure 3 Another way to calculate the color space based on luminance / chrominance of D.
[0067] Now, all possible color gamuts, such as the color gamut of an HDR image to be optimized for a 650 nit target display (corresponding to a practically available TV), are superimposed on the primary color gamut in a scaled manner. The normalized luminance L_nPQ of this, for example, perceptual quantizer color gamut still ends at 1.0. However, the absolute luminance corresponding to this relative maximum is the maximum luminance associated with the photoelectric transfer function (ML_PQ), which for PQ is defined as 10,000 nit (typically 5,000 nit for Philips' primary OETF). How high any color gamut is depends on the maximum luminance of the chosen target display. For example, if a video produced for ML_T needs to be 2,000 nit (e.g., starting from a 5,000 nit primary grading, which would again require R-type luminance degradation or relative upmapping), the maximum value will fall at the PQ luminance corresponding to a normalized luminance of 0.2 (2,000 / 10,000), which is calculated using the PQOETF. A similar situation of out-of-gamut mapping may exist if something is mapped outside the intended output image gamut (in the example, a 2000 nit target display gamut), as shown. An absolute luminance EOTF associates (as defined in the corresponding standard defining the EOTF) some maximum luminance value with its maximum luminance code (e.g., 1023 in 10 bits). Some EOTFs may use a variable (configurable) maximum luminance, but for example, a perceptual quantizer uses a fixed value of 10,000 nits.
[0068] It is advantageous if the method for transforming the input pixel color calculates an aggregate fifth value, obtained as the maximum of the consecutively calculated fifth values for a set of pixels in the input image. The aforementioned method is very simple because it can be calculated pixel by pixel over a run of consecutive pixels (e.g., from a zigzag scan of the input image). The attribute only needs to be determined based on the colorimetry of the pixel currently being processed, so many practical calculation schemes can be designed for algorithms running on, for example, an ASIC or a CPU or GPU. A disadvantage of the simple version is that all out-of-gamut pixels of the same chromaticity are projected onto the same upper gamut boundary point, resulting in some loss of internal detail. This may not be a problem for bright logos consisting of only a few consistently present colors (e.g., the TV station logo in the upper left corner of the image, which consists of, for example, monochromatic red text on a yellow background). The human eye will simply see a slightly darker version of the same color. It can also work for objects in the image, such as red TL tubes. However, some objects consisting of critically bright HDR pixels (e.g., sunlit clouds) may not be as beneficial. Several alternatives exist, including creating a better version or shutting down the processing and potentially applying another out-of-gamut processing strategy. This embodiment allows the fifth values of a plurality of pixels in a definable area (e.g., a rectangle covering the upper left marker) to be collected and the maximum of the fifth values selected for all pixels in the rectangle. This allows for a more differentiated color structure in the output colors returned to the gamut mapping. The pixel processing pipeline must be delayed until the aggregated final values for all or some of the pixels in the image area have been applied to the pixel scan. This can be implemented in various ways, for example, with two passes over the entire image (once to determine the fifth values and once to apply them to create the in-gamut output color), a smaller portion of memory, and a separate dedicated scan of the rectangle, followed by a normal scan of the rest of the image, maintaining a pixel mask to decide which pixels to apply a single fifth value to based on the pixel's own color (i.e., the pixel's color before it was mapped to the gamut) versus applying an aggregated fifth value to other pixels, or even no processing, etc., depending on whether the hardware circuit is designed for an ASIC or software to run on a mobile phone, etc.
[0069] It is advantageous if the method for transforming the input pixel colors includes a color detail analyzer 640, which is arranged to determine a detail metric (DM) for the pixel to be transformed and, if the detail metric is less than a threshold (ThM), apply the transformation. More complex detail analyzers can be designed, but a simple one is usually sufficient. The goal is to examine whether this method (particularly with a single, natively color-determined fifth value) creates (truly) objectionable results, in the sense that they are worse than doing nothing at all. Alternatively, in some scenes, it may be decided to use highly advanced out-of-gamut processing, such as strategies optimized and communicated by human color graders, but this will only be done in areas where carefully optimized out-of-gamut compensation is truly necessary, such as balancing brightness and saturation deviations, and perhaps even some minor hue errors, such as for sunlit cloudscapes. However, while simple algorithms can work reasonably well automatically, they are not favored by human graders. This method can, for example, be run on the creation (encoding) side to indicate where errors occur on the decoding side, allowing the video creator to take action. He can then, for example, define secondary areas in the image to which his advanced processing will be applied (e.g. as pixel masks or rectangles, etc.). On the other hand, the detail analyzer can also work only on the receiving side, in an image processing device such as a consumer television, to decide whether or what post-correction should be performed. An example of a detail analyzer would be to see that in a color histogram there are, for example, two color lobes, around which several varying colors with small color differences (e.g., Δbrightness or Δsaturation) are spread. This would constitute a typical two-dominant color case from practice, as in the example of red on a yellow sign. In this case, a simple algorithm with a fifth value associated with each pixel's own color can be well applied. If a sunset cloudscape is seen with a lot of yellow and pink, then this method would not be desirable and something else could be done. Thus, a simple detail measure could be a color quantity, calculated, for example, as the maximum deviation from a mean color, etc. The creator of the algorithm can determine the threshold ThM, for example, by studying a number of typical images on which he will exclude the algorithm that needs to work (for example, discotheque images need to be brightened so that the dancing people can still be more or less seen despite the black, but all common monochromatic light sources such as flashlights and lasers should undergo out-of-gamut correction, and some of those images can be placed in a test set to see how the system performs on average). For example, if the blue laser is set at a color difference of delta_Lab of, for example, 10 (as a threshold) (or based on the error in YCbCr), a simple method can be applied. Advanced detail analyzers can, for example, use trained neural networks, but generally require more chip surface and power, at least for training.
[0070] It is further advantageous to implement a practical variant as an apparatus (500) for transforming an input pixel color (C_Hb) of a high dynamic range input image into an output color (C_Sb) of an output image, wherein the input color is represented as a gamma-logarithmically defined triplet comprising an initial first color component, an initial second color component and an initial third color component according to a photoelectric transfer function having a gamma-logarithmic shape, wherein the apparatus comprises a color transform circuit (501) arranged to transform the initial first color component, the initial second color component and the initial third color component into an intermediate red component (R″_im), an intermediate green component (G″_im) and an intermediate blue component (B″_im) by mapping the initial color components using a function that is the same for each color component, wherein the function has a slope greater than 1 for a sub-range comprising the darkest color component values of the initial color components, characterised in that the apparatus comprises:
[0071] - a first maximizing circuit (551) arranged to calculate a first value (Mx) which is, for each input color, the maximum of the three intermediate color components;
[0072] a first subtractor (552) arranged to subtract a second value (V2) from the first value, the second value (V2) being equal to a value obtained from applying the optoelectronic transfer function to a three-level value, the three-level value being equal to a target display maximum luminance (ML_To) of the output image divided by a maximum luminance (ML_PQ) associated with the optoelectronic transfer function, the subtraction producing a fourth value (V4);
[0073] - a second maximizing circuit (553) arranged to calculate a fifth value, the fifth value being the highest of the fourth value and zero;
[0074] - a second subtractor (554), a third subtractor (555) and a fourth subtractor (556) arranged to calculate three output color components (R"_out, G"_out, B"_out) by subtracting the fifth value from the corresponding intermediate color component, and arranged to set the output color component equal to zero if the subtraction produces a negative value.
[0075] Alternatively, the apparatus (500) for transforming the input pixel color (C_Hb) of a high dynamic range input image according to claim 4 comprises a sixth maximization circuit (665), which is arranged to generate an aggregate fifth value (VF), the aggregate fifth value (VF) being obtained as the maximum one of consecutively calculated fifth values of a group of pixels of the input image.
[0076] Alternatively, the device (500) for transforming the input pixel color (C_Hb) of a high dynamic range input image according to claim 5 or 6 comprises a color detail analyzer circuit (640) which is arranged to determine the amount of color detail and thereby determine a control signal (CTRL) which controls whether, for at least one area of the input image, the processing starting from the calculation of the first value is applied or intermediate color components are generated as alternative output color components (R"_out2, G"_out2, B"_out2). BRIEF DESCRIPTION OF THE DRAWINGS
[0077] These and other aspects of the method and apparatus according to the present invention will become apparent and elucidated from the embodiments and examples described below with reference to the accompanying drawings, which serve only as non-limiting illustrations of the more general concepts and in which dashed lines are used to indicate that a component is optional, while non-dashed lines are not necessarily required. Dashed lines can also be used to indicate elements that are interpreted as required but are hidden within an object, or for intangible things such as the selection of an object / region.
[0078] In the attached figure:
[0079] Figure 1 Schematically illustrated as a technical background reference the classic so-called low dynamic range or standard dynamic range (SDR) video processing (creation, encoding and communication as well as display), which is still the main approach to video on the market at this time (i.e. it is the "normal video technology"), but is rapidly being replaced by high dynamic range (HDR) video technology;
[0080] Figure 2 Typical aspects of modern and future-oriented HDR video processing are schematically illustrated at a bird's-eye overview level, from the creation of new HDR looks to optimal display across the spectrum for variable maximum brightness capable end-user displays;
[0081] Figure 3 Schematically illustrates two alternative approaches to color processing, namely 3 color dimensions, i.e. two additional color dimensions in addition to the achromatic axis of brightness, which can be used for HDR image processing (note that e.g. image brightening on a color image is ultimately 3D color processing even if it is primarily defined by 1D brightness processing requirements, e.g. a luminance mapping function);
[0082] Figure 4 schematically illustrates brightness processing of HDR images with different associated maximum brightness, and corresponding re-grading of at least one pixel color (e.g., relative brightening so that the actual (absolute) displayed pixels appear the same brightness regardless of which image variant is used to drive the display);
[0083] Figure 5 Schematically shows the basic configuration of the sub-circuits of an apparatus for carrying out a simple embodiment of the present invention;
[0084] Figure 6 The schematic shows how several more advanced computational components can be added to the basic circuit to form more complex but potentially better quality variants. DETAILED DESCRIPTION
[0085] Figure 5 An example is shown, for illustration only and without undue limitation, of typical components that may be present in some receiving-side devices for receiving and processing (e.g., for display) HDR video, such as sports broadcasts or television programs (or movies, etc.). Without wishing to be limiting, we will assume that we are describing a consumer television that receives PQ-defined YCbCr pixel colors in the input image and maintains internal operation in the PQ domain (indicated by the double primes). The left subcircuit is the color conversion circuit 501. Because our innovation should work with many such embodiments, we will keep its description as brief as possible.
[0086] In many practical embodiments (although the innovation may be applied to devices which receive or determine their luminance mapping function differently, for example), the receiver 502 is configured to parse (extract and distribute to appropriate computational circuitry, and sometimes properly configure such circuitry, e.g., create a LUT, although this may also be performed internally in the computational circuitry) an HDR video signal (SH_in), typically broadcast or privately transmitted. Figure 2As explained, this will typically include the image essence—a matrix of YCbCr color pixels that will be sequentially scanned for color processing, along with metadata. Depending on the HDR standard, metadata can take on various flavors, but without limitation, we will assume SL-HDR metadata, which includes, for example, a dedicated (optimal) luminance mapping function for each time-sequential image, defined in the luminance portion of the metadata METL. We assume, but are not limited to, this is defined in the PQ domain, FL_ref_PQ. Another piece of metadata for one or more images is color metadata METC, which includes, for example, a lookup table of chroma multipliers for all possible gamma-log luminance values (ms[Y_lg]). The saturation processing circuit 505 selects a corresponding multiplier from the saturation multiplier LUT based on the actual luminance (Y_lg_in) of the pixel currently being processed. The two chroma components Cb and Cr of the current pixel are multiplied by the selected multiplier ms, which constitutes color processing, producing, at least at that moment, the "output chroma," namely, the intermediate chroma Cbim and Crim. Via the matrixing circuit 508, normalized versions of the red, green and blue components defined by the gamma logarithm are calculated. This inverse RGB-YCbCr equation is well known to the skilled person and will not be explained further (see equation 2 above). After the track has performed its desired calculations on the brightness processing, these normalized components will be scaled to their correct amplitudes (by the first, second and third rescaling multipliers 509, 510, 511). The reference brightness mapping function FL_ref_PQ extracted from the HDR input signal metadata is typically a mapping function from the HDR input to the SDR 100nit signal. Therefore, if a 100nit maximum brightness output is desired, it can be used directly in the brightness mapper 504, but otherwise (if, for example, the optimal image for a 350nit display should be calculated based on the input 1000nit image) an adaptation function (FL_d_PQ) should be calculated. Similarly, various display optimization algorithms can be used, for example also taking into account the optimization required for atypical viewing environments. Therefore, for the present explanation of the actual components of the present innovation (which reside in the return to gamut mapping subcircuit 550), it is sufficient to say that the display adaptation function calculation circuit 503 calculates a slightly darkened version of FL_ref_PQ that has been compressed towards the diagonal of the graph normalized to 1.0 input and output luminance. For more details on possible technical approaches, the interested reader is directed to WO2017108906. For the purpose of illustrating the present invention, even suboptimal display adaptation that only creates a function FL_d_PQ with half the orthogonal distance from any point on the diagonal compared to FL_ref_PQ will work: it may not produce perfect colors within the gamut, but the out-of-gamut correction will work regardless.
[0087] The luminance mapper applies the function: luma_out = FL_d_PQ(Y_lg_in). Furthermore, in this example, it converts the result into a multiplier mLum, which produces the same result as applying this function when the normalized RGB components R"_n, G"_n, and B"_n are each multiplied by the same multiplier mLum, thereby producing the color-processed intermediate color components R"_im, G"_im, and B"_im.
[0088] How this can be achieved follows. Assume we brighten the value V_in to V_out = F(V_in). Assume the brightness boost (mLum) is 2xV_in. Then we can find this mLum as V_out / V-in = F(V_in) / V_in. Therefore, the brightness mapping circuit knows its input and only needs to divide the output of the function by its input to obtain the desired multiplier mLum as output.
[0089] Continuing now to a typical embodiment for simple out-of-gamut processing (or into-gamut projection), we first determine in a first maximization circuit (551) which of the three input color components is the highest one (which can be above 1.0 in the extended range internal representation used for calculations). For example, if the middle color is R=0.9; G=0.1; B=1.2, then Mx will be 1.2 (the blue component). If we need, for example, an output image for a 350nit display, we calculate V2=OETF_PQ(350 / 10000), or typically OETF_lg(0.035) if we have a different gamma logarithm OETF. A subtractor 552 subtracts V2 from Mx.
[0090] The second maximization circuit (553) ensures that only out-of-gamut colors are down-mapped. The second subtractor (554), the third subtractor (555), and the fourth subtractor (556) subtract the value required to project toward gamut from at least one of the intermediate color components that is too high. It can happen that one of the color components, in the example above, which could be as little as 0.1 for the green component, goes below zero. To maintain positive output color components (R"_out, G"_out, B"_out), the third maximization circuit (557), the fourth maximization circuit (558), and the fifth maximization circuit (559) will equalize any negative values to zero.
[0091] Figure 6Shown are some possibilities of making a more advanced device (600) based on the basic concept that can work separately but also work well together. A detail analyzer 640 can apply some analysis of how many colors are present and how well or poorly each out-of-gamut compensation algorithm will perform. It can control switches 641, 642, and 643, which in this example completely bypass the out-of-gamut algorithm and output intermediate color components as alternative output components (R"_out2, G"_out2, B"_out2). This can still be processed by another TV color circuit, or not, in the latter case, the "odd" interior details of the false colors are considered better than those produced by, for example, a simple out-of-gamut processing embodiment. The switches are set to their appropriate positions by a control signal (CTRL), which is derived from, for example, the detail metric (DM) of the pixel to be transformed (e.g., the standard deviation from some main image color) being less than a threshold (ThM). The detail analyzer is in principle capable of performing any pre-analysis buffered image processing, as embodiments of the present invention can be preceded or followed by several other image processing blocks, although it may be desirable if the simple processing itself performs all the required satisfactory corrections. More advanced control signals from a detailed analysis of the input image colors can switch between several processing options (or no processing), the simple embodiment typically being one of them.
[0092] In subcircuit 650, we also show some interesting variations. A number of pixels, for example, controlled by a counter CNT_h, can be buffered in memory 651 (this works particularly well if a specific area is scanned first; those skilled in the art will appreciate alternative hardware implementations) until the final aggregate fifth value VF is determined and can be processed. Region management circuit 654 will arrange for this. For example, if the flag resides in a 100-pixel by 100-pixel flag area, it can count up to 10,000. Aggregation factor calculation circuit 655 will calculate the aggregation factor VF for that area. The simplest variant is to calculate which of the fifth values is the highest and use that as the VF, but other algorithms are possible. Control by region management circuit 654 will also be done on this circuit 655, for example by maintaining a maximum until CNT_h is exhausted. More complex region management could, for example, determine which pixels in a rectangle or other flag area receive which bit pattern. If some processing is not done or changed, these embodiments can work well with the detail analyzer circuit (640) to determine the best case from the total information (but the detail analyzer and aggregation can also be used separately). The out-of-gamut processing based on the gamma logarithm can be combined with other techniques (such as pre-processing or post-processing), or alternated with other out-of-gamut techniques such as pixel-by-pixel or region-by-region processing. For example, the creator of the video or image content (human or automatic) can indicate, for example, in the metadata of the video signal, a rectangle where the algorithm will not be applied (or will be applied), where it can be applied to the rest of the image. This can be used to protect, for example, a logo that gets projected onto a color that is too similar to the upper gamut color, while still seeing the original logo. In this case, it is sometimes better to see the logo with the deformed color, but at least its geometric structure is easier to perceive. This can be useful if, for example, the logo is not post-applied to the video, but is already present in the video captured by the camera, such as on a T-shirt or a decorative background. Sometimes, the person in charge can decide this in the acceptance round. For example, he can quickly click (or drag) on the diagonal points that define the upper left and lower right points of a rectangle, and these coordinates can be transmitted in the metadata as a protected area where the present method should not be applied. In more advanced cases, the video creator can indicate another algorithm to be used by the receiver, or the receiver can choose its own processing, or not to perform color gamut correction (i.e., color clipping), etc. A bitmap can be generated to perform more precise per-pixel processing (e.g., 1 = apply; 0 = not apply), and such a bitmap can also be transmitted in the metadata of the transmitted video signal, or generated internally by the receiver by applying some pixel differentiation algorithm (which can also be selected by transmitting a determiner value that specifies which of several preconfigured algorithms to apply), etc.
[0093] The algorithm components disclosed herein may in practice be implemented (completely or partially) as hardware (eg, part of a dedicated IC) or as software running on a special digital signal processor or a general purpose processor or the like.
[0094] From what we have said, a person skilled in the art will understand which components may be optional modifications and can be implemented in combination with other components, and how the (optional) steps of the method correspond to the corresponding modules of the device, and vice versa. The word "device" in this application is used in its broadest sense, i.e. a group of modules that allow a specific goal to be achieved, and can therefore be, for example, (a small circuit part of) an IC or a dedicated appliance (such as an appliance with a display) or part of a networked system, etc. "Arrangement" is also intended to be used in the broadest sense, so it can especially include a single device, a part of a device, a collection of (parts of) cooperating devices, etc.
[0095] The term "computer program product" should be understood to mean any physical implementation comprising a set of commands such that, after a series of loading steps (which may include intermediate conversion steps, such as translation into an intermediate language and a final processor language), a general-purpose or special-purpose processor can input the commands into the processor and perform any characteristic functions of the present invention. In particular, a computer program product can be implemented as data on a carrier such as a disk or tape, data present in a memory, data transmitted via a network connection (wired or wireless), or program code on paper. In addition to the program code, characteristic data required by the program can also be embodied as a computer program product.
[0096] Some steps required for the operation of the method may already exist in the functionality of the processor rather than being described in the computer program product, such as data input and output steps.
[0097] It should be noted that the above embodiments illustrate rather than limit the present invention. While a skilled person can readily map the examples presented to other areas of the claims, for the sake of brevity, we have not discussed all of these options in depth. In addition to the combinations of elements of the present invention combined in the claims, other combinations of elements are possible. Any combination of elements can be implemented in a single dedicated element.
[0098] Any reference signs between brackets in a claim are not intended to limit the claim. The word "comprising" does not exclude the presence of elements or aspects not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements.
Claims
1. A method for converting an input pixel color (C_Hb) of a high dynamic range input image into an output color (C_Sb) of an output image, wherein: The input color is represented as a gamma-logarithmically defined triplet according to a photoelectric transfer function having a gamma-logarithmic shape, the gamma-logarithmically defined triplet comprising an initial first color component, an initial second color component, and an initial third color component, wherein the initial first color component, the initial second color component, and the initial third color component are transformed into an intermediate red component (R″_im), an intermediate green component (G″_im), and an intermediate blue component (B″_im) by mapping the initial color components using a function that is the same for each color component, wherein the function has a slope greater than 1 for a sub-range of color component values that includes the darkest color components of the initial color components, wherein the intermediate color components are transformed into corresponding red (R″_out), green (G″_out), and blue (B″_out) output color components by applying: - calculating a first value (Mx) which, for each input color of the pixel being processed, is the maximum of the three intermediate color components; - subtracting a second value from the first value, the second value being equal to a value obtained from applying the optoelectronic transfer function to a three-level value, the three-level value being equal to the target display maximum luminance (ML_To) of the output image divided by the maximum luminance (ML_PQ) of the optoelectronic transfer function, the subtraction producing a fourth value; - calculating a fifth value by determining the highest of said fourth value and zero; - calculating three output color components by subtracting said fifth value from the corresponding intermediate color component and setting said output color component equal to zero if said subtraction results in a negative value.
2. The method for converting the color of an input pixel according to claim 1, wherein: The fifth value is obtained as a maximum one of consecutively calculated fifth values of a group of pixels of the input image.
3. A method for transforming the color of an input pixel according to claim 1 or 2, comprising a color detail analyzer, which is arranged to determine a detail metric (DM) of the pixel to be transformed and to apply the transformation if the detail metric is less than a threshold value (ThM).
4. An apparatus (500) for converting an input pixel color (C_Hb) of a high dynamic range input image into an output color (C_Sb) of an output image, wherein: The input color is represented as a gamma-logarithmically defined triplet according to a photoelectric transfer function having a gamma-logarithmic shape, the gamma-logarithmically defined triplet comprising an initial first color component, an initial second color component and an initial third color component, wherein the apparatus comprises a color conversion circuit (501) arranged to convert the initial first color component, the initial second color component and the initial third color component into an intermediate red component (R"_im), an intermediate green component (G"_im) and an intermediate blue component (B"_im) by mapping the initial color components using a function that is the same for each color component, wherein the function has a slope greater than 1 for a sub-range of the darkest color component values including the initial color components, characterised in that the apparatus comprises: - a first maximizing circuit (551) arranged to calculate a first value (Mx) which is, for each input color, the maximum of the three intermediate color components; a first subtractor (552) arranged to subtract a second value (V2) from the first value, the second value being equal to a value obtained from applying the optoelectronic transfer function to a three-level value, the three-level value being equal to a target display maximum luminance (ML_To) of the output image divided by a maximum luminance (ML_PQ) associated with the optoelectronic transfer function, the subtraction producing a fourth value (V4); - a second maximizing circuit (553) arranged to calculate a fifth value, said fifth value being the highest of said fourth value and zero; - a second subtractor (554), a third subtractor (555) and a fourth subtractor (556) arranged to calculate three output color components (R"_out, G"_out, B"_out) by subtracting the fifth value from the corresponding intermediate color component, and arranged to set the output color components equal to zero if the subtraction produces a negative value.
5. The device (500) for transforming the input pixel color (C_Hb) of a high dynamic range input image according to claim 4, comprising a sixth maximization circuit (665) arranged to generate an aggregate fifth value (VF), the aggregate fifth value being obtained as the maximum one of consecutively calculated fifth values of a set of pixels of the input image.
6. The device (500) for transforming the input pixel color (C_Hb) of a high dynamic range input image according to claim 4 or 5, comprising a color detail analyzer circuit (640), which is arranged to determine the amount of color detail and thereby determine a control signal (CTRL) which controls whether, for at least one area of the input image, the processing starting from the calculation of the first value is applied or the intermediate color component is generated as an alternative output color component (R"_out2, G"_out2, B"_out2).
Citation Information
Patent Citations
Pixel processing with color component
US20190206360A1
Image processing apparatus, image processing method, and non-transitory computer-readable storage medium storing program
US20210241055A1
Optimized decoded high dynamic range image saturation
US20210272497A1
Methods and apparatuses for creating code mapping functions for encoding an HDR image, and methods and apparatuses for use of such encoded images
WO2015007505A1
Methods and apparatuses for encoding HDR images
WO2016020189A1