Brightness range adaptation for computers
The method of remapping HDR pixel colors to a common color space addresses the challenge of displaying HDR images on displays with lower maximum luminance, ensuring accurate and detailed image representation across various display types.
Patent Information
- Application Number
- PCT/EP2025/053542
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-14
- Filing Date
- 2025-02-11
- Publication Date
- 2025-08-21
AI Technical Summary
Existing technologies struggle to effectively display and communicate High Dynamic Range (HDR) images on displays with lower maximum luminance capabilities, leading to clipping and loss of detail in brighter areas.
A method for brightness range adaptation that involves remapping HDR pixel colors to a common color space, using secondary pixel colors and brightness mapping functions, allowing for the creation of a communication image with a lower dynamic range that can be displayed on legacy displays, while maintaining color accuracy and detail.
Enables the display of HDR images on a wide range of displays with varying maximum luminance capabilities, preserving image quality and detail, and allowing for seamless communication between HDR and legacy systems.
Smart Images

Figure EP2025053542_21082025_PF_FP_ABST
Abstract
Description
[0001] BRIGHTNESS RANGE ADAPTATION FOR COMPUTERS
[0002] FIELD OF THE INVENTION
[0003] The invention relates to coordinating the brightnesses of pixels in various visual assets for display, in particular specifically for computer environments, in which various visual assets of different maximum luminance, or maximum brightness relative to a reference level, are combined in a total view. These assets may lie in different windows, and be generated or managed by different applications which run concurrently.
[0004] BACKGROUND OF THE INVENTION
[0005] For more than half a century, an image representation / coding technology which is now called Low Dynamic Range (LDR) or. Standard Dynamic Range (SDR) worked perfectly fine for creating, communicating and displaying electronic images such as videos (i.e. temporally successive sequences of images), for e.g. pre-recorded movies or live broadcasts, or still images, such as e.g. graphics, e.g. in games. Colorimetrically, i.e. regarding the specification of the pixel colors, it was based on the technology which already worked fine decades before for photographic materials and paintings: one merely needed to be able to define, and display, most of the colors projecting out of an axis of achromatic colors (a.k.a. greys) which spans from black at the bottom end to the brightest achromatic color giving the impression to the viewer of being white. An example of such a color gamut of all possible colors is the “Munsell tree”, which distributes colors of various hues around a vertical axis of achromatic greys. The Munsell tree can be used to characterize the color of an object one finds in the world, e.g. a colored stone. The real world is different: from the darkest comer at night, till the middle of a supernova, colors (and their brightnesses) can be almost anything, which would be represented in an infinite cylinder rather than a limited gamut like e.g. a diamond-shaped gamut. If one represents a single relatively uniformly lit environment, a cut from that cylinder will be well-mappable to the limited (closed) gamut like e.g. the Munsell tree. But in general one may have differently lit environments side by side, and then a white color indoors may be a darker white color than e.g. a white car outdoors being strongly lit by the sun. Human eyes may sometimes see the former as some grey, but in general all those colors will look white (though differently bright whites). If one now wants to represent or display such an environment (ideally), more colors need attention than just the absolute brightest white of any representation of colors in a scene (which is an important characteristic of the representation, but merely indicates the brightest color one still wants to faithfully code and / or process).
[0006] Apart from theoretical colorimetry concerns, one needs technically simple systems, which not only are capable of relatively faithfully displaying the required colors, but also automatically capture the colors of a scene in the world into such a displayable representation, respectively one would like human operators to be able to change the colors to their liking. For television communication, which relies on an additive color creation mechanism at the display side, a triplet of red, green and blue color components needed to be communicated for each position on the display screen (pixel), since with a suitably proportioned triplet (e.g. 60%, 30%, 25%) one can make almost all colors, and in practice all needed colors (white being obtained by driving the three display channels to their maximum, with driving signal Rmax=Gmax=Bmax).
[0007] The earliest television standards (NTSC, PAL) communicated the color components as three voltage signals (which defined the amount of a color component between 0 and 700m V), where the time positions along the voltage signal corresponded by using a scan path with pixels on the screen.
[0008] The control signals generated at the creation side, directly instructed what the display should make as proportion (apart from their being an accidental fixed gamma pre-correction at the transmitter, because the physics of the cathode ray tube took approximately a square power of the input voltage, which would have made the dark colors much blacker than they were intended e.g. as seen by a camera) at the creation side). So a 60%, 30%, 25% color (which is a dark red) in the scene being captured, would look substantially similar on the display, since it would be re-generated as a 60%, 30%, 25% color (note that the absolute brightness didn’t matter much, since the eye of the viewer would adapt to the white and the average brightness of the colors being displayed on the screen). One can call this “direct link driving”, without further color processing (except for arbitrary and unnecessary processing a display maker might still perform to e.g. make his sky look more blue-ish). For reasons of backwards compatibility with the older black and white television broadcasts, instead of actually communicating a red, green and blue voltage signal, a brightness signal and two color difference signals called chroma were transmitted (in current nomenclature the blue chroma Cb, and the red Cr). The relationship between RGB and YCbCr is an easy one, namely they can be calculated into each other by using a simple fixed 3x3 matrix (the coefficients of which depend on the emission spectra of the three primaries, and are standardized, i.e. also to be emulated electronically internally by LCDs which actually may have different optical characteristics, so that from image communication point of view all SDR displays are alike).
[0009] These voltage signals were later for digital television (MPEG-based et al.) digitized, as prescribed in standard Rec. 709, and one defined the various amounts of e.g. the brightness component with an 8 bit code word, 0 coding for the darkest color (i.e. black), and 255 for white. Note that coded video signals need not be compressed in all situations (although they oftentimes are). In case they are we will use the wording compression (e.g. by MPEG-HEVC, AVI etc.), not to be confused with the act of compressing colors in a smaller gamut respectively range. With brightness we mean the part of the color definition that will impact upon the colors (to be) displayed the visual property of being darker respectively brighter. In the light of the present technologies, it is important to correctly understand that there can be two kinds of brightnesses: relative brightness (as a percentage of something, which may be undefined until a choice is made, e.g. by the consumer buying a certain display, and setting its brightness setting to e.g. 120%, which will make e.g. the backlight emit a certain amount of light, and so also the white and colored pixels), and on the other hand absolute brightness. The latter can be characterized by the universal physical quantity luminance (which is measured technically in the unit nit, which is also candela per square meter). The luminance can be stated as an amount of photons coming out of a patch on an object, such as a pixel on the screen, towards the eye (and it is related to the lighting concept of illuminance, since such patch will receive a certain illuminance, and send some fraction of it towards the viewer).
[0010] Recently two unrelated technologies have emerged, which only came together because people argued that one might as well in one go improve the visual quality of images on all aspects, but those two technologies have quite different technical aspects.
[0011] On the one hand there was a strive towards wide gamut image technology. It can be shown that only colors can be made which lie within the triangle spanned by the red, green and blue primaries, and nothing outside. But one chose primaries (originally phosphors for the CRT, later color filters in the LCD, etc.) which lay relatively close to the spectral locus of all existing colors, so one could make sufficiently saturated colors (saturation specifies how far away a color lies from the achromatic colors, i.e. how much “color” there is). However, recently one wanted to be able to use novel displays with more saturated primaries (e.g. DCI_P3, or Rec. 2020), so that one also needed to be able to represent colors in such wider color spaces (a color space is the mathematical 3D space to represent colors (as geometric positions of coordinate numbers), the base of which being defined by the 3 primaries; for the technical discussion we may better use the word color gamut, which is the set of all colors that can be technically defined or displayed (i.e. a space may be e.g. a 3D coordinate system going to infinity whereas the gamut may be a cube of some size in that space); for the brightness aspect only, be will talk about brightness or luminance range (more commonly worded as “dynamic range”, spanning from some minimum brightness to its maximum brightness). The present technologies will not primarily be about chromatic (i.e. color per se, such as more specifically its saturation) aspects, but rather about brightness aspects, so the chromatic aspects will only be mentioned to the extent needed for the relevant embodiments.
[0012] A more important new technology is High Dynamic Range (HDR). This should not be construed as “exactly this high” (since there can be many variants of HDR representations, with successively higher range maximum), but rather as “higher than the reference / legacy representation: SDR”. Since there are new coding concepts needed, one may also discriminate HDR from SDR by aspects from the technical details of the representation, e.g. the video signal. One difference of absolute HDR systems is that they define a unique luminance for each image pixel (e.g. a pixel in an image object being a white dog in the sun may be 550 nit), where SDR signals only had relative brightness definitions (so the dog would happen look e.g. 75 nit (corresponding to 94%) on somebody’s computer monitor which could maximally show 80 nit, but it would display at 234 nit on a 250 nit SDR display (yet the viewer would not typically see any difference in the look of the image, unless having those displays side by side). The reader should not confuse luminances as they exist (ultimately) at the front of any display, with luminances as are defined (i.e. establishable) on an image signal itself, i.e. even when that is stored but not displayed. Other differences are metadata that any flavor of HDR signal may have, but not the SDR signal.
[0013] Colorimetrically, HDR images can represent brighter colors than SDR images, so in particular brighter than white colors (glowing whites e.g.). Or in other words, the dynamic range will be larger. The SDR signal can represent a dynamic range of 1000: 1 (how much dynamic range is actual visible when displaying, will depend inter alia on the amount of surround light reflecting on the front of the display). So if one wants to represent e.g. 10,000: 1, one must resort to making a new HDR image format definition (we may in general use the word signal if the image representation is being or to be communicated rather than e.g. merely existing in the creating IC, and in general signaling will also have its own formatting and packaging, and may employ further techniques depending on the communication mechanism such as modulation).
[0014] Depending on the situation, the human eye can easily see (even if all on a screen in front corresponding to a small glare angle) 100,000: 1 (e.g. 10,000 nit maximum and 0.1 minimum, which is a good black for television home viewing, i.e. in a dim room which only support lighting of a relatively low level, such as in the evening). However, it is not necessary that all images as created by a creative go as high: (s)he may elect to make the brightest image pixel in an image or the video e.g. 1000 nit.
[0015] The luminance of SDR white for videos (a.k.a. the SDR White Point Luminance (WP) or maximum luminance (ML)), is standardized to be 100 nit (not to be confused with the reference luminance of white text in 1000 nit HDR images being 200 nit). I.e., a 1000 nit ML HDR image representation can make up to lOx brighter (glowing) object colors. What one can make with this are e.g. specular reflections on metals, such as a boundary of a metal window frame: in SDR the luminance has to end at 100 nit, making them visually on slightly brighter than the e.g. 70 nit light gray colors of the part of the window frame that does not specularly reflect. In HDR one can make those pixels that reflect the light source to the eye e.g. 900 nit, making them glow nicely giving a naturalistic look to the image as if it were a real scene. The same can be done with fire balls, light bulbs, etc. This regards the definition of images; how a display which can only display whites as bright as 650 nit (the display maximum luminance ML_D) is to actually display the images is an entirely different matter, namely one of display adaptation a.k.a. display tuning, not of image (de)coding. Also the relationship with how a camera captures HDR scene colors may be tight or lose: we will in general assume that HDR colors have already been defined in the HDR image when talking about such technologies as coding, communication, dynamic range conversion and the like. In fact, the original camera-captured colors or specifically their luminances may have been changed into different values by e.g. a human color grader (who defines the ultimate look of an image, i.e. which color triplet values each pixel color of the image(s) should have), or some automatic algorithm. So already for the fact that the chromatic gamut size may stretch with less than a factor 2, whereas the brightness range e.g. luminance range may stretch by a factor 100, one expects different technical rationales and solutions for the two improvement technologies.
[0016] A HDR image may be associated with a metadatum called mastering display white point luminance (MDWPL), a.k.a. ML V. This value, which is typically communicated in metadata of the signal, and is a characterizer of the HDR video images (rather than of a specific display, as it is an element of a virtual display associated specifically with the video, being some ideal intended display for which the video pixel colors have been optimized to be conforming). This is an electable parameter of the video, which can be contemplated as similar to the election of the painting canvas aspect ratio by a painter: first the painter chooses an appropriate AR, e.g. 4: 1 for painting a landscape, or 1: 1 when he wants to make a still life, and thereafter he starts to optimally position all his objects in that elected painting canvas. In an HDR image the creator will then, after having established that the MDWPL is e.g. 5000 nit, make his secondary elections that in a specific scene this lamp shade should be at 700 nit, the flames in the hearth distributed around 500 nit, etc.
[0017] The primary visual aspect of HDR images is (since the blacks are more tricky) the additional bright colors (so one can define ranges with only a MDWPL value, if one assumes the bottom luminance to be fixed to e.g. 0.1 nit). However, HDR image creation can also involve deeper blacks, up to as deep as e.g. 0.0001 nit (although that is mostly relevant for dark viewing environments, such as in cinema theatres).
[0018] The other objects, e.g. the objects which merely reflect the scene light, will be coordinated to be e.g. at least 40x darker in a 5000 nit MDWPL graded video, and e.g. at least 20x darker in a 2000 nit video, etc. So the distribution of all image pixel luminances will typically depend on the MDWPL value (not making most of the pixels very bright). Relative brightness systems can code brighter pixels compared to some reference relative brightness or luminance level. E.g., one may use the level 100% to indicate the classical Lambertian reflection base lighting white of legacy LDR image representations, and denote brighter whites or colors with higher percentages (e.g. 1000% white is lOx brighter than the normal LDR white; and if one were to map the normal LDR white to 100 nit, one would typically map the 1000% white to 1000 nit).
[0019] It may be advantageous for technical systems to not encode the brightness in its native manner, i.e. e.g. a float number, or some N bit representation which varies linearly with the brightness values, e.g. if 255 codes 100 nit, then 128 codes 50 nit instead of e.g. 25 nit. The digital coding of the brightness, involves a technical quantity called luma (Y). We will give luminances the letter L, and (relative) brightnesses the letter B. Note that technically, e.g. for ease of definition of some operations, one can always normalize even 5000 nit WPDPL range luminances to the normalized range [0,1], but that doesn’t detract from the fact that these normalized luminances still represent absolute luminances on a range ending at 5000 nit (in contrast to relative brightnesses that never had any clear associated absolute luminance value, and can only be converted to luminances ad hoc, typically with some arbitrary value).
[0020] For SDR signals the luma coding used a so-called Opto-electronic Transfer Function (OETF). between the optical brightnesses and the electronic typically 8 bit luma codes which (approximately) by definition was:
[0021] Y float = sqrt (B relative). If B relative is a float number ranging from 0 to 1.0, so will Y float.
[0022] Subsequently that signal value Y_float is quantized, because we want 8 bit digital representations, ergo, the Y_dig value that is communicated to receivers over e.g. airways DVB (or internet-supplied video on demand, or blu-ray disk, etc.) has a value between 0 and 255 (i.e. power(2;8)- 1).
[0023] One can show the gamut of all SDR colors (or similarly one can show gamuts of HDR colors, which would if using the same RGB primaries defining the chromatic gamut have the same base, but after renormalization stretch vertically to a larger absolute gamut solid) as in Fig. IB. Larger Cb and Cr values will lead to a (psychovisually more relevant color characterizer) larger saturation (sat), which moves outwards from the unsaturated or colorless colors vertical axis showing the achromatic colors in the middle, towards the maximally saturated colors on the circle (which is a transformation of the usual color triangle spanned by the RGB primaries as vertices). Hues h (i.e. the color category, yellows, versus greens, versus blues) will be angles along the circle. The vertical axis represents normalized luminances (in a linear gamut representation), or normalized lumas (in a non-linear representation (coding) of those luminances, e.g via a psychovisually uniformized OETF or its inverse the EOTF). Since after normalization (i.e. division by the respective MDWPL values, e.g. 2000 for an HDR image of a particular video and 100 for an SDR image), the common representation will become easy, and one can define luminance (or luma) mapping functions on normalized axes as shown in Fig. ID (i.e. which function F_comp re-distributes the various image object pixel values as needed, so that in the actual luminance representation e.g. a dark object looks the same, i.e. has the same luminance, but a different normalized luminance since in one situation a pixel normalized luminance will get multiplied by 100 and in the other situation by 2000, so to get the same end luminance the latter pixel should have a normalized luminance of l / 20thof the former). When down-grading to a smaller range of luminances, one will typically get (in the normalized to 1.0 representation, i.e. mapping a range of normalized input luminances L in between zero and one to normalized output luminances L out) a convex function which everywhere lies above the diagonal diag, but the exact shape of that function F comp, e.g. how fact it has to rise at the blacks, will depend typically not only on the two MDWPL values, but (to have the most perfect re-grading technology version) also on the scene contents of various (video or still) images, e.g. whether the scene is a dark cave, and there is action happening in a shadowy area, which must still be reasonably visible even on 100 nit luminance ranges (hence the strong boost of the blacks for such a scenario, compared to a daylight scene which may employ a near linear function almost overlapping with the diagonal).
[0024] In the representation of Fig. IB one can show the mapping from a HDR color (C H) to an LDR color (C L) of a pixel as a vertical shift (assuming that both colors should have the same proper color, i.e. hue and saturation, which usually is the desired technical requirement, i.e. on the circular ground plane they will project to the same point). Ye means the color yellow, and its complementary color on the opposite side of the achromatic axis of luminances (or lumas) is blue (B), and W signifies white (the brightest color in the gamut a.k.a. the white point of the gamut, with the darkest colors, the blacks being at the bottom). So the receiving side (in the old days, or today) will know it has an SDR video, if it gets this format. The maximum white (of SDR) will be by definition the brightest color that SDR can define. So if one now wants to make brighter image colors (e.g of real luminous lamps), that should be done with a different codec (as one can show the math of the Rec. 709 OETF allows only a coding of up to 1000: 1 and no more).
[0025] So one defined new frameworks with different code allocation functions (EOTFs, or OETFs). What is of interest here is primarily the definition of the luma codes.
[0026] For reasons beyond what is needed for the present discussion, most HDR codecs start by defining an Electro-optical transfer function instead of its inverse, the OETF. Then one can at least basically define brighter (and darker) colors. That as such is not enough for a professional HDR coding system, since because it is different from SDR, and there are even various flavors, one wants more (new compared to SDR coding) technical information relating to the HDR images, which will be metadata.
[0027] The property of those HDR EOTFs is that they are much steeper, to encode a much larger range of needed to be coded HDR luminances, and a significant part of that range coding specifically darker colors (relatively darker, since although one may be coding absolute luminances with e.g. the Perceptual Quantizer (PQ) EOTF (standardized in SMPTE 2084). one applies the function after normalization). In fact if one were to use exact power functions as EOTFs for coding HDR luminances as HDR lumas, one would have a power of 4, or even 7. When a receiver gets a video image signal defined by such an EOTF (e.g. Perceptual Quantizer) it will know it gets a HDR video. It will need the EOTF to be able to decode the pixel lumas in the plane of lumas spanning the image (i.e. having a width of e.g. 4000 pixels and a height of 2000), which will simply be binary numbers. Typically HDR images will also have a larger word length, e.g. 10 bit. However, one should not confuse the non-linear coding one can at will design by optimizing a non-linear EOTF shape with linear codings and the amount of bits needed for them. If one needs to drive, with a linear (bit-represented) code, e.g. a DMD pixel, indeed to reach e.g. 10000: 1 modulation darkest to brightest, one needs to take the log2 to obtain the number of bits. There one would need to have at least 14 bits (which may for technical reasons get rounded upwards to 16 bits), since power(2;14)= 16384 > 10000. But being able to smartly design the shape of the EOTF, and knowing that the visual system sees not all luminance differences equally, the present applicant has shown that (surprisingly) quite reasonable HDR television signals can be communicated with only 8 bit per pixel color component (of course if technically achievable in a system, 10 bits may be better and more preferable). So the receiving side may in both situations get as input a coded pixel color (luma and Cb, Cr; or in some systems by matrixing equivalent non-linear R’G’B’ component values) which lie between 0 and 255, or 0 and 1023, but it will know the kind of signal it is getting (hence what ought to be displayed) from the metadata, such as the metadata (e.g. MPEG VUI metadata) co-communicated EOTF (e.g. a value 16 meaning PQ; 18 means another OETF was used to create the lumas, namely the Hybrid LogGamma OETF, ergo the inverse of that function should be used to decode the luma plane), in many HDR codings the MDWPL value (e.g. 2000 nit), and in more advanced HDR codings further metadata (some may e.g. co-encode luminance -or luma- mapping functions to apply for mapping image luminances from a primary luminance dynamic range to a secondary luminance dynamic range, such as one function FL enc per image).
[0028] We detail the typical needs of an already more sophisticated HDR image handling chain with the aid of simple elucidation Fig. 1 (for a typical nice HDR scene image, of a monster being fought in a cave with a flame thrower, the master grading (Mstr HDR) of which is shown spatially in Fig. 1A, and the range of occurring pixel luminances on the left of Fig. 1C). The master grading or master graded image is where the image creator can make his image look as impressive (e.g. realistic) as desired. E.g., in a Christmas movie he can make a baker’s shop window look somewhat illuminated by making the yellow walls somewhat brighter than paper white, e.g. 150 nit (and real colorful yellow instead of pale yellow), and the light bulbs can be made 900 nit (which will give a really lit Christmas-like look to the image, instead of a dull one in which all lights are clipped white, and not much more bright than the rest of the image objects, such as the green of the Christmas tree).
[0029] So the basic thing one must be able to do is encode (and typically also decode and display) brighter image objects than in a typical SDR image.
[0030] Let’s look at it colorimetrically now. SDR (and its coding and signaling) was designed to be able to communicate any Lambertian reflecting color (i.e. a typical object, like your blue jeans pants, which absorbs some of the infalling light, e.g. the red and green wavelengths, to emit only blue light to the viewer or capturing camera) under good uniform lighting (of the scene where the action is camera- captured). Just like we would do on a painting: if we don’t add paint we get the full brightness reflecting back from the white painting canvas, and if we add a thick layer of strongly absorbing paint we will see a black stroke or dot. We can represent all colors brighter than blackest black and darker than white in a so- called color gamut of representable colors, as in Fig. IB (the “tent”). As a bottom plane, we have a circle of all representable chromaticities (note that one can have long discussions that in a typical RGB system this should be a triangle, but those details are beyond the needs of the present teachings). A chromaticity is composed of a certain (rotation angle) hue h (e.g. bluish-green e.g. “teal”), and a saturation sat, which is the amount of pure color mixed in a grey, e.g. the distance from the vertical axis in the middle which represents all achromatic colors from black at the bottom becoming increasingly bright till we arrive at white. Chromatic colors, e.g. a half-saturated purple, can also have a brightness, the same color being somewhat darker or brighter. However, the brightest color in an additive system can only be (colorless) white, since it is made by setting all color channels to maximum R=G=B=255, ergo, there is no unbalance which would make the color clearly red (there is still a little bit of bluishness respectively yellowishness in the elected white point chromaticity, but that is also an unnecessary further discussion, we will assume D65 daylight white). We can define those SDR colors by setting MDWPL a (relative) 100% for white W (n.b., in SDR white does not actually have a luminance, since legacy SDR does not have a luminance associated with the image, but we can pretend it to be X nit, e.g. typically 100 nit, which is good average representative value of the various legacy SDR tv’s).
[0031] Now we want to represent brighter than Lambertian colors, e.g. the self-luminous flame object (flm) of the flame thrower of the soldier (sol) fighting the monster (mon) in this dark cave. Let’s say we define a (video maximum luminance ML_V) 5000 nit master HDR grading (master means the starting image -most important in this case best quality grading- which we will optimally grade first, to define the look of this HDR scene image, and from which we can derive secondary gradings a.k.a. graded images as needed). We will for simplicity talk about what happens to (universal) luminances, then we can for now leave the debate about the corresponding luma codes out of the discussion, and indeed PQ can code between 1 / 10,000 nit and 10,000 nit, so there is no problem communicating those graded pixel luminances as a e.g. 10 bit per component YCbCr pixelized HDR image, if coding according to that PQ EOTF (of course, the mappings can also be represented, and e.g. implemented in the processing IC units, as an equivalent luma mapping).
[0032] The two dotted horizontal lines represent the limitations of the SDR codable image, when associating 100 nit with the 100% of SDR white.
[0033] Although in a cave, the monster will be strongly illuminated by the light of the flames, so we will give it an (average) luminance of 300 nit (with some spread, due to the square power law of light dimming, skin texture, etc.).
[0034] The soldier may be 20 nit, since that is a nicely slightly dark value, still giving some good basic visibility.
[0035] A vehicle (veh) may be hidden in some shadowy comer, and therefore in a archetypical good impact HDR scene of a cave e.g. have a luminance of 0.01 nit. The flames one may want to make impressively bright (though not too exaggerated). On an available 5000 nit HDR range, we could elect 2500 nit, around which we could still gradually make some darker and brighter parts, but all nicely colorful (yellow and maybe some oranges).
[0036] What would now happen in a typical SDR representation, e.g. a straight from camera SDR image capturing?
[0037] The camera operator would open his iris so that the soldier comes out at “20 nit”, or in fact more precisely 20%. Since the flames are much brighter (note: we didn’t actually show the real world scene luminances, since master HDR video Mstr HDR is already an optimal grading to have best impact in a typical living room viewing scenario, but also in the real world the flames would be quite brighter than the soldier, and certainly the vehicle), they would all clip to maximum white. So we would see a bright area, without any details, and also not yellow, since yellow must have a lower luminance (of course the cinematographer may optimize things so that there is still somewhat of a flame visible even in LDR, but then that is firstly never as impactful extra bright, and secondly at the detriment of the other objects which must become darker).
[0038] The same would also happen if we built a SDR (max. 100 nit) TV which would map equi-luminance, i.e. it would accurately represent all luminances of the master HDR grading it can represent, but clip all brighter object to 100 nit white.
[0039] So the usual paradigm in the LDR era was to relatively map, i.e. the brightest brightness (here luminance) of the received image to the maximum capability of the display. So as this maps 5000 nit by division by 50 on 100 nit, the flames would still be okay since the are spread as yellows and oranges around 50 nit (which is a brightness representable for a yellow, since as we see in Fig. IB the gamut tent for yellows goes down in luminance only a little bit when going towards the most saturated yellows, in contrast to blues (B) on the other side of the slice for this hue angle B-Ye, which blues can only be made in relatively dark versions). However this would be at the detriment of everything else becoming quite dark, e.g. the soldier 20 / 50 nit which is pure black (and this is typically a problem that we see in SDR renderings of such kinds of movie scene).
[0040] So, if having established a good HDR maximum luminance (i.e. ML_V) for the master grading, and a good EOTF e.g. PQ for coding it, we can in principle start communicating HDR images to receivers, e.g. consumer television displays, computers, cinema projectors, etc.
[0041] But that is only the most basic system of HDR.
[0042] The problem is that, unless the receiving side has a display which can display pixels at least as bright as 5000 nit, there is still a question of how to display those pixels.
[0043] Some (DR adaptation) luminance down-mapping must be performed in the TV, to make darker pixels which are displayable. E.g. if the display has a (end-user) display maximum luminance ML_D of 1500 nit, one could somehow try to calculate 1200 nit yellow pixels for the flame (potentially with errors, like some discoloration, e.g. changing the oranges into yellows).
[0044] This luminance down-mapping is not really an easy task, especially to do very accurately instead of sufficiently well, and therefore various technologies have been invented (also for the not necessarily similar task of luminance up-mapping, to create an output image of larger dynamic range and in particular maximum luminance than the input image).
[0045] Typically one wants a mapping function (generically, i.e. used for simplicity of elucidation) of a convex shape in a normalized luminance (or brightness) plot, as shown in Fig. ID. Both input and output luminances are defined here on a range normalized to a maximum equaling one, but one must mind that on the input axis this one corresponds to e.g. 5000 nit, and on the output axis e.g. 200 nit (which to and for can be easily implemented by division respectfully multiplication). In such a normalized representation the darkest colors will typically be too dark for the grading with the lower dynamic range of the two images (here for down-conversion shown on the vertical output axis, of normalized output luminances L out, the horizontal axis showing all possible normalized input luminances L_in). Ergo, to have a satisfactory output image corresponding to the input image, we must relatively boost those darkest luminances, e.g. by multiplying by 3x, which is the slope of this luminance compression function F comp for its darkest end. But one cannot boost forever if one wants no colors to be clipped to maximum output, ergo, the curve must get an increasingly lower slope for brighter input luminances, e.g. it may typically map input 1.0 to output 1.0. In any case the luminance compression function F comp for down-grading will he above the 45 degree diagonal (diag) typically.
[0046] Care must still be taken to do this correctly. E.g., some people like to apply three such compressive functions to the three red, green and blue color channels separately. Whilst this is a nice and easy guarantee that all colors will fit in the output gamut (an RGB cube, which in chromaticity-luminance (L) view becomes the tent of Fig. IB) especially with higher non-linearities it can lead to significant color errors. A e.g. reddish orange hue is determined by the percentage of red and green, e.g. 30% green and 70% red. If the 30% now gets doubled by the mapping function, but the red stays in the feeble-sloped part of the mapping function almost unchanged, we will have a 60 / 70, i.e. 50 / 50 i.e. a yellow instead of an orange. This can be particularly annoying if it depends on (in contrast to the SDR paradigm) non-uniform scene lighting, e.g. an sports car entering the shadows suddenly turning yellow.
[0047] Ergo, whilst the general desired shape for the brightening of the colors may still be the function F_comp (e.g. determined by the video creator, when grading a secondary image corresponding to his master HDR image already optimally graded), one wants a more savvy down-mapping. As shown in Fig. IB, for many scenarios one may desire a re-grading which merely changes the brightness of the normalized luminance component (L), but now the innate type of color, i.e. its chromaticity (hue and saturation). If both SDR and HDR are represented with the same red, green and blue color primaries, they will have a similarly shaped gamut tent, only one being higher than the other in absolute luminance representation. If one scales both gamuts with their respective MDWPL values (e.g. MDWPL1= 100 nit, and MDWPL2= 5000 nit), both gamuts will exactly overlap. The desired mapping from a HDR color C H to a corresponding output SDR color C L (or vice versa) will simply be a vertical shifting, whilst the projection to the chromaticity plane circle stays the same.
[0048] Although the details of such approaches are also beyond the need of the present application, we have thought examples of such color mapping mechanism before, where the three color components are processed coordinately, although in a separate luminance and chroma processing path, e.g. in WO2017157977.
[0049] If it is now possible to down-grade with one (or more) luminance mapping functions (the shape of which may be optimized by the creator of the video(s)), in case one uses invertible functions one can design a more advanced HDR codec.
[0050] Instead of just making some final secondary grading from the master image, e.g. in a television, one can make a lower dynamic range image version for communication, communication image Im comm. We have elected in the example this image to be defined with its communication image maximum luminance ML_C equal to 200 nit. The original 5000 nit image can then be reconstructed (a.k.a. decoded) as a reconstructed image Rec HDR (i.e. with the same reconstructed image maximum luminance ML REC) by receivers, if they receive in metadata the decoding luminance mapping function FL dec, which is typically substantially the inverse of the coding luminance mapping function FL enc, which was used by the encoder to map all pixel luminances of the master HDR image into corresponding lower pixel luminances of the communication image Im comm. So the proxy image for communicating actually an image a higher dynamic range (DR_H, e.g. spanning from 0.001 nit to 5000 nit) is an image of a different, lower dynamic range (DR_L).
[0051] Interestingly, one can even elect the communication image to be a 100 nit LDR (i.e. SDR) image, which is immediately ready (without further color processing) to be displayed on legacy LDR images (which is a great advantage, because legacy displays don’t have HDR knowledge on board). How does that work? The legacy TV doesn’t recognize the MDWPL metadatum (cos that didn’t exist in the SDR video standard, so the TV is also not arranged to go look for it somewhere in the signal, e.g. in a Supplemental Enhancement Information message, which is MPEG’s mechanism to introduce all kinds of pre-agreed new technical information). It is also not going to look for the function. It just looks at the YCbCr e.g. 1920x1080 pixel color array, and displays those colors as usual, i.e. according to the SDR Rec. 709 interpretation. And the creator has chosen in this particular codec embodiment his FL enc function so that all colors, even the flame, map to reasonable colors on the limited SDR range. Note that, in contrast to a simple multiplicative change corresponding to the opening or shutting of a camera iris in an SDR production (which typically leads to clipping to at least one of white and / or black), now a very complicated optimal function shape can be elected, as long as it is invertible (e.g. we have taught systems with first a coarse pre-grading and then a fine-grading). E.g. one can move the luminance (respectively relative brightness) of the car to a level which is just barely visible in SDR, e.g. 1% deep black, whilst moving the flame to e.g. 90% (as long as everything stays invertible). That may seem extremely daunting if not impossible at first sight, but many field tests with all kinds of video material and usage scenarios have shown that it is possible in practice, as long as one does it correctly (following e.g. the principles of WO2017157977).
[0052] How do we now know that this is actually a HDR video signal, even if it contains an LDR-usable pixel color image, or in fact that any HDR-capable receiver can reconstruct it to HDR: because there are also the functions FL_dec in metadata, typically one per image. And hence that signal codes what is also colorimetrically, i.e. according to our above discussion and definition, a (5000 nit) HDR image.
[0053] Although already more complex than the basic system which communicates only a PQ- HDR image, this per SDR proxy coding is still not the best future-proof system, as it still leaves the receiving side to guess how to down-map the colors if it has e.g. a 1500 nit, or even a 550 nit, tv.
[0054] Therefore we added a further technical insight, and developed so-called display tuning technology (a.k.a. display adaptation): the image can be tuned for any possible connected tv, i.e. any ML_D, because one can double the function of the coding function FL enc as some guidance function for the up-mapping from 100 nit Im_comm not to a 5000 nit reconstructed image, but to e.g. a 1500 nit image. The concave function, which is substantially the inverse of F comp (note, for display tuning there is no requirement of exact inversion as there is for reconstruction), will now have to be scaled to be somewhat less steep (i.e. from the reference decoding function FL_dec a display adapted luminance mapping function FL DA will be calculated), since we expand to only 1500 nit instead of 5000 nit. I.e. an image of tertiary dynamic range (DR T) can be calculated, e.g. optimized for a particular display in that the maximum luminance of that tertiary dynamic range is typically the same as the maximum displayable luminance of a particular display.
[0055] Techniques for this are described in W02017108906 (we can transform a function of any shape into a similarly-shaped function which lies closer to the 45 degree diagonal, by an amount which depends on the ratio between the maximum luminances of the input image and the desired output image, versus the ratio of the maximum luminances of the input image and a reference image which would here be the reconstructed image, by e.g. using that ratio to obtain closer points on for all points on the diagonal orthogonally projecting a line segment from the respective diagonal point till it meets a point on the input function, which closer points together define the tuned output function, for calculating the to be displayed image Im disp luminances from the hn comm luminances).
[0056] Not only did we get more kinds of displays even for basic movie or television video content (LCD tv, mobile phone, home cinema projector, professional movie theatre digital projector), and more video different sources and communication media (satellite, streaming over the internet, e.g. OTT, streaming over 5G), but also did we get more production manners of video.
[0057] Fig. 2 shows -in general, without desiring to be limiting- a few typical creations of video where the present teachings may be usefully deployed.
[0058] In a studio environment, e.g. for the news or a comedy, there may still be a tightly controlled shooting environment (although HDR allows relaxation of this, and shooting in real environments). There will be controlled lighting (202), e.g. a battery of base lights on the ceiling, and various spot lights. There will be a number of bulky relatively stationary television cameras (201). Variations on this often real-time broadcast will be e.g. a sports show like soccer, which will have various types of cameras like near-the-goal cameras for a local view, overview cameras, drones, etc.
[0059] There will be some production environment 203, in which the various feeds from the cameras can be selected to become the final feed, and various (typically simple, but potentially more complex) grading decisions can be taken. In the past this often happened in e.g. a production truck, which had many displays and various operators, but with internet-based workflows, where the raw feeds can travel via some network, the final composition may happen at the premises of the broadcaster. Finally, when simplifying the production for this elucidation, some coding and formatting for broadcast distribution to end (or intermediate, such as local cable stations) customers will happen in formatter 204. This will typically do the conversion to e.g. PQ YCbCr from the luminances as graded as explained with Fig. 1, for e.g. an intermediate dynamic range format, calculate and format all the needed metadata, convert to some broadcasting format like DVB or ATSC, packetize in chunks for distribution, etc. (the etc. indicating there may be tables added for signaling available content, sub-titling, encryption, but at least some of that will be of lesser interest to understand the details of the present technical innovations).
[0060] In the example the video (a television broadcast in the example) is communicated via a television satellite 250 to a satellite dish 260 and a satellite signal capable set-top-box 261. Finally it will be displayed on an end-user display 263.
[0061] This display may be showing this first video, but it may also show other video feed, potentially even at the same time, e.g. in Picture -in-Picture windows (or some data of the first HDR video program may come via some distribution mechanism and other data via another).
[0062] A second production is typically an off-line production. We can think of a Hollywood movie, but it can also be a show of somebody having a race through a jungle. Such a production may be shot with other optimal cameras, e.g. steadicam 211 and drone 210. We again assume that the camera feeds (which may be raw, or already converted to some HDR production format like HLG) are stored somewhere on network 212, for later processing. In such a production we may in the last months of production have some human color grader use grading equipment 213 to determine the optimal luminances (or relative brightnesses in case of HLG production and coding) of the master grading. Then the video may be uploaded to some internet-based video service 251. For professional video distribution this may be e.g. Netflix.
[0063] A third example is consumer video production. Here the user will have e.g. when making a vlog a ring lighter 221, and will capture via a mobile phone 220, but (s)he may also be capturing in some exterior location without supplementary lighting. She / he will typically also upload to the internet, but now maybe to YouTube, or TikTok, etc.
[0064] In case of reception via the internet, the display 263 will be connected via a modem, or router 262 or the like (more complicated setups like in-house Wi-Fi and the like are not shown in this mere elucidation).
[0065] Another user may be viewing the video content on a portable display (271), such as a laptop (or similarly other users may use a non-portable desktop PC), or a mobile phone etc. The may access the content over a wireless connection (270), such as Wi-Fi, 5G, etc.
[0066] So it can be seen that today, various kinds of video, in various technical codings, can be generated and communicated in various manners, and our coding and processing systems have been designed to handle substantially all those variants.
[0067] Fig. 3 shows an example of an absolute (nit-level-defined) dynamic range conversion circuit 300 for a (HDR) image or video decoder shown in a video processing circuit chain in Fig. 3C. (The encoder would typically work similarly but with inverted functions typically, i.e. the function to be applied being the function of the other side mirrored over the diagonal). It is based on coding a primary image (e.g. a master HDR grading) with a primary luminance dynamic range (DR Prim) as another (so- called proxy) image with a different secondary range of pixel luminances (DR Sec). If the encoder and all its supply-able decoders have pre-agreed or know that the proxy image has a maximum luminance of 100 nit, this need not be communicated as an SDR WPL metadatum. If the proxy image is e.g. a 200 nit maximum image, this will be indicated by filling its proxy white point luminance P WPL with the value 200, or similarly for 80 nit etc. The maximum of the primary image (HDR_WPL= 1000), to be reconstructed by the dynamic range conversion circuit, will normally be co-communicated as metadata of the received input image, or video signal, i.e. together with the input pixel color triplets (Y in, Cb_in, Cr in). The various pixel lumas will typically come in as a luma image plane, i.e. the sequential pixels will have first luma Y 11 , second Y21 , etc . (typically these will be scanned, and the dynamic range conversion circuit will convert pixel by pixel to output pixel color triplets (Y out, Cb out, Cr out). We will primarily focus on the brightness dimension of the pixel colors in this elucidation. Various dynamic range conversion circuits may internally work differently, to achieve basically the same thing: a correctly reconstructed output luminance L out for all image pixels (the actual details don’t matter for this innovation, and the embodiments will focus on teaching only aspects as far as needed). The mapping of luminances from the secondary dynamic range to the primary dynamic range may happen on the luminances themselves, but also on any luma representation (i.e. according to any EOTF, or OETF), provided it is done correctly, e.g. not separately on non-linear R’G’B’ components. The internal luma representation need not even be the one of the input (i.e. of Y in), or for that manner of whatever output the dynamic range conversion circuitry or its encompassing decoder may deliver (e.g. a format luma Y sigfin for a particular communication format or communication system, “communicating” including storage to a memory, e.g. inside a PC, a hard disk, an optical storage medium, etc.).
[0068] We have optionally (dotted) shown a luma conversion circuit 301, which turns the input lumas Y in into perceptionally uniformized lumas Y_pc.
[0069] Applicant standardized in ETSI 103433 a useful equation to convert luminances in any range to such a perceptual luma representation:
[0070] Y_pc=log_10{ l+[RHO(WPL_inrep)-l]*power(Ln_in; 1 / (2.4)) } / log_10{ RHO(WPL_inrep)}
[0071] [Eq. 1-1]
[0072] In which the function RHO is defined as
[0073] RHO(WPL_inrep) = l+32*power{( WPL_inrep / 10000); l / (2.4)} [Eq. 1-2]
[0074] The value WPL inrep is the maximum luminance of the range that needs to be converted to psychovisually uniformized lumas, so for the 100 nit SDR image this value would be 100, and for the to be reconstructed output image (or the originally coded image at the creation side) the value would be 1000.
[0075] Ln_in are the luminances along that whichever range which need to be converted, after normalization by dividing by its respective maximum luminance, i.e. within range [0,1],
[0076] Once we have an input and an output range normalized to 1.0, we can apply a luminance mapping function actually in the luma domain, as shown inside the luma mapping circuit 302, which does the actual luma mapping for each incoming pixel.
[0077] In fact, this mapping function had been specifically chosen by the encoder of the image (at least for yielding good quality reconstructability, and maybe also a reduced amount of needed bits when MPEG compressing, but sometimes also fulfilling further criteria like e.g. the SDR proxy image being of correct luminance distribution for the particular scene -a dark cave, or a daytime explosion- on a legacy SDR display, etc.). So this function F dec (or its inverse) will be extracted from metadata of the input image signal or representation, and supplied to the dynamic range conversion circuit for doing the actual per pixel luma mapping. In this example the function F dec directly specifies the needed mapping in the perceptual luma domain, but other variants are of course possible, as the various conversions can also be applied on the functions. Furthermore, although for simplicity of explanation, and to guarantee the teaching is understood, we teach here a pure decoder dynamic range conversion, but other dynamic range conversions may use other functions, e.g. a function derived from F_dec, etc. The details of all of that are not needed for understanding the present innovative contribution to the technology.
[0078] In general one will not only change the luminances, but there will be a corresponding change in the chromas Cb and Cr. That can also be done in various manners, from strictly inversely decoding, to implementing additional features like a saturation boost, since Cb and Cr code the saturation of the pixels. Thereto another function is typically communicated in metadata (recoloring specification function FCOL), which determines the chromatic recoloring behavior, i.e. the mapping of Cb and Cr (note that Cb and Cr will typically be changed by the same multiplicative amount, since the ratio of Cr / Cb determines the hue, and generally one does not want to have hue changes when decoding, i.e. the lower and higher dynamic range image will in general have object pixels of different brightness, and oftentimes at least some of the pixels will have different saturation, but ideally the hue of the pixels in both image versions will be the same). This color function will typically specify a multiplier which has a value dependent on a brightness code Y (e.g. the Y_pc, or other codes in other variants). A multiplier establishment circuit 305 will yield the correct multiplier m for the brightness situation of the pixel being processed. A multiplier 306 will multiply both Cb in and Cr in by this same multiplier, to obtain the corresponding output chromas Cb_out= m*Cb_in and Cr_out=m*Cr_in. So the multiplier realizes the correct chroma processing, therefore the whole color processing of any dynamic range conversion being correctly configurable in the dynamic range conversion circuit.
[0079] Furthermore, there may typically be (at least in a decoder) a formatting circuit 310, so that the output color triplet (Y out, Cb out, Cr out) can be converted to whatever needed output format (e.g. an RGB format, or a communication YCbCr format, Y sigfin, Cb_ sigfin, Cr sigfm). E.g. if the circuit outputs to a version of a communication channel 379 which is an HDMI cable, such cables typically use PQ-based Y CbCr pixel color coding, ergo, the lumas will again be converted from the perceptual domain to the PQ domain by the formatting circuit.
[0080] It is important that the reader well understands what is a (de)coding, and how an absolute HDR image, or its pixel colors, is different from a legacy SDR image. There may be a connection to a display tuning circuit 380, which calculates ultimate pixel colors and luminances to be displayed at the screen of some display, e.g. a 450 nit tv which some consumer has at home.
[0081] However, in absolute HDR, one can establish pixel luminances already in the decoding step, at least for the output image (here the 1000 nit image).
[0082] We have shown this in Fig. 3B, for some typical HDR image being an indoors / outdoors image (the geometry and comprised image objects of which are shown in Fig. 3A). Note that, whereas in the real world the outdoors luminances may typically be 100 times brighter than the indoors luminances, in an actual master graded HDR image it may be better to make them e.g. lOx brighter, since the viewer will be viewing all together on a screen, in a fixed viewing angle, even typically in a dimly illuminated room in the evening, and not in the real world.
[0083] Nevertheless, we find that when we look at the luminances corresponding to the lumas, e.g. the HDR luminances L out, we typically see a large histogram (of counts N(L_out) of each occurring luminance in an output image of this homely scene). This spans considerably above some lower dynamic range lobe of luminances, and above the low dynamic range 100 nit level, because the sunny outdoors images have their own histogram lobe. Note that the luminance representation is drawn non-linearly, e.g. logarithmically. We can also trace what the encoder would do at the encoding side, when making the 100 nit proxy image (and its histogram of proxy luminance counts N(L_in)). A convex function, as shown in Fig. 1, or inside luma mapper 302, is used which squeezes in the brighter luminances, due to the limitations of the smaller luminance dynamic range. There is still some difference between the brightness of indoors and outdoors, and still a considerable range for the upper lobe of the outdoors objects, so that one can still make the different colors needs to color the various objects, such as the various greens in the tree. However, there must also be some sacrifices. Firstly the indoors objects will display (assuming for the moment an 100 or 200 nit display would faithfully display those luminances as coded, and not e.g. do some arbitrary beautification processing which brightens them) darker, darker than ideally desired, i.e. up to the indoors threshold luminance T in of the HDR image. Secondly, the span of the upper lobe is squeezed, which may give the outdoors objects less contrast. Thirdly, since bright colors in the tentshaped gamut as shown in Fig. 1 cannot have large saturation, the outdoors colors may also be somewhat pastellized, i.e. of lowered saturation. But of course if the grader at the creation side has control over all the functions (F_enc, FCOL), hey may balance those features, so that some have a lesser deviation at the detriment of others. E.g. if the outdoors shows a plain blue sky, the grader may opt for making it brighter, yet less blue. If there was a beautiful sunset, he may want to retain all its colors, and make everything dimmer instead, in particular if there are no important dark comers in the indoors part of the image, which would then have their contents badly visible, especially when watching tv with all the lights on (note that there are also techniques for handling illumination differences and the visibility of the blacks, but that is too much information for this patent application’s elucidation).
[0084] The middle graph shows what the lumas would look like for the proxy luminances, and that may typically give a more uniform histogram, with e.g. approximately the same span for the indoors and outdoors image object luminances. The lumas are however only relevant to the extent of coding the luminances, or in case some calculations are actually performed in the luma domain (which has advantages for the size of the word length on the processing circuitry). Note that whereas the absolute formalism can allocate luminances on the input side between zero and 100 nit, one can also treat the SDR luminances as relative brightnesses (which is what a legacy display would do, when discarding all the HDR knowledge and communicated metadata, and looking merely at the 0-255 luma and chroma codes).
[0085] Fig. 4 elucidates how the computer world looked at HDR, in particular for the calculation of HDR scenes, e.g. by ray-tracing (i.e. the equivalent of actual camera capturing). This technology, which can be used in e.g. gaming, in which a different view on an environment has to be re-calculated each time a player moves, potentially with a specular reflection appearing that wasn’t in view when the player was positioned one game meter to the left, was not primarily geared for communication, e.g. from a broadcaster to a receiver, any receiver, with any of various kinds of displays with different maximum brightness characteristics. These calculations and representations originated for a typical application (tightly managed) inside a single PC, and were certainly not developed with a view on many future quite different applications. When one actually calculates a HDR image, by defining e.g. a very bright light source and (internally) physically modeling how it illuminates a pixel of a mirror, one can define just any HDR luminance. But one of the problems is already one can calculate it, but not necessarily display it. E.g., one would typically use a floating point representation (e.g. 16 bit half-float for representing the luminance, in a logarithmic format, i.e. able to represent almost infinite luminances, basically ridiculously large from the point of view of actually using for display). One would then consider all luminances lower than 1 (which can also be stated as the relative 100%) as the normal SDR luminances. It may not be the best way to treat, e.g. display, an HDR float image, but some systems or components, or software would indeed clip everything above 1 to 1.0, and just use the lower luminances. The outdoors objects, like the house, when being computationally generated, may have luminances a multiplication factor higher than 1.0, e.g. 5x or 20x (if one were to associate 1.0 with 100 nit, which was not necessarily done in the computer view, one could say these would be e.g. 500 nit). One could make the sun e.g. as bright as 1 million nit, which would need severe down-mapping to make it visible as a light yellow sphere, or, it would typically clip to white (which would not be an issue, since also in the real world the sun is so bright that it will clip to white in camera capturings). However, there may be issues with other objects in the down-mapping.
[0086] One difference one already sees with e.g. the PQ representation is that this representation is closed (one assumes every luminance one is ever going to need in practice may reasonably fall below its maximum being 10,000 nit), and therefore one can do e.g. luminance mappings based on this endpoint, whereas the computer log format is -pragmatically- open, in the sense that it can go almost infinitely above 1. This may involve some (undesirable) clipping. One could say that even for a 16 bit logarithmic format there is some end-point, but given that luminance or brightness will be extremely high, that is much less interesting to map, than mapping the values around 1.0 (compressing e.g. linearly from 1 million nit would make all indoors pixels pitch black). So open-ended representations need a somewhat different luminance mapping (called tone mapping in this sub-area), but on a more general level one can come to some common ground approach, at least in the sense that both sub-fields of technology have been able to create LDR images (and similarly, in the modem area, one can also create (common) HDR images for both).
[0087] Computer games, which were until recently played on SDR displays mostly anyway, despite making beautiful lifelike images, needed to down-map those to yield SDR images. One could argue that would not be so different from the down-mapping of natural, camera-captured HDR images, which was not typical until recently anyway. But the tone mappers were sometimes difficult to fathom, and somewhat ad hoc. Nevertheless, HDR gaming did become possible, even with some introductory pains, and people loved the look. But having things sorted out for, like HDR television, another tightly controlled application, doesn’t mean yet that all HDR problems are solved in the computer world.
[0088] US2017 / 0347113 describes a system wherein a HDR video supply apparatus (e.g. typically a blu-ray player) will communicate HDR video images to a display, and the display still has to do luminance mapping for dynamic range conversion (e.g. the video has pixel luminances up to 5000 nit, and the display can display only up to 1200 nit, i.e. all bright pixels in the video images need to be converted to below or equal to 1200 nit). That can be done optimally by co-communicating a downgrading function, which the video creator can optimize to do -for each image in his video- an optimal down-grading. That would work perfectly if it is a global luminance mapping function, i.e. wherein the output luminance only depends on the value of the luminance of an input pixel (i.e. in the 5000 nit ML_V image that gets communicated by the BD player to the display), and not on the position of the pixel in the image, e.g. whether it resides at the center, or somewhere in the top-left quadrant, and that position having also an impact on how a pixel of the same input luminance would get a different output luminance. The present applicant developed a method how one could easily specify such a locally variable luminance mapping. What the receiving apparatus has to establish is “how to map”, “where”. The how can be the same as for a global mapping: one can just define a function for any input luminance or luma a pixel could have, i.e. e.g. 0-1023, irrespective of whether there actually are pixels of luma e.g. 12 in the location to be processed. The where can be communicated by adding additional metadata, which specifies coordinates of an encompassing rectangle, and a selection criterion of pixels in that rectangle, which would get selected only if they are brighter than L below and darker than L upper. With this simple metadata one can define a surprising amount of image situations. Such a locally variant re-grading of a 5000 nit image from the BD player to a 1200 nit image suitable for the connected display is also not problematic, since the display can determine which pixels to process by the basic function F_G1, and which pixels to identify for processing with the (one or more) local function F_L1.
[0089] A problem is however that blu-ray players may want to implement (also in the much more complex HDR video era) such advanced features as e.g. a director’s comment in a picture -in-picture area. The BD player then composes and sends images where together with the action of the movie, the director is giving his comments about how he came up with or shot this scene. Furthermore, the director’s comments tracks may be stored on the BD in a different resolution than the resolution of the PIP in which they are going to be placed, and the user may even be able to select how large the PIP window should be and / or where it should be positioned. If you have two global luminance mapping functions for both the main and the auxiliary PIP-ed video in the metadata track, that may still not be very challenging, since at least some blu-ray players may use the local processing specification mechanism to indicate where the PIP is localized, and then on all those pixels the FL_1 function will be applied for down-grading instead of the FG1 function, as will be readily understood by the display when receiving a coded video image according to the normal above described local processing coding formalism. However, some blu-ray players may not be that savvy, and although making a composite image, they may just pass the original metadata as is. The display would then need to figure out as best what is going on. The simpler embodiments may just communicate a single bit that PIP-ing has occurred. Then displays at least know they may need to be careful in their processing. What one can also expect that a blu-ray player could easily implement without excessive cost, is that, if it has already defined a rectangle for PIP-ing, it can at least communicate that rectangle’s coordinates. Without this invention, especially when the auxiliary video has some local processing defined, e.g. make the pixels in some circle maximally bright cos that is the sun, the original creator of the BD disk would not know about the actual scaling, or repositioning of the PIP, so he would just define the encompassing rectangle in the original format. (Although director’s comments may have lower resolution) let’s suppose the director’s comment track is also 4K resolution, just like the main movie. If the sun is in the upper-left quadrant, its rectangle may be specified as e.g. (200,200)-(250, 250), being its top-left and bottom-right coordinates. But if one shows that video PIP-ed on the top-right of the original movie, the wrong pixels will get brightened. But having the coordinates of the PIP, the display can calculate where the new scaled sun in the PIP would be, and apply the correct functions to the correct pixels. This technique doesn’t teach or inspire anything regarding common color representation coordination of software applications with an operating system, even if the blu-ray disk were to be played on a computer (as in the present not yet highly HDR savvy computer systems the composed PIP-ed video would directly be sent to a dedicated video core of the GPU, which takes care of basic video processing according to basic HDR techniques, as well as the metadata in any form).
[0090] Television versus computing environment
[0091] Although there are still many flavors (and technical visions, e.g. the absolute PQ-based coding versus the relative HLG-based coding), as shown above at least for the simple video communication (e.g. for broadcast television services) the situation is relatively simple and akin to the “direct display control line” situation of the analog PAL era (which used the display paradigm: the brightest code in the image -white- gets displayed as the brightest thing on the display, and perceived by the viewer as the whitest white). It is generalized beyond the exact direct display (of square root mapping followed by inverse square power mapping of SDR systems), as there may be e.g. complex display adaptation involved, but it is still similar in the sense that one video takes total area of view of a display, and will in general be optimized for one display of an end-consumer, to be thereafter directly displayed as sole asset on the end-user display.
[0092] Computer systems may be more complex, as there may be various unrelated visual assets, and the viewer may also be using different displays (e.g. an old SDR one, and a new one, and show at least some of the content on one of the displays, but he may also change the display for that content by dragging e.g. its window to the other display). By visual asset we mean a set of pixels which together form a complete imagery for conveying some message to a viewer, complete in the sense that nothing still related to that imagery is missing or left out, e.g. an image of a movie or the widget components making up a scrolling window. Furthermore, whereas a video processor / processing is usually provided by one manufacturer, and usually follows well-standardized principles, a computer may be running applications from just about anybody, and those manufacturers (e.g. of software and its look and feel) may have different visions about higher luminance representation and / or display in more than just minor details.
[0093] Also, if this issue would have been handled in the previous century, one could still argue that the physics of the at the time universal Cathode Ray Tube monitors would put some limit on the wildless of higher brightness colors that various manufacturers / suppliers could be using, but we now also may want to cater for very dissimilar types of displays in very dissimilar viewing scenarios (e.g. one may want to display the same content on a small LCD-based mobile phone watched on the train, a projector in a darkened home cinema in a consumer’s attic, LED panels in a supermarket, etc.). It seems that also the underlying circuits and software may sometimes be multiplying instead of converging: in addition to the classical operating systems (abbreviated herein as OS) Linux, Windows, Android, one may now have proprietary operating systems (e.g. Tizen) that, even when derived from a basic operating system may have some differential behavior regarding high brightness colors at least in one sense or condition.
[0094] A computer may also be seen as a “kit of parts”, and that may be good for its general usability for a myriad of tasks, but that doesn’t necessarily mean that these parts would work together in a stable well-defined manner.
[0095] The problems described herein, and the solutions offered by the embodiments will be similar just as well for native applications (which are specifically written for a specific computer platform, and run on that computer platform) and for the currently popular web applications, which are served from a remote server, i.e. have a number of services from that remote server, yet may rely for some computations (e.g. execute downloaded JavaScript) on the local client computer, and may need to do so for certain steps of the program. To the customer it does not seem to make much difference whether his spreadsheet is installed locally, or running on the cloud (except maybe for the subscription fee). He may not even realize that if he is typing an email in Gmail running in a browser, that under the hood he is actually making use of internet technologies, as he is writing html lines. The difference between a local file browser and an internet file browser is becoming less. But although the user may not care or want to see what exactly is going on to produce a well-working and visually pleasing result, having various parties work (and decide) on assets makes for big question who is in charge of the colorimetry. When any service running over the internet, those systems may want to make use of ever more advanced features, which may need to take recourse to details of the client’s computing device, under the hood. E.g., a server-controlled internet game, may use a farm of processors to calculate the behavior of virtual actors, yet want to benefit from the hardware acceleration of calculations on the client’s GPU, e.g. for doing the final shading (with shading in the computer sense we mean the calculations needed to come to the correct colors including brightnesses of an object, e.g. using a simulation of illumination of some object with some physical texture like a tapestry, or interpolation of colors of vertices of a triangle, which pixel colors typically end up as red, green and blue color components in a color buffer; the name rendering can also be used, but means the higher level whole process of coming to an image of an object from a viewpoint, for any generation of a NxM matrix of pixels in a so-called canvas, which is a memory in which to put finally rendered pixels, taking into account also such aspects as visibility and occlusion). Whereas we do not want to imply any limitations, we will elucidate some technical details of our innovations with some more challenging web application embodiments.
[0096] There are particular problems when several visual assets, of different brightness dynamic range, need to be displayed in a coordinated manner, such as e.g. in a window system on a computer (and possibly more challenging on several screens, of possibly different brightness capability). This is already a difficult technical problem per se, and the fact that in practice many different components hence producers / companies are involved (various software or middleware applications, different operating systems controlled by companies like Microsoft, Apple or Google, different Graphics Processing Unit vendors, etc.) does not necessarily make things easier. In this text with visual asset we do not necessarily mean a displayable thing in an area that has been prepared previously, and e.g. stored in a memory, but also visual assets that can be generated on the fly, e.g. a text with HDR text colors generated as it is being typed, into some window, and to end up in some canvas which collects all the visual assets in the various windows, to ultimately get displayed after traveling through the whole processing pipe which happens to be in place in the computer being configured with that particular set of applications presenting their visual assets.
[0097] To be clear, what we mean by graphics processing unit is a circuit (or maybe in some systems circuits plural, if there are two or more separate GPUs being supplied with the visual assets to be displayed in totality) which contain the final buffering for such pixels that should be shown on a display, i.e. be scanned out to at least one display (and not processors which do not have this scanout buffer and circuitry for communicating the video signals out to display(s); this GPU is sometimes also called Display Processing Unit DPU).
[0098] So there is a need for a universal well-coordinatable approach of handling HDR (and possibly some SDR) visual assets on computers, which may have several applications running in parallel, none of them controlling the entire visible screen, and which may work (in the sense of at least sending their preferred pixel colors to the GPU scanout buffer) through various layers of software, middleware, APIs etc. (i.e. which may communicate directly to an operating system, or via other processes, and typically via calls which can contain and communicate configurable data).
[0099] SUMMARY OF THE INVENTION
[0100] The indicated problems are handled by a computer (700) arranged to run an operating system (504), wherein the operating system is arranged to manage color coordination of visual assets received from a first software application (701) and a second software application (702), wherein the operating system comprises a geometrical management unit (71 l)arranged to manage spatial positioning in a canvas of the visual assets and arranged to receive a first visual asset (Assl) from the first software application comprising pixels of which a first pixel has a first pixel color specifying a first original brightness (Lui) lying in a first brightness range (DR Hl) which has a first maximum brightness (ML_Vorl), and arranged to receive a second visual asset (Ass2) from the second software application comprising pixels of which a second pixel has a second pixel color specifying a second original brightness (Lu2) lying in a second brightness range (DR H2) which has a second maximum brightness (ML_Vor2), wherein the first maximum brightness is different from the second maximum brightness; wherein the first visual asset (Assl) is received comprising a set of first pixel color codes, which are defined in a first common color space pre-agreed between the first software application and the operating system, wherein a first pixel color code represents a first representative brightness (LuRl), which is derived from the first original brightness by applying a first brightness mapping function (TM1) to the first original brightness, which yields the first representative brightness as output of the first brightness mapping function, and wherein the second visual asset (Ass2) is received comprising a set of second pixel color codes, which are defined in a second common color space pre-agreed between the second software application and the operating system, wherein a second pixel color code represents a second representative brightness (LuR2), which is derived from the second original brightness by applying a second brightness mapping function (TM2) to the second original brightness, which yields the second representative brightness as output of the second brightness mapping function, wherein the first common color space may be the same as or different from the second common color space, wherein the first representative brightness lies in a third brightness range (DR_L1) which has a third maximum brightness (ML_passl) which is different from the first maximum brightness (ML_Vorl), and wherein the second representative brightness lies in a fourth brightness range (DR_L2) which has a fourth maximum brightness (ML_pass2) which is different from the second maximum brightness (ML_Vor2); wherein the first visual asset comprises first color transformation control data (CTctrl) which comprises at least one of the first brightness mapping function (TM1) or its inverse, and a first reference white level (WL1), which is lower than the third maximum brightness (ML_passl), and wherein the second visual asset comprises second color transformation control data (CTctrl) which comprises at least one of the second brightness mapping function (TM2) or its inverse, and a second reference white level (WL2), which is lower than the fourth maximum brightness (ML_pass2); wherein the operating system comprises a brightness mapping unit (710) arranged to receive the first visual asset and to apply a third brightness mapping function (FL opl) to the first representative brightness to obtain an output luminance (Lol), or a corresponding output luma code (Y’o) which codes the output luminance (Lol), wherein the output luminance (Lol)lies in a final brightness range (ML COMP) which is different from the third brightness range (DR L1), wherein the output luma codes a brightness of an output color, wherein the third brightness mapping function (FL opl) is based on at least one of the first brightness mapping function (TM1) and the first reference white level (WL1); wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (hnFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
[0101] By software application we mean a set of computer codes which comprises functionality to visually show something related to the software processing to a human (e.g. an output graph of some measurements, or an image of a piece of clothing one can shop for). A first and second software application running concurrently on a same processing system -in the sense that they will show their visual asset(s) together on a same one or more displays- may typically come from anywhere (e.g. some server), and be defined by any creator. E.g., one software application may be native i.e. fixedly installed and always running in the background on the computing system, and the second application may e.g. be a dynamic webpage being visited, which by itself could be presenting various sources of imagery from various sources (i.e. various internal assets), e.g. commercial pop-ups or in-line ads and the like).
[0102] The new technical effect hitherto unmanageable is that this approach of designing in particular the operating system construction and its configurable communication with (coordinated) applications and their visual color asset desiderata, by means of the new color transformation control data (CTctrl) related to common pixel color definitions pre-agreed or pre-agreeable by software applications and the operating system, makes it possible that the operating system creates better coordinated colors for the one or more visual assets, for ultimate display. So the final displayed look of all the visual assets together will be more controllable towards a better looking total output picture. Unit in terms of software code may typically mean a sub-process, which will run on an electronic circuit (it will have access to memory for temporarily storing the data, and may have access to reconfigurable processing, such as different algorithms for color transformation, and different manners of deciding the geometric positioning of the various transformed pixels on a canvas, such as a windowing system, the details of the latter being of lesser relevance for understanding the color coordination technology of the present technical system). The brightness mapping unit (710) is arranged to receive the pixel color codes of a visual asset (from the geometrical management unit 711) to be brightness mapped to a different brightness range, and apply e.g. a function-based transformation to at least the pixel’s input luma to obtain the output luma as needed for a particular brightness range which the operating system decides to use in its final presentation of the aggregate one or more visual asset (e.g. on a background). So in contrast to basic geometrical management units, which merely determine the spatial positioning of objects in a total geometrical composition, unit 710 and unit 711 cooperating together will also define the final colors of the various visual assets in the output composition, typically to be ultimately displayed on one or more connected displays. This suitable (final) brightness range may be determined based on various parameters, e.g. for which kind of display the composited canvas is generated, user preferences, the kind of total presentation (e.g. a user-interface-focused presentation), etc., but those details are not the core aspects of the present innovation. The original ranges of the assets may be determined by the various applications based on very different criteria. E.g., a game which expects lots of bright explosions to be displayed on the end-user display, may make its graphics (e.g. a notation of successful shots) equally super-bright.
[0103] A set of lumas respectively colors may for many situations simply mean one or more arrays. There will be an array of e.g. 200x100 luma code values (e.g. for a graphics to put on the top-left of the screen), and typically two more arrays for the chromas. In other situations the set of lumas (respectively colors) will be specified by drawing instructions (e.g. draw a line from beginning position (xl,yl) to end position (x2,y2) with luma 422 out of 1023). The skilled person understands how e.g. a PQ luma can directly code a unique corresponding luminance, but other color definitions may (although none of its three components is actually / directly a code for the luminance) also (together) code a luminance for a pixel, e.g. a non-linear R’G’B’ coding defined according to the perceptual quantizer. An important technical property is that the original asset colors will be communicated according to a (possibly a few selectable alternatives, but ideally a single) common format for a communication version of a pixelated image of an asset (example elucidated with Iml), or a procedural description of generation of an asset in a color space (example of Tx3), which the operating system can prescribe for all the communicating applications (i.e. the operating system may e.g. configure: “now communicate to me all your visual assets according to an API based on e.g. sRGB pixel color arrays plus corresponding re-grading metadata”). In some scenarios at least some of the applications may negotiate (pre-agree / configure) to use some alternative common color definition for the geometric pixel color component arrays for their communication of their visual assets for the ultimate color specification and geometrical composition by the operating system (OS).
[0104] A color code representing an input luma code means that, whichever the color specification / model used in any detailed embodiment, the coding of the color is specifying colors in the respective gamut (e.g. an SDR gamut, or some HDR gamut). Such gamut is constructed around a luma axis in the center, so a {Cl, C2, C3} color coding also corresponds to some projection (height) on the luma axis. Ergo, any processing working on the luma value of the pixel color (or a corresponding luminance) will know what input luma to use for its processing to obtain a corresponding output luma, or luminance. Since a luma is coding for a luminance via a pre-agreed OETF, the skilled person understands how luminance mappings can be performed either directly on pixel luminance values, or their corresponding luma values. So it may be that the color model (i.e. the color component triplet of the pixel) directly comprises the input luma code, as in the model which is popular for video: Y’CbCr, in which Y’ is the luma code of a pixel. However, all these colors (i.e. the e.g. SDR gamut) can also be represented in another additive color model. E.g., in the computer world the color code R’,G’,B’ is a popular coding. The pixel luma (and therefore in case of an absolute system its luminance) is then still uniquely represented (coded) since it follows from a universal fixed colorimetric equation: Y’=a*R’+b*G’+c*B’, in which a,b, and c are fixed constants, depending on the primaries of the gamut.
[0105] A suitable embodiment of a common format will be based on an SDR color gamut, e.g. a normalized Rec. 709 primaries gamut in which the brightest white luma code may represent 100 nit (e.g. the 8 bit luma code 255 dictates that pixels having these values, are to be displayed at 100 nit on any display, if not transformed into a tertiary brightness). Note that it is not obligatory that the (advantageously e.g. SDR) color representation for communication (i.e. of the Iml which codes the spatial and color structure of the first asset Assl, or a corresponding set defined as vector graphics) to the OS has an actual maximum luminance as a nit number. It is enough that the function Tml is able to allocate (during the reconstruction) an original maximum luminance (ML_V) to the largest possibly occurring code of the pixels in Im 1 (or equivalently a percentage of that largest HDR white if the asset only goes as bright as e.g. 40% grey). In other words it is sufficient if one can establish the dynamic range (preferably in absolute luminances, or alternatively relative brightnesses) in which the asset resides, with its HDR white if the asset actually has white pixels, or compared to its HDR white if it is darker (which can be achieved by giving it lower sRGB values in an Iml which forms a proxy for a e.g. 1000 nit original / (master) video: even if there is no pixel in the asset which actually reaches 100% sRGB, i.e. to be shown as 1000 nit; it suffices that TM1 instructs, via its shape, that pixels of 100% sRGB would be reconstructed to 1000 nit, rather than e.g. in another TM1 shape to 2000 nit, or in a third one as a maximum being 750 nit). In such a case the maximum brightness of the communicated image may be taken as fixed to 100% (the normal sRGB interpretation). Note that this 100%, after down-grading, will correspond to some potentially quite bright HDR maximum luminance, and the actual (scene) “SDR white” (i.e. a dimmer white that one would see in the scene as e.g. an averagely lit piece of paper (a Lambertian reflecting white), casu quo it would display on a display as atypical SDR-brightness range white, e.g. 200 nit on a display than can go as high as 600 nit) may be coded in the sRGB communicated image e.g. at luma 60%, or 30%. And that Lambertian white level (under average scene illumination) can be communicated as the reference white level WL1 (not to be confused with the peak or maximum white (ML_V) of a color representation, which may be as high as e.g. 6000 nit). Even though both the original and communicated representation may be using the same amount of bits per color component, e.g. 10 or 12 (or the original may be 3x12 bit and the communicated image Iml 3x8 bit), the difference in luminance range will be apparent from the squeezing together of at least some of the object colors, e.g. typically the brighter colors (this can be verified for relative communication images by allocating an actual maximum luminance to the maximum brightness, i.e. the relative brightness of 100% luma code, and looking at the histogram and inter- and intra-object contrasts, e.g. the ratios of the pixel brightnesses of two brighter pixels will be smaller in Iml than in the original image). This can best be seen -at least for a representation which codes actual pixel luminances- by looking at the values, e.g. a spread, of the luminances a set of pixels, e.g. of one or more image objects.
[0106] Several technical communication mechanisms may fulfil this property, but one can e.g. formulate TM1 as a function which maps the largest possible communicated pixel luma, i.e. input luma code = 255* 100% (or 255* 1) to the (normalized) level of 1.0 of the reconstructed HDR range, and then associate a maximum luminance ML_V with that maximum HDR output luminance or luma code, which ML_V will in such an embodiment get communicated as part of the definition of the first brightness mapping function TM1 (be it directly as the highest value of the function, or as a separate related metadatum for a normalized version of the luminance mapping function). This function could in some embodiments directly map input luma codes to normalized output luminances, and one need then only multiply 1.0 (or any value below) by ML_V (which is e.g. 4000 nit) to obtain the absolute output luminance of the pixel of the reconstructed output color. In case the brightness mapping unit embodiment directly produces relative brightnesses or especially when it produces absolute luminances, the output luma code (Y’o) may be coded as a native, linear coding of said brightness respectively luminance. In case the communication to the OS communicates the (typically down-grading) original function TM1 determined by the software application, the OS will invert that received function (possibly scaled for display adaptation) before doing its brightness range conversion. In some embodiments a flag may indicate whether the direct or already pre-inverted form of the brightness mapping function is put in the metadata (or the system may work in a pre-agreed manner). It may alternatively be more pragmatic to calculate in luma code domain also for the output, e.g. PQ lumas. Since the PQ definition already has a universally recognized maximum of 10,000 nit, if one then communicates a relative function TM1 which maps from e.g. sRGB normalized lumas to PQ lumas, then if 100% SDR luma maps to e.g. 75% PQ luma, we know the brightest pixels of the communicated image is ideally to be displayed as 1000 nit, because 0.75 in PQ means 1000 nit (unless the display or receiving side apparatus still needs to display optimize for a display of lower maximum display luminance, in which case it will use the color transformation control data CTctrl, and specifically TM1 to guide how this down-grading should ideally happen). It should be noted that a space for calculation and a space for color pixel array representation are not technically a similar component. One should also differentiate between communications which need the metadata, from mere changes of one image into another, or mere change of color space representation of an image only.
[0107] The original colors and their lumas and the luminances (as an amount of nit a.k.a. Cd / m2) or relative brightnesses they code (relative to e.g. the 100% level of SDR white), i.e. of the visual asset as the creating / communicating application ideally wants to see it displayed, may be considerably different than the coded colors as communicated: i.e. the normally decoded colors of the communicated coded proxy colors (using the normal sRGB definition instead of the brightness mapping function TM1 for the decoding) will usually be in a much smaller secondary brightness range, which ends at e.g. 100 nit instead of the original e.g. 20,000 nit. Note that the secondary range as communicated may also be larger than the original range of the visual asset, in particular it may end at a larger maximum luminance than the original maximum luminance of the asset pixels, but usually a smaller brightness range ending at a lowered maximum luminance will perform sufficiently well for communicating assets to the operating system.
[0108] Because the operating system also receives the color transformation control data (CTctrl), it can perform suitable tertiary transformations in line with what the original asset’s colors were, i.e. are supposed to be displayed as. E.g., the operating system may chose the tertiary brightness range to be identical to the primary brightness range, and invert the brightness mapping function (TM1), and use inverted first brightness mapping function (ITM1) on the input luma codes to obtain output luma codes which are a reconstruction of the original lumas of the asset of the application.
[0109] But in general things won’t necessarily be so simple. If different applications (e.g. a game and a website, or an encyclopaedia) use assets of considerably different brightness range, and in particular maximum luminance or maximum relative brightness, the operating system may want to coordinate the output luma of the geometric composition of various assets together in a total canvas, into a tertiary brightness range which may end at lower luminance than a primary brightness range of at least one of the assets of at least one of the applications. The operating system may also take into account what display will be served by the GPU with the composite images of its scanout buffer (ScOBff), and if that display does not have a high displayable brightness range, or the viewer wants to see everything bright near the upper end of the displayable brightness range, the OS may take this into account when firstly electing a tertiary brightness range, and secondly deriving optimal brightness mapping functions (FL opl) for mapping the various luma code values of the various assets to that common range, i.e. the final luminance range. In general the mapping any luma code gets, which can equivalently be described as a multiplication by a multiplier which depends on the value of that luma code (a multiplier larger than 1 indicating a -relative if performed in a normalized to 1.0 representation of the lumas or absolute brightness boost, and a multiplier smaller than one performing a brightness dimming), will depend on where in the input range the luma code falls. E.g., as one may want to squeeze most or all of the possible input lumas into the tertiary / output range, the amount of boost a luma say halfway gets may depend on how much the darker lumas get brightened in their mapping. So in general each pixel luma will get an optimal mapping so that the range of input brightnesses (e.g. luminances in some embodiments) is well- represented in the output range, but the many shapes of brightness mapping function that any application or the operating system may chose is also a detail we need not dive into, since the new technical system construction must be able to function which substantially each desired function.
[0110] It will be shown that some embodiments of the operating system’s common mapping may depend only on the communicated at least one brightness mapping function, whilst others may work solely on the communicated at least one reference white level (WL1), whereas other more sophisticated tertiary mappings may design their tertiary mapping function (or algorithm, e.g. taking into account the spatial nature of an asset, such as geometrically non-uniform shading) on both of those communicated color transformation control data elements.
[0111] Whereas the brightness mapping function essentially communicates how many luminances respectively lumas of the original asset’s colors are squeezed into the typically smaller brightness range of the pixel color array (Iml) that gets communicated, one may know where the brightest street lamp falls in that lower brightness range (namely near 1.0, or 255 in an 8 bit coding of the brightnesses respectively luminances), but one doesn’t know yet where the reference level of the uniformly lit (under the average base lighting of the scene, usually the bigger area of the geometrical frame of the images) Lambertian diffusive white object falls. That reference white level (WL1) of the asset might (depending on how much brighter the brightest HDR objects in the original asset representation are) e.g. fall at 128 for a not so impressive brightness dynamic range (also depending on which EOTF is used for the lumas, e.g. PQ being pre-agreed between operating system and applications, or HLG), but it may also fall on luma value 55. So that 100% Lambertian reference white level (or e.g. 90% of that level if one desires), may also be communicated, and be used to the benefit by the OS when determining its optimal tertiary mapping fimction(s).
[0112] So the first mapping (with TM1) will map the original pixel colors and their lumas, originally lying in the first luminance dynamic range which is determined by whatever the asset was created to be (e.g. a games designer may create a blue laser beam which is as bright as 8000 nit). The secondary brightnesses (i.e. the LuRl) may be coded as lumas which either code absolute nit secondary color lumas, but which may pragmatically end at a lower maximum luminance, e.g. 100 nit for SDR (reversible) common asset communication, or relative / percentual brightnesses. This first mapping establishes the relationship between the original asset colors, and the common interface colors (in the pixel color component arrays) which actually get communicated to the OS. So conversely, for the OS only getting the interface colors, it establishes what the original colors were, and were supposed to be in a displaying. So the maximum brightness of the secondary colors will typically be lower than that of the original colors of the various assets (transforming HDR assets into SDR assets, but in a smart invertible manner by using typically invertible or largely invertible first mapping function(s) TM1). But as regards the third (tertiary) mapping the OS can basically do what it wants (except for it should normally try to fulfil the desired look of the original colors, by at least taking into account the original first mapping function and or reference white level, and typically using functions which keep the color look of most of the colors, such as their differences, still reasonably similar as far as the OS-side, e.g. output-side desiderata enable (which can be elegantly realized e.g. by giving the tertiary mapping function a shape which is similar to the shape of the first brightness mapping, e.g. a weakened down version of that function). E.g. it may lift primarily the darkest colors, and shift the whole secondary brightness range (or even when considering from the original brightness range) upwards to basically much brighter colors for display.
[0113] A geometrical management unit (711) is a unit which manages the geometrical aspects of the assets, and their basic ingestion. So e.g. it will determine the position and possibly scale of assets in the total canvas (but not solely as a set of parameters, such as a top-left (x,y) coordinate pair, but generically also with its content filling, e.g. an image (i.e. the which asset to go where)), potentially in windows and the like. In this innovation, it will also take care of the control of luma processing on demand (to the brightness mapping unit functionality), so they become already of the correct color, and only the geometric aspects of pixel placement are in order. E.g. this unit may determine a window size and position for a window in the total canvas showing a video, and it may then scale (by geometric interpolation) the various correctly mapped pixel colors received from the brightness mapping unit, so that the video fits the window. Brightness mapping unit means any unit that can do brightness processing for one of more pixels of an asset, so that ultimately these pixels will have the desired brightness, either relative to some intermediate or maximum value, or absolute in nits. We show just a simple version for understanding where a single brightness mapping function will be applied merely on the value of the input luma, irrespective where it resides in the asset, but more complex scenarios could involve a central mapping function for mapping the pixels in say a circle in the middle of the asset, and a surrounding mapping function for the surrounding pixels. In that case the application will, for this asset, communicate two brightness mapping function, and instructions where to apply which function (e.g. with a bitmap, where 1 means first function and 0 means the second function should be applied by the operating system, or more precisely, its tertiary function should be based on that second function). Any geometric information needed may be communicated by the application in a geometric data section (Geol).
[0114] Finally, once the assets have all been suitably mixed, the problem is less difficult. It may determine an output pixelated image format (ImFinFmt) for writing the geometric composition in the total canvas of the one or more assets into one or more portions of the scanout buffer, e.g. first memory portion 720 and second memory portion 721. We will assume -without wanting to be limiting- a common format which is also useful for further communication by the GPU to a display, such as a perceptual quantizer luma based format (such a 10 bit format may represent pixel luminances up to 10,000 nit, and even if the original asset had some higher pixel luminances, this may be satisfactory; for relative brightness communications one can still use this coding, by assuming, or explicitly communicating to a display by a further metadatum, that some value is the 100% value, e.g. 100 or 200 nit).
[0115] A software application is a computer program designed to carry out a specific task other than one relating to the operation of the computer itself, i.e. other than the basic control of the computer hardware. It may implement various functionalities for the user, e.g. online shopping, presentation of visual media, etc. In the present discussion we need not formulate differences between applications for e.g. classical personal computers, or computer-orchestrated professional systems, and apps for portable apparatuses (the shorthand app can mean all of those).
[0116] Basically, at least one of the assets will be luminance mapped to the final range (as the OS may elect the luminance range of one of the assets to be a good final range), but oftentimes in addition to mapping the image array of the e.g. luma codes of the first representative coding (or the original first asset’s colors) to the final range with the function FL_opl, the second representative lumas of the second asset may oftentimes also be mapped to the final range, by their own optimized mapping function (FL_op2).
[0117] It may be advantageous in practice to have most or all software applications (being directed to) communicate to the operating system with color codes of the various pixels of their one or more visual assets represented in an sRGB color representation. This is a well-understood reference color system, but for SDR, but now with the additional color transformation control data CTctrl it can be used to also communicate a myriad of different kinds of HDR assets.
[0118] Advantageously the operating system will work in a manner enabling software applications to specify and communicate the original brightness of any pixel of any asset specifying a luminance of a pixel to be displayed as an amount of nits. This means that such an embodiment will communicate, even if the image array of Iml itself codes only relative (0-100%) lumas or normalized brightnesses, or only relative non-linear R’G’B’ values are communicated, the totality with the color transformation control data allows the receiving operating system to establish for each pixel (i.e. its reconstruction, or some derived re-graded image of pixels) an absolute luminance. This can be performed e.g. by specifying the brightness mapping function TM1 in a format which established or allows to establish absolute nit outputs, such as Perceptual Quantizer EOTF luma output domain values for the function (or equivalently in other embodiments one could add additional data to the color transformation control data CTctrl enabling e.g. an absolute scaling of the normalized to 1.0 brightnesses, such as a common multiplier, typically a maximum luminance value for the image or video ML_V).
[0119] Corresponding to the color-coordinating operation of the operating system, there will be software applications (e.g. web application 502, or native application 503) arranged to create at least one visual asset (Assl) comprising pixels wherein a pixel has an original color code which specifies an original brightness and to communicate such visual asset to an operating system (504), wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_V); wherein the original color code is transformed into a secondary color code for communication to the operating system, wherein the secondary color code represents a second luma code, which codes a secondary brightness which lies in a secondary brightness range which is different from the first brightness range, wherein the secondary brightness is derived from the original brightness based on application of a brightness mapping function (TM1) to the original brightness; characterized in that the software application communicates color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Assl), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1).
[0120] The application may elect what original HDR format it will use for its asset’s colors, but PQ luma-based colors will be a useful manner. E.g., the original asset pixel colors may be 0-2000 nit absolute luminance colors, represented as equivalent PQ Y’CbCr or PQ R’G’B’ values. Actually, it doesn’t matter if the color representation for unified communication is e.g. sRGB, since then the receiving operating system need only understand the original colors per se from the proxy sRGB color, and it may internally generate its output colors (of the output intermediate pixelated image format ImFinFmt) in e.g. Philips EOTF-based format, a logarithmic representation of the luminances or color components, etc. Especially absolute systems will be understandable because the luminance is an absolute optical quantity (and the other two color components, e.g. Cb and Cr, establish what are also universal color properties, namely a hue like e.g. Chartreuse, and a saturation). But also the relative systems can be sufficiently unique, since then one has a percentage of some maximum, e.g. a display maximum, or the maximum of a composited total canvas presentation decided by the operating system.
[0121] Just like the original color codes (whatever their actual codification) will represent original brightnesses, the secondary color codes for actual communication of the asset to the operating system will represent a secondary brightness, along a different range of brightnesses, due to the brightness mapping (note that often advantageously the actual mapping processing may be applied to a brightness component per se, but one can also map on other representations, e.g. the RGB components, equivalently, so that the brightness mapping is achieved, in the whichever representation). The important point is the common interfacing, allowing the coordinated use (typically re-grading) of any application’s asset by the operating system. In case the brightnesses are represented as lumas according to some elected EOTF, e.g. PQ, the derivation of secondary brightness from the original brightness based on application of a brightness mapping function (TM1) to the original brightness may involve first converting e.g. absolute nit values to input (original) luma codes (e.g. PQ lumas) and then applying the function in the PQ domain to obtain e.g. output PQ lumas (in case output luminances are required, the PQ EOTF can be applied to those output lumas).
[0122] The color transformation control data (CTctrl) is associated with the asset (Assl), which can happen in many manner, but typically the API will e.g. first communicate the color code array(s), or the data of the procedure to generate a set of pixel colors at a geometrical management unit of the OS, and thereafter the color transformation control data (CTctrl), or vice versa, it first sends the control data and then the pixel color data. The first brightness mapping function (TM1) or its inverse, and a reference white level (WL1) are for use by the operating system to understand which exactly of the many possible HDR colors of the asset the coding of the asset as received (e.g. image array(s) Iml) originally represented, and therefore also how it should ideally present (re-grade) such asset colors in any of the many possible final composited canvases it may want to generate, and send to the GPU for ultimate display. As a first application may send an asset of a very different maximum brightness (e.g. ML_V1= 8000 nit) than a second application (e.g. ML_V2= 1000 nit), the operating system can understand this, and then also in any embodiment of its brightness mapping unit 710 decide how to coordinate those two assets in the final canvas (e.g. it may dim the first asset, or boost the second one somewhat, etc.).
[0123] The applications (or even the operating system) may be supplied to the computer, via some software communication mechanism.
[0124] The techniques may be embodied as a method of communicating a visual asset having pixel colors to an operating system, comprising the steps of: creating at least one visual asset (Assl) comprising pixels wherein a pixel has an original color code which specifies an original brightness, wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_V); transforming the original color code into a secondary color code for communication to the operating system, wherein the secondary color code represents a second luma code, which codes a secondary brightness which lies in a secondary brightness range which is different from the first brightness range, wherein the secondary brightness is derived from the original brightness based on application of a brightness mapping function (TM1) to the original brightness; communicating the secondary color code and color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Assl), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1).
[0125] The techniques may be embodied as a method of operating a computer, comprising a step of running an operating system, wherein the operating system is arranged to manage color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Assl) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness (ML_V), wherein the visual asset (Assl) is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness; wherein the visual asset in addition comprises color transformation control data (CTctrl) which comprises at least one of first brightness mapping function (TM1) or its inverse, and a reference white level (WL1); wherein the operating system comprises a brightness mapping unit (710) arranged to receive the visual asset and to apply a tertiary brightness mapping function (FL opl) to the input luma code to obtain an output luma code (Y’o), wherein the output luma code specifies a tertiary brightness which lies in a tertiary brightness range which is different from the secondary brightness range, wherein the output luma codes a brightness of an output color, wherein the tertiary brightness mapping function (FL_opl) is based on at least one of the first brightness mapping function (TM1) and the reference white level (WL1); wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
[0126] Interesting embodiments under the present innovative technical color coordination approach are inter alia:
[0127] An operating system (504) for being supplied to and for operation on a computer, wherein the operating system is arranged to manage color coordination of visual assets received from a first software applications (701) and a second software application (702), wherein the operating system comprises a geometrical management unit (711) arranged to manage spatial positioning in a canvas of visual assets, and arranged to receive a first visual asset (Assl) from the first software application which has a pixel which has a pixel color which has a first original brightness (Lui) lying in a first brightness range which has a first maximum brightness (ML_Vorl), wherein the first visual asset (Assl) is received comprising a set of pixel color codes which are defined in a first common color space pre-agreed between the operating system and the first software application, wherein a first pixel color code represents a first representative brightness (LuRl), which is derived from the first original brightness by applying a first brightness mapping function (TM1) to the first original brightnessand wherein the first representative brightness is defined on a range which ends at a third maximum brightness (ML_passl) which is different from the first maximum brightness (ML_Vorl), and arranged to receive a second visual asset (Ass2) from the second software application which has a pixel which has a pixel color which has a second original brightness (Lu2) which lies in a second brightness range which has a second maximum brightness (ML_Vor2), wherein the second visual asset (Ass2) is received comprising a set of pixel color codes which are defined in a second common color space pre-agreed between the operating system and the second software application, wherein a second pixel color code represents a second representative brightness (LuRl), which is derived from the second original brightness by applying a second brightness mapping function (TM2) to the second original brightness, wherein the second representative brightness is defined on a range which ends at a fourth maximum brightness (ML_pass2) which is different from the second maximum brightness (ML_Vor2), wherein the first maximum brightness is different from the second maximum brightness, wherein the first visual asset comprises first color transformation control data (CTctrl) which comprises at least one of the first brightness mapping function (TM1) or its inverse, and a first reference white level (WL1), which is lower than the third maximum brightness (ML_passl), and wherein the second visual asset comprises second color transformation control data (CTctrl2) which comprises at least one of the second brightness mapping function (TM2) or its inverse, and a second reference white level (WL2), which is lower than the fourth maximum brightness (ML_pass2); wherein the operating system comprises a brightness mapping unit (710) arranged to receive the first visual asset and the second visual asset and to apply a third brightness mapping function (FL opl) to the first representative brightness to obtain an output luminance (Lol), or a corresponding output luma code (Y’o), wherein the output luminance lies in a final brightness range which is different from the third brightness range (DR L1), wherein the output luma codes a brightness of an output color, wherein the third brightness mapping function (FL opl) is based on at least one of the first brightness mapping function (TM1) and the first reference white level (WL1); wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
[0128] The common color space / representation may advantageously be a standard dynamic range color space, e.g. an RGB representation, preferably quantized in more bits than 8, and the original brightnesses (and the output brightnesses) may be specified as luminances.
[0129] A method of operating a computer, comprising a step of running an operating system, wherein the operating system is arranged to manage color coordination of a first visual asset (Assl) received from a first software application (701) and a second visual asset (Ass2) received from a second software application(702), wherein the operating system comprises a geometrical management unit (711) arranged to manage spatial positioning in a canvas of visual assets, and arranged to receive the first visual asset (Assl) comprising a pixel color having a first original brightness (Lui) lying in a first brightness range which has a first maximum brightness (ML_Vorl), and arranged to receive the second visual asset (Ass2) comprising a second pixel color having a second original brightness (Lu2) lying in a second brightness range which has a second maximum brightness (ML_Vor2), wherein the first maximum brightness is different from the second maximum brightness, wherein the first visual asset (Assl) is received comprising a set of pixel color codes, which are defined in a first common color space preagreed between the first software application and the operating system, wherein a first color code represents a first representative brightness (LuRl), which is derived from the first original brightness by applying a first brightness mapping function (TM1) to the first original brightness, wherein the first representative brightness is comprised in a third brightness range which has a third maximum brightness (ML_passl) which is different from the first maximum brightness (ML_Vorl), and wherein the second visual asset (Ass2) is received comprising a set of pixel color codes, which are defined in a second common color space pre-agreed between the second software application and the operating system, wherein a second color code represents a second representative brightness (LuR2), which is derived from the second original brightness by applying a second brightness mapping function (TM2) to the second original brightness, wherein the second representative brightness is comprised in a fourth brightness range which has a second maximum brightness (ML_pass2) which is different from the second maximum brightness (ML_Vor2); wherein the first visual asset comprises first color transformation control data (CTctrl) which comprises at least one of the first brightness mapping function (TM1) or its inverse, and a first reference white level (WL1), which is lower than the third maximum brightness (DR L1), and wherein the second visual asset comprises second color transformation control data (CTctrll) which comprises at least one of the second brightness mapping function (TM1) or its inverse, and a second reference white level (WL2), which is lower than the fourth maximum brightness (DR L2); wherein the operating system comprises a brightness mapping unit (710) arranged to receive the first and second visual asset, and to apply a third brightness mapping function (FL opl) to the first representative brightness (LuRl) to obtain an output luminance (Lol), or a corresponding output luma code (Y’o), wherein the output luminance lies in a final brightness range which is different from the third brightness range (DR L1), wherein the output luma codes a brightness of an output color, wherein the third brightness mapping function (FL opl) is based on at least one of the first brightness mapping function (TM1) and the first reference white level (WL1); wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
[0130] A software application (502, or 503) arranged to create at least one visual asset (Assl) comprising pixels wherein a pixel has an original color code which specifies an original brightness (Lui) and to communicate such visual asset to an operating system (504), wherein the original brightness lies in a first brightness range (DR Hl) which has a first maximum brightness (ML Vorl); wherein the original color code is transformed into a secondary color code for communication to the operating system, wherein the secondary color code is represented in a common color space pre-agreed with the operating system, which codes a secondary brightness (LuRl) which lies in a secondary brightness range (DR L1) which is different from the first brightness range and which ends at a third maximum brightness (ML_passl), wherein the secondary brightness is derived from the original brightness based on application of a first brightness mapping function (TM1) to the original brightness; characterized in that the software application communicates color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Assl), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1), which has a lower value than a third maximum brightness (ML_passl). The software application wherein the common color space is demanded by the operating system for all communications of visual assets by the software application to the operating system.
[0131] The software application comprising an application-side color representation determining system (1310) which is arranged to coordinate with or receive from the operating system a specification of the common color space.
[0132] A method of communicating a visual asset having pixel colors to an operating system, comprising the steps of: creating at least one visual asset (Assl) comprising pixels wherein a pixel has an original color code which specifies an original brightness (Lui), wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_Vorl); transforming the original color code into a secondary color code for communication to the operating system, wherein the second color code is defined in a common color space pre-agreed between the method and the operating system, wherein the secondary color code represents a secondary brightness (LuRl) which lies in a secondary brightness range which is different from the first brightness range and ends at a third maximum brightness (ML_passl), wherein the secondary brightness is derived from the original brightness based on application of a brightness mapping function (TM1) to the original brightness; communicating the secondary color code and color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Assl), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1), which has a lower value than the third maximum brightness (ML_passl).
[0133] BRIEF DESCRIPTION OF THE DRAWINGS
[0134] These and other aspects of the method and apparatus according to the invention will be apparent from and elucidated with reference to the implementations and embodiments described hereinafter, and with reference to the accompanying drawings, which serve merely as non-limiting specific illustrations exemplifying the more general concepts, and in which dashes are used to indicate that a component is optional, non-dashed components not necessarily being essential. Dashes can also be used for indicating that elements, which are explained to be essential, but hidden in the interior of an object, or for intangible things such as e.g. selections of objects / regions.
[0135] In the drawings:
[0136] Fig. 1 schematically explains various luminance dynamic range re-gradings, i.e. mappings between input luminances (typically specified to desire by a creator, human or machine, of the video or image content) and corresponding output luminances (of various image object pixels), of a number of steps or desirable images that can occur in a HDR image handling chain;
[0137] Fig. 2 schematically introduces (non-limiting) some typical examples of use scenarios of HDR video or image communication from origin (e.g. production) to usage (typically a home consumer); Fig. 3 schematically illustrates how the brightness conversion and the corresponding conversion of the pixel chromas may work and how the pixel brightness histogram distributions of a lower and higher brightness, here specifically luminance, dynamic range corresponding image may look;
[0138] Fig. 4 shows schematically how computer representations of HDR images would represent a HDR range of image object pixel luminances;
[0139] Fig. 5 shows on a high level the concepts of the present application, how various applications may be dealing with the display of pixel colors via and / or in competition with all kinds of other software processes, which may not result in good or stable display behavior, certainly when several different displays are connected or connectable;
[0140] Fig. 6 shows an archetypical example of a user scenario, the user dealing with and looking at several HDR-capable web-applications in different windows occupying different areas of a same screen, may benefit from the current new technical approach and elements;
[0141] Fig. 7 schematically shows an embodiment of a computer according to the present innovations running a few applications which coordinate their assets with the innovative operating system according to the innovative technical specification format for the asset colors of arbitrary higher brightness dynamic range;
[0142] Fig. 8 schematically illustrates one embodiment of an algorithm by which the OS can beneficially use the new color transform control data to come to a total canvas for showing together several of the visual assets in a we 11 -coordinated manner, respectful of both the original intended color and the appearance of the totality;
[0143] Fig. 9 schematically illustrates how the OS can derive secondary brightness mapping functions from any brightness mapping function it receives in the color transform control data;
[0144] Fig. 10 discusses (without intending to be limited) so more details on how the present coordination framework concepts can be mapped internally in an OS;
[0145] Fig. 11 shows an example of how one can luminance re-grade a visual asset (in this case an image, such as a photo of a bam) between two representations of the same scene having different luminance dynamic range;
[0146] Fig. 12 shows how an operating system of a computer (e.g. PC or mobile phone) may typically transform two different visual assets from two different software applications to come to its tertiary image representation having luminances along a final luminance range, which image is for writing in a screen buffer for displaying, as an example of the present innovation’s technical approach;
[0147] Fig. 13 shows an embodiment showing a negotiation between an application and an operating system about which form of API to use, in particular which common color representation for the pixel color arrays (e.g. sRGB).
[0148] DETAILED DESCRIPTION OF THE EMBODIMENTS
[0149] In Fig. 5 we show generically (and schematically to the level of detail needed) atypical application scenario in which the current innovation embodiments would work. As this is currently becoming ever more popular, we show -from the client side system (500), i.e. e.g. running on a personal computer or mobile phone, how a web-app may work via a native app and finally to the operating system (OS), which may coordinate the ultimate communication with the GPU and management of ultimate display. A mobile phone, unless when in a screen casting application, will typically have one display only (second (HDR) display 522), but other client side systems may communicate with several displays. We have shown some options in dotted, to indicate optionality, e.g. the web application (502) may be communicating directly with an operating system (504) without an intermediate native application (503), or there may only be a locally installed and running (specially written for some system) native app involved, at least for some of the visual assets being prepared for display, and no contacting with anything over the web regarding those visual assets.
[0150] The web application could be e.g. an internet banking website, remote gaming, a video on demand site of e.g. Netflix, etc. The web application will run as software on typically a central processing unit (CPU) 501. E.g., the composition of a web presentation may be received as HTME code. The web application will be to a large extent (for the user interaction) based on visual assets, such as e.g. text, images, or video.
[0151] In Fig. 6 we have generically (without wanting to be limiting) shown what the user would see on his one or more display screens, displaying a total viewable area or canvas (the screen buffer 601), containing two windows showing two such applications. The web application may have needs to show its assets in high dynamic range (i.e. better visual quality, in particular a range of brighter pixel luminances than SDR, and correct / controlled pixel luminances). E.g., it may show HDR still pictures showing more beautifully than with a plain SDR JPEG to let the customer browse through new available movies. Or it may want to present commercial material which is already in some HDR format (or vice versa, an old not yet HDR commercial asset, i.e. still in Rec. 709). The first web application (or in general software application) may in first area 630 of displayable pixels (i.e. the canvas) show at least one visual asset, e.g. the first visual asset may be the HDR video of the cowboy (612). That video may comprise colors defined on a first HDR range ending at a first maximum luminance ML V vl i.e. it may contain pixels as bright as (e.g. the sun) e.g. 3000 nit. The operating system may need to (during its geometric composition phase) coordinate the colors of the first video with that of a second visual asset, being e.g. the HDR 3D graphics 622 (or another natural, e.g. originally camera-captured image or video), which may have its pixel colors defined on i.e. lying within a second HDR range of brightnesses or preferably luminances, e.g. ending at a second maximum luminance ML_V_v2 having the pre-fixed value of e.g. 4000 nit (i.e. as defined by the creator of this secondary visual asset, be that a human or autonomously operating computer program).
[0152] E.g., a first web application may present its content in a first window 610. The content of this window, may be on the one hand a first plain text 611 (i.e. graphics colors for standard text, which is well-representable in SDR, e.g. the colors black for the text on an “ivory” background). On the other hand, this webpage, which gets composed in the first window for the viewer, may also present a HDR video 612 (e.g. a video streamed in real time from a file location on another server, coded in HLG, and assuming 1000 nit would be a good level for presenting its brightest video pixel colors). This window could be side by side on a total canvas of the windowing server, which we will here call “screen buffer” (to discriminate from local canvases that windows may have). Canvas is a nomenclature from the technical field, which we shall use in this text to denote some geometrical area, of a set of N horizontal by M vertical pixels (or similarly a general object like an oval), in which we can specify (“write” or “draw”) some final pixels, according to a pixel color representation (e.g. 3x8 bit R,G,B). Even if there was only one such window on the screen buffer, it could already partially project on a first display 520, and partially on a second display 522, e.g. when dragging that window across screens. If the screen buffer were to represent the pixels in one kind of color representation, e.g. merely basic SDR (which is HDR unsavvy), then one of the challenging questions would already be how the second display, which is a HDR display, which expects PQ-defined Y CbCr pixel color representations, and can display pixel colors as bright as its second display maximum luminance ML_D2 capability being 1000 nit (compared to first display maximum luminance ML D1 capability being 100 nit) would display its part of the window. The viewer may find it at least weird or annoying if a white text background suddenly jumps from 100 nit to 1000 nit. Note that there will also be Graphical User Interface (GUI) widgets, such as first window top bar 613, and first window scroll bar 614. The operating system (OS, 504) may deal with such issues (although it sometimes also expects applications to deal with at least part of the GUI widgets, e.g. remove buttons when the application is in full screen, and re-display them upon an action such as clicking or hovering at the bottom of the screen, or showing the widgets partially transparent, etc.; the applications may also indicate how they prefer one or more UI (typically graphically representable) elements to behave, in particular as regards their primary HDR color(s) and possibly also a preferred manner to regrade them to secondary colors for a lower or higher dynamic range, even if the OS may have the final say on how the colors would come to look in the ultimate representation in the screen buffer ready for display). Usually, at least for the moment, such widget graphics may be solely in SDR, but that may change in the near future.
[0153] In the second window 620 we show another web application. The two web applications may not know about each other (and then cannot coordinate colors or their brightnesses), but the local computer will (e.g. the windowing server application of the operating system). The second window, apart from its second window top bar 623 and second window scroll bar 624, may e.g. comprise a HDR 3D graphics 622 (say fireworks generated by some physical model, with very bright colors, illuminating a commercial), and in another area second text 621, which may now e.g. comprise very bright HDR text colors (e.g. defined on the Perceptual Quantizer scale, say non-linear R_PQ, G_PQ, B_PQ values). E.g., it may be specified that the text is white, with white pixels to be displayed at 500 nit, on a dark background, e.g. to be displayed at 0.1 nit. The various parts of the screen (windows, or parts of windows, background), may correspond to areas and their pixel sets, e.g. area tightly comprising all pixels related to the first window and nothing from the screen background (shown slightly larger to be visible).
[0154] Applications nowadays call functionality of other software (i.e. calculations or other actions that the other software processes can do) or hardware via Application Programming Interface Specifications (API). This is a very powerful way to communicate typically only the bare essentials, to secondary / receiving components that may be of various construction doing similar things in a somewhat different manner. This hides the very different deep details of e.g. another software, or a hardware component like a GPU (like its specific machine code instructions that must be done). This works because there is universality in the tasks. E.g., the end user, or any upper layer application, doesn’t care about how exactly a GPU gets the drawing of a line around a scroll bar done, just that it gets done. Also for HDR pixel colors, although there may be many flavors, at least the way in which e.g. a particular pastel yellow color of 750 nit gets defined and created (for ultimate display) at a certain pixel, just that this elementary action gets done. The same is true for a rendering pipeline, which usually always consists of e.g. defining triangles, putting these triangles at the correct pixel coordinates, interpolating the colors within the triangle, etc. (and it doesn’t matter whether one processor does this, even running on the CPU, or 20 parallel processing cores on the GPU, and where and how they cache certain data, etc.).
[0155] At least in the SDR era this didn’t matter, because SDR was just what it was, a simple universal manner to represent all colors between, and projected around, black and white (approximately 0% and 100% light output or reflection). And a certain blue was a certain blue (it had a certain hue and saturation, and a certain amount of that chromaticity being reflected i.e. a certain lightness, and when shown on a 100 nit maximum white system a certain luminance), and therefore GUI widget blue could be easily mixed (composed) with movie or still image blue, etc. (because all those colors were already defined in the same manner, i.e. almost as if ready for mixing).
[0156] So e.g. the web application may specify some intentions regarding colors to use (e.g. for unimportant text, important text, and text that is being hovered over by the cursor), e.g. by using Cascaded Style Sheets (CSS), but it may rely for the rendering of all the colors, in a canvas, on a native application 503. Since applications may desire a well-contemplated and consistent design, CSS is a manner to separate the content (i.e. e.g. what is said in a news article’s text) from the definition and communication of its presentation, such as its color palette. For web browsing this native application would be e.g. Edge, or Firefox, etc. So it may call some standard functionality via a first API (API1). The native application may do some preparing or processing on the colored text, but it may also rely on some functionality of the operating system by calling a second API (APE). But the operating system may, instead of doing some preparatory calculations on the CPU, and then simply write the resultant pixel colors in the VRAM of the GPU, also rely on the GPU to calculate the colors in the first place (e.g. calculate the fireworks by using some compute -shader). In any case, it may issue one or more function calls via third API (API3). The various APIs may also query e.g. what the preferred color format of a canvas is (colors may be compressed according to some format to save on bus communication, needed VRAM), etc. The final buffer for colors of the screen to be shown, in the GPU 510, is the so-called ScanOut Buffer 511. From there the pixels as needed to be communicated to the display, e.g. over HDMI, are to be scanned, and formatted into the needed format (timing, signaling, etc.).
[0157] A first problem is that, although one may want to rely on other (layer) software for e.g. positioning and drawing a window, or capturing a mouse event, dynamic range conversion is far too complex, and quality-critical, to rely on other software components to do some probably fixed and simplistic luminance or luma mapping. A first intermediate application may first down-grade the colors, and then a second software component may up-grade them again, and the final result may be both unpredictable and ugly. The creator of the web page, or the commercial content running on it, may have spent far too much time making beautiful HDR colors, only to see them mapped with a very coarse luminance mapping, potentially even clipping away relevant information. If it was a simple color operation, such as the correction of a slight bluish tint, which would also not involve too severe differential errors or too different a look if not done appropriately, or at all, one could rely on a universal mechanism for doing that task, which could then be done anywhere in the chain. But color-critical HDR imagery has too many detail aspects to just handle it ad hoc.
[0158] Indeed, until recently, as a second problem the various applications might have defined various HDR colors (internally in their application space), but the windowing servers (e.g. the MS windows system) typically considered everything to be in the standard SDR color space sRGB. Because that is what they understood and had been using for decades, and HDR was complex, multi-variant, and ill-understood nor agreed. So you had HDR colors, but you needed to work with SDR colors, or that was what was going to come out to be able to see anything. The initial research of HDR video coding and handling, managed to define in a (be it multi-flavored) relatively stable manner what each successive component should do, but that only regards a sole visual asset, in which all the pixel colors are defined in the same HDR range, and also in the same output range (e.g. to be directly displayed on connected output display range), even if parts of the image would map from the primary dynamic range of the original video to the secondary / output range using different locally optimized functions (the output video is still considered as one video, having the range of all its pixel luminances falling below e.g. display maximum luminance ML_D_dl= 600 nit). To be clear the reader is not confused, the concept of local luminance mapping functions for secondary (proxy) images defining the colors of primary images is elucidated by aid of Fig. 11 (also, the parts of the video are typically not considered as separate visual assets, but the video is treated as a single visual asset, even if it is dynamic range processed in a more complicated manner). Fig. 11 shows an example of a relatively difficult HDR image: a sunny outdoors seen from a dark bam, which contains a vending machine which has a bright display showing a beverage can. The text (obj txt) “CO” on the can and the house (obj hou) may both be white objects, visually. But they would not have the same luminance in the real scene, and also typically not in the 1000 nit master grading. If we want to represent (e.g. code) the 1000 nit HDR image as a 100 nit (proxy) image, one could map the various image objects by a single “squeezing” function. This would have the physical property (already for needing to guaranteeing reversibility for the reconstruction from the received proxy video) that a pixel which is brighter in the master image would also be brighter in the proxy image. Whereas this is a good simple approach for some situations, this does not necessarily yield visually the best proxy corresponding to the master. The white text on the vending machine may become inconveniently dark. Even if one relies on the visual fact that in the surrounding darker pixels it will psychovisually somewhat brighter than if it were presented to the viewer under standardized presentation conditions (e.g. the sole color on a neutral background), one may still want to optimize those specific colors’ luminance in a separately controllable manner. One (e.g. the color grader creating the visual asset being this video image, and its coding comprising the luminance mapping function for HDR image reconstruction from the SDR proxy in metadata) can do that by finding that most of the other colors (e.g. the dark indoors objects like obj drk, or the tree obj_tr may well be gradable together by a sole global luminance mapping function (here first local mapping function FL locl), i.e. almost everything inside with everything outside. But for the display of the vending machine (obj dis), one may define a second local luminance mapping function FL_loc2 for finetuning the glowing look of the display, in that dark interior, as desired. This is a technical choice of standards orthogonal to the present teachings. The best quality HDR systems would according to the opinion of one of the inventors allow local mapping, but for practical reasons some implementers like to stick with only one sole function per image (i.e. per display time instant). In any case, the creator would define this sole visual asset and it colors and its re-grading needs as one or more luminance mapping functions, and bundle the lot as a coded visual asset, for storage or reception and possible further color processing. In this text a visual asset is to be understood as some total set of pixels which belong together and represent something (together), fully made and coded.
[0159] That was good for compositing any kind of SDR visual asset together. But an application having generated e.g. very colorful 750 nit fireworks, would need to luma map those to dull SDR firework colors, and then the windowing server would compose everything together on the screen buffer. Even if somewhere in between there would be some re-mapping to another dynamic range, that would be very ad hoc.
[0160] So at best any application would need to convert its beautiful HDR visual content to sRGB visual assets, and also there would be no coordination with other application’s assets. Furthermore, worst case there could be a cascade of various luma mappings in the various layers of software acting on top of each other, e.g. some of those involving SDR-to-HDR stretching, and all in an uncoordinated manner, ergo, one would have no idea of what came out displayed at the end, to the detriment of visual quality and pleasure of the viewer (e.g. (parts of) windows could be too dark, of insufficient contrast, if some pixels colors lay outside expected boundaries of some software some parts of the visual asset may have disappeared when displaying, etc.). Recently some windowing systems have introduced that applications can communicate to the window composition process canvases that can contain 16 bit floating point (so-called “half float”) color components (and consequently luminances of those colors), however this only allows to basically communicate a large amount of colors (per se). but doesn’t guarantee anything how these will be treated down the line, e.g. which luma mapping will be involved when the operating system makes the final composition of all visual assets, to correctly make the screen buffer, and ultimately the scanout buffer of the GPU. Note that how the screen buffer (especially if several displays are involved in an extended view) is exactly managed in one or more memory parts, with one or more API calls etc., is not critical to the present innovation embodiments. E.g., one display may refresh at higher rate than the other, and then its part of the total canvas will be read out more frequently. The innovation is about allowing the optimal coordination of the various parts the user ultimately gets to see, via the careful communications and attuned handling of the coordinated reference image representations (e.g. sRGB) and associated re-grading functions for defining HDR assets, and the coordinated use of it all when e.g. display adapting the final composition or part thereof.
[0161] In Fig. 7, which shows a computer 700 (which could also be e.g. a mobile phone) connected to a (external or internal) display, a first software application 701 may be showing e.g. images of a motorcycle (first asset). Say this motorcycle is a 2000x1000 pixel image, which may get re-scaled to the total output canvas by the OS. Its asset may be defined primarily by an image Iml, which may be e.g. 3 pixel color component arrays (MPEG compressed, or non-compressed), e.g. an array for the lumas Y’, and one for the blue chromas Cb and the red chromas Cr. This asset image (i.e. pixel colors arrays) is supplemented by associated metadata, e.g. in this first example at least a first brightness mapping function TM1. The ellipses around the first reference white level WL1 in the schematic Figure indicate that this data may be present in the communication to the OS (e.g. API call), or not. There may be geometric information Geol, which can e.g. indicate preferred size, position, whether the application would like to be on top at least partly (or transparency information), etc. Especially if this application communicates multiple assets, the geometric information may convey the relative positioning of these assets. A second software application 702 may be showing e.g. an encyclopedic article about bats (either from an internal memory of the computer device, potentially composed on the fly and upon request of a user, or from internet, etc.). It may communicate to operating system 504, again according to the innovative stable mechanism of the present application, two assets. Second asset Ass2 may e.g. be the image (potentially a computer-generated graphic instead of a photo) of the bat. In this case this graphic is communicated also as a (second) image hn2 (again one or more color component arrays coding the respective magnitude of the color component for any pixel position in the array(s)). In addition the second color transform control data now has both a second brightness mapping function TM2 and a second reference white level WL2. The third asset (the text below the bat graphic) may be procedural in coding. E.g. third pixel set color specification procedure Tx3 may give a HDR color to ASCII-coded text characters, and another HDR color for the background, both of which will be converted to corresponding sRGB color codes (or the like in similar embodiments), yet those can be re-graded to the HDR colors (preferably both with the same inverse function; though in general they could have their separate transformation). In this example (non-limiting), we have shown that we could also merely give a third reference white level WL3 (no luminance mapping function, luma mapping function or brightness mapping function). Based on this level, the operating system can then judge how bright or dark the creating software desired the colors to be, and take this into account when positioning the colors in the tertiary brightness range (e.g. it may boost colors to the level of an explosion, yet still, compared to that new reference level, keep dark colors of the text or background sufficiently dark).
[0162] The brightness mapping unit 710 of the OS can, as shown symbolically, both apply upgrading (convex; i.e. typically starting with a moderate slope and a higher slope for the brighter subrange) or downgrading (concave; typically bending towards the horizontal axis for brighter inputs) function to the input sRGB lumas, as the need may be. It will produce the correct tertiary range output lumas Y’o, e.g. typically in a re-graded image hnScal. Finally the GPU can communicate the total composited image to a display 750, via some image communication path 751, e.g. a HDMI cable, Wi-Fi screencasting, etc. The geometrical management unit (711), e.g. a window compositing unit of the OS, is shown as the unit which does the overall asset management, i.e. receives or retrieves all necessary information, does the brightness mapping via call to brightness mapping unit 710, and ultimately sends the composited image to one or more scanout buffers or portions of buffers (first scanout buffer portion 720 and second scanout buffer portion 721) of the GPU 510. Whether the brightness mapping unit is integrated within or cooperating with the geometrical mapping unit is a matter of embodiment.
[0163] Fig. 8 elucidates an embodiment of how the OS can use the new color transformation control data, namely a first brightness mapping TM1 for the video image Iml (of the monster in the cave), and a first reference white level WL1 (here e.g. luma 64 out of 255 luma codes), and simply a third reference white level (WL3=100) for some graphics text, in deciding its final (original color-attentive) tertiary brightness range composition of the total assets canvas.
[0164] The output (vertical axis) is specified as normalized luminances. In this example the OS has chosen it needs a 2000 nit final brightness range maximum (ML_COMP) for the geometric composition (i.e. the finally decided colors, good for displaying), ergo the linear scale endpoint means 2000 nit. The horizontal input axis are in this elucidating example just sRGB input lumas Y’ in SDR, i.e. approximately the square root of the linear relative brightnesses, and again normalized so that 1023 in 10 bit is represented as 1.0. The first reference white level WL1, can be used by the OS to e.g. construct a first, lower brightness part TMd opt of the mapping curve (from input lumas of pixels in Iml, to output luminances). E.g. the OS can determine (non-limiting) a “straight” line allocation (we have symbolically drawn a straight line, but since the input is square root, the shape should be approximately power 2), which maps the identified first reference white level WL1 to a first common white level WcomLl of the composite canvas, and all input linear brightnesses of the Iml correspondingly linearly proportional below this juncture point Pj 1. This would be an equiluminance mapping, but scaled (e.g. brightened or darkened) linearly proportional mappings and non-linear mappings below the juncture point would equally be possible, the juncture point primarily functioning for positioning a level of brightness of something in the visual asset. This becomes elegantly configurable: the OS can establish e.g. 250 nit to be a good first common white level. The brighter input lumas Y’ in SDR will follow a shape-conforming optimal mapping function TMl opt, which in this example largely follows the shape guidance of the original communicated brightness mapping function TM1 of the first asset (see further elucidated in Fig. 9 how this can be achieved, when mapping to a different maximum luminance of the composition compared to the maximum luminance of an original asset). Essentially, this means that primarily the flames should be well-boosted (note that the diagonal should be interpreted with a brighter output maximum luminance: if we allocate e.g. 100 nit to the input normalized maximum, mapping according to the diagonal already corresponds to a 20-fold brightening of all pixel luminances, at least if the input axis was linearly represented). We however also see that the creator of the TM1 function, i.e. the first software application, and potentially any human behind it, has also created a steep slope where the monster is coded, ergo, this may mean that he wanted the monster to be rather contrasty in any re-grading, ergo this guiding shape should be followed at least as far as achievable (as we will also see in Fig. 9, the shape of a function, i.e. basically how the output varies for increasing inputs, can elegantly be described by a set of distances of successive points on the diagonal to the locus of points of the function, e.g. Pj 1). So the shape of TMl opt, will at least be based on the communicated TM1 for asset Assl (Iml), and often also based on WLl.
[0165] This constitutes already a good final coloring (re-grading) of the first asset, on the tertiary range image (i.e. ImFinFmt, or its internal intermediate precursor of different colorimetric definition), for the GPU, and ultimate display. Note that the first reference white level WL1 can be communicated even though there is no object in the current image(s) that actually has this value, as it can still be used as important reference value by the OS, or the instructing software application.
[0166] Now the second asset is placed in the total canvas, after establishing suitable luma regrading. We show an example what can be done if the text is encoded in a third sRGB image Im3 (we now assume without limitation that the text is not ASCII, but already communicated as a pixelated sRGB image comprising a rendering of that text), together with only a third reference white level WL3.
[0167] The OS sees that the text (e.g. yellow text with luma 175) is actually supposed to be brighter than the communicated third reference white level WL3. So the second software application (or the first SA for this third asset), wanted to show above averagely bright text, which can be characterized by a first contrast Conti . The OS can decide to respect this, and respect this compared to the tertiary brightness range of the total composition. The OS may also want to e.g. lower or increase the contrast somewhat.
[0168] In order to do this processing, the OS may e.g. establish a characteristic brightness level CHRbriLev COMP of the brightest regions of the composited canvas, which will contain the flames. E.g. areas or objects can be extracted, and an average output luminance can be determined. This may be the brightness the ultra-bright text of the third asset (i.e. Im3) has to compete with. Ergo, the OS can decide to derive one or more of a second contrast Cont2 of a to determine second common white level WcomL2 to the characteristic brightness level CHRbriLev COMP, and a third contrast Cont3 of that second common white level WcomL2 to the first common white level WcomLl. E.g. WcomL2 can be lowered below CHRbriLev COMP the smaller Conti is (and e.g. in a linear or non-linear proportion of CHRbriLev COMP compared to WcomLl), etc. Alternatively, other embodiments can ignore the maximum areas of the video, and merely raise the amount of Cont3 as a linear or non-linear function of how much Conti is above WL3.
[0169] Fig. 9 shows generically an example of an algorithm that the OS can use to map brightnesses to a different brightness range (e.g. 100 nit luminances that were to be reconstructed to 2000 nit luminances, will actually be re-graded to a final range ending at maximum 1000 nit) using an essentially shape preserving tertiary mapping function TMl_opt. The principle is elucidated with a function that is already composed of three parts, but the OS can segment the function in parts that behave essentially similarly in their mapping, e.g. relative brightening (partitioning algorithms are known, e.g. based on extent of deviation from a common joining line, but the change of derivative may also be a good candidate for partitioning; note that the portioning is not necessary in all embodiments, but will be used if some partitions converge or diverge more strongly from the diagonal than others).
[0170] E.g. for all the darkest lumas, up to first function point Pfl, we can take at least one distance of the input function (TM1) to be shape-adapted, to the diagonal. This initial distance dl can be calculated by establishing a direction of projection (the OS can use a fixed direction, e.g. 80 degrees i.e. 10 degrees more slanted than vertical down-projection). A final distance is determined, e.g. 80% of the initial distance for this bottom part of the curve (this ratio will depend, usually in a non-linear manner, corresponding with visual appearance, on the difference between the maximum luminance the original function TM1 was intended for, e.g. 2000 nit, and the current situation maximum, of the composite canvas; e.g. 2000 / 1000 is 2x more, but visually a factor 2 is not a large amount, so one need not deviate the distance d2 by a factor 2, but can keep it close to the initial distance). This corresponds to optimal mapping first point Pci. All intermediate points can be established correspondingly, meaning, if it is a line then scaled, and if the curve has non-linear shape, the corresponding differing respective distances can all be similarly scaled to 80%. For the middle part up to optimal mapping second point Pc2 respectively second function point Pf2, the OS could elect not to scale to 80%, but e.g. to 90%. The third function point Pf3 lies below the diagonal, meaning the output range should not be used in its entirety (because for this image, or these images, or this asset, 2000 nit is too much). The OS could in principle deviate to an optimal mapping third point Pc3 which lies deeper than Pf3, but then the re-optimized curve is only partially shape preserving. In general, it will scale upwards, since for a lower output maximum luminance (1000 nit) one does not want to make the brightest pixels too dim. These procedures together form the optimal brightness mapping curve TMl_opt for optimizing the lumas of the asset for the different maximum luminance situation of the tertiary brightness range of the total composite canvas. As can be seen, often this function will lie everywhere closer to the diagonal, but is still essentially shape preserving, meaning at least the variations of mapping over the input range, not of course the exact output values for any input (e.g. the middle part is still essentially a large contrast part, because usually there is some object of particular interest there, and the brightness of the brightest object pixels is still essentially kept under moderation). Note that any variants of operating system mapping behavior are not quintessential to the technical contribution of the present application, and merely introduced to elucidate some possible uses of the innovative computer systems or its components. So one example of application of the two color transform control data elements, is that all assets that are to be juxtaposed will be mapped to a common tertiary brightness range by using the reference white level to map such level to one or more final white levels, with an equi -luminance ratio (i.e. 60% brightness of that white in the original asset becomes 60% of the chosen final white level, or at least close to that value) for the darker colors, and the (typically HDR effect) brighter colors are mapped by a final brightness mapping function, which distributes the remaining lumas in each asset above its white reference level over the remaining colors in the tertiary brightness range for output to the display, and in a manner which tries to follow the shape of the communicated brightness mapping function for the asset. If the asset’s maximum luminance would fall above the maximum luminance of the tertiary range chosen by the OS, the function can be scaled so that the maximum of such a brighter asset maps to the final maximum brightness of the tertiary range. But other manners of processing guided by the color transform control data are also possible.
[0171] Fig. 10 gives some further detail on how the geometric graphical presentation of visual assets typically happens internally. A human user 1000 typically interacts with some user app / application (1001). Although some apps could be talking with deeper levels more directly, typically they may be using a more generic graphical user interface language 1002 (a.k.a. shell), like e.g. KDE Plasma or GNOME (Gnu Network Object Model Environment). This will contain a graphics protocol GRPROT, i.e. a set of API calls to interact with the operating system 1010. An example of a Linux graphics protocol is Wayland.
[0172] The operating system will typically contain (at least) three parts. Besides the kernel 1005, it may typically a window manager 1004 and a so-called display manager 1003, which managers can talk to each other (e.g. the window manager can call functionality of the display manager). The window manager will take care of the user interaction (e.g. mouse focus), and the state of the windows (position, size, transparency and z-order, window shadow). An example of a window manager for ChromeOS is Ash. An example of a dynamic window manager for the X window system is Awesome. The display server takes care of the low level drawing capabilities, e.g. it can draw a line or an area. So the above described luminance mapping processing may typically be performed by the display manager (or by some capability on request of that display manager, e.g. in a calculation circuit of a GPU). The API calls of the protocol will according to the present innovation communicate via the common pixel color array representation and one or more of the luminance mapping function (which can be formulated as a luma mapping function) and the reference white level, so this will typically be communicated over the graphics protocol GRPROT, but in complex operating system behavior it may also communicate in between modules of the OS, and even back to and back from another application, and even to the GPU, etc.
[0173] MacOS and iOS can use a display server like e.g. the Quartz display server (they can then use a somewhat different window manager). Android can use SurfaceFlinger.
[0174] The operating system can talk with the hardware (the GPU 1006 and its connection to a display 1007, which may also be bidirectional in case properties of the display need to be polled like its maximum displayable luminance, for optimizing the tertiary range of brightnesses, i.e. of relative brightnesses or absolute luminances), via a GPU API, e.g. use the more generic APIs like e.g. Vulkan (which is an open standard cross-platform API for 3D graphics and computing), or OpenGL, etc. For request of specific luminance versions of parts of assets also the protocol of communication with the GPU can use the present data formulation, but that is another concept that the communication by apps to, and coordination of luminance distributions of assets by, the OS.
[0175] The present innovative communication of HDR asset and its luminance re-grading desiderata metadata may also be incorporated into generic OS display server / window server abstraction languages.
[0176] Fig. 12 elucidates with a clearer graph a basic typical operation when two assets of different dynamic range from two different software applications are communicated (in the new manner according to an advantageous commonly agreed color representation) to an operating system, for it to compose a final image of the two assets together for display, with the asset pixel colors optimally determined on a final luminance dynamic range (which in this example the OS elected to go as bright as ML_COMP = 2000 nit). A creator of a first asset (Assl) made a movie of a nightly street scene. The application providing it may e.g. be an internet video streaming service. The creator chose to make a nice quality HDR movie, with pixels potentially becoming as bright as first maximum brightness ML Vorl being 5000 nit (e.g. in a later daytime scene of the movie), which is the maximum of the first brightness range DR_H1 of the first application’s one (or several) assets. A street light (obj_lmp) may have been originally master graded with pixels as bright as 2000 nit. The moon crescent (obj_moo) may be less excessive, e.g. at 200 nit. In that situation, that level may also be elected as a good Lambertian white point level, i.e. the metadata WL1, when represented in the original range (in practice it will typically be defined in the representative range, because those pixel colors is what the application will communicate to the OS, but that corresponding position can be calculated by a first brightness mapping function TM1 being here a luminance mapping function (to the third brightness range DR L1 that ends at 100 nit being the third maximum brightness ML_passl). A dark house (obj_hou) may have pixels spread e.g. around 10 nit. So a first pixel on the lamp may have a first original brightness Lui, and (like similarly all other pixels in the Asset (note, we are assuming global mapping functions for simplicity of elucidation)) via function TM1 it gets mapped to a corresponding first representative brightness (LuRl), on the range of the agreed common representation for this application’s assets. When the operating system receives this first representative brightness or in this example luminance LuRl, it can determine its optimal luminance mapping function to map it to a nicely bright luminance on the final range (output luminance Lol being 1500 nit). We have for understanding also shown the corresponding arrows of a first original luminance mapping FL*_opl, which the OS would do if it were to start from the first original brightness Lui. A second application serves weather maps (second asset Ass2, originally defined on second brightness range DR H2 ending at second maximum brightness ML_Vor2) from a totally different server, but currently also subscribed to by the computer’s configuration. Ergo, the OS needs to place that weather map somewhere besides the movie. The principle works similarly, since it also works via a pre- established common color representation en lieu the original asset HDR colors for communication to the OS. E.g., the creator may have considered that 600 nit is bright enough for a weather map, but that 150 nit is a good value for its Lambertian second reference white level (to have a nicely bright map), and, he may have specified flashy arrows indicating the wind direction, e.g. arrow obj_arr being 400 nit (e.g. uniformly for all its arrow pixels). That has to be fitted in a corresponding SDR common pre-agreed representation, which we have shown here to be agreed with the OS as having a fourth maximum brightness ML_pass2 being 120 nit, but in many scenarios both applications may communicate with the operating system according to the same common pre-agreed color space, i.e. both ending at 100 nit (or both ending at 120 nit). A corresponding second representative brightness LuR2 for the arrow resulting from a good second brightness function TM2 which relates it with second original brightness Lu2 (and therefore defines the original asset's colors by the common representation) may be 90 nit. Fourth brightness range DR L2 ends at ML_pass2. An optimal fourth brightness mapping function (FL_op2) will map the representative brightnesses (luminances) of the second asset to the fmal / output brightnesses as desired by the OS. E.g. 450 nit will be a good value for the arrows in the final range. The second white reference level WL2 may project to a good value for a common white reference level WLC being 180 nit. The second optimized original brightness mapping function FL*_op2 corresponding to the fourth brightness mapping function FL_op2 is also shown for convenience.
[0177] Some software applications may function simply by using one pre-fixed common color representation for the pixel color arrays (upon which they will base their metadata to define their original / master versions of their visual assets), e.g. always sRGB, or another standard dynamic range color format. That may be e.g. because they are written for some OS that is known to support such a common representation. As elucidated with Fig. 13, other, more advanced applications, in particular if they are to communicate with various operating systems, may agree a common color representation for the pixel color arrays with one or more operating systems. So application 1301 has an application-side color representation determining system 1310, and operating system 1302 has an operating system-side color representation system 1311 (or process). The determination of the commonly agreed common color representation may work in different manners, e.g. operating system forced (i.e. the operating system prescribes an API each time an application starts, or presents itself to the operating system). In the example we show a polling method. The application-side color representation determining system 1310 issues a request (PLL) spectify common color representation. The operating system-side color representation determining system 1311 may respond to that in various manners in various embodiments of communicating back a common color representation specification (COD API). E.g. some embodiments may send back a number, or a mnemonic or the like, and the application will know that e.g. “3” means that it should specify its color arrays as sRGB. It may also communicate back as specification COD API a name of the color space, e.g. “sRGB”. It may also send back a full-write out of the API, e.g. Asset_Colorimetry(sRGB, 10 bit, ...). The application will thereafter send all its visual assets according to the commonly agreed common color representation (at least until a new determination is performed, if any). More advanced versions may determine a common color representation depending on desiderata by the application and / or the OS.
[0178] According to some principles of the present approach an embodiment may be a computer (700), or a corresponding operating system (or corresponding color-communicating applications or methods), is arranged to run an operating system (504), wherein the operating system is arranged to manage the color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Assl) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness, wherein the visual asset is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness; wherein the visual asset in addition comprises color transformation control data (CTctrl) which comprises at least one of first brightness mapping function (TM1) or its inverse and a reference white level (WL1); wherein the operating system comprises a brightness mapping unit (710) arranged to receive the visual asset and to apply a tertiary brightness mapping function (FL opl) to the input luma code to obtain an output luma code, wherein the output luma code specifies a tertiary brightness which lies in a tertiary brightness range which is different from the secondary brightness range, wherein the output luma codes a brightness of an output color, wherein the tertiary brightness mapping function (FL opl) is based on at least one of, or both of, the first brightness mapping function (TM1) and the reference white level (WL1); wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (hnFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format (hnFinFmt).
[0179] To enable controlled display of various visual assets of different brightness dynamic range, a computer (700) and in particular its operating system pre-agrees with various software applications (typically from different sources and / or different creators) a method for communicating high dynamic range colors which involves re-mapping the original HDR pixel colors of the visual assets to one or more pre-agreed common color spaces or representations for the operating system, using secondary pixel colors e.g. defined in sRGB space for the pixel color arrays defining the geometrical structure of the assets, which assets communicate such arrays together with color transformation control data (CTctrl) which comprises at least one of a first brightness mapping function (TM1) or its inverse, and a reference white level (WL1) based on such secondary colors and linking the original HDR colors to the secondary colors, so that the operating system may elect a common color space for various software applications to communicate all their visual assets, and that the operating system can manage color coordination of visual assets received from one or more software applications (701, 702) during a geometrical composition of the assets in a total canvas, for a screenout buffer for providing pixels for one or more connected displays.
[0180] The algorithmic components disclosed in this text may (entirely or in part) be realized in practice as hardware (e.g. parts of an application specific integrated circuit) or as software running on a special digital signal processor, or a generic processor, etc. At least some of the elements of the various embodiments may be running on a fixed or configurable CPU, GPU, Digital Signal Processor, FPGA, Neural Processing Unit, Application Specific Integrated Circuit, microcontroller, SoC, etc. The images may be temporarily or for long term stored in various memories, in the vicinity of the processor(s) or remotely accessible e.g. over the internet.
[0181] It should be understandable to the skilled person from our presentation which components may be optional improvements and can be realized in combination with other components, and how (optional) steps of methods correspond to respective means of apparatuses, and vice versa. Some combinations will be taught by splitting the general teachings to partial teachings regarding one or more of the parts. The word “apparatus” in this application is used in its broadest sense, namely a group of means allowing the realization of a particular objective, and can hence e.g. be (a small circuit part of) an IC, or a dedicated appliance (such as an appliance with a display), or part of a networked system, etc. “Arrangement” is also intended to be used in the broadest sense, so it may comprise inter aha a single apparatus, a part of an apparatus, a collection of (parts of) cooperating apparatuses, etc.
[0182] The computer program product denotation should be understood to encompass any physical realization of a collection of commands enabling a generic or special purpose processor, after a series of loading steps (which may include intermediate conversion steps, such as translation to an intermediate language, and a final processor language) to enter the commands into the processor, and to execute any of the characteristic functions of an invention. In particular, the computer program product may be realized as data on a carrier such as e.g. a disk, data present in a memory, data travelling via a network connection -wired or wireless- . Apart from program code, characteristic data required for the program may also be embodied as a computer program product. Some of the technologies may be encompassed in signals, typically control signals for controlling one or more technical behaviors of e.g. a receiving apparatus, such as a television. Some circuits may be reconfigurable, and temporarily configured for particular processing by software. Some parts of the apparatuses may be specifically adapted to receive, parse and / or understand innovative signals.
[0183] Some of the steps required for the operation of the method may be already present in the functionality of the processor instead of described in the computer program product, such as data input and output steps.
[0184] It should be noted that the above-mentioned embodiments illustrate rather than limit the invention. Where the skilled person can easily realize a mapping of the presented examples to other regions of the claims, we have for conciseness not mentioned all these options in-depth. Apart from combinations of elements of the invention as combined in the claims, other combinations of the elements are possible. Any combination of elements can in practice be realized in a single dedicated element, or split elements.
[0185] Any reference sign between parentheses in the claim is not intended for limiting the claim. The word “comprising” does not exclude the presence of elements or aspects not listed in a claim. In several situations the word “portion” of a set of elements is not intended to exclude that portion may also cover the totality of the elements, because that may function equally in a same manner. The word “a” or “an” preceding an element does not exclude the presence of a plurality of such elements, nor the presence of other elements. “And / or” means that both options may be present together, or one of them may be present alone. The word “e.g.” is typically used to indicate that we mean that something else is also belonging to the possibilities, e.g. a similar element, example, or teaching, “i.a.” means inter alia, or among others. An element between ellipses will normally be used to indicate that something is optional, i.e. also possible as a variant of a more general concept, rather than necessary, e.g. (local) luminance boosting is intended to say, (primarily, as main level teaching) “luminance boosting” in general, which may be for all pixels the same, but may also be different, i.e. of the “local luminance boosting” variant, e.g. only applied to some locality of the image.
Claims
CLAIMS:
1. A computer (700) arranged to run an operating system (504), wherein the operating system is arranged to manage color coordination of visual assets received from a first software application (701) and a second software application (702), wherein the operating system comprises a geometrical management unit (71 l)arranged to manage spatial positioning in a canvas of the visual assets and arranged to receive a first visual asset (Assl) from the first software application comprising pixels of which a first pixel has a first pixel color specifying a first original brightness (Lui) lying in a first brightness range (DR Hl) which has a first maximum brightness (ML Vorl), and arranged to receive a second visual asset (Ass2) from the second software application comprising pixels of which a second pixel has a second pixel color specifying a second original brightness (Lu2) lying in a second brightness range (DR H2) which has a second maximum brightness (ML_Vor2), wherein the first maximum brightness is different from the second maximum brightness; wherein the first visual asset (Assl) is received comprising a set of first pixel color codes, which are defined in a first common color space pre-agreed between the first software application and the operating system, wherein a first pixel color code represents a first representative brightness (LuRl), which is derived from the first original brightness by applying a first brightness mapping function (TM1) to the first original brightness, which yields the first representative brightness as output of the first brightness mapping function, and wherein the second visual asset (Ass2) is received comprising a set of second pixel color codes, which are defined in a second common color space pre-agreed between the second software application and the operating system, wherein a second pixel color code represents a second representative brightness (LuR2), which is derived from the second original brightness by applying a second brightness mapping function (TM2) to the second original brightness, which yields the second representative brightness as output of the second brightness mapping function, wherein the first common color space may be the same as or different from the second common color space, wherein the first representative brightness lies in a third brightness range (DR_L1) which has a third maximum brightness (ML_passl) which is different from the first maximum brightness (ML_Vorl), and wherein the second representative brightness lies in a fourth brightness range (DR_L2) which has a fourth maximum brightness (ML_pass2) which is different from the second maximum brightness (ML_Vor2); wherein the first visual asset comprises first color transformation control data (CTctrl) which comprises at least one of the first brightness mapping function (TM1) or its inverse, and a first reference white level (WL1), which is lower than the third maximum brightness (ML_passl), andwherein the second visual asset comprises second color transformation control data (CTctrl) which comprises at least one of the second brightness mapping function (TM2) or its inverse, and a second reference white level (WL2), which is lower than the fourth maximum brightness (ML_pass2); wherein the operating system comprises a brightness mapping unit (710) arranged to receive the first visual asset and to apply a third brightness mapping function (FL opl) to the first representative brightness to obtain an output luminance (Lol), or a corresponding output luma code (Y’o) which codes the output luminance (Lol), wherein the output luminance (Lol)lies in a final brightness range (ML COMP) which is different from the third brightness range (DR L1), wherein the output luma codes a brightness of an output color, wherein the third brightness mapping function (FL opl) is based on at least one of the first brightness mapping function (TM1) and the first reference white level (WL1); wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
2. The computer as claimed in claim 1, wherein the first visual asset is received from a software application by first pixel color codes which are represented in a standard dynamic range color representation, such as e.g. sRGB.
3. The computer as claimed in claim 1 or 2, wherein the first original brightness specifies a luminance of a pixel to be displayed as an amount of nits.
4. An operating system (504) for being supplied to and for operation on a computer, wherein the operating system is arranged to manage color coordination of visual assets received from a first software applications (701) and a second software application (702), wherein the operating system comprises a geometrical management unit (711) arranged to manage spatial positioning in a canvas of visual assets, and arranged to receive a first visual asset (Assl) from the first software application which has a pixel which has a pixel color which has a first original brightness (Lui) lying in a first brightness range which has a first maximum brightness (ML_Vorl), wherein the first visual asset (Assl) is received comprising a set of pixel color codes which are defined in a first common color space pre-agreed between the operating system and the first software application, wherein a first pixel color code represents a first representative brightness (LuRl), which is derived from the first original brightness by applying a first brightness mapping function (TM1) to the first original brightnessand wherein the first representative brightness is defined on a range which ends at a third maximum brightness (ML_passl) which is different from the first maximum brightness (ML_Vorl), and arranged to receive a second visual asset (Ass2) from the second software application which has a pixel which has a pixel color which has a second original brightness (Lu2) which lies in a second brightness range which has a second maximum brightness (ML_Vor2), wherein the second visual asset (Ass2) is received comprising a set of pixel color codes which are defined in a second commoncolor space pre-agreed between the operating system and the second software application, wherein a second pixel color code represents a second representative brightness (LuRl), which is derived from the second original brightness by applying a second brightness mapping function (TM2) to the second original brightness, wherein the second representative brightness is defined on a range which ends at a fourth maximum brightness (ML_pass2) which is different from the second maximum brightness (ML_Vor2), wherein the first maximum brightness is different from the second maximum brightness, wherein the first visual asset comprises first color transformation control data (CTctrl) which comprises at least one of the first brightness mapping function (TM1) or its inverse, and a first reference white level (WL1), which is lower than the third maximum brightness (ML_passl), and wherein the second visual asset comprises second color transformation control data (CTctrl2) which comprises at least one of the second brightness mapping function (TM2) or its inverse, and a second reference white level (WL2), which is lower than the fourth maximum brightness (ML_pass2); wherein the operating system comprises a brightness mapping unit (710) arranged to receive the first visual asset and the second visual asset and to apply a third brightness mapping function (FL opl) to the first representative brightness to obtain an output luminance (Lol), or a corresponding output luma code (Y’o), wherein the output luminance lies in a final brightness range which is different from the third brightness range (DR L1), wherein the output luma codes a brightness of an output color, wherein the third brightness mapping function (FL opl) is based on at least one of the first brightness mapping function (TM1) and the first reference white level (WL1); wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (hnFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
5. The operating system as claimed in claim 4, wherein the first visual asset is received from a software application by color codes represented in a standard dynamic range color representation.
6. The operating system as claimed in claim 4, wherein the first original brightness specifies a luminance of a pixel to be displayed as an amount of nits.
7. A tangible data source, comprising code enabling execution on a computer of the operating system as claimed in claim 4.
8. A method of operating a computer, comprising a step of running an operating system, wherein the operating system is arranged to manage color coordination of a first visual asset (Assl) received from a first software application (701) and a second visual asset (Ass2) received from a second software application(702), wherein the operating system comprises a geometrical management unit (711) arranged to manage spatial positioning in a canvas of visual assets, and arranged to receive the first visualasset (Assl) comprising a pixel color having a first original brightness (Lui) lying in a first brightness range which has a first maximum brightness (ML_Vorl), and arranged to receive the second visual asset (Ass2) comprising a second pixel color having a second original brightness (Lu2) lying in a second brightness range which has a second maximum brightness (ML_Vor2), wherein the first maximum brightness is different from the second maximum brightness, wherein the first visual asset (Assl) is received comprising a set of pixel color codes, which are defined in a first common color space preagreed between the first software application and the operating system, wherein a first color code represents a first representative brightness (LuRl), which is derived from the first original brightness by applying a first brightness mapping function (TM1) to the first original brightness, wherein the first representative brightness is comprised in a third brightness range which has a third maximum brightness (ML_passl) which is different from the first maximum brightness (ML_Vorl), and wherein the second visual asset (Ass2) is received comprising a set of pixel color codes, which are defined in a second common color space pre-agreed between the second software application and the operating system, wherein a second color code represents a second representative brightness (LuR2), which is derived from the second original brightness by applying a second brightness mapping function (TM2) to the second original brightness, wherein the second representative brightness is comprised in a fourth brightness range which has a second maximum brightness (ML_pass2) which is different from the second maximum brightness (ML_Vor2); wherein the first visual asset comprises first color transformation control data (CTctrl) which comprises at least one of the first brightness mapping function (TM1) or its inverse, and a first reference white level (WL1), which is lower than the third maximum brightness (DR L1), and wherein the second visual asset comprises second color transformation control data (CTctrll) which comprises at least one of the second brightness mapping function (TM2) or its inverse, and a second reference white level (WL2), which is lower than the fourth maximum brightness (DR L2); wherein the operating system comprises a brightness mapping unit (710) arranged to receive the first and second visual asset, and to apply a third brightness mapping function (FL opl) to the first representative brightness (LuRl) to obtain an output luminance (Lol), or a corresponding output luma code (Y’o), wherein the output luminance lies in a final brightness range which is different from the third brightness range (DR L1), wherein the output luma codes a brightness of an output color, wherein the third brightness mapping function (FL opl) is based on at least one of the first brightness mapping function (TM1) and the first reference white level (WL1); wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (hnFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
9. The method of operating a computer as claimed in claim 8, wherein the first visual asset is received from a software application by color codes represented in a standard dynamic range color representation.
10. The method of operating a computer as claimed in claim 8, wherein the first original brightness specifies a luminance of a pixel to be displayed as an amount of nits.
11. A software application (502, or 503) arranged to create at least one visual asset (Ass 1) comprising pixels wherein a pixel has an original color code which specifies an original brightness (Lui) and to communicate such visual asset to an operating system (504), wherein the original brightness lies in a first brightness range (DR Hl) which has a first maximum brightness (ML Vorl); wherein the original color code is transformed into a secondary color code for communication to the operating system, wherein the secondary color code is represented in a common color space pre-agreed with the operating system, which codes a secondary brightness (LuRl) which lies in a secondary brightness range (DR L1) which is different from the first brightness range and which ends at a third maximum brightness (ML_passl), wherein the secondary brightness is derived from the original brightness based on application of a first brightness mapping function (TM1) to the original brightness; characterized in that the software application communicates color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Assl), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1), which has a lower value than a third maximum brightness (ML_passl).
12. The software application as claimed in claim 11 wherein the common color space is demanded by the operating system for all communications of visual assets by the software application to the operating system.
13. The software application as claimed in claim 12 comprising an application-side color representation determining system (1310) which is arranged to coordinate with or receive from the operating system a specification of the common color space.
14. The software application (502, or 503) as claimed in claim 11, arranged to communicate the asset to the operating system in a standard dynamic range representation.
15. A method of communicating a visual asset having pixel colors to an operating system, comprising the steps of: creating at least one visual asset (Assl) comprising pixels wherein a pixel has an original color code which specifies an original brightness (Lui), wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_Vorl); transforming the original color code into a secondary color code for communication to the operating system, wherein the second color code is defined in a common color space pre-agreed betweenthe method and the operating system, wherein the secondary color code represents a secondary brightness (LuRl) which lies in a secondary brightness range which is different from the first brightness range and ends at a third maximum brightness (ML_passl), wherein the secondary brightness is derived from the original brightness based on application of a first brightness mapping function (TM1) to the original brightness; communicating the secondary color code and color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Assl), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1), which has a lower value than the third maximum brightness (ML_pass 1 ) .
16. The method of communicating a visual asset to an operating system as claimed in claim15, wherein the secondary color code is a standard dynamic range color code.
17. A computer program product comprising code enabling a computer to execute the software application (502, or 503) as claimed in claim 11.
Citation Information
Patent Citations
Local dynamic range adjustment color processing
US20170347113A1
Optimizing high dynamic range images for particular displays
WO2017108906A1
Encoding and decoding HDR videos
WO2017157977A1
Systems and methods for appearance mapping for compositing overlay graphics
US20180322679A1
Handling multiple HDR image sources
US20180336669A1