Improved hdr color processing for saturated colors
By using color lookup tables and luminance mapping functions in high dynamic range video coding circuits, the problem of tonal inhomogeneity in highly saturated color regions is solved, thus improving image quality.
Patent Information
- Application Number
- CN202180016828.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-25
- Filing Date
- 2021-02-12
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2041-02-12
AI Technical Summary
Existing technologies suffer from tonal inhomogeneity and artifacts when processing high dynamic range images, especially highly saturated colors, leading to a decrease in output image quality.
By introducing color lookup tables and luminance mapping functions into the high dynamic range video coding circuit, the luminance and chrominance values of pixels are adjusted to ensure color consistency and realism during dynamic range conversion.
It effectively reduces artifacts in highly saturated color areas, improves the visual quality of the output image, and ensures the uniformity and naturalness of the color tone.
Smart Images

Figure CN115176469B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a method and apparatus for performing a dynamic range conversion involving a change of luminance of pixels from an input image to an output image, in particular when the image contains highly saturated pixel colors. BACKGROUND
[0002] High dynamic range video processing (encoding, display adaptation, etc.) is a recent technical field which still comes with several unsolved problems and challenges. Although HDR televisions are now sold for several years (typically with a maximum displayable luminance, also called peak luminance, of 1000 nit or Cd / m^2 or less), the technologies which appear before the pure display (content production, encoding, color processing) still have many technical solutions to invent / improve and deploy. Many movies have been created and distributed and some early broadcast, and although the results are generally good, there is still a possibility to further improve some aspects, so the technology is not currently in a phase where it is massively solved.
[0003] A high dynamic range image is defined as an image compared to a traditional low dynamic range image, also called standard dynamic range image. Standard dynamic range images (which were generated and displayed in the second half of the 20th century and are still the mainstream of most video technologies, e.g. through television or movie distribution by any technology, from terrestrial broadcast to youtube video supply via internet, etc.) generally have more impressive colors, which means that generally they can have brighter pixels (encoded and displayed).
[0004] Formally, one can define the luminance dynamic range as the span of all luminances from the minimum black (MB) to the peak white or peak luminance (PB), so in principle one can have a HDR movie with very dark blacks. In practice, one can define and process (e.g. color process) a HDR image mainly based on the unique value, i.e. the higher peak luminance or more formally luminance (typically this is what the user is most interested in, whether it is a bright explosion or just a more realistic specular reflection spot on metals and jewels, etc., and one can practically state that the minimum black is the same for SDR and HDR images).
[0005] In practice, one can declare that the 1000:1 luminance dynamic range of SDR images will drop below 100 nit (and above 0.1 nit) and that HDR images will typically have at least 5 times brighter PBs, thus 500 nit or more. More precisely, one will display a legacy SDR image on an end-user display that displays a peak brightness (PB D) of about 100 nit even if the legacy SDR image does not have a well-defined maximum luminance. And one can reinterpret a SDR image in the novel HDR framework so that its brightest pixel luminance corresponds exactly to 100 nit.
[0006] Note that without going into many details that can not be necessary here, one can indicate which HDR an image has by associating a peak luminance number as metadata with the HDR image. This can be seen as the luminance of the brightest pixel present in the image, or more precisely, the luminance of the brightest encodable image pixel in the image video, and is typically formally defined by associating a reference display with the image (i.e. one associates a virtual display with the image whose peak luminance corresponds to the brightest pixel present in the image or video, i.e. the brightest pixel that needs to be displayed, and then one encodes this PB C-C for the encoding of this virtual display as metadata outside the image pixel color matrix). The skilled person can understand that whatever the encoded peak luminance PB C of an HDR image is, one can create normalized pixel luminances by dividing each pixel luminance by PB C. Normalized brightness can also be done by dividing by the highest possible code (e.g. 1023 in 10 bits, etc.).
[0007] In this way, one does not need to "ugly" encode HDR images with an excess of bits (e.g. 16 bits per color component) and can simply reuse existing technologies with 10 bits word length for color components (several technology suppliers have already considered 12 bits or more as quite heavy in the whole video processing chain for various reasons, at least in the near future, and typically only considering professional non-consumer usage; and PB C metadata allows to easily upgrade the HDR framework for future scenarios).
[0008] One aspect of the reuse of pixel colors from classical video engineering is that pixel colors are typically always transmitted as YCbCr colors, where Y is the so-called luma component (which is a technical encoding for luminance) and Cb and Cr are the blue and red chroma components, completing the trichromatic additive color definition.
[0009] Luminance is defined from the non-linear R'G'B' color components (the superscript'indicates the difference to linear RGB components of a color, one could say the amount of red, green and blue photons coming out of a display pixel of any particular color), and via the following equation: Y = a1*R' + a2*G' + a3*B'. [Equation 1]
[0010] These three constants depend on the particular primary colors defining the RGB color (obviously, if one defines a color by mixing 50% of a very saturated strong red to the red primary color, one will get a more red color than in the same R = 0.5; G = X; B = X color definition in a set of red primary colors where red is a weak pink primary color).
[0011] In fact, for each triad color defining the primaries, these constants a1, a2, a3 are uniquely determined as the constants that yield exact (absolute or typically with 1.0 relative to peak white) luminance in case of linear RGB components filling Equation 1.
[0012] The non-linear (original) components are derived from the linear components by applying an opto-electric transfer function. For HDR, this is a function with a steeper slope for the darkest linear input components than about square root (Rec. 709 OETF of SDR), e.g. the perceptual quantizer function standardized in SMPTE 2084.
[0013] In the LDR era, there was not much debate, as for high-end video, one only used Rec. 709 primaries, but now also with e.g. Rec. 2020, wide color gamut primaries.
[0014] Then, Cb is derived from Cb = b1*(B' - Y) and Cr = c1*(R' - Y), where b1 and c1 are again constants fixed from the chosen master system (so that there is no overflow in the normalized to 1.0 representation).
[0015] The second question one needs to ask is how the non-linear R'G'B' relates to (is actually defined for) the linear RGB components. For each color to be created by adding (at least within the color gamut of reproducible colors, e.g. within the triangle of the three chroma positions of the chosen primaries, in e.g. the standard CIE 1976 u'v' color plane; note that in principle one can choose the primaries arbitrarily, but it makes sense to define images in primaries that can also be displayed by actual displays on the market, i.e. not e.g. single-wavelength laser primaries), this comes from basic chroma metrics (for any color - e.g. single-wavelength laser primaries) CIE XYZ or hue, saturation, luminance - one can calculate the corresponding metameric RGB triad that a particular display with such primaries should display to display the same color to an average observer).
[0016] Definition of the encoding system, i.e. the so-called electro- optical transfer function, or its inverse optical-electrical transfer function, which calculates the non-linear components from the corresponding linear components, for example:
[0017] R' = OETF PQ(R) [Equation 2],
[0018] is a technical problem.
[0019] Note that one can also mathematically prove that even if the OETF is applied to the red, green and blue pixel color components, the luminance and the luminance are related via the same OETF (the achromatic gray has R=G=B and R'=G'=B', thus the luminance L=(a1+a2+a3=1.0)*0.x and the luminance Y=1.0*OETF[0.x]=OETF[L]).
[0020] For example, while in the LDR era there was only the Rec.709 OETF, which (simplifying some irrelevant details of the present patent application) is the inverse of the EOTF, and this is a quite good approximation of the simple square root.
[0021] Then, when the technical problem arose of encoding a large range of HDR luminances (e.g. 1 / 10,000 nit - 10,000 nit) in only 10 bits (which is not possible for the square root function), a new EOTF was invented, the so-called perceptual quantizer EOTF (US 9077994).
[0022] Therefore, this different EOTF definition and any input YCbCr color defined by it will clearly have different normalized components depending on whether it is YCbCr_Rec709 or YCbCr_PQ (in fact, one can see this by turning the R'G'B' cube on its black top end, where the y-axis of the different luminances of the achromatic gray now forms a vertical: the various image pixel colors will then have different extensions along this vertical depending on whether these same image object pixel colors are represented in the 709-based color cube or in the PQ-based color cube).
[0023] When combined with the HDR codec that the present applicant has created together with Technicolor, this given technical fact of YCbCr_PQ encoding will lead to new interesting experimental facts (see WO2017157977 and ETSI standard TS 103433-2 V1.1.1 "High performace Single Layer High Dynamic Range [SLHDR], Part 2"; and here in the present patent application with Figure 1 re-summarized).
[0024] Figure 1 Our decoder is shown, which can take a pixel color input image (IM_1) defined in YCbCr_PQ, where the pixels are processed one after the other in a scan line. This image can for example come from a Blu-ray disc, where the PQ normalized luminance corresponds to a PQ peak luminance of 10000 nit. The shape of the function needed to degrade to for example an SDR image can come from a grader, a human or an automatic machine, for example (non-limiting) computed when the movie on the disc is transcoded.
[0025] The upper track of the processing circuit (which is typically a hardware block on an integrated circuit, and can have various configurations, for example the whole luminance processing sub-circuit 101 can actually be embodied as a LUT, a so-called P_LUT (working in a perceptually equalized luminance), but also sub-blocks can be embodied as separate transistor computation engines of an integrated circuit, etc., but we need to technically interpret what in the block produces a technical benefit to the appearance of the image and any possible technical problems, and in particular the one processed below) is concerned with the luminance processing part of the pixel color. This is of course the most interesting part of the dynamic range adjustment: when for example computing a 600 nit output image from an input (for example a 1000 nit PB_C HDR image), the normalized pixel luminances of various image objects need to be redistributed along the Y axis of the YCbCr color space, i.e. typically the darker pixel luminances of the image will be shifted upwards towards the brighter (normalized) gray scale of the Y axis, compressing the brighter pixel luminances into a smaller sub-range of the Y axis. Figure 1 The bottom of the circuit of contains the color processing, in particular of the color saturation of the pixels, because we want to preserve the same color hue in the input and output image, which should also be noted for a good quality dynamic range conversion.
[0026] In this particular illustrated embodiment, the output pixel color is also defined with a PQ OETF (or technically speaking an inverse EOTF) non-linearity, yielding non-linear R", G" and B" components (but with the correct color mapping, also called re-grading, according to the Applicant's method) that can be directly transmitted to for example an HD RTV expecting such an input.
[0027] Some integrated circuit units are static in the sense that they apply fixed curves, matrix multiplications, etc., and some blocks are dynamic in the sense that they can adjust their behavior. In particular, they can potentially adapt the shape of their function to each successive image of a video - e.g. the degree to which the mapping function rises on the darkest inputs and the degree to which the mapping function compresses the brightest inputs. Indeed, the problem of image processing is that the content is always almost infinitely variable (the input video can be a dark Hollywood movie of a theft that took place at night, or a youtube movie in which the bright colors of a computer graphics following someone playing an online game, etc.; i.e. the input, although all coded in YCbCr, can vary greatly in any image about its pixel colors and luminances, whether the content is generated from a natural scene by any camera, professional or low quality, or generated as graphics, etc.). Thus, a night scene can require a different mapping of the darkest luminances or indeed luminances of the input image Im_in than e.g. a sunny beach scene image.
[0028] Thus, some circuits (rough dynamic range converter 112, customizable converter 113 and color lookup table 102) can apply functions of variable shape, the shape of which is controlled by parameters from an image-related information source 199. Typically, this source is fed with the image (Im_in) and metadata (MET) which includes parameters determining the shape of various luminance mapping functions applied by those dynamic circuits. This regrading function metadata can be transmitted from the content creation (encoder) side, transmitted through any video communication network (cable, internet, etc.) and temporarily stored in a memory in the decoder or connected to the decoder. Then, the decoder can apply those mappings prescribed by the video creator.
[0029] The pixel colors in the image appear in the form of YCbCr_PQ, i.e. all three non-linear R', G' and B' components are encoded with a perceptual quantizer function. The luminance component is sent through the PQ EOTF, the luminance computation circuit 129 being configured to compute the PQ EOTF (on achromatic grays, this can be done without problem since those luminance codes uniquely represent various normalized luminances).
[0030] Subsequently, this normalized luminance L is perceptually uniformly perceived by the perceptualizer 111, yielding a perceptually uniform luminance PY as output. This is performed by the following formula:
[0031] RHO(PB_C_HDR) = 1 + 32*power[PB_C_HDR / 10,000; 1 / (2.4)]
[0032] PY = log10 [1 + (RHO(PB_C_HDR) - 1) * power[L; 1 / (2.4)]] / log10[RHO(PB_C_HDR)] [Equation 3]
[0033] We see in this perception a dependency on the maximum value that can be encoded in the HDR input image, i.e. the maximum (absolute) pixel luminance PB_C_HDR that can occur in the image (i.e. to be displayed), which can also be transmitted as part of the metadata (MET) in general. Indeed, this already gives a quite reasonable distribution of the image pixel luminances and eventually at display gives a luminance from a certain input peak luminance to an output peak luminance, but in order to have a graded image of better quality, the content creator (encoding side) can continue to remap the luminances in further configurable circuits (i.e. 112 and 113).
[0034] Subsequently, the coarse dynamic range converter 112 applies a coarse good dynamic range reduction function, the shape of which is symbolically shown as a square root (the actual function that the applicant found useful consists of two linear segments with slopes configurable by metadata, in between with a parabolic segment, and the interested reader can find those details in the SLHDR ETSI standard as they are not relevant to the present discussion), and this function computes from the input perceived uniform luminance PY of the current processed set of pixels along the scan of the image a corresponding coarse luminance CY. The idea is, and this is already sufficient for many cases, to make relatively brighter the darker pixels of the HDR compared to the brighter ones, so that the overall fits in the smaller dynamic range.
[0035] Subsequently, the customizable converter 113 applies another selectable (e.g. determined under control of a color grader or shader that encodes the original captured video or movie at the content creation side, then communicates its chosen function as metadata to the regrading of this particular image or running of successive images from the video) luminance mapping function to the previously obtained coarse luminance CY to obtain a perceived output luminance PO. The shape of this mapping function can be determined at will (e.g. contrast in a subset of luminances corresponding to a certain important object or area in the image can be boosted by lifting this part of the function above the trend of the rest of the function).
[0036] As illustrated, we will assume (non limitatively) that this function is a sigmoid function, i.e. with a relatively weak slope at the lower and upper input luminance sub-ranges and a steep slope in the middle. The invention is not limited to a sigmoid function applied by the customizable converter 113 and is particularly suitable for any scenario where there is a varying slope behavior, in particular in the lower sub-range of input luminance (e.g. starting at 0.0 input value and ending at some value 0.x, typically lower than 0.5), e.g. of the type having a relatively low or average slope for very dark pixels compared to the slope of brighter pixels of the lower sub-range of input luminance. The slope can be considered as formulated from the starting point (0.0, 0.0).
[0037] Finally, the perceived output luminance PO, i.e. which has been fully luminance re- graded as the current image content would require, i.e. giving the best re-graded image corresponding to the input master HDR image (e.g. being output as an output SDR image with PB_C = 100 nit) (e.g. as far as the reduced dynamic range allows, so that luminance and contrast look as close as possible to their visual impact as they had in the master HDR image, or alternatively at least produce a good looking or well viewable image, etc.) is again linearized into normalized output luminance L_out, the linearizer 114 using the inverse of formula 1, but now with the PB_C_OUT value. For the sake of explanation, we will assume that the decoder downgrades the input HDR image to an SDR image, i.e. PB_C_OUT = 100 nit (as acknowledged by the standardization), but it could also downgrade to e.g. a 650 nit HDR output image, etc. (then the function would have a slightly different shape, but showing adaptation details is also useless for understanding the new technical principles of the application and can be found in said ETSI standard).
[0038] The lower track concerns chrominance processing, or in other words, the completion of the 3D color processing belonging to the luminance processing (as a matter of fact, any luminance processing on a color image is actually a 3D color processing, whether or not the chrominance components are carefully processed). This is technically more interesting than the dynamic range change behavior for non-color colors on the gray scale axis.
[0039] In the case of luminance, which quantifies colors in general, Cr and Cb quantify their hue (ratio of Cb / Cr) and their saturation (magnitude of Cb and Cr). Formally, Cb and Cr are chrominances.
[0040] Moreover, our codec processing circuit has a color lookup table 102 that specifies a function of Y_PQ that has a shape selected and again configurable by metadata. Thus, the processing of chroma is defined by two chroma multiplied by the same constant B, and this constant is a function of the luminance Y_PQ of any processed pixel, and the shape of this function, i.e. each respective value of B(Y_PQ) for any possible Y_PQ, is also configurable.
[0041] The main purpose of this LUT will be to correct over high or low amounts of Cb and Cr components at least for the altered (dynamic range adjusted) normalized luminance Y, since in fact Cb and Cr evolve together with the luminance in this color space representation YCbCr (the saturation of colors is in fact related to Cb / Y and Cr / Y), but other chroma aspects of the image can change with it, for example the saturation of a bright blue sky can be adjusted. In any case, given that the color LUT TCL(Y_PQ) is being loaded for dynamic range processing at least the current video image, the color lookup table 102 will produce a multiplier B for the Y_PQ value of each pixel running / being color processed. The input chroma components Cb_PQ and Cr_PQ of the pixel will be multiplied by this value B by the multiplier 121 to produce the corrected chroma components CbCr_COR. Then a matrixing operation is applied by the matrixer 122, the details of which can be found in the ETSI standard, to obtain normalized R’G’B’_norm components. These components are not normalized in the usual sense that they lie within a non-linear color cube between 0.0 and 1.0, but they are normalized in the sense that they are dynamic range free (all clustered together, just “chroma”, as a triplet instead of the usual pair, i.e. not yet any luminance or brightness). Thus, to obtain actual correct 3D colors, they must be multiplied by a luminance, and precisely the luminance that we have best determined in the upper luminance processing sub-circuit. For this, the normalized output luminance L_out needs to be PQ OETF mapped by the PQ OETF circuit 115. The multiplier 123 corrects multiplying scaling of the three normalized R’G’B_norm components output by the chroma sub-circuit.
[0042] Finally, the red, green and blue components R" G" B" _PQ defined by the perceptual quantizer can be sent directly for example to an HDR display, which is configured to understand such input, and it can directly display this input (e.g. a 650 nit PB_C output image can have been optimized by our technology to drive a 650 nit actual display peak luminance PB_D display, so that this display does not need to figure out by itself how to handle the difference between the luminance range encoded in the provided image and its display capabilities). Note that we are here setting out the basic core technology element behavior to explain the current technology innovation. In a real SLHDR decoder standardized by ETSI, there is also a black limiter circuit subunit, but this is not essential to the present method, and therefore not essential to its explanation (the black limiter can or can not be present in various embodiments, so discussing it in the introduction would create more message diffusion than the basic teaching, but for the interested reader we show this variant in Figure 5
[0043] This chroma optimization of various images, while slightly more complex because our various units have been designed over the years to give the best approach for all situations and applications (e.g. offline graded movie content versus real time broadcast etc.), can now be understood and works very satisfactorily in most of the many different input image situations.
[0044] However, as explained by means of Figure 2 there are some somewhat exotic situations where color decoding can still be improved. This problem will be addressed with the new technology elements taught.
[0045] Figure 2 a shows the ideal YCbCr color space, which can be formed using normalized luminance L instead of luminance Y as the vertical axis. A certain hue of blue lies in the vertical cross-section through the triangle of this diamond. More saturated colors have a greater angle with the vertical, so color 201 is a light blue of the same hue, but less saturated, and color 202 is a more saturated blue. These two colors can have the same luminance, and this can easily be read off the central vertical axis luminance scale (e.g. luminance levels LI vs. L2). We can verify that the two different colors indeed have the same luminance, as the luminance of (pixel) color 201 and (pixel) color 202 is the same, i.e. LI.
[0046] However, luminance was never designed to be a perfect (easy to use, uniform, etc.) color representation for human vision, compared to luminance that is unique in chrominance, but was designed so that it is easy to encode a triplet of colors as a 3x8-bit system, or now typically 3x10-bit for HDR (the reader is informed that this 3x10-bit can encode not only several times brighter, but up to 100 times brighter, and in addition also much darker, HDR!). Hence, luminance and in particular the YCbCr color representation is mainly designed for (invertible) encoding, not for color processing.
[0047] This encoding is in principle easy to mathematically invert, so it is always possible - because the encoding system should have its main property - to derive the original intended (linear) RGB components from the received YCbCr pixel color encoding, and then display these components. And one can also derive the luminance.
[0048] However, if one starts with a (needed) more advanced decoding of the video of the current HDR era, some inconvenient problems can occur. In fact, as Figure 2 B shows that for a color (i.e. it is not a pure gray color, whose pure gray color falls on the vertical axis with Cb=Cr=0 component values), the luminance Y no longer correctly or uniquely encodes the pixel luminance (i.e. the luminance immediately after applying equation 1 on the original linear RGB components), i.e. it cannot be converted to luminance by calculating L=EOTF(Y). In fact, for any chosen fixed luminance value L of a color, the higher the saturation S of the color, the lower the corresponding luminance Y. In fact, as Figure 2The typical case illustrated by D is that if we have more or less the same uniform illumination and thus more or less the same luminance for all pixels of a flower, but an object with different saturation for pixels (like a purple flower), then when mapped on the normalized level axis of luminance (corresponding to the vertical axis count N(L_in)), the tight almost single value histogram of pixel luminance 205 results in a corresponding luminance histogram 206 (count N(PY)) which is not only at lower normalized (input) values, but also more spread out. We can map the normalized (or absolute) luminance to the position on the luminance axis defined by the given OETF, because we know that for achromatic pixels, luminance Y uniquely corresponds to the normalized luminance (by remapping with the OETF shape: Y = "L_in" = OETF (pixel luminance)).
[0049] This creates several problems. First, we do not exactly know what luminance we have by just looking at the luminance values Y (i.e. by applying the EOTF to Y), and this can already give problems in a framework that processes 3D colorimetry in a natural 3D way, but processes it in a 1D+2D way for practical reasons, like the decoder of Figure 1 In fact, theoretically, one could do everything right if one introduced a correction dependency of the specific values of the pixels of Cb and Cr, but this can not be possible in some pre-designed color processing topologies. Moreover, one can see that we actually apply a luminance or luminance re- scaling function to the (perceptually uniformized) luminance, because the input image is already encoded by pixel luminance and not by luminance. So we can do something wrong, at least in theory.
[0050] However, this is usually not a (real) problem, as tested with many different kinds of image content over the last years, and it is reasonable to see that the technology does not see enough error magnitude to be a concern in practice. So it can be argued that most of the time it is fine to simply do the processing of Figure 1 and if occasionally the output image is not perfect, then one would solve that problem.
[0051] However, in very specific cases, image artifacts can be observed that are so large that they are necessary to guarantee an improved solution, as given below. This can occur, on the one hand, when there are very saturated colors (e.g., in Rec.2020, highly saturated colors can be produced; this usually doesn't happen because such colors are rarely seen in real life, as most normal image colors would have more desaturated CbCr values on the saturated basis of such color primaries, but it is at least theoretically possible), and on the other hand, when not only the smooth coarse illuminance mapping function of the coarse dynamic range converter 112 is applied, but also some more advanced function in which some inconvenient bending is applied, similar to, for example, the S-curve subsequently in the customizable converter 113, which has a low value corresponding to the low input luminance CY of the output PO value (which is similar to, for example, the S-curve in the customizable converter 113). Figure 2 The illuminance comparison shown in D would be incorrectly too low, and furthermore, the differences are potentially different for individual points on the object. That is, in particular, the variation in output color (especially illuminance) along the object caused by this overall processing can be considered problematic. Unfortunately, the patent system does not allow the inclusion of actual color images in the patent application specification, but... Figure 2 C symbolizes what will be seen, for example, in a saturated purple flower: there may be over-visibility of dark spots 203 or microstructures 204 in certain locations within the flower, which are normally invisible in the flower; not in the input image and not in the real world. A well-looking flower will appear fairly smooth, or at least have gradually changing illumination, rather than having such sharp and significant variations, because some colors that should be color-processed jump equally to different PY values, which span points where such a significant change in total illumination re-grading occurs.
[0052] exist Figure 2In d, the difference L_in between the normalized (chrominance dependent) luminance (PY) of the pixels (as they will be received) and the normalized illumination (or more precisely, illumination-luminance, i.e. the luminance that would be obtained on the achromatic axis, i.e. the correct illumination of the pixels of luminance PY) is shown by their respective luminance histogram 206 and illumination histogram 205 (N indicates the count of each value). For a single pixel, there will of course be one luminance PY and one corresponding correct illumination L_in, but we show the extension of multiple pixels of an object (e.g. a flower) with a certain average illumination and saturation. We see that if e.g. a saturated purple flower is uniformly illuminated, there can only be a few illumination values, but because there can be different color saturations in the flower pixels, the luminance histogram can be more spread out. In any case, they are "incorrect" - a one-dimensional representation with respect to illumination. As will be shown below, this can lead to processing errors in our SLHDR decoding method, which will need to be corrected backwards, and specifically as taught and claimed.
[0053] One could theoretically consider building a better, different 3D based decoding system that could easily solve the problem, but in the practically technically limited already deployed video ecosystem (with separate illumination processing and chrominance processing), this is not easily done in practice. In any case, given that we have such an encoder and corresponding decoder, there is not much that can be changed to improve it also for the last tiny inconvenience, specifically some incorrect illumination in saturated chrominance objects, so trying to find a practical solution can involve quite a bit of thinking and experimentation.
[0054] The following prior art documents have been found and briefly discussed for their tangential relevance.
[0055] Some background teachings of interest to the present applicant are as follows.
[0056] EP3496028 teaches another improvement of the applicant's basic luminance processing and parallel chrominance processing encoding or decoding pixel pipeline, namely that under certain luminance mappings, some output colors or specifically their luminance can fall above the upper limit of the encodable color gamut. Therefore, a strategy is needed to keep them within the color gamut without preferably affecting the darker colors too much. Furthermore, a specific mapping of the luminance inwards to the highest area within the 3D color gamut is taught. Some luminance down-mappings can be interchanged to reduce saturation so that excessive down-mapping is needed, but for a certain percentage of the needed luminance down-mapping, it is also possible to map towards the neutral gray axis, thereby slightly unsaturating the brightest colors. In any case, this does not happen through the G_PQ calculation as in the present invention, and also not on the different color problem in the other range of the color gamut.
[0057] WO2017 / 157977 relates to specific luminance processing of the darkest luminance in an image. For example, a video content creator can want the colors of a secondary image determinable from a primary master HDR image to be too dark for some dark pixels, which can impact the quality of the inverse luminance mapping used for decoding. Moreover, two parallel strategies of alternative luminance mapping are taught, wherein a second, safer luminance mapping can be launched against the problematic desired luminance mapping curve of the content creator. The invention does not deal with specific details of chrominance mapping as in the present application, except for mentioning standard means of chrominance mapping only.
[0058] US2018 / 0005356 is a specific generally good luminance mapping by means of configurable linear slopes of the darkest and brightest sub-ranges of the HDR pixel luminance, with parabolic smooth mapping in between. This luminance mapping can or can not be applied in the luminance processing track of our method, but is not the teaching about chrominance processing with color LUTs. SUMMARY
[0059] The remaining problem of SLHDR systems for infrequently occurring image pixel colors can be mitigated by a high dynamic range video encoding circuit (300) configured to encode a high dynamic range image (IM HDR) having a first maximum pixel luminance (PB CI) together with a second image (Im LWRDR) having a lower dynamic range and a corresponding lower second maximum pixel luminance (PB C2),
[0060] wherein the second image is functionally encoded as a luminance mapping function (400) for a decoder to be applied to pixel luminances (Y PQ) of the high dynamic range image to obtain corresponding pixel luminances (PO) of the second image,
[0061] the encoder comprising a data formatter (304) configured to output the high dynamic range image and metadata (MET) encoding the luminance mapping function (400) to a video communication medium (399),
[0062] the functional encoding of the second image is further based on a color lookup table (CL(Y PQ)) encoding a multiplier constant (B) for all possible values of pixel luminances of the high dynamic range image, the multiplier constant (B) to be used for multiplication with chrominances (Cb, Cr) of the pixels, and the formatter is configured to output this color lookup table in the metadata,
[0063] characterized in that the high dynamic range video encoding circuit comprises:
[0064] a gain determination circuit (302) receiving pixels of the high dynamic range image (IM HDR), each pixel having a luminance and two chrominances, the gain determination circuit being configured to determine, for each pixel, a luminance gain value (G PQ) quantifying a ratio of a first output luminance for a luminance equal to a normalized illuminance of the luminance of each pixel divided by a second output luminance for the luminance of the pixel, wherein the first output luminance is obtained by applying the luminance mapping function to the normalized illuminance and the second output luminance is obtained by applying a luminance mapping function to the luminance of the pixel;
[0065] wherein the high dynamic range video encoding circuit comprises a color lookup table determination circuit (303) configured to determine the color lookup table (CL(Y PQ)) based on values of the luminance gain values for various luminances of pixels present in the high dynamic range image, wherein values of the multiplier constant (B) for luminances from at least a subset of possible values of the luminance (Y PQ) are related to values of the determined luminance gains (G PQ) for those luminances.
[0066] For clarity, the encoding circuit not only conveys the HDR image, but also at least a corresponding secondary graded image, typically having a lower dynamic range, and typically a standard dynamic range image of PB_C_SDR = 100 nit. This ensures that the receiver does not have to guess what to do in terms of luminance mapping when receiving a high quality HDR image (IM HDR) having a peak luminance PB C much higher than the peak luminance capability (PB D) of the display on the receiving side. In fact, with the luminance mapping function, the content creator (e.g. his color grader) can specifically determine how the receiver should map all luminances in the input HDR image to corresponding SDR luminances according to his wishes (e.g. a shadow area in a dark cave should also remain bright enough in the SDR image so that a person with a knife hidden there can still be seen sufficiently, not too obviously, and not almost not at all, etc.). By conveying at least one luminance mapping function (400) for each image (or a plurality of temporally consecutive images), the reader understands that with this function, one can calculate for each input pixel luminance a corresponding output luminance (or luminance after linearization from any luminance encoding system specified by the corresponding EOTF luminance mapping). That is, for example, the input pixel comes from a 2000 nit PB C graded and encoded image (encoding typically involves at least a mapping to YCbCr), and the output pixel is an SDR image pixel (note that there will typically be metadata specifying which image the function refers to, which will typically be at least PB C HDR of the HDR image, i.e. 2000 nit for example, according to SL HDR). In fact, by conveying these two reference gradings, the receiving side will also understand that the image with PB C between 100 nit and PB C HDR is made by a process called display adaptation or display tuning, which involves specifying another luminance mapping function from the one received as from the content creator (see ETSI standard), but those details are not highly relevant to the present discussion, as the principles set forth below will work with the necessary modifications. That is, the metadata are the only information functionally required to encode a different dynamic range (typically SDR) image secondary to the primary one, as the receiver can compute its pixel colors by applying the function conveyed in common in the metadata to the pixel colors of the actual received YCbCr pixel color matrix image IM HDR.
[0067] The inventors realized that as a first step, one can compute in the PQ domain a luminance gain value (G PQ) quantifying how much the luminance value has dropped for a particular chrominance compared to the value on the achromatic axis (i.e. where the luminance of the current color can be uniquely determined by applying the PQ EOTF to this achromatic luminance value), because the colors should have the same luminance, and luminance- chrominance is readable on the achromatic axis, so the corresponding gray has chrominance set to zero, see Figure 2 b.
[0068] The inventors further experimentally realized that quite satisfactory results can be achieved by applying this same luminance gain value to the chrominance of the pixels. This can give some saturation errors in color, but this is far less noticeable than the original illumination errors, and is in fact acceptable for images with such high color saturation object issues.
[0069] Correlation means that if for any particular value of Y_PQ (which is usually normalized, falls between 0 and 1, with some precision step, corresponding to the number of color LUT entries or the number of luminance entries) one sees a scattered value of luminance gain value (G_PQ), the value of B in the color LUT for that location will be roughly the same as the location where the G_PQ value lies. Various algorithms can determine how the shape of the color LUT (i.e. the function of the variable value of B as a function of Y_PQ (B[Y_PQ])) will be roughly laid out in space with the scatter cloud of G_PQ values. For example, on the one hand, only a (geometric) subset of the pixels of the image, and their corresponding Y_PQ and G_PQ values (for example, there can be detectors that specifically detect the type of significant artifacts that occur in the image as with Figure 2 c) can contribute to the scatter plot, and only those pixels can contribute, and the other can be discarded, or on the encoding side, the selection of contributing pixels can also be determined by the human content creator. On the other hand, the entire color LUT shape does not necessarily have to be determined this way, for example, only the entire color LUT shape can be determined for a luminance subset, which is the lower half [0, 0.5], or where a lot of scattering occurs, and the other half can be determined by other algorithms, for example, is fixed. In the case where not all pixels are used, the subset of luminances that are changed can be related to the geometric subset of pixels that contribute. On the decoding side, since the determination of the color LUT can correspond to a deviation from the original co-transmitted color LUT, the correlation can also take into account this original color LUT, i.e. the final determined LUT to be applied to all input HDR image pixels to obtain the output image can deviate from the average position in the cloud of (Y, PQ, G_PQ) points, for example. The determination can also be limited to some subset of chrominance, in which case another well-working correction color LUT can be obtained.
[0070] Chrominance is sometimes also called color purity. The (correct) illumination corresponding to any luminance can not only be calculated, but can also be positioned on the luminance range, since both are normalized to 1.0.
[0071] In the case of an encoder, usually a well-working color LUT will be transmitted for the receiving side decoder to simply apply it. The content creator can check for example against one or more of his displays whether the correction is sufficiently working.
[0072] A useful embodiment of the high dynamic range video encoding circuit (300) has the color lookup table determination circuit (303) determining values (B) of a color lookup table based on a best fit function that generalizes a scatter plot of values of luminance gain values (G_PQ) versus corresponding values of high dynamic range image pixel luminance (Y_PQ). Multiple fits can be equally applied, for example, a function that goes roughly through the middle of the point cloud of (Y_PQ, G_PQ) points would do satisfactorily. Other functions can be designed according to the same innovative principle, for example, different functions for different tone regions in case the decoder has such a tone classification mechanism, etc.
[0073] Advantageously, the high dynamic range video encoding circuit (300) further comprises a perceptual slope gain determination circuit (301) configured to compute a perceptual slope gain (SG_PU) corresponding to a percentage increase of an output luminance (Y_o1) obtained when applying a luminance mapping function (400) to a luminance of a high dynamic range image pixel (Y1) in a coordinate system of perceptually uniformized luminances of horizontal high dynamic range image luminance (PY) and vertical output image luminance (PO) to reach a corrected output luminance (Y_corr1) of the high dynamic range image pixel (Y1), the corrected output luminance (Y_corr1) corresponding to a division of an output luminance (L_o1) obtained by applying the luminance mapping function (400) to an input luminance (L_e1) equal to a position of a normalized luminance corresponding to the high dynamic range image pixel (Y1) by the input luminance (L_e1) and multiplied by the high dynamic range image pixel (Y1). This is a way of linearizing the correction behavior of a non-linear luminance mapping function by giving at least dark luminance colors a slope that would have the same slope as would be obtained by drawing a line from the corresponding correct luminance input and output points to the (0,0) point.
[0074] The correction principle can also be embodied in a high dynamic range video decoding circuit (700),
[0075] The high dynamic range video decoding circuit is configured to color map a high dynamic range image (IM HDR) having a first maximum pixel luminance (PB CI) to a second image (Im LWRDR) having a lower dynamic range and a corresponding lower second maximum pixel luminance (PB C2), comprising a luminance mapping circuit (101) arranged to map luminances (Y PQ) of pixels of the high dynamic range image (IM HDR) with a luminance mapping function (400),
[0076] and a color mapping circuit comprising a look-up table (102) arranged to output a multiplier value (B) for each pixel luminance (Y_PQ) to a multiplier (121) multiplying a chrominance component (CbCr_PQ) of the pixel of the high dynamic range image (IM HDR) by the multiplier value (B), characterized in that the video decoding circuit comprises:
[0077] a gain determination circuit (302) configured to determine, for each pixel of a set of pixels of the high dynamic range image, a luminance gain value (G_PQ) quantifying a ratio of an output luminance for a luminance equal to a normalized luminance of a respective pixel's luminance (Y_PQ) divided by an output luminance for the respective pixel's luminance (Y_PQ); and
[0078] a color look-up table determination circuit (303) configured to determine a color look-up table of the look-up table (102) specifying a multiplier constant (B) as an output for various input luminance values, wherein values of at least one subset of the multiplier constants (B) for at least one subset of possible values of the luminances (Y_PQ) are determined based on values of luminance gains (G_PQ) for those luminances.
[0079] In case the encoder has not performed error mitigation in its jointly transmitted CL(Y_PQ) LUT for any reason, the same principle can be applied in the decoder. Moreover, the decoder will compute from the received YCrCb_PQ colors the corresponding luminances and luminance-luminances. It can then decode the received HDR image into any functionally jointly encoded secondary image of a different dynamic range, for example an SDR 100 nit PB_C_SDR image. The decoder will use the same new technology element to fill the appropriate luminance error mitigation CL(Y_PQ) look-up table in the corresponding LUT 102 of the color processing circuit of the SL HDR decoder.
[0080] A method of encoding a high dynamic range video is also useful, said method being configured to encode a high dynamic range image (IM HDR) having a first maximum pixel luminance (PB CI) with a second image (Im LWRDR) having a lower dynamic range and a corresponding lower second maximum pixel luminance (PB C2),
[0081] wherein the second image is functionally encoded as a decoder luminance mapping function (400) to be applied to pixel luminances (Y_PQ) of the high dynamic range image to obtain corresponding pixel luminances (PO) of the second image,
[0082] outputting said high dynamic range image and metadata (MET) encoding said luminance mapping function (400) to a video communication medium (399),
[0083] said functional encoding of said second image is also based on a color lookup table (CL(Y_PQ)) encoding a multiplier constant (B) for all possible values of luminance of pixels of said high dynamic range image, said multiplier constant (B) being used to be multiplied by chrominance (Cb, Cr) of said pixels, and said color lookup table being also outputted in said metadata,
[0084] characterized in that said high dynamic range video encoding comprises:
[0085] receiving pixels of said high dynamic range image (IM HDR) each having a luminance and two chrominances, and determining for each pixel a luminance gain value (G PQ) quantifying a ratio of a first output luminance for a normalized luminance equal to the luminance of each pixel divided by a second output luminance for the luminance of said pixel, wherein said first output luminance is obtained by applying said luminance mapping function to said normalized luminance and said second output luminance is obtained by applying a luminance mapping function to the luminance of the pixel;
[0086] wherein said high dynamic range video encoding comprises determining said color lookup table (CL(Y_PQ)) based on values of luminance gain values for various luminances of pixels present in said high dynamic range image, wherein values of said multiplier constant (B) for luminances from at least a subset of possible values of said luminance (Y PQ) are related to values of determined luminance gains (G PQ) for those luminances.
[0087] A method of high dynamic range video decoding is also useful, said method for color mapping a high dynamic range image (IM HDR) having a first maximum pixel luminance (PB CI) to a second image (Im LWRDR) having a lower dynamic range and a corresponding lower second maximum pixel luminance (PB C2), comprising: forming said second image by mapping luminances (Y PQ) of pixels of said high dynamic range image (IM HDR) with a luminance mapping function (400) to obtain luminances of said second image as output of said function,
[0088] wherein said color mapping further comprises obtaining chrominances of said second image by multiplying chrominance components (CbCr PQ) of pixels of said high dynamic range image (IM HDR) by a multiplier value (B),
[0089] The multiplier value (B) is specified in the look-up table (102) for various values of luminance (Y_PQ) of the pixels of the high dynamic range image (IM HDR),
[0090] characterized in that the video decoding comprises:
[0091] determining, for each pixel of a set of pixels of a high dynamic range image, a luminance gain value (G_PQ) quantifying a ratio of an output luminance equal to a normalized luminance of luminance (Y_PQ) of the respective pixel divided by an output luminance of luminance (Y_PQ) of the respective pixel; and
[0092] determining a color look-up table by computing values of at least one subset of multiplier constants (B) of possible values of luminance (Y_PQ) based on values of luminance gains (G_PQ) of those luminances.
[0093] In case the video content creator makes some problematic luminance mapping, the present novelty allows an improved re-grading (i.e. a down-mapping in particular for lower peak luminance) of the image while not creating problems for other cases. The computed G_PQ values can be implemented to a large extent as an additional saturation handling in the color path, by associating values of the B constant, defined as a Y_PQ dependent output of the color LUT, with values obtained from G_PQ computation. I.e. a single B value per Y_PQ roughly follows the position of dispersion of the G_PQ values, by some pre-designed metric co-located in the encoder or decoder. BRIEF DESCRIPTION OF DRAWINGS
[0094] These and other aspects of the method and apparatus according to the present application will become apparent from the following description and examples, and reference to the figures, which are presented solely for illustration of the specific embodiments and are not intended to limit the general concept in any way. Dashed lines or dots are used to indicate that a component is optional, non-dashed components are not necessarily required. Dashes or dots can also be used to indicate elements that are interpreted as necessary but hidden inside the object, or for intangible things, such as selection of objects / areas (and how they can be displayed on a display).
[0095] In the drawings:
[0096] Figure 1An applicant's HDR image decoder according to the so-called single-layer HDR image decoder (SLHDR) as standardized in ETSI TS 103 433-2 V1.1.1 is schematically illustrated; note that an encoder can typically have the same image processing circuit topology as the corresponding decoder, possibly with adjusted mapping function shapes (as decoding proceeds in the same degradation direction as how the encoder determines such degradation function to be applied by the decoder, so shape aspects like convexity can be the same for encoding and decoding, but depending on what output peak luminance the luminance is wanted to be mapped to, e.g. a commonly conveyed luminance mapping function of 1 / 7 power can be applied as 1 / 4 power in decoding).
[0097] Figure 2 Some problems that can occur for odd images with highly color-saturated objects in case of luminance usage in this SLHDR decoder (or corresponding encoder) are schematically illustrated;
[0098] Figure 3 A new way of SLHDR encoding according to the present inventive principle and a corresponding HDR video encoder 300 are schematically illustrated; in particular, it shows a part of the encoder topology for determining a color LUT in order to determine a constant B for multiplication with chroma similar to e.g. Figure 1
[0099] Figure 4 An embodiment of a perceptual slope gain determination circuit 301 according to the present inventive principle is schematically illustrated, and in particular the internal functioning of the same is illustrated;
[0100] Figure 5 An embodiment of a gain determination circuit 302 is schematically illustrated, and in particular how it works, i.e. sending the pixel color twice through at least some of all the illustrated luminance processing circuits, with the multiplier 401 having different multiplier constants;
[0101] Figure 6 A color lookup table determination circuit 303 can how it determines a function encoding the B multiplier for each Y_PQ value to be loaded in the LUT 102 from the various (Y_PQ, G_PQ) values computed by the circuit 302 for the pixels in the input HDR image is schematically illustrated (there are other ways of determining the LUT);
[0102] Figure 7 An HDR video decoder 700 is schematically illustrated, configured to apply the same inventive processing as the encoder 300 on the receiving side of the SLHDR encoded HDR video data; and
[0103] Figure 8 The apparatus in a typical deployment of this technology (i.e., real-time television broadcasting and reception) is shown (the same principle can also be used in other video communication technologies, such as non-real-time digital cinema production and rating, and later communication with cinemas, etc.). Detailed Implementation
[0104] Figure 3 This clarifies how a modified encoder conforming to our SLHDR HDR video encoding method works. On the encoding side, the input is a usable raw master HDR image (e.g., a 5000nit PB_C image graded by a human color grader, making all objects in various video images (e.g., bright explosions and dark shadow areas in caves) appear optimal) and thus has both usable pixel illuminance (L) and corresponding luminance (Y). A luminance mapping function determination circuit 350 is included (or connected), which can determine, for example... Figure 4 The best example shown and discussed is the S-shaped brightness mapping curve (i.e., Figure 1 (Circuit 113 in the diagram). Note that although we have illustrated the principle with this S-shaped mapping function, the processing is general regardless of the function (it can even work without improving or worsening the situation). We do not need to elaborate on the many ways this can be accomplished through various embodiments developed by the applicant, but for example, this can be done by an automaton that analyzes the luminance present in the master HDR input image (which will eventually be transmitted to the receiver as IM_HDR) and determines the optimal function (e.g., an S-curve) for this luminance distribution, or, in other physical embodiments of the encoding device, it can be configured by a human color grader via a user interface (such as a coloring console). The three arrows in the luminance mapping function determination circuit 350 pass the determined luminance mapping function 400 to the corresponding other circuits, for example, it will be used by the perceptual slope gain determination circuit 301, a typical operating embodiment of which is... Figure 4The teaching. Note that while the typical embodiment can work in the Philips perceptual domain defined by equation 3, this is not an essential element of our innovation, in some embodiments this can be expressed directly with PQ gains, hence the circuit 301 is drawn as a dashed line, indicating optionality. Eventually, the data formatter 304 will transmit to the receiver all data encoding the images at at least two different gradings (one of them as actual image of pixel colors IM HDR, e.g. DCT compressed in MPEG or similar video encoding standard, e.g. AV1), based on the master HDR image. That is, it will output to any video communication medium 399 (which can be e.g. a wired connection, such as e.g. a serial digital interface (SDI) cable or an HDMI cable, or an internet cable, a wireless data connection, such as e.g. a terrestrial broadcast channel, a physical medium for transmitting video data, such as a Blu-ray disc, etc.) an encoded and typically compressed version of the master HDR image, i.e. the high dynamic range image IM HDR, at least one luminance mapping function for the receiver to apply (when re-grading to a different, typically lower peak luminance), and the color lookup table (CL_(Y_PQ)), the latter two as metadata (MET), depending on the video communication protocol used, e.g. SEI messages. The color lookup table is used to provide for each lookup location Y_PQ (of a pixel) a value B to multiply with the two chroma values Cb and Cr of any pixel.
[0105] Figure 4 A mapping applied to the HDR input luminance is shown (as represented in the perceptually uniformized domain, where the SLHDR luminance mapping is specified), to obtain in this case the SDR output luminance. The perceptually uniformized luminance is normalized, i.e. values 1.0, on the vertical axis, corresponding to the SDR peak luminance PB_C2 = 100 nit, and on the horizontal axis, in this example, 1.0 corresponds to the master HDR image peak luminance PB_C1 of 1000 nit of the high dynamic range image (IM HDR). This is a useful way to describe a luminance mapping, e.g. it gives the human color grader better control than when specifying the mapping natively in the luminance domain (the skilled person understands that only the shape of any luminance mapping function will change in a predictable way if the axes are quantized differently). In this illustrative example, we skipped the coarse luminance mapping of the coarse dynamic range converter 112 (and other processing allowed by our codec, such as black / white stretching and gain limiting as shown in Figure 5 in the middle). The luminance mapping consists entirely of this S-curve (e.g. this curve can occur if there are quite bright scenes, where one wants to do contrast stretching in the darker parts, which makes the output SDR image look better).
[0106] Let us see the color of a particular pixel, for example saturated magenta, which has a luminance Y1. If this luminance is mapped through the luminance mapping curve 400, then an output luminance Y_01 is produced (similar to another pixel with input luminance Y2). This can be a relatively dark output result (as one should remember, the axis of perceptual uniformization is somewhat logarithmic in nature). If one wants to map "illuminance" instead of luminance, or technically more precisely, map the luminance (or position on the horizontal axis) that uniquely corresponds to the illuminance L_e1 on the encoder side (which can be the exact illuminance-luminance [which is the luminance that uniquely corresponds to the illuminance via the OETF] because the encoder has all the information available), then it would actually be better to obtain an output luminance L_01, i.e. brighter than Y_01. That is, if there is no strong illuminance leakage in the luminance representation (e.g. PQ or Philips perceptual uniformization luminance), then that L_e1 is the luminance axis value that the full-pixel of that color of that image object should have instead of its actual luminance Y1.
[0107] In fact, for achromatic grays, since there is no chroma-dependent luminance loss on that axis, and the EOTF (luminance) is equal to the illuminance, therefore uniquely and exactly, one would indeed map according to that illuminance-luminance L_e1 (i.e. one would follow the preferred re-grading of neutral / achromatic and near-neutral colors in the image, such as to make some good deep blacks). Moreover, as Figure 2 D, there can be quite a large expansion of luminance values (Y1 vs Y2) for initially similar colors, because they initially have almost the same illuminance, and then additionally for such a highly non-linear S curve, one of the luminances can jump to a highly different mapping part of the function, e.g. a small-slope linear part vs a high-slope middle part that maps the darkest input luminances.
[0108] Thus, on one hand, one can see that there is a problem of too dark output colors in highly saturated objects, so some brightening should be applied (at least for certain image colors). Moreover, the inventors found by experiment that for a linear curve, there is no significant visual disturbance (even if there is some error, it is not visually important one).
[0109] But one cannot simply change that luminance mapping function 400, because this is the key metadata that determines how the secondary image needs to be re-graded (according to the content creator) from the received image (IM HDR) to e.g. a 100 nit PB_C_SDR SDR image. As said, the luminance mapping function is perfectly correct for achromatic and near-neutral colors, so if one wants to change it, those colors will suddenly be incorrectly luminance re-graded, and those are often the more critical colors, e.g. face colors. On the other hand, in many cases, the function also works sufficiently well on non-achromatic colors as needed.
[0110] Returning to Figure 3 , the perceived slope gain determination circuit 301 simply determines a "case characterizer value", namely the perceived slope gain SG_PU, which is determined as:
[0111] SG_PU = (L_ol / L_el) / (Y_ol / Yl) [Equation 4]
[0112] i.e. it is the ratio of two slopes: first, the output luminance divided by the luminance - luminance L_el as obtained by applying the luminance mapping curve to the (substantially or exactly) correct luminance representing the luminance value L_el; and second divided by the ratio of the output luminance of that particular color pixel Yl (of the input image) divided by that input luminance Yl. Substantially or exactly, the correct luminance position refers to the fact that in some embodiments the luminance can be determined exactly, and in other embodiments it can be desired to estimate it (which can be done in various ways), but we will continue to "normalize luminance" (and represent in the particular Philips perceived uniformization range), and the reader can assume it is the exact correct luminance value (e.g. one of the video creator's grades, and can be calculated e.g. by applying Equation 1 in a computer to the linear RGB coefficients of the image pixels).
[0113] In Figure 5 , we show an embodiment of the gain determination circuit 302 in which this perceived slope gain SG_PU is used. What is ultimately needed is a correction in the PQ domain (of US9077994 and standardized as SMPTE ST.2084) as can be seen in Figure 1 , while all luminance mapping of at least the SLHDR decoder takes place in the perceived uniform (PU) domain (of Equation 3 above). The inventor realized that he can solve this problem quite adequately by adjusting the chroma (and in the encoder determining the color lookup table CL(Y_PQ) to be communicated to the receiving side decoder). I.e. when the grader has defined an initial color lookup table, this can be the altered color lookup table. I.e. everything in the luminance mapping sub-circuit 101 takes place in the PU domain, but the inputs Y_PQ and CbCr_PQ are in the PQ domain.
[0114] Thus, in practice a correction has to be determined for the errors of the saturated image objects in the PQ domain, which are blobs and other too dark parts, according to the present method, which is performed by the circuit 302. In an exemplary embodiment, it runs two passes (for each pixel color of the image), controlled by a multiplier with switched multiplication values. Note that the implementation details, e.g. buffering, providing image delays, etc. can be understood by the skilled person himself.
[0115] First, the value is set to 1 and then the input luminance L (which can be computed from the input luminance Y_PQ by applying the PQ EOTF to the input luminance Y_PQ) is converted again to a perceptually uniform luminance domain, i.e. to an equivalent, representative perceptual luminance value PY. Thus, the total mapping behavior will be applied to the pixel(s) of the process, which in Figure 2 d is shown as PY value.
[0116] In fact, the block of this unit corresponds to the Figure 1 and also Figure 4 explained with respect to the ETSI TS 103 433-2 V1.1.1 standard. The newly drawn block is a black-white level shifter 402, which can shift a certain constant value of the input HDR luminance to the 0 value and the 1 value of the perceptually uniformized representation, respectively. It is generally advantageous, for example, if the HDR is not darker than 0.1 nit (or any luminance representation value equivalent thereto), to have pre-mapped this value to the lowest SDR value zero and the same as the highest value before applying further luminance mappings. The gain limiter 403 is a circuit which is activated in certain cases and determines the maximum value between the maximum value resulting from the application of the various processing steps and alternative luminance mapping strategies in the upper track (i.e. circuits 402, 112, 113, etc.), so that the output luminance in the image resulting from the application of this luminance mapping to all pixels of the input image does not become too low (for details, see the ETSI standard; all that needs to be known for the present invention is that all the processing that would be applied in the normal case of the classic SLHDR luminance mapping without the present invention is applied, i.e. if a rough mapping is applied, the circuit 302 also applies a rough mapping, and if no rough mapping is applied, the circuit 302 also does not apply a rough mapping in exactly the same way).
[0117] Thus, in the first pass, when the multiplier is set to 1, the normal output luminance Y_oi occurs for the pixel luminance input Yi (or similarly for any other input luminance's corresponding output luminance such as Y2). Note that this Y_oi value is still in the perceptually uniformized luminance domain (PO), so it must still be linearized (by the linearization circuit 114) and then (by the circuit 115) PQ domain converted.
[0118] In the second pass, the multiplier multiplies the perceptual luminance PY by the perceptual slope gain value SG_PU (as determined by the circuit 301) and then derives the corrected value as shown in Y_corr1 in Figure 4 This corresponds to applying our (true) normalization of the luminance calculation, finding what slope is present in the total mapping curve for this position (L_e1, L_oi) and using it to locally increase the slope and output Yi.
[0119] In fact, the gain determination circuit (302) computes the PQ domain luminance gain value (G_PQ) as:
[0120] G_PQ = OETF_PQ[Inv_PU(Y_corr1)] / OETF_PQ[Inv_PU(Y_o1)] [Equation 5]
[0121] where Inv_PU is the inverse of the perceptual uniformization of Equation 3, i.e. the computation of the corresponding (linear) normalized luminance.
[0122] i.e. G_PQ is the ratio of the output luminance (in the PQ domain) obtained when sending the uncorrected (i.e. as received) luminance values through our whole SLHDR luminance processing chain, to the corrected PQ domain luminance, which is the luminance-luminance, i.e. EOTF(Y1), through the luminance processing chain. Note that at the encoding side, there can be several alternatives to build the encoder, e.g. the G_PQ value can be computed directly from the actual chroma-dependent luminance (i.e. chroma-dependent luminance). Because it will be encoded in the output image IM HDR of the encoder, and the correct (achromatic) luminance representation is the luminance-luminance, i.e. the circuit 301 is not a necessary core technical element of our innovation, and can not exist in some encoder embodiments (see the dashed line in Figure 3 which shows that in this case, the circuit 302 would do its two computations with luminance Y and luminance-luminance L as inputs respectively, and without multiplier or it set to 1x in both cases).
[0123] This G_PQ value will now be used by the lookup table determination circuit (303) in its determination of the corresponding color LUTCL(Y_PQ). There can be several ways to do this (depending on what one wants to achieve: there is only one shape of LUT function that can be determined for the case, but not focusing on e.g. the overall behavior, but one can focus in particular on certain aspects, like specific colors in the image, etc.), which can be illustrated by one prototype example in Figure 6 In fact, ideally one would like to correct specifically each pixel, but this is not how the SLHDR method works (due to the simplification by the 1D-2D processing split). The idea is to want to increase the saturation of the pixels where the luminance processing artifacts occur (at least). Then, the lower luminance corresponding to this higher chroma will be reconverted into a higher luminance in the output R" G" B" colors. The drawback is that there is some saturation change between the input and output colors, but this is acceptable, also because the eye is more critical on the luminance than on the saturation change, and also because the luminance readjustment is a more important visual property in the dynamic range adjustment. In any case, the regularization (although it can tweak the colors somewhat "globally") removes or mitigates the annoying local artifacts likeFigure 2 the spots in the flower of c, and this is why the present computational arrangement is performed.
[0124] Figure 6 It is shown how we (on average) determine the correction strategy by determining a fitting function to the data points computed by the circuit 302 (i.e. for any pixel in at least one specific image of the video, having Y_PQ, but also Cb and Cr values, the G_PQ value occurs). I.e. the circuit 302 obtains all (or some; e.g. in case only certain affected pixels are detected, and the algorithm, i.e. the color LUT determination, is facilitated) pixels of e.g. an image, and computes the G_PQ values by the above described technical procedure. When those are organized in a 2D plot, it can be seen that for any single value, several needed (for the respective pixel) G_PQ values can be generated (especially for lower Y_PQ values). Figure 6 The plot of G_PQ vs. Y_PQ shows the various (Y_PQ, G_PQ) values coming out of the circuit 302. One generally sees no spread for high luminance Y_PQ (there is the multiplier constant 1 equal to no correction). In the lower region, which usually corresponds to saturated colors (or dark unsaturated colors), there is a spread of G_PQ values per Y_PQ position, since a pixel with a specific Y_PQ can have various Cb, Cr values. From the plot of Figure 1 The color LUT, CL(Y_PQ), also has a behavior which can be shown in this plot: it determines a boost value per Y_PQ value, so if B is equal to G_PQ as computed above for the luminance handling track (but at the same time computable unchangeable in the luminance track), a function can be drawn which corresponds to the values filled in the 1D color LUT CL(Y_PQ). The skilled person is aware that there are various ways on how to fit a function to a cluster of points. E.g. a minimization of the mean square error can be used.
[0125] If this functionality is applied at the encoder side, basically only the color LUT (as shown in Figure 3 ) has to be determined, then this color LUT is transmitted to the receiving side, and then the decoder can simply apply it. This is an elegant way of embodying the present principle.
[0126] However, the same procedure can also be applied at the decoding side (when not applied at the encoding side), but it will have some small differences. However, what will remain the same is that whenever an illumination re-scaling needs to occur, the multiplier 121 will de- correct the chroma components for the luminance issue by computing the following:
[0127] Cb_COR = CL(Y_PQ) * Cb_PQ and Cr_COR = CL(Y_PQ) * Cr_PQ [Equation 6]
[0128] Figure 7 A decoder 700 according to the application is shown. Basically, almost all blocks are the same as in the encoder 100 of the application. Figure 1 In particular, the gain determination circuit 302 has to determine again a set of luminance gain values (G_pq) for at least the artifact corrupted pixels, which again quantifies the ratio of the output image luminance divided by the chroma dependent luminance of the output luminance of the pixels of the high dynamic range image, which is found on the vertical axis of the achromatic gray color (i.e. the luminance position which is equal to the (correct) normalized luminance position, i.e. as Figure 2 b shows, the non-reduced luminance position found on the vertical axis of the achromatic gray color). And now the high dynamic range video decoding circuit comprises a similar color lookup table determination circuit (303) which uses these G_PQ values to determine the color LUTs (Y_PQ) to be loaded in the pixel color processor pipeline of the decoder (i.e. in the LUT 102), but now all determined at the decoder side. There can still be detailed embodiments which do not determine and load color LUTs, but change the transmitted color LUTs, wherein they balance the behavior of the original saturation processing LUT (in circuit 303 or equivalent separate circuit) and the required correction according to the application, but those are details. For example, both LUTs can be interpolated, and the weight factor can depend on a measure of the severity of the artifacts (e.g. the severity of the artifacts). How much darker the average pixel is compared to its surroundings, or a measure based on accumulated edge strength or texture measure, etc.), and then for example for larger errors, the circuit 303 can instruct that the LUT determined according to the method obtains for example 80% weight, i.e. the other LUT can be pulled towards it with 20% strength, etc. It is very typical for the technology that the solution is then not 100% perfect, but the problem is clearly mitigated. Figure 6 at the top of the graph of the application, but only pulled towards it with 20% strength, etc. It is very typical for the technology that the solution is then not 100% perfect, but the problem is clearly mitigated.
[0129] What is new in this decoder is that the decoder in principle does not have the luminance L (which is typically easily available at the encoding side), but only Y_PQ.
[0130] There can be several ways in which L can be determined, which is performed by the luminance calculator 701, but the simplest one is to just do the matrixing from YCbCr to non-linear R’G’B’_PQ, then linearize via EOTF_PQ, and then calculate the luminance via the weighting of the R, G and B components as explained with formula 1. Note that the luminance calculator 701 typically calculates the luminance-luminance, which enters the circuit 302 in the second calculation through the whole luminance mapping chain, as explained with respect to Figure 5 .
[0131] Note that the decoder can also include a metadata checker configured to check an indicator in the metadata that indicates whether the encoder has applied the necessary corrections in its communicated color lookup table. Typically, the decoder will then not apply any corrections to said lookup table, but some embodiments can still do their own calculations to verify and / or fine-tune the received color lookup table. Another mechanism for creating / encoding a protocol between the encoding side and the consumption side can also be employed, e.g. a prefix case for a certain HDR video supply path, or something configured via a software update or user control, etc.
[0132] Figure 8 A typical (non-limiting) example of a video communication ecosystem is shown. In terms of video creation, we see a live studio where a presenter 803 is for example demonstrating a science program. He can walk through various parts of the studio, which can be very creatively (freely) lit. For example, there can be a bright area 801 of the demonstration area, which is lit by many studio lights 802. There can also be a darker area 804. The video image is captured by one or typically multiple cameras (805, 806). Ideally, these cameras are of the same type and are color coordinated, but for example one of the cameras can be a drone, etc. There is a final responsible person 820 who decides on the final product, and in particular its colorimetry (as the skilled person knows, depending on the kind of production, several systems and even people can be involved, but for this explanation we will call this person 820 the color grader). He can have means like a panel for switching, directing, coloring, etc. He can watch one or more live products on an HDR reference monitor and / or an SDR monitor, etc. For example, the display can also include an auxiliary video 810 for teaching, like an underwater scene. This can for example be an SDR video or an HDR video of a different illumination dynamic range characteristic than the live camera(s) feed. Looking at the highest level of discussion, there can be two cases: the secondary video 810 can actually be displayed at a specific light level, let us assume quite bright, or it can be a green screen and the video only exists in the production room of the final responsible person 820.
[0133] The discussion is interested in determining the luminance mapping in the production studio. For example, one can optimize the mapping function before the live action starts, which works well or quite well for all lighting areas after setting the HDR camera to a good iris level. But one can also determine several functions for the various areas, e.g. camera zoom on dark areas, and switch functions when calculating the final HDR image output in the video encoder 822, or use these functions in a per-image automatic machine-optimized version of the function without or with manual intervention, etc. Finally, in this example we have a satellite link via satellite dish 830 to a consumer or professional intermediate station, but the encoded HDR video can also be output as a video communication technology over the Internet. On the receiving side, we typically have the consumer's home, e.g. his living room 850. We illustrate that the consumer gets the broadcasted HDR video via his local satellite antenna 851 and a satellite TV set-top box 852, and the image is finally viewed on an HDR TV 853 or other display. The HDR video decoder can be included in the set-top box or the display. There can be a further image illumination optimization of the as-received HDR image (with its encoded peak luminance PB_C) to the maximum illumination PB_D of the consumer's TV. This can happen in the set-top box, where the optimized image goes over e.g. an HDMI cable or a wireless video link etc. to the TV, or the set-top box can simply be a data passer, and the TV can include any of the decoder embodiments as described above.
[0134] Technical means of any whole or part of the innovations taught above can be implemented in practice (completely or partially) as hardware (e.g. part of a special IC), or software running on a special digital signal processor or a general purpose processor, FPGA, etc. Any processor, part of a processor, or complex of connected processors can have internal or external data buses, on-board or on-board memories, such as caches, RAM, ROM, etc. A device comprising a processor can have special protocols for special connections to hardware, such as image communication cable protocols, e.g. HDMI for connecting to a display, or internal connections for connecting to a display panel in case the device is a display. A circuit can be configured by dynamic instructions before it performs its technical actions. Any element or device can form part of a larger technical system, such as a video creation system at any content creation or distribution site, etc.
[0135] The skilled person will understand which technical means can be used or can be used for conveying (or storing) the images, whether to another part of the world, or between two adjacent devices, such as a HDR suitable video cable, etc. The skilled person will understand in which cases which form of video or image compression can be used. The skilled person can understand that signals can be mixed, and that it is not absolutely necessary to first mix before applying some calculations on the images (e.g. determination of the optimal S-curve), but that this can be advantageous, although even when applying this determination on pre-mixed images, the determination can still weight various image aspects of e.g. the camera feed image and the secondary image in various specific ways. The skilled person understands that various implementations can work in parallel, e.g. by using several video encoders to output several HDR videos or video streams.
[0136] From our introduction, the skilled person should be able to understand which means can be optional improvements and can be implemented in combination with other means, and how the (optional) steps of the method correspond to the corresponding units of the apparatus, and vice versa. The word "apparatus" in this application is used in its broadest sense, i.e. a set of technical elements allowing to achieve a certain goal, and thus, for example, can be an IC (small circuit part), or a dedicated household appliance (e.g. a household appliance with a display), or a part of a network system, etc. "Arrangement" or "system" are also intended to be used in the broadest sense, thus, it can in particular comprise a single apparatus, a part of an apparatus, a set of (partly) cooperating apparatuses, etc.
[0137] Some of the steps needed to operate the method can already be present in the functionality of the processor instead of being described in the computer program. Similarly, some of the aspects with which the application cooperates can be present in well-known technical circuitry or elements or separate apparatuses, e.g. the functionality of a display panel that will display a corresponding display color in front of a screen when driven by some color-coded digital values, and these existing details will not be discussed exhaustively in order to make the teaching clearer by focusing on the exact contribution made to the technical field.
[0138] It should be noted that the above-mentioned embodiments illustrate rather than limit the application. Wherever possible, the same reference numerals and characters are used in the attached drawings and the following description to
[0139] Any reference signs in the claims should not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements.
Claims
1. A high dynamic range video encoding circuit (300) configured to encode a high dynamic range image (IM HDR) having a first maximum pixel luminance (PB CI) with a second image (Im LWRDR) having a lower dynamic range and a corresponding lower second maximum pixel luminance (PB C2), wherein said second image being functionally encoded as a luminance mapping function (400) for a decoder to be applied to pixel luminances (Y PQ) of said high dynamic range image to obtain corresponding pixel luminances (PO) of said second image, said encoding circuit comprising a data formatter (304) configured to output said high dynamic range image and metadata (MET) encoding said luminance mapping function (400) to a video communication medium (399), said functional encoding of said second image being further based on a color lookup table (CL(Y PQ)) encoding a respective multiplier constant (B) for each of all possible values of pixel luminances of said high dynamic range image, said multiplier constant being for multiplication with chroma (Cb, Cr) of said pixels, and said formatter being configured to output this color lookup table in said metadata, characterized in that said high dynamic range video encoding circuit comprises: a gain determination circuit (302) receiving pixels of said high dynamic range image (IM HDR), each pixel having a luminance and two chroma, said gain determination circuit being configured to determine, for each pixel, a luminance gain value (G PQ) equal to a ratio of a first output luminance for a normalized luminance equal to the luminance of each pixel divided by a second output luminance for the luminance of said pixel, wherein said first output luminance is obtained by applying said luminance mapping function to said normalized luminance and said second output luminance is obtained by applying said luminance mapping function to the luminance of said pixel; wherein said high dynamic range video encoding circuit comprises a color lookup table determination circuit (303) configured to determine said color lookup table (CL(Y PQ)) based on values of said luminance gain values for various luminances of pixels present in said high dynamic range image, wherein values of said multiplier constant (B) for luminances from at least a subset of possible values of said luminances (Y PQ) are related to values of determined luminance gains (G PQ) for those luminances.
2. The high dynamic range video encoding circuit (300) of claim 1, wherein, said color lookup table determination circuit (303) being arranged to determine values (B) of the color lookup table based on a best fit function that generalizes a scatter plot of values of said luminance gain values (G PQ) versus corresponding values of said high dynamic range image pixel luminances (Y PQ).
3. The high dynamic range video encoding circuit (300) of claim 1 or 2, further comprising a perceptual slope gain determination circuit (301) configured to compute a perceptual slope gain (SG_PU) corresponding, in a coordinate system of perceptually uniformized luminances of a horizontal high dynamic range image luminance (PY) and of a vertical output image luminance (PO), to a percentage increase of an output luminance (Y_o1) obtained when applying the luminance mapping function (400) to a luminance of a high dynamic range image pixel (Y1) to reach a corrected output luminance (Y_corr1) of the high dynamic range image pixel (Y1), the corrected output luminance corresponding to a division of an output luminance (L_o1) obtained by applying the luminance mapping function (400) to an input luminance (L_e1) by the input luminance (L_e1) and by the high dynamic range image pixel (Y1), the input luminance being equal to a position of a normalized luminance corresponding to the high dynamic range image pixel (Y1).
4. A high dynamic range video decoding circuit (700) configured to color map a high dynamic range image (IM HDR) having a first maximum pixel luminance (PB CI) to a second image (Im LWRDR) having a lower dynamic range and a corresponding lower second maximum pixel luminance (PB C2), comprising: a luminance mapping circuit (101) arranged to map luminances (Y_PQ) of pixels of the high dynamic range image (IM HDR) with a luminance mapping function (400), and a color mapping circuit comprising a lookup table (102) arranged to output, to a multiplier (121), a respective multiplier value (B) for each of all possible pixel luminances (Y_PQ), the multiplier multiplying chroma components (CbCr_PQ) of the pixels of the high dynamic range image (IM HDR) by the respective multiplier value (B), characterized in that the video decoding circuit comprises: a gain determination circuit (302) configured to determine, for each pixel of a set of pixels of the high dynamic range image, a luminance gain value (G_PQ) quantifying a ratio of an output luminance for a luminance equal to a normalized luminance of a luminance (Y_PQ) of the respective pixel divided by an output luminance for the luminance (Y_PQ) of the respective pixel; and a color lookup table determination circuit (303) configured to determine a color lookup table of the lookup table (102) specifying multiplier constants (B) as outputs for various input luminance values, wherein values of at least one subset of the multiplier constants (B) for at least one subset of possible values of the luminances (Y_PQ) are determined in relation to values of luminance gains (G_PQ) for those luminances.
5. A method of encoding a high dynamic range video, configured to encode a high dynamic range image (IM HDR) having a first maximum pixel luminance (PB_C1) together with a second image (Im LWRDR) having a lower dynamic range and a corresponding lower second maximum pixel luminance (PB_C2), wherein the second image is functionally encoded as a luminance mapping function (400) for a decoder to apply to pixel luminances (Y_PQ) of the high dynamic range image to obtain corresponding pixel luminances (PO) of the second image, the method comprises outputting the high dynamic range image and metadata (MET) encoding the luminance mapping function (400) to a video communication medium (399), the functional encoding of the second image is also based on a color lookup table (CL(Y_PQ)) encoding a respective multiplier constant (B) for each of all possible values of pixel luminances of the high dynamic range image, the multiplier constant being for multiplication with a chroma (Cb, Cr) of the pixel, and the color lookup table being also output in the metadata, characterized in that encoding the high dynamic range video comprises: receiving pixels of the high dynamic range image (IM HDR), each pixel having a luminance and two chromas, and determining, for each pixel, a luminance gain value (G_PQ) quantifying a ratio of a first output luminance for a luminance equal to a normalized illuminance of the luminance of each pixel divided by a second output luminance for the luminance of the pixel, wherein the first output luminance is obtained by applying the luminance mapping function to the normalized illuminance and the second output luminance is obtained by applying the luminance mapping function to the luminance of the pixel; wherein the encoding of the high dynamic range video comprises determining a color lookup table (CL(Y_PQ)) based on values of luminance gain values for various luminances of pixels present in the high dynamic range image, wherein values of a multiplier constant (B) for luminances from at least a subset of possible values of the luminances (Y_PQ) are related to values of the determined luminance gains (G_PQ) for those luminances.
6. A method of high dynamic range video decoding that color maps a high dynamic range image (IM HDR) having a first maximum pixel luminance (PB CI) to a second image (Im LWRDR) having a lower dynamic range and a corresponding lower second maximum pixel luminance (PB C2), comprising: forming the second image by mapping luminances (Y_PQ) of pixels of the high dynamic range image (IM HDR) with a luminance mapping function (400) to obtain luminances of the second image as an output of the function, wherein the color mapping further comprises obtaining chromas of the second image by multiplying chroma components (CbCr_PQ) of pixels of the high dynamic range image (IM HDR) by respective multiplier values (B), the multiplier values (B) being specified in a color lookup table (102) containing respective multiplier values for various values of luminances (Y_PQ) of pixels of the high dynamic range image (IM HDR), characterized in that the video decoding comprises: determining, for each pixel of a set of pixels of the high dynamic range image, a luminance gain value (G_PQ) quantifying a ratio of an output luminance for a luminance equal to a normalized illuminance of the luminance (Y_PQ) of the respective pixel divided by an output luminance for the luminance (Y_PQ) of the respective pixel; and The color lookup table is determined by calculating values of at least one subset of multiplier constants (B) for at least one subset of possible values of luminance (Y_PQ), the calculation being made by calculating a respective value of a multiplier value (B) for a luminance gain (G_PQ) associated with any luminance.
Citation Information
Patent Citations
Simple but versatile dynamic range coding
US20180005356A1
Device and method of improving the perceptual luminance nonlinearity-based image data exchange across different display capabilities
US9077994B2
Encoding and decoding HDR videos
WO2017157977A1
Graphics processing for high dynamic range video
CN103597812A
Improved HDR image encoding and decoding methods and devices
CN104471939A