Image brightness adjustment for perceiving uniform electro-optical transfer functions

By using a uniformly perceived electro-optical transfer function (EOTF) and logarithmic mapping in the image processing device, the problem of inconsistent display of HDR images on different displays is solved, and a higher quality and consistent display effect is achieved.

CN120359753APending Publication Date: 2025-07-22KONINKLIJKE PHILIPS NV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380084115.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-17
Filing Date
2023-12-05
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art, when processing high dynamic range (HDR) images, is difficult to effectively achieve optimal display on different displays, especially on low dynamic range (SDR) displays, resulting in image quality degradation and inconsistent display.

Method used

Using an image processing device, the pixel brightness encoded by a uniform electro-optical transfer function (EOTF) is received through the receiver, converted into a logarithmic-mapping EOTF-encoded second pixel brightness, and applied brightness adjustment processing, including first and second logarithmic functions and multiplication scaling, to generate the final second pixel brightness.

Benefits of technology

Improves the display quality and consistency of images on displays of different dynamic ranges, reduces processing complexity and resource usage, and achieves more accurate brightness adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359753A_ABST
    Figure CN120359753A_ABST
Patent Text Reader

Abstract

An image processing apparatus includes a receiver (401) that receives an image including a first pixel brightness encoded according to a perceptually uniform electro-optical transfer function (EOTF). A converter converts this to a second pixel brightness encoded according to an EOTF representing a logarithmic mapping from an optical light value to the second pixel brightness, and an image processor circuit (411) applies a brightness adjustment process to the second pixel brightness. The converter (410) includes a first converter (601, 605) that generates an intermediate pixel brightness that applies a logarithmic function to the first pixel brightness. A second converter (603) generates an intermediate pixel brightness by applying a second logarithmic function to an output value of the perceptually uniform EOTF for the first pixel brightness divided by a divisor equal to an exponent of the first pixel brightness, the exponent having a value greater than 1. A combiner generates the second pixel brightness by combining the intermediate pixel brightness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing, and more particularly to brightness adjustment processing, such as, but not exclusively, tone mapping for high dynamic range (HDR) images. Background Art

[0002] A few years ago, novel techniques for high dynamic range (HDR) video coding were introduced, including those of the applicant (see, for example, WO2017157977).

[0003] The encoding of video generally mainly or only involves creating or more precisely defining color codes (e.g., luminance information and two chromaticities for each pixel) to represent an image. This is different from knowing how to optimally display an HDR image (e.g., the simplest method simply uses a highly non-linear opto-electronic transfer function OETF to convert the desired luminance into, for example, a 10-bit luminance information code, and vice versa, and can map a 10-bit electrical luminance information code to the optical pixel luminance to be displayed by using an inverse-shaped electro-optical transfer function EOTF, converting those video pixel luminance information codes into the luminance to be displayed, but more complex systems may deviate in several directions, especially by decoupling the encoding of the image from the specific use of the encoded image).

[0004] The encoding and processing of HDR video is largely in contrast to how traditional video techniques are used. According to traditional video techniques, until recently all videos were encoded, which is now referred to as standard dynamic range (SDR) video encoding (also known as low dynamic range video encoding; LDR). This SDR started as PAL or NTSC in the analog era and evolved to Rec.709-based encoding, such as MPEG2 compressed in the digital video era.

[0005] Although a satisfactory technology for transmitting moving pictures in the 20th century, the advancement of display technologies beyond the physical limitations of the 20th century CRT's electron beam or the global TL backlit LCD has made it possible to display images with pixels that are significantly brighter (and possibly also darker) than those of traditional displays, which urgently requires the ability to encode and create such HDR images.

[0006] In fact, starting from brighter and possibly also darker image objects that cannot be encoded using the SDR standard (8-bit Rec.709), for various reasons, ways to technically represent those increased luminance range colors were first invented, and thus all the rules of video technology were revisited one by one and generally had to be reinvented.

[0007] The luminance information code definition of the Rec.709 SDR can only encode (using 8 or 10-bit luminance information) a luminance dynamic range of approximately 1000:1. Due to its approximate square root OETF function shape, the luminance information: Y_code = power(2,N)*sqrt(L_norm), where N is the number of bits of the luminance information channel, and L_norm is the version of the physical luminance normalized between 0 and 1.

[0008] In addition, in the SDR era, there was no defined absolute luminance to be displayed. Therefore, in practice, the maximum relative luminance L_norm_max = 100% or 1 was mapped via the square root OETF to the maximum normalized luminance information code Yn = 1 corresponding to, for example, Y_code_max = 255. Compared to creating an absolute HDR image, this has several technical differences, i.e., where an image pixel encoded to be displayed as 200 nits ideally (i.e., when possible) is displayed as 200 nits on all displays rather than being displayed as a completely different display luminance. In the relative paradigm, a 200 nit encoded pixel luminance can be displayed as 300 nits on a brighter display (i.e., a display with a brighter maximum displayable luminance PL_D (also known as the display peak luminance)) and, for example, as 100 nits on a display with less capability. Note that absolute encoding can also act on the normalized luminance representation or the normalized 3D color gamut, but 1.0 means, for example, uniquely 1000 nits.

[0009] At the display, such relative images are typically displayed somewhat heuristically by mapping the brightest luminance of the video to the brightest displayable pixel luminance (which occurs automatically without further luminance mapping via electrically driving the display panel with the maximum luminance information Y_code_max). So if you purchase a 200 nit PL_D display, your white will look 2 times brighter than on a 100 nit PL_D display, but given factors such as eye adaptation that are considered not too important, it gives a brighter, better viewable, and slightly more beautiful version of the same SDR video image.

[0010] Conventionally, if one now talks about an SDR video IMAGE in the (absolute framework), it typically has a video peak luminance of PL_V = 100 nits (e.g., agreed upon according to a standard). Therefore, in this application, we consider the maximum luminance (or SDR grading) of an SDR image to be exactly this value or around this value to be prevalent.

[0011] The grading in this application is intended to represent an activity or the resulting image, where pixels have been given the desired brightness, e.g., by a human color grader or an automaton. If one views an image (e.g., a design image), there will be several image objects, and it may be desirable to give the pixels of those objects a brightness distribution around an average brightness that is optimal for that object, taking into account the overall image and scene. For example, if the image has the available capability such that the brightest encodable pixel of the image is 1000 nit (the maximum brightness of the image or video PL_V), one grader may choose to give pixels an explosion brightness value between 800 and 1000 nit to make the explosion look rather impactful, while another filmmaker may choose an explosion that is not brighter than 500 nit so as not to suppress too much of the rest of the image at that moment (and of course, the technology should be able to handle both cases).

[0012] The maximum brightness of an HDR image or video can vary significantly and is typically communicated together with the image data as metadata about the HDR video or image (typical values can be, for example, 1000 nit, or 4000 nit, or 10000 nit, these are non - restrictive; when PL_V is at least 600 nit, it is usually said to have an HDR image). If a video creator chooses to define their image as PL_V = 4000 nit, they can of course choose to create a brighter explosion, although relatively it will not reach the 100% level of PL_V but, for example, only reach 50% for such a high PL_V definition of the scene.

[0013] An HDR display can have a maximum capability, i.e., for example (starting from the lower - end HDR display) 600 nit, or 1000 nit, or N times 1000 nit for the highest displayable pixel brightness. This display maximum or peak brightness PL_D is a value separate from the video maximum brightness PL_V, and the two should not be confused. A video creator usually cannot make an optimal video for every possible end - user display (i.e., where the capabilities of the end - user display are best utilized by the video, where the maximum brightness of the video (ideally) never exceeds the maximum brightness of the display but is also not lower, i.e., there should be at least some pixels in some video images with pixel brightness L_p = PL_V, and the continuous optimization for any particular display will also involve PL_V = PL_D).

[0014] The creator will make some of their own decisions (e.g., what type of content they are capturing and in what way), and will usually produce a video with PL_V so high that they can at least serve the highest PL_D displays of their intended audience today and possibly in the future when higher PL_D displays may emerge.

[0015] Then there is a second problem, which is how to optimally display an image with peak luminance PL_V on a display with a lower (usually much lower) display peak luminance PL_D, which is called display adaptation. Even in the future, there will still be displays that require a lower dynamic range image than, for example, a 2000 nit PL_V image created and received via a communication medium. In theory, a display can always re-scale, i.e., map the luminance of the image pixels so that they can be displayed by its own internal heuristics, but if the video creator is very careful in determining the pixel luminance, it may also be beneficial for him to indicate how his image should be displayed to adapt to the lower PL_D value, and ideally, the display largely follows these technical expectations.

[0016] Regarding the darkest displayable pixel luminance BL_D, the situation is more complex. Some of these may be fixed physical characteristics of the display, such as LCD cell light leakage, but even with the best displays, what the viewer can ultimately distinguish as different darkest blacks also depends on the lighting of the viewing room, which is not a well-defined value. This lighting can be characterized, for example, as an average luminance level in lux, but for video display purposes, it is more elegantly characterized as the minimum pixel luminance. This also typically involves the human eye, in a stronger way than the appearance of bright or medium luminance, because if the human eye is viewing many high-luminance pixels, the darker pixels, and especially their absolute luminance, may become less relevant. But it can be assumed that the eye is not a limiting factor, for example, when viewing a largely dark scene image but still shielded by ambient light in front of the display. If it is assumed that a person can see a just noticeable difference of 2%, there are some darkest drive levels (or luminance information) b above which the next darker luminance information level can still be seen (i.e., the display luminance is X% higher, e.g., 2% more luminance level is displayed).

[0017] In the LDR era, people simply did not care about the darkest pixels. The main consideration was the average luminance, about 1 / 4 of the maximum PL_V = 100 nit. If the image was exposed near this value, everything in the scene looked brightly and vividly good, except for the clipping of the bright parts of the scene above the maximum value by 100%. For the darkest parts of the scene, in cases where they were important enough, capture images with a sufficient amount of base lighting were created in the recording studio or shooting environment. If some scenes were not seen well, for example, because it was drowned into code Y = 0, this was considered normal.

[0018] If nothing further is specified, it can be assumed that the darkest black is zero, or in practice something like 0.1 or 0.01 nit. In this case, the technician is more concerned with the brighter pixels above average in the HDR image as encoded and / or displayed.

[0019] Regarding encoding, the differences between HDR and SDR are not only physical (more pixel brightnesses are to be displayed on a display with a greater dynamic range capability), but also technical in terms of different luminance information code assignment functions (for which the OETF is used; or the inverse of the EOTF in the absolute method), and may also be further technical HDR concepts, such as additional dynamic (per image or a set of temporally consecutive images) changing metadata that specifies how to re-grade the pixel luminance of various image objects to obtain an image with a secondary dynamic range different from the starting image dynamic range (the two luminance ranges typically end with peak luminances differing by at least 1.5x), etc.

[0020] Simple HDR codecs have been introduced on the market, namely the HDR10 codec, which is used, for example, to create the recently emerged black jewel box HDR Blu-ray. This HDR10 video codec uses a more logarithmically shaped function than the square root as the OETF (inverse EOTF), namely the so-called perceptual quantizer (PQ) function standardized in SMPTE 2084. Unlike the Rec.709 OETF which is limited to 1000:1, this PQ OETF allows luminance information to be defined for a greater (ideally to be displayed) luminance, namely between 1 / 10000 nit and 10000 nit, which is sufficient for practical HDR video production.

[0021] Note that the reader should not simply confuse HDR with a large number of bits in the luminance information code word. This may be true for linear systems (such as the amount of bits of an analog-to-digital converter), where, in fact, the amount of bits follows the base-2 logarithm of the dynamic range. However, since the code assignment function can have a rather non-linear shape, theoretically, one hopes to be able to define an HDR image with only 10 bits of luminance information (and even 8 bits for each color component of the HDR image), which results in the advantage of reusability of already deployed systems (for example, an IC can have a certain bit depth or a video cable, etc.).

[0022] After calculating the luminance information, a 10-bit plane with pixel luminance information Y_code has two chrominance components Cb and Cr added to each pixel as chrominance pixel planes. This image can classically be further processed along the line, "as if" it were a mathematically SDR image, such as MPEG-HEVC compression, etc. The compressor actually does not need to care about the pixel color or luminance.

[0023] However, the receiving device (such as a display (or actually its decoder)) usually needs to correctly interpret the pixel color of {Y, Cb, Cr} to display an image that looks correct, rather than an image with, for example, faded colors.

[0024] This is typically handled by co - communicating further image - defining metadata along with the three pixelized color - component planes, which define the image encoding, such as an indication of which EOTF to use. We will assume, but not be limited to, using the PQ EOTF (or OETF), and values such as PL_V.

[0025] More complex codecs can include further image - defining metadata, such as processing metadata, e.g., a function specifying how a normalized version of the luminance of a first image is mapped to a normalized luminance of a secondary reference image, e.g., a PL_V = 100 nit SDR reference image (as we will elucidate in more detail). Figure 2 as elucidated in more detail.

[0026] To bring less HDR - savvy readers up to speed, we quickly elucidate some interesting aspects in Figure 1 which show several prototype illustrative examples of many possible HDR scenarios that a future HDR system (e.g., connected to a 1000 nit PL_D display) may need to be able to handle correctly. The actual technical processing of pixel colors can occur in various ways in various color - space definitions, but the desired re - grading can be shown as an absolute luminance mapping between luminance axes spanning different dynamic ranges. Figure 1 For example, ImSCN1 is a clear outdoor image from a western movie, most of which has bright areas. The first thing that should not be misconstrued is that the pixel luminance in any image is not typically the luminance that can be actually measured in the real world.

[0027] Even without further human involvement in creating the output HDR image (which can serve as a starting image, which we will call the primary HDR grading or image), no matter how simple the act of fine - tuning a single parameter, the camera always measures the relative luminance in the image sensor due to its iris. Thus, in the available encoded luminance range of the primary HDR image, there are always at least some steps involved in the position where the brightest image pixel terminates.

[0028] Even without further human involvement in creating the output HDR image (which can serve as a starting image, which we will call the primary HDR grading or image), no matter how simple the act of fine - tuning a single parameter, the camera always measures the relative luminance in the image sensor due to its iris. Thus, in the available encoded luminance range of the primary HDR image, there are always at least some steps involved in the position where the brightest image pixel terminates.

[0029] For example, in the real world, one can measure the specular reflection of the sun on a sheriff's star badge to be above 100000 nits, but this is neither possible to display on a typical near-future display nor would it make viewers watching an image in a movie (e.g., in a dimly lit room at night) happy. Instead, the video creator can decide that 5000 nits is bright enough for the pixels of the badge, so if this is the brightest pixel in the movie, the video creator can decide to produce a PL_V = 5000 nit video. Although it is a relative pixel luminance measurement device only for the RAW version of the main HDR grading, the camera should also have a high enough native dynamic range (full pixels far exceeding the background noise) to produce good images. The pixels of the graded 5000 nit image are typically derived from the RAW image captured by the camera in a non-linear manner, where, for example, a color grader will consider aspects such as typical viewing situations, which will be different when standing at the actual shooting location (i.e., a hot desert). The best (highest PL_V) image is selected for this scene ImSCN1, i.e., in this example, the 5000 nit image is the main HDR grading. This is the minimum required HDR data to be created and transmitted, but not in all codecs that only transmit data, nor at all in some codecs that even transmit images.

[0030] Having such an encodable high luminance range DR_1 available (e.g., between 0.001 nits and 5000 nits) will allow content creators to provide viewers with a better experience of bright exteriors, but also provide darker night scenes (when well graded throughout the movie), provided, of course, that the viewer also has a corresponding high-end PL_D = 5000 nit display. A good HDR movie not only balances the luminance of various image objects in one image, but also balances the luminance of various image objects over time in a movie story or in generally created video material (e.g., a well-designed HDR football program).

[0031] Some (average) object luminances are shown on the leftmost vertical axis of Figure 1 as one would like to see them in a 5000 nit PL_V main HDR grading, ideally for a 5000 nit PL_D display. For example, in a movie, one might want to show a brightly sunlit cowboy with a pixel luminance of approximately 500 nits (i.e., typically 10x brighter than LDR, although another creator might expect a slightly less HDR impact, e.g., 300 nits), which will depend on how the creator composes the best way to display this western image to provide the best possible look for the end consumer.

[0032] The need for a higher dynamic luminance range can be more easily understood by considering an image that has very dark regions in the same image (such as the shadow corners of the cave image ImSCN3) but also relatively large regions with very bright pixels (such as the sunlight of the outside world seen through the cave entrance). This results in a different visual experience compared to, for example, the night image of ImSCN2, where only the streetlights contain high-luminance pixel regions.

[0033] The problem now is that it is necessary to be able to define a PL_V_SDR = 100 nit SDR image that best corresponds to the main HDR image, because there are still many consumers with LDR displays at this time, and even in the future, there are good reasons to grade movies twice rather than encoding only the prototype's unique HDR image itself. This is a technical expectation separate from the technical choices regarding the encoding itself. For example, it is proven that if one knows how to (reversibly) create one of the main HDR image and this secondary image from the other, then one can choose to encode and transmit either one of the pair (effectively transmitting two images at the cost of one, i.e., only transmitting one image and the pixel color component plane per video moment).

[0034] In such a reduced-dynamic-range image, of course, it is not possible to define a 5000 nit pixel luminance object such as a truly bright sun. The minimum pixel luminance or deepest black can also be as high as 0.1 nit instead of the more preferred 0.001 nit.

[0035] Therefore, in some way, it should be possible to make this corresponding SDR image have a reduced luminance dynamic range DR_2.

[0036] This can be done by some automatic algorithm in the receiving-side display. For example, a fixed luminance mapping function can be used, or it can be an algorithm adjusted by simple metadata such as the PL_V_HDR value and possibly one or more other luminance values.

[0037] However, more complex luminance mapping algorithms can generally be used. But for this application, without loss of generality, we assume that the mapping is defined by some global luminance mapping function F_L (for example, one function per image), which defines for at least one image how to map all possible luminances (i.e., for example, 0.0001 - 5000) in the first image to the corresponding luminances in the second output image (for example, 0.1 to 100 nit for the SDR output image). The normalization function can be obtained by dividing the luminance along both axes by their respective maximum values. In this context, global means that the same function is used for all pixels of the image, regardless of other conditions such as their position in the image (more general algorithms use several functions for pixels that can be classified according to some criteria).

[0038] Ideally, it should be determined by the video creator how all the luminance should be redistributed along the available range of the secondary image (SDR image), as he best knows how to sub-optimize for the reduced dynamic range such that, given the limitations, the SDR image still looks at least as good as the expected primary HDR image. The reader can understand that actually defining (locating) the object luminance in this way corresponds to defining the shape of the luminance mapping function F_L, the details of which are beyond the scope of this application.

[0039] Ideally, the shape of the function should also change according to different scenes, i.e., a cave scene versus a sunny western scene that is later in the movie or typically according to the time image. This is called dynamic metadata (F_L(t), where t indicates the image moment).

[0040] Now, ideally, the content creator would produce the best image for each case, i.e., for each potential end-user display of a service, e.g., a display with PL_D_MDR = 800 nit, which would require a corresponding PL_V_MDR = 800 nit image, but this would generally take too much effort for the content creator, even in the most expensive offline video creation.

[0041] However, the applicant has previously demonstrated that it is sufficient to make (only) two different dynamic range reference gradings of the scene (usually at the extremes, e.g., 5000 nit is the highest necessary PL_V, and 100 nit is usually sufficient as the lowest required PL_V), because then all other gradings can be automatically derived from these two reference gradings (HDR and SDR) via some (usually fixed, e.g., standardized) display adaptation algorithms, such as those applied in the end-user display that receives the information of the two gradings. Usually, the calculation can be done in any video receiver, such as a set-top box, TV, computer, movie equipment, etc. The communication channel for the HDR image can also be any communication technology, such as terrestrial or cable broadcast, a physical medium such as a Blu-ray Disc, the Internet, a communication channel to a portable device, professional inter-site video communication, etc.

[0042] This display adaptation typically also applies a luminance mapping function, e.g., to the pixel luminance of the main HDR image. However, the display adaptation algorithm needs to determine a luminance mapping function different from F_L_5000 to 100 (which is the reference luminance mapping function connecting the luminance of two reference levels), i.e., the luminance mapping function FL_DA of the display adaptation, which may not be very related to the original mapping function between the two reference levels F_L (there can be several variants of the display adaptation algorithm). The luminance mapping function between the main luminance defined on the 5000 nit PL_V dynamic range and the 800 nit medium dynamic range will be written as F_L_5000 to 800 herein.

[0043] We have symbolically shown the display adaptation (for only one of the average object pixel luminances) by an arrow that does not map to the following position: “simply” expecting the F_L_5000 to 100 function to cross the 800 nit MDR image luminance range, but, for example, slightly higher (i.e., in such an image, the denim must be slightly brighter at least according to the selected display adaptation algorithm). Thus, some more complex display adaptation algorithms may place the denim at the indicated higher position, but some customers may be satisfied with the simpler position where the connection between 500 nit HDR denim and 18 nit SDR denim crosses the 800 nit PL_V luminance range.

[0044] Generally, the display adaptation algorithm calculates the shape of the luminance mapping function FL_DA of the display adaptation based on the shape of the original luminance mapping function F_L (or the reference luminance mapping function, also called the reference regrading function).

[0045] This description based on Figure 1 constitutes the technical expectation of any HDR video coding and / or processing system. In Figure 2 we show some exemplary technical systems and their components to achieve the expectation (non-limiting) according to the applicant's codec method. Those skilled in the art should understand that these components can be embodied in various devices, etc. Those skilled in the art should understand that this example is only presented as part of various HDR codec frameworks to have a background understanding of some operating principles and is not intended to particularly limit any embodiment of the innovative contribution presented below.

[0046] Although possible, the technical communication of two actually different images at each moment (the HDR and SDR gradings are communicated as their respective three color planes) is especially expensive in terms of the required data volume.

[0047] This is also not necessary because if it is known that all corresponding secondary image pixel luminances can be calculated based on the luminance in the main image and the function F_L, then it can be decided to transmit only the main image and the function F_L as metadata at each moment (and it can be chosen to transmit the main HDR or SDR image as a representative of both). Since the receiver knows its usually fixed display adaptation algorithm, it can determine the FL_DA function at its end based on this data (there may be metadata that controls or guides further transmission of display adaptation, but it is not currently deployed).

[0048] There can be two modes of communicating the unique image and the function F_L at each moment.

[0049] In the first backward-compatible mode, an SDR image is transmitted ("SDR communication mode"). This SDR image can be directly displayed on a traditional SDR display (without further luminance mapping), but an HDR display needs to apply the F_L or FL_DA function to obtain an HDR image from the SDR image (or its inverse, depending on which variant of the transfer function, upscaling or downscaling variant). Interested readers can find all the details of the applicant's exemplary first-mode method standardized in the following:

[0050] ETSI TS103 433-1V1.2.1(2017-08): High-Performance Single Layer High Dynamic Range System for use in Consumer Electronics devices; Part1: Directly Standard Dynamic Range (SDR) Compatible HDR System (SL-HDR1).

[0051] Another mode transmits the main HDR image itself ("HDR communication mode"), i.e., for example, a 5000 nit image, and the function F_L that allows the calculation of a 100 nit SDR image from it (or any other lower dynamic range image via display adaptation). The main HDR transmitted image itself can be encoded, for example, by using the PQ EOTF.

[0052] Figure 2 A full video communication system is also shown. On the transmission side, it starts with the image source 201. Depending on whether one has, for example, an offline-created video from an Internet delivery company or a real-life broadcast, this can be anything from a hard disk to a wired output from, for example, a TV studio, etc.

[0053] This results in a master HDR video (MAST_HDR), such as colors graded by a human color grader, or a shadow version captured by a camera, or by means of an automatic brightness redistribution algorithm, etc.

[0054] In addition to grading the master HDR image, a set of color transformation functions F_ct that are generally reversible is defined. Without loss of generality, we assume this includes at least one luminance mapping function F_L (however, there can be further functions and data, e.g., specifying how the saturation of a pixel should change from HDR to SDR grading).

[0055] This luminance mapping function defines the mapping between the HDR and SDR reference gradings as described above ( Figure 2 the latter of which is the SDR image Im_SDR to be transmitted to the receiver; with or without data compression via, e.g., MPEG or other video compression algorithms).

[0056] Any color mapping of the color transformer 220 should not be confused with anything applied to the original camera feed to obtain the master HDR video, where it has been assumed that the input is the master HDR video, since this color transformation is used to obtain the image to be transmitted and, at the same time, the desired re-grading, as technically formulated in the luminance mapping function F_L.

[0057] For an exemplary SDR communication type (i.e., SDR communication mode), the master HDR image is input to the color transformer 202, which is configured to apply the F_L luminance mapping to the luminance of the master HDR image (MAST_HDR) to obtain all corresponding luminances written to the output image Im_SDR. For the sake of illustration, assume that a human color grader adjusts the shape of this function for each shot of an image of a similar scene in a movie by using color grading software. In the example MPEG supplementary enhancement information data SEI (F_ct), the applied function F_ct (i.e., at least F_L) is written into (dynamic, processed) metadata to be co-transmitted with the image, or into a similar metadata mechanism in other standardized or non-standardized communication methods.

[0058] After the HDR image to be transmitted is correctly redefined as the corresponding SDR image Im_SDR, these images are generally (at least for broadcasting to end users, for example) compressed using existing video compression techniques (e.g., MPEG HEVC or VVC or AV1, etc.). This is performed in the video compressor 203, which forms part of the video encoder 221 (which in turn can be included in various forms of video creation devices or systems).

[0059] The compressed image Im_COD is transmitted to at least one receiver via some image communication medium 205 (e.g., satellite or cable or Internet transmission, e.g., according to ATSC 3.0 or DVB, etc.; however, the HDR video signal can also be transmitted between two video processing devices via cable, for example).

[0060] Generally, before communication, the transmission formatter 204 can perform some further conversions, which can be applied according to the system, such as techniques like packetization, modulation, transmission protocol control, etc. This usually involves application-specific integrated circuits.

[0061] At any receiving site, the corresponding video signal de-formatter 206 applies the necessary de-formatting methods to regain the compressed video as a set of, for example, compressed HEVC images (i.e., HEVC image data), such as demodulation, etc.

[0062] The video decompressor 207 performs, for example, HEVC decompression to obtain a stream of pixelated uncompressed image Im_USDR, which is an SDR image in this example, but will be an HDR image in other modes. The video decompressor will also unpack the necessary luminance mapping function F_L, or generally the color transformation function F_ct, from, for example, the SEI message. The image and the function are input to the (decoder) color transformer 208, which is arranged to transform the SDR image into an image with any non-SDR dynamic range (i.e., PL_V is higher than 100 nit and is usually at least several times higher, e.g., 5x).

[0063] For example, by applying the inverse color transformation IF_ct of the color transformation F_ct used on the encoding side to generate Im_LDR from MAST_HDR, the 5000 nit reconstructed HDR image Im_RHDR can be reconstructed as a close approximation of the main HDR image (MAST_HDR). Then this image can be sent to, for example, the display 210 for further display adaptation. However, during decoding, by using the FL_DA function (determined in an offline loop, e.g., in the firmware) instead of the F_L function in the color transformer, it is also possible to produce the image Im_DA_MDR for display adaptation in one go. Therefore, the color transformer can also include a display adaptation unit 209 to derive the FL_DA function.

[0064] If the video decoder 220 is included in, for example, a set-top box or a computer, etc., the optimized image Im_DA_MDR with, for example, 800 nit display adaptation can be sent to, for example, the display 210. Or if the decoder resides in, for example, a mobile phone, it can be sent to the display panel. Or if the decoder resides in, for example, some Internet-connected server, etc., it can be transmitted to a cinema projector.

[0065] Figure 3 Shows a useful variant of the internal processing of the color converter 300 of an HDR decoder (or encoder, which generally can have largely the same topology but uses inverse functions and generally does not include display adaptation), namely corresponding to Figure 2 , 208.

[0066] The luminance of a pixel (in this example, an SDR image pixel) is input as the corresponding luminance information Y’SDR. The chrominance, also known as the chrominance components Cb and Cr, is input to the lower processing path of the color converter 300.

[0067] The luminance mapping circuit 310 maps the luminance information Y'SDR to the desired output luminance L'_HDR, such as the primary HDR reconstruction luminance, or some other HDR image luminance. It applies a suitable function obtained from the display adaptation function calculator 350, such as the luminance mapping function FL_DA(t) for display adaptation for a specific image and the maximum display luminance PL_D. The display adaptation function calculator 350 uses the reference luminance mapping function F_L(t) co-transmitted with the metadata as input. The display adaptation function calculator 350 can also determine a suitable function for processing chrominance. For now, we will only assume that a set of multiplication factors mC[Y] for each possible input image pixel luminance information Y is stored, for example, in the color LUT 301. The exact nature of the chrominance processing can vary. For example, it may be desired to keep the pixel saturation constant by first normalizing the chrominance by the input luminance information (the corresponding hyperbola in the color LUT) and then correcting the output luminance information, but any differential saturation processing can also be used. Generally, the hue will be maintained because both chrominance components are multiplied by the same multiplier.

[0068] When indexing the color LUT 301 with the luminance information value Y of the pixel currently undergoing color transformation (luminance mapping), the desired multiplication factor mC is output from the LUT. This multiplication factor mC is used by the multiplier 302 to multiply it by the two chrominance values of the current pixel, i.e., to produce the color-transformed output chrominance

[0069] Cbo = mC * Cb,

[0070] Cro = mC * Cr

[0071] Via a fixed color matrix processor 303 applying standard colorimetry, the chrominance can be converted to the normalized non-linear R’G’B’ coordinates R’ / L’, G’ / L’, and B’ / L’ lacking luminance.

[0072] The R’G’B’ coordinates giving the appropriate luminance of the output image are obtained by the multiplier 311, which calculates:

[0073] R’_HDR = (R’ / L’)*L’_HDR,

[0074] G’_HDR = (G’ / L’)*L’_HDR,

[0075] B’_HDR = (B’ / L’)*L’_HDR, which can be summarized in the color triplet R'G'B'_HDR.

[0076] Finally, the display mapping circuit 320 can further map to the format required by the display. This results in the display drive color D_C, which can be formulated not only in the color metrics expected by the display (e.g., even in the HLG OEFT format), but in some variants, the display mapping circuit 320 can be arranged to perform some specific color processing on the display, i.e., it can, for example, further remap some pixel luminances.

[0077] Some examples are taught in WO2016 / 091406 or ETSITS103 433-2V1.1.1 (2018-01), which illustrate some suitable display adaptation algorithms to derive the corresponding FL_DA functions for creating any possible F_L functions that the side scorer may have determined.

[0078] As mentioned before, the luminance information code representing HDR video / image data is usually represented in a perceptually uniform representation. In such a perceptually uniform representation, the mapping between luminance and the luminance information code (or vice versa) is such that the perceived impact of the luminance change corresponding to a change in one least significant bit of the luminance information code is the same, regardless of the absolute value of the luminance / luminance information code.

[0079] The captured luminance can be mapped to the luminance information code using a suitable OETF (inverse EOTF), which results in such a perceptually uniform representation. Similarly, the luminance information code of such a perceptually uniform representation can be mapped to the luminance of the display output using a suitable OETF. In fact, many different transfer functions have been standardized, such as EOTF_ST2084, which has been standardized by the Society of Motion Picture and Television Engineers SMPTE.

[0080] However, a perceptually uniform domain can provide advantageous quantization where the perceptual impact of the quantization remains constant. However, a perceptually uniform representation is not ideal for all operations. In particular, when using a perceptually uniform representation, performing brightness adjustment processes (such as tone mapping or grading) tends to be sub-optimal. In particular, such operations tend to be complex, resource-demanding, and / or do not provide optimal results. Additionally, conversion to other representations tends to be disadvantageous and can generally increase complexity and introduce errors and distortions in the mapping between the luminance information code and the luminance. This is especially true for perceptually uniform representations that often have extreme behavior towards zero luminance, as the required conversions often do not correspond to well-behaved functions and in fact often approach (positive) infinite first derivative values for luminances close to zero.

[0081] Accordingly, an improved method would be advantageous. In particular, a method that allows for increased flexibility, improved performance, increased image / video quality, reduced error or inaccuracy, convenient processing, reduced complexity and / or resource usage, convenient implementation, and / or an improved spatial audio experience would be advantageous.

[0082] US2020 / 0035198 teaches receiving a first luminance encoded, for example, with the inverse of SMPTE 2084 EOTF, reconstructing the first luminance by applying the EOTF to various pixel luminance values, applying some multiplicative (brightening or darkening) scaling in the linear domain, and converting to a secondary luminance using a different EOTF (such as an SDR EOTF (or more precisely, its inverse opto-electronic transfer function OETF), which would be Rec.709). It does not imply that an EETF, as in the present innovation, directly converts any perceptually uniform luminance to a logarithmic luminance, which may be more suitable for certain kinds of processing, such as multiplicative luminance regrading or relighting processing (in the linear domain). SUMMARY OF THE INVENTION

[0083] Accordingly, the present invention seeks to preferably alleviate, mitigate, or eliminate one or more of the above disadvantages, either singly or in any combination.

[0084] According to one aspect of the present invention, there is provided an apparatus according to claim 1, namely an image processing apparatus, comprising:

[0085] a receiver (401) arranged to receive an image comprising a first pixel luminance (Y’_HDR_PQ) encoded according to a perceptually uniform electro-optical transfer function EOTF;

[0086] A converter (410) arranged to convert the first pixel luminance to a second pixel luminance (Y’_LOG), the second pixel luminance being encoded according to an EOTF defined by a logarithmic mapping from luminance (i.e., optical value) to the second pixel luminance;

[0087] An image processor circuit (411) arranged to apply a luminance adjustment process to the second pixel luminance;

[0088] wherein the converter (410) comprises

[0089] A first converter circuit (601, 605) arranged to generate a first intermediate pixel luminance by applying a first logarithmic function to the first pixel luminance and applying a multiplicative scaling to the first intermediate pixel luminance to produce a scaled first intermediate pixel luminance, wherein the scaling multiplier is an exponent having a value greater than 1;

[0090] A second converter circuit (603) arranged to generate a second intermediate pixel luminance by applying a second logarithmic function to the output value of the perceptually uniform EOTF for the first pixel luminance divided by a divisor, the divisor being equal to the power function of the exponent of the first pixel luminance;

[0091] An adder (607) arranged to obtain the second pixel luminance by adding the scaled first intermediate pixel luminance and the second intermediate pixel luminance.

[0092] This process generally runs pixel by pixel along a scan through the image.

[0093] Since luminance is uniquely defined, when an EOTF or OETF is specified, so is the luminance (e.g., 10-bit quantization). Although the output of this innovation (overall an EETF before the luminance change process, from luminance defined e.g. by SMPTE 2084 to a logarithmically defined output luminance) is always logarithmic, the input EOTF can be configured (preceded in e.g. the processing circuit or even selected immediately before processing), and can for example have various partial logarithmic and partial non-logarithmic function shapes.

[0094] This method can provide improved performance and operation in many scenarios and embodiments. It can in particular allow for a more accurate and / or convenient luminance adjustment process. This method can allow the luminance adjustment process to be performed in the logarithmic domain, thus allowing addition and subtraction to be used to perform luminance scaling and multiplication. Such operations are widely used in most luminance adjustment processes, thus allowing a significant reduction in complexity and resource usage in many applications.

[0095] Certain methods of converting from a perceptually uniform domain for pixel luminance representation to a logarithmic domain can allow for accurate conversion and / or convenient implementation. In particular, the often poorly performing (and having a limit of -∞ for luminance levels approaching complete darkness) portions of the required conversion / transformation function can be addressed by dividing the conversion into different parts that are performed separately and then combined. This method allows the poorly performing parts to be primarily handled by a standard logarithmic function for which practical implementations capable of handling very small input values have been developed.

[0096] Pixel luminance can be represented by a luminance information code / value. Pixel luminance encoded according to an EOTF reflects that there is a mapping between the value of the pixel luminance ( / luminance information code) and the light value, and specifically, the optical light value in the linear domain (e.g., measured in nits). The EOTF can represent this mapping in the direction from the luminance information value / code to the optical light value. The inverse EOTF (EOTF -1 ) represents the same mapping in the direction from the optical light value to the luminance information value / code. Thus, both the EOTF and the matching OETF (=EOTF -1 ) represent the same mapping between the optical light value and the luminance information value / code. Pixel luminance encoded according to a perceptually uniform EOTF is also inherently encoded according to a perceptually uniform OETF (=EOTF -1 ), and vice versa. Pixel luminance encoded according to an EOTF representing a logarithmic mapping from the optical light value to a second pixel luminance is inherently also encoded according to an OETF (=EOTF -1 ) representing the same logarithmic mapping from the optical light value to the second pixel luminance.

[0097] The first logarithmic function and the second logarithmic function can have the same base. The combination of the first and second intermediate pixel luminances can be a (possibly weighted) sum. The pixel luminance can be a luminance information value. The pixel luminance can be a normalized value relative to a reference luminance, e.g., in the range from 0 to 1, where 1 can correspond to, for example, 10,000 nits.

[0098] The luminance adjustment process can be a tone mapping with the second pixel luminance as the input. This method can allow for improved and / or less complex tone mapping.

[0099] According to an optional feature of the present invention, claim 2 is provided.

[0100] This can allow for improved operation and / or performance in many scenarios and embodiments, and can particularly allow for more accurate conversion of pixel luminance to the logarithmic domain. The scaling factor can be a design parameter that can be optimized for a specific application.

[0101] According to an optional feature of the present invention, claim 3 is provided.

[0102] This can allow for improved operation and / or performance in many scenarios and embodiments. In many embodiments, the scaling factor can be set, for example, such that the first pixel luminance maps to the same value as the second pixel luminance of a reference luminance. For example, for a reference luminance of, for example, 10,000 nits, the first pixel luminance Ein = 1 can map to the same value of the corresponding second pixel luminance Eout = 1.

[0103] According to an optional feature of the present invention, claim 4 is provided.

[0104] This can allow for improved operation and / or performance in many scenarios and embodiments, and can particularly allow for a more accurate conversion of pixel luminance to the logarithmic domain.

[0105] The parameters γ and x can be selected to provide the specific performance and operation required for a particular implementation / embodiment.

[0106] According to an optional feature of the present invention, claim 5 is provided.

[0107] This can allow for improved operation and / or performance in many scenarios and embodiments, and can particularly allow for a more accurate conversion of pixel luminance to the logarithmic domain.

[0108] The parameters γ and x can be selected to provide the specific performance and operation required for a particular implementation / embodiment.

[0109] According to an optional feature of the present invention, claim 6 is provided.

[0110] This can allow for improved operation and / or performance in many scenarios and implementation schemes, such as specifically when encoding the first pixel luminance according to the SMPTE ST 2084 EOTF.

[0111] According to an optional feature of the present invention, claim 7 is provided.

[0112] This can allow for improved operation and / or performance in many scenarios and implementation schemes, such as specifically when encoding the first pixel luminance according to the SMPTE ST 2084 EOTF.

[0113] According to an optional feature of the present invention, claim 8 is provided.

[0114] This can allow for improved operation and / or performance in many scenarios and embodiments.

[0115] According to an optional feature of the present invention, claim 9 is provided.

[0116] The method can be particularly suitable for converting pixel luminance encoded according to SMPTE ST 2084 EOTF.

[0117] According to an optional feature of the present invention, claim 10 is provided.

[0118] The method can be particularly suitable for converting pixel luminance encoded according to SMPTE ST 2094-20 EOTF. Another useful EOTF is the applicant's perceptually uniform EOTF as defined in ETSI TS 103 433.

[0119] According to an optional feature of the present invention, claim 11 is provided.

[0120] In many embodiments, this can allow for convenient and / or improved conversion.

[0121] According to one aspect of the present invention, a method according to claim 12 is provided.

[0122] According to an optional feature of the present invention, claim 13 is provided.

[0123] According to an optional feature of the present invention, claim 14 is provided.

[0124] These and other aspects, features and advantages of the present invention will become apparent from and be elucidated with reference to the (one or more) embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0125] These and other aspects of the method and apparatus according to the present invention will become apparent from and be elucidated with reference to the embodiments and examples described hereinafter and with reference to the drawings, which are only non-limiting specific illustrations for exemplifying more general concepts, and in which, dashed lines are used to indicate that a component is optional and non-dashed components are not necessarily essential. Dashed lines can also be used to indicate elements that are interpreted as essential but hidden inside an object, or for intangible things such as the selection of an object / area.

[0126] In the drawings:

[0127] Figure 1Schematically shows a number of typical color transformations that occur when optimally mapping a high - dynamic - range image to a corresponding best - color - graded and similarly - looking lower - dynamic - range image (e.g., a standard - dynamic - range image with a maximum brightness of 100 nits) that is similar to the desired and feasible one given the difference DR_1 resp. DR_2 between the first and second dynamic ranges. In the case of reversibility, it also corresponds to the mapping of the received SDR image (which actually encodes an HDR scene) to a reconstructed HDR image of that scene. Brightness is shown as the position on the vertical axis from the darkest black to the maximum brightness PL_V. The brightness - mapping function is symbolically shown by an arrow that maps the average object brightness from their brightness on the first dynamic range to the second dynamic range (a person skilled in the art knows how to equivalently plot it as a classical function, e.g., on axes normalized to 1, which is normalized by dividing by the corresponding maximum brightness);

[0128] Figure 2 Schematically shows an example of a high - level view of a technique for encoding a high - dynamic - range image, which is an image that the applicant has recently developed and typically can have a brightness of at least 600 nits or greater (usually 1000 nits or greater). It can actually transmit the HDR image itself or an SDR image that is re - graded as the corresponding brightness plus metadata of an encoded color - transformation function, which includes at least an appropriate and determined brightness - mapping function (F_L) of pixel colors to be used by a decoder to convert the received (one or more) SDR images into (one or more) HDR images;

[0129] Figure 3 Shows the details inside an image decoder, particularly the pixel - color processing engine;

[0130] Figure 4 Shows some elements of an example of an image - processing apparatus according to some embodiments of the present invention;

[0131] Figure 5 Shows an example of the conversion from a perceptually - uniform brightness - information domain to a logarithmic - brightness - information domain;

[0132] Figure 6 Shows for Figure 4 an example of elements of a converter of an image - processing apparatus;

[0133] Figure 7 Shows that can be in Figure 6 an example of a conversion function that can be used in a converter;

[0134] Figure 8 Shows that can be in Figure 6 an example of a conversion function that can be used in a converter;

[0135] Figure 9 Shows some elements of a possible arrangement of a processor for implementing elements of an image processing apparatus according to some embodiments of the present invention; and

[0136] Figure 10 Shows some elements of another example of an image processing apparatus of another illustrative embodiment according to the EETF split embodiment. Detailed Description

[0137] Figure 4 Shows an image processing apparatus that can be used to adjust the brightness of pixels of an image, such as a color converter for an HDR decoder or encoder. The apparatus can correspond to Figure 3 the method, and the notes and descriptions provided for Figure 3 can be applied to Figure 4 the corresponding features. The apparatus can be arranged to provide brightness adjustment of the received image.

[0138] Figure 4 The signal processing apparatus of

[0139] includes a receiver 401 arranged to receive an image, which can be an isolated image or can be, for example, an image / frame of a video sequence. The image is formed by pixels represented by values indicating the brightness of the corresponding pixels. For a color image, a suitable color coding / representation can be used, and the following description will focus on a color representation including one brightness / luminance information value and two chrominance values. Specifically, the description will refer to a received image encoded according to the Y’CbCr color space / representation, where Y’ is the luminance information component and CB and CR are the blue-difference and red-difference chrominance components. However, it should be understood that in other embodiments, other color representations can be used, such as other color representations that can have one luminance information component and multiple chrominance components, or color representations using separate color channels, such as RGB-encoded data. In appropriate cases, such color representations can be converted, for example, by the receiver 401 into, for example, the Y’CbCr color space / representation.

[0140] In Figure 4 the method, the image processing circuit 411 is arranged to perform luminance processing. In the method shown, the upper path is thus arranged to perform luminance processing, while the lower path can perform chrominance (CbCr) processing.

[0141] The input data Y’_HDR_PQ of the upper path is specifically HDR pixel luminance information (e.g., an HDR graded image, such as having bright explosion pixels up to 4000 nit), for example, represented as PQ (Perceptual Quantization) luminance information.

[0142] The image processor circuit 411 can perform any luminance processing function, but this is performed in the logarithmic domain, thereby allowing linear light domain scaling / multiplication to be implemented by addition (e.g., the output luminance of brightening can be (linearly, typically) achieved by luminance multiplication by, for example, 3x, and can thus now be implemented logarithmically as addition, etc.). Thus, the converter 410 generates a luminance information value Y’_LOG, which represents luminance as a log(luminance) [e.g., log_10 or log_2] value.

[0143] In many embodiments, the luminance of a pixel is thus given by the luminance information code / value of the pixel.

[0144] The pixel luminance of the image is encoded according to a perceptually uniform EOTF / OETF. Thus, the mapping between the luminance information value and the optical luminance (true linear optical luminance, e.g., measured in nits) is a perceptually uniform mapping. Thus, a difference / change in the step size of one LSB in the luminance information code / value is perceived to result in the same luminance change, regardless of the absolute value (i.e., the step size represented by the luminance information code is considered to be the same perceptually across the entire range). It should be understood that different models or metrics of the perceived impact of luminance change are known and can be the basis for a perceptually uniform representation.

[0145] A perceptually uniform EOTF typically starts with a gamma, also known as a power function of the darkest luminance to be encoded, and gradually changes to a logarithmic mapping of the brightest luminance. Thus, a perceptually uniform EOTF is typically close to a gamma / power function for luminances close to zero luminance and close to a logarithmic function for luminances close to the maximum luminance.

[0146] It should be understood that the EOTF represents the mapping from the luminance information code to the optical luminance, and the OETF represents the mapping from the optical luminance to the luminance information code. Thus, both the EOTF and the OETF represent the mapping between the optical luminance and the luminance information code. In fact, in the sense that a given OETF will also have a corresponding EOTF given as the inverse OETF, i.e., the OETF will have a corresponding EOTF = OETF -1 . Similarly, a given EOTF will also have a corresponding OETF given as the inverse EOTF, i.e., the EOTF will have a corresponding OETF = EOTF -1Therefore, the luminance information code encoded according to the EOTF is inherently also encoded according to the corresponding OETF. In fact, both represent the mapping between the luminance information code and the optical luminance (but in opposite directions). In a practical implementation, the OETF is used to convert the optical luminance value into a luminance information code (i.e., on the capture / input side), and the EOTF is used to convert the luminance information code into an optical luminance value (i.e., on the rendering side). In a practical system, the EOTF used can actually be the inverse of the OETF applied to convert the captured pixel values, but in some practical systems, the OETF can be different, i.e., the inverse of the EOTF used is not necessarily the same as the OETF used (and similarly, the OETF used is not necessarily the same as the EOTF used).

[0147] In Figure 4 a system, the pixel luminance of the received image is represented by a luminance information code encoded according to a perceptually uniform EOTF and thus equivalently encoded according to a perceptually uniform OETF given as the inverse function of the perceptually uniform EOTF. Therefore, the luminance information code represents a quantized value that is designed to have the same relative perceptual impact across the range.

[0148] According to the mapping defined by the Society of Motion Picture and Television Engineers SMPTE as ST 2084 EOTF, the luminance information code can specifically be represented by a value mapped to the optical luminance. In some embodiments, the luminance information code can specifically be represented by a value mapped to the optical luminance according to the mapping defined by SMPTE as ST 2094-20 EOTF.

[0149] The EOTF is generally considered a decompression function (i.e., it is a concave function that stretches or brightens higher luminances compared to lower luminances, as mapped to the output luminance), and conversely, the EOTF -1 (which can be considered the OETF that matches the EOTF) is the inverse and thus the compression function under consideration (i.e., it is convex and has a decreasing slope for higher luminances). Therefore, the perceptually uniform domain can be considered connected to the EOTF -1 because this function converts linear light into perceptually uniform data values. This function generally has and approximate characteristics, and in between, moves from gamma to log behavior with increasing x.

[0150] Any EOTF that provides a one-to-one mapping from luminance information values to linear optical light values inherently also defines an inverse mapping from linear optical light values to luminance information values, i.e., it has a corresponding OETF. Thus, luminance information values inherently represent linear optical light values, and thus the EOTF / OETF defining the mapping is equivalently encoded by definition.

[0151] Contrary to Figure 3 the method of Figure 4 the device does not continue to perform a brightness adjustment process on the received luminance information values. Instead, Figure 4 the device includes a converter 410 arranged to convert a perceptually uniform luminance information code into a luminance information code encoded according to an EOTF defined by a logarithmic mapping from optical light values to second pixel brightness (and equivalently an OETF given by the inverse EOTF). This will also be referred to as a logarithmic EOTF / OETF. The pixel value / luminance information code is thus converted into a representation in which the luminance information code has a logarithmic mapping to optical brightness values. This representation will also be referred to as being in the logarithmic domain. Thus, in this representation, doubling the luminance information code value corresponds to an increase in optical brightness (to which the luminance information code is mapped) by a factor equal to the base of the logarithm.

[0152] Thus, the converter 410 provides an electro-electrical transfer function EETF that maps a luminance information code encoded according to a perceptually uniform mapping to a luminance information code encoded according to a logarithmic mapping. Thus, the converter maps the first / input pixel brightness / luminance information code to the second / output pixel brightness / luminance information code. The converter 410 attempts to generate second pixel brightness values such that they map to the same optical brightness as the first pixel brightness (i.e., for a pixel, the amount of nits represented by the output value is the same as the amount of nits of the input value), but the mapping is different (logarithmic rather than perceptually uniform), and thus the optical brightness is (usually) represented by different values / luminance information codes.

[0153] The converter 410 is coupled to an image processor circuit 411 arranged to apply a brightness adjustment process to the second pixel brightness. The brightness adjustment process can specifically be a tone mapping or grading process in many embodiments and can be a process that selectively adjusts the brightness of individual pixels. However, it should be understood that in other embodiments, other brightness adjustment processes can be performed.

[0154] The luminance information code / value resulting from the brightness adjustment process is output from the image processor circuit 411 and can be further processed or used in any suitable manner. For example, in Figure 4In the example of [[ID=]], the resulting pixel values are fed to a mixer 311, where they are used to scale the relative color channel values derived from the input chrominance values to provide an RGB output of color channels that can be further used, for example, to display an image. The output from the image processor circuit 411 is preferably PQ.

[0155] In other embodiments, the resulting luminance information code can be converted back to a perceptually uniform representation / mapping / encoding and, for example, transmitted or distributed to a remote source for rendering / display.

[0156] The method can offer many advantages, including in particular generally significantly reduced complexity and / or resource requirements. Luminance adjustment processing typically involves scaling of pixel values. For example, during tone mapping, the luminance values are scaled by a given factor, which can be applied to a large (sometimes all) of the pixels of an image. However, scaling is typically associated with an optically linear color representation, and determining the corresponding values in a perceptually uniform domain is complex and resource-demanding because it depends on the absolute value of the luminance and is thus different for different pixels. However, performing such a scaling operation in the logarithmic domain can be efficient and much lower in resource requirements. In particular, complex multiplication operations can be replaced by simpler addition / subtraction operations (since multiplication is converted to addition / subtraction of logarithmic values).

[0157] Therefore, using the EETF to convert from a perceptually uniform representation to a logarithmic representation can significantly reduce the use of computational resources and generally relax the processing requirements for luminance adjustment processing. This can be highly advantageous in many applications and scenarios. However, to ensure a sufficiently high processing quality, it is crucial that the conversion of the EETF is accurate and does not introduce much distortion. However, this is often a very difficult challenge to solve because the required EETF is often not a well-behaved function.

[0158] As a specific example, the first pixel luminance can be represented by a perceptually uniform luminance information code corresponding to the mapping provided by the ST2084 EOTF. The mapping for converting it to a logarithmic mapping can be calculated by the conversion factors required for different luminance information codes.

[0159] Figure 5 An example of such a conversion that performs a conversion from a luminance information code mapping based on the ST2084 EOTF to a luminance information code mapping based on log2 is shown. A particular problem with implementing such a conversion is that for low luminance information values (E in close to zero), the required conversion factor approaches -∞. The EETF mapping required for low luminance information values cannot be easily and accurately represented. It is particularly unsuitable for implementation in a lookup table (LUT) because for low luminance information values, this would require a very high number of entries.

[0160] Note that this problem is not only related to the ST2084 EOTF-based mapping of the received luminance information code, but is generally the case for most perceptually uniform mappings.

[0161] In fact, most perceptually uniform representations are based on the Barten luminance and contrast sensitivity function (CSF) function [see Peter Barten, “Formula for the contrast sensitivity of the human eye,” Proc. SPIE 5294, Image Quality and System Performance, (18 December 2003).]. For such mappings, the inverse EOTF (i.e., the OETF corresponding to the EOTF) behaves like a log function for high luminance levels and like a gamma function for low luminance (dark) levels. Such mappings inherently tend to lead to the problems described as well as difficulties in converting from a perceptually uniform representation to a logarithmic representation.

[0162] To apply tone mapping to an ST2084-based video signal (e.g., HDR10 or SL-HDR2), a choice can be made to apply the transformation in the linear light domain. Since tone mapping involves the application of a tone mapping gain as a function of one or more video components or their derivatives (such as maxRGB), this gain can be very easily implemented in the log domain. Multiplication and division are simply addition or subtraction in the log domain respectively. Another aspect of the log domain is that fewer bits are required in the log domain than in the linear light domain to achieve the same subjective perceived picture quality. For example, in the linear light domain, approximately 28 (= ceil(log2(10 4 / EOTF_ST2084(2 -10 )))) bits are required to represent the entire range of HDR10, while in the logarithmic domain, this is 12 bits (11 + s bits). Operating in the log domain can lead to a smaller, faster, and thus more efficient tone mapping implementation compared to implementing such operations in the linear light or perceptually uniform domain.

[0163] The problem of EETF EOTF-ST2084-to-log is indicated in Figure 5 The indicated region 501 reflects a particularly critical region that is difficult to implement by linear or non-linear LUTs, sub-functions, or (a set of) polynomials due to singularities

[0164] In Figure 4In the system, the converter 410 uses a specific method for converting luminance information values / pixel luminance levels, which allows for low complexity and / or improved conversion in many scenarios. In many embodiments and scenarios, this method may allow for improved and / or convenient implementation and / or operation. In particular, this method allows for the practical implementation of the desired EETF with high accuracy (including for low luminance information values).

[0165] Figure 6 Elements of an exemplary implementation of the converter are shown. In this method, the converter 410 not only applies a conversion function, but also takes a very specific approach of applying two transformations / conversions and then combining them into a suitable luminance information value in the logarithmic domain.

[0166] The converter 410 receives a luminance information value representing the pixel luminance of an image (also referred to as the first pixel luminance). The luminance information value / code is referred to as E in , and is provided in the perceptually uniform domain as previously described.

[0167] The first pixel luminance is fed to a first sub - converter 601, which is arranged to apply a first logarithmic function to the first pixel luminance, for example, a logarithm to the base 2. This results in a modified value that can be referred to as the first intermediate pixel luminance. It should be understood that in other embodiments, other bases of logarithms may be used. In many embodiments, the first sub - converter 601 may also apply a scaling factor to the logarithm. The scaling factor and / or the base of the logarithm may be implementation / design values that can be optimized for specific desired performance and operation of a particular embodiment.

[0168] The first pixel luminance is also fed to a second sub - converter 603, which is arranged to generate a second intermediate pixel luminance from the first pixel luminance. The second sub - converter 603 may specifically apply a perceptually uniform EOTF to the first pixel luminance and then divide the result by a divisor that is equal to an exponential power of the first pixel luminance, where the exponent is greater than 1. The perceptually uniform EOTF is specifically an EOTF that encodes the luminance information value of the first pixel luminance. Thus, the luminance information value of the first pixel luminance is mapped to a linear - light - domain optical luminance (e.g., measured in nits), and the perceptually uniform EOTF is an EOTF that maps from the luminance information value to the optical luminance. Therefore, the perceptually uniform EOTF used by the second sub - converter 603 is the inverse transfer function of the perceptually uniform OETF for converting from the captured optical luminance to the luminance information value. The second sub - converter 603 is arranged to apply a second logarithmic function to the output value from the division.

[0169] Specifically, the second sub-converter 603 may apply the following function to the first pixel luminance represented by the luminance information code to generate a second intermediate pixel luminance:

[0170]

[0171] where E in represents the first pixel luminance, and EOTF(E in ) represents the output value of a perceptually uniform EOTF; γ represents an exponent, and x represents the base of the logarithm.

[0172] The first sub-converter 601 may apply the following function to the first pixel luminance represented by the luminance information code to generate a first intermediate pixel luminance:

[0173] E 1stmod (E in ) = log x (E in )

[0174] where E in represents the first pixel luminance; and x represents the base of the logarithm.

[0175] In many embodiments, the second intermediate pixel luminance is scaled by a scaling factor equal to the exponent γ. In Figure 6 this is shown by a separate multiplier 605 for emphasis, but it should be understood that this may be considered part of the first sub-converter 601 shown.

[0176] Thus, the first sub-converter 601 may apply the following function to the first pixel luminance represented by the luminance information code to generate a first intermediate pixel luminance:

[0177] E 1stmod (E in ) = γlog x (E in )

[0178] where γ is the exponent used by the divisor also applied by the second sub-converter 603.

[0179] The first sub-converter 601 (including the multiplier 605) and the second sub-converter 603 are coupled to a combiner 607, which is arranged to combine the first intermediate pixel luminance and the second intermediate pixel luminance. The combiner 607 may specifically be an adder / summing circuit that adds the first intermediate pixel luminance and the second intermediate pixel luminance together to generate the second pixel luminance.

[0180] Overall, the converter 410 may accordingly provide the following EETF applied to the first pixel luminance to generate the second pixel luminance:

[0181]

[0182] The values of x and γ are design parameters that can be adapted to the specific preferences and requirements of each embodiment. The base of the logarithm is the same for the first sub-converter 601 and the second sub-converter 603 in the described examples, but it should be understood that in some embodiments, different bases may be used.

[0183] In the described examples, the first luminance value and the second luminance value are normalized luminance values relative to a given optical luminance value. The pixel luminance is specifically provided in the range of [0, 1], where the upper end of the range (i.e., the luminance information value 1) corresponds to a predetermined luminance level. The predetermined luminance level can specifically be the maximum / peak luminance level representing the highest luminance that can be represented by the luminance information code (in other embodiments, for example, luminance information codes higher than 1 may be allowed). For example, in many embodiments, the luminance information code can be a normalized code with a value of 1 corresponding to a predetermined luminance level of, for example, 10,000 nit.

[0184] In many embodiments, the converter 410 can also be arranged to scale the second pixel luminance by a gain / scaling factor α. In some embodiments, the scaling factor α can be applied to the combined value, for example, at the output of the combiner 607. However, in many embodiments, it can be performed separately in both paths, i.e., by both the first sub-converter 601 and the second sub-converter 603. Thus, the converter 410 can, for example, perform the following operations:

[0185]

[0186] Such an implementation may be particularly advantageous in many embodiments. In particular, the scaling can be effectively built into the logarithmic function implemented by the first sub-converter 601 or can be easily included by modifying the multiplication performed by the multiplier 605. Similarly, the scaling can be included as a component of the function implemented by the second sub-converter 603. For example, this can be implemented as a look-up table with a stored value including the scaling factor α.

[0187] In many embodiments, a common / total scaling factor α can be used to scale the output such that it has an appropriate relationship between the luminance information value of the second pixel luminance and the optical luminance level. In particular, it can be set such that the maximum value of the first pixel luminance (such as in a specific case, the value of E in = 1 corresponding to the maximum pixel luminance (such as 10,000 nit)) produces an output value E out = 1, and this also corresponds to the maximum pixel luminance, i.e., it also corresponds to, for example, 10,000 nit.

[0188] The scaling factor α can scale the pixel luminance such that the maximum input luminance information value / code is mapped to the (same) output luminance information code and to the same maximum pixel luminance. Specifically, in many embodiments, the scaling factor α can be set such that for E in = 1, the converter 410 generates E out = 1.

[0189] In some embodiments, the input (first) pixel luminance can be in a range that does not include the zero value. For example, for a 12-bit representation, the lowest input luminance can be 2 -12 value. However, in many embodiments, an input value / luminance of zero is also possible and is accommodated. However, this can be considered a separate / isolated case where the output (second) pixel luminance in the logarithmic domain is determined not based on the above method but by direct separate mapping. For example, an input value of zero can be directly mapped to an output value of zero, and in fact, the zero value can be preserved for zero luminance in different functions, processes, and domains. In fact, in some embodiments, a zero (luminance) input value can be directly mapped to a zero (luminance) output value (possibly throughout the signal processing chain). Thus, the described apparatus can treat the value as a special case that is handled differently from other values. This method can provide an effective EETF that can convert from pixel luminance and luminance information codes that provide a perceptually uniform representation to pixel luminance and luminance information codes that provide a logarithmic representation. This method is capable of providing an EETF that accurately converts between domains, and in fact, it can be shown that this method can provide a theoretically accurate conversion in many cases. Therefore, this method can implement a substantially convenient luminance adjustment process, where the scaling of luminance levels can be performed by lower-complexity addition / subtraction operations rather than multiplication.

[0190] The subsequent use of combined individual functions further allows for more practical and convenient operation. In fact, it can particularly alleviate the problem that the required conversion factor performs poorly for very low luminance information values and luminance values close to complete darkness. In fact, this method allows the required transfer function to be divided into a well-behaved function (implemented by the second sub-converter 603) and a logarithmic function (implemented by the first sub-converter 601).

[0191] The well-behaved function can be relatively easily implemented, for example, using a LUT with appropriate interpolation, or as a combination of multiple partial functions, each partial function covering a range of luminance information values, and each partial function being an approximation of the well-behaved function within the range.

[0192] For luminance values approaching full black (luminance information values approaching zero), the logarithmic function may not perform well in the sense that it still has the limitation of -∞. However, the logarithm is not a specialized function but a standard mathematical function for which effective implementations have been developed that provide accurate output values for input values approaching zero. In particular, very effective hardware (e.g., VLSI or ASIC logic structures) or signal processing algorithms have been developed that specifically address the problems of logarithmic functions approaching -∞.

[0193] Accordingly, the particular method allows for convenient implementation and improved accuracy in many scenarios and embodiments.

[0194] In many embodiments, the logarithm can be the base-2 logarithm x = 2. This is a particularly advantageous method because it provides a suitable conversion to the logarithmic representation, which can be implemented with a suitable / low complexity. Thus, due to the implementation advantages in VLSI as well as computer programs, log x can preferably be log2. The advantage is characterized, for example, by fewer tables in iterative implementations. In particular, in VLSI, many specific implementations of the log2 function have been developed that provide relatively accurate results for relatively low complexity.

[0195] In many embodiments, the first pixel luminance can be encoded according to a mapping defined by SMPTE, ST 2084 EOTF, or ST 2094-20 EOTF (or equivalently decoded according to the inverse of these EOTFs).

[0196] These perceptually uniform domains are widely used, and current conversion methods provide a particularly effective way to convert these specific perceptually uniform domains to the logarithmic domain with high accuracy, while allowing practical and generally low-complexity operations.

[0197] As previously mentioned, the exponent γ is a design parameter that can be carefully adjusted to provide desired results and effects. For γ values below 3, particularly advantageous performance can be achieved for many embodiments, and more preferably with values between 1.1 and 4; and generally even more preferably with values between 1.6 and 1.8. Particularly advantageous performance has been found for γ values substantially equal to 1.695.

[0198] These values allow for easy / practical implementations that allow for particularly accurate conversions between specific perceptually uniform mappings, particularly those represented by the ST 2084 EOTF or the ST 2094-20 EOTF. Specifically, values between 1.6 and 1.8 can in many cases be particularly suitable for the perceptually uniform mapping represented by the ST 2084 EOTF, and values between 2 and 3 (or more preferably, in many embodiments, between 2.3 and 2.5) can in many cases be particularly suitable for the perceptually uniform mapping represented by the ST 2094-20 EOTF. In fact, for ST 2094-20, a value of exactly 2.4 can be used and can be shown to provide a favorable conversion.

[0199] Similarly, the overall scaling factor α is a design parameter that can be carefully adjusted to provide desired results and effects. For example, in many embodiments, it can be adapted to provide a relative scaling between an input first pixel luminance and an output second pixel luminance for a given reference luminance (e.g., to ensure that the converter 410 for E in = 1, generates E out = 1).

[0200] In many embodiments, particularly favorable performance can be achieved for α below 0.5, and more preferably for values of the scaling factor between values in the range of 0.05 to 0.1; and generally even more preferably for values between 0.07 and 0.08. Particularly favorable performance has been found for α essentially equal to 0.075.

[0201] In many embodiments, the scaling factor can advantageously be determined as:

[0202]

[0203] The above values allow for easy / practical implementations that allow for particularly accurate conversions between specific perceptually uniform mappings, particularly those represented by the ST 2084 EOTF or the ST 2094-20 EOTF. In fact, the indicated favorable values of γ and α can work together advantageously to provide a particularly favorable conversion of pixel luminance, resulting in improved overall performance.

[0204] In many embodiments, the converter 410 can advantageously provide the following function in many embodiments:

[0205]

[0206] Or specifically for the ST 2084 EOTF mapping:

[0207]

[0208] where γ > 1. γ ≈ 1.695, and x = 2.

[0209] This method has been found to provide near-optimal results in terms of conversion accuracy while allowing for practical and efficient implementations that can use standard known methods to determine the logarithm of the second term (implemented by the first sub-converter 601) and that allow for the implementation of the first term (by the second sub-converter 603) using relatively low-complexity means (e.g., by using a LUT). In fact, the function of the first term is shown in Figure 7 . It can be seen that the function performs well and is suitable for implementation using a LUT, possibly in combination with suitable interpolation. The function can also be implemented, for example, by a series of sub-functions, each covering a range of input luminance values and each approximating the Figure 7 function within a specific range.

[0210] In particular, the function implemented by the second sub-converter 603 can be made self-facilitating to implement because the critical section of the conversion (for E close to zero, in close to -∞) has been excised and is handled by the second part / term of the first sub-converter 601. However, since this is a conventional log function, efficient implementations (especially the log2 function) have been developed and are known in the art.

[0211] The first term includes dividing the EOTF, which is the argument of the log function, by the gamma function E γ , where γ > 1. In the case of EOTF ST2084, γ ≈ 1.695 is near-optimal. As mentioned above, the choice in combination with x = 2 is preferred in the case of ST2084. The reason is that when the inverse EOTF of ST2084 is to be extrapolated beyond its 10 4 nit limit, the almost perfect approximation of this extrapolation will be E = 0.25 log 10 (L), where L ≥ 10 4 (in nits). This results in f(E) = E for E ≥ 1, which is advantageous for some applications when exceeding the limit of E = 1.

[0212] It should also be noted that when the mapping between the perceptually uniform pixel luminance / brightness information to the linear light function is according to ST2094 - 20 at L = 10000 nits, gamma γ = 2.4 can result in a substantially perfect and accurate conversion. The function that needs to be implemented by the second sub-converter 603 tends to be smoother (or more mathematically perfect) for ST2094 - 20 than for ST2084.

[0213] Both the ST2084 and ST2094-20 EOTFs are derived from the Barten luminance and CSF functions. The methods and functions described tend to be particularly advantageous for any mapping / EOTF / OETF that is derived from the Barten luminance and CSF and is expected to be transformed to the log domain. The inverse of this EOTF features gamma behavior for dark / black / low luminance and logarithmic behavior for peak white / high luminance values.

[0214] As previously mentioned, in many embodiments, the function implemented by the second sub-converter 603 can be implemented as a LUT. In some embodiments, the LUT can include values indicative of each possible luminance information value of the first pixel luminance. In other embodiments, a smaller LUT can be used, and interpolation between stored values can be used to determine the appropriate output value.

[0215] In some embodiments, the input range, i.e., the range of possible values of the first pixel luminance, can be divided into multiple sub-ranges / segments. For example, as Figure 8 shown, the function can be divided into 23 sub-ranges / segments. For each individual range, a function providing a transformation mapping from the input pixel luminance to the second intermediate pixel luminance can be stored. The function can specifically be a polynomial that has been adapted / optimized to approximate the best function of the second sub-converter 603 in a particular segment. Thus, in Figure 8 the example, a polynomial can be stored for each segment, thus storing a total of 23 polynomial functions.

[0216] For a given first pixel luminance value, the second sub-converter 603 can accordingly determine the sub-range / segment into which the value falls. Then, it can proceed to select / retrieve the corresponding polynomial function and evaluate the function for the first pixel luminance value. The result is then used as the second intermediate pixel luminance value.

[0217] It should be understood that different methods are known for adapting polynomials to approximate a given function (for a given range of values), and any suitable method can be used. It should also be understood that many such embodiments include adapting the ranges of the respective functions. It should also be understood that in some scenarios, a completely manual method can actually be used to determine the appropriate functions and ranges.

[0218] For example, these methods can be implemented, for example, based on an implementation of a non-uniform segmentation method for interpolating a LUT, such as described in D.H. Douglas and T.K. Peucker's "Algorithms for the Reduction of the Number of Points Required to Represent a Line or Its Caricature" (The Canadian Cartographer, Vol. 10, No. 2, pp. 112-122, 1973) or Tsutomu Sasao, Shinobu Nagayama, and Jon T. Butler's "Numerical Function Generators Using LUT Cascades" (IEEE Trans. Computers, Vol. 56, No. 6, June 2007).

[0219] (One or more) devices can be specifically implemented in one or more appropriately programmed processors. In particular, an artificial neural network can be implemented in one or more such appropriately programmed processors. Different functional blocks can be implemented in separate processors and / or can be implemented, for example, in the same processor. Examples of suitable processors are provided below.

[0220] Figure 9 is a block diagram showing an example processor 900 according to an embodiment of the present disclosure. Processor 900 can be used to implement one or more processors that implement the devices or their elements as described above (specifically including one or more artificial neural networks). Processor 900 can be of any suitable processor type, including but not limited to a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) (where the FPGA has been programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (where the ASIC has been designed to form a processor), or a combination thereof.

[0221] Processor 900 can include one or more cores 902. Core 902 can include one or more arithmetic logic units (ALUs) 904. In some embodiments, in addition to or instead of ALU 904, core 902 can include a floating point logic unit (FPLU) 906 and / or a digital signal processing unit (DSPU) 908.

[0222] Processor 900 may include one or more registers 912 communicatively coupled to core 902. Registers 912 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 912 may be implemented using static memory. The registers may provide data, instructions, and addresses to core 902.

[0223] In some embodiments, processor 900 may include one or more levels of cache memory 910 communicatively coupled to core 902. Cache memory 910 may provide computer-readable instructions for execution to core 902. Cache memory 910 may provide data for processing by core 902. In some embodiments, the computer-readable instructions may have been provided to cache memory 910 by local memory (e.g., local memory attached to external bus 916). Cache memory 910 may be implemented using any suitable cache memory type, e.g., metal-oxide semiconductor (MOS) memory such as static random-access memory (SRAM), dynamic random-access memory (DRAM), and / or any other suitable memory technology.

[0224] Processor 900 may include a controller 914 that may control input from other processors and / or components included in the system to processor 900 and / or output from processor 900 to other processors and / or components included in the system. Controller 914 may control the data path in ALU 904, FPLU 906, and / or DSPU 908. Controller 914 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of controller 914 may be implemented as discrete gates, FPGA, ASIC, or any other suitable technology.

[0225] Registers 912 and cache 910 may communicate with controller 914 and core 902 via internal connections 920A, 920B, 920C, and 920D. The internal connections may be implemented as a bus, multiplexer, crossbar, and / or any other suitable connection technology.

[0226] The input and output of processor 900 may be provided via bus 916, which may include one or more wires. Bus 916 may be communicatively coupled to one or more components of processor 900, such as controller 914, cache 910, and / or registers 912. Bus 916 may be coupled to one or more components of the system.

[0227] The bus 916 may be coupled to one or more external memories. The external memory may include a read-only memory (ROM) 932. The ROM 932 may be a mask ROM, an electronically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include a random access memory (RAM) 933. The RAM 933 may be a static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memory may include an electrically erasable programmable read-only memory (EEPROM) 935. The external memory may include a flash memory 934. The external memory may include a magnetic storage device such as a disk 936. In some embodiments, the external memory may be included in the system.

[0228] Figure 10 is a block diagram showing another illustrative embodiment of the present disclosure, which is a max RGB based tone mapper for high dynamic range to higher, medium or lower dynamic range conversion or vice versa. Other brightness metrics besides brightness can be used. For example, even when three color components are processed in parallel, when processed similarly together, R, G and B can vary with respect to brightness rather than more general color changes. For example, the largest of the three R, G and B color components can be used as a brightness metric. For example, in a color quadrant dominated by a red component, the red component will be used as a brightness metric (and an EETF splitting embodiment is used for it). RGB component video is defined in the ST2084 domain in this illustrative embodiment and is represented as R"G"B". The pixel color is the input to the input EETF 1001, which converts the ST2084 video components to the logarithmic domain according to the EETF splitting principle as described above. The output of the EETF 1001 is connected to the input of the input component maximum calculator 1002 to ensure that the maximum color component will be used as brightness (for the brightness processing track), which outputs its three input R log , G log and B log The output of maxRGB 1002 (we will refer to its value as maxRGB) is connected to the input of tone mapper 1003. Tone mapper 1003 can do any tone mapping as needed based on maxRGB (e.g., typically it applies a tone mapping function to maxRGB as input, such as the tone mapping transmitted by the creator or determined by the device itself) and outputs a gain g TM The gain g TM All video components R are added to the lower branch by adder 1004 log , G log and B logNote that, as described above, in the logarithmic domain, addition represents multiplication in the linear domain. The output of the adder 1004 is input to the output EETF 1005, which performs the conversion from the logarithmic domain to the preferred ST2084 for all video components (i.e., the inverse or substantially inverse of the input EETF). Thus, the output video is the tone mapping result of the input video, all represented in ST2084. Note that this principle can be varied, for example, by including another input to the component maximum calculator 1002 (luminance fourth input). Those skilled in the art will understand that some variations illustrated with this figure, such as addition instead of multiplication, can also be accomplished in other embodiments, as Figure 4 shown. The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The present invention may optionally be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In fact, the functions can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention can be implemented in a single unit, or can be physically and functionally distributed among different units, circuits, and processors.

[0229] Although the present invention has been described in connection with some embodiments, the present invention is not intended to be limited to the specific forms set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Additionally, although features may seem to be described in connection with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0230] Furthermore, although listed separately, multiple modules, elements, circuits, or method steps can be implemented by, for example, a single circuit, unit, or processor. Additionally, although the individual features may be included in different claims, these features may be advantageously combined, and inclusion in different claims does not imply that the combination of features is not feasible and / or advantageous. Moreover, including a feature in one category of claims does not imply a limitation to that category, but rather indicates that the feature is equally applicable to other claim categories. Additionally, the order of features in the claims does not imply any particular order in which the features must operate, and in particular, the order of the individual steps in method claims does not imply that the steps must be performed in that order. Instead, the steps can be performed in any suitable order. Moreover, singular references do not exclude pluralities. Thus, references to "a", "an", "first", "second", etc. do not exclude pluralities. The reference numerals in the claims are provided only as clarifying examples and should not be construed as limiting the scope of the claims in any way.

Claims

1. An image processing apparatus, comprising: a receiver (401) arranged to receive an image including a first pixel luminance (Y’_HDR_PQ) encoded according to a perceptually uniform electro - optical transfer function, EOTF; a converter (410) arranged to convert the first pixel luminance into a second pixel luminance (Y’_LOG), the second pixel luminance being encoded according to an EOTF defined by a logarithmic mapping from luminance to the second pixel luminance; an image processor circuit (411) arranged to apply a luminance adjustment process to the second pixel luminance; wherein the converter (410) includes a first converter circuit (601, 605) arranged to generate a first intermediate pixel luminance by applying a first logarithmic function to the first pixel luminance and applying a multiplicative scaling to the first intermediate pixel luminance to produce a scaled first intermediate pixel luminance, wherein the scaling multiplier is an exponent having a value greater than 1; a second converter circuit (603) arranged to generate a second intermediate pixel luminance by applying a second logarithmic function to the output value of the perceptually uniform EOTF for the first pixel luminance divided by a divisor, the divisor being equal to the power function of the exponent of the first pixel luminance; an adder (607) arranged to obtain the second pixel luminance by adding the scaled first intermediate pixel luminance and the second intermediate pixel luminance.

2. The image processing apparatus according to claim 1, wherein The converter (410) is arranged to scale the second pixel luminance by a scaling factor such that the converter (410) maps the maximum pixel luminance in a series of possible values of the first pixel luminance to the maximum pixel luminance in a series of possible values of the second pixel luminance.

3. The image processing apparatus according to any one of the preceding claims, wherein, The second converter circuit (603) is arranged to generate the second intermediate pixel luminance using the following equation: where E in represents the first pixel luminance, and EOTF(E in ) represents the output value of the perceptually uniform EOTF; γ represents the exponent, and x represents the base of the logarithm.

4. The image processing apparatus according to any one of the preceding claims, wherein, The first converter circuit (601, 605) is arranged to generate the first intermediate pixel luminance using the following equation: E 1stmod (E in ) = γ log x (E in ) where E in represents the first pixel luminance; γ represents the exponent, and x represents the base of the logarithm.

5. The image processing apparatus according to any one of the preceding claims, wherein, The exponent γ has a value between 1.1 and 4.

6. The image processing apparatus according to any one of claims 3 to 5, wherein, The converter (410) is arranged to scale the second pixel luminance by a scaling factor having a value in the range from 0.05 to 0.

1.

7. The image processing apparatus according to any one of the preceding claims, wherein, The logarithm is a base - 2 logarithm.

8. The image processing apparatus according to any one of the preceding claims, wherein, The perceptually uniform EOTF is the Society of Motion Picture and Television Engineers SMPTE, ST2084 EOTF.

9. The image processing apparatus according to any one of the preceding claims, wherein, The perceptually uniform EOTF is the Society of Motion Picture and Television Engineers SMPTE, ST2094 - 20 EOTF.

10. The image processing apparatus according to any one of the preceding claims, wherein, The perceptually uniform EOTF is the EOTF defined in ETSI TS103433.

11. The image processing apparatus according to any one of the preceding claims, wherein, The second converter circuit (603) is arranged to: store at least two functions that map the first input luminance to the second intermediate luminance for sub - ranges of all possible values that the first input luminance can span, each of the at least two functions being a polynomial function; select a first function from the at least two functions as the function for the range in which the first pixel luminance falls; and determining a second intermediate pixel value for the first pixel luminance by applying the first function to the first pixel luminance.

12. An image processing method, comprising:[[]] receiving an image including pixels having a first pixel luminance encoded according to a perceptually uniform electro-optical transfer function, EOTF; converting the first pixel luminance to a second pixel luminance, the second pixel luminance being encoded according to an EOTF defined by a logarithmic mapping from luminance to corresponding values of the second pixel luminance; applying a luminance adjustment process to the second pixel luminance; wherein the conversion comprises generating a first intermediate pixel luminance by applying a first logarithmic function to the first pixel luminance and multiplying and scaling by an exponent having a value greater than 1; generating a second intermediate pixel luminance by applying a second logarithmic function to an output value of the perceptually uniform EOTF for the first pixel luminance divided by a divisor equal to the exponentiated power of the first pixel luminance; obtaining the second pixel luminance by adding the first intermediate pixel luminance and the second intermediate pixel luminance.

13. The method according to claim 12, wherein the second intermediate pixel luminance is determined using the following equation: Among them, E in represents the first pixel luminance, and EOTF(E in ) represents the output value of the perceptually uniform EOTF; γ represents the exponent, and x represents the base of the logarithm.

14. The method according to claim 12 or 13, wherein the first intermediate pixel luminance is determined using the following equation: E 1stmod (E in ) = γ log x (E in ) Among them, E in represents the brightness of the first pixel; γ represents the exponent, and x represents the base of the logarithm.

15. A computer program product comprising computer program code modules which, when the program is run on a computer, are adapted to perform all the steps of any one of claims 11 to 13.

Citation Information

Patent Citations

  • Adjusting device, adjusting method, and program

    US20200035198A1

  • Optimizing high dynamic range images for particular displays

    WO2016091406A1

  • Encoding and decoding HDR videos

    WO2017157977A1