Method and apparatus for inverse tone mapping - Patents.com
The method addresses the lack of control in existing dynamic range extension techniques by using a search process within the inverse tone mapping operator to ensure compliance with light energy constraints, resulting in improved HDR image appearance and display capabilities.
Patent Information
- Application Number
- JP2022551402
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-02
- Filing Date
- 2021-02-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-02-22
AI Technical Summary
Existing methods for extending the dynamic range of low or standard dynamic range images to high dynamic range images lack control over the appearance of the HDR images and often rely on assumptions that are not universally applicable, particularly for extreme increases in dynamic range.
A method that involves obtaining a histogram of a low dynamic range image, applying an inverse tone mapping operator with a gain function dependent on pixel values, and using a search process to identify areas that produce bright areas in the HDR image, while ensuring compliance with predefined light energy constraints such as MaxFall and diffuse white.
This method provides improved control over the appearance of HDR images generated from conventional images, ensuring that they comply with predefined light energy constraints, thereby enhancing the display capabilities of HDR devices.
Smart Images

Figure 0007689537000033 
Figure 0007689537000034 
Figure 0007689537000035
Abstract
Description
[Technical field]
[0001] At least one of the present embodiments relates generally to the field of high dynamic range imaging, and more particularly to methods and apparatus for extending the dynamic range of low or standard dynamic range images. [Background technology]
[0002] Recent advances in display technology are beginning to allow for an expanded dynamic range of color, brightness, and contrast in displayed images. The term image, as used herein, refers to image content, which may be, for example, video or a still picture or image.
[0003]
[0003] Technologies that allow for an extended dynamic range in the luminance or brightness of an image are known as high dynamic range (HDR) imaging. Although several HDR display devices, as well as image cameras capable of capturing images with enhanced dynamic range, are emerging, the amount of HDR content that is available remains very limited. A solution is needed that extends the dynamic range of existing content, thereby enabling these contents to be efficiently displayed on HDR display devices.
[0004] To create conventional (hereafter referred to as LDR for low dynamic range or SDR for standard dynamic range) content for HDR display devices, a reverse or inverse tone mapping operator (ITMO) can be employed. ITMO allows the generation of HDR images from conventional (LDR or SDR) images by using algorithms that process the luminance information of pixels in the image with the aim of recovering or recreating the appearance of the corresponding original scene. Typically, ITMO accepts a conventional image as input and globally expands the luminance range of colors in this image, followed by locally processing highlights or bright areas to enhance the HDR appearance of colors in the image.
[0005] Although several ITMO solutions exist, they generally focus on perceptually reproducing the appearance of the original scene and rely on strict assumptions about the content. Moreover, most enhancement methods proposed in the literature are optimized for extreme increases in dynamic range.
[0006] Typically, HDR imaging is defined by expanding the dynamic range between dark and light values of color luminance, combined with an increased number of quantization steps. To achieve a more extreme increase in dynamic range, many methods combine global expansion with local processing steps that enhance the appearance of highlights and other bright areas of the image. Known global expansion steps proposed in the literature vary from inverse sigmoidal to linear or piecewise linear.
[0007] To enhance bright local features in an image, it is known to create a brightness dilation map that associates to each pixel of the image a dilation value to apply to the brightness of this pixel. In the simplest case, cropped regions of an image can be detected and dilated using a steep dilation curve, but such a solution does not provide sufficient control over the appearance of the image.
[0008] It would be desirable to overcome the above deficiencies.
[0009] It is particularly desirable to improve inverse tone mapping methods that allow improved control over the appearance of HDR images generated from conventional (LDR or SDR) images, and it is also particularly desirable to design new ITMOs with reasonable complexity. Summary of the Invention
[0010] In a first aspect, one or more of the present embodiments provides a method, the method comprising: obtaining a histogram representative of a low dynamic range image, called an LDR image; obtaining an inverse tone mapping operator, called an ITMO function, that makes it possible to obtain pixel values of a high dynamic range image, called an HDR image, from pixel values of the LDR image and a gain function dependent on said pixel values of the LDR image; applying a search process using the obtained histogram to identify areas of the LDR image that, when the ITMO function is applied to said LDR image, produce bright areas in the HDR image, the search process comprising: defining sub-portions of the histogram, called bands, and calculating the contribution of each band and a number of pixels, called the population, each contribution representing the light energy emitted by the pixels represented by said band after application of the ITMO function; and calculating at least one local maximum within the contribution. and for each local maximum, integrating a corresponding band with adjacent bands, referred to as a candidate; identifying at least one local maximum in the population, and for each local maximum, integrating a corresponding band with adjacent bands, referred to as a candidate population; creating integration candidates from each integration candidate population independent of any integration candidate; selecting at least one final integration candidate from the integration candidates in function of information representative of each integration candidate, the information including information representative of light energy emitted by pixels represented by the integration candidate and representative of a number of pixels represented by the integration candidate; and applying a decision process to determine when to modify the gain function using the final integration candidates to ensure that the HDR image complies with at least one predefined light energy constraint.
[0011] In an embodiment, the pixel values are luminance values.
[0012] In an embodiment, the at least one predefined light constraint includes a MaxFall constraint and / or a diffuse white constraint.
[0013] In an embodiment, selecting the at least one final merging candidate includes selecting a subset of merging candidates associated with a highest value of information representing light energy, wherein the at least one final merging candidate is selected from the subset of merging candidates representing a highest number of pixels.
[0014] In an embodiment, the determination process comprises: The method includes determining a pixel value, referred to as a final pixel value, which represents at least one final integration candidate, calculating a value representing an extended pixel value from the final pixel value using an ITMO function, and performing a modification process adapted to modify a gain function when the extended pixel value is higher than a light energy constraint representing a predefined diffuse white constraint value.
[0015] In an embodiment, the SDR image is a current image in a sequence of images, and the final pixel value is temporally filtered using at least one final pixel value calculated for at least one image preceding the current image in the sequence of images.
[0016] In an embodiment, the determination process comprises: It comprises executing a modification process adapted to modify the gain function when a value representing the MaxFall of the HDR image is higher than a light energy constraint representing a predefined MaxFall constraint.
[0017] In an embodiment, the value representing the MaxFall of the HDR image is the sum of the calculated contributions.
[0018] In a second aspect, one or more of the present embodiments provides a device, the device comprising: The invention relates to an electronic circuit adapted to: obtain a histogram representative of a low dynamic range image, called an LDR image; obtain an inverse tone mapping operator, called an ITMO function, making it possible to obtain pixel values of a high dynamic range image, called an HDR image, from pixel values of the LDR image and from a gain function dependent on said pixel values of the LDR image; and apply, using the obtained histogram, a search process for identifying areas of the LDR image which, when the ITMO function is applied to said LDR image, produce bright areas in the HDR image, the search process comprising: defining sub-portions of the histogram, called bands, and calculating for each band a contribution and a number of pixels, called a population, each contribution representing the light energy emitted by the pixels represented by said band after application of the ITMO function; and calculating at least one of the contributions, identifying at least one local maximum in the population and, for each local maximum, integrating the corresponding band with adjacent bands, called a candidate; identifying at least one local maximum in the population and, for each local maximum, integrating the corresponding band with adjacent bands, called a candidate population; creating integration candidates from each integration candidate population independent of any integration candidate; selecting at least one final integration candidate from the integration candidates in function of information representative of each integration candidate, the information representing the light energy emitted by the pixels represented by the integration candidate and including information representing the number of pixels represented by the integration candidate; and applying a decision process to determine when to modify the gain function using the final integration candidate to ensure that the HDR image complies with at least one predefined light energy constraint.
[0019] In an embodiment, the pixel values are luminance values.
[0020] In an embodiment, the at least one predefined light constraint includes a MaxFall constraint and / or a diffuse white constraint.
[0021] In an embodiment, for selecting the at least one final merging candidate, the device is further adapted to select a subset of the merging candidates associated with a highest value of the information representing the light energy, wherein the at least one final merging candidate is selected from the merging candidates of the subset representing a highest number of pixels.
[0022] In an embodiment, for applying the determination process, the device is further adapted to determine a pixel value, called a final pixel value, representing at least one final integration candidate, calculate a value representing a consumed pixel value from the final pixel value using the ITMO function, and perform a modification process adapted to modify the gain function when the expanded pixel value is higher than a light energy constraint representing a predefined diffuse white constraint value.
[0023] In an embodiment, the SDR image is a current image in a sequence of images, and the final pixel value is temporally filtered using at least one final pixel value calculated for at least one image preceding the current image in the sequence of images.
[0024] In an embodiment, for applying the determination process, the device is further configured to execute a modification process adapted to modify the gain function when a value representing the MaxFall of the HDR image is higher than a light energy constraint representing a predefined MaxFall constraint.
[0025] In an embodiment, the value representing the MaxFall of the HDR image is the sum of the calculated contributions.
[0026] In a third aspect, one or more of the present embodiments provides an apparatus comprising a device according to the second aspect.
[0027] In a fourth aspect, one or more of the present embodiments provide a signal generated by the method of the first aspect, or by the device of the second aspect, or by the apparatus of the third aspect.
[0028] In a fifth aspect, one or more of the present embodiments provide a computer program comprising program code instructions for implementing a method according to the first aspect.
[0029] In a sixth aspect, one or more of the present embodiments is an information storage means for storing program code instructions for implementing a method according to the first aspect. [Brief description of the drawings]
[0030] [Figure 1] 1 illustrates example scenarios in which the embodiments described below may be implemented; [Diagram 2] 1 illustrates a schematic diagram of an example of a hardware architecture of a processing module in which various aspects and embodiments can be implemented; [Diagram 3] FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments may be implemented. [Figure 4] 1 illustrates a schematic high-level representation of a first embodiment of a method for improving inverse tone mapping; [Diagram 5] 2 illustrates a schematic diagram of details of a first embodiment of a method for improving inverse tone mapping; [Figure 6] 4 illustrates a schematic diagram of details of a second embodiment of a method for improving inverse tone mapping; [Figure 7] 4 illustrates a schematic diagram of details of a third embodiment of a method for improving inverse tone mapping; [Figure 8] 4 illustrates a schematic diagram of details of a third embodiment of a method for improving inverse tone mapping; [Figure 9A] Three different ITM curves are shown. [Figure 9B] Three different ITM curves are shown. [Figure 9C] Three different ITM curves are shown. [Figure 10] 10 illustrates a schematic diagram of a high-level representation of a second embodiment of a method for improving inverse tone mapping; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0031] There are different kinds of inverse tone mapping methods. For example, in the field of local tone mapping algorithms, WO 2015 / 096955 describes the steps of obtaining, for each pixel P of an image, a pixel expansion exponent value E(P), and then multiplying the luminance Y(P) of pixel P by an expansion luminance value Y(P). exp (P) inverse tone mapping as follows: Y exp (P)=Y(P) E(P) ×[Y enhance (P)] (Equation 1) During the ceremony, ●Y exp (P) is the extended luminance value of pixel P. ●Y(P) is the luminance value of pixel P in the SDR (or LDR) input image. ●Y enhance (P) is the luminance enhancement value of pixel P in the SDR (or LDR) input image. ● E(P) is the pixel expansion exponent value of pixel P.
[0032] The set of values E(P) of all pixels of an image forms an expansion exponent map or expansion function or gain function for the image. This expansion exponent map can be generated by different methods. For example, a method may be to low-pass filter the luminance value Y(P) of each pixel P to obtain a low-pass filtered luminance value Y low (P) and applying a quadratic function to the low-pass filtered luminance values, said quadratic function being defined by parameters a, b and c according to the following equation: E(P)=a[Y low (P)] 2 +b[Y low (P)]+c
[0033] Another method based on WO 2015 / 096955 that facilitates hardware implementation uses the following formula:
[0034]
number
[0035] The above equation can be expressed as follows:
[0036]
number
[0037] Document ITU-R BT.2446-0 publishes a method for converting SDR content to HDR content with a similar formula.
[0038]
number
[0039]
number
[0040] As can be seen above, the expansion is based on a power function whose exponent depends on the luminance value of the current pixel, or on a filtered version of this luminance value.
[0041] More generally, any global expansion method can be expressed as an ITM function of the following form for all input values different from zero (for zero at the input, the output is logically zero): Y exp =Y G(Y) (Formula 2) where G() is the gain function of Y.
[0042] Similarly, all local expansion methods can be expressed in the following way for all input values different from zero:
[0043]
number
[0044] In either case (global or local), the expansion function is monotonic to ensure consistency with the input SDR image.
[0045] Some inverse tone mapping methods use a gain function G() (also called an expansion function) that is based on predefined expansion parameters (e.g. as described in the ITU-R BT.2446-0 document) without adapting to the image content. EP 3249605 discloses a method for inverse tone mapping of an image that can automatically adapt to the image content to be tone mapped. The method uses a set of profiles that form templates. These profiles are pre-determined in a learning phase, which is an offline process. Each profile is defined by visual features, such as a luminance histogram, to which a gain function is associated.
[0046] In the learning phase, profiles are determined from a number of reference images that have been manually graded by a colorist, for which the colorist manually sets the inverse tone mapping parameters and generates gain functions. The reference images are then clustered based on these generated gain functions. Each cluster is processed to extract a representative histogram of luminance and a representative gain function associated with it, thus forming a profile resulting from that cluster.
[0047] When new SDR content is received, a histogram is determined for the SDR image of the SDR content. Each calculated histogram is then compared to each histogram stored in the template resulting from the learning phase to find the histogram that best matches the template. For example, the distance between the calculated histogram and each of the histograms stored in the template is calculated. The gain function associated with the histogram of the template that best matches the calculated histogram is then selected and used to perform inverse tone mapping on the image (or images) that corresponds to the calculated histogram. In this way, the best gain function of the template that is adapted to the SDR image is applied to output the corresponding HDR image.
[0048] Nevertheless, even with the best gain function, as well as with a fixed gain function for more important reasons, bad grading is obtained in some luminance ranges in some specific images. In particular, highlights or bright areas over a large area in an SDR image may result in too bright areas in the HDR image. As a result, some HDR display devices cannot display these HDR images correctly because their power capacity is exceeded. To deal with such HDR images, some display devices apply more or less efficient algorithms to reduce the brightness of the HDR image locally or globally. This capacity of a display is called the MaxFall of the display and is expressed in nits (i.e., candela / m 2 (cd / m 2 MaxFall is expressed as the maximum frame average light level (i.e. the maximum average luminance level of an image) and can be defined as the maximum frame average light level (i.e. the maximum average luminance level of an image). MaxFall can also be thought of on the viewer's side, where large bright areas can dazzle the viewer or at least make the viewing experience of the HDR image unpleasant.
[0049] In another aspect, some recommendations have been made, such as one in document ITU-R BT.2408-1, that introduce the notion of a reference level or diffuse white of "203" nits, especially for PQ (perceptual quantization) method-based production and HLG (hybrid log-gamma) method-based production on a "1000" cd / m2 (nominal peak luminance) display under controlled studio lighting. The reader may refer to recommendation ITU-R BT.2100 for details on the HLG and PQ methods. The signal level of the HDR reference white is specified so as not to be related to the signal level of the SDR "peak white". In another aspect, annex "2" of document ITU-R BT.2408-1 is aimed at the analysis of the reference levels in a first set of images extracted from HLG-based live broadcasts and a second set of test images, concluding that:
[0050] "The HDR reference white level of 203 cd / m2 in Table 1 of this report is consistent with the average diffuse white measured for the content analyzed in this appendix. However, the standard deviation of the diffuse white for the two different content sources is large, indicating a significant spread of diffuse white around the average value."
[0051] These standard deviations translate (assuming a signal of 1000 cd / m2) into a range of about 123 to 345 cd / m2 (i.e., the mean ± one standard deviation) for the first set, and about 80 to 700 cd / m2 (i.e., the mean ± one standard deviation) for the second set. This notion of diffuse white is thus a difficult concept to address, and its level values can vary widely depending on the content.
[0052] At least one of the following embodiments comprises: 1. Ensuring that the MaxFall of the enhanced output HDR image, or at least the bright area portion of the MaxFall, does not exceed (or is close to) a predefined MaxFall value; and / or 2. To improve the inverse tone mapping of at least one input SDR image by tracking in an output HDR image large bright areas that are likely to be diffuse white areas and, if so, ensuring that their average luminance value, depending on their size, is close to a predefined target diffuse white value, provided that this average luminance value is higher than the predefined target diffuse white value.
[0053] The MaxFall constraint and the diffuse white constraint can be seen as light energy constraints for the output HDR image.
[0054] As a result, the present invention aims to reduce the overall brightness of an enhanced HDR image, depending on its content, and not increase it.
[0055] FIG. 1 illustrates an example of a situation in which the embodiments described below can be implemented.
[0056] In Figure 1, device 1, which may be a camera, a storage device, a computer, or any device capable of delivering SDR content, transmits SDR content to system 3 using communication channel 2. Communication channel 2 may be a wired (e.g., Ethernet) or wireless (e.g., WiFi, 3G, 4G, or 5G) network link.
[0057] SDR content includes fixed images or video sequences.
[0058] System 3 converts the SDR content into HDR content, i.e. applies inverse tone mapping to the SDR content to obtain the HDR content.
[0059] The acquired HDR content is then transmitted using a communication channel 4, which may be a wired or wireless network, to a display system 5. The display device then displays the HDR content.
[0060] In an embodiment, the system 3 is included in a display system 5 .
[0061] In an embodiment, device 1, system 3 and display device 5 are all included in the same system.
[0062] In an embodiment, the display system 5 is replaced by a storage device that stores HDR content.
[0063] FIG. 2 illustrates generally an example of a hardware architecture of a processing module 30 included in the system 3 and capable of implementing different aspects and embodiments. The processing module 30 includes a processor or central processing unit (CPU) 300, which may include, by way of non-limiting examples, one or more microprocessors, general purpose computers, special purpose computers, and processors based on multi-core architectures, connected by a communication bus 305; a random access memory (RAM) 301; a read only memory (ROM) 302; and a storage device 303, which may include non-volatile and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), read only memory (ROM), Programmable Read-Only Memory (PROM), random access memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drives, and / or optical disk drives, or a secure digital (SD) card reader and / or a hard disk drive. The module 300 may include a storage media reader and / or a network accessible storage device, such as a hard disk drive (HDD), and at least one communication interface 304 for exchanging data with other modules, devices, systems or equipment. The communication interface 304 may include, but is not limited to, a transceiver configured to transmit and receive data over a communication channel. The communication interface 304 may include, but is not limited to, a modem or a network card.
[0064] The communications interface 304, for example, enables the processing module 30 to receive SDR content and provide HDR content.
[0065] The processor 300 can execute instructions loaded into the RAM 301 from a ROM 302, an external memory (not shown), a storage medium, or a communication network. When the processing module 30 is powered on, the processor 300 can read instructions from the RAM 301 and execute them. These instructions form a computer program that causes the processor 300 to implement, for example, the inverse tone mapping method described below in relation to FIG. 4.
[0066] All or part of the algorithms and steps of the inverse tone mapping method may be implemented in software form by execution of a set of instructions by a programmable machine such as a DSP (Digital Signal Processor) or a microcontroller, or in hardware form by machines or dedicated components such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0067] FIG. 3 illustrates a block diagram of an example of a system 3 in which various aspects and embodiments are implemented. The system 3 may be embodied as a device including various components described below and configured to perform one or more of the aspects and embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected appliances, and servers. The elements of the system 3 may be embodied, alone or in combination, in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one embodiment, the system 3 comprises a processing module 30 that implements an inverse tone mapping method. In various embodiments, the system 3 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports.
[0068] Input to processing module 30 may be provided through various input modules, as shown in block 32. Such input modules may include, but are not limited to, (i) a radio frequency (RF) module, for example, to receive RF signals transmitted over the air from a broadcast station, (ii) a component (COMP) input module (or set of COMP input modules), (iii) a Universal Serial Bus (USB) input module, and / or (iv) a High Definition Multimedia Interface (HDMI) input module. Other examples include composited video, not shown in FIG. 3.
[0069] In various embodiments, the input modules of block 32 have associated respective input processing elements as known in the art. For example, the RF module may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band, which in certain embodiments may be referred to (for example) as a channel, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF module of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include, for example, a tuner that performs various of these functions, including down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In one embodiment of a set-top box, the RF module and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and refiltering to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF module includes an antenna.
[0070] Additionally, the USB and / or HDMI modules may include respective interface processors for connecting system 3 to other electronic devices via USB and / or HDMI connections. It should be appreciated that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or within processing module 30, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented in a separate interface IC or within processing module 30, as desired. The demodulated, error corrected and demultiplexed stream is provided to processing module 30.
[0071] The various elements of system 3 may be provided within a unitary housing, where the various elements may be interconnected and transmit data between them using any suitable connection arrangement, such as an internal bus known in the art, including an Inter-IC (I2C) bus, wires, and printed circuit boards. For example, in system 3, processing module 30 is interconnected to the other elements of system 3 by bus 305.
[0072] The communication interface 304 of the processing module 30 enables the system 3 to communicate over a communication channel 2. The communication channel 2 may be implemented, for example, in a wired and / or wireless medium.
[0073] Data is streamed or otherwise provided to system 3 in various embodiments using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received via a communication channel 2 and a communication interface 304 adapted for Wi-Fi communication. The communication channel 3 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, streaming data is provided to system 3 using a set-top box that delivers data via an HDMI connection of input block 32. In yet other embodiments, streaming data is provided to system 3 using an RF connection of input block 32. As indicated above, various embodiments provide data in a manner other than streaming. In addition, various embodiments use wireless networks other than Wi-Fi, e.g., a cellular network or a Bluetooth network.
[0074] System 3 can provide output signals to various output devices including displays 5, speakers 6, and other peripheral devices 7. Display 5 in various embodiments includes, for example, one or more of a touch screen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 5 can be for a television, a tablet, a laptop, a mobile phone, or other device. Display 5 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). Display device 5 is compatible with HDR content. Other peripheral devices 7 include, in various example embodiments, one or more of a standalone digital video disc (or digital versatile disc) (DVR for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 7 that provide functionality based on the output of system 3. For example, a disc player performs the function of playing the output of system 3.
[0075] In various embodiments, control signals are communicated between system 3 and display 5, speakers 6, or other peripheral devices 7 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that allow control between devices with or without user intervention. The output devices can be communicatively coupled to system 3 via dedicated connections through respective interfaces 33, 34, and 35. Alternatively, the output devices can be connected to system 3 using communication channel 2 via communication interface 304. The display 5 and speakers 6 can be integrated in a single unit with other components of system 3 in an electronic device such as a television. In various embodiments, the display interface 5 includes a display driver, such as a timing controller (T Con) chip, or the like.
[0076] Alternatively, the display 5 and speakers 6 may be separate from one or more of the other components, for example if the RF module of input 32 is part of a separate set-top box. In various embodiments in which the display 5 and speakers 6 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0077] Various implementations involve applying an inverse tone mapping method. As used in this application, inverse tone mapping may encompass all or part of the process performed on a received SDR image or video stream to produce a final HDR output suitable for a display, for example. In various embodiments, such a process includes one or more of the processes typically performed by an image or video decoder, for example, a JPEG decoder or H.264 / AVC (ISO / IEC14496-10-MPEG-4 Part10, Advanced Video Coding), H.265 / HEVC (ISO / IEC23008-2-MPEG-H Part2, High Efficiency Video Coding / ITU-T H.265), or H.266 / VVC (Versatile Video Coding), which are being developed by a joint collaborative team of ITU-T and ISO / IEC experts known as the Joint Video Experts Team (JVET) decoder.
[0078] Where a figure is presented as a flow diagram, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that the figure also provides a flow diagram of the corresponding method / process.
[0079] The implementations and aspects described herein may be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single type of implementation (e.g., discussed only as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in a processor, where a processor refers to a general processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes, for example, a communication device such as a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate communication of information between end users.
[0080] References to "one embodiment" or "embodiment" or "one implementation" or "implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" appearing in various places throughout this specification, as well as any other variations thereof, do not necessarily all refer to the same embodiment.
[0081] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, retrieving information from a memory, or retrieving information from, for example, another device, module, or user.
[0082] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0083] Additionally, the application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" generally involves in some way, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0084] Use of any of " / ", "and / or", "at least one of", "one or more", e.g., "A / B", "A and / or B", "at least one of A and B", "one or more of A and B" is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, "A, B, and / or C" and "at least one of A, B, and C", "one or more of A, B, and C" are intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This may be expanded to include as many of the items listed as would be apparent to one of ordinary skill in this and related arts.
[0085] As will be apparent to one skilled in the art, implementations or embodiments can produce various signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data produced by one of the described implementations or embodiments. For example, a signal can be formatted to convey an HDR image or video sequence of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding the HDR image or video sequence into an encoded stream and modulating a carrier with the encoded stream. The information that the signal carries can be, for example, analog information or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored in a processor-readable medium.
[0086] 4 illustrates a schematic high-level representation of a first embodiment of a method for inverse tone mapping. The objective of the embodiment of FIG. 4 is to respect the diffuse white constraint.
[0087] In FIG. 4, the inverse tone mapping method is performed by a processing module 30 .
[0088] In step 40, the processing module 30 obtains a current input SDR image. The current input SDR image is either a still image or an image of a video sequence. In the following, it is assumed that a gain function G() (as shown in Equations 2 and 3) has been defined for the current input SDR image. The goal of at least one of the following embodiments is to modify the gain curve G corresponding to this gain function G() (or equivalently modify the gain function G()) to comply with the diffuse whiteness constraint.
[0089] In step 41, processing module 30 calculates a histogram representing the current input SDR image (e.g., the histogram is calculated directly on the current input SDR image or on a filtered version of the current input SDR image). The histogram of the current input SDR image is used to identify bright regions of interest, hereinafter referred to as lobes. The current input SDR image is assumed to be gamma-ized (non-linear).
[0090] In an embodiment, the histogram includes a number of bins, nbOfBins, where nbOfBins is an integer multiplier of, for example, 64. For example, nbOfBins=256.
[0091] As an example, in the rest of the document, it is assumed that the target LMax (i.e., the target maximum luminance value) of ITMO is 1000 nits, and the current input SDR image is a 10-bit image, where the value 1023 corresponds to 100 nits. Note that 10 bits is chosen to illustrate the method, but if an 8-bit image is used, a simple scaling by 4 must be applied.
[0092] In that case, the ITMO function can be written as follows:
[0093]
number
[0094] Y SDR ' and Y HDR Both ' are gamma-ized and Y SDR and Y HDRBoth are linear, for example: Y SDR =(Y SDR ' / 255) 2.4 ×100 Y HDR =(Y HDR ' / 1000) 2.4 ×1000
[0095] In step 42, the processing module 30 obtains a gain function G(). Once obtained, the gain function G() makes it possible to obtain the ITM curve from the ITMO function of Equation 4.
[0096] 9A, 9B and 9C show three examples of ITM curves obtained with Equation 4 targeting a 1000 nit display. The SDR input is in the range [0...255] (which means that if it is not an 8-bit image, it needs to be normalized in the range [0;255]) and the output is in the range [0;1000]. The curve in FIG. 9A shows that a maximum input value of 255 (corresponding to an SDR of 100 nits) creates an output equal to 1000, which is 1000 nits when linearized. The curve in FIG. 9B shows that the maximum value is about 700, which corresponds to 425 when linearized, and the curve in FIG. 9C shows a maximum value of about 1200, which corresponds to 1550 nits when linearized.
[0097] From Figures 9A, 9B and 9C, it can be noted that if the SDR image is white (if the SDR image is a 10-bit image, the brightness values of all pixels are 1023), the MaxFall of the HDR image corresponding to Figure 9A is 1000 nits, the MaxFall of the HDR image corresponding to Figure 9B is 425 nits, and the MaxFall of the HDR image corresponding to Figure 9C is 1550 nits.
[0098] In step 43, processing module 30 applies a search process aimed at identifying areas of the current input SDR image that create lobes (i.e. bright areas of interest) in the output HDR image generated by the inverse tone mapping method.
[0099] It can be observed that a large amount of light in the output HDR image can be produced by a large amount of input pixels of the current input SDR image that have intermediate luminance values (e.g., 180 vs. maximum input luminance 255), or a smaller amount of input pixels that have high luminance (i.e., around 250). For example (using LMax = 1000 nits), assuming a gain function G() = 1.25 whatever the input luminance: ● Y HDR (180)=368nits; ● Y HDR (255)=1046nits.
[0100] Then the area centered on "255" will produce the same amount of light as the area centered on "180", but the number of pixels is almost three times larger (actually 1046 = 2.8 * 368). This means that the method needs to search for lobes in the output HDR image and also for lobes (or population lobes) of pixels in the histogram.
[0101] Step 43 is described in more detail below in connection with FIG.
[0102] In step 400, the processing module determines whether the diffuse white constraint is respected. Step 400 includes sub-steps 44, 45, and 46.
[0103] In step 44, processing module 30 calculates a luminance value YposDW that represents at least one of the identified regions of the current input SDR image. An embodiment of step 44 is described in more detail below in conjunction with FIG.
[0104] In step 45, the processing module 30 calculates an extended luminance value YexpDW for the luminance value YposDW. Step 45 is described in more detail below in conjunction with FIG.
[0105] In step 46, the processing module uses the extended luminance value YexpDW to determine whether the diffuse white constraint is respected. To do so, it compares the extended luminance value YexpDW with a diffuse white constraint value DWTarget, which is a predefined value that depends, for example, on the display device intended for displaying the HDR image corresponding to the current input SDR image, and / or on the ambient light of the room in which said HDR image is displayed, and / or on user-supplied parameters intended for displaying the displayed HDR image.
[0106] If the extended luminance value YexpDW is less than or equal to the diffuse white constraint value DWTarget, then the gain curve corresponding to the gain function G() (or equivalently the gain function G()) is not modified and the ITMO function of Equation 4 is used to generate the output HDR image in step 47. Otherwise, if the extended luminance value YexpDW is greater than the diffuse white constraint value DWTarget, then the modified gain curve
[0107]
number
[0108]
number
[0109]
number
[0110]
number
[0111] FIG. 5 illustrates diagrammatically a detailed embodiment of step 43 of the method of inverse tone mapping.
[0112] In step 4300, the processing module calculates the contribution (hereafter also called contrib or energy) of each band of the histogram. A band is a group of contiguous bins of the histogram. In the example of a histogram containing 256 bins, each band contains 4 bins, and the histogram is therefore divided into 64 bands of equal size. If the current input SDR image is coded with 10 bits, each bin will contain 4 luminance values Y SDR The first band of the histogram contains the input luminance values Y from "0" to "15". SDR ', the 32nd band collects input brightness values Y from "496" to "511". SDR ', and the 64th band is input brightness values Y from '1008' to '1023' SDR The contribution of a band represents the light energy emitted by the pixel corresponding to that contribution after application of the ITMO function of Equation 4, i.e., the energy emitted by the pixel represented by that band. Next, define the nth contrib.
[0113]
number
[0114] If the input is encoded into "8" bits, then Equation 5 above can be simplified to:
[0115]
number
[0116] In a variation of step 4300, the contributions are calculated only for a subset of the bands of the histogram, for example, one out of two bands or one out of three bands.
[0117] In step 4301, the processing module 30 calculates a population (contribPop) for each contribution:
[0118]
number
[0119] In step 4302, the processing module 30 identifies maxima that represent regions of high energy within the set of contributions. A contribution contrib[n] of band n is considered a maxima if its value is greater than the values of the contributions of its two adjacent bands (i.e., typically band n-1 and band n+1). In other words, a contribution contrib[n] is a maxima if: contrib[n]>contrib[n-1] and contrib[n]>contrib[n+1]
[0120] Note that for the highest-ranked contribution (herein contrib
[63] ), the condition is contrib
[63] >contrib
[62] . In the following, the bands corresponding to the contributions that represent the local maxima are called candidates.
[0121] In step 4303, the processing module 30 applies a merging process to the candidates. During the merging process, the band corresponding to each candidate is merged with up to LN adjacent bands on the left and up to RN adjacent bands on the right (totaling a maximum width LN+RN+1) as follows: For the LN adjacent bands to the left, the processing module 30 searches for the smallest ln value in [1;LN] such that: ○ contrib[n-ln]>0.2×contrib[n] and ○contrib[n-ln]<1.025×contrib[n-ln+1]. If neither is found, then ln=0. Similarly, for the RN adjacent bands to the right, the processing module 30 searches for the smallest rn value in [1;RN] that satisfies: ○ contrib[n+rn]>0.2×conrib[n] and ○contrib[n+rn]<1.025×contrib[n+rn-1]. with ln+rn≦63. If neither is found, then rn=0.
[0122] The integration is performed between n-ln and n+rn.
[0123] In an embodiment, LN=RN=5.
[0124] The integration of bands around candidates is motivated by the fact that large bright regions (such as clouds in the sky, the proximity of pages in a book, characters in transparent clothing, white animals, etc.) are not uniformly white, but generally exhibit some dispersion around a median that can be characterized by lobes with constant widths.
[0125] Note that the values "0.2" and "1.025" are examples and can be changed to other approximate values.
[0126] In step 4304, the processing module 30 stores each candidate information representative of this candidate. In an embodiment, the stored information includes the lowest and highest positions of the integrated bands in the histogram and the sum of the contributions of the integrated bands, termed the energy of the candidate.
[0127] In step 4305, the processing module 30 searches for a local maximum in the population contribPop associated with the contribution (calculated in step 4301). A population contribPop[n] is a local maximum if it respects the following conditions: contribPop[n]>contribPop[n-1] and contribPop[n]>contribPop[n+1]
[0128] For the highest ranked population (herein contribPop
[63] ), the condition is contribPop
[63] >contribPop
[62] .
[0129] In step 4306, the processing module 30 applies an integrating process to the bands (called candidatePop) corresponding to the identified maxima in the population contribPop. The integrating process applied to candidatePop during step 4306 is identical to the integrating process applied to the candidates during step 4303.
[0130] In step 4307, the processing module 30 verifies whether at least one merging candidate population is independent from any identified merging candidates, where independent means that this merging candidate population does not share any contributing positions with any identified merging candidates.
[0131] If there is at least one independent merge candidate population candidatePop, in step 4308, the processing module 30 finishes identifying the merge candidates by identifying each independent merge candidate population candidatePop as a merge candidate, i.e., for each independent merge candidate population, the processing module 30 creates a merge candidate from the independent merge candidate population. This may occur when the energy increases slightly to a maximum energy that is located further away in terms of higher brightness values. Nevertheless, this region may surround the candidate population candidatePop and embed a large number of pixels potentially representing a large amount of energy.
[0132] In step 4309, once all merging candidates have been identified, the processing module 30 selects from the identified merging candidates a number N_cand of merging candidates with the highest energy. In other words, in step 4309, the processing module selects the N_cand merging candidates consisting of pixels emitting the most light energy. In an embodiment, the number N_cand=5. For each of the N_cand selected merging candidates, the processing module 30 stores information representative of this selected merging candidate, for example, the selected merging candidate is: the lowest position of the consolidation candidate (i.e., corresponding to the value n-nr defined above), named loPos; the highest position of the integration candidate (i.e., corresponding to the value n+h defined above), named hiPos; The position of the center of the merged candidate, named maxEnergyPos, i.e., the position of the candidate itself (i.e., the value n defined above); ● information named "energy" representing the integrated candidate energy (i.e., the sum of all contributions in the integrated candidate), and ● a population named "population" of the integrated candidate (i.e., the sum of all bins of the histogram included in the integrated candidate from position loPos to position hiPos divided by sumOfBins).
[0133] In step 4310, the processing module 30 selects N_cand_Max_Pop (N_cand_Max_Pop > 1 and N_cand_Max_Pop < N_cand) integrated candidates having the maximum population in the set of N_cand selected integrated candidates.
[0134] In an embodiment, N_cand_Max_Pop = 2. In that case, if the two integrated candidates (hereinafter referred to as CP1 and CP2) selected in step 4310 overlap, i.e., ● loPos[CP1]<hiPos[CP2] if maxEnergyPos[CP1]>maxEnergyPos[CP2], or ● loPos[CP2]<hiPos[CP1] if maxEnergyPos[CP1]<maxEnergyPos[CP2], the two integrated candidates are merged (i.e., the energy and population are merged), and the processing module 30 selects an integrated candidate having the third population among the N_cand selected integrated candidates.
[0135] As can be seen, when applying the process of FIG. 5, the processing module 30 selects the N_cand_Max_Pop largest regions in terms of population (i.e., number of pixels) among the N_cand largest regions in terms of energy, which follows the above idea that a large amount of light in the output HDR image may be created by a large amount of input pixels. By finding the N_cand first integration candidates in terms of energy, the processing module 30 obtains the N_cand main lobes, and then selects the N_cand_Max_Pop main lobes that collect the largest number of pixels. Note that the N_cand first energy lobes as well as the N_cand_Max_Pop first populations may be small (or even non-existent), which will be the case if the input image is dark.
[0136] As described below, the N_cand_Max_Pop merging candidates are used to determine whether there is a risk that the diffuse white constraint will not be respected by the output HDR image when applying the ITMO function with the gain function G() to the current input SDR image. Thus, the N_cand_Max_Pop merging candidates are used to determine when to modify the gain function G() (or equivalently, the gain curve G obtained using the gain function) so that the output HDR image respects the predefined light energy constraint (diffuse white).
[0137] FIG. 6 illustrates diagrammatically a detailed embodiment of step 44 of the method of inverse tone mapping.
[0138] In step 4400, the processing module 30 determines the luminance value YposDW that represents the selected N_cand_Max_Pop candidates with the maximum population.
[0139] If N_cand_Max_Pop=2, the processing module 30 applies the following algorithm to determine the luminance value YposDW: The input luminance value Y whose histogram bin is the highest in the band corresponding to the position maxEnergyPos of CP1 (which contains 4 bins in the previous example where the histogram contains 64 consecutive bands of the same size and 256 bins). in and name this input luminance value as firstPopY. ● The input luminance value Y with the highest bin in the histogram in the band corresponding to the CP2 position maxEnergyPos in and name it as the input luminance value secondPopY. ● Determine the luminance value YposDW as follows: YposDW=firstPopY+EnByPopCoef×(secondPopY-firstPopY) (Formula 7) During the ceremony, ○EnByPopCoef=secondEnByPop / (firstEnByPop+secondEnByPop), ○firstEnByPop=(energy[CP1]×population[CP1] 0.5 , ○secondEnByPop=(energy[CP2]×population[CP2]) 0.5 ,
[0140] Therefore, the luminance value YposDW lies somewhere between the two largest pixel regions between the pixels with the highest energy.
[0141] In the above embodiment of step 4400, the parameter EnByPopCoef (and therefore the luminance value YposDW) depends on the square root of the product of the energy by the population of N_cand_Max_Pop (=2) candidates. In equation 7 above, the population and energy have the same weight.
[0142] In another embodiment of step 4400, more weight is given to the populations as follows: ●firstEnByPop=energy[CP1] 0.25×population[CP1] 0.75 , ●secondEnByPop=energy[CP2] 0.25 ×population[CP2] 0.75 ,
[0143] In another embodiment of step 4400, more weight is given to the energy, as follows: ●firstEnByPop=energy[CP1] 0.75 ×population[CP1] 0.25 , ●secondEnByPop=energy[CP2] 0.75 ×population[CP2] 0.25 ,
[0144] If N_cand_Max_Pop=1 with the maximum population, i.e. only one candidate CP1 is determined in the set of N_cand selected candidates, the processing module 30 applies the following algorithm to determine the luminance value YposDW: ● The input luminance value Y whose bin in the histogram is the highest in the band corresponding to the position maxEnergyPos of CP1. in and name it as the input luminance value firstPopY. ● Determine the luminance value YposDW as follows: YposDW=firstPopY (Formula 7bis)
[0145] Optionally (e.g. in an embodiment adapted to a current input SDR image extracted from a video sequence), in step 4401, the processing module 30 applies a temporal filter to the luminance value YposDW. The purpose of the optional step 4401 is to dampen (or even cancel) small luminance variations (or oscillations) between two successive HDR images. The temporal filtering process consists of calculating a weighted average between the luminance value YposDW and a luminance value, denoted recursiveYposDW, which represents the luminance value YposDW calculated for an image preceding the current input SDR image. The luminance values YposDW and recursiveYposDW are calculated as follows: recursiveYposDW=DWFeedBack×recursiveYposDW+(1-DWFeedBack)×YposDW; (Formula 8) YposDW=recursiveYposDW; where DWFeedBack is a weight in the range [0;1]. In an embodiment, DWFeedBack=0.9. In another embodiment, DWFeedBack depends on the frame rate of the video sequence. The higher the frame rate, the higher the weight DWFeedBack. For example, for a frame rate of "25" images per second (Im / s), DWFeedBack=0.95, while for a frame rate of "100" Im / s, DWFeedBack=0.975.
[0146] Note that no filtering is applied if the current input SDR image corresponds to a scene cut (i.e. a part of a video sequence that is not homogenous in terms of content with the image preceding the current input SDR image).
[0147] Note that the luminance value YposDW obtained in step 4400 (or step 4401, if applicable) is a floating point number in the range [0;255] and can be easily scaled to the bit depth of the current input SDR image: YposDW = YposDW × (Ymax / 255) where Ymax=2^n-1 of the input video is encoded with n bits (i.e., Ymax=1023 for n=10).
[0148] FIG. 7 illustrates diagrammatically details of step 45 of the inverse tone mapping method.
[0149] In step 4500, the processing module 30 calculates the extension value Y HDR ' is calculated as follows:
[0150]
number
[0151]
number
[0152] 8 illustrates in schematic detail step 48 of the method of inverse tone mapping. As a reminder, step 48 is executed when the processing module determines in step 46 that the inverse tone mapping applied to the current input SDR image using the gain function G() risks generating an output HDR image containing regions that are too bright. The purpose of step 48 is to reduce the brightness of such regions of the HDR image.
[0153] In step 4800, the processing module 30 determines the population DWpopulation of the selected N_cand_Max_Pop candidates having the largest population.
[0154] If N_cand_Max_Pop=2, then DWpopulation=population(CP1)+population(CP2) (Equation 10).
[0155] When N_cand_Max_Pop = 1, DWpopulation = population(CP1) (Equation 10bis).
[0156] The higher DWpopulation is, the closer YexpDW must be to DWTarget. For example, an empty bright sun with a size of "1%" of the image size should be ignored (in fact, in that case, there is no risk of dazzling the eyes of the user watching the video, for example), while a bright skating rink with a size of "60%" of the image size can be locked at DWTarget.
[0157] Next, two parameters are introduced: ● Parameter loThresholdPop: When DWpopulation < loThresholdPop, the gain at YposDW is maintained as it is. That is, when DWpopulation represents less than loThresholdPop% of the total number of pixels, the gain curve
[0158]
Number
[0159]
Number
[0160] The variables loThresholdPop and DWsensitivity are used to calculate the variable modDWpopulation: modDWpopulation = (DWpopulation - loThresholdPop) / (0.65 - 0.4 × DWsensitivity) (Equation 11) modDWpopulation is limited to the range [0;1].
[0161] As can be seen, when DWpopulation < loThresholdPop, modDWpopulation = 0. As a result, modDwpopulation = 0 indicates that there is no need to modify the gain curve
[0162]
Number
[0163] Regarding DWsensitivity: ● DWsensitivity = 0 induces hiThresholdPop = 0.7: DWpopulation should represent at least 70% of the pixels in the image so that ultimately YexpDW = DWTarget. In that case, the inverse tone mapping method, and especially the process of modifying the gain curve (or equivalently, the gain function G()), becomes less responsive to DWpopulation. ● DWsensitivity = 0.5 induces hiThresholdPop = 0.5. DWpopulation should represent at least 50% of the pixels in the image so that YexpDW = DWTarget. ● DWsensitivity = 1 induces hiThresholdPop = 0.3. DWpopulation should represent at least 30% of the pixels in the image so that ultimately YexpDW = DWTarget. In that case, the inverse tone mapping method, and especially the gain curve
[0164]
number
[0165] In step 4801, the processing module 30 derives a variable DWrate from the variable modDWpopulation:
[0166]
number
[0167] In step 4802, the processing module 30 determines YexpDWTarget as follows: YexpDWTarget=DWrate * DWTarget+(1-DWrate) * YexpDW (Equation 13) where YexpDWTarget and YexpDW are linear values.
[0168] As a result, the higher the DWpopulation, the closer the extended value of the luminance value YposDW is to the diffuse white target DWtarget.
[0169] When N_cand_Max_Pop=2, if the population of CP1 is much higher than the population of CP2, the luminance value YposDW will approach the position firstPopY, and then the pixels in lobe CP1 will have their HDR values approach the diffuse white target DWtarget depending on the size of lobe CP1.
[0170] As you can see: When DWrate=0, YexpDWTarget is equal to the luminance value YexpDW, which is the gain curve
[0171]
number
[0172]
number
[0173] In step 4803, the processing module 30 converts the value YexpDWTarget into a gammadized value. YexpDWTarget'=(YexpDWTarget / LMax) 1 / 2.4 ×LMax
[0174] In step 4804, the processing module calculates the gain gainAtDWTarget corresponding to the luminance value YposDW: gainAtDWTarget=log(YexpDWTarget') / log(255 / Ymax×YposDW) (Equation 14)
[0175] In step 4805, the processing module calculates the gain curve obtained using the gain function G() as a whole.
[0176]
number
[0177]
number
[0178]
number
[0179]
number
[0180]
number
[0181] Variation 1 is Y in Gain curve at high levels
[0182]
number
[0183]
number
[0184] gainMod(Y') means "modified gain of gamma-ized input luminance value Y'", HlExp is called high-level exponent, and HlCoef is called high-level coefficient.
[0185] In an embodiment, the high level exponent HlExp=6, but can be lowered while remaining above "2". In an embodiment, the high level coefficient HlCoef is in the range [0;0.3]. HlCoef=0 means that no compression is applied. HlCoef is calculated using equation 15 with HlExp=6 and known values of gain G(YposDW) and gainMod(YposDW).
[0186] Then, a contrast consideration is introduced into the method of compression of the gain curve. In variant 1, the contrast is "0" (minimum contrast). After compression, for the current input SDR image of "10" bits, Equation 15 is applied under the condition that the expanded output of the luminance value "1023" is higher than the expanded output of the luminance value "1020":
[0187]
number
[0188] Modification Example 3 can be implemented using both contrasts "0" and "1" by newly introducing a value of the contrast contrast within the range ]0...1[. Then, c and Hlcoef are calculated as follows (Equation 20): c = contrast × c1 + (1 - contrast) × c0 Hlcoef = contrast × Hlcoef1 + (1 - contrast) × Hlcoef0 Then, HlExp can be obtained using Equation 17 at the position YposDW: gainMod(YposDW) = G(YposDW) - HlCoef × (YposDW / Ymax) HlExp + C And then: c = log((G(YposDW) + c - gainMod(YposDW)) / HlCoef) / log(YposDW / Ymax) (Equation 21) Provided that YposDW is not equal to Ymax. When YposDW = Ymax, a value of Ymax - 1 can be obtained.
[0189] To avoid pumping effects, HlCoef, HlExp and c can be temporally filtered in the same way as YposDW (i.e., in step 4401), but possibly with different feedback values. An alternative and even better way of temporal filtering is to apply feedback to the gain function G() itself. In either case, the temporal filtering is reset at each cut position, similar to YposDW.
[0190] So far, an embodiment dealing with the objective of respecting the diffuse white constraint has been described in relation to Figures 4 to 8. In the following, we show that a similar embodiment can also address the objective of respecting the MaxFall constraint.
[0191] The MaxFall of an image, MF, can be defined as follows: MF=(Σmax(Rp,Gp,Bp)) / nbOfPixelsInTheImage where Rp, Gp and Bp are the three linear color component values of pixel P, and nbOfPixelsInTheImage is the number of pixels in the image.
[0192] In the following, this definition is approximated by the following formula (which is quite correct since the invention deals with bright areas, then areas with high or very high luminance, and we assume that at least two of the three RGB color component values are close): MF=(ΣY P ) / nbOfPixelsInTheImage In the formula, Y P is the gammadized luminance Y of pixel P P ' is the linearized luminance Y of pixel P obtained from P ' is. Y P =Y P ' 2.4 ×LMax In the formula, Y P ' is in the range [0;1] and LMax is the peak nits of the target display. P' can also be a code value encoded on n bits, in this case:
[0193]
number
[0194] For example, SDR input gamma luminance value Y P '= 100 coded on 8 bits has a linear value of 10.6 nits. HDR output luminance value Yexp' P = 400 coded with '10' bits and produced for a '1000' cd / m2 (nominal peak luminance) display device has a linear value of '105' nits.
[0195] In the following, the value MFTarget represents the maximum value that the MaxFall of the augmented image can take. This means that the calculated MaxFall of the augmented image must be lower than MFTarget. If the calculated MaxFall is higher than MFTarget, the gain function is modified to reach this value. The MaxFall constraint value MFTarget is a predefined value that depends for example on the display device intended for the display of the HDR image corresponding to the current input SDR image, and / or on the ambient light of the room in which said HDR image is displayed, and / or on user-supplied parameters intended for the display of the HDR image to be displayed.
[0196] The calculation of MaxFall for the dilated image, i.e. MaxFallOut, is simplified by using the contents of the histogram:
[0197]
number
[0198] As can be seen, MaxFallOut is equal to the sum of the “64” contrib[n] (denoted as Σcontrib[n]) defined in Equation 5 (i.e., MaxFallOut = Σcontrib[n]).
[0199] Figure 10 illustrates a schematic high-level representation of a second embodiment of the method of inverse tone mapping. The aim of the embodiment of Figure 10 is to respect the MaxFall constraint.
[0200] The method of FIG.
[0201] The method of Figure 10 differs from the method of Figure 4 in that steps 43, 400, and 48 are replaced by steps 43bis, 400bis, and 48bis, respectively, while the other steps remain the same.
[0202] Compared to step 43, in step 43bis, steps 4301, 4305, 4306, 4307, 4308 and 4310 are not performed. In step 4309, the processing module 30 selects from the identified merging candidates the merging candidate with the highest energy.
[0203] In step 400bis, the processing module 30 compares the value representative of the MaxFall of the extended HDR image with the MaxFall constraint MFTarget. During step 400bis, the sum Σcontrib[n] of the contributions contrib[n] calculated during step 4300 is considered to represent the MaxFall of the extended HDR image.
[0204] If the sum of all contributions Σcontrib[n] is less than or equal to MFTarget, then the gain function G() (or equivalently, the gain curve G) does not need to be modified (ie, step 47 is performed).
[0205] On the other hand, if the sum of all the contributions Σcontrib[n] is higher than MFTarget, then during step 48bis the processing module 30 uses the candidate CP1 with the highest energy (as defined in relation to FIG. 5 ) and selects the luminance value Y in (In the following, Y pos ), and the brightness value Y pos Modify the corresponding expansion gain as follows: gainAtMFTarget=G(Y pos )+log(MFTarget / MaxFallOut) / (2.4 * log(255×Y pos / Ymax))(Equation 23)
[0206] Y pos and Ymax is the codeword, which is then gamma-ized. If the histogram contains 256 bins, then Y pos is the highest input luminance value Y in the bin maxEnergyPos of CP1 (containing the maximum energy) in It is.
[0207] The processing module 30 then uses Equation 15 to calculate Y pos Modified gain at position gainModMF(Y pos ) to find the gainAtMFTarget=G(Y pos )-HlCoef×(Y pos / Ymax) GlExp =gainAtMFTarget
[0208] The processing module then applies the same strategy (i.e., using the same contrast view) for modifying the gain curve, or equivalently the gain function G(), as used to respect the diffuse white constraint, replacing gainMod and YposDW with gainModMF and Ypos, respectively, using equations 16-22, and finding the three values of HlCoef, HlExp and c.
[0209] In an embodiment, during inverse tone mapping of the current input SDR image, only the method of FIG. 4 that focuses on the diffuse white constraint is applied.
[0210] In an embodiment, during inverse tone mapping of the current input SDR image, only the method of FIG. 10 that focuses on the MaxFall constraint is applied.
[0211] In an embodiment, both the methods of Figure 4 and Figure 10 are applied during the inverse tone mapping of the current input SDR image. In an embodiment, the method of Figure 4 is applied before the method of Figure 10. Indeed, if the method of Figure 4 intended to respect the diffuse white constraint is applied to the current input SDR image, the resulting image may have a MaxFall higher than MFTarget. In this case, a triplet (HlCoef, HlExp, c) is found, which follows: gainAtDWTarget=G(YposDW)-HlCoef×(YposDW / Ymax) HlExp +c Then modify (HlCoef, HlExp, c) to (HlCoef', HlExp', c') to have: gainAtMFTarget=G(Y pos )-HlCoef * (Y pos / Ymax) HlExp '+c' Follow the rules above regarding contrast.
[0212] The above embodiment deals with bright areas. Nevertheless, the MaxFall detection can be extended to large saturated blue and red areas, which can produce large MaxFall values while the corresponding luminance value Y is relatively low. This can be solved by using blue and red histograms (blue and red values can be calculated from Y, U, V values). When using the G709 color space (as defined in recommendation UIT-R BT709), the distribution of R (red), B (blue), and G (green) components is done as follows: ●R is 21% of Y. ●G is 72% Y. ●B is 7% of Y.
[0213] Large whitish areas result in lobes in the luminance (Y) histogram. They also result in lobes that are located at approximately the same position in the R and B histograms. Conversely, if there are large lobes at high values of the R and / or B histograms and none at the same place in the Y histogram (meaning there are Y lobes of comparable size for low values of Y), the value of RGB MaxFall will be large (RGB MaxFall is the actual definition of MaxFall here), but the MaxFall calculated on Y will be small, or at least smaller. The method used to find population candidates on Y can be used for R and B, and then the Y integrated population can be matched with R and / or B at different positions. The energy of those matched population candidates can then be overestimated by considering the lack of green, and then the whole set of standard Y candidates and overestimated Y candidates can be classified. Then, the calculation of MaxFall is performed on the candidate with the highest energy.
[0214] Several embodiments have been described above. The features of these embodiments may be provided alone or in any combination. Furthermore, the embodiments may include one or more of the following features, devices, or aspects, alone or in combination, across various claim categories and types. A television, set-top box, mobile phone, tablet, or other electronic device that executes at least one of the described embodiments. ●A television, set-top box, mobile phone, tablet, or other electronic device that performs at least one of the described embodiments and displays the resulting images (e.g., using a monitor, screen, or other type of display). ●A television, set-top box, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to receive a signal including an encoded video stream and performs at least one of the described embodiments. ●A television, set-top box, mobile phone, tablet, or other electronic device that receives a signal containing an encoded video stream wirelessly (e.g., using an antenna) and executes at least one of the described embodiments.
Claims
1. 1. A method comprising: obtaining a gain function used to obtain an inverse tone mapping operator function that enables obtaining the high dynamic range image from the low dynamic range image by applying a search process to identify areas of the low dynamic range image that produce bright areas in the high dynamic range image when the inverse tone mapping operator function is applied to a low dynamic range image, the search process comprising: - defining bands representing sub-portions of a histogram of said low dynamic range image and obtaining a population representing a contribution and a number of pixels of each band, each contribution representing light energy emitted by pixels represented by said band after application of said inverse tone mapping operator function; generating at least one candidate from at least one maximum in the contribution and at least one maximum in the population; - selecting at least one final candidate from said at least one candidate in function of information representative of each candidate, said information including information representative of the light energy emitted by the pixels represented by said candidate and representative of the number of said pixels represented by said candidate; and applying a decision process using the final candidates to determine modifications to the gain function to ensure that the high dynamic range image complies with at least one light energy constraint.
2. generating at least one candidate from at least one maximum in the contribution and at least one maximum in the population; identifying at least one maximum within said contributions and, for each maximum, integrating a corresponding band with adjacent bands to generate candidates; identifying at least one maximum in said population and, for each maximum, merging a corresponding band with adjacent bands to create a candidate population; and generating a candidate from each candidate population that is independent of any candidate.
3. The method of claim 1 , wherein the pixel values are luminance values.
4. The method of claim 1 , wherein the at least one light energy constraint is at least one of a MaxFall constraint and a diffuse white constraint.
5. said selecting said at least one final candidate 2. The method of claim 1, comprising: selecting a subset of candidates associated with a highest value of information representing light energy, wherein the at least one final candidate is selected from the candidates of the subset representing a highest number of pixels.
6. The determination process comprises: determining a final pixel value representing said at least one final candidate; calculating a value representing an extended pixel value from the final pixel value using the inverse tone mapping operator function; and executing a modification process adapted to modify the gain function in response to determining that the expanded pixel value is higher than a light energy constraint representing a predefined diffuse white constraint value.
7. 7. The method of claim 6, wherein the standard dynamic range image is a current image in a sequence of images, and the final pixel value is temporally filtered using at least one final pixel value calculated for at least one image preceding the current image in the sequence of images.
8. The determination process comprises:
2. The method of claim 1, comprising: executing a modification process adapted to modify the gain function in response to determining that a value representing MaxFall of the high dynamic range image is higher than a light energy constraint representing a predefined MaxFall constraint.
9. The method of claim 8 , wherein the value representing MaxFall of the high dynamic range image is the sum of the obtained contributions.
10. A device, comprising: obtaining a gain function used to obtain an inverse tone mapping operator function that enables obtaining the high dynamic range image from the low dynamic range image by applying a search process to identify areas of the low dynamic range image that produce bright areas in the high dynamic range image when the inverse tone mapping operator function is applied to a low dynamic range image, the search process comprising: defining bands corresponding to sub-portions of a histogram of said low dynamic range image and obtaining a population representing a contribution and a number of pixels of each band, each contribution representing light energy emitted by pixels represented by said band after application of said inverse tone mapping operator function; generating at least one candidate from at least one maximum in the contribution and at least one maximum in the population; - selecting at least one final candidate from said at least one candidate in function of information representative of each candidate, said information including information representative of the light energy emitted by the pixels represented by said candidate and representative of the number of said pixels represented by said candidate; and applying a decision process using the final candidates to determine a modification of the gain function to ensure that the high dynamic range image complies with at least one light energy constraint.
11. generating at least one candidate from at least one maximum in the contribution and at least one maximum in the population; identifying at least one maximum within said contributions and, for each maximum, integrating a corresponding band with adjacent bands to generate candidates; identifying at least one maximum in said population and, for each maximum, merging a corresponding band with adjacent bands to create a candidate population; and generating a candidate from each candidate population independent of any candidate.
12. The device of claim 10 , wherein the pixel values are luminance values.
13. The device of claim 10 , wherein the at least one light energy constraint is at least one of a MaxFall constraint and a diffuse white constraint.
14. To select at least one final candidate, the electronic circuitry The device of claim 10, further adapted to select a subset of candidates associated with a highest value of information representing light energy, wherein the at least one final candidate is selected from the candidates of the subset representing a highest number of pixels.
15. To apply the determination process, the electronic circuit determining a final pixel value representing said at least one final candidate; calculating a value representing an extended pixel value from the final pixel value using the inverse tone mapping operator function; 11. The device of claim 10, further adapted to: execute a modification process adapted to modify the gain function when the expanded pixel value is higher than a light energy constraint representing a predefined diffuse white constraint value.
16. 16. The device of claim 15, wherein the standard dynamic range image is a current image in a sequence of images, and the final pixel value is temporally filtered using at least one final pixel value calculated for at least one image preceding the current image in the sequence of images.
17. To apply the determination process, the device The device of claim 10, further configured to execute a modification process adapted to modify the gain function when a value representative of MaxFall of the high dynamic range image is higher than a light energy constraint representative of a predefined MaxFall constraint.
18. The device of claim 17 , wherein the value representing MaxFall of the high dynamic range image is the sum of the obtained contributions.
19. 13. An information storage medium storing program code instructions for implementing the method of claim 1.
Citation Information
Patent Citations
Image processor, its method and computer readable storage medium
JP2000101840A
Method and apparatus for generating HDR image with reduced clipped area
JP2019165434A
Reshaping curve optimization in HDR coding
US20180007356A1
Method and apparatus for generating HDR images with reduced clipped areas
US20190236761A1