Method and apparatus for inverse tone mapping
By identifying bright areas during the inverse tone mapping process and applying histograms and gain functions for adjustment, the problem of overly bright highlights is solved, effectively expanding high dynamic range images and controlling the brightness of display devices, thereby improving image quality and viewing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INTERDIGITAL CE PATENT HOLDINGS SAS
- Filing Date
- 2021-02-22
- Publication Date
- 2026-07-03
AI Technical Summary
Existing inverse tone mapping methods cannot effectively control the image appearance when extending low dynamic range images to high dynamic range images. In particular, when processing highlight areas, they can easily lead to overly bright areas that exceed the power capacity of the display device and affect the viewing experience.
By identifying bright areas in low dynamic range images, a search process using histograms and gain functions is employed, combined with MaxFall and diffuse white constraints, to adjust the gain function to ensure that the brightness of high dynamic range images remains within a predefined range, avoiding overly bright areas.
It effectively controls the brightness of high dynamic range images, ensuring that the MaxFall and diffuse white constraints of the display device meet the standards, thereby improving the display quality and viewing experience of the image.
Smart Images

Figure CN115335853B_ABST
Abstract
Description
1. Technical Field
[0001] At least one embodiment of the present invention relates generally to the field of high dynamic range imaging, and more particularly to a method and apparatus for extending the dynamic range of low or standard dynamic range images. 2. Background Technology
[0002] Recent advances in display technology have begun to allow for an expanded dynamic range of color, brightness, and contrast in images to be displayed. The term "image" here refers to image content, such as video, still pictures, or images.
[0003] The technology that allows for an extended dynamic range of an image's brightness or luminance is called High Dynamic Range (HDR) imaging. Although many HDR display devices and imaging cameras capable of capturing images with an increased dynamic range have emerged, the amount of available HDR content remains very limited. Solutions are needed that allow for an extended dynamic range of existing content so that it can be efficiently displayed on HDR display devices.
[0004] To prepare conventional (here referred to as Low Dynamic Range LDR or Standard Dynamic Range SDR) content for HDR display devices, an Inverse or Inverse Tone Mapping (ITMO) operator can be used. ITMO allows the generation of HDR images from conventional (LDR or SDR) images using algorithms that process the luminance information of pixels in an image, with the aim of recovering or recreating the appearance of the corresponding original scene. Typically, ITMO takes a conventional image as input, globally expands the luminance range of the image's colors, and then locally processes highlights or bright areas to enhance the HDR appearance of the colors in the image.
[0005] While several ITMO solutions exist, they typically focus on perceptually reproducing the appearance of the original scene and rely on strict presentation of the content. Furthermore, most of the extension methods proposed in the literature are optimized for extreme increases in dynamic range.
[0006] Typically, HDR imaging is limited by expanding the dynamic range between the dark and bright values of a color, combined with an increase in the number of quantization steps. To achieve a more extreme increase in dynamic range, many methods combine global expansion with local processing steps that enhance the appearance of highlights and other bright areas of the image. Known global expansion steps proposed in the literature range from inverse S-shaped changes to linear or piecewise linear ones.
[0007] To enhance bright local features in an image, it is known to create a brightness expansion map, where each pixel of the image is associated with an expansion value to be applied to that pixel's brightness. In the simplest case, cropped regions in the image can be detected, and then expanded using a steeper expansion curve; however, such solutions do not provide sufficient control over the image's appearance.
[0008] The aim is to overcome the above shortcomings.
[0009] In particular, there is a desire to improve inverse tone mapping methods, thereby allowing for better control over the appearance of HDR images generated from conventional (LDR or SDR) images. There is also a particular desire to design novel ITMOs with reasonable complexity. 3. Summary of the Invention
[0010] In a first aspect, one or more embodiments of the present invention provide a method comprising:
[0011] Obtain a histogram representing a low dynamic range image, referred to as an LDR image; obtain an inverse tone mapping operator function, referred to as the ITMO function, thereby allowing the pixel values of a high dynamic range image, referred to as an HDR image, to be obtained from the pixel values of the LDR image and a gain function dependent on those pixel values; apply a search process using the obtained histogram to identify regions of the LDR image that produce bright areas in the HDR image when the ITMO function is applied to the LDR image, the search process including: defining sub-parts of the histogram, referred to as bands, and calculating the contribution and number of pixels of each band, referred to as a group, each contribution representing the light energy emitted after the ITMO function is applied to the pixel represented by the band; in the contribution Identify at least one local maximum in the group, and for each local maximum, aggregate the corresponding band, referred to as a candidate, with the adjacent band; identify at least one local maximum in the group, and for each local maximum, aggregate the corresponding band, referred to as a candidate group, with the adjacent band; create aggregated candidates from each aggregated candidate group, independent of any aggregated candidate; select at least one final candidate from the aggregated candidates based on information representing each aggregated candidate, including information representing the light energy emitted by the pixel represented by the aggregated candidate and information representing the number of pixels represented by the aggregated candidate; and apply a determination process using the final aggregated candidate to determine when to modify the gain function to ensure that the HDR image complies with at least one predefined light energy constraint.
[0012] In the implementation scheme, the pixel value is the brightness value.
[0013] In the implementation, the at least one predefined light constraint includes a MaxFall constraint and / or a diffuse white constraint.
[0014] In one implementation, selecting the at least one final aggregation candidate includes: selecting a subset of aggregation candidates associated with the highest value of information representing light energy, and selecting the at least one final aggregation candidate from the aggregation candidates of the subset representing the highest number of pixels.
[0015] In the implementation plan, the determination process includes:
[0016] The process involves determining a pixel value, referred to as the final pixel value, representing the at least one final aggregation candidate; calculating a value representing an extended pixel value from the final pixel value using the ITMO function; and performing a modification process suitable for modifying the gain function when the extended pixel value is higher than a light energy constraint representing a predefined diffuse white constraint value.
[0017] In the implementation, the SDR image is the current image in the image sequence, and the final pixel value is filtered temporally using at least one final pixel value calculated for at least one image preceding the current image in the image sequence.
[0018] In the implementation plan, the determination process includes:
[0019] Perform a modification process that adapts the gain function when the value of MaxFall representing the HDR image is higher than the light energy constraint representing the predefined MaxFall constraint.
[0020] In the implementation, this value representing the MaxFall of the HDR image is the sum of the calculated contributions.
[0021] In a second aspect, one or more embodiments of the present invention provide an apparatus comprising an electronic circuit system adapted to:
[0022] Obtain a histogram representing a low dynamic range image, referred to as an LDR image; obtain an inverse tone mapping operator function, referred to as the ITMO function, thereby allowing the pixel values of a high dynamic range image, referred to as an HDR image, to be obtained from the pixel values of the LDR image and a gain function dependent on those pixel values; apply a search process using the obtained histogram to identify regions of the LDR image that produce bright areas in the HDR image when the ITMO function is applied to the LDR image, the search process including: defining sub-parts of the histogram, referred to as bands, and calculating the contribution and number of pixels of each band, referred to as a group, each contribution representing the light energy emitted after the ITMO function is applied to the pixel represented by the band; in the contribution Identify at least one local maximum in the group, and for each local maximum, aggregate the corresponding band, referred to as a candidate, with the adjacent band; identify at least one local maximum in the group, and for each local maximum, aggregate the corresponding band, referred to as a candidate group, with the adjacent band; create aggregated candidates from each aggregated candidate group, independent of any aggregated candidate; select at least one final candidate from the aggregated candidates based on information representing each aggregated candidate, including information representing the light energy emitted by the pixel represented by the aggregated candidate and information representing the number of pixels represented by the aggregated candidate; and apply a determination process using the final aggregated candidate to determine when to modify the gain function to ensure that the HDR image complies with at least one predefined light energy constraint.
[0023] In the implementation scheme, the pixel value is the brightness value.
[0024] In the implementation, the at least one predefined light constraint includes a MaxFall constraint and / or a diffuse white constraint.
[0025] In an implementation, in order to select at least one final aggregation candidate, the device is further adapted to: select a subset of aggregation candidates associated with the highest value of information representing light energy, and select the at least one final aggregation candidate from the aggregation candidates of the subset representing the highest number of pixels.
[0026] In an implementation scheme, in order to apply the determination process, the device is further adapted to: determine a pixel value, referred to as a final pixel value, representing the at least one final aggregation candidate; calculate a value representing an extended pixel value from the final pixel value using the ITMO function; and perform a modification process adapted to modify the gain function when the extended pixel value is higher than a light energy constraint representing a predefined diffuse white constraint value.
[0027] In the implementation, the SDR image is the current image in the image sequence, and the final pixel value is filtered temporally using at least one final pixel value calculated for at least one image preceding the current image in the image sequence.
[0028] In the implementation scheme, in order to apply the determination process, the device is also configured to perform a modification process suitable for modifying the gain function when the value of MaxFall representing the HDR image is higher than the light energy constraint representing the predefined MaxFall constraint.
[0029] In the implementation, this value representing the MaxFall of the HDR image is the sum of the calculated contributions.
[0030] In a third aspect, one or more embodiments of the present invention provide an apparatus comprising the device according to the second aspect.
[0031] In a fourth aspect, one or more embodiments of the present invention provide a signal generated by the method of the first aspect or by the device of the second aspect or the apparatus of the third aspect.
[0032] In a fifth aspect, one or more embodiments of the present invention provide a computer program comprising program code instructions for implementing the method according to the first aspect.
[0033] In a sixth aspect, one or more embodiments of the present invention provide an information storage device that stores program code instructions for implementing the method according to the first aspect. 4. Description of the attached drawings
[0034] Figure 1 An example of the context in which the following implementation scheme can be achieved is shown;
[0035] Figure 2 An example of a hardware architecture for a processing module capable of implementing various aspects and implementation schemes is illustrated schematically;
[0036] Figure 3 A block diagram of an example system in which various aspects and implementation schemes are implemented is shown;
[0037] Figure 4 A high-level representation of various first embodiments of a method for improving inverse tone mapping is schematically shown;
[0038] Figure 5 The details of the first aspect of the method for improving inverse tone mapping are schematically illustrated;
[0039] Figure 6 The details of a second aspect of the method for improving inverse tone mapping are schematically illustrated;
[0040] Figure 7 The details of a third aspect of the method for improving inverse tone mapping are illustrated schematically;
[0041] Figure 8 The details of a third aspect of the method for improving inverse tone mapping are illustrated schematically;
[0042] Figure 9A , Figure 9B and Figure 9C This represents three different ITM curves; and,
[0043] Figure 10 A high-level representation of various second embodiments of the method for improving inverse tone mapping is schematically shown; 5. Detailed Implementation
[0044] Different types of inverse tone mapping methods exist. For example, in the field of local tone mapping algorithms, patent application WO2015 / 096955 discloses a method that, for each pixel P of an image, includes the step of obtaining the pixel expansion index value E(P) and then inversely toning the brightness Y(P) of pixel P to the expanded brightness value Y. exp (P) steps:
[0045] Y exp(P) =Y(P) E(P) ×[Y enhance (P)] (Equation 1)
[0046] in:
[0047] ·Y exp (P) is the extended brightness value of pixel P.
[0048] Y(P) is the brightness value of pixel P in the SDR (or LDR) input image.
[0049] ·Y enhance(P) It is the brightness enhancement value P of the pixels in the SDR (or LDR) input image.
[0050] E(P) is the pixel expansion index value of pixel P.
[0051] A set of values E(P) for all pixels in an image forms an exponentially extended graph, or an extended graph, or an extended function, or a gain function of the image. This exponentially extended graph can be generated by various methods. For example, one method involves low-pass filtering the luminance value Y(P) of each pixel P to obtain a low-pass filtered luminance value Y. low (P) and apply the quadratic function to the low-pass filtered brightness value, which is defined by parameters a, b, and c according to the following equation:
[0052] E(P)=a[Y low (P)] 2 +b[Y low (P)]+c
[0053] Another approach to facilitating the concrete implementation of hardware, based on WO2015 / 096955, uses the following equation:
[0054]
[0055] The above equation can be expressed as follows:
[0056]
[0057] The parameter d can be set, for example, to d = 1.25. Y enhance(P) In this case, it is the image brightness value Y(P) and its low-pass version Y. low(P) The functions of both.
[0058] Document ITU-R BT.2446-0 discloses a method for converting SDR content to HDR content using the same type of formula:
[0059] Y′ exp (P)=Y″(P) E(Y″(P))
[0060] in
[0061] Y′ is in the range [0……1]
[0062] ·Y″=255.0×Y′
[0063] When Y″≤T, E=α1Y″ 2 +b1Y″+c1,
[0064] When Y″>T, E=α2Y″ 2 +b2Y″+c2
[0065] ·T=70
[0066] ·α1=1.8712e-5, b1=-2.7334e-3, c1=1.3141
[0067] ·α2=2.8305e-6, b2=-7.4622e-4, c2=1.2528
[0068] As can be seen from the above, the expansion is based on a power function, the exponent of which depends on the brightness value of the current pixel or on the filtered version of that brightness value.
[0069] More generally, all global expansion methods can be expressed as ITM functions of the following form for all input values other than zero (for inputs that are zero, the output is logically zero):
[0070] Y exp =Y G(Y)(Equation 2)
[0071] Where G() is the gain function of Y.
[0072] Similarly, all local expansion methods can be expressed for all input values other than zero as follows:
[0073]
[0074] Where Y F It is the filtered version of Y, G() is the filter of Y F The gain function, and Y enhance It is Y and its surrounding pixels Ys i The function.
[0075] In both cases (global or local), the spread function is monotonic in order to match the input SDR image.
[0076] Some inverse tone mapping methods use a gain function G() (also known as an extension function) based on predetermined extension parameters (such as those described in ITU-R BT.2446-0 document, for example) without any adaptation to the image content. Patent application EP3249605 discloses a method for inverse tone mapping of images that automatically adapts the image content to a tone map. This method uses a set of profiles forming templates. These profiles are predetermined during a learning phase as offline processing. Each profile is defined by visual features such as a luminance histogram associated with the gain function.
[0077] During the learning phase, a profile is determined from a large number of reference images that are manually graded by colorists who manually set the inverse tone mapping parameters and generate gain functions for these images. The reference images are then clustered based on these generated gain functions. Each cluster is processed to extract a representative brightness histogram and its associated representative gain function, thus forming a profile emanating from that cluster.
[0078] Upon receiving new SDR content, a histogram of the SDR image is determined. Each calculated histogram is then compared to each histogram in the template stored from the learning phase to find the best-matching histogram for the template. For example, the distance between the calculated histogram and each histogram in the template is calculated. A gain function associated with the best-matching histogram for the template is then selected and used to perform inverse tone mapping on the image (or images) corresponding to the calculated histogram. In this way, the best-matching gain function for the template suitable for the SDR image is applied to output the corresponding HDR image.
[0079] Nevertheless, even with the optimal gain function, and not to mention with a fixed gain function, undesirable grading of certain brightness ranges in some specific images can still occur. Specifically, bright or highlighted areas on wide regions of an SDR image can create overly bright areas in an HDR image. Therefore, some HDR display devices may be able to display these HDR images correctly because they exceed their power capacity. To handle such HDR images, some display devices apply more or less efficient algorithms to locally or globally reduce the brightness of the HDR image. This capacity of the display is called the display's MaxFall, expressed in nits (i.e., candela / m² (cd / m²)), and can be defined as the maximum frame average luminous level (i.e., the maximum average brightness level of the image). MaxFall can also be considered on the viewer side: large bright areas can dazzle the viewer or at least make viewing HDR images unpleasant.
[0080] On the other hand, several recommendations have been made, such as those in document ITU-R BT.2408-1, which in particular introduce the concept of a reference level or diffuse white of 203 nits for production using the PQ (Perceptual Quantization) method and the HLG (Mixed Log-Gamma) method on a display with a nominal peak brightness of 1000 cd / m² under controlled studio lighting. Readers may refer to the recommendation ITU-R BT.2100 for higher precision on the HLG and PQ methods. The signal level specification for HDR reference white is independent of the signal level for SDR "peak white". On the other hand, the purpose of document ITU-R BT.2408-1 is to summarize in Annex "2" the analysis of the reference level in the first set of images extracted from HLG-based live broadcasts and in the second set of test images:
[0081] "The HDR reference white level of 203 cd / m2 in Table 1 of this report is consistent with the average diffuse white measured in the content analyzed in this appendix. However, the large standard deviation of diffuse white in the two different content sources indicates significant diffusion of diffuse white around the mean."
[0082] These standard deviations (for the assumed “1000” cd / m² signal) shift to a range of approximately “123” to “345” cd / m² (i.e., mean ± one standard deviation) for the first group, and to a range of approximately “80” to “700” cd / m² (i.e., mean ± one standard deviation) for the second group. This means that the concept of diffuse white is a difficult one to resolve, and its level can vary considerably depending on the content.
[0083] At least one of the following implementation schemes aims to improve the inverse tone mapping of at least one input SDR image in the following ways:
[0084] 1. Ensure that the MaxFall or at least the bright portion of MaxFall in the extended output HDR image will not exceed (or will approach) the predefined MaxFall value; and / or
[0085] 2. Continuously track large bright areas that may be diffuse white areas in the output HDR image, and (if so) ensure that their average brightness value is close to the predefined target diffuse white value based on their size, as long as the average brightness value is higher than the predefined target diffuse white value.
[0086] MaxFall constraints and diffuse white constraints can be viewed as light energy limitations for outputting HDR images.
[0087] Therefore, the present invention aims to reduce the overall brightness of an extended HDR image based on its content without increasing the overall brightness.
[0088] Figure 1 An example of the context in which the following implementation scheme can be achieved is shown.
[0089] exist Figure 1 In the system 3, a device 1, which can be a camera, a storage device, a computer, or any device capable of delivering SDR content, uses communication channel 2 to transmit the SDR content to the system 3. Communication channel 2 can be a wired (e.g., Ethernet) or wireless (e.g., WiFi, 3G, 4G, or 5G) network link.
[0090] SDR content includes fixed image or video sequences.
[0091] System 3 converts SDR content to HDR content; that is, it applies inverse tone mapping to SDR content to obtain HDR content.
[0092] The acquired HDR content is then transmitted to the display system 5 using communication channel 4, which can be a wired or wireless network. The display device then displays the HDR content.
[0093] In the implementation scheme, system 3 is included in display system 5.
[0094] In the implementation plan, device 1, system 3 and display device 5 are all included in the same system.
[0095] In the implementation scheme, display system 5 is replaced by a storage device that stores HDR content.
[0096] Figure 2An example of the hardware architecture of a processing module 30 included in system 3 and capable of implementing different aspects and implementations is illustrated schematically. As a non-limiting example, the processing module 30 includes the following items connected by a communication bus 305: a processor or CPU (Central Processing Unit) 300 containing one or more microprocessors, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture; random access memory (RAM) 301; read-only memory (ROM) 302; a storage unit 303, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives and / or optical disk drives, or storage media readers such as SD (Secure Digital) card readers and / or hard disk drives (HDDs) and / or network-accessible storage devices; and at least one communication interface 304 for exchanging data with other modules, devices, systems, or equipment. Communication interface 304 may include, but is not limited to, a transceiver configured to transmit and receive data through a communication channel. Communication interface 304 may include, but is not limited to, a modem or network interface card.
[0097] The communication interface 304 enables, for example, the processing module 30 to receive SDR content and provide HDR content.
[0098] Processor 300 is capable of executing instructions loaded into RAM 301 from ROM 302, external memory (not shown), storage media, or a communication network. When processing module 30 is powered on, processor 300 is capable of reading instructions from RAM 301 and executing these instructions. These instructions form a computer program that causes, for example, processor 300 to perform the following... Figure 4 The inverse tone mapping method is described.
[0099] All or part of the algorithm and steps of this inverse tone mapping method can be implemented in software by executing a set of instructions by a programmable machine such as a DSP (Digital Signal Processor) or a microcontroller, or in hardware by a machine or dedicated component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit).
[0100] Figure 3A block diagram illustrating an example of system 3 in which various aspects and embodiments are implemented is shown. System 3 may be embodied as a device including the various components described below and configured to perform one or more aspects and embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 3 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, system 3 includes a processing module 30 that implements an inverse tone mapping method. In various embodiments, system 3 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.
[0101] Inputs to processing module 30 may be provided by various input modules as shown in box 32. Such input modules include, but are not limited to: (i) a radio frequency (RF) module that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) a component (COMP) input module (or a set of COMP input modules); (iii) a universal serial bus (USB) input module; and / or (iv) a high-definition multimedia interface (HDMI) input module. Figure 3 Other examples not shown include composite video.
[0102] In various embodiments, the input module of block 32 has associated corresponding input processing elements as known in the art. For example, the RF module may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) re-band-limiting the signal to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF module of various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions among these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box implementation, the RF module and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various implementations rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various implementations, the RF module includes an antenna.
[0103] Additionally, the USB and / or HDMI modules may include corresponding interface processors for connecting System 3 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processing module 30. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within processing module 30. Demodulation, error correction, and demultiplexing streams are provided to processing module 30.
[0104] Various components of System 3 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected using suitable connection arrangements (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards) and data can be transferred between these components. For example, in System 3, processing module 30 is interconnected with other components of System 3 via bus 305.
[0105] The communication interface 304 of the processing module 30 allows the system 3 to communicate on the communication channel 2. For example, the communication channel 2 can be implemented in a wired and / or wireless medium.
[0106] In various implementations, data is streamed or otherwise provided to system 3 using a wireless network such as Wi-Fi, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). In these implementations, the Wi-Fi signal is received via a communication channel 2 and a communication interface 304 suitable for Wi-Fi communication. The communication channel 3 in these implementations is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other cloud-based communications. Other implementations use a set-top box that transmits data via an HDMI connection to input block 32 to provide streaming data to system 3. Still other implementations use an RF connection to input block 32 to provide streaming data to system 3. As mentioned above, various implementations provide data in a non-streaming manner. Additionally, various implementations use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0107] System 3 can provide output signals to various output devices, including a display 5, speakers 6, and other peripheral devices 7. The display 5 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 5 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. The display 5 can also be integrated with other components (e.g., in a smartphone) or standalone (e.g., an external monitor for a laptop). The display device 5 is HDR content compatible. In various examples of embodiments, other peripheral devices 7 include one or more of a standalone digital video disc (or digital universal disc, both terms being DVR), an optical disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 7 that provide functionality based on the output of System 3. For example, an optical disc player performs the function of playing the output of System 3.
[0108] In various embodiments, control signals are transmitted between system 3 and display 5, speaker 6, or other peripheral devices 7 using signaling protocols such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices are communicatively coupled to system 3 via dedicated connections through corresponding interfaces 33, 34, and 35. Alternatively, output devices can be connected to system 3 via communication interface 304 using communication channel 2. Display 5 and speaker 6 may be integrated into a single unit with other components of system 3 in electronic devices such as televisions. In various embodiments, display interface 5 includes a display driver, such as, for example, a timing controller (TCon) chip.
[0109] For example, if the RF section of input 32 is part of a separate set-top box, then display 5 and speaker 6 may optionally be separate from one or more other components. In various embodiments where display 5 and speaker 6 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0110] Various implementations involve the application of inverse tone mapping methods. As used in this application, inverse tone mapping may encompass all or part of a process performed, for example, on a received SDR image or video stream, to produce a final HDR output suitable for display. In various implementations, such a process includes one or more processes typically performed by an image or video decoder, such as a JPEG decoder or an H.264 / AVC (ISO / IEC 14496-10 – MPEG-4 Part 10, Advanced Video Coding), H.265 / HEVC (ISO / IEC 23008-2 – MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265) or H.266 / VVC (Various Video Coding) decoder.
[0111] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding devices are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.
[0112] The specific embodiments and aspects described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of specific embodiment (e.g., discussed only as a method), specific embodiments of the discussed features may be implemented in other forms (e.g., apparatus or program). Apparatus may be implemented, for example, in suitable hardware, software, and firmware. Methods may be implemented, for example, in a processor, which generally refers to a processing device.
[0113] The processing device includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0114] The reference to "an implementation scheme" or "implementation scheme" or "a specific implementation" or "specific implementation," and other variations thereof, means that the specific features, structures, characteristics, etc., described in connection with the implementation scheme are included in at least one implementation scheme. Therefore, the appearance of the phrase "in an implementation scheme" or "in an implementation scheme" or "in a specific implementation" or "in a specific implementation," and any other variations appearing throughout this application, do not necessarily refer to the same implementation scheme.
[0115] Additionally, this application may involve "determining" various types of information. Determining information may include, for example, estimated information, calculated information, predicted information, information retrieved from memory, or information obtained, for example, from another device, module, or user, one or more of these.
[0116] Furthermore, this application may relate to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or more of these.
[0117] Furthermore, this application may relate to "receiving" various types of information. Like "access," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or more. Moreover, "receiving" typically involves one or more of the following during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0118] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” “at least one of A and B,” and “one or more of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one,” “one or more” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As a further example, in the cases of “A, B, and / or C,” “at least one of A, B, and C,” and “one or more of A, B, and C,” such phrases are intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many of the listed items as possible.
[0119] It will be apparent to those skilled in the art that specific embodiments or implementations can produce various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the embodiments or implementations. For example, a signal may be formatted to carry an HDR image or video sequence of the embodiment. Such signals may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the HDR image or video sequence as an encoded video stream and modulating a carrier using the encoded stream. The information carried by the signal may be, for example, analog or digital information. It is known that signals can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.
[0120] Figure 4 A high-level representation of various first implementations of the inverse tone mapping method is schematically shown. Figure 4 The purpose of the implementation plan is to comply with the diffuse white constraint.
[0121] exist Figure 4 In this process, the inverse tone mapping method is executed by the processing module 30.
[0122] In step 40, processing module 30 obtains the current input SDR image. The current input SDR image is a still image or a photograph of a video sequence. Hereinafter, it is assumed that the gain function G() (as shown in equations 2 and 3) has been defined for the current input SDR image. The goal of at least one of the following embodiments is to modify the gain curve corresponding to this gain function G(). (or equivalently modify the gain function G()) to suit the diffuse white constraint.
[0123] In step 41, processing module 30 calculates a histogram representing the current input SDR image (e.g., on the current input SDR image or directly on a filtered version of the current input SDR image). The histogram of the current input SDR image is used to identify regions of interest, hereinafter referred to as lobes. The current input SDR image is assumed to be γ-polarized (non-linear).
[0124] In the implementation scheme, the bar chart includes a number of bins called nbOfBins, where nbOfBins is an integer, such as a multiple of "64". For example, nbOfBins = 256.
[0125] As an example, in the remainder of the document, ITMO's target LMax (i.e., the target maximum brightness value) is "1000" nits, and the current input SDR image is assumed to be a "10"-bit image, where the value "1023" corresponds to 100 nits. Note that "10" bits have been chosen to illustrate the method, but if an "8"-bit image were used, a simple scaling of "4" should be applied.
[0126] In this case, the ITMO function can be written as follows:
[0127]
[0128] Where Y′ SDR Y′ is the brightness value of the current input SDR image, and Y′ is the brightness value of the input SDR image. HDR This is the brightness value of the output HDR image. Brightness value Y′ SDR (Regardless of its bit depth) it is normalized within the range [0; 255]. Similarly, if LMax is “1000” nits, then the luminance value Y′ HDR (Regardless of its number of bits) are normalized within the range [0; 1000].
[0129] Y′ SDR and Y′ HDR Both are gamma-ized, and Y SDR and Y HDR Both are linear, for example:
[0130] Y SDR =(Y SDR ' / 255) 2.4 ×100
[0131] Y HDR =(Y HDR ( / 1000) 2.4 ×1000
[0132] In step 42, processing module 30 obtains the gain function G(). Once obtained, the gain function G() allows the ITM curve to be obtained from the ITM function in Equation 4.
[0133] Figure 9A , Figure 9B and Figure 9C Three examples are shown, representing the ITM curve for a “1000” nit display as obtained using Equation 4. The SDR input is in the range [0…255] (meaning it must be normalized within the range [0; 255] if it is not an “8” bit image), and the output is in the range [0; 1000]. Figure 9AThe curve shows that the maximum input value "255" (corresponding to "100" nits SDR) produces an output of "1000" nits, which is equal to "1000" nits when linearized. Figure 9B The curve shows a maximum value of approximately "700", which corresponds to "425" nits when linearized, and Figure 9C The curve shows a maximum value of approximately "1200", which corresponds to "1550" nits when linearized.
[0134] Available from Figure 9A , Figure 9B and Figure 9C As can be seen, if the SDR image is a white image (if the SDR image is a "10" bit image, then the brightness value of all pixels is "1023"), then it corresponds to Figure 9A The MaxFall of the HDR image is "1000" nits, corresponding to Figure 9B The MaxFall value of the HDR image is "425" nits, and corresponds to... Figure 9C The MaxFall value for the HDR image is “1550” nits.
[0135] In step 43, the processing module 30 applies a search process designed to identify regions of lobes (i.e., bright areas of interest) in the output HDR image generated by the inverse tone mapping method from the current input SDR image.
[0136] It can be observed that a significant amount of light in the output HDR image can be generated by a large number of input pixels with medium brightness values (e.g., "180" versus maximum input brightness "255") or a smaller number of input pixels with high brightness values (i.e., approximately "250"). For example (using LMax = "1000" nits), if we assume the gain function G() = 1.25 (regardless of input brightness):
[0137] ·Y HDR (180) = 368 nits;
[0138] ·Y HDR (255) = 1046 nits.
[0139] Then, the region centered at "255" will produce the same amount of light as the region centered at "180", but with up to three times more pixels (in fact, 1046 = 2.8 * 368). This means that the method must search the histogram for lobes and pixel lobes (or group lobes) in the output HDR image.
[0140] The following text is about Figure 5 Step 43 is described in detail.
[0141] In step 400, the processing module determines whether the diffuse white constraint is observed. Step 400 includes sub-steps 44, 45, and 46.
[0142] In step 44, processing module 30 calculates the luminance value YposDW representing at least one of the identified regions in the currently input SDR image. The following section discusses... Figure 6 The implementation plan for step 44 is described in detail.
[0143] In step 45, processing module 30 calculates the extended luminance value YexpDW of the luminance value YposDW. The following section discusses... Figure 7 Step 45 is described in detail.
[0144] In step 46, the processing module uses the extended luminance value YexpDW to determine whether the diffuse white constraint is observed. To do this, the extended luminance value YexpDW is compared with the diffuse white constraint value DWTarget. The diffuse white constraint value DWTarget is, for example, a predefined value that depends on the display device intended to display an HDR image corresponding to the currently input SDR image and / or the ambient light in the room where the HDR image is displayed and / or parameters given by the user of the expected HDR image to be displayed.
[0145] If the extended luminance value YexpDW is less than or equal to the diffuse white constraint value DWTarget, the gain curve corresponding to the gain function G() (or equivalently the gain function G()) is not modified, and the output HDR image is generated using the ITMO function of Equation 4 in step 47. Otherwise, if the extended luminance value YexpDW is greater than the diffuse white constraint value DWTarget, the modified gain curve is defined in step 48. (That is, a modified version of the gain curve G obtained by the gain function G() or equivalently a modified gain curve G′ obtained by the modified gain curve G'()) and using the modified gain curve Modify the luminance value Y′ obtained through the ITMO function in Equation 4. HDR To generate the output HDR image. The following section discusses... Figure 8 Step 48 is described in detail.
[0146] Figure 5 A detailed implementation of step 43 of the inverse tone mapping method is illustrated schematically.
[0147] In step 4300, the processing module calculates the contribution (hereinafter also referred to as contrib or energy) of each band in the histogram. A band is a set of consecutive bins in the histogram. In the example of a histogram containing “256” bins, each band includes four bins, so the histogram is divided into “64” bands of equal size. If the current input SDR image is encoded in “10” bits, each bin contains four luminance values Y′.SDR The first band of the histogram collects the input brightness values Y′ from “0” to “15”. SDR The 32nd band collects the input luminance values Y′ from “496” to “511”. SDR Furthermore, the 64th band collects the input luminance values Y′ from “1008” to “1023”. SDR The contribution of the band represents the light energy after applying the ITMO function of Equation 4 corresponding to that contribution, i.e., the energy emitted by the pixel represented by that band. The nth contrib is then defined as:
[0148] Contrib[n]=(∑ i (∑ j Y HDR (j))×histo[i]) / (sumOfBins×A) (Equation 5)
[0149] Where n varies from "0" to "63", and where:
[0150] • i is an integer that varies from n×nbOfBins / 64 to (n+1)×nbOfBins / 64-1.
[0151] • A = (Ymax + 1) / (nbOfBins), where Ymax is the maximum possible code value of the input image;
[0152] • j is within the range [A*i; A*(i+1)–1];
[0153] • nbOfBins is the number of bins in the bar chart (in the example, = 256).
[0154] ·Y HDR The output brightness value Y′ is calculated using Equation 4. HDR A linearized version.
[0155] ·sumOfBins=∑histo[k] where k is an integer ranging from 0 to nbOfBins-1.
[0156] sumOfBins represents the total number of pixels in the histogram that has been calculated.
[0157] When the input is encoded in "8" bits, equation 5 above can be simplified as follows:
[0158] Contrib[n]=∑Y HDR (j)×histo[i] / sumOfBins (Equation 5')
[0159] The number n varies from "0" to "63".
[0160] In a variant of step 4300, contributions are calculated only for a subset of the bands in the histogram (e.g., for one band out of two or one band out of three).
[0161] In step 4301, processing module 30 calculates the population (contribPop) for each contribution:
[0162] contribPop[n]=∑histo[i] (Equation 6)
[0163] Where i is an integer ranging from n×nbOfBins / 64 to (n+1)×nbOfBins / 64-1.
[0164] In step 4302, processing module 30 identifies local maxima representing high-energy regions within a set of contributions. A contribution contrib[n] of band n is considered a local maxima if its value is greater than the values of the contributions of its two adjacent bands (i.e., typically band n-1 and band n+1). In other words, a contribution contrib[n] is a local maxima if the following conditions are met:
[0165] contrib[n]>contrib[n-1] and contrib[n]>contrib[n+1]
[0166] It should be noted that for the highest-ranked contribution (here, contrib
[63] ), the condition is contrib
[63] > contrib
[62] . In the following text, the band corresponding to the contribution representing the local maximum is called the candidate.
[0167] In step 4303, processing module 30 applies an aggregation process to the candidates. During the aggregation process, the band corresponding to each candidate is aggregated with the maximum value of the LN neighboring bands on the left and the maximum value of the RN neighboring bands on the right (making the maximum total width LN+RN+1):
[0168] For the LN adjacent bands on the left, processing module 30 searches for the minimum ln value in [1; LN].
[0169] It satisfies:
[0170] οcontrib[n-ln]>0.2×contrib[n], and;
[0171] οcontrib[n-ln]<1.025×contrib[n-ln+1].
[0172] If the value is not found, then ln = 0.
[0173] Similarly, for the RN adjacent bands on the right, processing module 30 searches for the minimum rn value in [1; RN] that satisfies:
[0174] οcontrib[n+rn]>0.2×contrib[n], and;
[0175] οcontrib[n+rn]<1.025×contrib[n+rn-1].
[0176] Where ln + rn ≤ 63. If this value is not found, then rn = 0.
[0177] Aggregation is performed between n-ln and n+rn.
[0178] In the implementation plan, LN = RN = 5.
[0179] The aggregation of bands around the candidate is facilitated by the fact that large bright areas (such as clouds in the sky, close-ups of pages in a book, people wearing bright clothes, white animals, etc.) are not uniformly white, but usually exhibit some dispersion around the middle value, which can be characterized by lobes with a certain width.
[0180] Note that the values “0.2” and “1.025” are examples and can be changed to other close values.
[0181] In step 4304, processing module 30 stores information representing each candidate. In one embodiment, the stored information includes the lowest and highest positions of the aggregation bands in a histogram, and the sum of the energy contributions of the aggregation bands named as candidates.
[0182] In step 4305, processing module 30 searches for local maxima in the population contribPop (calculated in step 4301) associated with the contribution. The population contribPop[n] is a local maximum if it satisfies the following condition:
[0183] contribPop[n] > contribPop[n-1] and contribPop[n] > contribPop[n+1]
[0184] It should be noted that for the highest-ranked group (here, contribPop
[63] ), the condition is that contribPop
[63] > contribPop
[62] .
[0185] In step 4306, processing module 30 applies the aggregation process to the band (referred to as candidatePop) corresponding to the local maximum value identified in the population contribPop. The aggregation process applied to candidatePop during step 4306 is the same as the aggregation process applied to the candidate during step 4303.
[0186] In step 4307, processing module 30 verifies whether at least one aggregation candidate group, candidatePop, is independent of any identified aggregation candidate. Independence means that the aggregation candidate group does not share any contribution position with any identified aggregation candidate.
[0187] If at least one independent aggregated candidate group (candidatePop) exists, in step 4308, the processing module 30 terminates aggregated candidate identification by designating each independent aggregated candidate group (candidatePop) as an aggregated candidate; that is, for each independent aggregated candidate group, the processing module 30 creates an aggregated candidate from that independent aggregated candidate group. This can occur when the energy increases slightly to a maximum energy value that is further away from higher brightness values. However, this region can be embedded around the candidate group (candidatePop) and potentially represent a large number of pixels with a large amount of energy.
[0188] In step 4309, once all aggregation candidates have been identified, processing module 30 selects N_cand aggregation candidates with the highest energy from the identified aggregation candidates. In other words, in step 4309, the processing module selects N_cand aggregation candidates that include the pixels emitting the most light energy. In an embodiment, the number N_cand = 5. For each of the N_cand selected aggregation candidates, processing module 30 stores information representing that selected aggregation candidate, including, for example:
[0189] • The lowest position of the aggregation candidate is named loPos (i.e., corresponding to the value n-nr defined above);
[0190] • The highest position of the aggregation candidate is named hiPos (i.e., corresponding to the value n+h defined above);
[0191] • The position of the center of the aggregated candidates, named maxEnergyPos, which is the position of the candidate itself (i.e., the value n as defined above);
[0192] • Information representing the energy of the aggregation candidate, which is named energy (i.e., the sum of all contributions in the aggregation candidate);
[0193] · Aggregate the candidates of the population (i.e., divide the sum of all the bins of the histogram included in the aggregated candidates from position loPos to position hiPos by sumOfBins), which is named population.
[0194] In step 4310, the processing module 30 selects N_cand_Max_Pop (N_cand_Max_Pop > 1 and N_cand_Max_Pop < N_cand) aggregated candidates with the largest population among a set of N_cand selected aggregated candidates.
[0195] In an embodiment, N_cand_Max_Pop = 2. In this case, if the two aggregated candidates (hereinafter denoted as CP1 and CP2) selected in step 4310 overlap, i.e.:
[0196] · loPos[CP1] < hiPos[CP2] if maxEnergyPos[CP1] > maxEnergyPos[CP2]; or,
[0197] · loPos[CP2] < hiPos[CP1] if maxEnergyPos[CP1] < maxEnergyPos[CP2];
[0198] Merge the two aggregated candidates (i.e., merge the energy and the population), and the processing module 30 selects the aggregated candidate with the third population among the N_cand selected aggregated candidates.
[0199] As can be seen, the processing module 30 selects N_cand_Max_Pop largest regions in terms of population (i.e., the number of pixels) among the N_cand largest regions in terms of energy when applying the Figure 5 procedure: This follows the above concept that a large amount of light in the output HDR image can be generated by a large number of input pixels. By finding the N_cand first aggregated candidates in terms of energy, the processing module 30 obtains N_cand main lobes, and then selects N_cand_Max_Pop main lobes that collect the largest number of pixels. Note that the N_cand first energy lobes can be small (and even non-existent) and the N_cand_Max_Pop first populations: this would be the case if the input image is a dark image.
[0200] As will be described below, if a TMO function with a gain function G() is applied to the current input SDR image, the N_cand_Max_Pop aggregated candidates are used to determine whether the output HDR image does not comply with the risk of the diffuse white constraint. Therefore, the N_cand_Max_Pop aggregated candidates are used to determine when to modify the gain function G() (or equivalently, the gain curve obtained using the gain function) This ensures that the output HDR image adheres to a predefined light energy constraint (diffuse white).
[0201] Figure 6 A detailed implementation of step 44 of the inverse tone mapping method is illustrated schematically.
[0202] In step 4400, the processing module 30 determines the luminance value YposDW representing the selected N_cand_Max_Pop candidate with the largest population.
[0203] When N_cand_Max_Pop = 2, the processing module 30 applies the following algorithm to determine the luminance value YposDW:
[0204] • In the band corresponding to position maxEnergyPos (where the histogram includes “64” consecutive bands of equal size and “256” bins, as in the previous example containing four bins), search for the highest input brightness value Y in the bin of the histogram. in And name the input brightness value firstPopY.
[0205] • In the band corresponding to the position maxEnergyPos of CP2, search for the highest input brightness value Y in the bin in the histogram. in And name the input brightness value secondPopY.
[0206] The brightness value YposDW is determined as follows:
[0207] YposDW=firstPopY+EnByPopCoef×(secondPopY-firstPopY) (Equation 7)
[0208] in
[0209] οEnByPopCoef=
[0210] secondEnByPop / (firstEnByPop+secondEnByPop);
[0211] οfirstEnByPop=(energy[CP1]×population[CP1]) 0.5 ;
[0212] οsecondEnByPop=(energy[CP2]×population[CP2]) 0.5 ;
[0213] Therefore, the luminance value YposDW is located somewhere between the two largest pixel regions in the pixel region with the highest energy.
[0214] In the above embodiment of step 4400, the parameter EnByPopCoef (and therefore the luminance value YposDW) depends on the square root of the product of energy and N_cand_Max_Pop (=2) candidate populations. In Equation 7 above, the population and energy have the same weight.
[0215] In another embodiment of step 4400, the group is given more weight as follows:
[0216] ·firstEnByPop=energy[CP1] 0.25 ×population[CP1] 0.75 ;
[0217] ·secondEnByPop=energy[CP2] 0.25 ×population[CP2] 0.75 ;
[0218] In another embodiment of step 4400, energy is given more weight as follows:
[0219] ·firstEnByPop=energy[CP1] 0.75 ×population[CP1] 0.25 ;
[0220] ·secondEnByPop=energy[CP2] 0.75 ×population[CP2] 0.25 ;
[0221] When N_cand_Max_Pop = 1, that is, when only one candidate CP1 with the largest population is determined from a set of N_cand selected candidates, the processing module 30 applies the following algorithm to determine the luminance value YposDW:
[0222] • In the band corresponding to the position maxEnergyPos of CP1, search for the highest input brightness value Y in the bin in the histogram. in And name the input brightness value firstPopY.
[0223] The brightness value YposDW is determined as follows:
[0224] YposDW = firstPopY (Equation 7bis)
[0225] Optionally (e.g., in an implementation suitable for extracting the current input SDR image from a video sequence), in step 4401, processing module 30 applies a temporal filter to the luminance value YposDW. The purpose of optional step 4401 is to attenuate (or even cancel) small luminance variations (or oscillations) between two consecutive HDR images. The temporal filtering process includes calculating a weighted average between the luminance value YposDW and a luminance value representing the luminance value YposDW calculated for an image preceding the current input SDR image and labeled as recursiveYposDW. The luminance values YposDW and recursiveYposDW are calculated as follows:
[0226] recursiveYposDW=DWFeedBack×recursiveYposDW+(1–DWFeedBack)×YposDW; (Equation 8)
[0227] YposDW = recursiveYposDW;
[0228] DWFeedBack is a weight within the range [0; 1]. In one implementation, DWFeedBack = 0.9. In another implementation, DWFeedBack depends on the frame rate of the video sequence. The higher the frame rate, the higher the weight of DWFeedBack. For example, for a frame rate of "25" frames per second (Im / s), DWFeedBack = 0.95, while for a frame rate of "100" Im / s, DWFeedBack = 0.975.
[0229] Note that if the current input SDR image corresponds to a scene change (i.e., a portion of the video sequence that is not uniform in terms of content compared to the image preceding the current input SDR image), then filtering should not be applied.
[0230] It should be noted that the luminance value YposDW obtained in step 4400 (or step 4401, if applicable) is a floating-point number in the range [0; 255] and can be easily scaled to the bit depth of the current input SDR image.
[0231] YposDW = YposDW × (Ymax / 255)
[0232] The input video's Ymax = 2 n -1 is encoded in n bits (i.e., if n = 10, then Ymax = 1023).
[0233] Figure 7 The details of step 45 of the inverse tone mapping method are illustrated schematically.
[0234] In step 4500, processing module 30 calculates an extended value Y′ corresponding to the luminance value YposDW as follows HDR :
[0235]
[0236] And, in step 4501, linearization is applied to the obtained value:
[0237] YexpDW = (Y′ HDR (YposDw) / LMax) 2.4 ×LMax
[0238] Figure 8 FIG. schematically shows the details of step 48 of the inverse tone mapping method. As a reminder, step 48 is executed when the processing module determines in step 46 that there is a risk of generating an output HDR image including too bright regions when applying the inverse tone mapping to the current input SDR image using the gain function G(). The purpose of step 48 is to reduce the luminance of such regions in the HDR image.
[0239] In step 4800, processing module 30 determines the selected N_cand_Max_Pop candidate population DWpopulation with the largest population.
[0240] When N_cand_Max_Pop = 2, DWpopulation = population(CP1)+population(CP2) (Equation 10).
[0241] When N_cand_Max_Pop = 1, DWpopulation = population(CP1) (Equation 10bis).
[0242] The higher DWpopulation is, the closer YexpDW must be to DWTarget. For example, a bright sun in the sky that is "1%" of the image size must be ignored (in fact, in this case, for example, there is no risk of dazzling the user watching the video), while a bright ice rink that is "60%" of the image size can be locked to DWTarget.
[0243] Then two parameters are introduced:
[0244] · Parameter loThresholdPop: If DWpopulation < loThresholdPop, the gain at YposDW remains as it is. This means that if DWpopulation represents less than loThresholdPop% of the total number of pixels, the gain curve is not modified (or equivalently the gain function G()) and apply step 47. In an embodiment, loThresholdPop is a predefined value. For example, loThresholdPop = 5%.
[0245] · A parameter Dwsensitivity within the range [0…1]: Dwsensitivity defines another threshold hiThresholdPop above which the gain curve at the luminance value YposDW is modified (or equivalently the gain function G()) such that YexpDW = DWTarget.
[0246] The variables loThresholdPop and Dwsensitivity are used to calculate the variable modDWpopulation:
[0247] modDWpopulation = (DWpopulation – loThresholdPop) / (0.65 – 0.4 × Dwsensitivity) (Equation 11) where modDWpopulation is limited to the range [0;1].
[0248] As can be seen, if DWpopulation < loThresholdPop, then modDWpopulation = 0. Thus, modDwpopulation = 0 indicates that there is no need to modify the gain curve (or equivalently the gain function G()) to obtain the HDR image from the current input SDR image.
[0249] Regarding Dwsensitivity:
[0250] · Dwsensitivity = 0 causes hiThresholdPop = 0.7: DWpopulation must represent at least 70% of the pixels of the image to ultimately achieve YexpDW = DWTarget. In this case, the inverse tone mapping method and especially the process of modifying the gain curve (or equivalently the gain function G()) has low reactivity to DWpopulation.
[0251] · Dwsensitivity = 0.5 causes hiThresholdPop = 0.5. DWpopulation must represent at least 50% of the pixels of the image to ultimately achieve YexpDW = DWTarget.
[0252] • DWsensitivity = 1 results in hiThresholdPop = 0.3. DWpopulation must represent at least 30% of the image's pixels to ultimately achieve YexpDW = DWTarget. In this case, the inverse tone mapping method is used, and specifically the gain curve is modified. The process (or equivalently the gain function G()) is highly responsive to DW population.
[0253] In step 4801, processing module 30 derives variable DWrate from variable modDWpopulation:
[0254] DWrate = mod DWpopulation 1 / p Where p≥1. (Equation 12)
[0255] In step 4802, processing module 30 determines the value YexpDWTarget as follows:
[0256] YexpDWTarget=DWrate*DWTarget+(1–DWrate)*YexpDW (Equation 13)
[0257] Where YexpDWTarget and YexpDW are linear values.
[0258] Therefore, the higher the DWpopulation, the closer the expanded value of the luminance value YposDW is to the diffuse white target DWtarget.
[0259] When N_cand_Max_Pop = 2, if the population of CP1 is much larger than that of CP2, the luminance value YposDW is closer to the position firstPopY, and then the HDR value of the pixel inside lobe CP1 is closer to the diffuse white target DWtarget, depending on the size of lobe CP1.
[0260] As can be seen:
[0261] • If DWrate = 0, then YexpDWTarget equals the luminance value YexpDW, which means there is no need to modify the gain curve. (or equivalently, the gain function G()). In this case, step 47 is performed by processing module 30.
[0262] • If DWrate = 1, then YexpDWTarget equals the diffuse white target DWTarget, which means the gain curve must be modified. (or equivalently, the gain function G()) is required to obtain the luminance value YposDW's DWTarget.
[0263] In step 4803, processing module 30 converts the value YexpDWTarget into a γ-value:
[0264] YexpDWTarget′=(YexpDWTarget / LMax) 1 / 2.4 ×LMax
[0265] In step 4804, the processing module calculates the gain gainAtDWTarget corresponding to the luminance value YposDW:
[0266] gainAtDWTarget=log(YexpDWTarget') / log(255 / Ymax×YposDW) (Equation 14)
[0267] In step 4805, the processing module uses the gain curve obtained by the gain function G() Modify the entire curve to obtain the new gain curve. Several variations of step 4805 are possible:
[0268] 1. Modify the gain curve only for high input brightness values. At the same time ensure Y′ HDR The curve (obtained using Equation 4) is monotonic. However, if multiple high input brightness values take the same output brightness value, this implementation can achieve a higher output brightness value in Y′. HDR A clipping mechanism is introduced into the high brightness value of the curve;
[0269] 2. By analyzing the gain curve Subtract a constant to modify the gain curve for all input values. At the same time ensure Y′ HDR The curve does not produce overly dark images;
[0270] 3. A hybrid of the two primary solutions.
[0271] Variant 1 can be considered as a high level of Y in For the gain curve Compression, thus producing only the highest Y′ HDR Horizontal compression is generated. In embodiment 1 of step 4805, the modified gain curve can be obtained using the following equation.
[0272] gainMod(Y′)=G(Y′)-HlCoef×Y′ / (Ymax) HlExp (Equation 15)
[0273] gainMod(Y’) means “modified gain of γ - transformed input luminance value Y’”, HlExp is called the high - level exponent, and HlCoef is called the high - level coefficient.
[0274] In an embodiment, the high - level exponent HlExp = 6, but it can be decreased while remaining higher than “2”. In an embodiment, the high - level coefficient HlCoef is within the range [0; 0.3]. HlCoef = 0 means no compression is applied. In the case of HlExp = 6, HlCoef is calculated by using the known gain value G(YposDW) and gainMod(YposDW) in Equation 15.
[0275] Then, the concept of contrast is introduced in the form of a compressed gain curve.
[0276] · In Variant 1, the contrast is “0” (minimum contrast). Equation 15 is applied on the condition that after compression, and for a “10” - bit current input SDR image, the extended output of luminance value “1023” is higher than the extended output of luminance value “1020”:
[0277]
[0278] If this is not the case, the HlExp in Equation 15 is recursively decreased using a given reduction parameter RedParam (e.g., RedParam = 0.05), which results in a higher value of HlCoef until Equation 16 is true. Then, if HlExp < 2, HlExp is clipped to “2”, or if HlCoef < 0.3, HlCoef is clipped to “0.3”, and the value c0 is calculated such that:
[0279] gainMod(YposDW)=G(YposDW)-HlCoef×(YposDW / Ymax)^HlExp + C0 (Equation 17)
[0280] c0 must be negative. If HlExp ≥ 2 and HlCoef ≥ 0.3, then c0 is equal to “0”. Note that other values can be chosen for the maximum value of HlCoef and for the starting and minimum values of HlExp.
[0281] · In Variant 2, the contrast is 1 (maximum contrast). Equation 15 is applied on the condition that after compression, and for a “10” - bit current input SDR image, the gain of luminance value “1023” is higher than the extended output of luminance value “1020”:
[0282] gainMod(1020)<gainMod(1023) (Equation 18)
[0283] If this is not the case, use Equation 15 to find the new value HlCoef1, for which gainMod(1020) = gainMod(1023). HlCoef1 is lower than the previous value HlCoef. Then use HlCoef1 and YposDW in the following Equation 19 to calculate the value c1 (HlExp1 = 6):
[0284] gainMod(YposDW)=G(YposDW)-
[0285] HlCoef1×(YposDW / Ymax) HlLExp1 +C1 (Equation 19)
[0286] c1 is a negative number. If equation 18 is verified, then c1 equals "0".
[0287] Variant 3 can be implemented using contrast "0" and "1" by introducing new values for contrast in the range [0...1]. Then, c and Hlcoef are calculated as follows (Equation 20):
[0288] c=contrast×c1+(1-contrast)×c0
[0289] Hlcoef=contrast×Hlcoef1+(1-contrast)×Hlcoef0
[0290] And then Equation 17 can be used to find HlExp at position YposDW:
[0291] gainMod(YposDW)=G(YposDW)-HlCoef×(YposDW / Ymax) HlExp +C
[0292] And then:
[0293] c=log((G(YposDW)+c-gainMod(YposDW)) / HlCoef) / log(YposDW / Ymax) (Equation 21)
[0294] The prerequisite is that YposDW is not Ymax. If YposDW = Ymax, then the value of Ymax-1 can be found.
[0295] To avoid the pumping effect, HlCoef, HlExp, and c can be filtered in time in the same way as YposDW (i.e., in step 4401) (but possibly with different feedback values). Another, and even better, approach to time filtering is to apply feedback to the gain function G() itself. In both cases, the time filtering is reset at each shearing position in the same way as YposDW.
[0296] So far, there has been information about Figures 4 to 8 An implementation scheme for handling targets that comply with the diffuse white constraint is described. A similar implementation scheme for handling targets that comply with the MaxFall constraint is shown below.
[0297] The MaxFall MF of an image can be defined as follows:
[0298] MF=(∑max(Rp,Gp,Bp)) / nbOfPixelsInTheImage
[0299] Rp, Gp, and Bp are the three linear color component values of pixel P, and nbOfPixelsInTheImage is the number of pixels in the image.
[0300] In the following text, this definition is approximated by the following formula (which is quite accurate because the present invention deals with bright areas, followed by areas of high or very high brightness, which assumes that at least two of the three RGB color component values are close):
[0301] MF=(∑Y P ) / nbOfPixelsInTheImage
[0302] Where Y P It is the γ-luminance Y′ of pixel P. P The obtained linearized brightness:
[0303]
[0304] Where Y′ P It is a value within the range [0; 1], and LMax is the peak nit of the target display. Y′ P It can also be a code value encoded in n bits, and in this case:
[0305] Where Y′ max =2 n-1
[0306] For example, the SDR input gamma-enhanced luminance value Y′ encoded in "8" bits. P=100 has a linear value of "10.6" nits. The HDR output brightness value Yexp′ is encoded in "10" bits and generated for a display device with "1000" cd / m2 (nominal peak brightness). P =400 has a linear value of "105" nits.
[0307] In the following text, the value MFTarget represents the maximum value that MaxFall can take for the extended image. This means that the calculated MaxFall of the extended image must be lower than MFTarget. If the calculated MaxFall is higher than MFTarget, the gain function is modified to reach that value. The MaxFall constraint value MFTarget is, for example, a predefined value that depends on the display device intended to display an HDR image corresponding to the current input SDR image and / or the ambient light in the room where the HDR image is displayed and / or parameters given by the user who expects the HDR image to be displayed.
[0308] The calculation of MaxFall for the extended image is simplified by using the content of a bar chart, i.e., MaxFallOut:
[0309] MaxFallOut=(∑ i (∑ j Y HDR (j))×histo[i]) / (sumOfBins×A) (Equation 22)
[0310] in:
[0311] • i is within the range [0; nbOfBins];
[0312] • A = (Ymax + 1) / (nbOfBins), where Ymax is the maximum possible code value of the input image;
[0313] • j is within the range [A*i; A*(i+1)–1];
[0314] ·Y HDR(j) It is linear brightness (expressed in nits).
[0315] As can be seen, MaxFallOut equals the sum of “64”contrib[n] (denoted as (∑contrib[n]) as defined in Equation 5) (i.e. MaxFallOut=∑contrib[n]).
[0316] Figure 10 A high-level representation of various second implementations of the inverse tone mapping method is schematically shown. Figure 10 The purpose of the implementation plan is regarding the MaxFall constraint.
[0317] Figure 10 The method is executed by the processing module 30.
[0318] Figure 10 Methods and Figure 4 The difference in the method is that steps 43, 400, and 48 are replaced by steps 43bis, 400bis, and 48bis, respectively. The other steps remain the same.
[0319] Compared to step 43, steps 4301, 4305, 4306, 4307, 4308, and 4310 are not executed in step 43bis. In step 4309, processing module 30 selects the aggregation candidate with the highest energy from the identified aggregation candidates.
[0320] In step 400bis, processing module 30 compares the value of MaxFall representing the extended HDR image with the MaxFall constraint MFTarget. In step 400bis, it is assumed that the sum of the contributions contrib[n] calculated during step 4300, ∑contrib[n], represents the MaxFall of the extended HDR image.
[0321] If the sum of all contributions ∑contrib[n] is less than or equal to MFTarget, then there is no need to modify the gain function G() (or equivalently the gain curve). (That is, perform step 47).
[0322] In contrast, if the sum of all contributions ∑contrib[n] is higher than MFTarget, then during step 48bis, processing module 30 uses the candidate CP1 with the highest energy (as per [reference]). Figure 5 (As defined) Search for the luminance value Y with the highest energy among candidate CP1. in (hereinafter referred to as Y) pos And modify the value Y corresponding to the brightness value as follows. pos Extended gain:
[0323] gainAtMFTarget=G(Y pos )+log(MFTarget / MaxFallOut) / (2.4*log(255×Y pos / Ymax)) (Equation 23)
[0324] Y pos Ymax is the codeword that is then gamma-ized. If the histogram contains "256" bins, then Y... pos The maximum input luminance value Y in the bin is CP1's maxEnergyPos. in (It contains the most energy.)
[0325] Processing module 30 then uses equation 15 in Y pos The modified gain gainModMF(Y) was found at the location. pos ):
[0326] gainAtMFTarget=G(Y pos )-HlCoef×(Y pos / Ymax) GlExp =gainAtMFTarget
[0327] Then, the processing module applies the same strategy as that used to comply with the diffuse white constraint regarding the gain curve or equivalently the gain function G() (i.e., using the same concept of contrast), using equations 16 to 22 and replacing gainMod and YposDW with gainModMF and Ypos respectively to obtain the values of HlCoef, HlExp, and c.
[0328] In the implementation scheme, during the inverse tone mapping of the current input SDR image, only... Figure 4 The method focuses on diffuse white constraints.
[0329] In the implementation scheme, during the inverse tone mapping of the current input SDR image, only... Figure 10 The method focuses on MaxFall constraints.
[0330] In the implementation scheme, during the inverse tone mapping of the current input SDR image, the following is applied: Figure 4 and Figure 10 Two methods. In the implementation plan, in Figure 10 Methods applied before Figure 4 The method. In fact, if applied to the current input SDR image... Figure 4 If the method adheres to the diffuse white constraint, the resulting image can still have a MaxFall higher than MFTarget. In this case, the following triplet (HlCoef, HlExp, c) is obtained:
[0331] gainAtDWTarget=G(YposDW)–HlCoef×(YposDW / Ymax) HlExp +c
[0332] Then, change (HlCoef,HlExp,c) to (HlCoef',HlExp',c') to get:
[0333] gainAtMFTarget=G(Y pos )–HlCoef'*(Y pos / Ymax)HlExp′ +c'
[0334] At the same time, follow the rules above regarding contrast.
[0335] The above implementation scheme handles bright areas. Nevertheless, MaxFall detection can be extended to large saturated blue or red areas, which can produce large MaxFall values while their corresponding luminance values (Y) are relatively low. This can be addressed by using blue and red histograms (the blue and red values can be calculated from the Y, U, and V values). When using the 709 color space (as defined in the recommended UIT-R BT 709), the distribution of the R (red), B (blue), and G (green) components is as follows:
[0336] R is 21% of Y
[0337] G is 72% of Y
[0338] B is 7% of Y.
[0339] Large white areas produce lobes in the luminance (Y) histogram. They should also produce lobes located approximately at the same positions in the R and B histograms. In contrast, having large lobes at high values in the R and / or B histograms but no lobes at the same positions in the Y histogram (meaning there are Y lobes of equivalent size for lower Y values) will produce large values of RGB MaxFall (RGB MaxFall is the actual definition of MaxFall here), while the MaxFall calculated on Y will be small or at least smaller. The method used to find candidate populations in Y can be used on R and B, and then matching can be performed between Y aggregates and R and / or B aggregates located at different positions. The energy of those matched population candidates can then be overestimated by taking into account their lack of green, and then the entire set of standard Y candidates and the overestimated Y candidates can be sorted. MaxFall is then calculated at the candidate with the highest energy.
[0340] The foregoing describes several embodiments. Features of these embodiments may be provided individually or in any combination. Furthermore, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination, across various claim classes and types:
[0341] • A television set, set-top box, mobile phone, tablet computer or other electronic device that performs at least one of the described implementation schemes.
[0342] • A television, set-top box, mobile phone, tablet computer, or other electronic device that performs at least one of the described implementation schemes and displays the resulting image (e.g., using a monitor, screen, or other type of display).
[0343] A television, set-top box, mobile phone, tablet computer, or other electronic device that tunes (e.g., uses a tuner) a channel to receive signals including encoded images or video streams and performs at least one of the described embodiments.
[0344] A television, set-top box, mobile phone, tablet computer, or other electronic device that receives signals including encoded images or video streams over the air (e.g., using an antenna) and performs at least one of the described embodiments.
Claims
1. A method for inverse tone mapping, wherein the method comprises: Obtain a gain function for obtaining the inverse tone mapping operator function, thereby allowing a high dynamic range image to be obtained from a low dynamic range image by applying a search procedure to identify regions of the low dynamic range image that produce bright areas in the high dynamic range image when the inverse tone mapping operator function is applied to the low dynamic range image, the search procedure including: Define bands representing sub-parts of the histogram of the low dynamic range image, and obtain the contribution and group of the number of pixels represented by each band, each contribution representing the light energy emitted after applying the inverse tone mapping operator function to the pixels represented by the band; At least one candidate is created based on at least one local maximum in the contribution and at least one local maximum in the population, each candidate corresponding to one or more continuous bands, the one or more continuous bands including a band corresponding to a local maximum in the contribution or the population; At least one final candidate is selected from the at least one candidate based on information representing each candidate, the information including information representing the light energy emitted by the pixel represented by the candidate and information representing the number of pixels represented by the candidate, the information of the selected at least one final candidate representing both the highest light energy and the maximum number of pixels; Furthermore, the final candidate application determination process is used to determine the modification of the gain function to ensure that the high dynamic range image complies with at least one light energy constraint, wherein the modification of the gain function is determined in response to determining that the light energy of the final candidate in the high dynamic range image will be higher than the at least one light energy constraint.
2. The method of claim 1, wherein creating at least one candidate based on at least one local maximum in the contributions and at least one local maximum in the population comprises: At least one local maximum is identified in the contributions, and for each local maximum, the corresponding band is aggregated with the adjacent bands to create a candidate; In the population, at least one local maximum is identified, and for each local maximum, the corresponding band is aggregated with the adjacent bands to create a candidate population; as well as, Create candidates from each candidate group that are independent of any other candidate.
3. The method according to claim 1, wherein the histogram is a histogram of brightness values.
4. The method according to claim 1, wherein the at least one light energy constraint is at least one of a MaxFall constraint and a diffuse white constraint.
5. The method of claim 1, wherein selecting the at least one final candidate comprises: Select a candidate subset associated with the highest value representing light energy, and select at least one final candidate from the candidates of the subset representing the highest number of pixels.
6. The method according to claim 1, wherein the determining process comprises: Determine the final pixel value representing the at least one final candidate; The inverse tone mapping operator function is used to calculate a value representing the extended pixel value from the final pixel value; as well as, Perform a modification process adapted to modify the gain function in response to determining that the extended pixel value is higher than the light energy constraint representing a predefined diffuse white constraint value.
7. The method of claim 6, wherein the low dynamic range image is the current image in the image sequence, and the final pixel value is filtered temporally using at least one final pixel value calculated for at least one image preceding the current image in the image sequence.
8. The method according to claim 1, wherein the determining process comprises: Perform a modification process adapted to modify the gain function in response to determining that the value of MaxFall representing the high dynamic range image is higher than the light energy constraint representing a predefined MaxFall constraint.
9. The method of claim 8, wherein, The value of MaxFall representing the high dynamic range image is the sum of the obtained contributions.
10. An apparatus for inverse tone mapping, wherein the apparatus includes an electronic circuit system adapted to: Obtain a gain function for obtaining the inverse tone mapping operator function, thereby allowing a high dynamic range image to be obtained from a low dynamic range image by applying a search procedure to identify regions of the low dynamic range image that produce bright areas in the high dynamic range image when the inverse tone mapping operator function is applied to the low dynamic range image, the search procedure including: Define bands corresponding to sub-parts of the histogram of the low dynamic range image, and obtain contributions and groups representing the number of pixels, each contribution representing the light energy emitted by the inverse tone mapping operator function after applying the pixel represented by the band; At least one candidate is created based on at least one local maximum in the contribution and at least one local maximum in the population, each candidate corresponding to one or more continuous bands, the one or more continuous bands including a band corresponding to a local maximum in the contribution or the population; At least one final candidate is selected from the at least one candidate based on information representing each candidate, the information including information representing the light energy emitted by the pixel represented by the candidate and information representing the number of pixels represented by the candidate, the information of the selected at least one final candidate representing both the highest light energy and the maximum number of pixels; Furthermore, the final candidate application determination process is used to determine the modification of the gain function to ensure that the high dynamic range image complies with at least one light energy constraint, wherein the modification of the gain function is determined in response to determining that the light energy of the final candidate in the high dynamic range image will be higher than the at least one light energy constraint.
11. The device of claim 10, wherein creating at least one candidate based on at least one local maximum in the contributions and at least one local maximum in the population comprises: At least one local maximum is identified in the contributions, and for each local maximum, the corresponding band is aggregated with the adjacent bands to create a candidate; In the population, at least one local maximum is identified, and for each local maximum, the corresponding band is aggregated with the adjacent bands to create a candidate population; as well as, Create candidates from each candidate group that are independent of any other candidate.
12. The device of claim 10, wherein the histogram is a histogram of brightness values.
13. The device of claim 10, wherein the at least one light energy constraint is at least one of a MaxFall constraint and a diffuse white constraint.
14. The device of claim 10, wherein, in order to select at least one final candidate, the electronic circuitry is further adapted to: Select a candidate subset associated with the highest value representing light energy, and select at least one final candidate from the candidates of the subset representing the highest number of pixels.
15. The device of claim 10, wherein, for applying the determining process, the electronic circuitry is further adapted to: Determine the final pixel value representing the at least one final candidate; The inverse tone mapping operator function is used to calculate a value representing the extended pixel value from the final pixel value; and... Perform a modification process that adapts the gain function to the light energy constraint when the extended pixel value is higher than the light energy constraint representing a predefined diffuse white constraint value.
16. The device of claim 15, wherein the low dynamic range image is the current image in an image sequence, and the final pixel value is temporally filtered using at least one final pixel value calculated for at least one image preceding the current image in the image sequence.
17. The device of claim 10, wherein, in order to apply the determining process, the device is further configured to: Perform a modification process that is suitable for modifying the gain function when the value of MaxFall representing the high dynamic range image is higher than the light energy constraint representing the predefined MaxFall constraint.
18. The device according to claim 17, wherein, The value of MaxFall representing the high dynamic range image is the sum of the calculated contributions.
19. A computer program product comprising program code instructions for implementing the method according to claim 1.
20. An information storage device, the information storage device storing program code instructions for implementing the method according to claim 1.
Citation Information
Patent Citations
Inverse tone mapping method and corresponding device
EP3249605A1
Method for inverse tone mapping of an image
WO2015096955A1
Computational complexity adaptive HDR image conversion method and its system
CN106603941A
Method and apparatus for generating HDR images with reduced clipped areas
CN110097493A