Low complexity deep neural network for inverse tone mapped image generation using mixed data
By optimizing parameters through cascaded neural networks and statistical representations, the problem of high complexity in inverse tone mapping in existing technologies is solved, enabling efficient conversion from SDR to HDR and generating high-quality HDR images.
Patent Information
- Application Number
- CN202510467705.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-17
- Filing Date
- 2025-04-15
- Publication Date
- 2025-10-24
AI Technical Summary
Existing inverse tone mapping methods based on deep neural networks are highly complex and difficult to achieve real-time conversion of standard dynamic range video to high dynamic range video.
By employing a cascaded neural network structure and combining statistical representations such as histograms, the neural network parameters are optimized through the training process to achieve inverse tone mapping from standard dynamic range images to high dynamic range images.
It reduces the complexity of inverse tone mapping, improves conversion efficiency, makes real-time implementation possible, and generates high-quality HDR images.
Smart Images

Figure CN120833286A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] At least one of the embodiments in this document relates to the field of generation of high dynamic range content, and more particularly, to a method and apparatus for applying inverse tone mapping to standard or medium dynamic range content to generate HDR content. BACKGROUND
[0002] Recent advances in display technology allow for an extended dynamic range of colors, luminance and contrast in pictures to be displayed. The term "picture" here refers to picture content, which for example can be a video or a still picture or image.
[0003] High dynamic range video (HDR video) describes a video having a dynamic range greater than standard dynamic range video (SDR video) or medium dynamic range video (MDR video). HDR technology provides a better viewer experience (or quality of experience (QoE)) of video content.
[0004] Since in the past, most video content was produced in SDR, in order to benefit from the improvements brought by HDR technology when displaying these video contents, it is necessary to convert these video contents from SDR to HDR. Methods allowing the conversion of SDR content in HDR content are generally called inverse tone mapping (ITM) methods and use an inverse tone mapping operator (ITMO).
[0005] Most ITMOs use a mapping function to convert samples of SDR content into samples of HDR content (see document Y. Kinoshita. 2017. Fast Inverse Tone Mapping with Reinhard's Global Operator.).
[0006] In video production professional studios, ITM is described by one-dimensional (1D) or three-dimensional (3D) LUTs, which can be static (as proposed in document BBC. 2021. Release Notes for HLG Format Conversion LUTs vl.5.) or dynamic (as proposed in SL-HDR1 HDR distribution technology (ETSI TS 103 433-1) and document WO2021175633A1).
[0007] Recently, ITMOs based on deep neural networks have been proposed (see document Kinoshita, Y. 2019. "iTM-Net: Deep Inverse Tone Mapping Using Novel Loss Function Based on Tone Mapping Operator." (hereinafter Kinoshita 19) or document Kim, Soo Ye. 2019. "Deep SR-ITM Joint Learning of Super-Resolution and Inverse Tone-Mapping for 4K UHD HDR Applications.") Document Kinoshita 19 shows that the HDR pictures obtained using ITMOs based on deep neural networks have a higher quality than the HDR pictures obtained applying traditional ITMOs based on mapping functions. However, the complexity of these algorithms is quite high and their real-time implementation in products is problematic.
[0008] It is desirable to reduce the complexity of ITMOs based on deep neural networks in order to make them usable to replace traditional mapping functions. SUMMARY
[0009] In a first aspect, one or more of the embodiments herein provide a method comprising:
[0010] obtaining standard dynamic range (SDR) picture data; and,
[0011] applying a neural network implementing an inverse tone mapping process to the SDR picture data to obtain high dynamic range (HDR) picture data, wherein the neural network comprises a concatenation of a sample array representing the SDR picture data and a sample array representing at least one statistical representation of the SDR picture data.
[0012] In one embodiment, the SDR picture data is a full SDR picture and the HDR picture data is a full HDR picture.
[0013] In a second aspect, one or more of the embodiments herein provide a method comprising:
[0014] obtaining a pair of picture data comprising a standard dynamic range (SDR) version and a high dynamic range (HDR) version of the same picture data;
[0015] applying a neural network implementing an inverse tone mapping process to the SDR version to obtain a prediction of the HDR version, the neural network comprising a concatenation of a sample array representing the SDR version and a sample array representing at least one statistical representation of the SDR version.
[0016] computing an error measure between the prediction of the HDR version and the HDR version; and
[0017] using the error measure to update the parameters of the neural network during backpropagation.
[0018] In a third aspect, one or more of the embodiments provide a method comprising:
[0019] obtaining a database of pairs of pictures, each pair of pictures comprising a standard dynamic range (SDR) version and a high dynamic range (HDR) version of the same picture;
[0020] dividing each picture in the database into sample blocks to obtain pairs of SDR versions and HDR versions of the same sample block; and,
[0021] applying the method of the second aspect iteratively to each pair of the pairs of SDR versions and HDR versions of the same sample block to determine the parameters of the neural network, the parameters of the neural network updated in an iteration being used for the next iteration.
[0022] In one embodiment, the computation of the error measure involves samples of a sub- portion of the prediction of the HDR version and samples of a corresponding sub-portion of the HDR version, each sub-portion depending on characteristics of at least one of the at least one convolution process and the at least one subsampling process comprised in the neural network.
[0023] In one embodiment, the at least one statistical representation comprises a histogram of SDR pictures containing SDR picture data.
[0024] In a fourth aspect, one or more of the embodiments provide an apparatus comprising electronic circuitry configured for:
[0025] obtaining standard dynamic range (SDR) picture data; and,
[0026] applying a neural network implementing an inverse tone mapping process to the SDR picture data to obtain high dynamic range (HDR) picture data, wherein the neural network comprises a concatenation of a sample array representing the SDR picture data and a sample array representing at least one statistical representation of the SDR picture data.
[0027] In one embodiment, the SDR picture data is a full SDR picture and the HDR picture data is a full HDR picture.
[0028] In a fifth aspect, one or more of the embodiments provide an apparatus comprising electronic circuitry configured for:
[0029] obtaining a pair of picture data comprising a standard dynamic range (SDR) version and a high dynamic range (HDR) version of a same picture data;
[0030] applying a neural network implementing an inverse tone mapping process to the SDR version to obtain a prediction of the HDR version, the neural network comprising a concatenation of a sample array representative of the SDR version and a sample array representative of at least one statistical representation of the SDR version;
[0031] computing an error measure between the prediction of the HDR version and the HDR version; and
[0032] using the error measure to update parameters of the neural network in a backpropagation process.
[0033] In a sixth aspect, one or more of the embodiments herein provide a device comprising electronic circuitry configured for:
[0034] obtaining a database of pairs of pictures, each pair of pictures comprising a standard dynamic range (SDR) version and a high dynamic range (HDR) version of a same picture;
[0035] dividing each picture in the database into sample blocks to obtain pairs of SDR versions and HDR versions of a same sample block; and
[0036] applying an iterative process to the pairs of SDR versions and HDR versions of the same sample block to determine parameters of a neural network implementing an inverse tone mapping process, the parameters of the neural network updated in an iteration being used in the next iteration, the iterative process comprising, in an iteration, for a pair of the pairs of SDR versions and HDR versions of the same sample block:
[0037] applying a neural network implementing an inverse tone mapping process to the SDR version of the pair to obtain a prediction of the HDR version of the pair, the neural network comprising a concatenation of a sample array representative of the SDR version of the pair and a sample array representative of at least one statistical representation of the SDR version of the pair,
[0038] computing an error measure between the prediction of the HDR version of the pair and the HDR version of the pair, and
[0039] using the error measure to update parameters of the neural network in a backpropagation process.
[0040] In one embodiment, the computation of the error measure involves samples of a sub- portion of the prediction of the HDR version and samples of a corresponding sub-portion of the HDR version, each sub-portion depending on characteristics of at least one among at least one convolution process and at least one subsampling process comprised in the neural network.
[0041] In one embodiment, the at least one statistical representation comprises a histogram of the SDR picture including the SDR picture data.
[0042] In a seventh aspect, one or more of the embodiments herein provide a non-transitory information storage medium storing program code instructions for implementing the methods according to the first, second and third aspects.
[0043] In an eighth aspect, one or more of the embodiments herein provide a computer program comprising program code instructions for implementing the methods according to the first, second and third aspects. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 An example of a context in which various embodiments are implemented is schematically illustrated;
[0045] Figure 2A An example of a hardware architecture of a processing module capable of implementing various embodiments is schematically illustrated;
[0046] Figure 2B A block diagram illustrating an example of a first system in which various aspects and embodiments are implemented;
[0047] Figure 2C A block diagram illustrating an example of a second system in which various aspects and embodiments are implemented;
[0048] Figure 3 An inverse tone mapping process for generating an HDR picture from an SDR picture based on a neural network is schematically illustrated;
[0049] Figure 4 A training process is schematically illustrated;
[0050] Figure 5A A neural network for implementing the inverse tone mapping process is schematically illustrated;
[0051] Figure 5B A block diagram of a process implementing a neural network implementing an inverse tone mapping process is illustrated;
[0052] Figure 6A An impact of an input sample on a 16x16 pixel block is schematically illustrated; and
[0053] Figure 6B Processing of a block by a neural network implementing an inverse tone mapping process is illustrated. DETAILED DESCRIPTION
[0054] Figure 1 A context in which embodiments are implemented is schematically illustrated.
[0055] In Figure 1In the middle, system 11 - which can be a camera, a storage device, a computer, a server, or any device capable of transmitting a video stream (i.e. video data) - transmits the video stream to system 13 using communication channel 12. The video stream is either encoded and transmitted by system 11, or received and / or stored by system 11, and then transmitted. The video stream represents, for example, standard dynamic range (SDR) content encoded using a compression method such as VVC (Versatile Video Coding (VVC), ITU-T H.266), HEVC (ISO / IEC 23008-2 - MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265), AVC (ISO / CEI 14496-10), EVC (Elementary Video Coding / MPEG-5), AV1, AV2, and VP9 or JPEG.
[0056] Communication channel 12 is a wired (e.g. Internet or Ethernet) or wireless (e.g. WiFi, 3G, 4G or 5G) network link.
[0057] System 13, which can be for example a set-top box, receives and decodes the video stream to generate, for example, SDR content. From the SDR content, system 13 generates high dynamic range (HDR) content using a neural network (NN) based inverse tone mapping (ITM) method, which is described later in connection with Figure 3 、 Figure 4 、 Figure 5A 、 Figure 5B 、 Figure 6A 、 Figure 6B .
[0058] The obtained HDR content is then transmitted to display 15 using communication channel 14, which can be a wired or wireless network. Display 15 then displays the HDR content.
[0059] In one embodiment, system 13 is included in display 15. In that case, system 13 and display 15 are included in a TV, a computer, a tablet, a smartphone, a head-mounted display, etc.
[0060] Figure 2AAn example of a hardware architecture of a processing module 200 used for example in system 11 or in system 13 is schematically illustrated. The processing module 200 comprises, connected by a communication bus 2005: a processor or CPU (Central Processing Unit) 2000, encompassing, as non-limiting examples, one or more microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, a Random Access Memory (RAM) 2001, a Read Only Memory (ROM) 2002, a storage unit 2003, which can include non-volatile memory and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, disk drives and / or optical disk drives, or a storage medium reader, such as an SD (Secure Digital) card reader and / or a Hard Disk Drive (HDD) and / or a network-accessible storage device, at least one communication interface 2004 for exchanging data with other modules, devices, systems or equipment. The communication interface 2004 can include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 12. The communication interface 2004 can include, but is not limited to, a modem or a network card.
[0061] For example, when implemented in system 13, the communication interface 2004 enables the processing module 200 to receive SDR content and output HDR content, for example.
[0062] The processor 2000 is capable of executing instructions loaded into the RAM 2001 from the ROM 2002, an external memory (not shown), a storage medium or a communication network. When the processing module 200 is powered on, the processor 2000 is capable of reading instructions from the RAM 2001 and executing them.
[0063] When the processing module 200 is included in system 11, the instructions form a computer program that enables, for example, a process of training parameters of a NN implementing an ITM process, implemented by the processor 400.
[0064] When the processing module 200 is included in system 13, the instructions form a computer program that enables, for example, a NN-based ITM process, implemented by the processor 2000.
[0065] All or some of the algorithms and steps of the above processes can be implemented in software form by execution of a set of instructions by a programmable machine such as a DSP (Digital Signal Processor) or a microcontroller, or in hardware form by a machine or a dedicated component such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit). Microprocessors, DSPs, FPGAs and ASICs are considered as electronic circuits.
[0066] Figure 2C FIG. 1 illustrates a block diagram of an example of a system 13 in which various aspects and embodiments are implemented.
[0067] System 13 can be embodied as a device that includes various components or modules, and is configured to generate HDR content from SDR content. Examples of such systems include, but are not limited to, various electronic systems, such as a personal computer, a laptop computer, a smartphone, a tablet computer, a TV, or a set-top box. The components of system 13 can be embodied individually or in combination as a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, system 13 includes at least one processing module 200 that implements an ITM module that implements a NN-based ITM process. In various embodiments, system 13 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.
[0068] As indicated in block 231, input can be provided to processing module 200 by various input modules. Such input modules include, but are not limited to, (i) a radio frequency (RF) module that receives RF signals transmitted, for example, over the air by a broadcaster, (ii) a component (COMP) input module (or a set of COMP input modules), (iii) a universal serial bus (USB) input module, and / or (iv) a high-definition multimedia interface (HDMI) input module. Figure 2C Other examples, not shown in FIG. 1, include composite video.
[0069] In various embodiments, the input modules of block 231 have associated respective input processing elements as known in the art. For example, the RF module can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower frequency band to, for example, select a signal frequency band which can be referred to in some embodiments as a channel, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF module of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various of these functions, including for example, downconverting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF module includes an antenna.
[0070] Additionally, the USB and / or HDMI modules can include respective interface processors for connecting system 13 to other electronic devices across USB and / or HDMI connections. It will be appreciated that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented as desired, for example, within a separate input processing IC or within processing module 200. Similarly, aspects of USB or HDMI interface processing can be implemented within a separate interface IC or within processing module 200 as desired. The demodulated, error corrected, and demultiplexed streams are provided to processing module 200.
[0071] Various elements of system 13 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using a suitable connection arrangement, for example, an internal bus as known in the art, including an Inter-IC (I2C) bus, wiring, and printed circuit boards. For example, in system 13, processing module 200 is interconnected with other elements of the system 13 by bus 405.
[0072] Communication interface 2004 of processing module 200 allows system 13 to communicate over communication channel 12. Communication channel 12 can be implemented, for example, within a wired and / or wireless medium.
[0073] In various embodiments, data is streamed or otherwise provided to system 13 using a wireless network such as a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). The Wi-Fi signals of these embodiments are received through communication channel 12 and communication interface 2004 adapted for Wi-Fi communication. Communication channel 12 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Still other embodiments provide streaming data to system 13 using an RF connection of input block 231. As indicated above, various embodiments provide data in a non-streaming manner, for example, when system 13 is a smart phone or tablet. Additionally, various embodiments use wireless networks other than Wi-Fi, for example, a cellular network or a Bluetooth network.
[0074] System 13 can provide output signals to various output devices using communication channel 12 or bus 2005. For example, system 13 can provide HDR content.
[0075] System 13 can provide output signals to various output devices including display 15, speakers 235, and other peripheral devices 236. Display 15 of various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 15 can be used in a television set, a tablet, a notebook computer, a cellular phone (mobile phone), or other device. HDR display 15 can also be integrated with other components (e.g., in a smart phone), or be separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 236 include one or more of a stand-alone digital video recorder (or digital versatile disk) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 236 that provide a function based on the output of system 13. For example, a disc player performs the function of playing the output of system 13.
[0076] In various embodiments, signaling such as AV.Link, Consumer Electronics Control (CEC), or other communications protocols that enable control of devices to devices with or without user intervention are used to communicate control signals between system 13 and display 15, speakers 235, or other peripheral devices 236. Output devices can be communicatively coupled to system 13 via dedicated connections through respective interfaces 232, 233, and 234. Alternatively, output devices can be connected to system 13 using the communication channel 12 via the communication interface 2004. Display 15 and speakers 235 can be integrated with other components of system 13 in a single unit, such as for example, a television. In various embodiments, display interface 232 includes a display driver such as for example, a timing controller (T Con) chip.
[0077] For example, if the RF module of block 231 is part of a standalone set-top box, display 15 and speakers 235 can instead be separate from one or more other components. In various embodiments where display 15 and speakers 235 are external components, output signals can be provided via dedicated output connections, including for example, HDMI ports, USB ports, or COMP outputs.
[0078] Figure 2B FIGURE 1 illustrates a block diagram of an example of a system 11 in which various aspects and embodiments are implemented.
[0079] System 11 can be embodied as a device including the various components and modules described above, and configured to perform one or more of the aspects and embodiments described in this document.
[0080] Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, cameras, smartphones, and servers. Elements or modules of system 11 can be embodied in a single integrated circuit (IC), multiple ICs, and / or a discrete set of components, singly or in combination. For example, in at least one embodiment, system 11 includes at least one processing module 200 that implements a training process for a NN that implements an ITM process. In various embodiments, system 11 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.
[0081] Input to processing module 200 can be provided through various input modules, as indicated in block 231 already described with reference to FIGURE 1. Figure 2C
[0082] The various elements of system 11 can be disposed within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangements, e.g., internal buses as known in the art, including Inter-IC (I2C) buses, wiring, and printed circuit boards. For example, in system 11, processing module 200 is interconnected with other elements of the system 11 by bus 2005.
[0083] Communication interface 2004 of processing module 200 allows system 11 to communicate over communication channel 12. Communication channel 12 can be implemented, for example, in a wired and / or wireless medium.
[0084] In various embodiments, data is streamed to or otherwise provided to system 11 using a wireless network such as a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). The Wi-Fi signals of these embodiments are received through communication channel 12 and communication interface 2004 adapted for Wi-Fi communication. The communication channel 12 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Still other embodiments provide streaming data to system 11 using the RF connection of input block 231. As indicated above, various embodiments provide data in a non-streaming manner.
[0085] When one of the drawings is presented as a flowchart, it should be understood that it also provides a block diagram for a corresponding apparatus. Similarly, when one of the drawings is presented as a block diagram, it should be understood that it also provides a flowchart for a corresponding method / process.
[0086] Implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation, a feature discussed can also be implemented in other forms (e.g., an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), smart phones, tablets, and other devices that facilitate communication of information between end-users.
[0087] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variants thereof, means that a particular feature, structure, characteristic, and so forth being described in connection with an embodiment is included in at least one embodiment. Therefore, the appearance of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well as any other variations thereof, throughout this application, and in each case does not necessarily refer to the same embodiment.
[0088] Additionally, the application can involve “determining” various pieces of information. Determining information can include one or more of, for example, estimating information, calculating information, predicting information, retrieving information from memory, or obtaining information from another device, module, or from a user.
[0089] Furthermore, the application can involve “accessing” various pieces of information. Accessing information can include one or more of, for example, receiving information, retrieving information (for example, from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0090] Additionally, the application can involve “receiving” various pieces of information. As with “accessing”, receiving is intended to be a broad term. Receiving information can include one or more of, for example, accessing information or retrieving information (for example, from memory). Furthermore, “receiving” is typically involved in some fashion during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0091] It will be appreciated that, in the case of “A / B”, “A and / or B”, and “at least one of A and B”, “one or more of A and B”, any such “ / ”, “and / or”, and “at least one of”, “one or more of” uses are intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the case of “A, B, and / or C” and “at least one of A, B, and C” or “one or more of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This can be extended, as is clear to one of ordinary skill in the art, for as many items in a similar manner.
[0092] It will be apparent to those skilled in the art that the embodiments or examples can produce a variety of signals including but not limited to digital signals, analog signals, wireless signals, digital data, analog data, wireless data, etc. The signals can be amplified, manipulated at the binary level, or otherwise processed, by various techniques long known to those of skill in the art. It will be apparent to those of skill in the art that the embodiments or examples can produce a variety of signals including but not limited to digital signals, analog signals, wireless signals, digital data, analog data, wireless data, etc. The signals can be amplified, manipulated at the binary level, or otherwise processed, by various techniques long known to those of skill in the art.
[0093] Figure 3 An inverse tone mapping process for generating an HDR picture from an SDR picture based on a neural network is schematically illustrated.
[0094] Figure 3 The process of FIG. 13 is for example implemented by the processing module 200 of the system 13.
[0095] In step 1330, the processing module 200 obtains SDR content. The SDR content is for example a 2K (1920x1080), 4K (3840x2160) or 8K (7680x4320) SDR picture.
[0096] In step 1331, the processing module 200 applies a NN-based ITM process to the SDR picture to obtain an HDR picture. In Figure 5A The NN implementing ITM is illustrated in FIG. 13. The NN is based on Figure 5A The NN-based ITM process based on the NN of FIG. 13 is detailed below in connection with Figure 5B The NN comprises a concatenation of a sample array representative of the SDR picture and a sample array representative of at least one statistical representation of the SDR picture.
[0097] In step 1332, the processing module 200 of the system 13 provides the HDR picture to the display 15 so that the display 15 displays the HDR picture.
[0098] Figure 4 A training process for determining parameters of a NN implementing an ITM process is schematically illustrated.
[0099] The training of the NN is performed during an offline process, for example, by the processing module 200 of the system 11. During the training process, a database comprising a plurality of pairs of picture data is obtained, each pair of picture data comprising an SDR version and an HDR version of the same picture data. Figure 4 The process is applied iteratively to each pair in the database. Each iteration uses the parameters of the NN updated in the previous iteration, thereby computing the prediction of the HDR version of the pair from the SDR version using the NN-based ITM process during training.
[0100] In step 1140 , the processing module 200 obtains a pair of HDR versions of the database.
[0101] In step 1141 , the processing module 200 obtains the corresponding SDR version of the same pair from the database.
[0102] In step 1142, the processing module 200 applies the NN-based ITM process to the SDR version, where the NN parameters are updated in the previous iteration, to obtain the prediction of the HDR version. Figure 5A The process for updating NN parameters is shown in Figure 5B Explained below. The NN comprises a concatenation of a sample array representing the SDR version and a sample array representing at least one statistical representation of the SDR picture comprising the SDR version.
[0103] In step 1143, the processing module 200 calculates an error metric between the predicted and HDR versions of the pair.
[0104] In step 1144, the processing module 200 uses the error metric in a back-propagation process to update the parameters of the NN for the next iteration.
[0105] Figure 5A A NN for implementing the inverse tone mapping process is schematically illustrated.
[0106] Figure 5A The NN consists of various modules described below.
[0107] Module 5043 implements a convolution stage based on K number of 3x3 convolution kernels on an input two-dimensional (2D) array of samples of size N x M, allowing K filtered samples to be obtained for each 2D sample position of the input 2D array. In the following, we refer to a "2D sample position" as the position of a sample in a plane parallel to the plane of the SDR pictures input to the NN. Thus, all output samples resulting from the convolution of the same input sample by several convolution kernels have the same 2D sample position, but in different planes parallel to the plane of the SDR pictures. The convolution stage is followed by a batch normalization and a rectified linear unit (ReLu) function. Batch normalization (also referred to as batch normalization) is a method (which is presented in Ioffe, Sergey; Szegedy, Christian (2015). "Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift") that is used to normalize the input of a layer by recentering and rescaling, making the training of the NN faster and more stable.
[0108] The ReLu function is defined as follows:
[0109]
[0110] where x is the input of the function.
[0111] Module 5044 implements a 2x2 max pooling. 2x2 max pooling is a pooling operation that computes the maximum value of a 2x2 patch of samples of the input 2D array and uses it to create an output array corresponding to a down-sampled (pooled) version of the input 2D array. Thus, for an input array of size N x M x K, the output array has size (N / 2) x (M / 2) x K.
[0112] Module 5045 implements a convolution stage based on K x 2 number of 3x3xK convolution kernels, allowing K x 2 filtered samples to be obtained for each 2D sample position of the input array, followed by a batch normalization and a ReLu function. Thus, the output array has size (N / 2) x (M / 2) x (K x 2).
[0113] Module 5046 implements a 2x2 max pooling, whose output is an array of size (N / 4) x (M / 4) x (K x 2).
[0114] Module 5047 implements a convolution stage based on K x 4 number of 3x3x(K x 2) convolution kernels, allowing K x 4 filtered samples to be obtained for each 2D sample position of the input array, followed by a batch normalization and a ReLu function. Thus, the output array has size (N / 4) x (M / 4) x (K x 4).
[0115] Module 5049 generates at least one statistical representation of the picture. For example, module 5049 generates a histogram of the picture and represents this histogram with a 1D array whose size is equal to the number of bins of the histogram.
[0116] Module 5050 implements a fully connected layer followed by batch normalization and ReLu. Here, the fully connected layer simply connects all inputs of module 5050 to all outputs of this module without reducing the number of samples per bin of the histogram.
[0117] Module 5051 implements a replication process. The replication process generates a 2D array from each sample of the input array output by module 5050 by replicating the value of the sample at each 2D sample position of the 2D array. If the input array is a histogram comprising B bins, then B 2D arrays are generated. We then obtain a 3D array of size (N / 4)x(M / 4)x B.
[0118] Module 5052 concatenates the output of module 5047 and the output of module 5051. The output of module 5052 is a 3D array of size (N / 4)x(M / 4)x(Kx4+B).
[0119] Module 5053 implements a convolution stage based on a number Kx8 of 3x3x(Kx4+B) convolution kernels, allowing to obtain Kx8 samples for each 2D sample position of the input array, followed by a batch normalization and a ReLu function. The output array is thus of size (N / 4)x(M / 4)x(Kx8).
[0120] Module 5054 implements a transpose convolution based on a 4x4 convolution kernel, followed by a batch normalization and a ReLu function. The output array is thus of size (N / 2)x(M / 2)x(Kx4).
[0121] Module 5055 implements a convolution stage based on a number Kx4 of 3x3x(Kx4) convolution kernels, allowing to obtain Kx4 samples for each 2D sample position of the input array, followed by a batch normalization and a ReLu function. The output array is thus of size (N / 2)x(M / 2)x(Kx4).
[0122] Module 5056 implements a transpose convolution based on a 4x4 convolution kernel, followed by a batch normalization and a ReLu function. The output array is thus of size NxMx(Kx2).
[0123] Module 5057 implements a convolution stage based on a K×2 number of 3×3×(K×2) convolution kernels, allowing K×2 samples to be obtained for each 2D sample position of the input array, followed by batch normalization and ReLu functions. Therefore, the output array size is N×M×(K×2).
[0124] Module 5058 implements a convolution stage based on a 1×1×(K×2) convolution kernel, allowing to obtain a sample for each 2D sample position of the input array, followed by a ReLu function. Therefore, the output 2D array size is N×M
[0125] Figure 5B The diagram shows the Figure 5A Block diagram of an example of the process of implementing the inverse tone mapping process by a NN.
[0126] Figure 5B The diagram shows the Figure 5A The NN implements two aspects of the ITM process: a process for updating the NN parameters used during training of the NN, and applying the ITM to SDR picture data using the trained NN.
[0127] In one embodiment, the training process and the ITM process are applied to the luma component Y of the YUV picture. The chroma components U and V of the HDR picture are derived from the luma component Y. For example, if the luma component of the sample of the HDR picture is Y HDR , and the brightness component of the corresponding sample of the SDR picture is Y SDR , then the chrominance component U of the sample of the HDR picture HDR and V HDR The calculation is as follows:
[0128]
[0129] Among them U SDR and V SDR They are the chroma components of the corresponding samples of the SDR image.
[0130] NN training:
[0131] When explaining the training of NN, Figure 5B The process corresponds to Figure 4 In one embodiment, each luminance component Y of the pictures of the database is divided into blocks of size N×M=64×64. Thus, Figure 4The picture data is a 64x64 block of the luminance component of the SDR picture. Thus, the training process is applied to each pair of 64x64 luminance component blocks of the SDR picture and 64x64 luminance component blocks of the HDR picture of the database. Processing 64x64 blocks instead of the complete picture allows better capturing of local characteristics of the picture, and processing of multiple blocks of a picture in parallel. Moreover, in the example, the number of bins B of the histogram is equal to the number of bins of the histogram used in the training process. In the example, the number of bins B of the histogram is equal to 64. Figure 5B In the example, the initial number of convolutional kernels K = 32, and the number of bins B of the histogram is 64.
[0132] In step 5141, the processing module 200 obtains the luminance component of the input SDR picture.
[0133] In step 5142, the processing module 200 crops a 64x64 block of the luminance component of the input SDR picture. During the training process, all 64x64 blocks of the luminance component of the input SDR picture are cropped and processed successively or in parallel.
[0134] In step 5143, the processing module 200 applies a convolutional stage to the 64x64 block, followed by a batch normalization and a ReLu function implemented by the module 5043. The output of step 5143 is an array of size 64x64x32.
[0135] In step 5144, the processing module 200 applies a 2x2 max pooling process implemented by the module 5044 to the output of step 5143. The output of step 5144 is an array of size 32x32x32.
[0136] In step 5145, the processing module 200 applies a convolutional stage to the array resulting from step 5144, followed by a batch normalization and a ReLu function implemented by the module 5045. The output of step 5145 is an array of size 32x32x64.
[0137] In step 5146, the processing module 200 applies a 2x2 max pooling process implemented by the module 5046 to the array resulting from step 5145. The output of step 5146 is an array of size 16x16x64.
[0138] In step 5147, the processing module 200 applies a convolutional stage to the array resulting from step 5146, followed by a batch normalization and a ReLu function implemented by the module 5047. The output of step 5147 is an array of size 16x16x128.
[0139] In step 5149, the processing module 200 obtains the statistical representation of the luminance component of the input SDR picture. In Figure 5BIn the example of , the processing module 200 obtains a histogram of an input SDR picture comprising B=64 bins. During step 5149 , the histogram is represented by a 1D array of size “64”.
[0140] In step 5150, the processing module 200 applies the fully connected layer implemented by module 5050 to the array resulting from step 5149. The output of step 5150 is a 1D array of size "64".
[0141] In step 5151, processing module 200 applies the copy process implemented by module 5051 to the array resulting from step 5150. The output of step 5151 is a 3D array of size 16x16x64.
[0142] In step 5152, the processing module 200 applies the process implemented by module 5052 to concatenate the array output by module 5047 with the array output by module 5051. In other words, the processing module 200 concatenates the sample array representing the 64×64 block of the luma component of the input SDR picture cropped in step 5042 to the sample array representing the statistical representation of the luma component of the input SDR picture. The output of step 5152 is an array of size 16×16×(128+64).
[0143] In step 5153, processing module 200 applies a convolution stage to the array resulting from step 5152, followed by batch normalization and a ReLu function implemented by module 5053. The output of step 5153 is an array of size 16x16x256.
[0144] In step 5154, processing module 200 applies a transposed convolution followed by batch normalization and a ReLu function implemented by module 5054 to the array resulting from step 5153. The output of step 5154 is an array of size 32x32x128.
[0145] In step 5155, processing module 200 applies a convolution stage to the array resulting from step 5154, followed by batch normalization and a ReLu function implemented by module 5055. The output of step 5155 is an array of size 32x32x128.
[0146] In step 5156, processing module 200 applies a transposed convolution followed by batch normalization and a ReLu function implemented by module 5056 to the array resulting from step 5155. The output of step 5156 is an array of size 64x64x64.
[0147] In step 5157, the processing module 200 applies a convolution stage to the array resulting from step 5156, followed by a batch normalization and a ReLu function implemented by module 5057. The output of step 5157 is an array of size 64x64x64.
[0148] In step 5158, the processing module 200 applies a convolution stage to the array resulting from step 5157, followed by a ReLu function implemented by module 5058. The output of step 5158 is an array of size 64x64, which represents the prediction of the 64x64 HDR block corresponding to the 64x64 SDR block obtained in step 5141.
[0149] In step 5159, the processing module 200 computes an error measure between the 64x64 HDR block and the prediction of the 64x64 HDR block. For example, the error measure is the sum of absolute differences (SAD) or the sum of squared differences (SSD). The error measure is used in the backpropagation process to update the NN parameters of the next 64x64 SDR block of the database.
[0150] The output of the training process is the trained NN implementing the ITM process. The parameters of the NN thus trained are provided to the system 13 so that the system 13 can implement the NN-based ITM process.
[0151] In a variant of the training process illustrated by Figure 5B In this case, N and M represent the dimensions of the SDR picture.
[0152] Since the size of the convolutional filters is 3x3, as represented in Figure 6A the sample values after two convolutional layers followed by a two-times pooling and a final convolutional layer depend on the samples of a 10x10 region (also called receptive field) in the original input block, regardless of its size. Figure 6A It is shown that the sample of row #5 in the last stage depends on the input samples of row #0 to row #10 in the original input. Thus, as represented in Figure 6B the internal samples located in a 54x54 block in the center of the original 64x64 block are perfectly computed without border effects. In a variant, in the case of a 64x64 input block, in step 5159, the error measure is computed only on the internal 54x54 block. In this variant, the internal samples used for computing the error measure thus depend on the characteristics of the convolutional filters and on the characteristics of the subsampling process included in the NN (here implemented by a 2x2 max pooling).
[0153] Inverse tone mapping of the SDR picture:
[0154] Using Figure 5Ainverse tone mapping of the pictures of the NN is performed during the streaming session, for example by the processing module 200 of the system 13. During the streaming session, the processing module receives the sequence of SDR pictures and applies the ITM process of the NN to each received SDR picture. Figure 5A
[0155] In one embodiment, the processing module 200 of the system 13 applies the NN to the luminance component Y of each received SDR picture. As can be seen, in this embodiment, while the NN is applied to blocks of SDR pictures during the training process, the full SDR picture is input to the NN when applying the ITM. Figure 5A
[0156] The ITM process applies the process of Figure 5B However, during the ITM process, the clipping step 5142 is not applied and in step 5149, the error measure is not computed. It can be noted that, in the case of the ITM process, the process of Figure 5B corresponds to steps 1330 and 1331 of the process of Figure 3
[0157] Therefore, when implementing the ITM process applied to the luminance component of the input SDR picture, Figure 5B the output of the process of is the luminance component of the HDR picture, the chrominance components being derived from the luminance component.
[0158] So far, we have considered that only the luminance component Y of the SDR picture is considered in the training process and in the ITM process, the chrominance components U and V being derived from the luminance component Y. In another embodiment, each component Y, U and V is successively input to the NN during the training process and during the ITM process. In this embodiment, the output of the training process is three sets of NN parameters, one set per component. The ITM process is successively applied on each component of the input SDR picture using the NN parameters trained specifically for this component.
[0159] Moreover, Figure 5A the NN of is designed to process a 2D array of SDR samples, which means that either the NN is applied to the luminance component of the SDR picture and the chrominance components of the HDR picture are derived from the luminance component of the HDR picture obtained using the NN, or the NN is applied independently to each component of the SDR picture. In one embodiment, the NN is designed to process a 3D array of SDR samples, which comprises three sample values for each 2D sample location. In this case, the 3x3 convolution kernel used in module 5043 is replaced by a 3x3x3 convolution kernel and the lxlx(Kx2) convolution kernel used by module 5058 is replaced by three lxlx(Kx2) convolution kernels, one per component.
[0160] We have also seen in the above embodiments that, during the training, each luminance component of the SDR picture of the database is divided into blocks of size N x M = 64 x 64. In another variant, other sizes of blocks can be used, such as 128 x 128, 32 x 32, 128 x 64, 64 x 128, etc.
[0161] We have also seen in embodiments of the training process that, instead of blocks, the full SDR picture can be processed (step 5042 is skipped).
[0162] Similarly, in another embodiment, instead of processing the full picture, during the ITM process, the SDR picture is divided into blocks of size N x M (step 5042 is enabled) and the ITM process is applied successively or in parallel to each block of the SDR picture.
[0163] In another embodiment, other values of K and B are possible, for example K = 64 or K = 16 and B = 128 or B = 256.
[0164] In the example of the NN of Figure 5A , the NN comprises: two analysis blocks consisting of a 2 x 2 max-pooling module and a module implementing a convolution stage followed by a batch normalization and a ReLu function (5044 + 5045 and 5046 + 5047); and two synthesis blocks consisting of a module implementing a transposed convolution followed by a batch normalization and a ReLu function and a module implementing a convolution stage followed by a batch normalization and a ReLu function (5054 + 5055 and 5056 + 5057). In one embodiment, the number of analysis and synthesis blocks is different from two and is for example equal to "4", "3" or "1".
[0165] Moreover, in the example NN of Figure 5A , the statistical representation of the SDR picture is a single histogram of the SDR picture. In other embodiments, several histograms of various granularities (i.e. various numbers of bins) can be used. Other types of statistical representations can also be used. For example, the statistical information can be a cumulative histogram of the SDR picture, or an array of variances or mean values of N x M blocks of the SDR picture, or any combination of several statistical representations.
[0166] In one embodiment, each 2 x 2 max-pooling step is replaced by a step of computing the average or median value of the samples of a 2 x 2 patch.
[0167] In one embodiment, both the training process and the ITM process are performed by the system 11. In this case, the HDR content generated by the system 11 from the SDR content is transmitted to the system 13.
[0168] We described several embodiments above. Features of the embodiments can be provided alone or in any combination. In addition, embodiments can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types:
[0169] • A TV, set-top box, cell phone, tablet, personal computer, or other electronic device that performs at least one of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display).
[0170] • A TV, set-top box, cell phone, tablet, personal computer, or other electronic device that tunes a channel (e.g., using a tuner) to receive a signal including encoded SDR video and performs at least one of the described embodiments.
[0171] • A TV, set-top box, cell phone, tablet, or other electronic device that receives a signal including encoded SDR video over the air (e.g., using an antenna) and performs at least one of the described embodiments.
[0172] • A server, camera, cell phone, tablet, personal computer, or other electronic device that performs at least one of the described embodiments to generate HDR content from SDR content and transmits a signal including the HDR content over the air (e.g., using an antenna).
[0173] • A server, camera, cell phone, tablet, personal computer, or other electronic device that tunes a channel (e.g., using a tuner) to transmit a signal including SDR video and metadata representing parameters of a NN implementing an ITM process and performs at least one of the described embodiments.
[0174] • A server, camera, cell phone, tablet, personal computer, or other electronic device that transmits a signal including SDR video and metadata representing parameters of a NN implementing an ITM process over the air (e.g., using an antenna) and performs at least one of the described embodiments.
Claims
1. A method comprising: obtaining (1330) standard dynamic range (SDR) picture data; as well as, A neural network implementing an inverse tone mapping process is applied (1331) to SDR picture data to generate high dynamic range (HDR) picture data, wherein the neural network comprises a concatenation of a first sample array representing the SDR picture data and a second sample array representing at least one statistical representation of the SDR picture data.
2. The method of claim 1, wherein, The SDR picture data is a full SDR picture, and the HDR picture data is a full HDR picture.
3. A method comprising: obtaining (1140, 1141) a pair of picture data comprising a standard dynamic range (SDR) version and a high dynamic range (HDR) version of the same picture data; applying (1142) a neural network implementing an inverse tone mapping process to the SDR version to generate a prediction of the HDR version, the neural network comprising a concatenation of a first sample array representing the SDR version and a second sample array representing at least one statistical representation of the SDR version; calculating (1143) an error metric between the HDR version of the prediction and the HDR version; as well as The (1144) error metric is used during back-propagation to update the parameters of the neural network.
4. A method comprising: Obtain a database of multiple pairs of images, each pair including a standard dynamic range (SDR) version and a high dynamic range (HDR) version of the same image; Divide each image in the database into sample blocks to obtain multiple pairs of SDR and HDR versions of the same sample block; and The method according to claim 3 is iteratively applied to each of multiple pairs of SDR and HDR versions of the same sample block to determine parameters of the neural network, and the parameters of the neural network updated in the iteration are used for the next iteration.
5. The method of claim 4, wherein, The calculation of the error measure involves samples of a sub-portion of the prediction of the HDR version and samples of a corresponding sub-portion of the HDR version, each sub-portion depending on characteristics of at least one of at least one convolution process and at least one subsampling process included in the neural network.
6. The method of claim 1, wherein, The at least one statistical representation comprises a histogram of the SDR picture comprising the SDR picture data.
7. A device comprising an electronic circuit, the electronic circuit being configured to: obtaining (1330) standard dynamic range (SDR) picture data; and, A neural network implementing an inverse tone mapping process is applied (1331) to SDR picture data to generate high dynamic range (HDR) picture data, wherein the neural network comprises a concatenation of a first sample array representing the SDR picture data and a second sample array representing at least one statistical representation of the SDR picture data.
8. The apparatus of claim 7, wherein, The SDR picture data is a full SDR picture, and the HDR picture data is a full HDR picture.
9. A device comprising an electronic circuit configured to: obtaining (1140, 1141) a pair of picture data comprising a standard dynamic range (SDR) version and a high dynamic range (HDR) version of the same picture data; applying (1142) a neural network implementing an inverse tone mapping process to the SDR version of the pair to generate a prediction of the HDR version of the pair, the neural network comprising a concatenation of a first array of samples representative of the SDR version of the pair and a second array of samples representative of at least one statistical representation of the SDR version of the pair; computing (1143) an error measure between the prediction of the HDR version and the HDR version of the pair; and using (1144) the error measure to update the parameters of the neural network in a backpropagation process.
10. A device comprising an electronic circuit configured for: obtaining a database of pairs of pictures, each pair of pictures comprising a standard dynamic range (SDR) version and a high dynamic range (HDR) version of a same picture; dividing each picture in the database into sample blocks to obtain pairs of SDR and HDR versions of the same sample block; and applying an iterative process to pairs of SDR versions and HDR versions of a same sample block to determine parameters of a neural network implementing an inverse tone mapping process, the parameters of the neural network updated in an iteration being used for the next iteration, the iterative process comprising, for a pair of a plurality of pairs of SDR versions and HDR versions of a same sample block in one iteration: applying (1142) a neural network implementing an inverse tone mapping process to the SDR version of the pair to generate a prediction of the HDR version of the pair, the neural network comprising a concatenation of a first array of samples representative of the SDR version of the pair and a second array of samples representative of at least one statistical representation of the SDR version of the pair; computing (1143) an error measure between the prediction of the HDR version and the HDR version of the pair; and using (1144) the error measure to update the parameters of the neural network in a backpropagation process.
11. The apparatus of claim 10, wherein, The computation of the error measure involves samples of a sub-portion of the prediction of the HDR version and samples of a corresponding sub-portion of the HDR version, each sub-portion depending on a characteristic of at least one among at least one convolution process and at least one subsampling process comprised in the neural network.
12. The apparatus of claim 7, wherein, The at least one statistical representation comprises a histogram of the SDR picture containing the SDR picture data.
13. A non-transitory information storage medium storing program code instructions for implementing the method of claim 1.
14. The method of claim 3, wherein, The at least one statistical representation comprises a histogram of the SDR picture containing the SDR picture data.
15. The method of claim 4, wherein, The at least one statistical representation comprises a histogram of the SDR picture containing the SDR picture data.
16. The apparatus of claim 9, wherein, The at least one statistical representation comprises a histogram of the SDR picture containing the SDR picture data.
17. The apparatus of claim 10, wherein, The at least one statistical representation comprises a histogram of the SDR picture containing the SDR picture data.
18. A non-transitory information storage medium storing program code instructions for implementing the method of claim 3.
19. A non-transitory information storage medium storing program code instructions for implementing the method of claim 4.
Citation Information
Patent Citations
Method and apparatus for inverse tone mapping
WO2021175633A1