Reconstructing HDR video with converted tone mapping
By dividing the luma mapping function into a horizontal stretch and scaling operation, the system addresses ITM-induced artifacts in HDR encoding, ensuring accurate HDR image reconstruction from lower dynamic range proxies with reduced compression artifacts.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-02
- Publication Date
- 2026-04-01
AI Technical Summary
Existing HDR video encoding systems face issues with luminance mapping that result in undesirable artifacts, particularly when using inverse tone mapping (ITM) automata, which can cause hard or soft clipping, leading to visible compression artifacts in high dynamic range images.
The system divides the luma mapping function into two operations: a horizontal stretch followed by a scaling operation, where the input coordinates where clipping occurs are mapped to the maximum normalized input coordinates, and the resulting extended luma mapping function is communicated as metadata, allowing for improved reversibility and reduced artifacts.
This approach enhances the encoding process by minimizing decoding artifacts, ensuring accurate reconstruction of HDR images from lower dynamic range proxies, maintaining artistic intent while reducing visible compression artifacts.
Smart Images

Figure 0007839304000001 
Figure 0007839304000002 
Figure 0007839304000003
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for reconstructing a high dynamic range (HDR) image from a received low dynamic range (LDR) image including information and a tone mapping function necessary for decrypting a faithful reconstructed image of a master HDR image created and encoded on a transmission side at a reception side of a video communication system, and a corresponding encoding method.
Background Art
[0002] Until the first surveys around 2010 (and before the introduction of the first purchasable HDR decryption TVs in 2015), at least with respect to video, all videos were made according to a common low dynamic range (LDR) encoding framework also known as standard dynamic range (SDR). This has several characteristics. First, there was only one video made, and this video was good for all displays. This system was a relative system, and white was the maximum (100%) signal encoded using the maximum luma code corresponding to the maximum non-linear RGB values R' = G' = B' = 255 (255 in 8-bit YCbCr encoding). There was nothing brighter than white, and everything was a typical reflected color (for example, paper either reflects all incident light maximally or absorbs some of the red wavelength and reflects blue and green towards the eye, resulting in a cyan solid color, which by definition is somewhat darker than the white of the paper). Each display was technically constructed to display this whitest white (as a "drive requirement") as the brightest color, for example, at 80 nits (a simplified notation name for the SI quantity Cd / m^2) on a computer monitor and at 200 nits on an LCD display with a TL backlight. The observer's eye compensated for differences in brightness immediately, and thus, all observers saw almost the same image at home (despite differences in display) when not beside a store.
[0003] The goal was to improve the perceptible appearance of an image by creating pixels that emit a much brighter, more realistic light than the "white of paper," also known as "diffuse white," rather than simply the color printed or painted onto the paper.
[0004] Some systems, like the BBC's HLG, achieve this by defining values above white (where white is given a baseline level of "1" or 100%) in the encoded image, for example up to 10 times white, which can be defined to display pixels that shine 10 times brighter.
[0005] Currently, most systems are moving towards a paradigm where video creators can define the absolute nit value of their image (i.e., a value that is neither 2 nor 10 times the undefined white level, which is converted to a variable actual nit output on each endpoint display) based on the dynamic range capabilities of the selected target display. The target display is the video creator's virtual (intended) display, e.g., a 4000nit (ML_C) target display to define a particular video of 4000nits, while the actual consumer endpoint display has a lower display maximum brightness (ML_D), e.g., 750nits. In such a scenario, the end display still needs to include hardware or software for luminance remapping, which is typically implemented as luminance mapping, where pixel luminances in the HDR input image that are too high in luminance dynamic range (specifically their maximum luminance) to be faithfully displayed are somehow matched to values on the end display's dynamic range. The simplest mapping simply clips all luminances above 750 nits to 750 nits, but in that case, the beautiful structure of a 4000-nit sunset image with sun-lit clouds in the 1000-2000-nit range is clipped and disappears, appearing as a uniform white 750-nit patch, making this the worst way to handle dynamic range mapping. A better luma mapping shifts the 1000-2000-nit subrange in an HDR input image to, for example, the 650-740 dynamic range of the end display by a appropriately determined function (the function may be automatically determined within the receiving device such as a TV or STB, or determined by the video producer as being best suited to the video producer's artistic video or program and communicated with the video signal). Luma means encoding luminance using, for example, 10 bits, with a function that assigns luma codes from 0 to 1023 to video luminances from, for example, 0.001 to 4000 nits, using a so-called electro-optical transfer function (EOTF).
[0006] The simplest system is to simply transmit the HDR image itself (along with a properly defined EOTF), for example, an image with a maximum brightness of 4000 nits. This is what is done in the HDR10 standard. More advanced systems, such as HDR10+, also communicate a function to downmap the 4000 nit image to a lower dynamic range, such as 750 nits. These systems make this easy by defining a mapping function between two different maximum brightness versions of the same scene image, and then using an algorithm to calculate an endpoint function for the maximum brightness of the other display, which in turn calculates a modified version of that function. For example, if an SDR image is newly interpreted as an absolute nit image rather than a relative image, and it is agreed that an SDR image should always have a maximum pixel brightness of 100 nits, then the video creator can communicate together a function that specifies how to map the initial reference image grading, which is brightness from 0.001 (or 0) to 4000 nits, to the corresponding desired SDR brightness from 0 to 100 nits (which is a secondary reference grading), and this is called display tuning or adaptation. For both a 4000nit ML_C input image (horizontal axis) and a 100nit ML_C secondary grading / reference image, if we define a function that increases the darkest 20% of colors by, for example, 3 times in a plot normalized to 1.0, i.e., when reducing from 4000nit to 100nit, if it is necessary to reduce to 750nit on a particular end-user TV, the required increase would be, for example, only 2 times. (This depends on which EOFT definition is used for luminance. As mentioned above, luminance mapping is usually actually implemented as luminance mapping in the color processing IC / pipeline, and for example, using a psychovisually uniform EOFT makes it possible to define the effects of luminance changes along the range in a more visually uniform way, i.e., in a way that is more relevant and visually impactful to humans.)
[0007] A third class of even more advanced HDR encoders takes this to the next level by re-imaging these two reference grading images in a different way. Limiting the use to functions that are mostly reversible, an LDR image, which can be calculated at the transmitting end by, for example, downmapping the luminance or luma of a 4000-nit HDR image to an SDR image, can actually be transmitted as a proxy for the actual master HDR image created by the video creator, such as a Hollywood studio or sports broadcaster for BD or OTT distribution. The receiving device can then apply the inverse function to reconstruct a faithful reconstruction of the master HDR image. A system that communicates the HDR image itself (as it was created) is called a "mode HDR," and a system that communicates an LDR image is called a "mode LDR coder."
[0008] Figure 1 shows a typical example (for example, summarizing a principle previously patented by the applicant in WO2015 / 180854), which includes the decoding function itself and the subsequent display adaptation as a block, and these two techniques should not be confused.
[0009] Figure 1 schematically illustrates a video coding, communication, and processing (display) system. On the creation side, a typical embodiment of an encoder (100) is shown. Those skilled in the art will understand that first, a pixel-by-pixel processing pipeline for luminance processing is shown (i.e., every pixel of the input HDR image Im_HDR (typically one of several master HDR images of a video created by a video creator, with creation details such as camera capture and shading, or offline color grading being understandable to those skilled in the art and therefore omitted as they do not add to this explanation) is processed sequentially through the processing pipeline), and then, video processing circuits operating on the entire image are shown. For example, a similar compressor such as MPEG's DCT compression or AV1 operates on blocks of pixels. Without wanting to limit the teaching, let us assume that the input image is a master HDR grading created by a human color grader. After the human color grader selects the maximum luminance of the master grading video, they determine where to position the various video objects (with respect to average luminance) within the available range so that the image has the best effect on the consumer. For example, conventionally averaged objects will have a brightness comparable to the brightness they would achieve with LDR grading, but various types of HDR effect objects (such as clouds illuminated by bright sunlight, explosion fireballs, lampshades, or objects illuminated by a flashlight or sunlight) will have a range of brightness levels in higher sub-ranges of the available brightness range, for example, beyond the white level of a local scene. Some of these brightness levels change dynamically in various images, for example, when a dark corridor is gradually illuminated by switching on a series of ceiling TL lamps.
[0010] Figure 10 shows two different scene images from the video (dragon and suk, see Figure 10A), which the creator selected as best defined within the master HDR maximum luminance (ML_C_MH) luminance range of 1500 nits (i.e., the master is the best representation of this video, which the creator judged to be best because it brings good HDR effects to things like the dragon's fire breath, although it may require secondary and not optimal (re)grading to be as similar as possible). This can be imagined as first selecting a canvas shape for drawing and then starting to draw with the optimal configuration. Thus, in practice, for the master HDR video (PRIM_GRAD), a human grader (or automaton) selects various average luminances where the pixel luminances of different image objects are distributed around them. For example, diffuse reflection objects are typically bright below the selected level LowH, e.g., 220 nits (or, if the corresponding secondary (re)grading (SEC_GRAD) ends at the secondary grading maximum brightness ML_C_secG, the corresponding secondary limit level (lowS) for typically bright objects is considered to be, for example, 120 nits (roughly, this level corresponds to most of the pixel colors present in the SDR image when reinterpreted within each brightness range). A dragon, for example, becomes black at an average brightness of 10 nits, while the trees are around 50 nits. Yes. To achieve maximum effect despite its large area, the grader gives the flame a brightness distribution of approximately 800 nits. Specular reflections on a metal vase in sunlight within the suk image reach pixel brightness of up to 1300 nits, for example. Hanging lamps should also be above average brightness, but since they are not meant to attract attention, they are more modestly rated at, for example, 350 nits (in contrast to dragon flames, which can be seen by the observer even when they are only 20 nits, but in that case the flames don't look very realistic or impressive).
[0011] As shown in the two luminance range projection representations in Figure 10B, when regrading a secondary image with a narrower luminance dynamic range, for example, the grader needs to find the corresponding luminance position for all objects (in any image). For example, a very impressive flame is desired even within a limited range, and coincidentally there are not many things brighter and more prominent than the flame in this video, so the flame is projected to a 400nit level by object projection Fo_regrad. The tree is projected with equal luminance, and the dragon is made somewhat brighter, for example, considering that secondary grading SEC_GRAD is expected to have less ability to produce deep blacks.
[0012] As can be seen in Figure 10C, for practical consumer video communications (i.e., broadcast, film, etc.) or video conferencing purposes, this regrading can typically be summarized in a 2D plot by the shape of some luminance mapping function FL_regrad (sometimes shown normalized to a 1.0 plot). The left side of Figure 10B also shows that it is not always necessary to start with a master HDR image created as the starting image by the video creator. In some cases, it is necessary to start with an SDR image called an SDR master (MST_SDR) and obtain a master HDR image (i.e., PRIM_GRAD in Figure 10B) by using some upgrading algorithm. For example, a strong augmentation function (MeffBoo) is used for the brightest object in the dragon SDR image. Note that, in principle, when assigning luminance within the master HDR luminance range, the SDR image does not have a maximum luminance, but it can be assumed that the SDR image goes up to, for example, 100 nits.
[0013] Returning to Figure 1, we assume that the luminance of the HDR pixels, L_HDR, is sent through the selected HDR inverse EOTF in the luminance conversion circuit 101 (note that some systems already start working with luminance), and that the corresponding HDR luminance Y_HDR is obtained. It should be noted that the EOTF is a (typically fixed) function for obtaining luminance from a luminance code that encodes luminance, and should not be confused with a regrading function that typically has different optimal shapes for various scene images. For example, a perceptual quantizer EOTF is used. This input HDR luminance is luminance-mapped by the luminance mapper 102 to obtain the corresponding SDR luminance Y_SDR. In this unit, colorimetric knowledge is applied; that is, there is an input connection UI to a well-determined shape of the luminance mapping function (LMF). Roughly speaking, there may be two classes. Offline systems employ human color graders who, via color grading software, determine the best LMF according to their artistic preferences. LMF is assumed to be defined as a LUT defined using the coordinates of several nodes (for example, (x1, y1) for the first node), but is not limited to this. For example, if there is dark content in an image, and a human grader wants that content to be sufficiently visible when displayed on a low dynamic range display, for example, specifically on a 100nit ML_C image for a 100nit ML_D LDR display, the human grader sets the slope of the first line segment, i.e., the position of the first node.
[0014] The second class of embodiments, presented below, utilizes automata, where current technologies are crucial. These automata analyze an image (what luminances are present, where they are present, and to what extent) and propose the best LMF function shape. A particularly interesting automaton, the so-called ITM ("Inverse Tone Mapping"), does not analyze a master HDR image, but rather analyzes an input LDR image on the creator side and creates a pseudo-HDR image from this LDR image. The term "pseudo" here is intended to mean not a low-quality HDR image, but rather one calculated from an LDR image (which may, in some cases, be a high-quality LDR image with not too much clipping and not too much digitization higher than 8 bits, e.g., 10-bit or 12-bit Luma, DCT, or other compression artifacts) rather than being directly generated as the original master HDR by, for example, grading the RAW coding of the original camera capture of digital video by a human color grader incorporating computer graphics objects, etc. This is extremely useful because the majority of video is presented as LDR and may be created as SDR today or in the near future (at least, for example, some cameras in multi-camera productions output SDR, and for example, a drone captures side footage of a sports game, and this side footage needs to be converted to HDR format for the main program). By appropriately combining the capabilities of the Mode LDR coding system with the ITM system, the inventors and their technical partners were able to define a system that can be double-inverted. That is, the upgrading function for the pseudo-HDR image produced by analyzing the original LDR input is essentially the inverse function of the LMF used when coding an LDR communication proxy, which is typically an SDR image broadcast or unicast to a receiver, and can be re-graded by some of the receivers to closely approximate the master HDR grading when a luminance mapping function, which is communicated together in the metadata, is applied.In this way, it is possible to create a system that not only transmits the original (master) LDR image, but also transmits information for creating a good HDR image from it (automatically, or in other versions, with human input, such as fine-tuning of automatic settings, if desired by the content-creating customer). Therefore, in practice, the HDR master image of this video is being transmitted to the receiver (albeit by transmitting a proxy SDR video image).
[0015] The automaton uses all sorts of rules, such as identifying the location of light sources in an image, but the exact details are irrelevant to the description of this invention; it is simply that the automaton can generate some function LUP (the inverse function of LMF). The automaton's function is similarly input via the connection UI and can be applied in the LumaMapper 102. Note that in the description of the simplest embodiment, there is only one (downgrading) LumaMapper 102 present. This does not have to be a limitation. Since both EOTF and LumaMapping typically map a normalized input domain [0,1] to a normalized output domain [0,1], there are one or more intermediate normalization mappings that (effectively) map 0 inputs to 0 outputs and 1s to 1s. In such cases, the former intermediate LumaMapping then functions as the base mapping, and (secondarily) LumaMapper 102 then functions as a corrective mapping based on the first mapping.
[0016] The encoder has a set of LDR images, LumaY_SDR, corresponding to the HDR image LumaY_HDR. For example, the darkest pixels in a scene are defined to appear with substantially the same luminance on both an HDR and an SDR display, while brighter HDR luminances are reduced to within the upper limit of the SDR image, as indicated by the convex shape of the LMF function displayed in LumaMapper 102, thereby reducing the slope (or reconverging toward the diagonal of the normalized axis system). It should be noted that normalization is readily understood by those skilled in the art; it simply involves dividing the Luma code by power(2; number_of_bits). The normalized luminance can also be normalized, if necessary, by dividing any pixel luminance by the maximum value of its associated target display ML_C, e.g., 4000 nits.
[0017] Therefore, we can imagine examples of indoor and outdoor scenes. In the real world, outdoor brightness is typically 100 times brighter than indoor pixels, so in older LDR images, indoor objects are shown brightly and vividly, while everything outside the window is hard-clipped to a uniform white (i.e., invisible). Now, when communicating HDR video with a lossless proxy image, the bright outdoor areas visible through the window are made brighter (and sometimes less saturated), but in a controlled manner, enough information is still available for reconstruction into HDR. This is advantageous for both outputs, as systems that want to use the LDR image as is will see a good depiction of the outdoor scene to the extent possible within the limited LDR dynamic range.
[0018] Therefore, a set of Y_SDR pixel luminances (along with their chrominances) forms a "conventional LDR image" in the sense that subsequent circuitry does not need to consider whether this LDR image was intelligently generated or simply captured directly from the camera, as in older LDR systems (details of chrominance are unnecessary for this explanation). Thus, the video compressor 103 applies an algorithm such as MPEG HEVC or VVC compression. This algorithm is a set of data reduction techniques that, in particular, uses the discrete cosine transform to convert, for example, an 8x8 pixel block into a limited set of spatial frequencies, a technique that does not require much information to represent the spatial frequencies. The amount of information required is adjusted by determining the quantization coefficients that determine how many DCT frequencies are retained and how accurately they are represented. The drawback is that the compressed LDR image (Im_C) is not as precise as the input SDR image (Im_SDR), and block artifacts occur in particular. Depending on the broadcaster's choice, block artifacts can become so severe that, for example, some blocks in the sky are represented solely by their average brightness, appearing as uniform rectangles. This block artifact is usually not a problem because the compressor determines all its settings (including quantization coefficients) in such a way that quantization errors are barely noticeable or at least not a problem for the human system.
[0019] The formatter 104 executes any signal format required for the communication channel (the signal format differs depending on whether the communication is via storage on a Blu-ray disc, for example, or in the case of DVB-T broadcasting). Generally, all modified forms have the characteristic that the compressed video image Im_C is combined into the output image signal S_im along with a luma mapping function LMF that changes (or does not change) for each image.
[0020] The deformator 151 reverses the format so that the compressed LDR image and function LMF can be executed in later circuits to reconstruct the HDR image or perform other useful dynamic range mapping processes. The decompressor 152 reverses the compression, for example, VVC or VP9, and obtains a sequence of approximate LDR luma Ya_SDR that is sent to the inverse HDR image reconstruction pipeline. In addition, the upgrading luma masher 153 converts the SDR luma to the reconstructed HDR luma YR_HDR (this conversion uses the inverse luma mapping function ILMF, which is (effectively) the inverse of LMF). One of the explanatory diagrams shows two possible receiver (150) devices, where the receiver (150) device exists as a dual function within a single physical device (the end user can choose which parallel processing to apply), or some devices have only one of the parallel processing (for example, some set-top boxes only perform the reconstruction of the master HDR image and store that image in memory 155, such as a hard disk).
[0021] If a display panel is connected to the receiver embodiment, for example, in the case of a 750nit ML_D end-user display 190, the receiver has a display adaptation circuit 180 that calculates a 750nit output image instead of a reconstructed image of, for example, 4000nit (this is shown with a dotted line to indicate that it is an optional component unrelated to the teachings of the present invention, but is often used in combination). We will not go into detail about the many variations that can achieve display adaptation, but typically there is a function determination circuit 157 which proposes an adapted luma mapping function F_ADAP (usually close to diagonal) based on the inverse shape of the LMF. This function is loaded into a display adaptive luma mapper 156 which calculates a lower intensity HDR luma L_MDR that typically ends at ML_D=750nit instead of ML_C=4000nit with a smaller dynamic range. [Overview of the project] [Problems that the invention aims to solve]
[0022] Problems arise when an ITM automaton determines an LMF function that has soft clipping, or worse, hard clipping, for the brightest luminance. Soft or hard clipping means that the slope of the highest line segment (or tangent to the curve) is small, especially when measured with respect to the typical settings of the compressor. For example, if the compressor blocks out the sky, which is the brightest object in the LDR image, the decoder's inverse LMF function will have a large slope for that highest line segment or part of the curve. This increases the visibility of artifacts that should be invisible according to normal compression principles, leading to undesirable artifacts. While it is possible to compel a human grader or automaton to use only functions with a sufficiently large slope for the brightest part of the curve, this compulsion is undesirable for some images from an artistic standpoint, or for automatons that are not easily programmed into the set of rules.
[0023] Therefore, a general solution to the problem, which the inventors have identified as a problem to address, is desired. This solution also has the advantage of not requiring the modification of existing algorithms for determining the optimal shape of the desired lumens mapping function (LMF) for any HDR scene image, as well as the brightness characteristics and composition of its constituent objects, using artificial intelligence, for example (especially when obtained from ITM).
[0024] EP3621307 defines a system that can encode a higher-quality master HDR image by calculating a proxy HDR image with a lower maximum coded brightness (for example, the proxy video sent to the receiver may reach or potentially reach a maximum pixel brightness of 800 nits, which is assumed to represent the original pixel brightness at the same location up to 2000 nits). Furthermore, a transformation based on some scale value of the brightness mapping function that determines how to downgrade the received 800-nit proxy image to an even lower dynamic range, i.e., the maximum brightness image, is stretched upwards, also defining a regrading relationship between two HDR images, namely the creator's master HDR image and the corresponding presentation-time proxy image. This deformation typically moves the curve (while retaining its shape) across the entire input and output range so that it approaches the diagonal of the Luma plot, i.e., so that the curve becomes more linear or approximates the identity curve (i.e., this deformation is unrelated to clipping behavior; note that whether or not any clipping behavior is observed on any receiving display is independent of the clipping behavior on the generating or encoding side).
[0025] WO2014 / 128586 describes one possible method of HDR video encoding, namely, creating a technically imperfect SDR proxy and adding a function to the metadata that is communicated together to convert it to a secondarily better-looking SDR video. [Means for solving the problem]
[0026] An encoder (100) for encoding an input high dynamic range image (Im_HDR_PSEU) as encoded data (S_im), wherein the encoded data includes, firstly, a matrix of pixel colors (Y_SDR, Cb_SDR, Cr_SDR) of an image (Im_SDR) with a lower dynamic range than the high dynamic range image, and secondly, image metadata (SEI) including a luma mapping function for calculating the high dynamic range pixel luma (Y_HDR) of the high dynamic range image by applying a luma mapping function to the pixel luma of the low dynamic range image, The encoder has an input (997) for receiving a luma mapping function (LMF) from a connectable or provided inverse tone mapping system (200), the inverse tone mapping system is configured to derive an upgrading function (LUP) for calculating a high dynamic range image (Im_HDR_PSEU) from a master low dynamic range image (Im_LDR_mastr) by applying an upgrading function to the luma of the master low dynamic range image (Im_LDR_mastr) based on the analyzed characteristics of the master low dynamic range image (Im_LDR_mastr), the inverse tone mapping system is configured to obtain a luma mapping function by inverting the upgrading function (LUP), and in the encoder (100) for encoding a high dynamic range image, The mapper splitting unit (901) is configured such that a symbolizer composes a luma mapping function (LMF) from two data items, where the first data item defines an extended luma mapping function (LMF_HS) by mapping a clipping point (PtCli) at which the luma mapping function (LMF) first reaches its maximum output value (Vomax) (in the immediate vicinity of the maximum output value, or in many cases exactly at the maximum output value) to an end point (Ptmax) having as coordinates the maximum value of the input values and the maximum value of the output values corresponding to the horizontal scaling factor, and by mapping all points of the luma mapping function having input coordinates lower than the input coordinates of the clipping point to extended input coordinates equal to the input coordinates multiplied by the horizontal scaling factor while maintaining the output coordinates of the respective curve points, the second data item is a scale factor preferably equal to the reciprocal of the horizontal scaling factor, and the symbolizer has a formatter (104) configured to output this extended luma mapping function (LMF_HS) and this scaling value (SCAL) as metadata of a low dynamic range image (Im_SDR) to be output as well, a symbolizer for encoding a high dynamic range image, and a corresponding decoder that uses the inverse operation of the first luma mapping and then scaling using the received parameters defining that operation When used, decoding artifacts are easily avoided.
[0027] When this operation handles lumens, this solution can work with both relative HDR systems, i.e., systems that encode a certain degree of excessive brightness as a multiple of the diffuse (LDR) white level by the maximum lumens code, and absolute HDR coding systems defined by the display, which encode the exact pixel brightness (by the corresponding lumens code) up to the maximum encoded brightness ML_C (e.g., ML_C = 5000 nits, or 1000 nits) that is intended to be displayed on some target display. Proxy SDR images are images of a nature that allows all HDR colors to be calculated with sufficient accuracy by using the lumens mapping function. Naturally, there are slight rounding errors, but when using, for example, 3 x 10 bits for the Y, Cb, and Cr pixel color images, the system has been demonstrated to function correctly, except for issues addressed by current improved techniques. Other communication standards or channels, such as HDMI®, have their own placeholders for communicating metadata instead of SEI messages, which are the standard MPEG mechanism for communicating specific metadata as needed. Encoders exist in professional systems such as television production studios, or in consumer devices such as those used to upload videos taken with mobile phones to social networking sites. Inverse tone mapping (ITM) systems are systems that perform the "inverse" (not strictly a mathematical inverse, but a functional inverse) of conventional tone mapping, which is defined to reduce high dynamic range images with a wide luminance range to standard or low dynamic range luminance. Thus, an ITM system creates some corresponding HDR image (also a pseudo-HDR image, since it was not originally created as an HDR image) from an input LDR image, which in this document is called a master LDR image (using the same nomenclature as a master HDR image). Therefore, a system that uses, for example, a heuristic rule program or other techniques such as machine learning techniques to maintain the luminance (or their coding luma) of reflective objects as is, but increases the luminance of luminescent objects such as the sun, or increases the luminance of pixels of outdoor objects compared to indoor pixels, may be called an ITM.In the present invention, the ITM is limited to a variant form that defines the upscaling to the HDR image by using a luma mapping function, typically by using only the luma mapping function. The extended function maintains the shape, i.e., the offset above or below, for example, compared to the diagonal, at least up to some selected endpoints (losing the clipping part), and is a function extended in the size of a certain dimension. For example, the extended function may be stretched horizontally.
[0028] Advantageously, an encoder for encoding a high-dynamic-range image comprises a clipping detection circuit (904) configured to detect whether (Im_HDR_PSEU) has a portion clipped to the maximum output in the input range. According to the core principle, the original LMF function is divided into two luma processing operations. That is, in one operation, a normal function that is mapped to the output maximum only when the input maximum is reached is used, whereby, as a secondary characteristic, it advantageously becomes closer to the diagonal, improving coding and reversibility. It is possible to construct a variant form that processes some bad LMF curves by checking whether there is clipping, or performs some processing on all curves without checking, but continues the mapping in the same way when encountering a function that has already mapped 1.0 to 1.0, for example.
[0029] Advantageously, an encoder for encoding a high-dynamic-range image functions for an image defined on an absolute nit dynamic range where the maximum luminance is the end.
[0030] Advantageously, an encoder for encoding a high-dynamic-range image functions in a coding system in which a low-dynamic-range image (Im_SDR) is predefined as a low-dynamic-range image having a maximum luminance equal to 100 nit.
[0031] Advantageously, an encoder for encoding high dynamic range images has a mapping division unit (901) which determines a stretched Luma mapping function (LMF_HS) by performing a horizontal stretch consisting of linear scaling such that the input coordinates (XC) where clipping first occurs are mapped to the maximum normalized input coordinates.
[0032] A useful new technical principle can also be embodied as a method for encoding high dynamic range images (Im_HDR_PSEU). The high dynamic range image is represented by image metadata (SEI), which includes, firstly, a matrix of pixel colors (Y_SDR, Cb_SDR, Cr_SDR) of images with a lower dynamic range than the high dynamic range image (Im_SDR), the low dynamic range image being compressed for communication as a compressed low dynamic range image (Im_C), and secondly, a lumen mapping function for calculating the high dynamic range pixel lumens (Y_HDR) of the high dynamic range image by applying a lumen mapping function to the pixel lumens of the low dynamic range image. In a method in which a luma mapping function (LMF) is received from an inverse tone mapping system (200), and the inverse tone mapping system (200) is configured to derive a luma mapping function for constructing a corresponding high dynamic range image (Im_HDR_PSEU) based on the analyzed characteristics of a master low dynamic range image (Im_LDR_mastr), The method comprises the step of dividing a luma mapping function into two consecutive operations, the second of which includes converting the luma mapping function into an expanded luma mapping function (LMF_HS) having a shape that maps a maximum normalized input to a maximum normalized output, the first of which includes linear scaling using a scaling value (SCAL), and the method is configured to output the expanded luma mapping function (LMF_HS) and the scaling value (SCAL) as metadata for a similarly output low dynamic range image (Im_SDR).
[0033] The method includes the step of detecting whether (Im_HDR_PSEU) has a portion of the input range that is clipped to the maximum output.
[0034] The method works on images defined by an absolute nit dynamic range where the maximum brightness is at the end, more specifically, a low dynamic range image (Im_SDR) is a low dynamic range image with a maximum brightness equal to 100 nits (and an HDR image has an ML_C equal to what is selected by the video creator, e.g., the person adjusting the ITM settings, to result in an HDR image of a quality / impression such as ML_C = 2000 nits or 10,000 nits). The method determines the stretched luma mapping function (LMF_HS) by performing a horizontal stretch consisting of linear scaling so that the input coordinates (XC) where clipping first occurs are mapped to the maximum normalized input coordinates.
[0035] In particular, those skilled in the art understand that these technical elements can be embodied in various processing elements such as ASICs (Application-Specific Integrated Circuits, i.e., typically IC designers causing an IC (or part of an IC) to perform a method), FPGAs, programmed processors, etc., and can reside in various consumer or non-consumer devices, whether they have a display (e.g., a mobile phone encoding consumer video) or a non-display device that can be externally connected to a display, and that images and metadata can be communicated via various image communication technologies such as wireless broadcasting and cable-based communications, and that devices can be used in various image communication and / or usage ecosystems such as television broadcasting, on-demand via the Internet, video surveillance systems, and video-based communication systems.
[0036] These and other aspects of the methods and apparatus according to the present invention will become apparent from and be described with reference to the implementations and embodiments described below and the accompanying drawings, the accompanying drawings serving only as non-limiting specific illustrations illustrating more general concepts, and dashed lines in the accompanying drawings are used to indicate that components are optional, with components without dashed lines not necessarily being required. Dashed lines may also be used to indicate elements that are described as required but are hidden within an object, or intangible things such as the selection of an object / region. [Brief explanation of the drawing]
[0037] [Figure 1] This diagram schematically illustrates two possible functionalities / embodiments of an HDR image (especially video) coding system that encodes HDR images by transmitting an LDR image that acts as a proxy, rather than transmitting the HDR image itself. The LDR image can be calculated from the HDR image using a substantially reversible luma mapping function, and vice versa. The LDR image is typical of a color quality that can be used directly (without requiring further color adjustment) on older LDR displays. [Figure 2] This figure shows an improved encoder for such an HDR video communication system. The encoder is coupled to or comprises an ITM circuit for generating a visually appealing pseudo-HDR image based on an input master LDR image and corresponding luma mapping curves for converting one image to a corresponding other image. [Figure 3] This diagram schematically illustrates the types of undesirable functions (in the form of HDR to LDR downgrading mapping) that are often determined by certain ITM (In-the-Motion) transformations. These functions have the undesirable characteristic of starting clipping too early, which leads to reversibility problems. [Figure 4] This diagram schematically illustrates the reversibility issues that arise. No matter how much one attempts to regularize the decoding, there is always a high slope to the luma of the brightest image, which is particularly problematic when used in high levels of video compression, as it results in strong block artifacts (e.g., in LDR images) that should normally be imperceptible, but become less imperceptible and less unpleasant when using better encoding ( / decoding) systems such as those taught in this patent application. [Figure 5] This figure shows the first potential improvement to the luma mapping curve from ITM. However, it is not yet perfect, as it itself introduces another color artifact regarding the balance of brightness of image objects. [Figure 6] This figure shows an embodiment in which the diagonals are stretched after rotation. This embodiment is a preferred method for forming a better reversible luma mapping curve that maps to the maximum output only when the maximum input is reached, and is used in conjunction with a rescaling operation in the decoder to obtain substantially the same grading effect as when using the luma mapping function LMF. [Figure 7] This diagram illustrates the concept of expansion (scaling and different / expanded Luma mapping functions) in an encoder. [Figure 8]This figure shows the behavior of the decoder reversed from that of the encoder. The decoder first applies a new, expanded luma mapping function as taught by embodiments of the present invention, and then applies a corrective scaling linked to the expansion of the original luma mapping function (so that the shapes of the functions overlap). [Figure 9] This is a schematic diagram of a unit typically found in a novel encoder that follows the principle of coding by the innovative ordered luma mapping of the present invention. [Figure 10] This figure shows several examples of regrading the typical (e.g., average) brightness of several image objects along a range of various possible brightness levels. [Modes for carrying out the invention]
[0038] In Figure 2, the encoding side is shown (as already explained in Figure 1), but the ITM (Inverse Tone Mapping) system (200) is also shown here. The ITM system (200), as explained, starts from a master LDR image (Im_LDR_mastr) rather than a master HDR image (i.e., an image originally created as, for example, 4000 nit HDR), and the ITM can create a good-looking HDR image (Im_HDR_PSEU) from that master LDR image. The arrow for the ITM system (200) is shown as a dotted line because the core system does not actually need to calculate Im_HDR_PSEU, but only needs to calculate the best-calculated upgrading function LUP. However, it should be noted that some variations of the ITM may actually calculate the HDR output image and calculate some of its properties (e.g., histograms of regions, relationships between brightness levels of different regions, texture measurements, outlines of illumination, etc.). The input device that provides the master LDR image is also shown with a dashed line, because there may be several systems in which the present invention should be incorporated into the input device. A typical application is to acquire real-time images from a camera (or more formally, one or more cameras, the images from which are mixed, for example, into a mixed master LDR video) and to roughly set their luminance characteristics (e.g., average luminance) by the shader 205. Another example is when an already completed LDR video is acquired from memory, for example, old video. The luma mapping function derivation unit 201 derives the optimal upgrading function (LUP) for the incoming LDR image. This function is inverted by the inverter 202, which yields the inverse function of the upgrading function LUP as the output function LMF_out, and this inverse function serves as the luma mapping function LMF for a mode LDR-based encoder. However, it should be noted that (again, in the case of 101 and 102, which are again dashed lines) actual HDR to LDR downgrading is not required.The reason is that the ITM system already handles this, with the master LDR image Im_LDR_mastr acting as a proxy image communicated in place of the native HDR image, and the inverse function of the LUP function being communicated together with metadata (e.g., SEI message) to reconstruct a close version of the (pseudo) HDR image IM_HDR_PSEU as the HDR image.
[0039] As schematically shown in Figure 3, the inventors found that some ITM variants give LMF curves with hard clipping to some input images. That is, when the normalized input of the luma modifier 102 is equal to xc (e.g., xc=0.85 for a perceptually homogenized luma that is approximately logarithmic), it maps to the maximum output of the LDR proxy image (i.e., normalized and represented as 1.0, which corresponds to the 8-bit code value Y=R'=G'=B'=255). In such a function, all inputs higher than, for example, 0.85 are written to an LDR image matrix with Y=255, and the communicated function LMF has a horizontal slope greater than XC, for example, 0.85. As shown in Figure 4, a typical mode LDR decoder will encounter the associated problems when reconstructing the HDR output image. Theoretically, there is an infinite slope with the LDR luma Y=255 or normalized to 1.0, which in principle means that the decoder cannot function correctly with such an ILMF input function received in the metadata of the image being processed. In practice, heuristic mitigation measures are employed in the decoder. For example, the decoder uses a quadratic curve ILMF2 derived from ILMF (or, in practice, often, LMF, which is the actual function communicated in the metadata). Such an ILMF2 curve typically has the same values (x,y) as ILMF for all points on its curve trajectory except for the highest value. One embodiment is shown in which the highest subrange of the curve is divided into two parts, one with a small slope and the other with a steep (but not infinite) slope. However, this still leads to significant visual artifacts, i.e., a lot of compression artifacts, which are significantly amplified, especially at the brightest ends of the LMF range, making them even more visually distracting.
[0040] As illustrated in Figure 5, the inventors realized that rotating the LMF function diagonally (beyond the angle ROTH) results in a clipping point where the horizontal segmentation begins, which lies diagonally and has equivalent x and y coordinates (resulting in the rotated luma mapping function LMF_R). This already yields a clearly good reversible function, as the curve now approaches the identity transformation. However, the problem is that its output is not exactly the output desired for this image or this particular HDR scene. For example, instead of the pseudo-HDR luma value XC becoming the whitest white on an LDR display (as it usually appears), a higher value XE simply becomes the same brightness as YE (e.g., 240 in 8 bits). Only a brighter value, i.e., 1.0 (which in some situations doesn't even exist in the pseudo-HDR image), becomes the whitest LDR color. Therefore, the resulting image is too dark. One of the advantages, or even the purpose, of a Mode LDR coding / decoding system is to enable consumers still viewing on older LDR displays to obtain a very suitable simulated LDR image corresponding to the HDR image being ultimately transmitted. However, at present, this purpose is hindered because consumers end up seeing an image that is too dark. We can see how the clipping point (the first point in time when the curve reaches 1.0, or the point when it reaches a value very close to 1.0 if some relaxation is allowed, such as removing some of the brightest values in a scenario where soft clipping is performed) can be projected onto the diagonal by rotation (ROTH) to obtain the diagonal point PtDi and ultimately the endpoint Ptmax (indicated by an extended arrow starting from a point, which can be equivalently represented by an arrow starting from (0,0)).
[0041] Therefore, as shown in Figure 6, we considered performing further stretching along the diagonal (DSTR) to brighten the resulting output LDR image again (thus approaching the originally determined optimal LMF function, but with a clipping portion at the upper limit). In fact, by further studying this system, we found that the rotation ROTH and diagonal stretching DSTR substantially form an equilateral triangle, so the stretching operation after rotation substantially reacquires the Y values of the curve's nodes, resulting in a rotated and stretched LMF mapping function LMF_RS. Only the X point is in a different position, which is still not perfect. Another characteristic of this curve is that, except for the horizontal clipping portion, it still retains its original shape (although stretched), which, as explained, is undesirable, at least for the ultimate goal of decoding HDR images from communicated LDR proxy images with good quality (it is more important to build a high-quality, innovative HDR codec that allows viewing of near-perfect HDR images on future top-quality HDR displays, rather than building a high-quality, innovative HDR codec just to be able to receive good quality LDR images).
[0042] However, this substantial equivalence of y-values led to further insight: eliminating the second dimension and performing everything in one dimension. In other words, horizontal stretching of the curve (HSTR) can be easily performed.
[0043] This results in the ability to use 1D mapping, as shown in Figures 7 and 8, and furthermore, improved coding techniques (for corresponding mirror decoding) can be designed from the 1D mapping.
[0044] As shown in Figure 7, the entire regrading operation is illustrated in the plot between the normalized input Norm_in (this normalized input has HDR characteristics; i.e., these lumas encode a 2000nit pseudo-HDR image that can be obtained, for example, from the corresponding master LDR image by ITM) and the normalized output (this normalized output is an equally normalized 8-bit or 10-bit LDR luma).
[0045] If you want the properties of this LMF_HS mapping curve (which corresponds to LMF_RS in the rotation and stretch embodiments, but in the preferred embodiment arises only from horizontal stretching that maps the XC point of the LMF curve obtained from ITM to 1.0 on the horizontal axis of the input range) to map 1.0 to 1.0 instead of, for example, XC=0.8 to 1.0, then you should define the point 0.8 as a new "1.0" in some way.
[0046] This is done by a pre-scaling operator in the novel encoder, which addresses improper ITM LMF functions by correcting them as needed (i.e., if there are hard-clipping portions in the curve). Norm_in_NEW=Norm_in*(1 / XC) [Formula 1] A linear scaling operation that defines this can be used.
[0047] This can be seen as two units of luma mapping when the decoder behaves accordingly (mirror-symmetrically), as shown in Figure 8.
[0048] If the decoder is aware of the encoding's scaling factor SCAL (or any value from which SCAL can be calculated, e.g., the horizontal coordinate XC of the LMF function's clipping start), the decoder can use the mapped block. Thus, the decoder first maps the 1.0 input to the 1.0 output and applies the inverted horizontal stretching luma mapping function ILMF_HS, which has the shape of the inverted LMF_RS. Then, the second block performs compression scaling using the value SCAL (or any value related to SCAL that enables the receiver to calculate SCAL) communicated in the metadata, resulting in a final value of 1 being mapped back to XC. This XC value is, for example, the HDR luma defined in the perceptual quantizer EOTF (a visually homogenized luma system, i.e., a system designed so that arithmetic steps of luma give humans substantially equal visual effects across the entire luma scale) encoding the intended brightness of 2500 nits. This value indicates how an HDR image would appear, for example, if the user had a 4000-nit or 5000-nit display, which could normally display such an image by directly showing the intent of the pixel brightness in its precise nit representation. In embodiments with display adaptation, the system takes this into account and calculates the best possible approximation of, for example, a 4000-nit image with the dynamic range capability of, for example, a 750-nit display. In fact, what is required as the decoder output is not a very high HDR output, but a value of DES or near that in Figure 4. It should be noted that if an intermediate luma mapping is present, this strategy will not yield exactly the same results as the original ITM clipping mapping function, except for correction of the infinite decoding slope behavior, and there will be slight differences. Studies have shown that in practice this difference is not a real problem of concern. If further improvement is still considered necessary, an innovative encoder may calculate a slightly different SCAL value, and as a result, even when an intermediate mapper is present, the LMF_HS curve will overlap more closely with the original LMF curve when used with scaling.In practice, once the concept of partitioning is formalized, the full shape details of the LMF_HS function are communicated to the decoder as metadata, so the encoder embodiment can also adjust the shape of the LMF_HS function to some extent as needed.
[0049] A preferred method for defining the SCAL value is one in which the SCAL value is used directly on the decoding side and can be scaled down based on some luma definition (e.g., PQ luma) (typically by upgrading the received SDR proxy image and then scaling it down to below the theoretical maximum brightness). Thus, for example, if the value 0.89 is mapped to 1.0 on the encoder side, the SCAL value can be defined as the value to which 1.0 should be (re)mapped on the decoding side, i.e., 0.89 (the x-coordinate of the first clipping point PtCli (typically a normalized input value, normalized to have a maximum value of 1.0)). Note that the clipping detector can function, for example, for a function by ensuring that all higher input values are mapped to the maximum normalized output value or y-coordinate having a value of 1.0. Note that since the function and the input SDR image define a pseudo-HDR image, i.e., it is coding a pseudo-HDR image, there is no need to verify anything about the pseudo-HDR image that can be generated by ITM. Expanded HSTR corresponds to multiplying the input value coordinate of each point on a curve, which has an output value (y-coordinate) fixed between 0 and 1, by a multiplier equal to 1 / SCAL. Thus, for example, an x-coordinate equal to 0.5, also called the input value of the original LMF curve, is moved to the input value of the LMF_HS curve, which is 0.5 * 1.1236 = 0.56 (by multiplication by the reciprocal 1 / SCAL, also called the reciprocal scaling factor), and that point has the same y-coordinate as the original point on the LMF curve.
[0050] Therefore, as shown in Figure 9, if the decoder uses these two consecutive luma mapping stages, the new ITM-improved encoder can handle this.
[0051] Figure 9 shows one embodiment of an improved encoder based on current innovative insights. In this case as well, most of the technical units are substantially as described above.
[0052] The innovative mapping splitting unit 901 inserts, instead of a simple luma mapping function, a luma mapping function adjusted using one of the embodiments described with Figures 3 to 8, for example, a horizontally stretched luma mapping function LMF_HS (or a rotated and stretched luma mapping function resulting from that embodiment), into the metadata associated with the image (or set of images to be mapped). The mapping splitting unit 901 also outputs a corresponding scale factor SCAL in the metadata, and as a result, the combined operation of these two functions substantially maps the result of mapping using only the original LMF obtained from the ITM, but with improved behavior in the decoder. In addition, the mapping splitting unit 901 includes a function reshaping circuit 902 configured to determine the stretched version of the LMF so that the maximum normalized input (1.0) maps to the maximum normalized output. The mapping splitting unit 901 also includes a scaling factor determination unit 903 configured to determine and output the scale factor SCAL. Cb_SDR and Cr_SDR are the blue and red chroma components of the pixel colors of the LDR proxy image Im_SDR in a typical color coding, respectively (those skilled in the art will recognize how the pixel colors can be reformatted using known colorimetric techniques, for example, to R'G'B' coding). The image signal output S_im for communication to the receiver includes the compressed LDR image Im_C and metadata (SEI) including LMF_HS and SCAL. Furthermore, the function reformatting circuit 902 is typically embodied to analyze whether the LMF function requires reformatting, for example, the function reformatting circuit 902 includes a clipping detection circuit 904 configured to detect whether the LMF function of ITM is clipped to a maximum value, i.e., typically, whether it is a function that is primarily strictly increasing, or whether input values less than 1.0 are already mapped to an output of 1.0. If function reformatting is not required, the function reformatting unit can be embodied to perform an identity transformation. Various units are physically combined.
[0053] The scaling factor relates to the starting point of clipping (up to the maximum output value) and should not be confused with other scaling factors in HDR technology. For example, it is a value that adaptively controls how much a curve should move toward or away from the diagonal in a normalized luma plot, and this corresponds to the identity transformation. (For example, a 1000 nit maximum brightness image is optimal for a 1000 nit maximum capability display, and therefore, there is no need to apply mapping to these image lumas to obtain a driving image for the display).
[0054] The algorithmic components disclosed herein are actually implemented (in whole or in part) as hardware (e.g., as part of an application-specific IC) or as software that runs on a specialized digital signal processor or general-purpose processor.
[0055] Those skilled in the art should be able to understand from this presentation which components are optional improvements and can be realized in combination with other components, and how the (optional) steps of the method correspond to each means of the apparatus, and vice versa. The term “apparatus” in this application is used in its broadest sense, namely a set of means that enable the achievement of a particular purpose, and therefore may be, for example, an IC (a small circuit portion thereof), a dedicated device (such as a device with a display), or part of a networked system. “Arrangement” is also intended to be used in its broadest sense, and therefore “arrangement” includes, among other things, a single device, a part of an apparatus, a collection of interconnected devices (or part of a collection of interconnected devices).
[0056] The term "computer program product" should be understood to encompass all physical realizations of a set of commands that enable a general-purpose or dedicated processor to execute one of the characteristic functions of an invention by inputting commands into the processor after a series of load steps (including intermediate translation steps such as translation into an intermediate language and a final processor language). In particular, computer program products can be realized as data on a carrier such as disk or tape, data residing in memory, data traveling over a network connection (wired or wireless), or program code on paper. Apart from the program code, characteristic data required for a program can also be realized as a computer program product.
[0057] Some of the steps required for the operation of the method, such as data input and output steps, are already present in the processor's functionality, rather than being described in the computer program product.
[0058] It should be noted that the embodiments described above are illustrative and not limiting to the present invention. For the sake of brevity, not all of these options are described in detail where a person skilled in the art can easily map the examples presented to other areas of the claims. Apart from the combinations of elements of the present invention combined in the claims, other combinations of elements are also possible. Any combination of elements may be realized in a single, dedicated element.
[0059] None of the parenthetical reference numerals in the claims are intended to limit the scope of the claims. The word “equipped with” does not preclude the existence of elements or aspects not listed in the claims. The singular form of an element does not preclude the existence of multiple such elements.
Claims
1. An encoder for encoding an input high dynamic range image (Im_HDR_PSEU) as encoded data (S_im), wherein the encoded data includes, firstly, a matrix of pixel colors (Y_SDR, Cb_SDR, Cr_SDR) of an image (Im_SDR) with a lower dynamic range than the high dynamic range image, and secondly, metadata (SEI) of the image, which includes a lumens mapping function for calculating the high dynamic range pixel lumens (Y_HDR) of the high dynamic range image by applying the lumens mapping function to the pixel lumens of the low dynamic range image. In an encoder for encoding a high dynamic range image, the encoder has an input for receiving a luma mapping function (LMF) from a connectable or provided inverse tone mapping system, the inverse tone mapping system derives an upgrading function (LUP) for calculating the high dynamic range image (Im_HDR_PSEU) from the master low dynamic range image (Im_LDR_mast) by applying an upgrading function to the luma of the master low dynamic range image (Im_LDR_mast) based on the analyzed characteristics of the master low dynamic range image (Im_LDR_mast), and the inverse tone mapping system obtains the luma mapping function by inverting the upgrading function (LUP), The encoder comprises a mapping division unit that comprises two data items for the luma mapping function (LMF), wherein the first data item is an expanded luma mapping function (LMF_HS), defined by mapping the clipping point (PtCli) where the luma mapping function (LMF) first reaches its maximum output value (Vomax) to an endpoint (Ptmax) having the maximum input value and the maximum output value as coordinates, corresponding to a horizontal scaling coefficient, and by mapping all points of the luma mapping function having input coordinates lower than the input coordinates of the clipping point to expanded input coordinates equal to the value obtained by multiplying each input coordinate by the horizontal scaling coefficient, while maintaining the output coordinates of each curve point. An encoder for encoding a high dynamic range image, characterized in that a second data item is a scale coefficient preferably equal to the reciprocal of the horizontal scaling coefficient, and the encoder has a formatter that outputs the expanded luma mapping function (LMF_HS) and scaling value (SCAL) as metadata for the low dynamic range image (Im_SDR) that is similarly output.
2. An encoder for encoding a high dynamic range image according to claim 1, comprising a clipping detection circuit for detecting whether the luma mapping curve (LMF) supplied to the input by the inverse tone mapping system has a portion of the input range that is clipped to the maximum output.
3. An encoder for encoding a high dynamic range image according to claim 1 or 2, wherein the image is defined on an absolute nit dynamic range where the maximum brightness is at the end.
4. An encoder for encoding a high dynamic range image according to claim 3, wherein the low dynamic range image (Im_SDR) is a low dynamic range image having a maximum brightness equal to 100 nits.
5. A step of encoding an input high dynamic range image (Im_HDR_PSEU) as encoded data (S_im), wherein the encoded data includes, firstly, a matrix of pixel colors (Y_SDR, Cb_SDR, Cr_SDR) of an image (Im_SDR) with a lower dynamic range than the high dynamic range image, and secondly, metadata (SEI) of the image, which includes a luma mapping function for calculating the high dynamic range pixel luma (Y_HDR) of the high dynamic range image by applying the luma mapping function to the pixel luma of the low dynamic range image. A step of receiving a Luma Mapping Function (LMF) from a connectable or equipped inverse tone mapping system, wherein the inverse tone mapping system derives an upgrading function (LUP) for calculating the High Dynamic Range Image (Im_HDR_PSEU) from the master low dynamic range image (Im_LDR_mast) by applying an upgrading function to the luma of the master low dynamic range image (Im_LDR_mast) based on the analyzed characteristics of the master low dynamic range image (Im_LDR_mast), and the inverse tone mapping system obtains the Luma Mapping Function (LMF) by inverting the upgrading function (LUP). In a method including, The method is a step of dividing the luma mapping function (LMF) into two data items, wherein the first data item is an expanded luma mapping function (LMF_HS), defined by mapping the clipping point (PtCli) where the luma mapping function (LMF) first reaches its maximum output value (Vomax) to an endpoint (Ptmax) having the maximum input value and the maximum output value as coordinates, corresponding to a horizontal scaling coefficient, and by mapping all points of the luma mapping function having input coordinates lower than the input coordinates of the clipping point to expanded input coordinates equal to the value obtained by multiplying each input coordinate by the horizontal scaling coefficient, while maintaining the output coordinates of each curve point. The second data item is a scale factor which is preferably equal to the reciprocal of the horizontal scaling factor, and the step is as follows: The steps include outputting the expanded luma mapping function (LMF_HS) and scaling value (SCAL) as metadata for the low dynamic range image (Im_SDR) that is output similarly, and A method characterized by including
6. The method according to claim 5, further comprising the step of detecting whether the luma mapping function has a portion of the input range that is clipped to the maximum output.
7. The method according to claim 5 or 6, wherein the image is defined on an absolute nit dynamic range where the maximum brightness is at the end.
8. The method according to claim 7, wherein the low dynamic range image (Im_SDR) is a low dynamic range image having a maximum brightness equal to 100 nits.
Citation Information
Patent Citations
Multi-range HDR video coding
EP3621307A1
Method and device for determining control parameters for mapping an input image with a high dynamic range to an output image with a lower dynamic range
EP3672219A1
Method and apparatus for encoding HDR images
JP2022058724A
Methods and apparatuses for encoding an HDR images, and methods and apparatuses for use of such encoded images
WO2015180854A1