HDR format conversion for processing
Patent Information
- Application Number
- JP2026510781
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-24
- Filing Date
- 2024-08-20
- Publication Date
- 2026-09-09
Smart Images

Figure 2026530596000001_ABST
Abstract
Description
[[Technical Field]]
[0001] The present invention relates to luminance-varying color processing for high dynamic range images, and particularly to narrow range coding. [[Background Art]]
[0002] For more than half a century, the creation, transmission and display of video (that is, temporally consecutive sequences of images) has been relatively straightforward. Silver-based film production could only send copies to a limited number of locations. In the first half of the 20th century, it became possible to capture images as analog voltage signals, transmit the signals via radio or cables, and finally display them on cathode-ray tube displays. Around the turn of the century, analog signals came to be digitally represented with 3×8 bits of color code per pixel, giving rise to the digital video that is today called low dynamic range (LDR), also known as standard dynamic range (SDR).
[0003] CRTs were uncharacterized displays with variable characteristics, but one aspect that remained constant was that the range of displayable brightness was relatively low. Although the luminance range of every CRT display differed somewhat, on average CRTs could display the brightest color (pure white) at approximately 100 nits. The darkest black depended on observation conditions, and might have been up to 10 nits, but could reach as low as 0.1 nits under ideal dim viewing room conditions.
[0004] As shown in Figure 1, although it has been considered practically sufficient for many years, a dynamic range of (at best) 1000:1 cannot realistically represent or display the luminance of objects existing in the real world. These differ greatly between daytime and nighttime scenes, and can even differ considerably within a single scene. For example, the luminance of an object in an indoor environment depends on the geometry and the amount of light present, whether as artificial light or entering from outdoors through a small window.
[0005] The scene's luminance dynamic range (RA_SCN) can range from 8000 nits for a light bulb 104, 150 nits for a well-lit indoor object 101, and 0.1 nits for an object in shadow (e.g., a robot vacuum cleaner 102). Objects outdoors in the sun, such as the wall of a house 103, may reach, for example, 3000 nits (on average).
[0006] The dynamic range of a variable is calculated by dividing the brightest value that the variable possesses or can possess by its smallest value. To faithfully represent all the tonal gradations of a scene, the dynamic range of that representation must be greater than or equal to the scene's dynamic range (i.e., 80,000:1).
[0007] Dynamic range is often given in terms of the number of stops (a preferred expression for camera operators), which is the number of times the minimum value (Qm) is doubled to reach the maximum value QM (in other words, the log_2 representation). In linear representations, such as the output of a camera's analog-to-digital converter, this also corresponds to the number of bits. However, generally, the number of bits used should not be confused with dynamic range, as bit codes can be assigned to relative or absolute luminance via a highly nonlinear electron-optical transfer function (in fact, as demonstrated in this field, even 8-bit image coding can represent high dynamic range and SDR images).
[0008] Dynamic range is often given in decibels in the field of electronics (e.g., sensor technology), and in linear terms, it is given by the equation dB = 20 * log_10(QM / Qm).
[0009] Furthermore, the human eye has evolved to "see" a much higher dynamic range than what older cameras could capture or what older LDR displays could display.
[0010] This results in some unavoidable artifacts; for example, outdoor scenes illuminated by sunlight are clipped to maximum white, meaning they display a completely white rectangle with nothing to see. Our eyes sometimes have difficulty adapting to HDR scenes, such as indoor scenes from an outdoor perspective, but we know that everyone can see things outside windows and doors perfectly normally. Thus, there is an unrealistic aspect to the image. While one could argue that an image is always just an interpretation of reality, this dynamic range limitation is not always a desirable aspect.
[0011] As a result, LDR video creators had to be extremely careful about how they created their video content to ensure that it would look very good on any LDR TV, even if the dynamic range was small due to imperfect viewing conditions. Filming real-world scenes could yield imperfect results. For example, at the 2023 King's Day parade in Rotterdam, dancers performed on a gray and white checkerboard floor. However, even the dark gray squares were brightly lit by the sun, so in LDR video, it was impossible to distinguish them from the white squares. A few minutes later, the director thought it would be a good shot to film the musicians from behind, with the sunlit audience in the background, but everything about the musicians in the shadows under the tent was almost completely lost in black.
[0012] In the era of high dynamic range (HDR) video (video with a higher dynamic range than LDR), all image processing chain scenarios (limited past, present, and future) can be represented using a simplified dynamic range remapping series as shown in Figure 1.
[0013] First, the brightness of the original object in the real world is irrelevant. This is true not only when the video creator is free to determine the brightness of objects in the video image according to their preference, but also when the mapping between the camera's sensor capture and the output image is direct with a simple technical mapping function. One reason is that the camera operator technically controls the exposure of the sensor pixels; that is, they measure not the amount of light emitted by the real-world object, but how much of that light actually enters the sensor pixels. This can be done by means of apertures, neutral density filters, etc. The amount of light that is actually converted into a measurable (digitalizable) voltage depends particularly on the quantum efficiency of the sensor.
[0014] A perfect sensor that has all the necessary features, including price, has not yet been developed or produced for any single application (you might wonder if future mobile phones will have sensors like those in professional cameras).
[0015] For example, R. Ikeno et al., "A 4.6um, 127-dB Dynamic Range, Ultra-low power stacked digital pixel sensor with overlapped triple quantization" (IEEE Tr. On Electron Devices, Vol. 69, No. 6, June 2022), describes a sensor with a dynamic range that captures 2,000,000:1 (127 dB). However, the sensor has relatively large pixels and a spatial resolution of only 512 × 512.
[0016] Nevertheless, even inexpensive mobile phones can soon be expected to be equipped with sensors capable of faithfully capturing the colors outside doors and windows in bright sunlight.
[0017] This raises questions about how to display HDR images on low dynamic range displays. The brightness and luminance of these outer pixels need to be mapped down (aka downgraded). A display with a maximum end-user display brightness (ML_D) of 100 nits cannot render an outdoor scene that appears much brighter than indoor pixels. However, one can choose to downmap them so that they display something better than a pure white pixel (e.g., a desaturated, whitish color). But if all these pixels receive an input signal with the same values (e.g., maximum brightness or lumacode 255), it can never estimate what was outside the door.
[0018] In fact, displays began to improve in the first two decades of the 21st century. CRTs had a technical limitation: the front light had to be made with an electron beam. While it was easy to make a strong electron beam, the electrons repelled each other, making the beam thicker and reducing spatial resolution. Liquid crystal displays initially had a single fixed backlight behind the liquid crystal material, such as a set of serpentine or linear TL tubes. This allowed the backlight's light output to be brighter, but the darkest black (i.e., the amount of light LCD pixel leakage when driven by the darkest code, zero) would be brighter by the same amount, meaning the dynamic range would not increase. In patterned, also known as 2D dimming backlights, a dedicated LED was placed behind each spatial group of pixels, brightening the light passing only through those pixels where they needed to show bright objects on the screen, and dimming the LEDs where they needed to show dark areas.
[0019] This has resulted in current television sets that can easily display pixels with the brightest brightness of 1000 nits (and some even exceeding, for example, 2000 nits) and the darkest pixels with, for example, 0.01 nits. Professional displays can reach up to 10k nits.
[0020] Therefore, in principle, it is now possible to capture (or generate with CG) natural images with various HDR effects (e.g., bright and colorful explosions, sunlight illuminating glossy objects, etc.) and display them in a way that is impressive to the end user. However, in the LDR era, there was one standard codec, Rec709, which was developed to code (match) for typical best LDR display scenarios, i.e., to code a (relative) luminance range of only 1000:1.
[0021] Therefore, in the 2010s, it was necessary to develop one or more HDR video coding standards, otherwise the created images could not be displayed where they were ultimately needed. Also, with HDR, engineers had to reconsider each technical aspect known from SDR video processing, while keeping in mind the need to maintain compatibility with existing technologies where possible.
[0022] Figure 1 briefly illustrates what the camera operator wants to achieve. The camera's ADC is typically positioned to match the sensor pixel design (e.g., LOFIC, dual conversion gain, etc.), but not necessarily to match the scene. What it has in common with LDR capture is to obtain the best possible basic (RAW) capture of the scene.
[0023] In LDR systems, this is extremely difficult. Therefore, the camera operator typically looks for where facial highlights begin to become too bright (above the ideal level of approximately 80%). On the darker side, they check whether at least the important parts of the dark scene are clearly visible, and add auxiliary lighting if not. All other objects will have to be clipped to white or black (or buried in noise level NOI). During production, especially in broadcast production that airs in real time, the scope is constantly monitored to ensure the signal remains (virtually) within the required levels. In real-world scenes, i.e., without controlled lighting, the results may not be optimal.
[0024] Modern (HDR) cameras generally have a 14-stop (16,000:1) camera ADC range (CAM_ADC) and a sufficient digital number MAXFW, which is (2;14). Therefore, even in this exemplary scene, there is generally more freedom in exposure tolerance than with an LDR camera, although some care should still be taken with exposure. Typically, you don't need to capture every detail (small brightness changes) of a light bulb, so you can expose it so that the pixels and ADC overflow and the maximum ADC digital number (DN) is output. In this case, the dimly lit vacuum cleaner still has enough detail to be brightened later with brightness regrading. Other objects will get a linear percentage (i.e., relative) digital number depending on this setting, that is, you get what brightness in the scene (e.g., 7000 nits, or actually the corresponding pixel exposure F_EXPO) just starts clipping.
[0025] From here, the approaches to representing (i.e., coding) HDR images diverge. One reason for this divergence is whether you want to code directly, for example, directly record from the camera to memory, or format it for some distribution method. However, there are other reasons, and in reality, there are many more types of HDR codecs than strictly necessary.
[0026] These can be classified into two categories: absolute or target display resolution, and relative or brightness resolution.
[0027] To ensure that the future use of video images is as professional and timeless as possible, regardless of whether it's for future use (e.g., optimizing for a specific display called display adaptation, or for the near future for transcoding, communication, regrading, etc.) or for the distant future such as long-term storage, we want to define the situation relatively accurately, but allow for a lot of variability that various technical applications may require.
[0028] Referring to FIG. 1, it can be seen by the reader that if the relative position of values can be adjusted completely freely at a later stage, the precision of CAM_ADC hardly has any meaning (this may include a complete inverse relationship with the original values, but it is less likely to invert the relative lightness such that in the obtained image, a first object that is originally darker than a second object becomes brighter than the second object). Nevertheless, it is found that many displays do not have the dynamic range present in received images, so it is necessary to take some measures.
[0029] In the first decade of the 21st century, all input videos are SDR, and therefore displays with a wider dynamic range calculate pseudo-HDR images (for example, identifying light objects in an image such as a set of pixels forming the solar surface, then arbitrarily increasing their lightness relative to other pixels of the image). For now, it seems that there is a tendency to reduce or keep low the dynamic range at various points in the image processing chain. After a few years, with more insight into the HDR aspect that leads to stabilization, some produce lower DR videos, while others produce higher DR videos (that is, videos with a higher maximum brightness ML_C), and it is expected that displays exist along a continuum of various maximum end-user display brightness (ML_D). Today, downgrading displays can use their own algorithms to downgrade the brightness or lightness of input pixels, but content creators still do not know what the final image will look like. Therefore, in order to achieve more consistency across variable scales, several approaches describe guidance on how end displays downgrade various images (there are other re-gradings depending on the scene present in the image, for example a dark cave, or an explosion in a daytime scene, and thus this intended re-grading behavior, which is applicable to each subsequent image of the video, can be specified by the video creator by putting information into metadata that is, for example, an SEI message).
[0030] To enable more control by the creator over how the creator's video will ultimately appear to end viewers, first a target display having a color gamut with a target display luminance dynamic range (TARG_DIS) is defined. In this case, the maximum value associated with the video is the maximum value of the target display luminance dynamic range, that is, ML_C. This allows a content creator to specify the intended luminance in a clear manner. For example, because a 3000 nit house would appear too bright when viewed in the evening, it may be specified with a luminance of 500 nits, which is still brighter than indoor objects. This process can be compared to selecting a canvas aspect ratio to optimize the spatial aspects of a painting: one may select 1:1 for a still life, but for a landscape one may first select that a 3:1 aspect ratio is better than 1:1. After setting this initial technical condition, one can optimize the settings for exactly which tree and which mountain to draw along the canvas. The same applies if it is selected that an ML_C of 5000 nits provides a good representation for a current film: an outer house is optimized to have pixel brightness spread around 500 nits, and a light bulb will likely be at the maximum value of 5000 nits or a lower 3500 nits, and so on. Having a target display associated with the image means that one can also create a dark scene to later harmonize with, for example, an image having a 5000 nit sun.
[0031] In the present patent specification, the resulting image, regardless of whether the mapping is performed by a human or by some automatic algorithm, is generally referred to as HDR master grading (mast_grad).
[0032] To specify how the creator sees the necessary regrading of the primary video of a master-grade image, the creator can specify a reference secondary grading. Typically, this is not communicated as an actual secondary image, but rather one or more functions that map the various luminances that an object can have in the range of 0 to 5000 to the corresponding luminances in a reference LDR range (RA_REF_LDR), usually ranging from 0 to 100 nits. Then any actual display can base itself on the communicated pixelated image (typically, e.g., a 10-bit / channel YCbCr image and a selected EOTF communicated in the metadata), and the reference regrading function calculates a display-optimized image (where the pixel luminances of the object are optimally rearranged along the dynamic range (RA_USR_DIS) of the user display).
[0033] Various codec versions transmit images of different grades to the receiver. For example, one codec can directly transmit HDR master grading, while another transmits a corresponding regraded reference LDR image. Yet another codec transmits some intermediate image with an intermediate dynamic range RA_INT_FRMT and ending at an intermediate maximum luminance ML_IM (e.g., 600 nits).
[0034] Finally, for the sake of simplicity, everything was explained using universally understandable luminance ranges and mappings (for now, only for absolute nit codecs (nit, also known as cd / m2, is the physical unit of luminance)). However, in reality, images are not coded, but rather transmitted directly as luminance values as (one of) their pixel values. Instead, some lumacode Y represents luminance, and an EOTF can be selected that links all lumacodes to their corresponding luminances (e.g., perceptual quantizer PQ). Thus, in practice, virtually any luminance range corresponds to a range of luma (e.g., coded luma range in intermediate format COD_INT_FRMT). In the example shown, a typical 10-bit coding was chosen, i.e., the luma range is 0 to 1023. The maximum luminance in this video corresponds to, for example, luma=820 (typically, ML_IM is transmitted natively in the metadata as the number of nits, and the corresponding lumacode can be evaluated at the receiving end by simultaneously specifying the EOTF). While LDR objects fall within the SR_LDR subrange of low scene brightness, which requires robust coverage by an LDR camera, they can be more freely placed within the wider HDR range of sensor capture by an HDR camera.
[0035] Furthermore, while various remappings have been conceptually ordered one after the other in a typical processing chain, in reality, what happens where also depends on the production (some examples of which are shown in Figure 2). For example, one production might choose to start almost entirely with SDR capture and create an HDR reference video from all contributions at some point in the future (e.g., just before broadcasting to end consumers), while another workflow might create a single HDR master but want derived SDR versions for the majority of viewers, which is done semi-automatically (i.e., an algorithm with several adjustable parameters that affect the manner of the automaton's mapping) and in parallel with master HDR production, etc. Other systems might involve conversions between relay stations, for example, a local cable station mixing an upgraded SDR commercial feed (national commercials may be produced in HDR, but local commercials are still SDR but need to be broadcast and displayed in HDR) into its program. In particular, in systems that select a fixed mapping function (i.e., the same function for all consecutive images in the program, rather than a function that can optimize the shape for each image), artistic grading choices regarding the visibility of blacks, for example, are made first in the SDR domain (i.e., viewed on a reference monitor), but the ripple effect spreads as corresponding choices for HDR masters or both HDR and SDR grading become more distinct.
[0036] There is a second class of HDR coding that operates differently in series. Hybrid log-gamma coding, developed by the BBC and standardized in ARIB STD-B67, is not all very well defined and does not work with absolute luminance values. In their philosophy, it simply creates relative brightness. To create a larger assignable range for video objects, this relative range goes up to 1000%, rather than ending at 100% (white) like SDR. However, until considering conversions of absolute codec material, such as from Blu-ray® movies, the BBC had not specified, or assumed, that the HLG range (RA_HLG) relative maximum value of 1000% was what exactly luminance was. We can agree that the 1000% level should correspond to 1000 nits. Furthermore, on the display side, no display typically displays 1000% as an arbitrary specific brightness, but it was intended to display it as its maximum capability (thus 1500 nits for a 1500 nit ML_D display, 550 nits for a 550 nit ML_D display). This creates images that look quite different, but one can justify this by making the effort to purchase a more expensive 1500 nit HDR display. However, in order to make dark colors appear with sufficient contrast (otherwise the eye will adjust much of what appears to be a higher dynamic range), the display needs to display the HLG signal with gamma correction that increases as the ML_D value increases.
[0037] Therefore, the dynamic range at the far right of Figure 1 (RA_REF_LDR) is the dynamic range of the SDR image. All other ranges are from the HDR image. SDR images were well known to video professionals (how to create, code, use, and display SDR images, etc.). In terms of brightness, an SDR image ends at Lambertian reflected white (i.e., white that evenly divides incident light in all directions; for clarity, one can think of a blank sheet of paper) under a uniform amount of illumination (the brightest color that an SDR image represents). If we want to relate luminance to this 100% level, this is the typical display luminance of an LDR display: 100 nits. HDR images are defined in comparison to SDR images in that they have a larger dynamic range (in brightness or luminance) than SDR images. There may also be deeper blacks, but in the video technology field, we always want (at least) brighter pixels. Therefore, if we define the black level for typical living room viewing (darkest black) as 0.1 nits, a wider range corresponds to higher percentage luminance (e.g., 1000% is 10 times Lambertian white 100%) or greater maximum luminance (e.g., ML_C = 800 nits). This HDR image definition allows for the creation of very bright image objects, such as white objects illuminated by lights greater than average scene lighting, specularly reflective objects, or self-illuminating objects.
[0038] In HDR technology, the brightness of the HDR signal for at least some image objects is significantly increased, not just in terms of how precisely the maximum brightness should be measured in nits (i.e., remember that a 2% difference is imperceptible to the human viewer, not even to five decimal places). For example, each 100% is at least twice as high as 100 nits. In the real world, specular reflection is thousands of times brighter than diffuse reflection, but in image representation, a 10-fold increase in brightness is sufficient (the goal is not to cause spots in the viewer's eyes due to bleaching of cone opsin molecules as in the real world, and technical limitations such as reasonable power consumption should be taken into consideration). When speaking of distribution containers for television broadcasts or video cable signals (e.g., serial digital interfaces), there may be some modifications to these basic principles, but their details are discussed only to the extent relevant to any given embodiment.
[0039] Not only have the types of displays (LCD TVs, mobile phones, home cinema projectors, professional cinema digital projectors), video sources, and communication media (satellite, streaming via the internet, such as OTT streaming via 5G) increased, but the means of video production have also expanded.
[0040] Figure 2 shows several typical video creation methods that can effectively implement this instruction without being restrictive in general.
[0041] In studio environments, for example, for news or comedy, there are still often tightly controlled shooting environments (although HDR mitigates this, allowing for shooting in real-world settings). There is controlled lighting (202), such as battery-powered base lights on the ceiling and various spotlights. There are numerous large, relatively stationary television cameras (201). Variations of this often real-time broadcast include, for example, sports programs like soccer, which have various types of cameras such as goal-scoring cameras for localized viewpoints, wide-angle cameras, and drones.
[0042] There are several production environments 203 where various feeds from the cameras can be selected for the final feed, and various (usually simple, but potentially more complex) grading decisions can be made. In the past, this was often done on a production track with many displays and various operators, for example, but in an internet-based workflow where raw feeds are transmitted over several networks, the final assembly takes place on the broadcast station's premises. Finally, to simplify production for the sake of this explanation, coding and formatting for broadcast distribution to the end customer (or intermediate customer such as a local cable station) is done in the formatter 204. This typically involves, for example, converting graded luminance to PQ YCbCr for intermediate dynamic range formats, as illustrated in Figure 1, calculating and formatting all necessary metadata, converting to broadcast formats such as DVB or ATSC, and packetizing into chunks, etc., for distribution (the "etc." indicates that tables to inform of available content, subtitles, encryption, etc., may be attached, but at least some of these will not be of much interest to understanding the details of this innovation).
[0043] In this example, video (television broadcast in this example) is transmitted via television satellite 250 to a satellite receiving antenna and a satellite signal-enabled set-top box 261. It is then displayed on an end-user display 263.
[0044] This display shows this first video, but can also display other video feeds, for example, in a picture-in-picture window, potentially even simultaneously (or some data from the first HDR video program may come through one distribution mechanism, and other data through another).
[0045] The second type of production is typically offline production. This could be a Hollywood movie or a show where someone runs through a jungle. Such productions are shot with other optimal cameras, such as Steadicam 211 and drone 210. Again, we assume that the camera feed (which may be a raw feed or a feed converted to an HDR production format such as HLG) is stored somewhere on the network 212 for later processing. In such productions, there may be several human color graders using grading equipment 213 to determine the optimal luminance (or relative luminance in the case of HLG production and coding) for master grading in the last few months of production. The video is then uploaded to an internet-based video service 251. In professional video distribution, this would be, for example, Netflix®.
[0046] A third example is consumer video production. Here, a user, for example, when creating a vlog, has a ring writer 221 and captures via a mobile phone 220, sometimes capturing in an outdoor location without auxiliary lighting. The user also typically uploads to the internet, but nowadays they may also upload to YouTube® or TikTok.
[0047] When receiving data via the internet, display 263 connects via router 262 (complex settings such as internal Wi-Fi® are not shown in this simplified explanation).
[0048] Therefore, today we can see that various types of videos on various technical coding are generated and transmitted through various means.
[0049] If HDR had been developed purely by computer engineers, it would still have been reasonably simple (at least in the aspects mentioned above). In the computer world, Lumacode is simply seen as a numerical value, that is, not necessarily optimized for any transmission standard, but simply used to code corresponding lightness (relative) or luminance (absolute). Three numerical values, for example, 8-bit or 10-bit, sufficiently specify the colors of a three-primary additive color system, such as nonlinear R'G'B' or YCbCr which can be calculated from R'G'B' by applying a fixed matrix, of which nine constant coefficients depend on the primary colors (e.g., EBU primary colors describing a standard set of CRT phosphors). The luminance channel, coded as LumaY, behaves like one of the nonlinear red, green, and blue components, to which the color difference, also known as chrominance coordinates Cb and Cr, are added.
[0050] In computer representation (or simply when representing the amount of silver measured from a scan of a wet photographic image), 0 codes for the darkest black, and 255 codes for the brightest possible white, respectively.
[0051] This is not the case with the broadcasting technology explained in Figure 3.
[0052] The concept of narrow-range (also known as legal range LEG_RA) coding has become somewhat confusing or misleading in the age of HDR video coding. It is one of two coding modes permitted for broadcast video, along with so-called full-range FU_RA (which does not actually use full range but is extended compared to narrow-range).
[0053] This concept stems from historical reasons when analog (NTSC, PAL, or SECAM) voltage video signals were digitized. While the white voltage was assumed to be 700mV, overflow in analog circuits such as filters could cause the voltage to extend "slightly" above 700mV. On the black side, the main mechanism for underflow would likely be noise.
[0054] In the digital age, one might consider whether we should simply clip all values above 700mV (with visual relevance also provided) to the digital maximum value (e.g., 1023 for a 10-bit codec).
[0055] In the broadcast world, a mix of digital values intended to contain special codes and values containing coded values representing pixel brightness was chosen. In the analog world, lower values indicate, for example, a synchronization pulse. In the full range, values less than 4 and greater than 1019 should not be used for video brightness coding. These are reserved for timing criteria (and should not be misunderstood). In the legal range, larger code sets are not permitted, particularly for small word lengths such as 8 bits.
[0056] This is a somewhat ambiguous concept. To be legal, it is required that the valid, i.e., the associated pixel brightness code, should not be outside the legal range. However, if absolutely nothing exists there, if nothing can happen outside the range, then the full range can also be used (with the exception of some reserved codes). Those grading with the full range on a computer should exercise caution. They can either keep all values within the range of 64-940, or freely grade "1.0" as some maximum floating-point number (which may overflow in some cases), and then "scale" the values to fit within the narrow range. Even if you end up with a signal that was generated as full range (for some reason), you can scale 1019 to the upper limit of the narrow range, NR_UL, which is 10 bits 940 (even if this introduces further quantization errors).
[0057] The idea was that these guard bands (headroom and footroom) should not be completely empty, but should only contain elements of secondary importance.
[0058] For example, if there is a small amount of ringing due to the filter, the receiver can choose to reduce the ringing by using both values close to and slightly above the upper limit within the legal range. The same applies to noise. Clipping the noise value to below a lower luma (NR_LL) in the narrow range, i.e., below 64, results in noise with different statistics, i.e., it is no longer Gaussian, like the original, for example, thermal noise. Using an averaging filter yields a somewhat different noise-reduced value. Whether that is truly important for viewing under viewing conditions where black cannot be seen well in any case can be debated, but the signal was defined that way nonetheless.
[0059] In the HDR era, the importance of asymmetrically optimized headroom and footroom is now even less significant, as the impact of visibility in these areas depends on luminance mapping.
[0060] Some have reinterpreted headroom and footroom as concepts that can be used to convey useful image information. That is, the range of brightness that can be coded can be extended by several additional luminance codes representing several additional luminances above and below the range of Rec.709, i.e., LDR, resulting in additional highlights and deeper blacks. Extending a relatively low curvature EOTF like Rec.709, which is roughly square root in shape, doesn't bring in much additional luminance or brightness to make it truly meaningful. However, modern EOTFs like PQ and HLG have a highly nonlinear (at least partially logarithmic) shape, which means that several additional codes can significantly extend the codeable dynamic range (for example, a real logarithm can be coded with, say, 255 codes, and 10 years per 50 codes, i.e., 200 to 250 codes, can be extended by more than 10 times, i.e., 100 to 1000, i.e., the last 50 codes get an additional 90% range). However, it's possible to use a trick that transmits something that doesn't even transmit a full-range normal LDR range or a full-range HDR range, but rather an extended narrow range. Thus, the system recognizes that 940 is a "normally bright" white, and the higher the code, the brighter something is encoded.
[0061] This can cause problems in receiving systems that don't expect anything useful to happen above 940 in the received narrow-range signal, and these receiving devices can be assumed to simply ignore everything in illegal headroom and footroom (e.g., if they don't perform ringing mitigation, or if they already do so in a different way). Therefore, if such a system performs luminance regrading or relative luminance regrading, it can be assumed that 940 is the maximum luminance, i.e., 100%. If some luminance is associated with its value, by definition, a coded pixel with 1000 nits is assumed to have a narrow-range code of 940. But also, we assume there is nothing else. Many luminance regradings, especially those that downgrade to lower dynamic ranges, map the highest luminance of the input image (or the luma, if the processing is done directly in the luma region) to the luma of each of the highest luminances in the output region defined for the LDR output image, e.g., Rec.709. And nothing would be expected above that (below the respective narrow-range lower limit). Therefore, if such a system exists, it cannot handle these luminance codes, nor can it handle them even by mistake. Those skilled in the art of colorimetric analysis should be well aware of methods of mapping luminance using software algorithms or corresponding ASIC processing circuits, for example, where images are supplied from some memory on a pixel-by-pixel basis. However, those skilled in the art, if interested, can find examples of the types of processing they wish to perform using HDR signals in the single-layer HDR standard of the present invention (ETSI TS 103 433-2 V1.1.1 January 2018).
[0062] Figure 3 shows a simpler of the possible narrow container definitions to give a brief overview of this approach. Instead of having only one narrower range within a larger range, one can also have a multi-range, for example, a dual narrow-range system. In this system, there is a narrowest range, beyond that (usually, but not exclusively, at both ends) some further codes (e.g., used in a manner that is used in some applications but not important in others), and beyond that a second duo of headroom and footroom lumens. This duo is intended to be used only, for example, for some overshoot (e.g., to reduce compression artifacts due to bitrate reduction). The maximum ("full") range can end with the maximum and minimum codes of bitwords (e.g., for internal representation in processing ICs or grading software) or it can have some reserved codes (e.g., for SDI communication).
[0063] Formally speaking, the narrow range is smaller than the full range, which extends from zero to the maximum code value CM, and this depends on the number of bits (aka word length) used to code the luma (even if some codes are reserved in the end). The equation is CM = power(2, N) - 1, where N is the number of bits, for example, values N = 8, 10, and 12 are frequently used. If there is an additional narrowed range, such as by having a secondary narrow range upper limit NR_UL2 between NR_UL and CM (for example, in the middle), then the following embodiments can generally be used by substituting the NR_UL2 value for the NR_UL value, and similarly for the dark edge. A container representation can be defined where only one of the two edges is extended, or further subdivided with an additional range limit (but not both).
[0064] Figures 4A and 4B illustrate an HLG to LDR conversion system (although the presented technology is also useful for other codecs such as narrow-range coded absolute luminance coding). The technology is useful for this system, and the system is easily explained to those skilled in the art.
[0065] HLG was initially presented as an automatically backward-compatible system. Its coded image is already a directly displayable LDR image. In other words, the image looks good, even though it's not a lumens representation of an actual HDR image, but merely "pretending" to be a regular LDR image. This offers a free and significant advantage: new HDR displays can decode into HDR images using the new HDR technological insights, while already deployed televisions simply display the image as usual (i.e., its YCbCr pixel code). Despite the fact that some assumptions and limitations must be made regarding the generation of the lumens signal (i.e., it cannot be generated as freely as the absolute HDR codec mentioned above), the system's simple rigidity has several drawbacks that prevent it from being a desirable backward-compatible system due to excessive color errors. Therefore, HLG becomes the only means of coding HDR images (with little backward compatibility). For example, with cameras that have a wide capture dynamic range, direct coding from the camera has become common (given the technical specifications of the cameras, how the camera generates an HLG image from a RAW capture may be a variable for the camera manufacturer, but in any case, this approach was considered sufficient).
[0066] One thing the HLG conversion does when driving a legacy LDR display using the HLG conversion directly (for example, from a 1000nit ML_C HDR master video) is fixed tone mapping (as end-to-end automatic behavior). However, that doesn't mean it's the tone mapping you want.
[0067] Figure 4A shows what happens when we encode HDR video of ML_C at 1000 nits as 10-bit lumens using the HLG photoelectron transfer function (assuming full range for now). The horizontal (input) axis is normalized luminance, defined as L_n = L / 1000, so 1.0 represents 1000 nits (when we first encode). The normalized lumens are shown on the vertical axis, i.e., 1.0 represents 1023 in 10-bit representation. On the same plot, we have added what happens when a “dumb” legacy display determines these are normal SDR lumens. This then decodes according to the reciprocal of the Rec 709 OETF, which is essentially near the 45-degree diagonal of the plot. This has the advantage that the resulting mapping (which is done somewhat automatically, i.e., not formally as an actual tone mapping processing step) is tone mapping, which works relatively well in that context, but is not perfect for all kinds of video content. For example, sports program producers might not like it. For example, during a soccer match, the ball may be in a sunny spot in the stadium (remember that ideally it should show a reasonably dark area), while half the spectators may be sitting in the shade. Therefore, the producer may find that, at least on that particular day, the spectators being filmed are either a little too dark or, conversely, too bright. This can be resolved by performing tone mapping (e.g., static, i.e., fixed) according to the producer's preference. An illustrative example is shown in Figure 4B.
[0068] Following the upward, rightward, and downward arrows in Figure 4A, we can see an increase in relative brightness of the decoded and ultimately displayed normalized luminance from 0.1 to 0.3. This is indeed what is expected from tone mapping to a lower dynamic range for darker colors within the range. However, it is important to understand that the input range is normalized by 1000, while the final output range of the displayed SDR image is shown superimposed on the horizontal axis. These normalized luminances need to be multiplied by 100 to obtain the final (absolute) luminance that is displayed. Thus, in reality, in absolute luminance representation, we are mapping from a coding-side input of 100 nits (i.e., the selection in the master HDR grading of the object in the image to which those pixels belong) to just 30 nits. Again, "generally" this is the behavior we want to see for downmapping, because we must make space in a more limited SDR range for ultra-high luminance pixels in the HDR image (a soccer ball under bright lighting, or LEDs on a commercial board in a stadium). However, there is nothing particularly correct about the value 30. Depending on what the original object is, or the original common subrange of brightness around 100 nits, the output value may be considered sufficiently reasonable, or it may not be good and may require further adjustment.
[0069] Typically (though not always), the selection of the HLG-to-SDR mapping curve (F_grad) has a shape similar to that described in Figure 4B (though usually more rounded and not composed of segments with sharp angles). From this, darker blacks than some blacks are considered too dark (due to division by 3) and it may be desirable to brighten them. Slightly brighter colors may be made even brighter, as can be seen from the higher slope of the third line segment (recall that the axis is in units of lumens and is therefore somewhat logarithmic, with the horizontal axis showing full-range HLG lumens and the output axis showing 10-bit SDR, i.e., lumens as defined in Rec709 (PRO_SDR)). The last (fourth) segment usually has a lower slope because it needs to cut into the ultra-high luminance colors of the master HDR image (some sub-optimization may be necessary, and these colors are usually of lower importance, such as stadium lighting or highlights rather than, for example, the faces of players).
[0070] We chose to continue the second segment through the lower limit of the narrow range, which has the property of darkening dark colors (for example, 64 for 10 bits and 16 for 8 bits, although these values may differ in both cases). Further segments (from the first) may be used to map the deepest black (a black that is theoretically irrelevant or doesn't even exist if we strictly adhere to the legal range, but which may still be present in the input signal and for which some processing is desired (rather than simply clipping or ignoring it)).
[0071] Therefore, when a processing element, such as dynamic range processing in a receiver, receives or requests a legal range signal as input, it is typical to map the legal range limit to the limit of its processing range (PROC_RA) (by remapping TCHRAMA). As shown on the right side of Figure 3, for example, when performing normalized luma mapping, the upper limit of the narrow range (NR_UL) is remapped to correspond to 1.0 and the lower limit of the narrow range (NR_LL) is remapped to 0 (values in between can be distributed linearly). Problems can arise if the input signal contains overflow or underflow values outside the narrow range. The mapping is 1023 on the input axis of the luma mapping for unnormalized 10-bit numbers and 255 for 8-bit luma. Out-of-range situations can occur, for example, a camera may output values outside the legal limit, and a broadcaster may decide to perform follow-up processing for these out-of-range values. While final conformity to the legal range occurs immediately before broadcasting the video signal onto the broadcast medium, processing prior to this may be performed using non-compliant legal ranges, i.e., outliers in headroom and footroom. [Overview of the project]
[0072] If values outside the narrow range cannot be processed (for example, because they fall outside the processing range, e.g., the normalized range 0-1.0), the solution is a luma mapping device (500). The device (500) luma maps a video image containing pixels (IM_IN), where the input luma of a pixel (Y_leg) is specified by a reduced range (RR) compared to the full range (FR), where the full range extends from zero to the code maximum (CM), where the code maximum is 2 to the power of N minus 1, where N is the number of bits representing the luma, and the reduced range is delimited by a narrow range lower limit (NR_LL) greater than zero and a narrow range upper limit (NR_UL) less than the code maximum, where some luma of some pixels have values outside the reduced range. This device includes a range compressor (102) that maps the luma to a renormalized luma (Y_norm), which includes mapping the narrow range lower limit (NR_LL) to the processing range lower limit (PR_LL), mapping the narrow range upper limit (NR_UL) to the processing range upper limit (PR_UL), and mapping all values in between by linear scaling. This device includes a luma mapping circuit (103) that maps a renormalized luma (Y_norm) to a remapped luma (Y_RM) by applying a fixed or configurable luma mapping function to the renormalized luma. The device comprises a premapper (101) that maps luma values below a fixed or configurable inflection point (IFP) to values above the inflection point by a mirroring operation, and outputs a sign bit having a first value for the mirrored pixel luma (this indicates a "negative original," i.e., below the IFP, which is coded as 1, for example, as an actual bit value, and all non-mirrored luma correspond to the other bit value 0, and vice versa), and a postmapper (104) that re-mirrors pixels where the sign bit is "negative" around the inflection point (i.e., the first value is set, which occurs only for mirrored luma, and non-mirrored luma obtain other values).
[0073] The primary requirement is that it be interpreted in accordance with the pre-established technical purpose of narrow range in a lumen container (e.g., broadcast television). Therefore, this means that the relevant gray and color values should be specified within the narrow range, and only incidentally, should there be some (insignificant) out-of-range values, such as ringing overflows from electronic components or processing, or some noise added to the system that causes at least some pixel values to be below the lower limit of the narrow range—i.e., below 64 in 10-bit coding, or slightly above the upper limit of the narrow range. However, this is usually not intentionally specified by the video creator. More specifically, the distinction of where all specified colors should be within the narrow range is interpreted in that the video creator should create a narrow-range container specification for lumens (or general color components) expecting a typical display to clip everything above NR_UL before display (i.e., display everything brighter as maximum display white) and everything below NR_LL to minimum display black. The display may have performed some processing on out-of-range values, such as noise reduction, before renormalizing the display. Therefore, the majority of pixel values are specified to fall within that narrow range.
[0074] By performing mapping, even if the values are within the processing range, a reasonable processing can be applied, and for example, good noise reduction can be guaranteed.
[0075] The inflection point value is typically set to the lower limit of the narrow range, i.e., 64.
[0076] It is advantageous for the luma mapping device to further include a control interface for setting the luma value of an inflection point, for example, 64. This allows the system to pre-determine the type of narrow-range container coding that will occur as input in an ecosystem where a single value is not applied. For example, for higher-bit luma words, this value may be set higher (e.g., twice as high for each additional luma coding bit), or for multiple sub-range containers, the device may set the optimal value, for example, by automatically identifying metadata or as a user setting.
[0077] An advanced embodiment includes a premapper (101). The premapper (101) treats lumas lower than an inflection point (IFP) as at least two segments, a first segment (SEG_1) containing lumas higher than a first lower luma (Yls1), and a second segment (SEG_2) containing lumas lower than the first lower luma (Yls1). The premapper (101) maps the lumas of the second segment onto a projected segmentation point (P_Yls1), which is a mirroring around the inflection point of the first lower luma, in a compressed manner by scaling the distance between the mapping point (pmap) of any point in the second segment and the projected segmentation point in a ratio corresponding to the angles of the first segment and the second segment.
[0078] This is useful for dealing with intended luma mappings where the shape of the luma mapping function below the inflection point is nonlinear or can be approximated by a single linear segment.
[0079] Alternatively, when performing the same mirroring on all lumas below the inflection point (i.e., regardless of the shape of the luma mapping function applied), the postmapper (103) identifies a projected segmentation point (P_Yls1) which is a mirroring around the inflection point of the first lower luma (Yls1), which is the lowest bound luma of the first segment of lumas below the inflection point, and maps lumas above the projected segmentation point to positions below the inflection point by an expression that includes the ratio of two distances, where the numerator is the distance between the projected segmentation point (P_Yls1) and the point where the luma-mapped value of the mapping of any input lumas around the inflection point is located as the remapped luma (Y_RM), and the denominator is the distance between the projected segmentation point (P_Yls1) and the reference value (cvm) of the remapped luma value.
[0080] Various embodiments of the technology can also be implemented as a method for luma mapping a video image containing pixels (IM_IN). The input luma of a pixel (Y_leg) is primarily specified by a narrowed range (RR) compared to the full range (FR), where the full range extends from zero to the code maximum (CM), and the code maximum is 2 to the power of N minus 1, where N is the number of bits representing the luma. The narrowed range is delimited by a narrow range lower limit (NR_LL) greater than zero and a narrow range upper limit (NR_UL) less than the code maximum, where some lumas of some pixels have values outside the narrowed range. This method performs range compression, which maps the luma to a renormalized luma (Y_norm), and this mapping includes mapping the narrow range lower bound (NR_LL) to the processing range lower bound (PR_LL), mapping the narrow range upper bound (NR_UL) to the processing range upper bound (PR_UL), and mapping all values in between by linear scaling. This method then performs luma mapping, which maps the renormalized luma (Y_norm) to the remapped luma (Y_RM) by applying a fixed or configurable luma mapping function to the renormalized luma. This method performs premapping before range compression, where luma values below a fixed or configurable inflection point (IFP) are mapped to values above the inflection point by a mirroring operation, and a "negative" sign bit is output for the mirrored pixel luma. Post-mapping is performed, and in post-mapping, the remapped luma of pixels whose sign bit is "negative" around the inflection point (for example, setting s=0 for pixel (0,0)) is remirrored back to below the inflection point. [Brief explanation of the drawing]
[0081] These and other aspects of the methods and apparatus according to the present invention are made apparent from and described with reference to the implementations and embodiments described below, as well as the accompanying drawings. These serve merely as non-limiting specific examples illustrating a more general concept, and dashed lines are used to indicate that components are optional, with components not using dashed lines not necessarily being absolutely essential. Although dashed lines are described as absolutely essential, they can also be used to indicate elements hidden inside an object, or for intangible things, such as the selection of an object / area.
[0082] [Figure 1] Figure 1 schematically illustrates the various possibilities for creating HDR images (and SDR images) by showing the luminance mapping of image objects along various possible HDR luminance ranges (i.e., dynamic ranges). [Figure 2] Figure 2 schematically illustrates several typical current video creation, transmission, and usage scenarios in which the following technologies can be used. [Figure 3]Figure 3 illustrates how some video coding technologies (subfields / markets) employ so-called narrow-range coding when embedding Lumacode into a container (e.g., broadcast or other transmitted content), and highlights some of the problems associated with mapping to a secondary range, such as a processing range for image processing, including Luma mapping for dynamic range conversion. [Figure 4A] Figure 4A illustrates the problems encountered when using HLG-encoded HDR video as an SDR signal for a legacy LDR display. [Figure 4B] Figure 4B shows how to improve a less-than-ideal direct display by using image enhancement lumina mapping. [Figure 5] Figure 5 schematically shows a possible hardware implementation of the device based on the principles of this teaching. [Figure 6] Figure 6 illustrates the processing principles of several techniques for performing mirror mapping of pre-mappers and post-mappers using this innovation. [Figure 7] Figure 7 shows how to perform mirroring on (individual) RGB components. [Modes for carrying out the invention]
[0083] Figure 5 shows the appearance of the apparatus for carrying out this embodiment. The luma mapper processes the input pixel lumas one by one from the scan of the input image IM_IN (for example, inputting the first luma Y1 of the first pixel image as the current Y_leg). For the basic concepts of range renormalization (for example, renormalization to a processing range that typically extends from the lower processing range limit PR_LL to the upper processing range limit PR_UL (for illustrative purposes, we use values of 0 and 1, though not limited)) and subsequent luma mapping (only lumas that fall within the narrow range used by the range compressor (102)), the new circuit unit based on this technical insight is shown as a thick rectangle.
[0084] Note that when processing luma (for example, normalizing it to 1.0 luma) (by Lumamappa 103), it does not matter whether the luma represents relative brightness (i.e., a percentage of the maximum value that is not numerically specified) or actual luminance value (i.e., the luminance displayed on the display). However, the display has a displayable luminance range that is greater than the luminance range specified (encoded) for the video.
[0085] It is advantageous for a luma mapping function to be constrained to use the same narrow-range lower bound for both the input (e.g., HLG) luma range and the output (e.g., Rec.709 SDR) luma range. That is, the function maps, for example, 64 to 64, regardless of its shape around this range.
[0086] When the mapping is nonlinear, approximation with a linear function or a set of connected segments is desirable. This linearization is only important in the region around the inflection point, i.e., below 64, and in the region where the darkest black luma mirrors, i.e., below 128, for example. For the 1023 luma function, it is not a major constraint if the shape of the function passing through the inflection point up to luma 128 is linear, and above 128, it still has any desired shape (e.g., the color of a light bulb is strongly compressed while maintaining contrast in a sunny house outside a door opening). In some situations, a double-segment approach may be desired for the darkest colors below the inflection point IFP (many segments are not necessary in many applications where the invention is useful).
[0087] In other words, the projection of colors above the inflection point, which is used by premapper 101 to position colors within the range so that they can be later lumen-mapped, can be formulated as follows:
[0088] For any luma Yx (below IFP), if it has a value of Yx = 64 - D (for example, on the horizontal axis), it is mapped to 64 + D. It can also be shown that this corresponds to the formula Y_bfli = 128 - Yx, where Y_bfli is the inflected luma if the original luma Y_leg was below the luma of the inflection point, and Y_leg otherwise. The premapper also outputs a sign bit for each mirrored luma. A simple way to actually do this (although other possibilities exist, such as storing the bits with the pixel position coding or a list of all signs s(0,0), s(0,1), ... where each "+1" is "-1" or each 0 is 1) is to send each bit of the pixel being processed directly to the postprocessor (possibly via a delay element that takes the same time as the processing delay in 102 and 103). This depends on whether the device is actually manufactured as, for example, a full ASIC color processing pipeline circuit, or whether only element 103, or 102 and 103, are present in the processing circuit, with the other units being management software (firmware) around the circuit. The (re)normalized luma (Y_norm) output from the range compressor becomes the remapped luma (Y_RM) after luma mapping. Note that a simple luma mapper is shown to illustrate the new technical concept, but actual mapping is more complex. For example, mapping can be performed with a different luma representation. For example, first convert the luma Y_norm to a perceptually homogenized luma (Y_perc), perform luma mapping in the perceptually homogenized luma region (this will result in luma mapping functions of different shapes, but ultimately yield the same result; therefore, those skilled in the art can calculate the shape of the perceptually homogenized luma region function when knowing the luma mapping function that will ultimately be applied), and then return to the output luma region as desired in the example Rec.709 luma described.
[0089] Finally, after re-mirroring by a post-processor, they become output LUMA Y_out (these are ready to be formatted, for example, as broadcast signals backward compatible with legacy LDR displays, or sent to a panel driver, for example, as LDR images on a display and / or further processed). The inflection point value (e.g., 64) is set by an external circuit and loaded into the various units that require it.
[0090] For a single segment, the postmapper 103 re-mirrors the point values after luma mapping using the same formula (but in the reverse direction, downwards toward darker areas than the IFP). In fact, the range compressor (102) also consists of several range mappings to intermediate ranges (e.g., full range) before mapping to the processing range, for example (although shown as a single mapping in the instruction).
[0091] Dual-segment Luma mapping allows you to choose to perform mirroring adjustments (as scaled / compressed mirroring) either in the post-processor or pre-processor, while standard mirroring (128-Yx) is applied in the other.
[0092] This is shown in Figure 6, where the principle of pre-mapping is shown on the left and the principle of post-mapping is shown on the right.
[0093] As explained in the graph on the left (which shows the luma mapping function F-Grad in detail around the inflection point), the luma in the first segment (SEG_1) below IFP is mirrored as usual (i.e., there is no slope adjustment compression). In other words, other values at the inflection point are mirrored using equation 128-Yx, etc. Note that in some videos, there may be no values in the second segment below Yls1, but the curve behavior may be defined, and in any case it is loaded into the luma mapping device (in which case there is no need to know if dark values exist, and anything that comes in is simply luma-mapped).
[0094] When processed with the intended luma mapping function (i.e., the guide function in scenarios where we simply map lumas below the IFP), we find that luma Yx lies only at a distance b*di, not a*di, the vertical distance below the lower endpoint of the second segment. Knowing that the mirroring of the lower endpoint of the first segment ends at P_Yls1, we can map it to a compressed mirrored position (pmap) that is precisely compressed by the ratio of the angles (or inclinations) of the two segments, i.e., b / a. Performing a normal mirroring back around the IFP, the pmap point moves to the correct position (pfin), which is correctly generated for all pixels of the second segment (SEG_2). (di is the distance from the lowest point of the first segment to any point, as the horizontal luma value distance to understand the principle).
[0095] On the right, we see that the premapper can initially map all points in the same way (by normal, non-compressor-up-mirroring above the IFP), but in this case, the postmapper needs to properly re-mirrore, taking into account the correct scaling of the second segment.
[0096] One implementation code for doing this is as follows: % Pre-mapping Converts negative values (<64) to positive values while preserving the sign bit. %input:hlgcv[] contains a 10-bit extended narrow-range HLG code value; %cvm contains the highest HLG code value available. That is, %cvm=min(128, "Highest R=G=B HLG code value, relative to R=G=B gamma code value ≤ 128") if(hlgcv[i]<64) sb[i]=1;%negative else sb[i]=0;%positive end if(hlgcv[i]>31)&hlgcv[i]<64)%code value 32..63 hlgcv[i] = 128 - hlgcv[i]; Mirrored map to value 65..96 end if(hlgcv[i]≦31)%code value 0..31 hlgcv[i]=97+(31 hlgcv[i])*(cvm97) / 31;% value 97..Map mirrored to cvm end %output:hlgcv[] contains only "positive" HLG code values ≥ 64. %sb[] stores the sign bit that originally indicated a value of "negative" < 64.
[0097] The cvm value is, for example, 113 (depending on the luma mapping being performed, such as the amount of emphasis on the darkest luma).
[0098] Figure 7 shows a preferred embodiment. Mirroring does not necessarily have to be performed on the luminous components of the pixel color. It is advantageous to do this for the three RGB components (which, according to some EOTFs, are typically nonlinear components R'G'B'). Typically, the pixel color comes in as Y_in, Cb_in, Cr_in, so the first matrix circuit 701 converts them to RGB, yielding inputs R_inp, G_inp, B_inp. The coefficients of this matrix are fixed for a selected set of RGB primary colors, and those skilled in video technology are well aware of how to do this or vice versa. Next, each of these three components is inverted separately by a similarly functioning premapper 702. Typically, there are no differential inflection points for these three components, but rather the three mappings use the same values. However, these three components can have different sign bits. For example, under color gamut overflow, it can happen that one of the components is negative (or more precisely, below the inflection point) while the other two are fine (automatically within the processing range). Therefore, for one pixel (or more pixels), a set of three sign bits, sR(0,0), sG(0,0), and sB(0,0), is transmitted to the re-mirroring postmapper 704.
[0099] Here too, the dynamic range mapping process performed on the luma is shown (without intending to limit it). Therefore, the red, green, and blue components (R_NL, G_NL, B_NL) within the mirrored processing range obtained by the narrow-range mapping (R_fli, G_fli, B_fli) are inversely converted by the matrix circuit 703 to become YCbCr again. However, in other embodiments, luminance processing can be performed directly on the RGB components, preferably keeping the ratios R / G, R / B, and G / B constant. That is, R_NL / G_NL = R_RM / G_RM, etc. Here, R_RM and G_RM are the remapped red and blue color components, similar to Y_RM as described above. The postmapper 704 is also generally slightly different because the three are potentially different sign bits (some +1 and some -1 for the pixel). The postmapper 704 is shown with its internal subcircuit (internal rectangle). Typically, the matrix is first converted from (Y_RM,Cb,Cr) to RGB, and the RGB components that require it are mirrored depending on the sign bit. For example, if sR(0,0) = "-1", the red component of the first pixel is mirrored, but if sG(0,0) = "1", the green component is not mirrored. In this example, a direct RGB output for driving a display has also been described, but in other embodiments there may be further color representation conversions, which do not need to be detailed in this description of the technique.
[0100] The algorithmic components disclosed herein are actually implemented (whole or in part) as hardware (e.g., as part of an application-specific integrated circuit) or as software running on a specialized digital signal processor or a general-purpose processor. Images may be stored temporarily or permanently in various memories located near the processor or accessible remotely, for example, via the Internet.
[0101] It should be evident from this presentation to those skilled in the art which components are optional improvements, which can be realized in combination with other components, and how the (optional) steps of the method correspond to each means of the apparatus (and vice versa). Some combinations are taught by dividing the overall teaching into partial teachings relating to one or more components. The word “apparatus” in this application is used in its broadest sense, that is, a group of means for achieving a particular purpose, and thus could be, for example, an IC (a small circuit portion thereof), a dedicated device (such as an electrical appliance with a display), or part of a networked system. “Arrangement” is also intended to be used in its broadest sense, and in particular could include a single apparatus, a part of an apparatus, a collection of cooperating apparatuses (or parts thereof), and so on.
[0102] The explicit meaning of a computer program product should be understood as encompassing any physical realization of a set of commands that, after a series of loading steps (which may include intermediate translation steps such as translation into an intermediate language or a final processor language), enable a general-purpose or special-purpose processor to input commands into the processor and execute any of the characteristic functions of the invention. In particular, a computer program product can be realized as data on a carrier such as disk or tape, data residing in memory, or data traveling over a wired or wireless network connection. Apart from the program code, characteristic data required for the program may also be realized as a computer program product. Some technologies can be encompassed in signals, typically control signals for controlling one or more technical behaviors of a receiving device such as a television. Some circuits are reconfigurable and are temporarily configured by software for specific processing.
[0103] Some of the steps necessary for operating a method, such as data input and output steps, may already exist in the processor's functionality rather than being described in the computer program product.
[0104] It should be noted that the embodiments described above are illustrative, not limiting, of the present invention. Those skilled in the art will readily be able to map the presented embodiments to other areas of the claims, but for the sake of brevity, not all of these options are described in detail. Apart from the combinations of elements of the present invention combined in the claims, other combinations of elements are possible. In practice, any combination of elements can be realized with a single dedicated element or divided elements.
[0105] Any reference numerals in parentheses in a claim are not intended to limit the claim. The word “includes” does not exclude the existence of elements or aspects not enumerated in the claim. A singular element does not exclude the existence of multiple elements or other elements.
Claims
1. A luma mapping device for luma mapping a video image (IM_IN) containing pixels, wherein the input luma (Y_leg) of the pixels is mainly specified by a reduced range (RR) compared to the full range (FR), the full range extends from zero to the maximum code value (CM), the maximum code value is 2 to the power of N minus 1, where N is the number of bits representing the luma, the reduced range is delimited by a narrow range lower limit (NR_LL) greater than zero and a narrow range upper limit (NR_UL) less than the maximum code value, and some of the input luma of some of the pixels have values outside the reduced range. The luma mapping device includes a range compressor that maps the luma to a renormalized luma (Y_norm), which includes mapping the narrow range lower limit (NR_LL) to the processing range lower limit (PR_LL), mapping the narrow range upper limit (NR_UL) to the processing range upper limit (PR_UL), and mapping all values in between by linear scaling. The luma mapping device includes a luma mapping circuit that maps the renormalized luma (Y_norm) to a remapped luma (Y_RM) by applying a fixed or configurable luma mapping function to the renormalized luma, The luma mapping device is characterized by including a premapper that maps luma values below a fixed or configurable inflection point to values above the inflection point by mirroring operation and outputs a sign bit having a first value for the mirrored pixel luma, and a postmapper that remirrarizes pixels around the inflection point such that the sign bit has the first value.
2. The luma mapping apparatus according to claim 1, further comprising a control interface for setting the luma value of the inflection point, for example, 64.
3. Luma mapping apparatus according to claim 1, wherein the premapper treats the luma lower than the inflection point (IFP) as at least two segments, a first segment (SEG_1) containing luma higher than a first lower luma (Yls1), and a second segment (SEG_2) containing luma lower than the first lower luma (Yls1), and the premapper maps the luma of the second segment onto the projected segmentation point (P_Yls1), which is a mirroring of the first lower luma around the inflection point, in a compressed manner by scaling the distance between the mapping point (pmap) of any point in the second segment and the projected segmentation point in a ratio corresponding to the angle of the first segment and the angle of the second segment.
4. Luma mapping apparatus according to claim 1, wherein the postmapper identifies a projected segmentation point (P_Yls1) which is a mirroring of a first lower luma (Yls1) around the inflection point, which is the lowest luma of a first segment of lumas below the inflection point, and the postmapper maps lumas above the projected segmentation point to a position below the inflection point by a formula including the ratio of two distances, the numerator being the distance between the projected segmentation point (P_Yls1) and the point where the luma-mapped value of the mapping of any input lumas around the inflection point is located as a remapped luma (Y_RM), and the denominator being the distance between the projected segmentation point (P_Yls1) and a reference value (cvm) of the remapped luma value.
5. A method for lumen mapping a video image containing pixels (IM_IN), wherein the input lumens (Y_leg) of the pixels are primarily specified by a reduced range (RR) compared to a full range (FR), the full range extending from zero to the code maximum (CM), the code maximum being 2 to the power of N minus 1, where N is the number of bits representing the lumens, the reduced range being delimited by a narrow range lower limit (NR_LL) greater than zero and a narrow range upper limit (NR_UL) less than the code maximum, and some of the lumens of some of the pixels among the input lumens have values outside the reduced range. The method described above involves range compression, which maps the luma to a renormalized luma (Y_norm), and this mapping includes mapping the narrow range lower limit (NR_LL) to the processing range lower limit (PR_LL), mapping the narrow range upper limit (NR_UL) to the processing range upper limit (PR_UL), and mapping all values in between by linear scaling. The method further performs luma mapping, and in the luma mapping, the renormalized luma (Y_norm) is mapped to the remapped luma (Y_RM) by applying a fixed or configurable luma mapping function to the renormalized luma, The method involves performing premapping before range compression, in which luma values below a fixed or configurable inflection point (IFP) are mapped to values above the inflection point by a mirroring operation, and a sign bit having a first value is output for the mirrored pixel luma. The method is characterized by performing post-mapping, wherein the post-mapping remirrates the remapped luma of the pixels where the sign bit has a first value around the inflection point, and returns it to below the inflection point.