Image display improvement in a brightened viewing environment

By using high dynamic range image display technology and monitor tuning, the problem of image display in non-dark environments has been solved, achieving image brightening and color accuracy in high-illuminance environments, and is suitable for various display devices.

CN122374813APending Publication Date: 2026-07-10KONINKLIJKE PHILIPS NV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2024-11-29
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively display high dynamic range images in non-dark viewing environments, especially in high-light conditions, where image details and colors are severely distorted, failing to meet the demands for high brightness and high dynamic range.

Method used

Employing high dynamic range image display technology, through color grading and metadata processing, combined with display tuning technology, it achieves image brightening and adaptation, ensuring that image details and color accuracy are maintained even in non-dark environments.

Benefits of technology

It achieves effective display of high dynamic range images in high-illuminance environments, maintaining image detail and color accuracy, and is suitable for various display devices without the need for additional color processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122374813A_ABST
    Figure CN122374813A_ABST
Patent Text Reader

Abstract

To better perform the adaptation of an image to be displayed under varying ambient lighting, it is proposed to use a method of controlled adjustment using a control processor (510) or a corresponding, arranged to adjust the luminance of at least a subset of dark colors in an input image (IMG_in) for display on a display (520) in response to a measurement of the illuminance (Lev_illumdrown) of the environment in which the display (520) is located, wherein the adjustment comprises establishing an initial adjustment function (F_adj) for shifting one or more color components in an input color component triplet (R'_PQ, G'_PQ, B'_PQ) of a pixel of the input image by at least one correction value (Adj_brightn) to obtain a corresponding output color component triplet (R'_out, G'_out, B'_out), wherein the adjustment comprises multiplying the initial adjustment function (F_adj) by a correction function (F_corr) for values of one or more color components in the input color component triplet (R'_PQ, G'_PQ, B'_PQ), which are below a threshold value (GAM1), to obtain a final version of the adjustment function for shifting the input color component triplet (R'_PQ, G'_PQ, B'_P Q), wherein the correction function (F_corr) satisfies: a first property that produces zero output for zero input; a second property that produces an output value equal to one for an input value equal to the threshold value (GAM1); and a third property having a first derivative equal to zero at the input value equal to the threshold value (GAM1).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the display of, in particular, high dynamic range images or videos (image sequences) that require a certain degree of brightening to adjust the display for non-dark viewing environments with illumination higher than the expected typical illuminance, and the images are prepared for non-dark viewing environments through color grading. Background Technology

[0002] For over half a century, an image representation / encoding technique now known as Low Dynamic Range (LDR) or Standard Dynamic Range (SDR) has worked flawlessly in creating, transmitting, and displaying electronic images (such as video, i.e., a temporally continuous sequence of images). In colorimetry (i.e., in terms of the specifications of pixel colors), this technique is based on methods already applied to photographic materials and paintings decades ago: it only requires defining and displaying most colors projected from the achromatic (also known as grayscale) axis, which ranges from the lowest black level to the brightest achromatic color, giving the viewer the impression of white. For television communications, which rely on display-side additive color creation mechanisms, a triplet of red, green, and blue color components needs to be transmitted for each location (pixel) on the display screen. This is because, using appropriate proportions of triplets (e.g., 60%, 30%, 25%), almost any color can be produced, and practically all the desired colors can be produced (the white level is obtained by driving the three display channels to their maximum values, where the drive signals Rmax = Gmax = Bmax).

[0003] The earliest television standards (NTSC, PAL) transmitted color components as three voltage signals (defined as the amount of color components between 0 and 700mV), where the time position on the voltage signal was mapped using a scan path with pixels on the screen.

[0004] The control signals generated on the creation side directly instruct the display to produce the proportions that should be reproduced on the creation side (except for occasional fixed gamma pre-correction at the transmitter, due to the physical characteristics of the cathode ray tube, which takes approximately the square power of the input voltage, making dark colors much darker than expected (e.g., as seen by a camera)). Therefore, 60%, 30%, and 25% of the color in a captured scene (which is deep red) will appear substantially similar on the display because it will be regenerated as 60%, 30%, and 25% of the color (it should be noted that absolute brightness is not very important because the viewer's eye adapts to the white level and the average brightness of the colors displayed on the screen). This can be called "direct-link drive" without additional color processing (except for arbitrary and unnecessary processing that display manufacturers may still perform, such as making the sky appear bluer). For backward compatibility with older black-and-white television broadcasts, red, green, and blue voltage signals are not actually transmitted; instead, a luminance signal and two color difference signals called chrominance (in current terminology, blue chrominance Cb and red chrominance Cr). The relationship between RGB and YCbCr is simple and can be understood using a simple fixed... The matrices (whose coefficients depend on the emission spectra of the three primary colors and have been normalized, i.e., also internally electronically simulated by LCDs that may actually have different optical properties, so that all SDR displays are similar from the perspective of image communication) calculate them together.

[0005] These voltage signals are later digitized according to the Rec. 709 standard for use in digital television (based on MPEG, etc.), and various quantities are defined, such as a luminance component with 8-bit codewords, 0 encoding for the darkest color (i.e., black level), and 255 encoding for white level. It should be noted that compression of the encoded video signal is not always necessary (although it usually is). When compression is performed, it will be (e.g., via MPEG-HEVC, AV1, etc.) worded compression, and should not be confused with the act of compressing colors to a smaller color gamut or range. Luminance refers to the portion of the color definition that will affect the darker or brighter visual characteristics of the (to be) displayed color. Given this technology, it is important to correctly understand that two types of luminance can exist: relative luminance (as a percentage of something, which may be undefined until, for example, a consumer purchasing a monitor makes a choice and adjusts its brightness setting to, for example, 120%, which will cause, for example, the backlight to emit a certain amount of light, and therefore the white level and color pixels to also emit a certain amount of light); and, on the other hand, absolute luminance. Absolute luminance can be characterized by the general physical quantity “luminance” (which is technically measured in nits, which is also a candela per square meter). Luminance can be expressed as the amount of photons emitted from a block on an object (such as pixels on a screen) toward the eye (and it is related to the lighting concept of illuminance, since such a block will receive a specific illuminance and direct a portion of it toward the viewer).

[0006] Two unrelated technologies have recently emerged, appearing together because people believe it's possible to comprehensively improve the visual quality of images all at once, but these two technologies have quite different technical aspects.

[0007] On the one hand, there has been a persistent pursuit of wide color gamut imaging technology. It can be demonstrated that only colors located within the triangle formed by the three primary colors of red, green, and blue can be reproduced; colors outside this triangle cannot. However, people have chosen primary colors (EBU, also re-normalized in Rec. 709) that are relatively close to the spectral locus of all existing colors (initially phosphors in CRTs, later color filters in LCDs, etc.), thus enabling the production of sufficiently saturated colors for most purposes. Saturation indicates how close a color is to an achromatic (colorless) color, i.e., how "colorful" a color is. However, recently there has been a desire to use new displays with more saturated primary colors (e.g., DCI_P3 or Rec. 2020), necessitating the ability to represent colors in such a wider color space. A color space is a mathematical 3D space representing colors (defined by coordinate values, e.g., the weighted combination of red, green, and blue values ​​of the primary color intensities in the generated total colors), typically presented in a shape where the base is defined by a triangle of the three primary colors. A color model refers to the choice of numerical categories used to define a space. For example, red, green, and blue are natural representations used to specify an additive color generator. However, the same color can be defined by a range of three other coordinates in a hue, saturation, and lightness model, which characterizes color in a more human-relevant way. For technical discussion, it is best to use the term "color gamut," which is the set of all colors that can be technically defined or displayed (i.e., a space can be, for example, a 3D coordinate system extending to infinity, while a color gamut can be a cube or tent shape of a certain size within that space). Regarding lightness alone, we will discuss the range of lightness or luminance (more commonly referred to as "dynamic range," which spans from a minimum lightness to its maximum lightness). This technique is not primarily concerned with chromaticity (i.e., the color itself, such as its saturation, for example), but rather with lightness; therefore, chromaticity will only be mentioned to the extent necessary for the relevant embodiments.

[0008] The more important new technology is High Dynamic Range (HDR). This should not be interpreted as "exactly this high" (because many variations of HDR representation can exist with successively higher maximum values), but rather as "higher than the reference / traditional representation: SDR". Due to the need for new coding concepts, HDR can also be distinguished from SDR by aspects of the technical details of the representation (e.g., the video signal). One difference between absolute HDR systems and SDR systems is that they define a unique luminance for each pixel in an image (e.g., a white dog in sunlight might have 550 nits in pixels), whereas SDR signals only have a relative luminance definition (so on a computer monitor that can display a maximum of 80 nits, the dog might look, for example, 75 nits (corresponding to 94%), but on a 250-nit SDR monitor, the dog would appear as 234 nits; however, viewers typically wouldn't see any difference in image appearance unless the monitors were placed side-by-side). Readers should not confuse the luminance that (ultimately) exists on the front of any display with the luminance defined (i.e., established) on the image signal itself, even when that image signal is stored but not displayed. Another difference is that any type of HDR signal can have metadata that an SDR signal does not.

[0009] In colorimetry, HDR images can represent brighter colors than SDR images, and therefore, colors brighter than white (e.g., luminous white). In other words, the dynamic range will be greater. An SDR signal can represent a dynamic range of 1000:1 (the actual visible dynamic range when displayed will depend particularly on the amount of ambient light reflected from the front of the display). Therefore, if a dynamic range of, for example, 10000:1 is desired, then a new HDR image format definition must be used (if the image representation is being or will be transmitted rather than simply existing within the IC, we can generally use the term "signal," and the signaling will typically also have its own format and encapsulation, and may employ additional techniques depending on the communication mechanism (e.g., modulation).

[0010] Depending on the situation (even if everything on the front screen corresponds to a small glare angle), the human eye can easily perceive a dynamic range of 100,000:1 (e.g., a maximum of 10,000 nits and a minimum of 0.1 nits, which is a good black level for home television viewing (i.e., in a dimly lit room that only supports relatively low levels of lighting, such as at night)). However, not all images created by creators need to reach such a high level: creators can choose to set the brightest image pixels in an image or video to, for example, 1,000 nits.

[0011] The SDR white level luminance of a video (also known as SDR white point luminance (WP) or maximum luminance (ML)) is normalized to 100 nits (not to be confused with the reference luminance of 200 nits for white text in a 1000-nit HDR image). That is, an HDR image with 1000 nit ML can represent objects that are up to 10x brighter (glowing) in color. This can be used to create, for example, specular reflections on metal, such as the edges of a metal window frame: in SDR, the luminance must end at 100 nits, making them visually slightly brighter than, for example, a 70-nit light gray portion of the window frame without specular reflections. In HDR, pixels that reflect light to the eye can be set to, for example, 900 nits, making them glow well, thus giving the image a natural look, as if it were a real scene. The same applies to fireballs, light bulbs, etc. This involves the definition of an image; however, how a display with a white level that can only show a maximum brightness of 650 nits (the maximum luminance of a display, ML_D) should actually display an image is a completely different issue—a problem of display adaptation (also known as display tuning), not image encoding (decoding). Furthermore, the relationship between how a camera captures the colors of an HDR scene can be close or loose: when discussing techniques such as encoding, communication, and dynamic range conversion, we often assume that HDR colors are already defined in the HDR image. In reality, the colors captured by the original camera, or specifically their luminance, may have been altered to different values ​​by, for example, human color graders (who define the final appearance of an image, i.e., which color triplets each pixel in one or more images should have) or some automatic algorithm. Therefore, given the fact that "the color gamut size may be stretched by less than 2 times, while the luminance range (e.g., luminance range) may be stretched by 100 times," it is expected that these two improvement techniques have different technical principles and solutions.

[0012] HDR images can be associated with metadata known as Master Production Display White Point Luminance (MDWPL) (also called ML_V). This value is typically conveyed in the signal's metadata and is a characteristic parameter of the HDR video image (not a characteristic parameter of a specific display, as it is an element of a virtual display specifically associated with the video, where the video pixel colors have been optimized to conform to some ideal intended display). This is a selectable parameter for the video, which can be imagined as similar to a painter choosing the aspect ratio of a canvas: first, the painter chooses an appropriate aspect ratio, for example, 4:1 when painting a landscape, or 1:1 when wanting to create a still life, and then he begins to optimally position all his objects within the chosen canvas. In an HDR image, the creator, after having already determined the MDWPL to be, for example, 5000 nits, then makes his secondary choices, such as the lampshade should be 700 nits in a particular scene, the flames in the fireplace should be distributed around 500 nits, etc.

[0013] The main visual aspect of HDR images lies in the additional brightness (due to the greater complexity of black levels) (therefore, if we assume the bottom luminance is fixed at, for example, 0.1 nit, then the range can be defined using only the MDWPL value). However, HDR image creation can also involve even lower black levels, as low as, for example, 0.0001 nit (although this is primarily relevant to dark viewing environments, such as in a movie theater).

[0014] Other objects (such as those that only reflect scene light) will be reconciled to be at least 40x darker in a 5000nit MDWPL-rated video, and at least 20x darker in a 2000nit video, etc. Therefore, the distribution of luminance across all image pixels will generally depend on the MDWPL value (rather than making most pixels very bright).

[0015] The digital encoding of luminance involves a technical quantity called luma (Y). We will use the letter L to represent luminance and the letter B to represent (relative) luminance. It is important to note that, technically, for example, for ease of defining some operations, luminance in the WPDPL range of even 5000 nits can always be normalized to a normalized range [0, 1], but this does not change the fact that these normalized luminances still represent absolute luminance in the range up to 5000 nits (in contrast to relative luminance, which never has any explicitly associated absolute luminance value and can only be temporarily converted to luminance that usually has some arbitrary value).

[0016] For SDR signals, luma coding uses a so-called photoelectric transfer function (OETF) that converts between optical brightness and typically 8-bit electronic luma code, as defined below: .

[0017] If B_relative is a floating-point number in the range of 0 to 1.0, then Y_float will also be.

[0018] Subsequently, the signal value Y_float is quantized. Because an 8-bit digital representation is desired, the Y_dig value transmitted to the receiver via, for example, terrestrial DVB (or video-on-demand or Blu-ray discs provided by the internet, etc.) has a range of 0 to 255 (i.e., The values ​​between ).

[0019] like Figure 1B As shown, the color gamut of all SDR colors can be represented (or similarly, the color gamut of HDR colors can be represented if the same RGB primary colors used to define the chromaticity gamut have the same base but are vertically stretched to a larger absolute gamut solid after renormalization). Larger Cb and Cr values ​​will result in (more psychologically visually relevant color characterization parameters) larger saturation (sat), which shifts from the unsaturated or colorless vertical axis shown in the middle towards the maximum saturated color on the circle (a transformation of the usual color triangle formed by the RGB primary colors as vertices). Hue h (i.e., color category, yellow, green, blue) will be along the angle of the circle. The vertical axis representation (e.g., via a psychovisually homogenized OETF or its inverse function EOTF) normalizes luminance (in linear color gamut representation) or (e.g., via a psychovisually homogenized OETF or its inverse function EOTF) normalizes luma (in non-linear representations (encoding) of these luminances). Since after normalization (i.e., dividing by the corresponding MDWPL value, e.g., dividing by 2000 for an HDR image of a particular video and by 100 for an SDR image), a common representation becomes readily available and can be expressed as follows: Figure 1DThe diagram shows a luminance (or luma) mapping function defined on the normalization axis. (That is, the function F_comp redistributes the pixel values ​​of various image objects as needed, so that in the actual luminance representation, for example, dark objects look the same (i.e. have the same luminance), but have different normalized luminances. This is because in one case, the pixel normalized luminance will be multiplied by 100, while in another case, the pixel normalized luminance will be multiplied by 2000. Therefore, in order to obtain the same final luminance, the normalized luminance of the later pixel should be 1 / 20 of the normalized luminance of the earlier pixel.) When downgrading to a smaller luminance range, you will typically get a convex function (in a representation normalized to 1.0, i.e., mapping the range of normalized input luminance L_in between zero and one to normalized output luminance L_out), which lies above the diagonal diag everywhere. However, the exact shape of the function F_comp (e.g., how fast it must rise at the black level) usually depends not only on the two MDWPL values, but also (for the most perfect version of the regrading technique) on the scene content of various (video or still) images, such as whether the scene is a dark cave and whether there is action taking place in the shadow areas that must be reasonably visible even in the 100 nit luminance range (therefore, for such scenes, a strong boost to the black level is required compared to daylight scenes that can take a near-linear function that almost overlaps the diagonal).

[0020] exist Figure 1B In the representation, the mapping of pixels from HDR color (C_H) to LDR color (C_L) can be shown as a vertical shift (assuming that the two colors should have the same appropriate hue (i.e., hue and saturation), which is generally the desired technical requirement (i.e., on a circular base, they will project to the same point)). Ye represents yellow, and its complementary color on the opposite side of the non-color axis of luminance (or luma) is blue (B), and W represents white (the brightest color in the color gamut, also known as the white point of the color gamut; the darkest color in the color gamut, with the black level at the bottom).

[0021] Therefore, the receiving side (past or present) will know that it has SDR video when it receives this format. By definition, the maximum white (in SDR) will be the brightest color that SDR can define. Therefore, if you now want to create brighter image colors (e.g., the colors of a real light source), you should use a different codec (because it can be shown that the mathematical operations of Rec. 709OETF only allow encoding up to 1000:1 and cannot encode a higher range).

[0022] Therefore, a new framework with different code assignment functions (EOTF or OETF) was defined. This paper is primarily interested in the definition of Luma codes.

[0023] For reasons beyond the scope of this discussion, most HDR codecs define the electro-optical transfer function (OETF) first, rather than its inverse function. Then, at least essentially, brighter (and darker) colors can be defined. This is insufficient for professional HDR coding systems, as it differs from SDR and even has various variations, so more (newer than SDR coding) technical information related to HDR images is desired—this information would be metadata.

[0024] The characteristic of those HDR EOTFs is that they are steeper to encode a wider range of HDR luminance, and a significant portion of this range is specifically for darker colors (relatively darker because, although absolute luminance can be encoded using, for example, a perceptual quantizer (PQ) EOTF (normalized in SMPTE2084), this function is applied after normalization). In fact, if a precise power function is used as the EOTF to encode HDR luminance into HDR luma, the power will be 4 or even 7. When a receiver receives a video image signal defined by such an EOTF (e.g., a perceptual quantizer), it will know that it has received HDR video. It will need the EOTF to decode pixel luma across the luma plane of the image (i.e., with, for example, a width of 4000 pixels and a height of 2000 pixels), which will simply be binary numbers. Typically, HDR images will also have a larger word length, such as 10 bits. However, non-linear encoding, which can be arbitrarily designed by optimizing the shape of the non-linear EOTF, should not be confused with linear encoding and the number of bits required for linear encoding. If linear (bit-represented) codes are needed to drive, for example, DMD pixels to achieve, say, a modulation ratio of 10000:1 from darkest to brightest, then log2 is needed to obtain the number of bits. There, at least 14 bits will be required (which may be rounded up to 16 bits for technical reasons), because... However, given the ability to cleverly design the shape of the EOTF and the knowledge that visual systems do not perceive all luminance differences equally, the applicant has demonstrated (surprisingly) that a fairly reasonable HDR television signal can be transmitted using only 8 bits per pixel color component (of course, 10 bits would likely be better and more preferred if technically feasible in the system). Therefore, in both cases, the receiving side can receive encoded pixel colors (luma and Cb, Cr; or in some systems, nonlinear R'G'B' component values ​​equivalent to matrix transformations) located between 0 and 255 or 0 and 1023 as input, but it will learn from metadata what type of signal it received (and thus what should be displayed), such as the EOTF (e.g., value 16 indicates PQ; value 18 indicates that another OETF was used to create the luma, i.e., a hybrid log-gamma OETF, so the inverse function of that function should be used to decode the luma plane), the MDWPL value (e.g., 2000 nit) in many HDR codes, and additional metadata in more advanced HDR codes (some metadata may, for example, co-encode a luminance (or luma) mapping function to be applied to map image luminance from the primary luminance dynamic range to the secondary luminance dynamic range, e.g., a function FL_enc per image).

[0025] We illustrate in detail the typical requirements of the increasingly complex HDR image processing chain using a simple illustrative Figure 1 (for a typical good HDR scene image (a scene where a monster is repelled by a flamethrower in a cave), its master grade (Mstr_HDR) is... Figure 1A The image is shown in spatial form, and the range of pixel luminance is within... Figure 1C (As shown on the left). Master grading or master-grading images are where an image creator can make their image look as impressive (e.g., realistic) as possible. For example, in a Christmas movie, he could make the bakery window look slightly brighter by making the yellow walls slightly brighter than paper white (e.g., 150 nits) (and presenting a true, vibrant yellow rather than a pale yellow), and he could set the light bulbs to 900 nits (which would give the image a truly bright Christmas look, rather than a dull look where all the lights are cropped to white and not much brighter than the rest of the image (e.g., the green of the Christmas tree)).

[0026] Therefore, the basic thing that must be able to do is to encode (and usually also decode and display) image objects that are brighter than in a typical SDR image.

[0027] Now, from a colorimetric perspective, SDR (and its encoding and signaling) is designed to transmit any Lambertian reflection color under good, uniform lighting (in a scene where camera capture occurs)—that is, a typical object, such as your blue jeans, absorbs some incident light (e.g., red and green wavelengths) to emit only blue light towards the viewer or the capturing camera. It's like what we do in painting: without paint, light reflects back to full brightness from a white canvas, while if we add a thick layer of highly absorbent paint, we will see black strokes or dots. We can represent all colors brighter than the darkest black and darker than white in what's called the representable color gamut, such as... Figure 1B (As shown in the “tent”). As the base, we have a circle representing all the representable chromaticities (note that it could be discussed at length that this should be a triangle in a typical RGB system, but such details are beyond the scope of this teaching). A chromaticity consists of a specific (rotation angle) hue h (e.g., turquoise, for example, “teal”) and a saturation sat, which is the amount of pure color mixed in gray, e.g., the distance from the central vertical axis representing all non-chromatic colors gradually brightening from black at the bottom to white. A chromatic color (e.g., a semi-saturated violet) can also have a brightness, meaning the same color can be slightly darker or slightly brighter. However, the brightest color in an additive color system can only be (achromatic) white, because it is produced by setting all color channels to the maximum R=G=B=255, so there is no imbalance that would make the color noticeably reddish (there is still a slight blue or yellow tint in the chosen white point chromaticity, but this is unnecessary for further discussion; we will assume D65 is daylight white). We can define these SDR colors by setting MDWPL to (relatively) 100% for white (it should be noted that in traditional SDR, white does not actually have luminance because traditional SDR does not have luminance associated with the image, but we can assume it to be X nit, for example, usually 100 nit, which is a good average representative value for various traditional SDR TVs).

[0028] Now we want to represent colors that are brighter than Lambert colors, for example, the color of the self-illuminating flame object (flm) of a soldier (sol) fighting a monster (mon) in a dark cave.

[0029] Suppose we define a master HDR grading of 5000 nits (maximum video luminance ML_V). (The master refers to the starting image (in this case, the most important and highest quality grading). We will first perform optimal grading on this image to define the appearance of this HDR scene image, from which secondary gradings (also called grading images) can be derived as needed.) For simplicity, we will discuss variations in (general) luminance, thus temporarily setting aside the discussion of encoding the corresponding luma. In fact, PQ can be encoded between 1 / 10000 nit and 10000 nits, so if encoded according to this PQ EOTF (or, of course, the mapping can be represented (and implemented, for example, in the processing IC unit) as an equivalent luma mapping), then transmitting those graded pixel luminances as, for example, YCbCr pixelated HDR images with 10 bits per component is not a problem.

[0030] The two horizontal dashed lines represent the limitations of SDR-encoded images when 100 nits is associated with 100% SDR white.

[0031] Although the monster will be strongly illuminated by the firelight in the cave, we will give it an average luminance of 300 nits (due to the square power law of light dimming, skin texture, etc., the monster's luminance will have a certain distribution range).

[0032] The soldier can be 20 nit, because this is a very slight dark value, and it still provides some good basic visibility.

[0033] The vehicle might be hidden in a dark corner, and therefore, in a typical, impactful cave HDR scene, its brightness would be, for example, 0.01 nits. The goal might be to make the flames impressively bright (but not overly so). Within the available 5000 nit HDR range, we could choose 2500 nits, and around 2500 nits we could still gradually create some darker and brighter areas, but all very colorful (yellow and possibly some orange).

[0034] What happens now in a typical SDR representation (such as an SDR image captured directly from a camera)? The camera operator will open his aperture so that the soldier appears at "20 nits," or more precisely, 20%. Since the flames are much brighter (note: we are not actually showing the brightness of the real-world scene, as the master HDR video Mstr_HDR is already optimally graded for best impact in a typical living room viewing scene, and in the real world, the flames would be much brighter than the soldier, and certainly much brighter than the vehicle), the flames will be cropped to maximum white. Therefore, we will see bright areas without any detail, and not yellow, because yellow must have lower luminance (of course, a cinematographer could optimize things so that the flames are still somewhat visible even in LDR, but then that would be far less impactful extra brightness first, and second, it would sacrifice other objects that would have to be darker).

[0035] The same thing would happen if we built an SDR (maximum 100 nits) TV that would perform isoluminance mapping (i.e., it would accurately represent all the luminance of the master HDR rating it could represent, but crop all brighter objects to 100 nits white).

[0036] Therefore, the common paradigm in the LDR era was to perform a relative mapping, that is, to map the brightest luminance (in this case, luminance) of the received image to the maximum capacity of the display. Thus, since this maps 5000 nits to 100 nits by dividing by 50, the flame will still be good because the flame is distributed in yellow and orange at around 50 nits (this is for the luminance that yellow can represent, as we...). Figure 1B As seen in the image, the yellow gamut tent only decreases in luminance slightly as it moves toward the most saturated yellow, which is the opposite of the blue (B) on the other side of the slice, where blue (B) can only be produced in a relatively darker version at that hue angle B-Ye. However, this sacrifices everything else, making everything else quite dark, for example, the soldier at 20 / 50 nits, which is pure black (and this is often the problem we see in SDR renderings of such movie scenes).

[0037] Therefore, if a good HDR maximum luminance (ML_V) for master grading and a good EOTF (e.g., PQ) for encoding that HDR maximum luminance have been established, then in principle, HDR images can be started to be transmitted to receivers such as consumer TV displays, computers, cinema projectors, etc.

[0038] But this is just the most basic HDR system.

[0039] The problem is that unless the receiving side has a display capable of showing pixels as bright as 5000 nits, how to display these pixels remains an issue.

[0040] Some (DR adaptive) luminance downmapping must be performed in the television to create darker, displayable pixels. For example, if the display has a (end-user) maximum luminance ML_D of 1500 nits, then one could somehow try to calculate the yellow pixels of a flame at 1200 nits (there may be errors, such as some color shifts, like changing orange to yellow).

[0041] This luminance downmapping is not an easy task, especially to do it very accurately rather than just well enough, and therefore various techniques have been invented (also for not-so-similar luminance upmapping tasks to create output images with a larger dynamic range and, in particular, a larger maximum luminance than the input image).

[0042] Typically, it is desirable to obtain a convex mapping function in the normalized luminance (or brightness) plot (usually, i.e., for simplicity), such as Figure 1D As shown. Both input and output luminance are defined here within a range normalized to a maximum value of one. However, it must be noted that this maximum value of one corresponds to, for example, 5000 nits on the input axis and, for example, 200 nits on the output axis (which can be easily implemented by performing division and multiplication respectively). In such a normalized representation, the darkest colors will typically be too dark for the gradations with lower dynamic range in both images (here, the downconversion of the normalized output luminance L_out is shown on the vertical output axis, and all possible normalized input luminances L_in are shown on the horizontal axis). Therefore, to obtain a satisfactory output image corresponding to the input image, those darkest luminances must be relatively boosted (e.g., by multiplying by 3x, which is the slope of the luminance compression function F_comp at its darkest end). However, if it is desired that no colors are cropped to the maximum output, then this boosting cannot be indefinite; therefore, for brighter input luminances, the curve must have a gradually decreasing slope, for example, which typically maps an input of 1.0 to an output of 1.0. In any case, the luminance compression function F_comp used for degradation will typically be located above the 45-degree diagonal (diag).

[0043] Care must be taken to do this correctly. For example, some people like to apply three such compression functions separately to the red, green, and blue color channels. While this is fine and simple, it guarantees that all colors will fit into the output color gamut (RGB cube, which becomes 1000 Hz in the chroma-luminance (L) view). Figure 1B(e.g., a tent), but (especially in cases with high nonlinearity) it can lead to significant color errors. For example, a red-orange hue is determined by the percentages of red and green (e.g., 30% green and 70% red). If 30% is now doubled by the mapping function, but the red remains almost constant in the gentler part of the mapping function, then you get 60 / 70, or 50 / 50, which is yellow instead of orange. This can be particularly annoying if it depends on non-uniform scene lighting (as opposed to the SDR paradigm) (e.g., a sports car suddenly turning yellow when entering shadow).

[0044] Therefore, while the overall desired shape for color brightening may still be the function F_comp (e.g., determined by the video creator when grading a secondary image corresponding to their already optimally graded master HDR image), a more subtle downmapping is needed. For example... Figure 1B As shown, for many scenarios, the following regrading might be desirable, which only changes the brightness of the normalized luminance component (L) but not the inherent type of the color, i.e., its chromaticity (hue and saturation). If both SDR and HDR use the same red, green, and blue primary colors for representation, they will have color gamut tents of similar shape, one higher than the other in absolute luminance representation. If the two color gamuts are scaled by their respective MDWPL values ​​(e.g., MDWPL1 = 100 nits and MDWPL2 = 5000 nits), then the two color gamuts will completely overlap. The desired mapping from the HDR color C_H to the corresponding output SDR color C_L (or vice versa) will be merely a vertical shift, while the projection onto the chromaticity plane circle remains unchanged.

[0045] While the details of such a method are beyond the scope of this application, we have previously considered examples of such color mapping mechanisms in which the three color components are processed in a coordinated manner, although in separate luminance and chrominance processing paths (e.g., in WO2017157977).

[0046] If we can now use one (or more) luminance mapping functions (whose shapes can be optimized by the creators of (one or more) videos) for degradation, then more advanced HDR codecs can be designed using reversible functions.

[0047] Instead of simply performing some final secondary grading from the master image (as in television), a lower dynamic range version of the image is created for communication (the communication image Im_comm). In this example, we choose to define this image as having a maximum luminance (ML_C) of 200 nits for its communication image. Then, if the receiver receives the decoding luminance mapping function FL_dec in the metadata (which is typically essentially the inverse of the encoding luminance mapping function FL_enc, used by the encoder to map all pixel luminances of the master HDR image to corresponding lower pixel luminances in the communication image Im_comm), the receiver can reconstruct (or decode) the original 5000-nit image into a reconstructed image Rec_HDR (i.e., with the same maximum reconstructed image luminance (ML_REC)). Therefore, the proxy image used to actually transmit the higher dynamic range (DR_H, e.g., spanning from 0.001 nits to 5000 nits) image is an image with a different lower dynamic range (DR_L).

[0048] Interestingly, it's even possible to select a 100-nit LDR (i.e., SDR) image for communication, which is immediately ready (without additional color processing) to be displayed on a traditional LDR image (this is highly advantageous because traditional displays lack HDR knowledge). How does this work? Traditional televisions don't recognize MDWPL metadata (because this metadata doesn't exist in the SDR video standard, and therefore televisions aren't configured to look for it somewhere in the signal (e.g., in supplemental enhancement information messages, which is the mechanism by which MPEG introduces all sorts of pre-agreed new technology information)). Nor does it look for that function. It simply looks at a YCbCr, for example, 1920×1080 pixel color array and displays those colors as usual (i.e., as interpreted according to SDR Rec. 709). And the creator has already selected its FL_enc function in that particular codec embodiment, such that all colors (even flames) map to a reasonable color within a limited SDR range. It should be noted that, contrary to the simple multiplicative changes corresponding to opening or closing the camera aperture in SDR production (which typically result in cropping to at least one of white and / or black), it is now possible to choose very complex optimal function shapes, provided they are reversible (e.g., we have already taught a system that first performs a coarse pre-grading and then a fine grading). For example, the luminance (or relative brightness) of a car can be shifted to a level that is just visible in SDR (e.g., 1% deep black), while a flame can be shifted to 90% (provided everything remains reversible). This may seem extremely difficult, even impossible, at first glance, but numerous field tests using a variety of video materials and use cases have shown that it is possible in practice, provided it is done correctly (following, for example, the principles of WO2017157977).

[0049] How do we now know that this is actually an HDR video signal, even though it contains pixel-rich color images available for LDR, or that any receiver with HDR capabilities can reconstruct it as HDR? Because the metadata also contains the function FL_dec, typically one FL_dec function per image. Therefore, the signal is also chromaticly encoded as an HDR image (5000 nits) according to the discussion and definitions above.

[0050] Although it is more complex than a basic system that only transmits PQ-HDR images, this per-SDR proxy encoding is still not the optimal future-proof system because it still leaves the receiving side to guess how to downmap colors when it has a TV with, for example, 1500 nits or even 550 nits.

[0051] Therefore, we added further technical insights and developed a technique known as display tuning (also called display adaptation): the image can be tuned for any TV that can be connected (i.e., any ML_D) because the power of the encoding function FL_enc can be doubled as some kind of guiding function for mapping upwards from 100nit Im_comm (but not to a reconstructed image of 5000nit). The concave function (which is essentially the inverse of F_comp) (it should be noted that for display tuning, a function as precise as the inverse function for reconstruction is not needed) will now have to be scaled to be slightly less steep (i.e., the luminance mapping function FL_DA for display adaptation will be calculated based on the reference decoding function FL_dec) because we are only scaling to 1500nit instead of 5000nit. That is, an image with a three-level dynamic range (DR_T) can be calculated (e.g., optimized for a specific display) such that the maximum luminance of that three-level dynamic range is generally the same as the maximum displayable luminance of that particular display.

[0052] The technique used for this is described in WO2017108906 (we can transform a function of any shape into a similar shape that is closer to the 45-degree diagonal, the extent of which depends on: the ratio between the maximum luminance of the input image and the maximum luminance of the desired output image, relative to the ratio between the maximum luminance of the input image and the maximum luminance of the reference image (here, the reconstructed image), for example, by using this ratio to obtain closer points on the diagonal from which line segments are orthogonally projected from the corresponding diagonal points until the line segments intersect with points on the input function, these closer points together define the tuned output function used to calculate the luminance of the image to be displayed, Im_disp, based on the Im_comm luminance).

[0053] We have gained access to a wider variety of displays for basic movie or television video content (LCD TVs, mobile phones, home theater projectors, professional cinema digital projectors) and more diverse video sources and communication media (satellite, internet streaming media (e.g., OTT), 5G streaming media), as well as more ways to produce videos.

[0054] Figure 2 Several typical video creations in which this teaching can be usefully deployed are shown (in general, and without limitation).

[0055] In a studio environment (e.g., for news or comedy), there will still likely be a strictly controlled shooting environment (although HDR allows for relaxation of this and allows shooting in real-world settings). There will be controlled lighting (202), such as baselights on the ceiling and various spotlights. There will be many bulky, relatively stationary television cameras (201). This variant, typically broadcast live, would be, for example, a football-like sports program, which would feature various types of cameras, such as cameras near the goal for close-up views, panoramic cameras, drones, etc.

[0056] There will be some production environment 203 where various feeds from cameras can be selected to become the final feed, and various (usually simple but potentially more complex) hierarchical decisions can be made. In the past, this typically occurred, for example, in a broadcast van with many monitors and various operators, but in an internet-based workflow, the raw feeds can be transmitted via some network, and the final assembly may occur at the broadcaster's office. Finally, to simplify the production process for this explanation, some encoding and formatting will occur in formatter 204 for distributing the broadcast to the final (or intermediate, such as a local cable television station) clients. This will typically involve converting the luminance of the hierarchical format as explained in Figure 1 to, for example, PQYCbCr for, for example, an intermediate dynamic range format, calculating and formatting all the necessary metadata, converting to some broadcast format (such as DVB or ATSC), packaging in chunks for distribution, etc. ("etc." indicates that there may be added for signaling available content, captions, encryption tables, but at least some of these are not important to understanding the details of this technological innovation).

[0057] In this example, video (in this case, a television broadcast) is transmitted via television satellite 250 to satellite antenna 260 and set-top box 261 capable of processing satellite signals. Ultimately, the video will be displayed on end-user display 263.

[0058] The display can show the first video, but it can also show other video feeds, and may even show them simultaneously, for example, in a picture-in-picture window (or some data of the first HDR video program may appear via one distribution mechanism, while other data may appear via another distribution mechanism).

[0059] The second stage of production is typically done offline. This could be a Hollywood movie, or even a short film about someone racing in the jungle. Such production could be shot with other optimal cameras, such as a Steadicam 211 and a drone 210. Again, we assume the camera feed (which could be raw or already converted to some HDR production format, such as HLG) is stored somewhere on a network 212 for later processing. In such a production, we might have a human colorist using grading equipment 213 to determine the optimal luminance for the master grading (or, in the case of HLG production and encoding, the relative brightness of the master grading) during the final months of production. The video can then be uploaded to an internet-based video service 251. For professional video distribution, this could be, for example, Netflix.

[0060] The third example is consumer video production. Here, the user (e.g., when creating a video blog) would have a ring light 221 and would be captured via a mobile phone 220, but the user could also be captured from an external location without auxiliary lighting. The user would typically also upload to the internet, but nowadays it might be to YouTube or TikTok, etc.

[0061] When receiving via the Internet, the display 263 will be connected via a modem or router 262, etc. (More complex setups, such as indoor Wi-Fi, are not shown in this simple description).

[0062] Another user can watch the video content on a portable display (271), such as a laptop (or similarly, another user can use a non-portable desktop PC) or a mobile phone. They can access the content via a wireless connection (270) (e.g., Wi-Fi, 5G, etc.).

[0063] Currently, professional reference monitors used to view created videos can reach up to ML_D = 10,000 nits. Even in the consumer market, 10,000-nit televisions have recently been showcased. Therefore, if things develop favorably, we may soon see high-quality HDR imagery becoming commonplace in various applications. It should be noted that cameras don't actually capture precise luminance, but rather the relative brightness of the scene (due to arbitrary aperture selection, etc.). However, today's cameras have no problem capturing high dynamic range. Even standard cameras can at least faithfully capture the scene's dynamic range. The difference in brightness, while more advanced cameras with dedicated sensor construction or techniques such as multiple exposures can far exceed this range (e.g., This is a fairly reasonable way to capture dynamic range in most cases.

[0064] Therefore, it can be seen that today, a wide variety of videos can be generated and transmitted in various ways, encoded using various technologies, and our encoding and processing systems have been designed to handle virtually all of these variations.

[0065] Figure 3 illustrates an example of an absolute (nit-level defined) dynamic range conversion circuit 300 for a (HDR) image or video decoder (the encoder works in a similar manner, but typically has an inverse function, i.e., the function to be applied is the function mirrored on the other side of the diagonal). It is based on encoding a master image (e.g., a master HDR grade) with a master luminance dynamic range (DR_Prim) into another (so-called surrogate) image with a different secondary pixel luminance range (DR_Sec). If the encoder and all its available decoders have pre-agreed upon or know that the surrogate image has a maximum luminance of 100 nits, then this does not need to be transmitted as SDR_WPL metadata. If the surrogate image is, for example, a maximum image of 200 nits, then this will be indicated by filling its surrogate white point luminance P_WPL with a value of 200, or similar processing for 80 nits, etc. The maximum value (HDR_WPL = 1000) of the master image to be reconstructed by the dynamic range conversion circuit is typically transmitted as metadata of the received input image or video signal (i.e., along with the input pixel color triplets (Y_in, Cb_in, Cr_in)). Various pixel luma are typically input as a luma image plane, i.e., sequential pixels will have a first luma Y11, a second luma Y21, etc. (These luma are typically scanned, and the dynamic range conversion circuit converts them pixel-by-pixel to the output pixel color triplets (Y_out, Cb_out, Cr_out)). In this description, we will focus primarily on the luminance magnitude of pixel colors. Various dynamic range conversion circuits may operate internally in different ways to achieve essentially the same purpose: the correct reconstruction of the output luminance L_out of all image pixels (the actual details are not important to this invention, and the embodiments will focus only on aspects necessary for teaching).

[0066] The mapping of luminance from the secondary dynamic range to the master dynamic range can be applied directly to the luminance itself or to any luma representation (i.e., according to any EOTF or OETF), as long as it is done correctly; for example, it does not have to be done separately on the nonlinear R'G'B' components. The internal luma representation does not even have to be the luma representation of the input (i.e., Y_in), nor does it have to be the luma representation of any output that the dynamic range conversion circuitry or its contained decoder might deliver (e.g., the format luma Y_sigfm for a specific communication format or communication system, where "communication" includes storage in memory (e.g., inside a PC, on a hard disk, on an optical storage medium, etc.)).

[0067] We have optionally shown (in dashed lines) the luma conversion circuit 301, which converts the input luma Y_in into a perceptually uniform luma Y_pc.

[0068] The applicant standardized a useful formula in ETSI 103433 for converting luminance in any range into such a perceived luma representation: The function RHO is defined as follows: The value WPL_inrep is the maximum luminance of the luminance range that needs to be converted to a psychologically visually uniform luma. Therefore, for a 100 nit SDR image, this value will be 100, while for the output image to be reconstructed (or the original encoded image on the creation side), this value will be 1000.

[0069] Ln_in is the luminance to be converted, which is converted along an arbitrary range after being normalized by dividing by its respective maximum luminance (i.e., within the range [0, 1]).

[0070] Once we have the input and output ranges normalized to 1.0, we can actually apply the luminance mapping function in the luma domain as shown inside the luma mapping circuit 302, which performs the actual luma mapping for each input pixel.

[0071] In fact, this mapping function has been specifically chosen by the image encoder (at least to produce good reconstruction quality and potentially reduce the number of bits required for MPEG compression, but sometimes also to meet other criteria, such as the correct luminance distribution of an SDR proxy image on a conventional SDR display for a specific scene (a dark cave or a daytime explosion). Therefore, the function F_dec (or its inverse) is extracted from the metadata of the input image signal or representation, and this function F_dec (or its inverse) is supplied to the dynamic range conversion circuitry for the actual per-pixel luma mapping. In this example, the function F_dec directly specifies the desired mapping in the perceptual luma domain, but other variations are certainly possible, as various transformations can also be applied to these functions. Furthermore, while this paper teaches a pure decoder dynamic range conversion for simplicity and to ensure understanding of this teaching, other dynamic range conversions can also use other functions, such as those derived from F_dec. Understanding the innovative contribution of this invention to the technology does not require all of these details.

[0072] Typically, not only luminance is altered, but corresponding changes are also made to chroma Cb and Cr. This can be done in various ways, from strict inverse decoding to implementing additional features (such as saturation boosting), since Cb and Cr encode the saturation of a pixel. Therefore, another function (the recoloring canonical function FCOL) is usually passed in the metadata, which determines the chroma recoloring behavior, i.e., the mapping of Cb and Cr. (It should be noted that Cb and Cr are usually changed by the same factor, because the Cr / Cb ratio determines the hue, and hue changes are generally not desired during decoding; i.e., lower dynamic range images and higher dynamic range images will typically have object pixels with different luminance, and typically at least some pixels will have different saturation, but ideally, the hue of pixels in both image versions will be the same.) This color function will typically specify a multiplier with a value dependent on the luminance code Y (e.g., Y_pc, or other codes in other variants). The multiplier establishment circuit 305 will generate the correct multiplier m for the luminance condition of the pixel being processed. Multiplier 306 multiplies both Cb_in and Cr_in by the same multiplier to obtain the corresponding output chromaticity Cb_out = m * Cb_in and Cr_out = m * Cr_in. Therefore, the multiplier achieves correct chromaticity processing, thus enabling proper configuration of overall color processing for any dynamic range conversion within the dynamic range conversion circuitry.

[0073] Furthermore, (at least in the decoder) a formatting circuit 310 can typically be present, allowing the output color triplets (Y_out, Cb_out, Cr_out) to be converted to any desired output format (e.g., RGB format or communication YCbCr format (Y_sigfm, Cb_sigfm, Cr_sigfm)). For example, if the circuit outputs to a version of communication channel 379 as an HDMI cable, such a cable typically uses PQ-based YCbCr pixel color encoding; therefore, the luma will again be converted from the perceptual domain to the PQ domain by the formatting circuit.

[0074] Importantly, the reader should fully understand what encoding (decoding) is, and how an absolute HDR image or its pixel colors differ from a traditional SDR image. A connection may exist with a display tuning circuit 380, which calculates the final pixel colors and brightness to be displayed on the screen of some displays (such as 450-nit televisions in some consumer homes).

[0075] However, in absolute HDR, pixel luminance can be established during the decoding step, at least for the output image (here, a 1000 nit image).

[0076] We are Figure 3B This is illustrated in the example, which uses HDR images of some typical indoor / outdoor images. It should be noted that while outdoor luminance can often be 100 times that of indoor luminance in the real world, it might be better to make them, for example, 10x brighter in actual master-grade HDR images because viewers will be viewing everything together on the screen at a fixed angle, even in what is typically a dimly lit room at night, rather than in the real world.

[0077] Nevertheless, we also found that when looking at the luminance corresponding to luma (e.g., HDR luminance L_out), a large histogram is typically seen (a histogram of the count N(L_out) of each luminance appearing in the output image of this home scene). This spans significantly beyond some lower luminance dynamic range lobes and beyond the low dynamic range 100 nit level, as outdoor images on sunny days have their own histogram lobes. It should be noted that the luminance representation is drawn in a non-linear manner (e.g., logarithmically). We can also trace back what the encoder does when creating a 100 nit proxy image (and its proxy luminance count N(L_in) histogram) on the encoding side. As shown in Figure 1 or inside the luma mapper 302, convex functions are used to compress the brighter luminance due to the limitations of the smaller luminance dynamic range. There are still some differences between indoor and outdoor luminance, and the upper lobe of outdoor objects still has a considerable range, allowing one to still create the different colors needed to color various objects (e.g., the different shades of green in a tree). However, there must also be some sacrifices. First, indoor objects will appear darker (assuming for now that 100 or 200 nit displays will faithfully display the encoded luminance, rather than undergoing some arbitrary beautification process to brighten them, for example) than ideally (i.e., up to the indoor threshold luminance T_in of the HDR image). Second, the upper lobe span is compressed, which can result in lower contrast for outdoor objects. Third, since bright colors in the tent-shaped color gamut shown in Figure 1 cannot have high saturation, outdoor colors may also have a slightly muted tone, i.e., reduced saturation. However, of course, if the colorist on the creation side has control over all functions (F_enc, FCOL), then he can balance these features, making some features slightly biased while sacrificing others. For example, if the outdoor display shows a pure blue sky, then the colorist can choose to make it brighter but with less blue. If there is a beautiful sunset, he might want to preserve all its colors, but make all the colors darker, especially if there are no significant vignettes in the indoor portion of the image, then this would make the contents of those vignettes difficult to see, especially when watching television with all the lights on (it should be noted that there are also techniques for handling lighting differences and black level visibility, but that is too informative for the description of this patent application).

[0078] The middle diagram illustrates what a luma for surrogate luminance might look like, and it typically provides a more uniform histogram; for example, the luminance of objects in an indoor image has roughly the same span as that in an outdoor image. However, luma is only relevant to the degree of coded luminance, or only if some calculations are actually performed in the luma domain (which has advantages for word length on processing circuitry). It should be noted that while the absolute form can assign luminance between zero and 100 nits on the input side, SDR luminance can also be considered as relative brightness (what conventional displays do when discarding all HDR knowledge and transmitted metadata and only looking at 0-255 luma and chroma code).

[0079] There are some points to note regarding the direct presentation paradigm of video decoding. In fact, as we have already explained, in the SDR era and at least for absolute nit HDR encoding, the idea was that the viewer would see exactly what the creator wanted him to see because (aside from all the offsetting hardware limitations) the display would render the various gray percentages (or more generally, all colors) exactly as created (and as can be seen on the reference monitor in the creative studio, which should behave roughly similarly to a home monitor, provided that both monitors have the same white point luminance; and the percentage luminance relationships roughly correspond to the actual luminance relationships of at least the important objects in the scene, and to how the camera captures these objects, at least in the SDR era, except for camera details).

[0080] This is where we are Figure 4The behavior observed (with the standard display driver TF_std) explains what was possible back in the CRT era (the typical monitor of the SDR video era). If a drive signal s_driv is applied to the monitor (for simplicity, assume only one color channel), temporarily ignoring the monitor's power-law behavior on that signal—that is, assuming it has been pre-squared into an equivalent linear drive signal—then we can see that the monitor can be driven to the maximum value of the drive signal, which can be called "1" or 100%, and this corresponds to the monitor's brightest display, as it cannot display anything brighter, which is determined by the fact that it is directly controlled by the monitor. If we then ask the monitor to display, for example, 25% of its maximum value to create some gray (i.e., s_driv = 0.25), then the displayed screen luminance (L_screen) will indeed be 25% of the brightest pixel the screen can see. And the eye and brain will adapt to the displayed image, recognizing the illuminance and the visual perception or perceptual quality of white, as well as all other colors. Therefore, a brightness (grayscale) drive signal s_driv "0" will display the darkest possible color on the display, i.e., black level (e.g., an OLED display will make a portion of the OLED not emit light), 10% should give 10% output, and 100% will give 100% output, whatever that may be (e.g., a relative display of a traditional SDR display will, for example, display white as 150 nits if that is a purchased display and is a custom brightness setting not retained as factory settings; and an absolute display should ideally display the maximum indicated pixel luminance of the image, or some reduced version thereof).

[0081] However, in a studio, people might view a monitor under ambient lighting that results in an average reflected gray object brightness of, for example, 10 nits around the monitor. This is not the case for consumers: their viewing environment can vary significantly (often being essentially static on average, but potentially dynamic if, for example, the sun is replaced by incoming storm clouds). One user might watch the same movie in near darkness, with only their device's LED indicators and some other stray light illuminating the room, while another viewer, or the same viewer at a different time, might watch the movie in bright sunlight during the day. While some extreme cases are difficult to fully compensate for and therefore technically beyond our reach (one may never be able to compensate for direct sunlight on a monitor screen that can only display a maximum brightness of 100 nits, in which case curtains should be drawn), it is advantageous if the monitor has some controllable mechanism to correct for situations such as a companion reading a book a little further away in the room. Therefore, CRTs already have two separately configurable control mechanisms. The first is called "contrast," and it sets the contrast, or difference between adjacent grayscale values, by acting on the amount of light emitted by the brightest white, through (multiplicative) luminance metric Adj_contrast added to the white Whi color (and linearly down to all colors in proportion to their respective values). Some displays (such as older LCDs with fixed backlights) cannot electronically adjust this. In any case, a second control setting called "brightness" (an interesting name, since contrast also acts on brightness, which primarily acts on the brightness of darker colors, hence also called "offset") adds a fixed amount as a black level offset (Adj_brightn) to all luminance values. The formula for modeling the controllable electronic behavior, especially for SDR displays, would be: [Formula 1] This will produce several controllable versions of the modified display driver function (TF_adj).

[0082] It should be noted that it may not be desirable to drive the maximum brightness of HDR in the same way (because there are more complex processes involved, both in terms of the luminance distribution of the received HDR image and in how it is adapted for display), nor may it be desirable to adjust the contrast of one or more intermediate areas in such a coarse manner. However, regarding black levels, especially since in many cases it may not be necessary to pay close attention to the ultra-deep HDR black levels, it may be a reasonable technical goal to make the HDR display scene with black levels the same or similar to the SDR scene.

[0083] Therefore, in HDR image reception and display scenarios, at least the additive constant Adj_brightn is used to boost all luminance (and possibly some multiplicative boosting using Adj_contrast, but assumed to be left at "1", i.e., a standard value). This will primarily affect the visual effect on darker colors in the image (it should also be noted that the visual system (i.e., processing in the eye and brain) makes it approximately logarithmic, so the difference in "black level" (i.e., the darkest color) is considered more significant than adding a small additive change to a bright color). Specifically, the darkest black level is boosted to a value close to the typical luminance level derived from the ambient lighting in the room. While some visual adaptation exists in the human eye, the main problem here will be reflections from brighter environments (e.g., walls) on the screen in front of the display, which will mask image details in the darkest screen luminance, at least above zero nits. For example, the flooding level Lev_illumdrown might average 5 nits, added to the displayed image according to the following formula: [Formula 2] Where L_screen_electrically_displayed is the amount of colored light and luminance electronically generated by the display based on the input image (i.e., Equation 1). L_screen_seed is the total luminance propagating to the viewer's eye. Therefore, a perfectly black pixel in the image (s_driv = 0) is still perceived as a slightly brighter black, i.e., the black expected to be seen in such a slightly brighter viewing environment. If the Adj_brightn value is set near the average Lev_luminhdrown value (e.g., slightly lower), then the user can clearly see all the different very dark blacks and grays driven slightly above the perfect black level (Blk), i.e., s_drive = 0 (i.e., s_driv = 1 or 2, etc.). That is, the user can still see all the image details in the darkest areas of the image, such as a battle taking place at night, or a monster lurking in the darkness.

[0084] US2015 / 213781 describes a control system that increases the brightness of a device (such as a laptop or smartphone) to offset increased ambient light, such as when moving from indoors to outdoors or near a window. If the device is plugged into a power outlet, the brightness of the backlight (now typically LED) can be increased. When battery powered, this may drain the battery more quickly. Another way to increase the brightness of visible image areas (pixels) on the screen is by processing the image itself, thus avoiding additional battery consumption without changing the backlight brightness. This is achieved by applying a gamma function (or a LUT implementing that gamma function), i.e., Here, gain is a power of the measured ambient illuminance (e.g., when normalized to 1.0 as the maximum luma representation, 1 / 3 increases pixel brightness more than 1 / 2). This increases the brightness of dark and mid-brightness pixels (at the cost of some compression of the brightness and contrast of the brightest pixels, so this operation is best performed on images with a natural dark appearance (e.g., images with few code values ​​higher than 80 out of 255 code values), which can be determined by histogram analysis of the input image.

[0085] Although this illumination countermeasure control method has worked well for many years, the inventors of this invention have found that it is not the best way to achieve high-quality display of HDR images. Summary of the Invention

[0086] The indicated problem is handled by a control processor (510), which is configured to adjust the brightness of at least a dark subset of the input image (IMG_in) displayed on the display (520) in response to a measurement of the ambient illuminance (Lev_illumdrown) of the display. This adjustment includes establishing an initial adjustment function (F_adj) for shifting one or more color components of the input color component triplets (R'_PQ, G'_PQ, B'_PQ) of the pixels of the input image by at least one correction value (Adj_brightn) to obtain the corresponding output color component triplets (R'_out, G'_out, B'_brightn). The adjustment includes multiplying the initial adjustment function (F_adj) by the correction function (F_corr) for the values ​​of one or more color components in the input color component triplet (R'_PQ, G'_PQ, B'_PQ) below a threshold (GAM1) to obtain a final version of the adjustment function for shifting the input color component triplet (R'_PQ, G'_PQ, B'_PQ), wherein the correction function (F_corr) satisfies: a first property of producing zero output for zero input; a second property of producing an output value equal to one for an input value equal to the threshold (GAM1); and a third property of having a first derivative of zero at the input value equal to the threshold (GAM1).

[0087] Standard ambient adjustment functions (such as) Figure 4The explanation provided applies to many types of image content because the distribution of values ​​indicates a reasonable level of black (black is a receptive quality in the brain that can vary to some extent: aside from absolute minimum black, if the brightness of dark grays drops below 5% of white, then they begin to appear black). Better or worse-looking blacks can be produced under good dim viewing conditions, but under significant ambient lighting, the visibility of all dark pixels is the dominant criterion for better-looking blacks, while slightly worse blacks are acceptable. This is also true when ambient lighting is not constant in space or time. For example, some parts of a movie might reflect slightly brighter ambient objects onto a specularly reflected screen, such as a lampshade in a room, and the brightness might change if someone walks past the lampshade. The initial adjustment function (F_adj) depends on the illuminance. The exact selection of the initial adjustment function is a technical guideline for the implementer (e.g., in a television display or mobile phone) and is beyond the scope of this patent application, as the invention lies in improving the brightness adjustment behavior of the darkest pixel in any image, and the resulting function shape (which can be defined as the distance from the horizontal axis in the function graph). That is, the concept works for any adjustment function that might be found in practice (i.e., at least increasing the brightness of the darkest pixel). This increase can, in principle, be large enough that the brightness of all pixels is raised to a level far exceeding the ambient lighting reflection level represented by a certain base black level luminance. However, since this would significantly reduce image contrast, the function can also increase a portion of that value (e.g., the estimated average luminance from the worst screen position divided by 2, 4, or 10, etc.), because only for certain applications is it acceptable for poor viewing of black details as long as it remains viewable. However, a common characteristic of the alternative initial adjustment functions is that they raise the zero level to some non-zero output (as an electronic luma code used to drive the display, and correspondingly, a non-zero pixel brightness amount, typically higher than the light leakage of some displays driven by zero luma inputs to image pixels). Some embodiments may boost the level of the darkest pixels based on the pixels in the image, but in other embodiments, this will not be done, and a general ambient brightness adjustment will be applied regardless of the content in one or more consecutive images. For some kinds of completely black and large image objects (e.g., black bars from spatial format conversion (e.g., a 21:9 movie shown on a 16:9 television)), conventional solutions may not be preferred because those objects may appear rather gray than true black.

[0088] Therefore, the inventors have devised a correction strategy that keeps most of the processing unchanged (because highly tuned ambient light compensation may exist in some embodiments), but at least ensures that true black objects defined in the input video appear true black. In a typical embodiment, the remainder of the initial adjustment function above a threshold (GAM1) will remain the same as in the initial adjustment function (i.e., it can be considered as continuing to multiply by 1.0 above GAM1).

[0089] Actual color processing can be performed in various color spaces. What they have in common is that additive color mixing mechanisms typically require three color components to produce virtually any desired color. However, the primary colors (also known as base colors) can vary. For example, EBU primary colors can be reused while maintaining SDR chromaticity and a so-called narrow color gamut; however, those EBU primary colors can be given a significantly increased maximum brightness (compared to a classic CRT, which could produce red pixels with a maximum brightness of around 30 nits and no higher, current OLEDs, for example, can produce much brighter red primary colors; therefore, in the color triplet that defines any color, the red component can contribute brighter than ever before, thus producing brighter colors with a larger color gamut, although the component is still defined in the RGB color model). However, the saturation (i.e., color gamut) of the colors that can be produced must still fall within the boundaries of the triangle spanned by the chosen EBU primary colors. If you want to be able to produce more saturated (still RGB) colors, you can use, for example, DCI_P3 or Rec. 2020 primary colors.

[0090] While the driving of additive color displays (whether OLED, LCD, or any other type) is always performed in RGB (because it must create the luminance of red, green, and blue subpixels in appropriate weighted proportions at their positions on the screen), the calculation of the desired driving color (and thus the adjustments in this embodiment) can be performed in different color representations or models. Since those skilled in optimizing color displays should know how to convert between color models (i.e., the selection of the three components) or color spaces (i.e., the quantization of the model, assigning specific values ​​to various definable colors), which typically involves formulas using the defined color space, this document will only illustrate two typical examples.

[0091] If color is processed in separate luma (Y) and chromaticity (Cb and Cr) processing, then the adjustment results and their novel correction functions can usually be applied only to the luma component (Cb and Cr can also be corrected in reverse, but can also be processed arbitrarily or even remain unchanged).

[0092] If processed in an RGB color model (with any nonlinear electro-optic transfer function specification, i.e., transforming the percentage or absolute contributions of linear red, green, and blue by, for example, an inverse EOTF using a perceptual quantizer normalized in SMPTE 2084), which is sometimes necessary in some color processing architectures, then the three color components will receive function-based adjustments as taught in this paper.

[0093] In some embodiments, there may be only one correction value (Adj_brightn), which is applied similarly to all possible input color component values, at least in the initial adjustment function (F_adj). However, a more advanced approach may be to apply differential brightening, in which the initial adjustment function (F_adj) shifts the input component values ​​by different amounts depending on the values ​​of the input components. That is, more than one correction value (Adj_brightn) will be applied in the initially envisioned ambient lighting adaptive processing.

[0094] This improved method can work in any scenario because it corrects the initial adjustment function (F_adj) for dark input values ​​(defining dark image colors) below a selectable threshold (GAM1). Therefore, in luma processing, this threshold will be used for the input luma, and RGB processing will use the same threshold three times, processing each color channel (e.g., the red channel) separately (typically in parallel). The threshold can be chosen as a pre-fixed small value (e.g., 0.1 nit or less, or a relative value of 0.0001), but some embodiments may choose a threshold equal to the measured ambient illuminance (Lev_illumdrown). This value can be evaluated once, for example, read at the beginning as a previous value while watching television, and then corrected to the new value viewed today after, for example, 5 minutes, and then held as this new value, or the value can be continuously adapted, but not on overly fine scales; because it is undesirable to see the black level fluctuate back and forth, it is performed, for example, every 10 minutes or more. Therefore, for all possible input values, the correction function is multiplied by the initial adjustment function (F_adj) to obtain the final adjustment function, which will be used to perform adaptive brightening of the ambient lighting mainly for the darker pixels (although, as mentioned above, if pure constant addition were used as the initial adjustment, this would also change the brightest pixel, but as mentioned above, this is done in a less visually noticeable way).

[0095] Advantageously, the control processor (510) applies a strictly monotonically increasing adjustment function (it should be noted that the initial adjustment function may have, for example, a pruning above the maximum value). In this paper, we will use the general definition that a monotonically increasing function, when making its input higher, cannot have a lower output for higher input values ​​than for lower input values ​​(although it can have a constant output for input values ​​increasing over a certain range). A strictly monotonically increasing function must always have a higher output value for higher input values ​​than for previously lower input values; that is, a strictly monotonically increasing function cannot have a flat region of output values.

[0096] Advantageously, the control processor (510) applies its illumination adjustment and thus its correction function to the input and output domains of the psychovisually homogenized color component values. An example of such a function is the perceptual quantizer function of SMPTE 2084, which, when processing only the luma component, pre-applies an inverse PQ EOTF to the pixel's luminance and then applies illumination adjustment to luma in the homogenization domain. When processing the three RGB components in parallel, an inverse PQ EOTF is applied to the three linear red, green, and blue component contributions of any input pixel color being processed, and then illumination adjustment is applied to the nonlinear representation of these three resulting color components (the nonlinearity is typically indicated, for example, by the superscript apostrophe R' rather than linear R). Another example of a psychovisually homogenized function is the Philips homogenization (PU) function as specified in ETSI 1034233 (which is incorporated herein by reference in its entirety), namely: The function RHO is defined as follows: The value WPL_inrep is the maximum luminance of the range of lumens that needs to be converted to a psychologically visually uniform lumens. Therefore, for a 100 nit SDR image, the value will be 100, and for a maximum image of 1000 nits, the value will be 1000.

[0097] Ln_in is the luminance to be converted, which is converted along an arbitrary range after being normalized by dividing by its respective maximum luminance (i.e., within the range [0, 1]).

[0098] Therefore, prior to actual processing, the inverse perceptual quantizer electro-optic transfer function is pre-applied to linear versions of one or more color components of the input color component triplet (in some embodiments, ambient lighting acts as a post-processor, and these linear versions themselves may have already undergone image processing (e.g., display adaptation), but in embodiments where all processing occurs in a single step, these one or more components may be, for example, input color components received in a broadcast video signal). Those skilled in the art will understand the psychovisually homogenized set of luma as a code such that differences in the code (e.g., 5) have approximately the same effect regardless of which starting level is applied (the effect being that the numerical representation of luma between their minimum and maximum values ​​represents a visually equal scale of brightness increment steps; for example, luma10 represents the darkest black level, luma20 represents a brightness increment above the darkest black level, luma30 represents two such steps up to the brightest white level). The transformation from a non-visually uniform representation to a visually uniform representation can be pre-applied by using a non-linear scaling function (e.g., if we know the first luminance Lx = 2x the second luminance Ly, and Ly = 2x Lz, then this function maps luminance Lx to luminance Yx, the second luminance Ly to Yx-step, and the third luminance Lz to the third luma Yz = Yx-2 * step). Since the uniformized representation does not need to be 100% perfectly uniform, but can be largely uniform, different functions with roughly the same mapping behavior from luminance to uniform luma can be used. This function can be applied to color components, such as the non-uniform luma Y or (non-uniform) luminance, and then typically the same mapping function can be applied to the other two color components, such as Cb and Cr. Alternatively, the same function can be applied to the red, green, and blue components to obtain uniform (i.e., substantially uniform) non-linear red, green, and blue components.

[0099] Advantageously, the control processor (510) uses an embodiment of a correction function that has multiple higher-order derivatives equal to zero at the input value equal to the threshold (GAM1).

[0100] Useful examples of correction functions applied in practice include: Where x is the input value between zero and a linear representation of a threshold (GAM1), which is normalized to the maximum value equal to 1, and y is the output value of the correction function for the input value; or .

[0101] The control processor can be located in any device prepared for display (such as a computer) or in the display itself.

[0102] The invention can also be embodied in a method that adjusts the brightness of at least a dark subset of an input image (IMG_in) displayed on a display (520) in response to a measurement of the illuminance (Lev_illumdrown) of the environment in which the display is located, wherein the adjustment includes establishing an initial adjustment function (F_adj) for shifting one or more color components of the input color component triplets (R'_PQ, G'_PQ, B'_PQ) of the pixels of the input image by at least one correction value (Adj_brightn) to obtain corresponding output color component triplets (R'_out, G'_out, B'_out). In this process, the adjustment includes: for the values ​​of one or more color components in the input color component triplet (R'_PQ, G'_PQ, B'_PQ) below a threshold (GAM1), multiplying the initial adjustment function (F_adj) by the correction function (F_corr) to obtain a final version of the adjustment function used to shift the input color component triplet (R'_PQ, G'_PQ, B'_PQ), wherein the correction function (F_corr) satisfies: a first property of producing zero output for zero input; a second property of producing an output value equal to one for an input value equal to the threshold (GAM1); and a third property of having a first derivative of zero at the input value equal to the threshold (GAM1).

[0103] In a useful embodiment of the method for adjusting the brightness of a display, the embodiment of the correction function is strictly monotonically increasing.

[0104] In a useful embodiment of the method for adjusting the brightness of a display, (e.g., by pre-applying the inverse perceptual quantizer electro-optic transfer function to a linear version of one or more color components in the input color component triplet) a correction function is applied to the input and output domains of the psychovisually uniform color component values ​​(i.e., both the input and output values ​​are pre-transformed from their reference linear representations to any such domain, typically defined by the corresponding EOTF or OETF).

[0105] Various methods can be embodied in computer program products that include software code used to instruct the processor to execute one of these methods, and the software code can be delivered, for example, via the Internet (e.g., for firmware upgrades of displays). Attached Figure Description

[0106] These and other aspects of the methods and apparatus according to the invention will become apparent and elucidated by referring to the embodiments and examples described below and by referring to the accompanying drawings. The drawings are provided only as non-limiting specific illustrations to illustrate a more general concept, and in the drawings, dashed lines are used to indicate components that are optional, while non-dashed components are not necessarily essential. Dashed lines can also be used to indicate elements that are interpreted as essential but are hidden inside an object, or for intangible things, such as the selection of objects / areas.

[0107] In the attached diagram: Figure 1 schematically illustrates the various luminance dynamic range regradings that can occur in multiple steps of the HDR image processing chain or the desired image, i.e., the mapping between the input luminance (typically specified as desired by the creator of the video or image content, either human or machine) and the corresponding output luminance (of various image object pixels). Figure 2 This paper illustrates (non-limiting) examples of some typical use cases for HDR video or image communication from the source (e.g., production) to the user (typically a home consumer); Figure 3 schematically illustrates how the corresponding conversion between luminance and pixel chrominance can work, and what the pixel luminance histogram distribution of the image can look like for the dynamic range of lower and higher luminance (here, especially luminance). Figure 4 This schematically illustrates how a display can be controlled to show the same input image in different ways to adapt to the amount of ambient lighting at the location where the image is displayed; Figure 5 A first circuit diagram for implementing the present invention is schematically shown (one or more circuits can operate under software control and may be at least partially configurable). Figure 6 A second circuit diagram for implementing the present invention is schematically shown (one or more circuits can operate under software control and may be at least partially configurable). Figure 7 The operation of an embodiment of the present invention is illustrated numerically. This embodiment of the present invention is used to achieve better ambient lighting adaptation for a scene with a maximum display of 1000 nits and an ambient illuminance of 1 nit (other initial adjustment function shapes are also possible). Figure 8 The same embodiment is illustrated, but in a scenario with a maximum display brightness of 1000 nits and an ambient illuminance of 10 nits; and Figure 9 The shapes of some embodiments of the correction function that can be used according to this innovative principle are shown. Detailed Implementation

[0108] Figure 5 A first embodiment of how improved ambient lighting compensation is implemented is shown. In this variant, the control processor (510) is implemented as a post-processor and acts on an input color component triplet (a red nonlinear component R'_PQ and green and blue nonlinear components G'_PQ and B'_PQ), which forms the pixels of the input image IMG_in in the ambient lighting adjustment stage of this invention, and is output from the preprocessing stage, where the nonlinear components in this example are nonlinearized according to the inverse perceptual quantizer EOTF. The control processor includes three function-based mapping circuits (a first function-based mapping circuit 511, a second function-based mapping circuit 512, and a third function-based mapping circuit 513) for the corresponding input components to produce the corresponding brightened output components, such as the nonlinear red output R'_out (and correspondingly G'_out, B'_out). In practice, this can be implemented, for example, as a lookup table after the function is computed. The shape of the adjustment function (F_adj) is (e.g., Figure 7 and Figure 8 The example illustrated depends on the illuminance measurement (Lev_illumdrown) of the illuminance meter 580. How the illuminance meter is used is a detail beyond the scope of this invention; for example, it could include a white light homogenizing dome before sending incident light to measuring electronics that captures the amount of incident photons and converts them into a voltage or current and a corresponding luminance (Lev_illumdrown), which is typically calibrated at the factory to a representative digital value. The illuminance meter can be located, for example, on the front of the display. The preprocessing circuitry 501 can typically implement, for example, image processing as described above, i.e., decoding and display optimization of the input HDR image or video, regardless of a specific (higher) level of ambient lighting (i.e., assuming a dark room or some typical reference viewing lighting, e.g., dim or dark lighting in a home movie-watching environment). Several systems will be simple and pragmatically refer to the display of the darkest pixel (in any environment) as zero, regardless of the actual amount of, for example, light leakage from the display, a small amount of reflected light from the front of the display, etc. For example, a more advanced method is taught in the applicant's WO2022 / 233612 (whose teachings are incorporated herein by reference in their entirety), which takes all these details into account to more accurately characterize the actual display output of black (and therefore more accurately characterize its appearance). This improved technique can work in both variations and in its various embodiments (i.e., the order of the processing circuitry, the selected color space used for calculation, etc.).

[0109] The preprocessing circuit 501 typically includes a luma processing circuit 502, which adapts the absolute luminance or relative percentage brightness of the input using some luma mapping function. This luma mapping function will typically depend on the characteristics of the display on which the image is to be displayed, such as the display's maximum luminance capability ML_D. The shape of the function may also depend on image attributes, such as its global attribute of average luma or average luminance, or more detailed attributes that can be transmitted as metadata (e.g., guiding the luma mapping function F_Lguid). A chroma processing circuit can perform chroma processing, ideally corresponding to luma processing, so that the colors do not appear oversaturated or undersaturated. A color transformation circuit 504 transforms the color representation, consisting of the mapped lumaY_LM and the processed chroma values ​​Cb_pr and Cr_pr, into the input format required by the control processor, i.e., PQ-based nonlinear R'G'B' components.

[0110] Figure 6Another possible embodiment of the control processor is shown, now as a total color mapping circuit 5101. This circuit can again perform some luma mapping in the luma mapping circuit 550. For example, some luma mapping can be performed using a luma mapping function received as an externally guiding luma mapping F_Lguid, which can be further optimized to another function shape (e.g., considering metadata MET, such as the maximum luminance ML_D of the connected display) before the application of this technique, or is already a fully optimized initial adjustment function F_adj for the connected display capabilities and the current viewing environment characterization (i.e., the currently stored or determined illuminance Lev_illumdrown). However, before actually performing luma processing, this function is multiplied by a correction function to produce a new bottom portion 555 of the final version of the adjustment function F_adj. The function can then be stored, for example, in a LUT, and subsequently applied to any input pixels scanned sequentially from the input image until, for example, a new scene appears, and the optimal luma mapping for the luma distribution of the new scene will require a new function shape, or until a new ambient lighting measurement requires a new bottom portion (and / or a new additive constant Adj_brightn), even when the static mapping curve is used for the entire movie to adapt the video display to the maximum luminance capability of the display. The chroma processing circuit 551 can similarly apply any of the chroma adaptation variants already mentioned above, which is either independent of the measured illuminance Lev_illumdrown or also depends on it. For example, preferably, chroma correction can take luminance dependence into account. For example, similar to RGB parallel processing, the same or similar correction functions can be used for the bottom portions of the Cb and Cr components. Finally, the color converter 552 will be configured to perform static color transformation again to the desired output format (R'_out, G'_out, B'_out) for driving the display (e.g., a PQ-based format for communication via an HDMI cable).

[0111] Figure 7 It shows in Figure 5The numerical representation of the processing in the example. Each of the three function-based mapping circuits (511, 512, and 513) applies a function where Ein is the normalized value of the function's input in the PQ domain, and Eout is the normalized value of the function's output. The initial adjustment function F_adj here is simply a linear compression function in the PQ domain that is psychologically visually homogenized (after all image objects in the scene have been redistributed to correspond to the capabilities of the connected displays in the preprocessor), which maps zero image pixel values ​​to output values ​​equal to the measured illuminance Lev_illumdrown (i.e., as the output on the vertical axis), and maps the maximum value to the maximum value of the image (i.e., retaining its value, here 1000 nit, represented by its equivalent PQ value 0.7518). F_std will again be a function that does not consider the specific ambient lighting of the viewing environment, i.e., it will be an identity transformation (up to 1000 nit). A simplified embodiment of the initial adjustment function F_adj in this case would be: [Formula 3] Where PQ represents the inverse perceptual quantizer EOTF, GAM is 1 nit in this example (which is relatively dark lighting), and the value of the display's maximum luminance ML_D is selected as 1000 nit, meaning the display can show pixels with a luminance of 1000 nits.

[0112] To obtain the final F_adj function for the various nonlinear red, green, and blue pixel color components to be applied to the pixels of the input image, this function in Equation 3 is multiplied by the correction function F_BCor for the darkest pixel values ​​(i.e., values ​​below GAM1 on the input axis).

[0113] Therefore, the final function is defined as: if ,So and if ,So You can also specify the shape of the function F_BCor so that it is normalized on the input axis, i.e., "1" corresponds to any GAM1 value selected.

[0114] The function (called h(x), where 0≤x≤1) will typically satisfy the following properties: h(0) = 0 The first derivative of h(1) = 0 h(1) = 1 And h(x) is usually strictly monotonically increasing.

[0115] (And generally, this is also advantageous for any polynomial of order n≥3 if its second, third, ... (n-1) derivatives are also zero.)

[0116] The inventors have identified the following function as a useful function for the correction function F_BCor (where the coefficients are the coefficients of powers, starting from the highest power n on the left and ending at the constant term on the right (i.e., the term multiplied by 1): That is to say, if we want to use a second-order polynomial (but it has been found that fifth-order or sixth-order polynomials are more preferred), then the second-order polynomial is: ; Other polynomials that can be considered are: Figure 9 The shapes of these polynomials are given (it is generally desirable to use polynomials of order 4 to 8). It can be seen that if fast convergence to the initial adjustment function F_adj is desired, then a higher-order polynomial (e.g., an 8th-order or higher) should be used.

[0117] Figure 8 Another example of the same procedure for obtaining a truly solid black in an image by multiplying by a correction function is given, but now for an illuminance of 10 nits (equivalent luminance) (i.e., a slightly brighter viewing environment). It now has a higher second threshold, GAM2.

[0118] The algorithmic components disclosed herein can be implemented in practice ( wholly or partially) as hardware (e.g., portions of an application-specific integrated circuit) or as software running on a dedicated digital signal processor or a general-purpose processor. At least some of the elements in the various embodiments can run on a fixed or configurable CPU, GPU, digital signal processor, FPGA, neural processing unit, application-specific integrated circuit, microcontroller, SoC, etc. Images can be temporarily or permanently stored in various memories, stored near one or more processors, or remotely accessible (e.g., via the Internet).

[0119] Those skilled in the art will understand from our specification which components may be optional modifications and can be combined with other components, and how the (optional) steps of the method correspond to the corresponding units of the apparatus, and vice versa. Some combinations will be taught by dividing the general teaching into partial teachings about one or more parts. The term "apparatus" in this application is used in its broadest sense, that is, a set of units that allows the achievement of a particular objective, and can therefore be, for example, an IC (a small circuit part), or a special electrical appliance (e.g., a device with a display), or a part of a networked system, etc. "Arrangement" is also intended to be used in its broadest sense, and therefore can include a single apparatus, a part of an apparatus, a collection of (partial) cooperative apparatuses, etc.

[0120] The representation of a computer program product should be understood as including any physical implementation of a set of commands that enables a general-purpose or special-purpose processor, after a series of loading steps (which may include intermediate conversion steps, such as translation into an intermediate language, and a final processor language), to input commands into the processor and execute any of the feature functions of the invention. Specifically, a computer program product can be implemented as data on a carrier (e.g., a disk), data residing in memory, or data transmitted via a (wired or wireless) network connection. In addition to program code, the feature data required by the program can also be embodied in the computer program product. Some technologies can be included in signals, which are typically control signals used to control one or more technical behaviors, such as those of a receiving device (e.g., a television). Some circuitry can be reconfigurable and temporarily configured for specific processing performed by software. Parts of the device can be specifically adapted to receive, parse, and / or understand innovative signals.

[0121] Some steps required for the operation of this method may already exist in the processor's functionality, rather than as described in the computer program product (e.g., data input and output steps).

[0122] It should be noted that the above embodiments are illustrative and not limiting of the invention. For the sake of brevity, we have not delved into all these options, as mapping the presented examples to other areas of the claims can be readily achieved by those skilled in the art. Other combinations of elements are possible besides those combined in the claims. In fact, any combination of elements can be implemented in a single dedicated element or separate elements.

[0123] Any reference numerals within parentheses in the claims are not intended to limit the claims. The word "comprising" does not exclude the presence of elements or aspects not listed in the claims. The words "a" or "an" preceding an element do not exclude the presence of a plurality of such elements, nor do they exclude the presence of other elements.

Claims

1. A control processor (510) configured to adjust the display brightness of at least a dark subset of an input image (IMG_in) on the display in response to a measurement of the illuminance (Lev_illumdrown) of the environment in which the display (520) is located, wherein, The adjustment includes establishing an initial adjustment function (F_adj) to shift one or more color components in the input color component triplets (R'_PQ, G'_PQ, B'_PQ) of the pixels of the input image by at least one correction value (Adj_brightn) to obtain the corresponding output color component triplets (R'_out, G'_out, B'_out). The initial adjustment function (F_adj) depends on the illumination. The adjustment includes: for input color component triplets (R'_PQ, G'_PQ, B'_PQ) below a threshold (GAM1),... The values ​​of one or more color components (R'_PQ, G'_PQ, B'_PQ) are used to multiply the initial adjustment function (F_adj) by the correction function (F_corr) to obtain a final version of the adjustment function for shifting the input color component triplet (R'_PQ, G'_PQ, B'_PQ), wherein the correction function (F_corr) satisfies: a first property of producing zero output for zero input; a second property of producing an output value equal to one for an input value equal to the threshold (GAM1); and a third property of having a first derivative of zero at the input value equal to the threshold (GAM1).

2. The control processor (510) according to claim 1, wherein, The correction function is strictly monotonically increasing.

3. The control processor (510) according to any one of the preceding claims, wherein, The correction function is applied to the input and output domains of the psychovisually homogenized color component values, for example, by pre-applying the inverse perceptual quantizer electro-optic transfer function to a linear version of one or more color components in the input color component triplet.

4. The control processor (510) according to any one of the preceding claims, wherein the threshold (GAM1) on the input axis is set to be equal to the illuminance (Lev_illumdrown).

5. The control processor (510) according to any one of the preceding claims, wherein, The correction function has multiple higher-order derivatives equal to zero at the input value equal to the threshold (GAM1).

6. The control processor (510) according to any one of the preceding claims, wherein, The correction function is equal to , where x is the input value between zero and a linear representation of the threshold (GAM1), which is normalized to a maximum value equal to 1, and y is the output value of the correction function for the input value.

7. The control processor (510) according to any one of claims 1 to 5, wherein, The correction function is equal to: , where x is the input value between zero and a linear representation of the threshold (GAM1), which is normalized to a maximum value equal to 1, and y is the output value of the correction function for the input value.

8. A display comprising a control processor according to any one of the preceding claims.

9. A method for adjusting the display brightness of at least a dark subset of an input image (IMG_in) on the display in response to a measurement of the illuminance (Lev_illumdrown) of the environment in which the display (520) is located, wherein, The adjustment includes establishing an initial adjustment function (F_adj), which is used to shift one or more color components in the input color component triplets (R'_PQ, G'_PQ, B'_PQ) of the pixels of the input image by at least one correction value (Adj_brightn) to obtain the corresponding output color component triplets (R'_out, G'_out, B'_out). The adjustment includes: for input color component triplets (R'_PQ, G'_PQ, B'_PQ) below a threshold (GAM1),... The values ​​of one or more color components in Q are used to multiply the initial adjustment function (F_adj) by the correction function (F_corr) to obtain a final version of the adjustment function for shifting the input color component triplet (R'_PQ, G'_PQ, B'_PQ), wherein the correction function (F_corr) satisfies: a first property of producing zero output for zero input; a second property of producing an output value equal to one for an input value equal to the threshold (GAM1); and a third property of having a first derivative of zero at the input value equal to the threshold (GAM1).

10. The method for adjusting the display brightness according to claim 9, wherein, The correction function is strictly monotonically increasing.

11. The method for adjusting the display brightness according to claim 9 or 10, wherein, The correction function is applied to the input and output domains of the psychovisually homogenized color component values, for example, by pre-applying the inverse perceptual quantizer electro-optic transfer function to a linear version of one or more color components in the input color component triplet.

12. The method for adjusting the display brightness according to any one of the preceding method claims, wherein, The threshold (GAM1) on the input axis is set to be equal to the illuminance (Lev_illumdrown).

13. A computer program product comprising software code for instructing a processor to perform the method according to any one of claims 9 to 12.

Citation Information

Patent Citations

  • Image processing circuit and method thereof

    US20150213781A1

  • Optimizing high dynamic range images for particular displays

    WO2017108906A1

  • Encoding and decoding HDR videos

    WO2017157977A1

  • Display-optimized HDR video contrast adaptation

    WO2022233612A1