Mixing secondary graphic elements in HDR images

The method addresses the challenge of blending secondary graphics in HDR images by determining a secondary graphics luma range and mapping luminance values, ensuring consistent rendering across displays with varying maximum luminances.

JP7752792B2Active Publication Date: 2025-10-10KONINKLIJKE PHILIPS NV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024569131
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-24
Filing Date
2023-05-09
Publication Date
2025-10-10
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

Existing systems struggle with effectively blending secondary graphics elements in high dynamic range (HDR) images, as the luminance values are undefined or ill-defined, leading to inconsistent and potentially flickering graphics across different displays with varying maximum luminances.

Method used

A method for determining a secondary graphics luma range by analyzing an HDR image signal, identifying a primary graphics luma range, and luminance mapping the secondary graphics elements within this range to ensure consistent blending across different displays.

Benefits of technology

Ensures stable and consistent rendering of secondary graphics elements across displays with varying maximum luminances, preventing flickering and maintaining the intended appearance of graphics in HDR images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007752792000001
    Figure 0007752792000001
  • Figure 0007752792000002
    Figure 0007752792000002
  • Figure 0007752792000003
    Figure 0007752792000003
Patent Text Reader

Abstract

To more appropriately adjust the pixel luminance of the primary and secondary graphics elements inserted at different points in time of the HDR video, and to realize a more predictable graphics insertion, the inventors propose a method (or apparatus) for determining the second luma of the pixels of the secondary graphics image element 216 mixed with at least one high-dynamic-range input image 206 in a circuit for processing digital images. The method includes receiving a high-dynamic-range image signal S_im including at least one high-dynamic-range input image, and the method further includes analyzing the high-dynamic-range image signal to determine a range R_gra of the primary graphics luma of the primary graphics elements of the at least one high-dynamic-range input image, where the range of the primary graphics luma is a partial range of the luminance range of the high-dynamic-range input image, and determining the range of the primary graphics luma includes determining a lower luma Y_low and an upper luma Y_high that specify the endpoints of the range R_gra of the primary graphics luma, luminance mapping the secondary graphics element using at least the brightest subset of its pixel luma included in the range of the graphics luma, and mixing the secondary graphics image element with the high-dynamic-range input image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for constructing high dynamic range (HDR) images, a novel image encoding technique compared to standard low dynamic range images, particularly those containing graphics elements. Specifically, the present invention may address more advanced HDR applications involving luminance or tone mapping, which may be specified differently for each video scene image. This allows for the computation of secondary color gradings of different luminance dynamic ranges ending at different maximum luminances from an input image having a primary grading. (Grading typically involves redistributing the normalized luminance of pixels of various objects in the image from their initial relative values ​​in the input image to different relative values ​​in the output image, where the relative redistribution also affects the absolute luminance values ​​of the pixels relative to their normalized luminance.) Specifically, the present invention may be useful when a previously created HDR image requires the blending of secondary graphics elements by some video processing device. (The primary graphics elements are already present in the image before blending at least one secondary graphics element, and must be distinguished from the secondary graphics elements.) [Background technology]

[0002] Until the first research around 2010 (and before the first HDR-decoded televisions hit the market in 2015), all video, at least for video, was created according to a common low-dynamic-range (LDR) or standard-dynamic-range (SDR) encoding framework. This had several properties. First, it created a single video suitable for all displays. This system was a relative system in which the maximum (100%) signal encoded with the maximum luma code (255 in 8-bit YCbCr encoding) corresponding to the maximum nonlinear RGB values ​​R' = G' = B' = 255 was white. Nothing brighter than white exists, and all typical reflected colors can be represented as colors darker than such a brightest white (for example, a piece of paper will at most reflect all incident light, or absorb some red wavelengths and reflect blue and green back to the eye, resulting in a local cyan color somewhat darker than the white of the paper). Each display will display (as a "drive request") this whitest white as the brightest color it is technically constructed to render, e.g., 80 nits (SI quantity Cd / m) for a computer monitor. 2 For a TL backlit LCD display, this was 200 nits. Viewers' eyes quickly compensate for brightness differences, so unlike people sitting side by side in a store, all viewers at home saw roughly the same image (despite display differences).

[0003] There has been a demand to improve the perceived appearance of images by making colors not just colors that can be printed or painted on paper, but pixels that actually glow much brighter than "paper white," also known as "diffuse white." In practice, this means, for example, that whereas in the past, photography took place in studios where everything important was brightly lit by overhead lights, it is now possible to take a good-looking photograph simply by shooting in strong backlight. Cameras continue to evolve, allowing them to simultaneously capture enough scene luminances for many scenarios and to further adjust or compensate for brightness in a computer. Displays, including consumer television displays, also continue to improve.

[0004] In some systems, such as the BBC's Hybrid Log-Gamma (HLG), this is achieved by defining values ​​in the coded image that exceed white (where white is given a reference level of "1", or 100%) (for example, up to 10 times white), thereby defining the display of pixels that shine 10 times brighter.

[0005] To date, most systems have also moved to a paradigm where video creators can define absolute nit values ​​for their images (i.e., not 2x or 10x relative to an undefined white level that translates to variable actual nit output on each endpoint display) based on a selected target display dynamic range capability. The target display is the video creator's virtual (intended) display (e.g., a 4000-nit (ML_C) target display to define a 4000-nit video), while the actual consumer endpoint display may have a lower display maximum luminance (ML_D), e.g., 750 nit. Even in such cases, the end display must be equipped with luminance remapping hardware or software, typically realized as luma mapping. Luminance remapping hardware or software somehow adapts pixel luminances in HDR input images whose luminance dynamic range (specifically, maximum luminance) is too high to be faithfully displayed to values ​​in the dynamic range of the end display. The simplest mapping is to clip all luminances above 750 nits to 750 nits, but this is the worst way to handle dynamic range mapping. This is because, for example, the beautiful structure of a sunset, including sunlit clouds in the 1000-2000 nit range, in a 4000 nit image is removed and displayed as a uniform white 750 patch. A better luma mapping would move the 1000-2000 nit subrange in the HDR input image into the end display dynamic range, e.g., 650-740, using an appropriately determined function. The function may be determined automatically within the receiving device, e.g., a TV or STB, or it may be determined by the video creator to best suit their artistic movie or program and communicated with the video signal, e.g., as metadata in a standardized format. Luma refers to any encoding of luminance, e.g., at 10 bits, using a function that assigns luma codes from 0 to 1023 to video luminances, e.g., 0.001 to 4000 nits, via the so-called electro-optical transfer function (EOTF).

[0006] The simplest system would be to simply transmit the HDR, say, maximum 4000 nit image itself (using a well-defined EOTF) (i.e., provide the image to the receiver without indicating how to downgrade it via luminance mapping for displays with lower maximum luminance capabilities, for example). This is what the HDR10 standard does. More advanced systems, such as HDR10+, might also communicate a function for downmapping the 4000 nit image to a lower dynamic range, such as 750 nit. These systems facilitate this by defining a mapping function between two different maximum brightness version images, or reference gradings, of the same scene image, and then using an algorithm to calculate modified versions of that reference brightness or luma mapping function to calculate other brightness mapping functions, such as endpoint functions for other display maximum brightnesses. (For example, the reference brightness mapping function might specify how to modify the normalized brightness distribution in a 2000-nit input image (e.g., a master HDR-graded video created by an author) to obtain a 100-nit reference video, and a display adaptation algorithm might calculate a final brightness mapping function based on the reference brightness mapping function so that a 700-nit TV can map a brightness between 0 and 2000 nits to a display brightness in the range of 0 to 700 nits.) For example, if we agree to define an SDR image to always have a 100 nit maximum pixel brightness when newly interpreted as an absolute rather than relative nit image, video creators can define and communicate together a function that specifies how to map a first reference image grading, 0.001 (or 0) to 4000 nit brightness, to the corresponding desired SDR 0 to 100 nit brightness (secondary reference grading). This is called display tuning or adaptation.If we define a function that boosts the darkest 20% of colors in the plot, normalized to 1.0, by, say, a factor of 3 for both the 4000-nit ML_C input image (horizontal axis) and the 100-nit ML_C secondary grading / reference image, i.e., when going from 4000 nit down to 100 nit, if a particular end-user's TV needs to go down to 750 nit, then the boost needed may only be, say, 2x. (This depends on the definition of luma used, EOTF, because, as mentioned above, in color processing ICs / pipelines, luma mapping is usually realized as a luma mapping (e.g., using a psychovisually uniform EOTF) because realizing it as a luma mapping allows the impact of changes in luma along the range to be defined more visually uniform, i.e., more relevant and visually impactful to humans.)

[0007] A third class of more advanced HDR encoders takes this principle to the next level by restating it in another way: defining regrading requirements based on two reference-graded images. By restricting ourselves to using mostly lossless functions, an LDR image, which can be calculated at the sender by downmapping the luminance or luma of, say, a 4000-nit HDR image to an SDR image, can actually be transmitted as a proxy for the actual master HDR image created by the video creator (e.g., by a Hollywood studio for Blu-ray or OTT distribution, or by a sports broadcaster). The receiving device can then apply the inverted functions to reconstruct an exact reconstruction of the master HDR image. We refer to systems that communicate the (as-created) HDR image itself as "modal HDR," and systems that communicate LDR images as "modal LDR coding frameworks."

[0008] A typical example is shown in Figure 1 (which summarizes the principles previously patented by the applicant, e.g., in WO2015 / 180854). Note that Figure 1 includes the decoding function itself and the subsequent display adaptation as blocks, and that these two technologies should not be confused.

[0009] 1 illustrates a schematic diagram of an exemplary video coding and communication system, as well as a processing (display) system. On the production side, a typical encoder (100) embodiment is shown. Those skilled in the art will appreciate that the pixel-by-pixel processing pipeline for luma processing (i.e., all pixels of the input HDR image Im_HDR (typically one of multiple master HDR images for a video created by a video creator; details of its production, such as camera capture and shading or offline color grading, are omitted as they are understandable to those skilled in the art and do not enhance this description) are processed sequentially) is shown first, followed by the video processing circuitry operating on the entire image.

[0010] Assume we start with an HDR pixel luminance L_HDR (although some systems may already start with luma), which is sent through a selected HDR inverse EOTF in luma conversion circuit 101 to obtain a corresponding HDR luma Y_HDR. For example, a perceptual quantizer EOTF may be used. This input HDR luma is luma mapped by luma mapper 102 to obtain a corresponding SDR luma Y_SDR. In this unit, the creating device applies appropriate color modifications, including appropriate luminance re-grading for any particular scene image in the output image, corresponding to the color details of the input image (e.g., maintaining an appropriate average brightness for a cave image). That is, an input connection UI exists to obtain an appropriately determined shape for the luminance mapping function (LMF). For example, if creating an output image with a low maximum luminance for a cave image, a normalized version of the luma mapping function may need to boost the darkest lumas of the input by a factor of 2.0, determined by, for example, a human color grader or an automaton. This means that the luma mapping function is concave, or so-called r-shaped (as shown within the box representing unit 102).

[0011] Broadly speaking, there are two classes. An offline system may use color grading software to have a human color grader determine the optimal LMF based on his or her artistic preferences. Without limitation, assume that the LMF is defined as a look-up table (LUT) defined using the coordinates of a small number of node points (e.g., (x1, y1) for the first node point). The human grader may set the slope of the first line segment, i.e., the location of the first node point, based on, for example, dark content in the image and the desire to maintain sufficient visibility when displayed on a display with a low dynamic range (e.g., a 100-nit ML_C image for a 100-nit ML_D LDR display).

[0012] The second class uses automata to obtain secondary grading from primary gradings with different maximum luminances. These automata analyze images and propose optimal LMFs (e.g., neural networks trained with image aspects from a set of training aspects generate normalized coefficients of some parametric function at the output layer, or any deterministic algorithm). One particularly interesting automaton, the so-called inverse tone mapping (ITM), analyzes an input LDR image rather than a master HDR image and creates a pseudo-HDR version of this LDR image. This is extremely useful because most video is displayed as LDR and may be generated as SDR now or in the near future (at least, for example, some cameras in a multi-camera production may output SDR (e.g., a drone capturing side footage of a sports game), and this side footage may need to be converted to the HDR format of the main program). By appropriately combining the capabilities of a modal LDR coding system with an ITM system, the applicant and its partners were able to define a system capable of double inversion. That is, the upgrade function for the pseudo-HDR image generated from the original LDR input is essentially the inverse of the LMF used when coding the LDR communication proxy. In this way, a system can be created that not only communicates the original LDR image, but also conveys information for creating a suitable HDR image for it (either automatically, depending on the desires of the content-creating customer, or in other versions with human input (e.g., adjusting automatic settings)). The automaton can use any kind of rule (e.g., it can determine the location of a light source in an image, which requires a certain relative brightness compared to the average brightness of the image), but the exact details are not relevant to the present description. What is important is that any system employing or interfacing with an innovative embodiment of the present invention can generate some kind of luminance (or luma) mapping function LMF. The luma mapping function determined by the automaton can likewise be input via a connected UI and applied to the luma mapper 102.Note that in the description of the simplest embodiment, there is only one (downgraded) luma mapper 102. This is not necessarily a limitation. Because both the EOTF and luma mapping typically map a normalized input domain [0,1] to a normalized output domain [0,1], there may be one or more intermediate normalization mappings that (effectively) map 0 input to 0 output and 1 to 1. In such cases, the former intermediate luma mapping acts as a base mapping, and the (second) luma mapper 102 acts as a correction mapping based on the first mapping.

[0013] The encoder obtains (via a mapping function) a set of LDR image lumas Y_SDR that correspond to the HDR image lumas Y_HDR. For example, the darkest pixels in a scene may be defined to appear as substantially the same luminance on both the HDR and SDR displays, while brighter HDR luminances may be pushed into the upper range of the SDR image. This is illustrated by the convex shape of the LMF function shown in luma mapper 102, which decreases in slope (or reconverges toward the diagonal of the normalized axis system). Those skilled in the art will readily understand normalization and simply divide the luma code by power(2; number_of_bits). Optionally, normalized luminance may be defined by dividing the pixel luminance by the maximum value ML_C (e.g., 4000 nits) of the associated target display. That is, the brightest possible pixel of the target display, fully utilizing the capabilities of the associated target display by driving it to the target display's maximum luminance, which is the upper limit of its luminance dynamic range, is represented as a normalized luminance of 1.0 (which serves as the input brightness value displayed on the horizontal axis of a graph such as unit 102). Those skilled in the art will understand how to express different normalizations to different maximum values ​​and how to specify a function that maps normalized input values ​​to normalized luma (for any bit length and any chosen EOTF) or any normalized luminance (for any associated maximum luminance) on the vertical axis.

[0014] So, we can picture in our mind the example of an indoor-outdoor scene. In the real world, outdoor luminance is typically 100 times higher than indoor pixels, so a traditional LDR image will show indoor objects nicely bright and colorful, but anything outside the window will be heavily clipped to a uniform white (i.e., invisible). Now, if we communicate HDR video using a lossless proxy image, we will ensure that the bright exterior areas visible through the window are brightened (and possibly desaturated), but in a controlled way that ensures there is enough information for HDR reconstruction. This benefits both outputs, because a system that only wants to use the LDR image as is will display a good rendering of the outdoor scene, to the extent that the limited LDR dynamic range allows.

[0015] Thus, the set of Y_SDR pixel intensities (together with chromaticity, details of which are unnecessary for this discussion) forms a "conventional LDR image." That is, subsequent circuitry does not need to constantly consider whether this LDR image was intelligently generated or simply acquired directly from the camera, as in conventional LDR systems. Therefore, the video compressor 103 applies algorithms such as MPEG HEVC or VVC compression. This is a set of data reduction techniques, including, among other things, the discrete cosine transform, which converts blocks of, say, 8x8 pixels into a limited set of spatial frequencies that require less information to represent. The amount of information required can be adjusted by determining the quantization factor, which determines the number of DCT frequencies that are retained and how accurately they are represented. The drawback is that the compressed LDR image (Im_C) is not as accurate as the input SDR image (Im_SDR), particularly resulting in block artifacts. Depending on the broadcaster's choice, the situation may be so severe that, for example, some blocks of the sky are represented only by their average luminance and appear as uniform rectangles. Normally, this is not a problem. This is because the compressor determines all settings (including quantization coefficients) so that quantization errors are barely visible to the human visual system, or at least not a problem.

[0016] The formatter 104 performs any signal formatting required for the communication channel (which may be different for example stored and communicated on a Blu-ray disc than for example DVB-T broadcast). In general, all variations have the property that the compressed video images Im_C are packed into an output image signal S_im using a luminance mapping function LMF (which may or may not vary from image to image).

[0017] Deformatter 151 performs the inverse process of formatting, obtaining a compressed LDR image and function LMF to reconstruct an HDR image, or other useful dynamic range mapping operations may be performed in subsequent circuitry (e.g., optimization for the specific connected display). Decompressor 152 removes, for example, VVC or VP9 compression, obtaining a sequence of approximate LDR luma Ya_SDR, which is sent to the HDR image reconstruction pipeline. To perform the essentially inverse process of coding the communicated HDR video as an LDR proxy video, upgrade luma mapper 153 converts the SDR luma to reconstructed HDR luma YR_HDR (which uses an inverse luma mapping function ILMF, which is (effectively) the inverse of LMF). One illustration shows two potential receiving devices (150), which may exist as dual functions in one physical device (e.g., the end user can choose which parallel processing to apply), or some devices may only have one of the parallel processing (e.g., some set-top boxes may only reconstruct the master HDR image and store it in memory 155, such as a hard disk).

[0018] If a display panel (e.g., a 750-nit ML_D end-user display 190) is connected to an embodiment of the receiver, the receiver may have a display adaptation circuit 180 that calculates, for example, a 750-nit output image instead of a 4000-nit reconstructed image (this is shown dotted to indicate that, while both techniques are often used in combination, this is an optional component unrelated to the teachings of the present invention). Without going into detail about the many variations in which display adaptation can be achieved, there will typically be a function determination circuit 157 that proposes an adapted luma mapping function F_ADAP based on the inverse shape of the LMF (usually close to the diagonal when the difference between the maximum luminance of the source image and the maximum luminance of the destination image is smaller than the maximum luminance difference between two reference gradings (e.g., a 4000-nit master HDR image and the corresponding 100-nit LDR image). This function is loaded into the display adaptive luminance mapper 156, which calculates a less-boosted HDR luma L_MDR, typically with a smaller dynamic range (ending at ML_D=750 nit instead of ML_C=4000 nit). Regrading refers to the luminance (or luma) mapping from a first image in a first dynamic range to a second image in a second dynamic range, where at least some pixels are given a different absolute or relative luminance (e.g., compressing the brightest luma to a smaller subrange to reduce the output dynamic range). Such a graded image may also be referred to as a graded image. If only reconstruction is required, the EOTF conversion circuit 154 is sufficient, and a reconstructed (or reconstructed) HDR image Im_R_HDR is generated, containing pixel colors including the reconstructed HDR luminance LR_HDR.

[0019] When video content consists solely of natural images (e.g., captured by a camera), high dynamic range techniques are already complex, considering the various ways of calculating re-graded pixel luminance relative to the input luminance. Additional complexity is added by the details of various coding standards, which may be intermixed; for example, picture-in-picture requires mixing HDR10 video data with HLG capture.

[0020] Additionally, content creators typically want to add graphics to images. The term graphics refers to non-natural image elements / areas, i.e., those that are typically simpler in nature and typically have a more limited, discrete set of colors (e.g., lacking photon noise). Another way to characterize graphics is that they do not form part of a similarly lit image, but other considerations, such as readability or attention-grabbing, may be important. Graphics are typically computer-generated and typically refer to graphics that are not natural. For example, CG furniture that is visually indistinguishable from real furniture in a scene captured by a camera. For example, there may be a bright company logo, or graphic elements such as plots, weather maps, or information banners.

[0021] Graphics began in the old LDR era, mimicking the way graphics elements were created on paper, typically with a limited set of tools such as pens or crayons. Indeed, early systems such as Teletext initially defined three primary colors (red, green, and blue), two secondary colors (yellow, cyan, and magenta), and black and white. Graphics, such as an information page about airport flight schedules, had to be generated from pixel primitives (e.g., pixel blocks) with one of these eight colors. However, "red" (though it must appear reddish to a human viewer) is not a uniquely defined color, and in particular does not have a predefined, unique intensity. More advanced color palettes can include, for example, 256 predefined colors. For example, graphical user interfaces such as the Unix X-windows system define a large number of selectable colors, such as "honeydew," "gainsboro," or "lemon chiffon" in X11. These colors had to be defined as representative codes that would produce an approximately specific color appearance when displayed, defined in the LDR color gamut as percentages of standard red, green, and blue mixed together. For example, honeydew is composed of 94% red, 100% green, and 94% blue (or hex code #F0FFF0), and appears as a bright greenish-white color.

[0022] The problem is that while these colors are stably defined in the limited (and universal) LDR gamut, they are undefined, or at least ill-defined, in HDR images. There are several reasons for this. First, in LDR, there is only one white, and all graphics colors can be defined relative to that white. Indeed, just as when drawing with a marker of the appropriate color on the white of a display that acts as a white canvas, code that mimics the absorption of a particular hue defines the creation of a particular chromaticity (or color nuance, such as a desaturated chartreuse). However, in HDR, there is no unique white. This is evident from the real world: when we look at a white-painted wall indoors, we are visually impressed by the overall appearance of white, although there may be a slight gray in the shadows. However, when we look at a white garage door outdoors with sunlight shining through a window, this also appears to be a type of "white," but this white is a different, much brighter white. A high-quality HDR processing system (where processing typically includes aspects such as optimal creation, encoding, and processing for proper display) may desire not only to extend the LDR appearance with several fine HDR aspects, but also to render various kinds of white on the display (although typically at lower luminances than in the real world). Technically, such considerations have led to the useful definition of a framework within which various kinds of HDR images can be defined with different coded maximum luminances ML_C, such as 1000 nits or 5000 nits.

[0023] Further complicating the issue is the existence of different types of HDR images (with different ML_Cs; even relative systems generally have the same issues), which necessitates luma mapping, mapping luminance along the input image dynamic range to luminance along the output dynamic range (e.g., of an end consumer display). Luma mapping is typically defined and calculated as a corresponding LUMA mapping function, which maps a luma code that uniquely represents luminance, e.g., via a perceptual quantizer EOTF. For example, suppose an input image luminance ranging from 0 to 4000 needs to be mapped to an output luminance ranging from 0 to 100 using the function F_H2S_L (i.e., L_out = F_H2S_L(L_in)). Because the perceptual quantizer EOTF (EOTF_PQ) can encode luminance up to 10,000 nits, both the input luminance and the output luminance can be expressed as PQ-defined luma: Y_out = OETF_PQ(L_out) and L_in = OETF_PQ(L_in). Thus, the luma mapping F_H2S_Y is: Y_out=F_H2S_Y(Y_in)=F_H2S_Y(OETF_PQ(L_in)). Alternatively, the luma mapping function and the luma mapping function are related by: L_out=EOTF_PQ(Y_out)=EOTF_PQ[F_H2S_Y(OETF_PQ(L_in))], and therefore F_H2S_L=EOTF_PQ(o)F_H2S_Y(o)OETF_PQ, where (o) denotes function composition (usually denoted by a small circle symbol). It is also possible to define a mapping to luma specified by another EOTF (e.g., a mapping from input PQ luma to output Rec. 709 SDR luma).

[0024] A possible graphics insertion pipeline is described with reference to FIG.

[0025] Figure 2 shows an HDR image communication and processing pipeline that illustrates a similar multi-video communication system. This is a classic sports program (horse racing) for a broadcast service. One or more cameras 201 capture the sporting event, and the captured images are mixed in a production booth 202. In the production booth 202, feeds from the various cameras are mixed to create a total HDR video that is graded (or shaded) as desired. The method for specifying pixel luma for successive video images is beyond the scope of this description. The broadcaster creates an original blended HDR image 205 composed of native video content 206 (camera feeds after appropriate shading) and the broadcaster's own graphics (an example of primary graphics). This graphics can vary depending on the application, but in this example, it is a list of the names of the two leading horses in the race. This original graphics, pre-mixed and communicated within the video image, is called a primary graphics element 207. It is then distributed, for example, by satellite dish 203, to a local distributor's equipment 210. This could be, for example, a Dutch broadcaster or redistributor distributing to consumers via terrestrial broadcast or cable 219. The local distributor might add their own secondary graphics, such as the broadcaster's logo ("NL6") within a sun. For example, assume that the pixel colors of the primary graphics are all white, and the pixel colors of the secondary graphics are white-dependent, such as a yellow sun with 90% white intensity and red text color within it at 60% brightness. The broadcaster might also mix in other types of secondary graphics, such as a teaser 217 for a later sports program. This further mixed video (215) (secondary mixed image / video) is further distributed to end users via television communications cable (CATV) 219. The end customer might have a set-top box (or computer, etc.) 220 that can mix tertiary graphics, etc.In this example, closed caption information is communicated along with the video signal (typically purely textual information that must be rendered within the video pixels that make up the generated video element), which is rendered by the set-top box as tertiary graphics 226 and mixed back into the video image to generate tertiary video 225, which is communicated to a consumer display 230, such as via an HDMI® cable 229. In practice, the end display may insert its own tailored graphics instead of (or in addition to) the set-top box, which may not be present in some processing pipelines.

[0026] In this text (when referring to specific graphics at a higher conceptual level), all additional graphics (in this example, secondary and tertiary graphics) will be referred to as "secondary graphics" to distinguish them from their preceding graphics context, primary graphics (often the first graphics inserted into a video by the original creator of the video). Typically, the image itself (i.e., the matrix of pixel colors, e.g., YCbCr) Im_HDR is communicated along with metadata (MET) in the HDR image or video signal S_im. This metadata can code several things in various HDR video coding variations, such as the coded maximum luminance of the HDR image (e.g., ML_C_HDR=1000 nits) and, in many cases, a luma mapping function for calculating the secondary colors of the regraded image.

[0027] It is undesirable for the various graphics to be untuned across the entire luminance dynamic range, or even worse, for their luminance to vary over time (e.g., compared to each other), but rather to have the same average luminance, as illustrated, for example, by the 950 nit ML_D luminance dynamic range of the end consumer display 230. This can improve the technical capabilities of, for example, a television display (or other device receiving the primary video graphics mix) for rendering secondary graphics. While this may not be very difficult or important to do if a simple system exists that pre-calculates a display-adapted video image with optimal pixel luminance and one simply adds primary graphics to that pre-established range, the situation can become more complicated, especially when adding different types of graphics by different devices at different points in the HDR video processing pipeline.

[0028] US20180018932 teaches some blending techniques for primary graphics elements (not blending secondary graphics if primary graphics are already blended), and does not establish the proper location of the graphics subrange of the primary graphics within the master HDR video (the graphics maximum value simply maps to the video maximum, i.e., the entire range).

[0029] The video-first mode described in US '932 - Figure 2A, which is the most complex variation to understand, is summarized in this application using Figure 8. If you want to display stable graphics with an image output adapted to the end user's display (e.g., if you want to display an image created for the end user's display's maximum luminance ML_endUSR_disp750), you can achieve this by assigning the same value to all pixel luminances of, for example, solid white text (TXT) across multiple consecutive images (e.g., across several scenes, or, in the case of user interface graphics, as long as the graphics are presented). The appropriate final value for the text luminance is assumed to be the value at which a video cloud object is projected within the display dynamic range. The dynamic range of graphics (especially the graphics' maximum luminance ML_gra) is typically not the same as the input video with which the graphics are to be mixed (the video's maximum luminance ML_vid). It is also typically different from the maximum luminance of the end user's display, since each user may have purchased a different display. In fact, associating graphics with some dynamic range is already a somewhat advanced HDR technique. This is because graphics in the LDR era typically had relative 3x8-bit codes (e.g., 255 / 255 / 255). The problem is that the dynamic metadata, which conveys the different luminance mapping functions applied by the end user's display for successive image shots, does not guarantee what pre-mixed text will do. When mixing original text (TXT_OR) with video at its original luminance, e.g., 390 nits, the display may increase its luminance with some mapping functions and decrease it with others, resulting in flicker. However, this can be compensated for in advance if we know exactly the dynamic mapping function (e.g., the first dynamic mapping function F_dyn1) that the display will apply to the video, regardless of whether the video contains graphics.In this case, the graphics can be pre-mapped (using a first pre-compensation function, F_precomp1) to the exact luminance required for the HDR image sent to the display, for example, from the set-top box, and then mapped to the desired final luminance of the displayed image (as bright as the clouds) by F_dyn1. In the next scene, the piece of paper can be up-mapped to the same final luminance for the TXT by a second dynamic mapping function, F_dyn2, in which case the second pre-compensation function, F_precomp2, is used. Typically, the function applied by the display depends on a maximum luminance, ML_endUSR_disp, and a dynamic, time-varying reference luminance mapping function that specifies how a 4000 nit video luminance should be mapped to, for example, a 100 nit reference image. A pre-fixed algorithm transforms this reference luminance mapping function into the final display-adaptive luminance mapping function. This function is used by the TV for subsequent images until a new, dynamically changing function is entered as metadata. However, if the set-top box is codec compliant, i.e., it knows the reference luminance mapping function and knows the algorithm to convert to the final display-adaptive luminance mapping function for any value of ML_endUSR_disp, and polls ML_endUSR_disp from the connected display, it can perform any such pre-correction.

[0030] The prior art provides several other examples of blending primary graphics, but these are simpler.

[0031] The graphics-first mode shown in US '932 - Figure 2B simply maps the graphics maximum to the display maximum. Then, video can be down-mapped to the range already present in the set-top box. Graphics mixing is relatively simple if everything is prepared within the STB, and an optimal display-adaptive image requiring no further processing is transmitted to the display via HDMI. The STB has everything at hand, and no variable mapping is performed later. However, these two options are not always available. For example, some codec technologies are expensive, and STBs are sold with very little profit, so STBs may not allow such codecs. In any case, displays such as consumer televisions must perform dynamic HDR luminance mapping. Pre-correction is also not always possible, for example, when a television refuses to communicate its ML_endUSR_disp value to the STB. While these methods may be useful for consumer electronics products such as an STB and a television, or a computer and a monitor, there may be many more places in the video communication pipeline from the "camera" to the end-use display (e.g., in a shopping mall) where graphics need to be inserted.

[0032] If the TV can perform graphics blending (as shown in US '932 - Figure 2D), things can be relatively simple. The TV can apply its normal display adaptive brightness mapping function to the video and then place the graphics at the desired brightness, e.g., the output brightness where the clouds end. This may be fine for pass-through DVB subtitles with standardized coding, but in that case the STB needs a mechanism to communicate the graphics of its own UI graphics elements over the HDMI interface.

[0033] US2020193935 statically switches the display mapping when mixing graphics, so that the graphics are always placed at the same (shifted) luminance position, but this also has advantages and disadvantages.

[0034] In the current field of high dynamic range video / television, it is clear that there is still a need for superior graphics processing technology. Summary of the Invention

[0035] The challenges present in simple approaches to graphics blending / insertion are addressed by a method for determining a second luma of pixels of a secondary graphics image element (216) to be blended with at least one high dynamic range input image (206) in a circuit for processing digital images, the method comprising: receiving a high dynamic range image signal (S_im) comprising at least one high dynamic range input image, the method further comprising: analyzing the high dynamic range image signal to determine a primary graphics luma range (R_gra) of at least one primary graphics element of the high dynamic range input image, the primary graphics luma range being a subrange of the luminance range of the high dynamic range input image, determining the primary graphics luma range including determining a lower luma limit (Y_low) and an upper luma limit (Y_high) that specify endpoints of the primary graphics luma range (R_gra); luminance mapping the secondary graphics element with at least the brightest subset of pixel lumas of the secondary graphics element that fall within a range of graphics lumas; and blending the secondary graphics image element with the high dynamic range input image.

[0036] Luminance mapping to graphics subranges within an input (video + primary graphics mix) HDR image, or within an output variation image that can be derived by mapping luma using some luma mapping function, may in principle be performed when creating (i.e., defining) the pixels of the graphics. However, in general, even if the graphics are not read from storage where they are stored in a predefined state, but rather generated by some method or apparatus embodiment, the graphics may be generated in some different format (e.g., an LDR format, or a format with a maximum luminance of 1000 nits), but the luma (or any code, e.g., a language code such as Chartreuse) of the pixels that make up the secondary graphics will be mapped to an appropriate position within (or slightly outside, typically above the lower / darker end of) the established graphics range, where it becomes the corresponding secondary luma of the various graphics pixels.

[0037] Blending often simply replaces the original video pixels with graphics pixels (but with adjusted color, especially adjusted luma). This method also works for more advanced blending, especially when more graphics are kept and less graphics are blended. For example, a linear weighting called blending may be used, where a certain percentage of a determined graphics luminance (Y_gra) is mixed with a complementary percentage of the video. Y_out=alpha*Y_gra+(1-alpha)*Y_HDR [Formula 1]

[0038] Alpha is typically higher than 0.5 (or 50%), so that the pixel essentially contains mostly graphics, and the graphics are well visible. In fact, ideally / preferably, such blending happens on the derived luminance itself, rather than on luma (especially for highly non-linear luma definitions like PQ). L_out=alpha*L_gra+(1-alpha)*L_HDR [Formula 2]

[0039] In embodiments in which a proxy SDR image is transmitted for an HDR image, such HDR video luminance L_HDR may be obtained by reconverting the SDR luma to an HDR reconstruction luminance, for example, using the maximum luminance ML_C transmitted with the metadata of the transmitted HDR representation. Optionally, display adaptation may also be used during the acquisition to downgrade this luminance to a corresponding luminance within the dynamic range of the display. Indeed, from FIG. 3 , it can be seen that the advantage of at least such an approach to constructing a luma mapping function (or a corresponding luminance mapping function) is that the curve maps the graphics range in a stable manner, allowing for mapping / blending of graphics in both the input and output ranges. More generally, this technique can take advantage of the fact that the content creator likely determined the optimal position of graphics, or at least approved the reasonable position of graphics, relative to the primary graphics. Therefore, secondary graphics may also contain the appropriate luminance, provided they are properly adjusted relative to the primary graphics.

[0040] When specific information about such primary graphics is not available, for example because the video creator does not want to go to the effort of coding and communicating it, a method or apparatus of the present invention can analyze the context of the input HDR video signal S_im and, according to various embodiments, determine an appropriate range R_gra for locating the luma (or luminance) of at least the majority of pixel colors within the secondary graphics elements.

[0041] Exactly how this is done depends on the nature of the graphics element.

[0042] If you have secondary graphics (e.g., subtitles) with only one or a few colors (e.g., the original colors before brightness mapping for adjustment), you can render the secondary graphics using any luma within the graphics range. In more complex images, the darkness of blacks can be limited (e.g., more than 10% of the brightness of white), or some dark colors can be set outside or below the range of the primary graphics (although the effect of stable graphics display is largely maintained when the luma of light colors is taken into account). For example, a slightly varying black perimeter rectangle at the top of an HDR video may not be too bothersome if the subtitle colors are stabilized relative to the brightness of one or more subsequent scene video objects. While you may want to avoid large fluctuations in the brightness of relatively large dark colors that are clearly below average, if this is a problem, you can address it separately, for example, by brightening the dark colors to approach the lower end of the appropriate graphics range (i.e., slightly above or below minimum brightness). It is desirable to keep the brightest colors of secondary graphics within the graphics range R_gra, for example, the lighter subrange being defined (e.g., by a television receiver's construction engineer) as all colors brighter than 50% of the luminance of the secondary graphics element (this percentage may depend on the type of graphics being added; for example, menu items may not require as much precision as some logos). Often, the lighter subrange will contain enough colors to represent (geometrically) most of the secondary graphics element, so that, for example, its shape can be recognized.

[0043] It is not necessary to completely know the entire luma distribution of the primary graphics; it is sufficient to at least roughly know the typical luma of the brightest colors in the primary graphics. Multiple approaches can be used, alone or in combination, to analyze the HDR video signal and its image context and the luminance or luma of its object pixels. In the latter case, some algorithm or circuit determines the final best graphics range based on the input of one or more circuits that apply the analysis method (if only one circuit / method exists in any device or method, such a final integration step or circuit is not necessary).

[0044] Preferably, the method for analyzing a high dynamic range image signal includes detecting one or more primary graphics elements in a high dynamic range input image, establishing lumas of pixels of the one or more primary graphics elements, and summarizing the lumas as a lower luma limit and an upper luma limit, where the lower luma limit is lower than all or a majority of the lumas of the pixels of the one or more primary graphics elements and the upper luma limit is higher than all or a majority of the lumas of the pixels of the one or more primary graphics elements.

[0045] When no further information is available or reliable, analyzing the image itself may be a viable option. For example, the algorithm (or circuit) might detect the presence of two primary graphics elements in the “original” version of video content. Consider subtitles displayed in three colors (white, light yellow, and light blue) for three speakers and a logo displayed in multiple colors (e.g., several primary colors and a black color that is 5% brighter than the white of the subtitles or logo itself). In many cases, the white of the logo and the white of the subtitles (or at least the brightness of the white predicted by extrapolating, for example, the yellow pixels of the logo) are expected to be the same or nearly the same. In particular, even when the two colors are not too different (e.g., the subtitles are five times brighter than the white of the logo, or vice versa), the method of the present invention allows for a relatively simple definition of the graphics range R_gr by treating the darker white as if it were a “gray” relative to the lighter white. That is, the top of the primary graphics range R_gra could be the white of the brightest pre-mixed graphics, and yet all graphics could be identified as one range, rather than applying twice the identified difference range (in which case, for example, secondary graphics could be mixed in either range, or the preferred range). If the brightness (intensity of white) of both graphics elements differs significantly, the brighter primary graphics element may be retained in the graphics analysis, the darker primary graphics element may be discarded, and R_gra may be calculated based on the brightest primary element. This embodiment may be useful when the luma mapping function does not have a safe graphics range as described in Figures 3 and 4, e.g., it has no special region and is simply a simple power function that acts as a continuous compression function.

[0046] Preferably, the method includes reading from metadata in the high dynamic range image signal the luma lower limit (Y_low) and / or luma upper limit (Y_high) values ​​written in the metadata of the high dynamic range image signal. Using this mechanism, the creator of at least one high dynamic range input image can communicate what they (or an automaton) consider to be the appropriate graphics range for this image or this video image shot (e.g., balancing the bright pixels of a nearby explosion that the human visual system should not perceive as negligible, given the perception of the white color of a nearby subtitle). The creator sets the primary graphics to such a luminance. Note that the human visual system can make all sorts of assumptions about what is in the image (e.g., local lighting), which can be quite complex, but the technical mechanism of the present disclosure provides a simple solution for those who should know (e.g., the human creator, color grader, or post-producer of the video). For example, the communicated range could be slightly darker than the luminance actually used in the primary graphics, thereby warranting adjustment and higher brightness of the primary graphics. In this case, the integration step or circuitry may simply discard or turn off other functionality, such as analysis of the HDR image itself, or any available mechanism may be used.

[0047] Preferably, the method includes analyzing the high dynamic range image signal, the analysis including reading two or more luma mapping functions associated with temporally consecutive images from metadata in the high dynamic range image signal, and establishing a range of graphics luma as a range that satisfies a first condition for the two or more mapping functions, where the two or more luma mapping functions map input lumas within the graphics range to corresponding output lumas, and the first condition is that, for each input luma, two or more corresponding output lumas obtained by applying a respective mapping function from the two or more luma mapping functions to the input luma are substantially the same. Figure 7 shows an example where two functions have subranges of the input range over which the two functions are (substantially) identical, i.e., for any input value within the range, both produce (almost) the same output value. This is detectable because the content creator intended to create a graphics-robust dynamic luma mapping function.

[0048] Thus, not only will that range be mapped to a similarly stable output range, but various lumas within it will be mapped to substantially the same output luminance by both functions (except for slight deviations that are typically not very noticeable, e.g., within 2% of the difference noticeable to typical human vision, or within 10% if some flicker is acceptable). It may be desirable to supplement such analysis with further analysis.

[0049] Specifically, this may include verifying the second condition: that the primary graphics luma determined based on the function falls within a secondary range of luma appropriate for graphics. Simply finding an area with a variety of functions across at least a series of shots (e.g., an indoor shot followed by an outdoor shot) does not necessarily mean that graphics should be placed there. For example, identifying only the darkest 10% of the total range (input or output) may lead to the conclusion that dark graphics are generally undesirable. For example, half to one-tenth of a high brightness dynamic range (e.g., maximum brightness of 4000 nits or greater) may be a suitable location for placing subtitles or other graphics. However, half of the range may not be optimal above 4000 nits. This produces maximum brightness for very bright graphics. For graphics such as text, where the brightest color is normal (diffuse) white, a 2000 nit level may be considered high, as opposed to more general-purpose graphics, where there may be several variations of white that constitute super white. However, the creator of the mapping function may have considered a stable graphics range of, say, 0.8*0.5 to 1.2*0.5 to be optimal for a 5000-nit master grade. Using this curve to downgrade the output to a maximum output of 1000 nits, 500-nit graphics, even subtitles, may be deemed reasonable—that is, meet the conformance criteria. If the second criterion is not met, the graphics blending device may, for example, decide to blend graphics with luma below the lowest value of the graphics range, attempting to approach the identified range but depending on the identified function shape (i.e., there may be subranges where two or more functions project the same input subrange to the same output subrange).Various secondary criteria can be used, such as the percentage of the maximum value of the output range that the upper limit of the stable graphics range identified by the shape of the function should be (i.e., this also depends on the absolute maximum value; i.e., the higher the maximum value of the output range, the lower the acceptable percentage may be, e.g., a 500 nit upper limit for a 1000 nit range, or a 1200 nit upper limit luma for an output range maximum of 4000 nit or greater), and / or the absolute luminance value at which this upper limit of the output range should appear. Thus, while the appropriateness of the range identified by the function typically involves checking values ​​along the range of the input luma (i.e., the received image), the appropriateness may also involve the output range and the maximum luminance of the expected output range (i.e., what the transcoding that will be mixed with the graphics will produce, for example), and it may be difficult to determine what is appropriate from the input luma value alone. For example, original graphics brighter than 2000 nits that exist as pre-mixed graphics in the input image will typically have a maximum output of less than 1000 nits. This means that, for example, using a leveling soft clipping mapping function starting with a maximum input of 4000 nits will result in an output of 900 nits (which is a fairly bright subtitle compared to the HDR effect of the brightest video objects, which should ideally be more impressive than the graphics white level, but 900 nits of graphics may still be acceptable in some cases and for some users, so it may be a level that should not be taken lightly, at least in itself).A suitable secondary criterion is that the maximum value of the secondary range should be lower than a pre-fixed percentage of the maximum value of the mixed image and graphics range.

[0050] Any device can be programmed with predictable graphics limits. For example, they can be expected to be the same as or somewhat above the expected subrange of regular ("LDR") image objects within an HDR image. For a 1000 nit ML_C defined HDR image, graphics whites can be expected to typically fall within the range of 80 nit to 350 nit, with blacks lower. This also depends on whether there are multiple primary graphics elements and whether their luminance (or luma) characteristics are identical.

[0051] Thus, an embodiment of the verification is as follows: If an analysis of the luma mapping curve determines that the upper luma limit for graphics is, for example, 150 nits, this certainly falls within the range of expected luma values ​​for bright graphics pixels, and also falls within the broader range [80.350]. Similar considerations can be made for the lower level limit, but as mentioned above, the lower level limit is less critical and may simply be set, for example, to 10% of the determined upper luminance limit (or the corresponding x% of the upper luma limit when examining the applicable EOTF for pixel color coding). For example, if this method estimates the upper luma limit luminance to be 950 nits, this may be an accident of the analysis, since ideally such overly bright graphics would be undesirable. However, it is also possible that the content provider actually created such bright subtitles. In such situations, rather than simply rejecting the values ​​and concluding that the graphics range could not be determined with sufficient certainty, further analysis may be performed, such as checking what graphics luma the creator actually sent together in the video signal, or further image analysis may be attempted by looking for such luma values ​​within the video signal and determining, for example, how connected or large those sets of pixels are.

[0052] It should be noted that an HDR image does not necessarily require that primary graphics actually be present at each appropriate luma location (e.g., a first shot of an image having a first luma mapping curve tuned to map a primarily bright HDR scene may actually contain graphics elements, while a subsequent second shot having a different luma mapping curve for mapping normal and bright image objects may not have actual graphics inserted, or if inserted, they will be set to approximately the same luma value).

[0053] Preferably, the method includes the step of luma mapping pixel lumas of at least one high dynamic range input image to corresponding output lumas of at least one corresponding output image, wherein the blending is performed within the at least one corresponding output image according to at least one luma mapping function obtained from metadata of the high dynamic range image signal.

[0054] This relationship, especially with functions like those in FIG. 3, can allow mixing in both the input and output luma domains. Whether mixing occurs in the input and output domains may not always be the same (although in some embodiments supporting some HDR codecs, it may be useful to know where in the input luma domain to mix). In any case, video pixels are often mapped to the output domain by a luma mapping function received in metadata. The method or device can use this mapping function to determine the appropriate luma location in the output domain (e.g., ML_C up to 350 nits) for mixing the output-determined graphics. This allows flexibility as to which devices can mix what and when. In particular, using stable subregion functions makes it easier to understand what will happen to graphics mixing already performed, even if those functions are distorted in display adaptation scenarios (i.e., changing the shape of the original function to something closer to a quadratic function will expand or compress its stable subrange, but if the adaptation is performed correctly and both the linear and basis functions respect that subrange, the same adjusted mapping characteristics will be maintained for luma within the range).

[0055] Preferably, the method includes designing the secondary graphics image element using a set of colors spanning a color scale of different lightness, some of the colors of different lightness appear lighter than average to the human eye and some appear darker than average, and at least the lighter than average colors are blended into at least one high dynamic range input image using a luma within the range of graphics luma.

[0056] A color scale is a set of colors of increasing (or decreasing) lightness, e.g., from dark to light, and often to colors of intermediate lightness. The scale need not necessarily include all steps of a particular primary color (e.g., blue), but a limited color palette might, for example, have a few light blue steps and a few dark green steps but no navy blue (in which case green defines the dark steps of the scale). When visually inspecting the approximate average lightness of graphic elements, some colors typically appear dark and some appear light. A lightness of approximately 50% or a brightness of approximately 25% may be typical values ​​that can be viewed or used as a midpoint. Thus, pixels with a brightness below, e.g., 25%, may be displayed as dark and no longer need to meet the criteria for the graphics range. If a more rigorous system requires more colors within the graphics range, secondary graphics may be designed with less dark colors.

[0057] Some embodiments of this method of secondary graphics insertion can perform the blending using pixel substitution (i.e., drawing graphics pixels of the appropriate established luma where video pixels were in the original video), which is a more predictable blending method, or by blending in a percentage of less than 50%, preferably 30% or less, of the luminance of at least one high dynamic range input image, so that the video shows through somewhat and is still visible, but the luminance of the graphics dominates and is largely preserved from the luma determination.

[0058] Preferably, the method for designing the color set for the secondary graphics elements includes selecting a limited set of relatively bright, below-average dark colors, where the below-average dark colors have a luma higher than the secondary luma lower limit, which is a certain percentage of the lower limit luma, preferably higher than 70% of the lower limit luma. Once the graphics range R_gra is determined, appropriate secondary graphics colors can be determined. For example, the method can mitigate the impact of potentially large variations in luma mapping by ensuring that the range does not fall too far below the determined safe lower limit, e.g., less than 30. This percentage may depend on aspects such as the number and location of dark pixels in the secondary graphics. For example, if a 100x100 pixel graphic has only five dark pixels, the average perceived lightness of the secondary graphics will not be significantly affected by these few dark pixels, regardless of how they ultimately appear in the output image.

[0059] The method may be implemented as an apparatus, one example of which is an apparatus (500) for determining a second luma of a pixel of a secondary graphics image element (216) and blending the secondary graphics image element with at least one high dynamic range input image (206) that includes a primary graphics element (207), the apparatus comprising: an input (501) for receiving a high dynamic range image signal (S_im) containing at least one high dynamic range input image; an image signal analysis circuit (510) that analyzes the high dynamic range image signal to determine a graphics luma range (R_gra) of at least one high dynamic range input image, the graphics luma range being a subrange of the luminance range of the high dynamic range input image, and determining the graphics luma range includes determining a lower luma limit (Y_low) and an upper luma limit (Y_high) that specify endpoints of the primary graphics luma range (R_gra); a graphics generation circuit (520) that generates or reads the secondary graphics elements and luminance maps the secondary graphics elements using at least the brightest subset of pixel lumas of the secondary graphics elements that fall within a range of graphics lumas; an image mixer (530) that mixes the secondary graphics image element with the high dynamic range input image to produce a pixel having a blended luminance (Lmax_fi); and an output unit (599) for outputting at least one blended image (Im_out) comprising pixels having the blended luminance (Lmax_fi).

[0060] Alternatively, the image signal analysis circuit (510) is a device including an image graphics analysis circuit (511), and the image graphics analysis circuit is Detecting one or more primary graphics elements in a high dynamic range input image; Establishing luma for pixels of one or more primary graphics elements; Summarizing the lumas as a lower luma (Y_low) and an upper luma (Y_high), where the lower luma is lower than all or most of the lumas of the pixels of the one or more primary graphics elements and the upper luma is higher than all or most of the lumas of the pixels of the one or more primary graphics elements.

[0061] Alternatively, the image signal analysis circuit (510) includes a metadata extraction circuit (513) that reads from the metadata of the high dynamic range image signal, for example, values ​​of lower luma (Y_low) and upper luma (Y_high) written into the high dynamic range image signal by the creator of at least one high dynamic range input image, in the device.

[0062] Alternatively, the image signal analysis circuit includes a luma mapping function analysis unit, wherein the luma mapping function analysis unit performs the steps of: reading two or more luma mapping functions of temporally consecutive images from metadata in the high dynamic range image signal; and establishing a range of graphics luma as a range that satisfies a first condition for the two or more mapping functions, wherein the two or more luma-luma mapping functions map input lumas within the graphics range to corresponding output lumas, and the first condition is that, for each input luma, corresponding two or more output lumas obtained by applying a respective mapping function from the two or more luma mapping functions to the input luma are substantially the same.

[0063] Alternatively, the image mixer includes a luma mapper that maps luma of pixels of at least one high dynamic range input image to corresponding output luma of at least one corresponding output image, the at least one corresponding output image having a dynamic range (e.g., maximum luminance) different from the dynamic range of the at least one high dynamic range input image, and the luma mapper further performs mixing within the at least one corresponding output image according to at least one luma mapping function obtained from metadata of the high dynamic range image signal.

[0064] Alternatively, the graphics generation circuitry designs the secondary graphics image elements using a set of colors spanning a color scale of different lightness, some of the colors of different lightness appear lighter than average to the human eye and some appear darker than average, and at least the lighter-than-average colors are mixed into at least one high dynamic range input image using a luma within the range of graphics luma.

[0065] Alternatively, the image mixer performs the mixing by replacing video pixels of the at least one high dynamic range input image with secondary graphics element pixels or blending at a percentage of luminance of the at least one high dynamic range input image that is less than 50%, preferably 30% or less.

[0066] In particular, those skilled in the art will understand that these technology elements can be embodied in various processing elements, such as ASICs (application-specific integrated circuits, i.e., IC designers typically configure (part of) an IC to perform a method), FPGAs, programmed processors, etc., and can reside in various consumer or non-consumer devices, whether they include a display (e.g., a mobile phone encoding consumer video) or a non-display device that can be externally connected to a display. Those skilled in the art will also understand that images and metadata can be communicated via various image communication technologies, such as wireless broadcasting and cable-based communication, and that these devices can be used in various image communication and / or usage ecosystems, such as television broadcasting, on-demand via the Internet, video surveillance systems, video-based communication systems, etc. An innovative coded HDR signal may correspond to the various methods described above, for example, at least one lower and upper luma value of at least one primary (pre-mixed) graphics range can be communicated. [Brief explanation of the drawings]

[0067] These and other aspects of the method and apparatus according to the present invention will be elucidated with reference to the implementations and embodiments described below and the accompanying drawings, which are merely non-limiting and illustrative diagrams illustrating general concepts, and in which dotted lines are used to indicate that a component is optional, but not necessarily that a component not shown in a dotted line is required. Dotted lines may also be used to indicate elements hidden inside an object, or intangible things such as object / region selection, although described as required.

[0068] [Figure 1] FIG. 1 illustrates schematically one example of an HDR processing chain (including encoding, decoding, and possibly display) in which the present innovations may provide benefits. [Figure 2] FIG. 2 shows a schematic representation of a typical situation where primary graphics may occur, but secondary graphics need to be inserted in one or more places. [Figure 3] FIG. 3 illustrates schematically one possibility for defining an image content adaptive luma mapping function that allows for more stable and determinable graphics insertion. [Figure 4] FIG. 4 illustrates schematically another example of the possibility of defining an image content adaptive luma mapping function that provides more stable and determinable graphics interpolation for multiple different luminance content scenes. [Figure 5] FIG. 5 shows a schematic diagram of typical components of an apparatus illustrating some of the principles of the present invention. [Figure 6] FIG. 6 illustrates schematically one possible embodiment that defines at least an upper luma limit, and typically also a lower luma limit, of an identified primary graphics range from an analysis of the primary graphics in at least one received high dynamic range image. [Figure 7] FIG. 7 illustrates schematically one possible embodiment of defining at least one upper luma limit, and typically also a lower luma limit, from at least two consecutive luma mapping functions formulated according to the principles described in FIG. 3 or FIG. 4. [Figure 8] FIG. 8 illustrates blending primary graphics and HDR video with pre-compensation luminance mapping according to the prior art. DETAILED DESCRIPTION OF THE INVENTION

[0069] Figure 3 illustrates in more detail how to optimally construct HDR video images containing graphics. The figure shows two typical HDR scene images (e.g., temporally adjacent image shots) for which specific grading of pixel brightness may be desired to achieve a striking HDR effect. In particular, HDR images allow for the creation of bright and saturated pixel colors, whereas LDR requires the desaturation of bright colors, resulting in a significant loss of aesthetic appeal. Additionally, color graders typically consider the psychovisual impact of various image objects, ensuring, for example, that a dragon's flames appear bright, but not too bright.

[0070] While not intended to be limiting, there are two scenes that a creator might typically process as follows. Both scenes (similar only in that they contain significant areas of high brightness) have very different content and brightness specifications (i.e., different grading), but both contain several regular reflective objects (i.e., objects that reflect a certain percentage of the light that locally falls on them). These reflect the average amount of light present in the scene, resulting in a relatively low brightness equivalent to that achieved with LDR. The idea is to create dark colors that are lower than the high brightness range of luminous colors, such as light bulbs and clouds that strongly reflect sunlight, when creating images to be displayed after or assuming human visual adaptation. In the case of image 301 of a fire-breathing dragon, this corresponds to the dragon's body. In image 302 of a sunset over the ocean, this corresponds to a boat. These regular objects fall within the dark subrange R_Lrefl, which covers the lowest brightness available within the input dynamic range (in this example, the full range is 0 to MaxL_HDR=3000 nit), and in these scenes, the maximum is, for example, 200 nit. If the dragon's body is red, it might have pixels that average 130 nits, for example, while if it's black, it might have pixels that average 20 nits, for example. Graphics pixels, such as subtitles, are typically placed slightly above this dark range. This has the advantage of being bright and noticeable (like white in LDR) while not overwhelming the HDR effect of the video itself. One could argue that the upper luminance limit (or luma coding of that luminance, e.g., perceptual quantizer luma definition) does not need to be the same for both ranges of normal scene colors, especially when two images are being mixed from separate HDR sources. However, adjusting the video images before mixing them can limit it to that vicinity (e.g., the dragon's range is up to 1.5 times brighter than the boat and ocean ranges at the upper end of its range). This can then be taken into account, for example, by setting the graphics range to start at the brighter end of the two (e.g., starting at 300 nits instead of around 200 nits, or 220 nits instead of around 150 nits).An HDR effect, for example, the orange / yellowish flames of a dragon, might be rendered at a pixel brightness of about 1500 nits (since this is a relatively large portion of the image, a brighter brightness would be advantageous). Sunlit clouds at sunset might reach, say, 2000 nits, and the sun disk itself might be about 3000 nits (e.g., 95% of 3000 nits for yellow). All of these HDR effects extend far beyond the graphics range. Video creators can specify that these should exist between, say, 400 nits and 500 nits (or 50 nits and 500 nits, given the small visual impact of dark graphics colors) within the master HDR range ML_C=3000 nit. For graphics that use only a limited subset of bright colors, they can focus on such a limited range and cover less of the colors in regular and effect video that may vary during grading, especially regrading with luminance mapping. If dark colors are required, the extent to which they fall below the 400 level can be adjusted, for example, based on the expected or reported variation in the bottom of the mapping function. (For example, offline, all curves occurring during the movie can be reviewed before processing, characteristic values ​​for the variation in the top of the lower curve (shown as the second line segment in this example), and such characteristics can be reported when displaying the reported video on the fly to take into account when determining where to map the deepest black in the graphics.) For example, if the curve begins to vary by more than 20% or 40% compared to the average curve starting downward from a fixed point, the low graphics range, additional luma or luminance values ​​(below L_low and Y_low, respectively) can be reported. For more critical applications (or types of critical graphics, such as black text boxes around monochrome subtitles), a variation point of 20% or less for the darkest colors might be used, while for less critical applications, the 40% point might be used. In many cases, the upper limit can be more important than the lower limit (for example, if bright pixels are heavily compressed during regrading).The 400-nit lower luminance limit L_low is expressed as the lower luma code limit Y_low (e.g., 0.65 multiplied by a bit-depth-dependent coefficient (e.g., 1023) for PQ). While not intended to be limiting, this specification assumes that all pixel luminance is always coded with a common perceptual quantization coding, although other EOTFs or OETFs, such as Hybrid Log-Gamma (HLG), may also be used. The inventive technique also works with mixed EOTF / luma code definitions (e.g., locating a first primary graphics element with a PQ luma-defined pixel color coding and locating a second primary graphics element as a secondary reference in an HLG-encoded image portion). The upper primary graphics range luminance L_high and upper primary graphics range luma Y_high are typically determined, for example, when significant HDR bright objects are not expected to occur frequently and therefore no particularly strong boosting or compression of local curve shapes is expected. The continuity of the curve starting from a fixed point may be defined as a percentage multiplier (e.g., 150%) of the lower limit, since it causes less visual mapping discrepancy issues if the graphics remain within that range. The amount of graphics typically depends on the maximum luminance of the input video, since it is expected that various HDR objects in the upper range will need to be compressed heavily. For example, if the input image has a maximum of 1000 nits, choosing an upper limit of 1.5x400 nits is likely undesirable because it would either leave little room for bright HDR objects (which may require grading away from regular objects, which is not always desirable), or, if the established primary and secondary graphics ranges coexist, there would be significant overlap between the luminance of the HDR video objects and, at least, the higher luminance possible within the secondary graphics range.

[0071] While the graphics problem is not so complicated when only the HDR image itself (i.e., the final image, displayed at equal luminance, i.e., each coded image pixel luminance is displayed exactly as coded) is available, the HDR input image typically requires luma mapping (or equivalently, luminance mapping) (as described in FIG. 1). For the Dragon image, a first luminance mapping function 310 is shown for obtaining a corresponding graded LDR image as output. (The corresponding luma mapping function can be obtained, for example, by converting both the input and output luminance ranges to the PQ luma range, or by converting the input as the PQ luma axis and the output as the conventional LDR luma of Rec. 709.) This function conceptually consists of three parts: a mapping of an appropriate shape for dark, regular objects in the input range R_Lrefl, a mapping for the HDR effect range R_Heffs, and a mapping for the intermediate part, i.e., the graphics subrange R_gra. The next scene in the video may be a sunset scene containing multiple consecutive sunset images (i.e., the shot). It may have a separate preferred regrading luma mapping function (315), which, as described in FIG. 1, is typically sent along as metadata with the video signal to guide the display optimization of the end device (e.g., a consumer TV display). However, the video creator may define the regrading function so that the graphics portion remains stable, i.e., all luminances are mapped to corresponding output luminances that are the same in both HDR scenes. In that case, the graphics will look the same in both scenes, in both HDR and LDR (or any other intermediate display-adaptive image, e.g., a 950-nit ML_C image optimized for a 950-nit ML_D display). Also, the video creator (i.e., creator of the luminance mapping function) can select the appropriate graphics subrange for both HDR (in this example, the 3000-nit master HDR) and LDR.If you need to adjust the graphics by mixing HDR videos with different maximum brightness levels, you can do so through the mapping output stable graphics range (75-80nit LDR range in this example). The graphics can be freely set, so it may depend on what is included in the HDR movie's effect range.

[0072] As illustrated in Figure 4, the remainder of the movie contains two essentially dark scenes. The first is a night scene 401 of a motorcycle riding through a city. The motorcycle may be rendered at, say, 30 nits to give a relatively dark impression while still remaining fully visible. (In LDR, tricks like turning pixels blue were necessary to simulate night, but HDR allows for more play with the darkness of the pixels themselves. However, one might choose to stay on the bright side of the subrange of the darkest pixels, assuming, for example, that HDR movies are often viewed in relatively brightly lit environments.) The only bright object in this image is a streetlight, so its pixel brightness may be set to 900 nits so that it isn't obnoxiously bright and distracting compared to the rest of the image. This is followed by another dark scene 402. While not strictly a night scene, it is still generally dark (it's a rainy day, so outdoor pixels appear dim). The very dark portion is a clown hiding in a sewer, with a brightness of, say, 10 nits or less to make the clown nearly invisible. For these two sets of images, a video creator might consider a lower typical level for the primary graphics more appropriate, such as an exemplary lower luminance limit L_low2=80 nit for HDR (represented as luma code Y_low2 in the HDR image) and an upper luminance limit L_high2=90 nit for HDR grading. Also shown are the corresponding appropriate LDR lower luminance limit (Ll_low=60) and upper LDR luminance limit Ll_high=65 nit. These depend on the graphics portion of the two luminance mapping functions for these two dark scenes (e.g., a third luminance mapping function 405 for the nighttime image 401 of the motorcycle). The dependency may be reversed. The darkest portion of the luminance mapping function 405, shown as a line segment for simplicity, can be adjusted so that the clown appears appropriately dark in a re-graded LDR image that could be used to drive a conventional LDR television, for example.The sudden change in luminance of the graphics with this scene change (e.g., from a bright fire-breathing dragon in daytime to a cityscape at night) is not a problem because it is what the video creator wanted, as long as the change is not "arbitrary." This mechanism allows the video creator to select all the luminances of the various different image objects as desired, rather than the arbitrary changes that would occur when using a simpler curve, such as a power function luminance mapping function with a variable power coefficient. Note that not all colors used in the graphics need to be within subranges that remain evenly mapped between the various luma mapping curves. In general, at least in some applications or systems, it is sufficient for the light colors of the graphics to remain stable, while the dark colors can vary somewhat. A device that mixes secondary graphics may decide to use a narrower range of graphics luma than the primary graphics. For example, only lumas that do not map differently across consecutive scenes with different regrading may be used.

[0073] FIG. 5 illustrates (schematically) an example of an apparatus embodying the concepts of the present invention.

[0074] In this exemplary embodiment, it is assumed (without limitation) that the HDR image signal input unit 501 receives not only the HDR image itself (i.e., the coded pixel color matrix and information for decoding (e.g., the selected EOTF) that is transmitted together or is known to the receiver), but also at least one luma mapping function LMF for a time instant t corresponding to one of the images in the video. This function may have a variable shape for consecutive images or shots of similar images of the same scene (e.g., a cave scene). Also, there may or may not be explicit metadata regarding the lower and upper luma limits of the primary graphics (MET(YLo, Yhi), rather than using L for luma and Y for luma). It is assumed that these are defined in the input domain, i.e., for the received HDR image. An image graphics analysis circuit (511) may be present within the device. How this works varies from device to device. For example, a simple device may only detect text and a graphics range from the detected text (e.g., a range from the darkest text pixel to the lightest pixel that includes all lumas used in all text found, or a portion of it if the range is larger (e.g., text pixels that are brighter than the average text luma)).

[0075] 6 illustrates an example of a suitable circuit for performing analysis of primary graphics within an image (although other analysis algorithms may be used). Those skilled in the art will appreciate that other equivalent graphics analysis techniques may exist, either alternatively or complementary to the various configurations of units taught to illustrate the present innovation.

[0076] The limited color variation detector (601) is configured to identify typical graphics colors within typical graphics low-level elements. While some graphics may be complex, many have a limited subset, e.g., only two distinct chromaticities, even when illuminated differently. For example, because text may be present in an HDR image (e.g., a horse's name may have the same white or single pixel color as the horse's colored head shape), a text detector 602 may typically be present to make an appropriate or initial determination of the primary graphics color. A segmentation map, such as the first segmentation map (SEG_MAP1), is a simple way to summarize areas identified as likely representative primary graphics pixels for easy later determination of the region / map color. Thus, for example, if the text character "G" is identified, the pixels that make up the character receive a value of, e.g., "1" in an initial matrix of zeros. A characteristic-based graphics analyzer 610 may also be present and may be used in an iterative manner for more complex graphics (e.g., the graphics color characteristics are identified, a set is determined, and the characteristics are re-identified). This can be done, for example, by using the Color Type Characterizer 611. For example, a highly saturated color like a strange purple often (except for flowers, etc.) suggests that the pixel is likely to be a graphic (graphics often contain at least some primary colors, with some RGB components high or maximum and one or two components at or near zero; natural image content ideally contains fewer such colors, but rather more muted colors). These candidates may be further verified by other units, such as the Basic Graphics Geometry Analyzer 612. For example, graphics with characterizable shapes may be matched against their shape characterizer.

[0077] A small set of multiple connected or nearby (e.g., repetitive) pixels with a particular color characteristic may be a graphics element. This is because, especially when there are multiple graphics, it is usually desirable to make them unobtrusive and sufficient for reading or viewing. Also, key objects in a movie are often enlarged and larger (e.g., a purple coat or a magic fireball may contain more pixels). Location can also be a heuristic; graphics typically reside near the image boundaries, such as a logo at the top or a ticker tape at the bottom. Such pixels may further be initial candidates in the first segmentation map SEGMAP_1 for validation, or vice versa, may be re-extracted through further analysis. Those skilled in the art will understand how to use, for example, neural networks trained using aspects such as the simplicity or frequency of change of graphics and natural video for contrast (e.g., texture metrics). A higher-level shape analysis circuit 620 may analyze further characteristics of the originally assumed graphics region (i.e., starting, for example, from a segmentation map that first identifies potential graphics pixels) to obtain a more stable set of pixels for summarizing luma. As noted above, it is not necessary for all primary graphics pixels to be accurately identified in detail. For example, there may be an edge detector 621 to detect edge pixels between the graphics shape and the surrounding video.

[0078] The G-criterion may be used to detect such boundaries (M. Mertens et al.: A robust nonlinear segment-edge finder, 1997 IEEE Workshop on Nonlinear Signal and Image Processing).

[0079] The G criterion may be used to define characteristics as needed (e.g., plainer natural object colors versus contrasting bright, saturated colors) and calculate the amount of their co-occurrence within two regions. Because counting is involved, the shape of the regions can also be adjusted as needed.

[0080] For example, define the first property using two saturations of the pixel color. P1=1000*Cb+Cr

[0081] This characteristic can be transformed by a further function, such as the deviation from a characteristic based on a locally determined average chromaticity, given by: Delta_P=P1_pixel - P1_determined P2=Function(Delta_P1)

[0082] This function may, for example, classify as 0 if Delta_value is below a first threshold, as 1-9 for intermediate delta values, and as 10 if the delta values ​​are sufficiently different (i.e., abs(1000*Cb-1000*Cb_reference)>1000*Threshold1 or abs(Cr-Cr_reference)>Threshold2).

[0083] Next, two regions, e.g., two adjacent rectangles, are defined that are moved across the image until they are positioned to the left and right of the horizontal boundary of the ticker tape. The amount of movement may depend on the amount of match.

[0084] G_criterion(G) is the sum of the absolute values ​​of the number of pixels in rectangle 1 with a given value P2_i (e.g., P2_0 means the red component of the pixel being counted = 0, P2_1 means red = 10) minus the number of pixels in rectangle 2 with the same value P2_i, for all possible different values ​​of P2. Finally, this sum of absolute differences is divided by a normalization factor, which is usually the number of pixels in both rectangles (or the number of pixels adjusted for area if the regions are of different sizes). When comparing two equally sized test rectangles in adjacent locations in the image, the above formula becomes: G=sum_i{abs[count(P2_i)_right_rectangle-count(P2_i)_left_rectangle]} / 2*L*W [Formula 3]

[0085] where L and W are the lengths and widths of the two rectangles.

[0086] If the detector is inside a graphic, it will detect that both rectangles have approximately the same color, all zeros, i.e., no edge. If one rectangle is on a graphic (e.g., a saturated yellow) and the other is on a video (e.g., a desaturated green (a small excess of green)), the moving average from the green side (e.g., on the green graphic) will be green, and the P2 value of the upper rectangle will be approximately zero. Compared to that continuous green, the P2 values ​​of the other sampled rectangle will all be lower, e.g., by approximately 10. In this case, the left rectangle will have L*W distinct pixels that are almost all zero, and the right rectangle will have L*W pixels with P2 characteristics (e.g., a value of 10 determined as input to the G criterion). The distinct characteristic color will not match on the left, i.e., the G criterion will approximate a value of 1 if it corresponds to an edge. An advantage of the G criterion is that any characteristic, such as a texture-based index, can be added for comparison. Other, more classical edge detectors may instead be used to find a set of candidate points that lie between the edge of the graphics region and the beginning of the natural video region. Edge detectors are often noisy, meaning that both gaps and false edge pixels can exist. Therefore, the shape analyzer 622 may include preprocessing circuitry to identify connected shapes in the HDR image from the detected edge pixels. Various techniques are known to those skilled in image analysis, such as using a Hough transform to detect lines, matching circles, or using splines or snakes. If the analysis reveals that four lines (or one line and the image boundary) form a rectangle with particular characteristics, such as being (exactly or approximately) the same width as the image and being at the bottom, the color inside is likely to be a graphics pixel (e.g., a news program ticker tape). Therefore, the entire rectangle may be added to a second segmentation map, SEGMAP_2. Determining symmetry based on the detected boundaries of what is likely to be graphics can confirm the actual presence of a primary graphics element.A simpler embodiment might focus on a few simple geometric elements, such as a rectangle detected at the bottom of the image (which is often sufficient for initial consideration of the graphics range R_gr). Then, only if no such elements are found, the algorithm can further explore more complex graphics objects, such as a star, a small area that appears and remains there for a few seconds across several consecutive video shots, and then disappears again (i.e., the time-varying behavior of the graphics is also taken into account). Alternatively, a graphics area / object identified as a candidate might remain particularly consistent across several different shots of a movie. The gradient analyzer circuit 623 may further analyze situations in which graphics do not consist of a fixed set of colors but have internal gradients (long-range gradients that typically change slowly over tens or hundreds of pixels). Such gradients include, for example, all yellow with saturation varying from left to right. The likelihood that a gradient is generated by a graphics source may be confirmed or rejected. For example, a blue gradient at the top of an image may be discarded as likely sky. In some embodiments, this may depend on further geometric characteristics of the gradient, such as the size of the gradient region, the steepness of the gradient, the amount of color the gradient spans, and especially whether the gradient terminates at a complex lower boundary (e.g., tree). In the case of a blue line that may distinguish water from sky, pixels may be discarded if the position of this line is too far below the upper boundary of the image.

[0087] The pseudo-image analysis circuit 630 may provide further certainty regarding the robustly determined graphics pixels in the third segmentation map SEGMAP_3. Modern, more attractive graphics may include, for example, a shape with (graphically generated) clouds in the lower banner. This may be confusing because it looks almost like a natural image. Even if correctly identified as a graphic, it may not provide significant new insight into the pixel luma of the graphics, identifiable from the rest of the graphics. If it is actually part of the graphics, it will have an adjusted color anyway. This may roughly overlap with the otherwise determined graphics range R_gra, resulting in a somewhat higher determined upper luma Y_high or a somewhat lower determined Y_low, but the method is not critical. Such regions attached to the identified graphics region, such as the lower left corner of a rectangular banner, may still be discarded (or, in advanced embodiments, retained if they are verified to be graphics elements associated with the rest of the banner, e.g., if they have an associated color; however, these pseudo-graphics may be complex and may be better discarded). The discarding can result in a segmentation map that is more reliable in determining the appropriate graphics range R_gra. More advanced embodiments can apply texture or pattern recognition algorithms to special regions of interest, i.e., regions with few colors and simple gradients (i.e., deviations to, e.g., 20% lower saturation) that are not geometrically simple and symmetrical, such as rectangular or substantially circular. For example, a busyness measure can be calculated that indicates how often and how quickly pixel colors change (e.g., which colors change) for each 10x10 pixel region.Calculating the angular extent of lines (e.g., the centroid of an object with small color variations) is another measure that helps distinguish between complex patterns in nature (e.g., leaves) and simple geometric patterns that appear in graphics (and text typically has only two or a few stroke directions). In a simpler embodiment, gradients may be calculated. If there are short-range gradients (i.e., large changes over a few pixels) rather than the long-range gradients (slow changes over tens or hundreds of pixels) that are common in graphics, such areas may be excluded from the map for robustness reasons. Also, if the color of an area of ​​large surrounding geometry of a graphics area (e.g., a text box) deviates significantly from the average color of the rest of the graphics area, even if it is a complementary color (i.e., a color that may have been intentionally chosen for this graphics, e.g., blue to complement orange), the processing logic of circuit 630 may remove those pixels from the set that ultimately determines the graphics range (i.e., remove those pixels from SEGMAP_3).

[0088] Regardless of the number of pixel region analysis subcircuits or processes (more or fewer than the illustrated three), the final processing is typically performed by a luma histogram analysis circuit 650. This circuit examines all lumas of pixels identified as primary graphics in the third segmentation map SEGMAP_3 (or an equivalent segmentation map or mechanism). It outputs a lower luma limit Y_low_ima based on image analysis that is at or near the lowest luma in the histogram of lumas of pixels identified as graphics in SEGMAP_3. If the luma histogram contains a large number of dark colors, the output may be higher than the minimum value, which is set, for example, to 10% or more of the highest luma to obtain a useful graphics range. Conversely, if only white graphics pixels are detected in SEGMAP_3, instead of setting the lower and upper limits to the same value, which is not useful, the luma histogram analysis circuit 650 can again output an image-analysis-based lower luma limit, Y_low_ima, such as 25% of the image-analysis-based upper luma limit, Y_high_ima (which typically determines more significant parameters first, and thus could be the luma of white text, for example). Also, Y_high_ima does not always need to be the same as the maximum luma found in SEGMAP_3. For example, if only yellows are detected as the brightest colors, knowing that their luminance is typically 90% of white, the upper luma limit can be set to 110% of the luminance of the maximum luma detected in the histogram. Alternatively, if a bright logo coexists with a dark white subtitle, a value corresponding to the white value appropriate for the brightest primary graphics element can be set.

[0089] Returning to Figure 5, there may also be a luma mapping function analysis unit 512 to analyze, for a series of (consecutive or adjacent) images of a video, regions of similarity in the mapping functions (i.e., regions where a subset of lumas is mapped to approximately the same corresponding output luma by all the different optimal image-dependent luminance mapping functions), an example algorithm for which is described in Figure 7. That is, any device will have unit 511, unit 512, or both, and will usually have some circuitry to summarize detected graphics areas, e.g., overlaps or junctions. Similarly for unit 513.

[0090] The various luma mapping functions (LMF(t)) valid at a particular time are input to the identity analysis circuit 701. One or more previous luma mapping functions LMF_p are retrieved from memory 702. If a function with a different shape is determined (i.e., the input LMF(t) has a different shape than the stored LMF_p, at least in some respects), the old LMF_p may be replaced or supplemented when comparing two or more functions. The identity analysis circuit 701 is configured to first check whether there is overall identity of the functions, rather than just a subregion of the input luma. Some HDR codecs send one function per image, but these functions should be discarded because they may all have the same shape for all images in the same scene. This is indicated by the establishment unit 703 detecting the ID Boolean FID as "yes" or "1." Then, the next function is simply loaded until a new function is actually loaded (e.g., the dragon function is the old one, and the sunlit sea 302 function in Figure 3 is the new one).

[0091] The subrange identity circuit 710 calculates the output value difference DEL for nearly all input luma values ​​Y_in. It is expected to find at least one outer region (e.g., below the graphics range) where there is a difference, and an intermediate range R_id where the output is identical. The lower and upper luma limits Yb and Yt of this range may be determined. These may be output as the final lower and upper luma limits of the identified graphics range R_gra, whether or not they are suitable for secondary graphics blending. Typically, the graphics range determination circuit 720 may perform further analysis before outputting the lower and upper luma limits Y_lowfu and Y_highfu determined by the function. As explained, this may be based, for example, on determining whether the upper luma limit is within a typical range, i.e., whether it is below a typical upper limit HighTyp from the range supply circuit 740 (dotted line, and therefore optional). The same can happen with respect to the typical lower luma limit LowTyp. Given a 5000 nit master HDR image, finding a graphics range near 5000 nit is likely not an appropriate graphics range because some viewers perceive the graphics as too bright. If the analysis fails, for example, if the requested Yt is significantly higher than HighTyp, the embodiment can suggest a value for Y_highfu that works well on average, but typically throws an error condition ERR2. It is also possible for a curve to be communicated that does not have an intermediate range for the identifier (R_id), which can also throw an error condition (first error condition ERR). For simplicity, the graphics range is shown here as being identical for both curves for all points within that range. However, in general, this range may be relaxed and deemed similar within a certain tolerance (e.g., a maximum 10% brightness deviation).

[0092] Returning to FIG. 5, in this situation, the integrated logic circuit 515 cannot use the inputs of the lower and upper luma limits from the luma mapping function analysis unit 512. In other cases, the similarity between Y_highfu and Y_high_ima can be checked, and if they are close, either value, or, for example, the average value, can be used as the final upper luma limit Y_high. Alternatively, the upper limit Y_hi based on the metadata provided by the metadata extractor 513 can be used directly (as well as the lower luma limit Y_lo from the metadata, if present; if not, it can be replaced by a percentage of Y_hi, such as 1 / 3 of the corresponding luminance L_hi_met). In either case, at least the upper luma limit Y_high and typically also the lower luma limit Y_low are communicated to the graphics generation circuit 520 so that these lumas can be taken into account when selecting colors for secondary graphics elements (or when converting pre-created graphics elements, if they need to be mixed). As explained above, the appropriate graphics luma Y_gra for each pixel is determined, typically within the graphics range (R_gra) or not significantly above or below it. The actual value of Y_gra depends on what the graphics pattern contains as pixel color and luma. For example, if the darkest color (Y_orig_graph_min) is mapped to Y_low and the lightest color (Y_orig_graph_max) of an existing or to-be-created graphic is mapped to Y_high, then the colors of pixels between Y_orig_graph_min and Y_orig_graph_max may be mapped proportionally (or non-linearly) between Y_low and Y_high.

[0093] Finally, the image mixer 530 blends the graphics color (using the graphics pixel luminance Lgra, since blending is often more sophisticated in the linear domain, or using the graphics pixel luminance Y_gra, for example, if the blending consists of simple pixel substitution) with the video pixel color for (substantially) every pixel in the image.

[0094] In some embodiments, such as those capable of blending in the output domain, a luma mapper (533) may be present (some device embodiments may be capable of blending only in the output domain, or only in the input domain (possibly before the final luma mapping), or both, and may switch between them as needed). Even when the appropriate graphics are blended in the input domain, a luma mapping function, such as the luma mapping function for Dragon, is applied to the pre-blended image of the secondary graphics. This function may be scaled to suit a variety of different end-user displays with different end-user display maximum luminance capabilities. However, with the innovations of this disclosure, the graphics remain relatively stable, even with luma mapping. However, as noted above, blending is also possible in the output domain (i.e., the vertical axis in FIG. 3) of any function. However, in that case, the luma mapper knows how to blend in the output graphics range if it knows the luma mapping function to be applied. The reader will still appreciate that this is a dynamic mapping (a mapping that can significantly change luminance or relative luminance) that has not yet been applied, which facilitates blending. In a typical application, the graphics are pre-scaled and sent to the output domain (e.g., they may be coded in PQ) for final blending by, for example, the end user's television.

[0095] The final mixed output image Im_out with the mixed luminance Lmix_fi of the pixels (or luma coded as they normally are) is provided via the image or video signal output 599.

[0096] The algorithmic components disclosed herein may in practice be realized (in whole or in part) as hardware (e.g., part of an application-specific IC) or as software running on a specialized digital signal processor or a general-purpose processor, etc.

[0097] Those skilled in the art will understand from this disclosure which elements are optional improvements and can be implemented in combination with other elements, and how the (optional) steps of a method correspond to the respective means of an apparatus, and vice versa. The term "apparatus" in this application is used in the broadest sense, i.e., as a group of means that enable the realization of a specific purpose. Thus, it may be, for example, (a small circuit portion of) an IC, a dedicated device (such as a device with a display), or a part of a network system. The term "configuration" is also used in the broadest sense and may include, inter alia, a single device, a part of a device, or a collection of (parts of) cooperating devices.

[0098] Reference to a computer program product should be understood to encompass any physical realization of a set of commands that, after a series of reading steps (which may include intermediate conversion steps such as conversion to an intermediate language or a final processor language), allows a general-purpose or special-purpose processor to input the commands into the processor and perform any of the characteristic functions of the invention. In particular, a computer program product may be realized as data on a carrier such as a disk or tape, data residing in a memory, data traveling via a wired or wireless network connection, or program code on paper. In addition to the program code, characteristic data required for the program may also be embodied as a computer program product.

[0099] Some of the steps required for the operation of the method, for example data input and output steps, may not be written in the computer program product but may already be present in the functionality of the processor.

[0100] It should be noted that the above embodiments are illustrative rather than limiting of the present invention. Those skilled in the art can easily map the presented examples to other areas of the claims, but for the sake of brevity, not all these options are detailed. In addition to the combinations of elements of the present invention combined in the claims, other combinations of elements are also possible. Any combination of elements can also be realized by a single dedicated element.

[0101] Reference signs in parentheses in the claims are not intended to limit the claims. The term "comprises" does not exclude the presence of elements or aspects other than those listed in a claim. The use of an "a" or "an" element in the singular does not exclude the presence of a plurality of such elements. In certain embodiments, each unit of an apparatus in the present teachings may be formed by a circuit on an application-specific integrated circuit (e.g., a color processing pipeline that applies technical color modifications to input pixel colors such as YCbCr), or a software-defined algorithm running on a CPU or GPU (e.g., in a mobile phone), or on an FPGA. Typically, the computing hardware, whether a general-purpose bit-processing computer or a specific digital processing unit, is connected under the control of operational commands to a memory unit, which may be on-board the specific circuit or off-board, connected via a digital bus or the like. Some of such computers may be directly connected to a larger device, such as a display panel controller or a hard disk controller for long-term storage, or to physical media such as a Blu-ray disc or USB stick. Some functions may be distributed to various devices over a network. For example, some computations may be performed on servers in the cloud.

Claims

1. 1. A method for determining a second luma of pixels of a secondary graphics image element to be mixed with at least one high dynamic range input image in a circuit for processing a digital image, the method comprising:

1. A method comprising receiving a high dynamic range image signal comprising the at least one high dynamic range input image; analyzing the high dynamic range image signal to determine a range of primary graphics luma for primary graphics elements of the at least one high dynamic range input image, the range of primary graphics luma being a subrange of the range of luminance of the high dynamic range input image, determining the range of primary graphics luma including determining a lower luma limit and an upper luma limit specifying endpoints of the range of primary graphics luma; luminance mapping the secondary graphics image elements with at least the brightest subset of their pixel lumas that fall within the range of the primary graphics lumas; and blending the secondary graphics image element with the high dynamic range input image.

2. 2. The method of claim 1 , wherein analyzing the high dynamic range image signal includes detecting the one or more primary graphics elements in the high dynamic range input image, establishing lumas of pixels of the one or more primary graphics elements, and summarizing the lumas as a lower luma limit and an upper luma limit, wherein the lower luma limit is lower than all or a majority of the lumas of pixels of the one or more primary graphics elements and the upper luma limit is higher than all or a majority of the lumas of pixels of the one or more primary graphics elements.

3. 2. The method of claim 1, wherein analyzing the high dynamic range image signal includes reading the luma lower limit and luma upper limit values ​​from metadata within the high dynamic range image signal that are written as metadata to the high dynamic range image signal.

4. 2. The method of claim 1, wherein analyzing the high dynamic range image signal includes reading two or more luma mapping functions associated with temporally consecutive images from metadata in the high dynamic range image signal; and establishing a range of graphics luma as a range that satisfies a first condition for the two or more mapping functions, the two or more luma mapping functions mapping input lumas within the range of graphics lumas to corresponding output lumas, the first condition being that, for each input luma, two or more corresponding output lumas obtained by applying a respective mapping function from the two or more luma mapping functions to the input luma are substantially the same.

5. 5. The method of claim 4, wherein the range of the graphics luma satisfies a second condition verification that the range of the graphics luma is contained within a secondary range of an appropriate graphics luma, and the maximum value of the secondary range must be lower than a predetermined percentage of the maximum value of a range of a blend of image and graphics.

6. 6. The method of claim 1, wherein the method comprises luma mapping pixel lumas of the at least one high dynamic range input image to corresponding output lumas of at least one corresponding output image, the blending being performed within the at least one corresponding output image by mapping the pixel lumas using at least one luma mapping function obtained from metadata of the high dynamic range image signal.

7. 6. The method of claim 1, further comprising determining the secondary graphics image elements using a set of colors spanning a color scale of different lightness, some of the different lightness colors appear lighter than average to the human eye and some appear darker than average, and at least the lighter-than-average colors are blended into the at least one high dynamic range input image using a luma within the graphics luma range.

8. 6. The method of claim 1, wherein the blending comprises pixel substitution or blending of a percentage of pixel intensities of the at least one high dynamic range input image that is less than 50%, the percentage preferably being equal to or less than 30%.

9. 8. The method of claim 7, wherein determining the set of colors includes selecting a limited set of relatively bright below-average dark colors, the below-average dark colors having luma higher than a secondary lower limit luma, which is a percentage of the lower limit luma, preferably higher than 70% of the lower limit luma.

10. 1. An apparatus for determining a second luma of pixels of a secondary graphics image element and blending the secondary graphics image element with at least one high dynamic range input image including a primary graphics element, the apparatus comprising: an input for receiving a high dynamic range image signal comprising said at least one high dynamic range input image; an image signal analysis circuit that analyzes the high dynamic range image signal to determine a graphics luma range of the at least one high dynamic range input image, the graphics luma range being a subrange of a luminance range of the high dynamic range input image, and determining the graphics luma range includes determining a lower luma limit and an upper luma limit that specify endpoints of a range of luma for pixels representing the primary graphics; a graphics generation circuit that generates or reads the secondary graphics elements and luminance maps the secondary graphics elements using at least a brightest subset of pixel lumas of the secondary graphics elements that fall within the range of graphics lumas; an image mixer that mixes the secondary graphics image elements with the high dynamic range input image to generate pixels having a mixed luminance; and an output section for outputting at least one blended image comprising pixels having said blended luminance.

11. The image signal analysis circuit includes an image graphics analysis circuit, the image graphics analysis circuit comprising: detecting the one or more primary graphics elements in the high dynamic range input image; establishing luma for pixels of the one or more primary graphics elements; and summarizing the lumas as a lower limit luma and an upper limit luma, the lower limit luma being lower than all or a majority of the lumas of pixels of the one or more primary graphics elements and the upper limit luma being higher than all or a majority of the lumas of pixels of the one or more primary graphics elements.

12. 11. The apparatus of claim 10, wherein the image signal analysis circuitry includes a metadata extraction circuitry that reads, from metadata of the high dynamic range image signal, the luma lower limit and the luma upper limit values ​​written to the metadata of the high dynamic range image signal.

13. 11. The apparatus of claim 10, wherein the image signal analysis circuit includes a luma mapping function analysis unit that is configured to read two or more luma mapping functions for temporally consecutive images from metadata in the high dynamic range image signal and establish a range of the graphics luma as a range that satisfies a first condition for the two or more luma mapping functions, the two or more luma mapping functions mapping input lumas within the graphics range to corresponding output lumas, and the first condition being that, for each input luma, two or more corresponding output lumas obtained by applying a respective luma mapping function from the two or more luma mapping functions to the input luma are substantially the same.

14. 11. The apparatus of claim 10, wherein the image mixer includes a luma mapper that maps luma of pixels of the at least one high dynamic range input image to corresponding output luma of at least one corresponding output image having a different dynamic range, e.g., maximum luminance, than a dynamic range of the at least one high dynamic range input image, and wherein the luma mapper further performs the mixing within the at least one corresponding output image according to at least one luma mapping function obtained from metadata of the high dynamic range image signal.

15. 11. The apparatus of claim 10, wherein the graphics generation circuitry designs the secondary graphics image elements using a set of colors spanning a color scale of different lightness, some of the different lightness colors appear lighter than average to the human eye and some appear darker than average, and at least the lighter-than-average colors are blended into the at least one high dynamic range input image using a luma within the graphics luma range.

16. 11. The apparatus of claim 10, wherein the image mixer performs mixing by replacing video pixels of the at least one high dynamic range input image with secondary graphics element pixels or blending at a percentage less than 50%, preferably 30% or less, of pixel luminance of the at least one high dynamic range input image.

Citation Information

Patent Citations

  • Method and device for improved HDR image encoding and decoding

    JP2018110403A

  • Graphics Blending for High Dynamic Range Video

    US20160080716A1

  • Transitioning between video priority and graphics priority

    US20180018932A1

  • Graphics-safe HDR image luminance re-grading

    US20200193935A1