Mixing of Secondary Graphic Elements in an HDR Image
The method addresses inconsistent luminance in HDR image mixing by determining a range of primary graphics luma and applying luminance mapping to secondary elements, ensuring stable and consistent brightness across varying display conditions.
Patent Information
- Application Number
- JP2024569131
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-24
- Filing Date
- 2023-05-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-05-09
AI Technical Summary
Existing technologies face challenges in accurately mixing secondary graphics elements with high dynamic range (HDR) images, as they often result in inconsistent luminance across different displays and scenes, leading to flickering or improper brightness levels.
A method and apparatus for determining a range of primary graphics luma in HDR images, using luminance mapping to adjust secondary graphics elements within this range, ensuring consistent luminance across varying display conditions.
Ensures stable and consistent luminance of secondary graphics elements across different HDR scenes and displays, maintaining visual coherence and reducing flickering issues.
Smart Images

Figure 2025519098000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for constructing a High Dynamic Range (HDR) image, which is a new image encoding method, especially as compared to a standard low dynamic range image including graphic elements. Specifically, the present invention may be applicable to more advanced HDR applications with luminance or tone mapping that may be specified differently for each scene image of a video. Thereby, from an input image having a primary grading, a secondary color grading of different luminance dynamic ranges ending at different maximum luminances can be calculated (grading generally includes redistributing the normalized luminance of pixels of various objects in the image from an initial relative value in the input image to different relative values in the output image, and the relative redistribution also affects the absolute luminance value of the pixels when related to the normalized luminance). Specifically, the present invention may be useful when it is necessary to mix secondary graphics elements by some video processing device in a previously created HDR image (the primary graphics elements already exist in the image before at least one secondary graphics element is mixed and need to be distinguished from the secondary graphics elements).
Background Art
[0002] Until the first studies around 2010 (and before the first HDR decoding TVs were sold in 2015), at least with respect to video, all videos were created according to a common low dynamic range (LDR), or rather standard dynamic range (SDR), encoding framework. This had several characteristics. First, a single video suitable for all displays was created. This system is a relative system in which the maximum (100%) signal encoded with the maximum luma code corresponding to the maximum non-linear RGB values R’=G’=B’=255 (255 in 8-bit YCbCr encoding) is white. There is nothing brighter than white, and all typical reflective colors can be represented as colors darker than such the brightest white (for example, a piece of paper reflects all incident light at most, or absorbs some of the red wavelengths and reflects blue and green back to the eye, resulting in a local cyan color that is somewhat darker than the white of the paper). Each display shows this brightest white as the brightest color technically constructed for that display to render (as a “drive request”), for example, 80 nits (abbreviation of SI quantity Cd / m 2 ), and 200 nits for an LCD display with TL backlight. Since the viewer's eyes quickly correct for differences in brightness, as long as they are not side by side in a store and are at home, all viewers saw approximately the same image (despite differences in display).
[0003] Rather than simply making the color printable or paintable on paper, there was a need to improve the perceptible appearance of an image by making the pixels actually glow, much brighter than the "white of the paper", also called "diffuse white". This actually means, for example, that whereas in the past, photography was done in studios, etc., and all important subjects were photographed while being brightly illuminated by ceiling lighting groups, nowadays, good-looking photos can be taken just by shooting with strong backlighting. Cameras have continued to evolve and can now simultaneously acquire a sufficient number of scene luminances for many scenarios and further adjust or complement the brightness with a computer. Displays, including consumer television displays, have also continued to improve.
[0004] In some systems, such as the BBC's Hybrid Log Gamma (HLG), this is achieved by defining values in the coded image that exceed white (white is given a reference level of "1", i.e., 100%), for example, up to 10 times white at most, thereby enabling the display of pixels that glow 10 times brighter.
[0005] To date, most systems have also shifted to a paradigm where video creators can define the absolute nit value of an image (i.e., not 2x or 10x compared to an undefined white level that is converted to the actual nit output variable at each endpoint display) on a selected target display dynamic range function. The target display is the virtual (intended) display on the video creation side (e.g., a 4000 nit (ML_C) target display to define a 4000 nit video), and the actual consumer endpoint display may have a lower display maximum luminance (ML_D), e.g., 750 nit. Even in such a case, the end display must typically be equipped with luminance remapping hardware or software, typically realized as a lumen mapping. The luminance remapping hardware or software adapts the pixel luminance in an HDR input image, whose luminance dynamic range (specifically the maximum luminance) is too high to be faithfully displayed, to the value of the dynamic range of the end display in some way. The simplest mapping is to clip all luminance above 750 nit to 750 nit, but this is the worst way to handle dynamic range mapping. This is because, for example, the beautiful structure of a sunset with clouds illuminated by the sun in the range of 1000 - 2000 nit in a 4000 nit image is removed and displayed as a uniform white 750 patch. A better lumen mapping moves a partial range of 1000 - 2000 nit in the HDR input image, for example, to the end display dynamic range of 650 - 740 by a properly determined function. The function may be automatically determined, for example, within a receiving device such as a TV or STB, or the video creator may determine the one most suitable for their artistic movie or program and communicate it with the video signal as metadata in a standardized format. Luminance refers to any encoding of luminance, for example, at 10 bits, using a function that assigns a luminance code of 0 - 1023 to a video luminance of, for example, 0.001 - 4000 nit by a so-called electro-optical transfer function (EOTF).
[0006] The simplest system is to simply transmit the HDR image itself (using an appropriately defined EOTF), for example with a maximum luminance of 4000 nits (i.e., provide the receiver with the image without showing a method of downscaling by luminance mapping, for example for a display with a low maximum luminance capability). This is what the HDR10 standard does. In more advanced systems such as HDR10+, it may also communicate a function for downmapping a 4000 nit image to a lower dynamic range such as 750 nits. These systems define a mapping function between two different maximum luminance versions of the same scene image, or between a reference grading, and use an algorithm to calculate a modified version of that reference luminance or luma mapping function to calculate other luminance mapping functions, for example endpoint functions for other display maximum luminances, to facilitate this (e.g., the reference luminance mapping function may specify how to modify the normalized luminance distribution within a 2000 nit input image (e.g., the master HDR grading video created by the author) to obtain a 100 nit reference video, and the display adaptation algorithm calculates the final luminance mapping function based on the reference luminance mapping function, enabling a 700 nit TV to map luminance from 0 - 2000 nits to display luminance within the range 0 - 700 nits). For example, if it is agreed to define the SDR image such that when newly interpreted as an absolute nit image rather than a relative image, it always has a maximum pixel luminance of 100 nits, the video creator can define and transmit together a function that specifies how to map the luminance from 0.001 (or 0) - 4000 nits, the first reference image grading, to the corresponding desired SDR 0 - 100 nit luminance (secondary reference grading). This is called display tuning or adaptation.For both the 4000 nit ML_C input image (horizontal axis) and the 100 nit ML_C secondary grading / reference image, when defining a function that boosts the darkest 20% of the colors in a plot normalized to 1.0, for example by a factor of 3, i.e., when reducing from 4000 nit to 100 nit, and when a particular end-user's TV needs to reduce down to 750 nit, the required boost can be, for example, only 2 times (this depends on the defined EOFT used for the lumen. Because, as described above, by implementing the luminance mapping as a lumen mapping, the influence of the luminance change along the range can be defined more visually uniformly, i.e., it can be defined to be more appropriate and visually impactful for humans, so luminance mapping is usually implemented as a lumen mapping (e.g., using a psychovisual uniform EOTF) in the color processing IC / pipeline).
[0007] A third class of more advanced HDR encoders takes this principle of defining the re-grading requirements based on two reference grading images to the next level, by rephrasing it in another way. If limited to the use of mostly reversible functions, for example, an LDR image that can be calculated on the transmitting side by down-mapping the luminance or lumen of a 4000 nit HDR image to an SDR image can actually be transmitted as a proxy for the actual master HDR image created by a video creator (e.g., by a Hollywood studio for BD or OTT delivery, or by a sports broadcaster). The receiving device can then apply the inverted function to reconstruct an accurate reconstruction of the master HDR image. A system that communicates the HDR image itself (in the created state) is called "mode HDR", and a system that communicates an LDR image is called "mode LDR coding framework".
[0008] A typical example is shown in FIG. 1 (for example, a summary of the principle for which the applicant of the present application previously obtained a patent in WO2015 / 180854 or the like). FIG. 1 includes the decoding function itself and subsequent display adaptation as a block, and it should be noted not to confuse these two technologies.
[0009] FIG. 1 schematically shows an exemplary video coding and communication, and processing (display) system. On the creation side, an embodiment of a typical encoder (100) is shown. Those skilled in the art will understand that the pixel-by-pixel processing pipeline of the luminance processing (that is, all pixels of the input HDR image Im_HDR (typically, one of a plurality of master HDR images of a video created by a video creator, and details of creation such as camera shooting and shading, or offline color grading are understandable to those skilled in the art and are omitted here as they do not improve this description) are sequentially processed) is first shown, and then a video processing circuit operating on the entire image is shown.
[0010] Here, it is assumed that we start with the HDR pixel luminance \(L_{HDR}\) (note that in some systems, work may already start from luma), which is transmitted through the selected HDR inverse EOTF in the luma conversion circuit 101, and the corresponding HDR luma \(Y_{HDR}\) is obtained. For example, a perceptual quantizer EOTF may be used. This input HDR luma is luma mapped by the luma mapper 102, and the corresponding SDR luma \(Y_{SDR}\) is obtained. In this unit, the creating device applies appropriate color changes, which includes performing a luminance regrading suitable for any specific scene image in the output image corresponding to the color details of the input image (e.g., maintaining an appropriate average brightness for a cave image). That is, there is an input connection UI for obtaining an appropriately determined shape of the luminance mapping function (LMF). For example, when creating an output image with a low maximum luminance for a cave image, the normalized version of the luma mapping function may need to boost the darkest luma of the input by a factor of 2.0 determined by, for example, a human color grader or an automaton. This means that the luma mapping function is concave, i.e., a so-called r-shaped (as shown within the rectangle representing unit 102).
[0011] Broadly speaking, there are two classes. An offline system may use color grading software to determine an optimal LMF according to his or her artistic preferences by a human color grader. Without limitation, assume that the LMF is defined as a LUT defined using the coordinates of a few node points (e.g., at the first node point \((x1,y1)\)). The human grader may set the slope of the first line segment, i.e., the position of the first node point, for reasons such as when there is dark content in the image and sufficient visibility is desired when displayed on a display with a low dynamic range (e.g., specifically, a 100 nit \(M_{L_C}\) image for a 100 nit \(M_{L_D}\) LDR display).
[0012] The second class uses automata to obtain a secondary grading from primary gradings with different maximum brightnesses. These automata analyze the image to propose an optimal LMF (for example, using a neural network trained with an image aspect of a series of training aspects to generate normalized coefficients of some parametric function in the output layer, or any deterministic algorithm may be used). A particularly interesting automaton, the so-called inverse tone mapping (ITM), analyzes the input LDR image on the creation side rather than the master HDR image and creates a pseudo-HDR image of this LDR image. This is very convenient. Because most videos are displayed as LDR and may be generated as SDR currently or in the near future (at least, for example, some cameras in multi-camera production may output SDR (for example, drones shooting side videos of sports games), and it may be necessary to convert this side video to the HDR format of the main program). By appropriately combining the functions of the mode LDR coding system with the ITM system, the applicant of the present application and its partners were able to define a system capable of double inversion. That is, the upgrade function of the pseudo-HDR image generated from the original LDR input is a substantially inverse function of the LMF used when coding the LDR communication proxy. By doing so, a system can be created that not only communicates the original LDR image but also conveys information for creating a suitable HDR image (automatically according to the wishes of the customer creating the content, or in other versions with human input (for example, adjustment of automatic settings)). The automaton can use any kind of rules (for example, the position of the light source in the image can be determined, which requires a certain relative brightness compared to the average brightness of the image), but the exact details are not relevant to the description of the present invention. What is important is that any system adopting or cooperating with the innovative embodiment of the present invention can generate some luminance (or lum) mapping function LMF. The lum mapping function determined by the automaton can also be input via the connection UI and applied to the lum mapper 102.Note that in the description of the simplest embodiment, there is only one (down - graded) lumamapper 102. This is not necessarily a limitation. Both the EOTF and the lumamapping usually map the normalized input domain [0,1] to the normalized output domain [0,1], so there may be one or more intermediate normalization mappings that map (substantially) 0 input to 0 output and 1 to 1. In such a case, the former intermediate lumamapping functions as a basic mapping, and the (second) lumamapper 102 functions as a correction mapping based on the first mapping.
[0013] The encoder obtained a set of LDR image luma Y_SDR corresponding to the HDR image luma Y_HDR (through the mapping function). For example, the darkest pixel in the scene can be defined to be displayed with substantially the same luminance on both the HDR display and the SDR display, but brighter HDR luminances can be pushed within the upper range of the SDR image. This is shown by the convex shape of the LMF function with decreasing slope (or converging back towards the diagonal of the normalized axis system) shown within the lumamapper 102. A person skilled in the art can easily understand the normalization by simply dividing the luma code by power(2; number_of_bits). Optionally, the pixel luminance can be normalized by dividing by the maximum value ML_C of the associated target display (e.g., 4000 nit), so that the normalized luminance is defined. That is, by driving up to the maximum luminance of the target display, which is the upper limit of the luminance dynamic range, the brightest pixel that can fully utilize the function of the associated target display functions as the input brightness value shown on the horizontal axis of the graph like unit 102 and is represented as a normalized luminance of 1.0. A person skilled in the art can understand the method of representing various normalizations to various maximum values and the method of specifying the function that maps the normalized input value to the normalized luma (any bit length and any selected EOTF) or any normalized luminance (any associated maximum luminance) on the vertical axis.
[0014] Thus, one can envision examples of indoor-outdoor scenes. In the real world, since outdoor luminance is typically 100 times higher than that of indoor pixels, conventional LDR images display indoor objects well, brightly and vividly colored, but everything outside the window is strongly clipped to uniform white (i.e., becomes invisible). Here, when communicating HDR video using reversible proxy images, the bright external areas visible through the window are made brighter (and in some cases, less saturated), but in a controlled way such that sufficient information for reconstruction into HDR is ensured. This has advantages for both outputs. Because systems that only wish to use the LDR image as is display excellent renditions of outdoor scenes within the limited LDR dynamic range possible.
[0015] Thus, the set of Y_SDR pixel luminances (along with chromaticities which are not necessary for this description in detail) forms a "conventional LDR image". That is, subsequent circuitry does not always need to consider whether this LDR image was cleverly generated or simply directly acquired from a camera like a conventional LDR system. Thus, video compressor 103 applies an algorithm such as MPEG HEVC or VVC compression. This is a bundle of data reduction techniques that, among other things, uses the discrete cosine transform to convert, for example, 8x8 pixel blocks into a limited set of spatial frequencies with less information required for representation. The amount of information required can be adjusted by determining quantization coefficients. The quantization coefficients determine the number of DCT frequencies retained and how accurately they are represented. The drawback is that the compressed LDR image (Im_C) is not as accurate as the input SDR image (Im_SDR), especially block artifacts occur. Depending on the broadcaster's choice, the situation can be very serious, for example, some blocks of the sky may be represented only by the average luminance and displayed as uniform rectangles. Usually, this is not a problem. Because the compressor determines all settings (including quantization coefficients) such that the quantization error is hardly visible or at least not problematic to the human visual system.
[0016] Formatter 104 performs all the signal formatting required for the communication channel (which may be different, for example, when stored and communicated on a Blu-ray disc and when, for example, in the case of DVB-T broadcast). Generally, all variations have the characteristic that the compressed video image Im_C is incorporated into the output image signal S_im using a luminance mapping function LMF (which may or may not vary from image to image).
[0017] De-formatter 151 performs the reverse process of formatting, obtains the compressed LDR image and the function LMF, and may reconstruct the HDR image or other useful dynamic range mapping processes may be performed in subsequent circuits (for example, optimization of a specific connected display). Decompressor 152, for example, releases VVC or VP9 compression and obtains a sequence of approximate LDR luma Ya_SDR that is sent to the HDR image reconstruction pipeline. To perform the substantially reverse process of coding the communicated HDR video as an LDR proxy video, upgrade luma mapper 153 converts the SDR luma to the reconstructed HDR luma YR_HDR (which uses an inverse luma mapping function ILMF that is (substantially) the inverse of the LMF). One explanatory diagram shows two potential receiving devices (150), which may exist as dual functions in one physical device (for example, the end user can select which parallel processing to apply), or some devices may have only one of the parallel processes (for example, some set-top boxes may only perform the reconstruction of the master HDR image and store it in a memory 155 such as a hard disk).
[0018] When a display panel (e.g., 750 nit ML_D end-user display 190) is connected to an embodiment of the receiver, the receiver may have a display adaptation circuit 180 that calculates an output image of 750 nit instead of a reconstructed image of, for example, 4000 nit (this is shown by the dotted line, the reason being that although both technologies are often used in combination, this is to show that this is an optional component not related to the teachings of the present invention). Although not described in detail for the many variations that can achieve display adaptation, typically there is a function determination circuit 157 that proposes an adapted luma mapping function F_ADAP based on the inverse shape of the LMF (when the difference between the maximum luminance of the starting image and the maximum luminance of the target image is smaller than the maximum luminance difference between two reference gradings (e.g., a 4000 nit master HDR image and a corresponding 100 nit LDR image), it will typically be closer to the diagonal). This function is loaded into the display adaptation luminance mapper 156 and typically has a smaller dynamic range (an HDR luma L_MDR with less boost that ends at ML_D = 750 nit instead of ML_C = 4000 nit is calculated. Re-grading means the luminance (or luma) mapping from a first image of a first dynamic range to a second image of a second dynamic range, and in the mapping, different absolute or relative luminances are given to at least some of the pixels (e.g., compressing the brightest luma into a smaller partial range to reduce the dynamic range of the output). Such a graded image may also be called a grading. If only reconstruction is required, the EOTF conversion circuit 154 is sufficient, and a reconstructed (or reconstructed) HDR image Im_R_HDR including pixel colors including the reconstructed HDR luminance LR_HDR is generated.
[0019] When video content consists only of natural images (e.g., those captured by a camera), considering various methods of calculating pixel luminance re-graded with respect to the input luminance, high dynamic range technology is already complex. The complexity increases depending on the details of various coding standards, and these may be mixed. For example, in a picture-in-picture, it is necessary to mix HDR10 video data with HLG capture.
[0020] Furthermore, usually, content creators want to add graphics to the image. The term "graphics" refers to things that are not natural image elements / regions, that is, things that usually have a simpler nature and usually have a more limited set of discrete colors (e.g., without photon noise). Another way to characterize graphics is that they do not form part of an image that is illuminated in the same way, although other considerations such as readability or drawing attention may be important. Graphics are usually computer-generated and usually mean graphics that are not natural. For example, CG furniture that is visually indistinguishable from actual furniture within a scene captured by a camera. For example, there may be graphics elements such as a company's bright logo, or a plot, weather map, information banner, etc.
[0021] Graphics began by mimicking how graphics elements were created on paper in the old LDR era. That is, there were usually only a limited set, such as pens or crayons. In fact, in initial systems such as Teletext, initially only three primary colors (red, green, blue), two secondary colors (yellow Ye, cyan Cy, magenta Mg), and black and white were defined. Graphics such as the information page about scheduled flights at an airport had to be generated from graphic elements (e.g., pixel blocks) of pixels with one of these 8 colors. However, "red" (although it had to look reddish to a human viewer) was not a uniquely defined color and, in particular, did not have a uniquely defined luminance pre-defined. More advanced color palettes can include, for example, 256 pre-defined colors. For example, a graphical user interface such as Unix's X-windows system defines a number of selectable colors in X11 such as "honeydew", "gainsboro", or "lemon chiffon". These colors need to be defined as representative codes that produce approximately a particular color appearance when displayed and are defined in the LDR color gamut as the percentages of standard red, green, and blue that are mixed. For example, honeydew is composed of 94% red, 100% green, and 94% blue (or hex code #F0FFF0) and is displayed as a light greenish white.
[0022] The problem is that while these colors are stably defined in the limited (and universal) LDR color gamut, they are not defined, or at least are ambiguously defined, in HDR images. There are several reasons for this. First, since there is only one type of white in LDR, all graphics colors can be defined in comparison to that white. In fact, similar to drawing with color markers of appropriate colors on the white of a display that functions as a white canvas, the code that mimics the absorption of a specific hue defines the creation of a specific chroma (or color nuance, such as a desaturated Chartreuse). However, there is no unique white in HDR. This is understandable from the real world. When looking at a wall painted white indoors, humans receive the visual impression that it looks white overall, although there may be a slight gray in the shadow areas. However, when looking at a white garage door outdoors with sunlight shining through the window, this also appears to be a color of the "white" type. However, this white is a much brighter white of a different kind. In high-quality HDR processing systems (processing usually includes aspects such as optimal creation, encoding, and processing for proper display), it may be desirable not only to extend the LDR look in some fine HDR aspects but also to render various types of white on the display (although usually at lower luminance than the real world). Technically, such considerations have led to a useful definition of a framework that can define various types of HDR images with different coded maximum luminances ML_C, such as 1000 nit or 5000 nit, for example.
[0023] What further complicates the issue is that there are different types of HDR images (with different ML_Cs; even relative systems generally have the same problem), and due to luma mapping, it is necessary to map the luminance along the dynamic range of the input image to the luminance along the output dynamic range (e.g., of an end - consumer display). Luminance mapping is typically defined and calculated as a corresponding luma mapping function, which maps a luma code that uniquely represents luminance, for example, via a perceptual quantizer EOTF. For example, assume it is necessary to use the function F_H2S_L to map the luminance of an input image from 0 to 4000 to an output luminance from 0 to 100 (i.e., L_out = F_H2S_L(L_in)). Since the perceptual quantizer EOTF (EOTF_PQ) can encode luminance up to 10,000 nits, both the input luminance and the output luminance can be represented as PQ - defined luma. Y_out = OETF_PQ(L_out) and L_in = OETF_PQ(L_in) are required. Therefore, the luma mapping F_H2S_Y is as follows. Y_out = F_H2S_Y(Y_in)=F_H2S_Y(OETF_PQ(L_in)). Or, the luminance mapping function and the luma mapping function are related by the following equation, L_out = EOTF_PQ(Y_out)=EOTF_PQ[F_H2S_Y(OETF_PQ(L_in))], so F_H2S_L = EOTF_PQ(o)F_H2S_Y(o)OETF_PQ. Here, (o) means function composition (usually denoted by a small circle symbol). It is also possible to define a mapping to a luma specified by another EOTF (e.g., a mapping from an input PQ luma to an output Rec.709 SDR luma).
[0024] Referring to FIG. 2, a possible graphics insertion pipeline is described.
[0025] Figure 2 shows an HDR image communication and processing pipeline that illustrates a plurality of similar video communication systems. This is a sports program (horse racing) for classical broadcast services. One or more cameras 201 capture a sports event, and the plurality of captured images are mixed at a production booth 202. At the production booth 202, feeds from various cameras can be mixed to produce an overall HDR video that is graded (or shaded) according to preference. The method of specifying the pixel luma of consecutive video images is outside the scope of the present invention. The broadcaster creates an original mixed HDR image 205 composed of native video content 206 (the feed from the camera after appropriate shading) and the broadcaster's own graphics (an example of primary graphics). This graphics can be various depending on the application, but in this example, it is a list of the names of the two horses at the head of the race. This original graphics pre-mixed and communicated within the video image is called a primary graphics element 207. This is distributed to a local distributor's device 210, for example, by a satellite antenna 203. This is, for example, a Dutch broadcaster or re-distributor that distributes to consumers via terrestrial broadcast or cable 219. The local distributor may add its own secondary graphics, for example, the broadcaster's logo ("NL6") inside the sun. For example, assume that all pixel colors of the primary graphics are white, and the pixel colors of the secondary graphics are white-dependent, for example, a yellow sun with a white luminance of 90% and a red text color with a brightness of 60% inside it. This broadcaster may also mix other types of secondary graphics, such as a teaser 217 for a later sports program. This further mixed video (215) (secondary mixed image / video) is further distributed to the end user via a telecommunications cable (CATV) 219. The end customer may have a set-top box (or a computer, etc.) 220 that can mix tertiary graphics, etc.In this example, subtitle information is communicated together with a video signal (usually pure text information that needs to be rendered within video pixels that make up the generated video elements). The video signal is rendered as tertiary graphics 226 by a set-top box, mixed back into the video image to generate tertiary video 225, and communicated to a consumer display 230 via an HDMI (registered trademark) cable 229 or the like. In practice, instead of (or in addition to) a set-top box that may not exist in some processing pipelines, the end display may insert its own adjusted graphics.
[0026] In this text (when showing certain graphics at a higher conceptual level), all additional graphics (in this example, secondary and tertiary graphics) will be referred to as "secondary graphics", distinguished from the primary graphics which is the previous graphics situation (usually the first graphics inserted into the video by the original creator of the video). Generally, the image itself (i.e., a matrix of pixel colors, e.g., YCbCr) Im_HDR is communicated together with metadata (MET) within an HDR image or video signal S_im. This metadata can code several things in various HDR video coding variations, for example, the coded maximum luminance of the HDR image (e.g., ML_C_HDR = 1000 nit), and often, a luma mapping function for calculating the secondary color of the re-graded image, etc.
[0027] It is undesirable for various graphics not to be adjusted across the entire luminance dynamic range, or even worse, for the luminance to change over time (e.g., when compared to each other). For example, as shown by the 950 nit ML_D luminance dynamic range of the end consumer display 230, it is desirable to have the same average luminance. By doing so, for example, the technical functions related to the rendering of secondary graphics of a TV display (or other devices that receive a primary video graphics mix) can be improved. There is a simple system that pre - calculates a video image adapted to the display at optimal pixel luminance, and it may not be very difficult or important to do this if only adding primary graphics to its pre - established range, but the situation can become more complex, especially when adding different types of graphics by different devices at different locations in an HDR video processing pipeline.
[0028] US20180018932 teaches some techniques for mixing primary graphics elements (do not mix secondary graphics if the primary graphics are already mixed). Also, the appropriate location of the graphics portion range of the primary graphics within the master HDR video is not established (the maximum value of the graphics is simply mapped to the maximum value of the video, i.e., the entire range).
[0029] The most complex variation to understand, the video priority mode described in US’932 - Figure 2A, is summarized in this application using Figure 8. When it is desired to display stable graphics with the output of an image adapted to the end - user's display (for example, when it is desired to display an image created for the maximum luminance ML_endUSR_disp = 750 of the end - user's display), this can be achieved, for example, by giving the same value to all pixel luminances of a single - color white text (TXT) over a plurality of consecutive images (for example, several scenes, or as long as the graphics are presented in the case of user - interface graphics). Assume that the appropriate final value of the text luminance is the value at which the video cloud object is projected within the display dynamic range. Usually, the dynamic range of the graphics (especially the maximum luminance ML_gra of the graphics) is not the same as the input video (the maximum luminance ML_vid of the video) with which the graphics should be mixed. Also, since each user may have purchased various displays, it is usually different from the maximum luminance of the end - user's display as well. In fact, associating graphics with some dynamic range is already a somewhat advanced HDR technology. Because graphics in the LDR era usually had a relative 3x8 - bit code (for example, 255 / 255 / 255). The problem is that in the dynamic metadata that conveys different - shaped luminance mapping functions for consecutive image shots applied by the end - user's display, it is not guaranteed what the pre - mixed text will do. When mixing the original text (TXT_OR) with the video at the original luminance, for example, 390 nits, the display may increase the luminance with one mapping function and decrease the luminance with another mapping function, that is, flicker may occur. However, if the display knows exactly the dynamic mapping function (for example, the first dynamic mapping function F_dyn1) that it applies to the video, regardless of whether graphics are mixed in the video, this can be corrected in advance.In that case, the graphics can be accurately pre-mapped (using the first pre-correction function F_precomp1) to the luminance required for an HDR image transmitted from, for example, a set-top box to a display, whereby F_dyn1 can map it to the desired final luminance (the same brightness as the cloud) of the displayed image. In the next scene, the piece of paper can be up-mapped to the same final luminance for TXT by the second dynamic mapping function F_dyn2, in which case the second pre-correction function F_precomp2 is used. Usually, the function applied by the display depends on the maximum luminance ML_endUSR_disp and a dynamic time-varying reference luminance mapping function that specifies how a video luminance of 4000 nits should be mapped to, for example, a reference image of 100 nits. By a pre-fixed algorithm, this reference luminance mapping function is transformed into the final display adaptation luminance mapping function. This function is used by the TV for subsequent images until a new dynamically changing function is input as metadata. However, if the set-top box complies with the codec, i.e., knows the reference luminance mapping function and the algorithm to convert it to the final display adaptation luminance mapping function for any value of ML_endUSR_disp and polls ML_endUSR_disp from the connected display, any such pre-correction can be performed.
[0030] In the prior art, several other examples of mixing primary graphics are shown, but these are simpler.
[0031] The graphics priority mode shown in US’932 - Figure 2B simply maps the graphics maximum value to the display maximum value. And the video can be down - mapped to the range that already exists within the set - top box. When an optimal display - adapted image that is all prepared within the STB and requires no further processing is transmitted to the display via HDMI (registered trademark), the graphics mixing is relatively simple. The STB has everything at hand and no variable mapping is performed later. However, these two options are not always available. For example, some codec technologies are expensive and the STB is sold at a very small profit, so the STB may not permit such a codec. In any case, displays such as consumer TVs need to perform dynamic HDR luminance mapping. Pre - correction is not always possible either. For example, the TV may reject communicating its ML_endUSR_disp value to the STB. These methods may be useful for home appliances such as STBs and TVs, or computers and monitors, but there may be more places where graphics need to be inserted in the video communication pipeline from the “camera” to the final consumer display (e.g., within a shopping mall).
[0032] If the TV can perform graphics blending (as shown in US’932 - Figure 2D), things can be relatively simple. The TV applies the normal display - adapted luminance mapping function to the video and then can place the graphics at the desired luminance, such as the output luminance when the cloud ends. This may be okay in the case of pass - through DVB subtitles with standardized coding, but in that case, the STB requires a mechanism to communicate the graphics of its UI graphics elements via the HDMI (registered trademark) interface.
[0033] US2020193935 statically switches the display mapping during graphics mixing. As a result, the graphics are always placed at the same (shifted) luminance position, which has both advantages and disadvantages.
[0034] In the field of recent high dynamic range video / TV technology, it is clear that excellent graphics processing technology is still needed.
Summary of the Invention
[0035] The problems existing in a simple approach to graphics mixing / insertion are addressed by a method for determining a second luma of pixels of a secondary graphics image element (216) that is mixed with at least one high dynamic range input image (206) in a circuit for processing digital images. The method includes receiving a high dynamic range image signal (S_im) including at least one high dynamic range input image, and the method further includes analyzing the high dynamic range image signal to determine a range (R_gra) of primary graphics luma of primary graphics elements of the at least one high dynamic range input image, wherein the range of primary graphics luma is a partial range of the luminance range of the high dynamic range input image, and determining the range of primary graphics luma includes determining a lower luma (Y_low) and an upper luma (Y_high) that specify endpoints of the range (R_gra) of primary graphics luma, luminance mapping the secondary graphics element using at least the brightest subset of pixel luma of the secondary graphics element included in the range of graphics luma, mixing the secondary graphics image element with the high dynamic range input image.
[0036] The luminance mapping to the graphics partial range within the input (video + primary graphics mix) HDR image, or within the output variation image derivable by mapping luma using some luma mapping function, may in principle be performed when creating (i.e., defining) the pixels of the graphics. However, generally, even when the graphics are not read from the storage where it is stored in a pre-defined state and are generated by any method or apparatus embodiment, the graphics are generated in some different format (e.g., LDR format, or a format with a maximum luminance of 1000 nits), but the luma (or any code, e.g., language code such as sRGB) of the pixels constituting the secondary graphics is mapped to an appropriate position within the established graphics range (or slightly outside, usually above the lower / darker side end), where it becomes the corresponding secondary luma of various graphics pixels.
[0037] Mixing often simply replaces the original video pixels with graphics pixels (however, with adjusted colors, especially adjusted luma). This method works even for more advanced mixing, especially when more graphics are retained and less graphics are mixed. For example, linear weighting called blending may be used, in which case a certain percentage of the determined graphics luminance (Y_gra) is mixed with the complementary percentage of the video. Y_out = alpha * Y_gra+(1 - alpha)*Y_HDR [Equation 1]
[0038] Since alpha is usually higher than 0.5 (or 50%), basically most of the pixels contain graphics and the graphics can be seen well. In fact, ideally / preferably, such mixing occurs with the derived luminance itself rather than with luma (especially in the case of a highly non-linear luma definition such as PQ). L_out = alpha * L_gra+(1 - alpha)*L_HDR [Equation 2]
[0039] In an embodiment where a proxy SDR image is transmitted for an HDR image, such HDR video luminance L_HDR may be obtained, for example, by re-converting the SDR luma to HDR reconstruction luminance using the maximum luminance ML_C transmitted together with the metadata of the transmitted HDR representation. Optionally, at the time of acquisition, this luminance may be downscaled to the corresponding luminance within the dynamic range of the display using display adaptation. In fact, from FIG. 3, it can be seen that an advantage of the approach of constructing at least such a luma mapping function (or corresponding luminance mapping function) is that graphics can be mapped / blended in both the input range and the output range because the curve maps the graphics range in a stable manner. More generally, this technique can utilize the fact that the content creator would have determined or at least approved the appropriate position of the graphics with respect to the primary graphics. Thus, if appropriately adjusted in relation to the primary graphics, the secondary graphics can also include appropriate luminance.
[0040] For example, if no specific information regarding such primary graphics is available because the video creator does not want to expend the effort to code and transmit it, the method or apparatus of the present invention analyzes the situation of the input HDR video signal S_im and, depending on various embodiments, can determine an appropriate range R_gra for locating the luma (or luminance) of at least most of the pixel colors within the secondary graphics elements.
[0041] How exactly this is done depends on the nature of the graphics elements.
[0042] If there is secondary graphics (e.g., subtitles) that has only one or a few colors (e.g., the original colors before luminance mapping for adjustment), any lumas within the graphics range can be used to render the secondary graphics. In more complex images, the darkness of black can be restricted (e.g., 10% or more of the luminance of white), or some of the dark colors can be set outside / under the range of the primary graphics (however, considering the lumas of bright colors, the effect of stable graphics display is mostly maintained). For example, a slightly changing black surrounding rectangle at the top of an HDR video may not be very bothersome if the color of the subtitles is adjusted and stable with respect to the luminance of the video objects in one or more subsequent scenes. It is also conceivable that one may want to avoid large fluctuations in the brightness of relatively large dark colors that are clearly below average. If there is a problem, it can be processed individually, for example, by lightening the dark colors to approach the lower limit point of the appropriate graphics range (i.e., slightly above or below the minimum luminance). It is desirable to maintain the brightest color of the secondary graphics within the graphics range R_gra. For example, the brighter partial range is defined as all colors that are brighter than 50% of the luminance of the secondary graphics elements (this percentage may depend on the type of additional graphics. For example, menu items may not require as much precision as some logos). In many cases, the brighter partial range contains enough colors to (geometrically) represent most of the secondary graphics elements, so that, for example, its shape can be recognized.
[0043] It is not necessary to fully know the entire luma distribution of the primary graphics. It is sufficient to roughly know at least the typical luma of the brightest color within the primary graphics. To analyze the HDR video signal, the situation of its image, and the luminance or luma of its object pixels, multiple approaches can be used alone or in combination. In the latter case, based on the input of one or more circuits to which the analysis method is applied, the final best graphics range is determined by some algorithm or circuit (if there is only one circuit / method in any device or method, such a final integration step or circuit is not necessary).
[0044] Preferably, a method for analyzing a high-dynamic-range image signal includes detecting one or more primary graphics elements in the high-dynamic-range input image, establishing the luma of the pixels of the one or more primary graphics elements, and summarizing the luma as a lower-limit luma and an upper-limit luma, where the lower-limit luma is lower than all or most of the luma of the pixels of the one or more primary graphics elements, and the upper-limit luma is higher than all or most of the luma of the pixels of the one or more primary graphics elements.
[0045] If no further information is available, or if there is no reliable information, analyzing the image itself may be a sure option. For example, this algorithm (or circuit) may detect that there are two primary graphics elements in the "original" version of the video content. Suppose there are subtitles displayed in three colors (white, light yellow, light blue) for three speakers, and a logo displayed in multiple colors (e.g., some primary colors and black with a luminance of 5% compared to the white of the subtitle or the logo itself). In many cases, it is expected that the white of the logo and the white of the subtitle (or at least the luminance of the white expected by extrapolating from, for example, the yellow pixels of the logo) will be the same or approximately the same. In particular, even if the two are not very different (e.g., the subtitle is five times brighter than the white of the logo, or vice versa), in the method of the present invention, by treating the darker white as if it were "gray" based on the brighter white, the graphics range R_gr can be defined relatively easily. That is, the upper part of the primary graphics range R_gra may be the brightest pre-mixed graphics white, yet all the graphics can be identified as one range, and it is not a case of applying twice the identified difference range (in which case, for example, the secondary graphics can be mixed in either range or the preferred range). If the brightness (luminance of white) of both graphics elements is significantly different, the brighter primary graphics element may be retained in the graphics analysis, the darker primary graphics element may be discarded, and R_gra may be calculated based on the brightest primary element. This embodiment may be useful when the lumamapping function does not have a safe graphics range as described in FIGS. 3 and 4, for example, does not have a special area, and is simply a simple power function that functions as a continuous compression function.
[0046] Preferably, the method includes reading from the metadata within the high-dynamic range image signal the values of the lower luma (Y_low) and / or the upper luma (Y_high) written to the metadata of the high-dynamic range image signal. The creator of at least one high-dynamic range input image can use this mechanism to convey what he (or an automaton) considered to be the graphics range suitable for this image or a shot of this video image (e.g., adjust with the bright pixels of an explosion nearby that should not be recognized as tiny by the human visual system in view of the recognition of the white of nearby subtitles). The creator sets the primary graphics to such luminance. Note that the human visual system can make all kinds of estimations about what is shown in the image (e.g., local illumination), which can be very complex, but the technical mechanism of the present disclosure provides a simple solution for those who should know about it (e.g., human creators of videos, color graders, post-producers). For example, the range transmitted may be slightly darker than the luminance actually used in the primary graphics, thereby ensuring adjustment of the primary graphics and higher brightness. In this case, the integration step or circuit can simply discard or turn off other functions, such as the analysis of the HDR image itself. Alternatively, any available mechanism may be used.
[0047] Preferably, the method includes analyzing a high-dynamic-range image signal, the analysis including reading two or more luma mapping functions associated with temporally consecutive images from metadata within the high-dynamic-range image signal, and establishing the range of the graphics luma as a range that satisfies a first condition with respect to the two or more mapping functions, the two or more luma mapping functions mapping an input luma within the range of the graphics to a corresponding output luma, the first condition being that for each input luma, the two or more corresponding output lumas obtained by applying each of the mapping functions from the two or more luma mapping functions to the input luma are substantially the same. FIG. 7 shows an example having a sub-range of the input range such that the two functions are (substantially) identical in that range, i.e., they create (nearly) the same output value for any input value within that range. This is detectable because the content creator intended to create a graphics robust dynamic luminance mapping function.
[0048] Accordingly, not only is the range similarly mapped to a stable output range, but also the various lumas within it are mapped by both functions to substantially the same output luminance (except for a slight deviation that is usually not very noticeable, e.g., within 2% of the difference noticeable to the average human vision or within 10% if some flicker is tolerated). It may also be desirable to supplement such analysis with further analysis.
[0049] Specifically, this may also include verification of a second condition that the primary graphics luminaire determined based on the function is within the secondary range of the luminaire suitable for graphics. Just finding areas with various functions for at least a series of shots (e.g., an indoor shot followed by an outdoor shot) does not necessarily mean that it is always desirable to place graphics there. For example, if only the darkest 10% of the entire range (input or output) is identified, generally an undesirable dark graphics conclusion is obtained. For example, 1 / 2 to 1 / 10 of a high brightness dynamic range (such as a maximum luminance of 4000 nits or more) may be a suitable position for placing subtitles or other graphics. However, in the case of 4000 nits or more, 1 / 2 may not be optimal. This generates the maximum brightness of very bright graphics. For graphics such as text, the brightest color is usually normal (diffused) white, in contrast to more general graphics where there may be several whites that make up super white. The 2000 nit level may be considered high. However, the creator of the mapping function may consider it optimal to set a stable graphics range, for example, to 0.8 * 0.5 to 1.2 * 0.5 in the case of a 5000 nit master grading. When using this curve to downscale the output to a maximum output of 1000 nits, a 500 nit graphics can be considered reasonable, i.e., meeting the compliance criteria, even if it is a subtitle. If the second criterion is not met, the graphics mixing device may, for example, attempt to approach the identified range but determine to mix graphics with a luminaire lower than the lowest value of the identified function shape-dependent (i.e., there is a partial range where two or more functions project the same input partial range to the same output partial range) graphics range.Various secondary criteria can be used. For example, a criterion such as what percentage of the maximum value of the output range is the upper limit of the stable graphics range identified by the shape of the function (i.e., it also depends on the absolute maximum value, i.e., the higher the maximum value of the output range, the lower the allowable percentage may be. For example, in the case of a range of 1000 nits, the upper limit is 500 nits, and when the maximum value of the output range is 4000 nits or more, the upper limit is 1200 nits), and / or a criterion such as the height of the absolute luminance value at which the upper limit of this output range is displayed can be used. Therefore, usually, regarding the validity of the identification of the range determined by the function, the values along the range of the input lumens (i.e., the received image) are checked. However, regarding appropriateness, the output range and the maximum luminance of the assumed output range (i.e., what is generated by transcoding mixed with graphics, for example) may also be involved, and it may be difficult to determine what is appropriate only from the input lumen values. For example, original graphics brighter than 2000 nits existing as pre-mixed graphics in the input image will usually have a maximum output of less than 1000 nits. This is, for example, when using a normalization soft-clipping mapping function starting from a maximum input of 4000 nits, the output will be 900 nits (this is a rather bright subtitle compared to the HDR effect of the brightest video object, which ideally should be more impressive than the white level of the graphics, but the 900 nit graphics may still be acceptable in some cases and to some users, so at least it may not be a negligible level in itself). An example of an appropriate secondary criterion is that the maximum value of the secondary range must be lower than a pre-fixed percentage of the maximum value of the range of the mixed image and graphics.
[0050] Any device can be programmed using predictable graphics limitations. For example, it can be expected that it is the same as, or somewhat higher than, the expected partial range of normal ("LDR") image objects within an HDR image. In the case of a 1000 nit ML_C defined HDR image, the white of the graphics typically falls within the range of 80 nit to 350 nit, and the black can be expected to be lower. This also depends on whether there are multiple primary graphics elements or whether their luminance (or luma) characteristics are the same.
[0051] Therefore, the verification embodiment is as follows. If it is determined by analyzing the luma mapping curve that the upper limit luma of the graphics is, for example, 150 nit, this certainly falls within the range of the expected luma values for bright graphics pixels and also within a wider range [80, 350]. Similar considerations can be made for the lower level, but as mentioned above, the importance of the lower level is not very high. For example, it can simply be set to 10% of the determined upper limit luminance (or the corresponding x% of the upper luma when examined on the applicable EOTF for pixel color coding). For example, if the luminance of the upper limit luma is estimated to be 950 nit in this way, ideally such overly bright graphics are not desirable, so this may be an accident of the analysis. However, it is also possible that the content provider actually created such bright subtitles. In such a situation, rather than simply rejecting the value and concluding that the graphics range could not be determined with sufficient certainty, further analysis such as checking the graphics luma actually sent together with the video signal by the creator can be performed, or such luma values can be searched for within the video signal, and further image analysis can be attempted, such as how much those pixel sets are combined and how large they are.
[0052] Note that in HDR images, it should be noted that primary graphics do not necessarily have to actually exist at their respective appropriate luma positions (for example, in the first shot of an image with a first luma mapping curve adjusted to map a basically bright HDR scene, there may actually be graphics elements included, but in subsequent second shots with different luma mapping curves for mapping normal image objects and bright image objects, there may be no actual graphics inserted, and if they are inserted, they are set to approximately the same luma value).
[0053] Preferably, the method includes the step of luma mapping the pixels' luma of at least one high dynamic range input image to the corresponding output luma of at least one corresponding output image, and the mixing is performed within at least one corresponding output image according to at least one luma mapping function obtained from the metadata of the high dynamic range image signal.
[0054] This relationship can potentially be mixed in both the input and output lumadomains, especially when having a function as shown in Figure 3. Whether to perform the mixing in the input domain or the output domain may not always be the same (in some embodiments corresponding to some HDR codecs, it may be convenient to know where in the input lumadomain to mix). In any case, video pixels are often mapped to the output domain by the lumamapping function received in the metadata. This method or apparatus can use this mapping function to determine an appropriate lumaposition (e.g., ML_C of up to 350 nits) within the output domain for mixing the output determination graphics. This creates flexibility regarding which device can mix what and when. In particular, using stable subregion functions makes it easier to understand what happens to the already executed graphics mixing even if those functions are distorted in a display adaptation scenario (i.e., when the shape of the original function is made closer to the shape of a quadratic function, its stable partial range expands or compresses, but if the adaptation is correctly executed and both the linear function and the reference function respect that partial range, the same adjusted mapping characteristics for the lumas within the range are maintained).
[0055] Preferably, the method includes designing secondary graphics image elements using a set of colors across a scale of colors of different brightnesses, some of the colors of different brightnesses appear brighter than average to the human eye, some appear darker than average, and at least the colors brighter than average are mixed into at least one high - dynamic range input image using the lumas within the range of the graphics luma.
[0056] A color scale is a set of colors in which the lightness increases (or decreases), for example, changing from a dark color to a light color and, in many cases, to intermediate lightness colors. This scale does not necessarily have to include all steps of a particular primary color (e.g., blue), but a limited color palette may have, for example, a few light blue steps and a few dark green colors but no navy blue (in which case the green color defines the dark steps of the scale). By visually checking the approximate average lightness of the graphics elements, some colors usually appear dark and some appear light. About 50% lightness, or about 25% luminance, can be typical values that are seen as or used as the midpoint. Thus, pixels with a luminance of less than, for example, 25% may be displayed as dark colors and no longer need to meet the criteria within the graphics range. If more colors are needed within the graphics range in a more rigorous system, the secondary graphics can be designed with colors of a lower degree of darkness.
[0057] Some embodiments of the secondary graphics insertion of this method can use pixel replacement, which is a more predictable mixing method (i.e., drawing appropriate established Luma graphics pixels at the positions where the video pixels of the original video were), or blend by, for example, less than 50%, preferably 30% or less of the luminance of at least one high-dynamic range input image. As a result, the video appears somewhat transparent and is still visible, but the luminance of the graphics is dominant and most of the Luma determination is retained.
[0058] Preferably, the method of designing the set of colors of the secondary graphics elements involves selecting a limited set of relatively bright and darker-than-average colors, where the darker-than-average colors have a luminance higher than a secondary lower luminance, which is a certain percentage of the lower luminance, preferably higher than 70% of the lower luminance. When the graphics range R_gra is determined, appropriate secondary graphics colors can be determined. For example, in this method, the influence of large changes in luminance mapping can be reduced by ensuring that the determined safe lower limit is not exceeded too much, for example, not exceeding less than 30. This percentage may depend on aspects such as how many dark pixels are in the secondary graphics and where the dark pixels are located. For example, if there are only 5 dark pixels in a 100x100 pixel graphics, regardless of how it is finally displayed in the output image, the average perceived brightness of the secondary graphics is not overly affected by these few dark pixels.
[0059] The method may be implemented as a device. One example is a device (500) that determines the second luminance of the pixels of a secondary graphics image element (216) and mixes the secondary graphics image element with at least one high-dynamic-range input image (206) including a primary graphics element (207), and the device includes an input unit (501) that receives a high-dynamic-range image signal (S_im) including at least one high-dynamic-range input image, an image signal analysis circuit (510) that analyzes the high-dynamic-range image signal to determine a range that specifies the luminance of the pixels representing the primary graphics and determines the graphics luminance range (R_gra) of at least one high-dynamic-range input image, where the graphics luminance range is a partial range of the luminance range of the high-dynamic-range input image, and determining the graphics luminance range includes determining a lower luminance (Y_low) and an upper luminance (Y_high) that specify the endpoints of the primary graphics luminance range (R_gra), the image signal analysis circuit, A graphics generation circuit (520) that generates or reads secondary graphics elements and luminance-maps the secondary graphics elements using at least the brightest subset of the pixel luminances of the secondary graphics elements included in the range of the graphics lumina. An image mixer (530) that mixes the secondary graphics image elements with the high-dynamic range input image to generate pixels having a mixed luminance (Lmax_fi). An output unit (599) that outputs at least one mixed image (Im_out) including pixels having the mixed luminance (Lmax_fi).
[0060] Alternatively, the image signal analysis circuit (510) is a device including an image graphics analysis circuit (511), and the image graphics analysis circuit detects one or more primary graphics elements in the high-dynamic range input image, establishes the lumina of the pixels of the one or more primary graphics elements, summarizes the lumina as a lower-limit lumina (Y_low) and an upper-limit lumina (Y_high), where the lower-limit lumina is lower than all or most of the lumina of the pixels of the one or more primary graphics elements, and the upper-limit lumina is higher than all or most of the lumina of the pixels of the one or more primary graphics elements.
[0061] Alternatively, the image signal analysis circuit (510) is a device including a metadata extraction circuit (513) that reads, from the metadata of the high-dynamic range image signal, values of a lower-limit lumina (Y_low) and an upper-limit lumina (Y_high) written to the high-dynamic range image signal, for example, by the creator of at least one high-dynamic range input image.
[0062] Alternatively, the image signal analysis circuit includes a luma mapping function analysis unit, and the luma mapping function analysis unit reads two or more luma mapping functions of temporally consecutive images from metadata in a high-dynamic range image signal, and establishes the range of the graphics luma as a range that satisfies a first condition regarding the two or more mapping functions. The two or more luma mapping functions map an input luma within the range of the graphics to a corresponding output luma, and the first condition is that for each input luma, the corresponding two or more output lumas obtained by applying the respective mapping functions from the two or more luma mapping functions to the input luma are substantially the same. Device.
[0063] Alternatively, the image mixer includes a luma mapper, and the luma mapper maps the luma of the pixels of at least one high-dynamic range input image to the corresponding output luma of at least one corresponding output image. At least one corresponding output image has a dynamic range (e.g., maximum luminance) different from the dynamic range of at least one high-dynamic range input image, and the luma mapper further performs mixing within at least one corresponding output image according to at least one luma mapping function obtained from the metadata of the high-dynamic range image signal. Device.
[0064] Alternatively, the graphics generation circuit designs secondary graphics image elements using a set of colors across a scale of colors of different brightnesses, and some of the colors of different brightnesses appear brighter than average to the human eye, and some appear darker than average. At least the colors brighter than average are mixed into at least one high-dynamic range input image using the luma within the range of the graphics luma. Device.
[0065] Alternatively, the image mixer is a device that performs mixing by replacing the video pixels of at least one high-dynamic range input image with secondary graphics element pixels, or by blending at a percentage less than 50%, preferably 30% or less, of the luminance of at least one high-dynamic range input image.
[0066] In particular, those skilled in the art will understand that these technical elements can be embodied in various processing elements such as ASICs (application-specific integrated circuits, i.e., typically, the IC designer causes the method to be executed on a part of the IC), FPGAs, programmed processors, etc., and can exist in various consumer or non-consumer devices, whether or not they include a display (e.g., a mobile phone that encodes consumer video) or a non-display device that can be externally connected to a display. Also, those skilled in the art will understand that images and metadata can be communicated via various image communication technologies such as wireless broadcasts, cable-based communications, etc., and that these devices can be used in various image communication and / or usage ecosystems such as television broadcasts, internet-based on-demand, video surveillance systems, video-based communication systems, etc. Innovative coding HDR signals may be applicable to the various methods described above. For example, at least one lower and upper luma value of at least one primary (pre-mixed) graphics range can be communicated.
Brief Description of the Drawings
[0067] The above and other aspects of the method and apparatus according to the present invention will be described and clarified with reference to the implementations and embodiments described below, as well as the accompanying drawings. The accompanying drawings are merely non-limiting and specific diagrams that illustrate general concepts. Dashed lines are used to indicate that a component is optional, but components that are not dashed lines are not necessarily essential. Dashed lines may also be used to indicate elements hidden inside an object or intangible things such as, for example, the selection of an object / region, even though they are described as being essential.
[0068]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
[0069] FIG. 3 shows in more detail a method for optimally constructing an HDR video image including graphics. This figure shows two typical HDR scene images (e.g., temporally adjacent image shots), and for these images, it may be desirable to specifically grade the pixel luminance to obtain an impressive HDR effect. In particular, while an HDR image enables the creation of bright and highly saturated pixel colors, in LDR, it is necessary to reduce the saturation of bright colors, greatly compromising the beauty. Further, a color grader usually takes into account the psychovisual effects of various image objects. For example, the dragon's flame appears bright but not too bright.
[0070] Although not intended to be limiting, there are two scenarios that can typically be processed on the creation side as follows. The content of both scenarios (similar only in that they contain important areas of high brightness) is very different, and although the brightness is specified (i.e., the gradation is also different), both contain some normal reflective objects (i.e., objects that reflect a certain percentage of the locally incident light). Since these reflect an average amount of light present in the scene, they obtain a relatively low brightness equivalent to that obtained with LDR. The idea is to generate darker colors lower than the range of high brightness of the emission color such as clouds that strongly reflect electric bulbs or sunlight when creating an image that is displayed after or on the premise of human visual adaptation. In the case of the image 301 of a fire-breathing dragon, it corresponds to the body of the dragon. In the case of the image 302 of a setting sun sinking into the sea, it is the boat. These normal objects are included within the dark portion range R_Lrefl that covers the lowest brightness available within the input dynamic range (in this example, the entire range is 0 to MaxL_HDR = 3000 nit), and in these scenes, it is, for example, at most 200 nit. If the body of the dragon is red, it can have pixels with an average of, for example, 130 nit, but if the body is black, it can have pixels with an average of, for example, 20 nit. Graphics pixels such as subtitles are usually appropriately placed slightly above this dark portion range. This has the advantage of being bright and noticeable (like LDR white), while not overwhelming the HDR effect of the video itself. In particular, when two images are mixed from separate HDR sources, it can also be argued that the upper limit brightness (or its luminance coding in terms of luma, e.g., the perceptual quantizer luma definition) does not need to be the same across the ranges of both normal scenes. However, by adjusting before mixing the video images, it can be restricted to its vicinity (e.g., the range of the dragon can be at most 1.5 times brighter than the range of the boat and the sea at the upper limit of the range). And this can be taken into account, for example, by setting the graphics range starting from the upper limit of the brighter of the two (e.g., starting from 300 nit instead of about 200 nit, or starting from 220 nit instead of about 150 nit).The HDR effect, for example, the orange / yellowish flames of a dragon, can be rendered at a pixel brightness of about 1500 nits (since this is a relatively large part of the image, a brighter brightness seems to be advantageous). Clouds illuminated by the setting sun can reach, for example, 2000 nits, and the sun disk itself can be about 3000 nits (for example, 95% of 3000 nits for yellow). All these HDR effects far exceed the graphics range. Video creators can specify that these exist, for example, between 400 nits and 500 nits (or 50 - 500 nits considering the small visual impact of dark graphics colors), on the master HDR range ML_C = 3000 nits. In the case of graphics that use only a limited subset of bright colors, focus on such a limited range and avoid covering the colors of normal videos and effect videos that may vary in grading, especially re - grading by luminance mapping. If dark colors are required, it is possible to adjust how far below the 400 level to go, for example, based on the prediction or reported variation at the lower part of the mapping function (for example, in the offline case, check all the curves that occur during the movie before processing, determine the characteristic value of the upper variation of the lower curve (shown as the second line segment in this example), and when displaying the transmitted video on - the - fly, transmit such characteristics and consider them when determining where to map the darkest black of the graphics). For example, if the curve starts to vary more than 20% or 40% compared to the average curve starting downward from a point in the low graphics range that is a fixed point, additional luma or brightness values (each less than L_low, Y_low) can be transmitted. For more important uses (or important types of graphics, such as the black text box around monochrome subtitles), the brightness or luma at the variation point of 20% or less can be used for the darkest color, and for less important uses, the 40% point can be used. In many cases, the upper limit point may be more important than the lower limit point (for example, when bright pixels are strongly compressed during re - grading).The lower limit luminance L_low of 400 nits is represented as a lower limit luma code Y_low (for example, in PQ, multiplied by a coefficient according to the bit depth (for example, 1023), which is 0.65). Although not intended to be limiting, in this specification, it is assumed that all pixel luminances are always coded in a common perceptual quantization coding, but other EOTFs or OETFs such as Hybrid Log Gamma (HLG) may be used. Also, the technology of the present invention functions with a mixture of EOTF / luma code definitions (for example, finding a first primary graphics element in a pixel color coding of a PQ luma definition and finding a second primary graphics element as a secondary reference in a part of an HLG-encoded image). The upper limit primary graphics range luminance L_high and the upper limit primary graphics range luma Y_high are usually determined, for example, when it is expected that there are not many bright objects in important HDRs, and thus it is not necessary to particularly strongly boost or compress the local curve shape. The continuity of the curve starting from a fixed point may be defined as a percentage multiplier of the lower limit point (for example, 150%) because the graphics staying within that range does not cause much problem of visual mapping mismatch. Since it is expected that various HDR objects in the upper range need to be compressed frequently, the amount of graphics usually depends on the maximum luminance of the input video. For example, when the maximum of the input image is 1000 nits, it is likely not desirable to select an upper limit point of 1.5x400 nits. This is because there is little room for bright HDR objects (it may be necessary to shade away from normal objects, which is not always desirable), or there is a significant overlap between the luminance of HDR video objects and at least the higher luminance possible within the secondary graphics range when coexisting the established primary graphics range and the secondary graphics range.
[0071] If there is only the HDR image itself (i.e., as the final image, displayed with equal luminance, i.e., each coded image pixel luminance is displayed as coded), the graphics problem is not yet that complex, but the HDR input image usually needs to be gamma mapped (or equivalent luminance mapping) (as explained in Figure 1). For the dragon image, a first luminance mapping function 310 is shown to obtain the corresponding graded LDR image as output (the corresponding gamma mapping function can be obtained, for example, by converting both the input luminance range and the output luminance range to the range of PQ gamma, or by taking the input as the PQ gamma axis and the output as the conventional LDR gamma of Rec.709). This function conceptually consists of the following three parts: an appropriate shaped mapping for dark normal objects included in the input range R_Lrefl, a mapping for the HDR effect range R_Heffs, and an intermediate part, i.e., a mapping for the graphics part range R_gra. The next scene in the video can be a sunset scene including a plurality of consecutive sunset images (i.e., that shot). This may have another preferred regrading gamma mapping function (315), which is usually sent together as metadata of the video signal and guides the display optimization of the end device (e.g., consumer television display) as explained in Figure 1. However, the video creator may define the regrading function so that the graphics part remains stable, i.e., all luminances are mapped to the corresponding output luminance that is the same in both HDR scenes. In that case, the graphics will look the same in both scenes for both HDR and LDR (or, any other intermediate display-adapted image (e.g., 950nit ML_C image optimized for a 950nit ML_D display)). Also, the video creator (i.e., the creator of the luminance mapping function) can select an appropriate graphics part range for both HDR (3000nit master HDR in this example) and LDR.If it is necessary to adjust the graphics by mixing HDR videos with different maximum brightness levels, it can be done through the output stable graphics range of the mapping (in this example, the 75 - 80 nit LDR range). Since the graphics can be freely set, it may depend on what is included within the effective range of the HDR movie.
[0072] As described with reference to FIG. 4, the remaining part of the movie basically includes two dark scenes. The first is a night scene 401 of riding a motorcycle through the city. The motorcycle may be rendered, for example, at 30 nits while giving a relatively dark impression but still being sufficiently visible (in LDR, tricks such as turning pixels blue were necessary to simulate night, but in HDR, more play can be had with the darkness of the pixels themselves. However, assuming that, for example, HDR movies are often viewed in a relatively bright lighting environment, it is also conceivable that the darkest pixel part ranges stay on the brighter side). Since the only bright objects in this image are streetlights, the pixel luminance may be set to 900 nits so as not to be unpleasantly bright and obstructive compared to other parts of the image. Then, another dark scene 402 follows. Strictly speaking, it is not a night scene, but generally it is still a dark scene (since it is a rainy day, outdoor pixels are displayed dimly). The very dark part is a clown hidden in the sewer, and the luminance is, for example, 10 nits or less so that the clown can hardly be seen. Regarding these two image sets, the video creator may consider that lower typical levels for primary graphics are more suitable, and examples of such levels are, in HDR, the other exemplary lower luminance L_low2 = 80 nits (represented as Y_low2 as a lumacode in the HDR image), (in the HDR grading) the other lower luminance L_high2 = 90 nits. Also shown are the corresponding appropriate LDR lower luminance (Ll_low = 60) and LDR upper luminance Ll_high = 65 nits. These depend on the graphics part of the two luminance mapping functions for these two dark scenes (for example, the third luminance mapping function 405 for the nighttime image 401 of the motorcycle). The dependency may also be reversed. The darkest part of the luminance mapping function 405 shown as a line segment for simplicity can be adjusted so that the clown appears appropriately dark in the re-graded LDR image that may be used to drive a conventional LDR TV, for example.With this scene change (e.g., a change from a bright fire-breathing dragon during the day to a night-time cityscape), a sudden change in the brightness of the graphics is not a problem as long as the change is not "arbitrary", which is what video producers want. This mechanism enables the video creator to select the brightness of all different image objects as desired, rather than an arbitrary change when using a simpler curve such as a power function brightness mapping function with a variable power coefficient. Note that it should be noted that not all colors used in the graphics need to be within a partial range where they are evenly mapped between various luma mapping curves. In general, for at least some applications or systems, it is sufficient if the bright colors of the graphics are stable, and the dark colors may vary somewhat. A device that mixes secondary graphics may decide to use a narrower range of graphics luma than the primary graphics. For example, only luma that does not have different mappings in a plurality of consecutive scenes with different re-gradings may be used.
[0073] FIG. 5 schematically illustrates an example of an apparatus embodying the concept of the present invention.
[0074] In this exemplary embodiment, it is (non-limitingly) assumed that the HDR image signal input unit 501 receives not only the HDR image itself (i.e., the coded pixel color matrix and the decoding information (e.g., the selected EOTF) that is transmitted together or known to the receiver), but also at least one luminance mapping function LMF for the time point t corresponding to one of the video images. This function may have a variable shape for successive images or shots of similar images of the same scene (e.g., a cave scene). Also, explicit metadata regarding the lower and upper lumas of the primary graphics may or may not be present (MET(YLo,Yhi), not necessarily using L for luminance and Y for luma). It is assumed that these are defined in the input domain, i.e., defined for the received HDR image. There may be an image graphics analysis circuit (511) within the device. How this operates may vary from device to device. For example, a simple device may detect only text and the graphics range from the detected text (e.g., the range including all the lumas used by all the text found from the darkest text pixel to the brightest pixel, or a part thereof if the range is wide (e.g., text pixels brighter than the average text luma)).
[0075] Figure 6 illustrates an example of a suitable circuit for performing an analysis of the primary graphics within an image (other analysis algorithms may be used). One of ordinary skill in the art will understand that alternative or supplementary other equivalent graphics analysis techniques may exist with respect to the various configurations of the units taught to explain this innovation.
[0076] A detector (601) of an area with limited color variation is configured to identify typical graphics colors within typical graphics low-level elements. While some graphics may be complex, many have a limited subset and may have only two different chromaticities even when illuminated differently. For example, since there may be text in an HDR image (e.g., the name of a horse may have the same white or single pixel color as the colored head shape of the horse), there may typically be a text detector 602 to make an appropriate determination of the primary graphics color or to make an initial determination. A segmentation map, such as a first segmentation map (SEG_MAP1), is a simple way to summarize areas identified as potentially representative primary graphics pixels for later easily determining the color of the area / map. Thus, for example, when the text character "G" is identified, the pixels constituting the character obtain a value of, for example, "1" within an initially zero matrix. There may also be a graphics analyzer 610 based on characteristics, which may be used in an iterative manner, for example, for more complex graphics (the color characteristics of the graphics are identified, a set is determined, and the characteristics are re-identified). For this, for example, a color type characterizer 611 is used. For example, highly saturated colors, such as strange purples, often (except in cases such as flowers) suggest that the pixels are likely to be graphics (often, graphics contain at least some primary colors, some of the RGB components are high or maximum, one or two components are zero or near zero, and such colors are not ideally included in natural image content, but rather subdued colors are included). These candidates may be further verified by other units, such as a basic graphics geometry analyzer 612. For example, graphics with characterizable shapes may be collated with their shape characterizer.
[0077] If a set of multiple connected or nearby (e.g., repetitive) pixels with specific color characteristics is small, it may be a graphics element. This is because, especially when there are multiple graphics, it is usually not obstructive and it is desirable to make it sufficient for reading or viewing. Also, the main objects in a movie often expand and become larger (e.g., a purple coat or a magic fireball may contain more pixels). Location can also be heuristic, and graphics usually exist near the boundaries of an image, such as a logo at the top or a ticker tape at the bottom. Such pixels may further be initial candidates in the first segmentation map SEGMAP_1 for verification, or conversely, may be retrieved again by further analysis. A person skilled in the art can understand how to use, for example, a neural network trained using aspects such as graphics and the simplicity or frequency of change of natural videos for comparison (e.g., texture metrics). The higher-level shape analysis circuit 620 may analyze further characteristics of the initially assumed graphics area (i.e., start based on, for example, a segmentation map that first identifies potential graphics pixels) and obtain a more stable set of pixels to summarize the lumen. As described above, it is not necessary to accurately identify all primary graphics pixels down to the details. For example, there may be an edge detector 621 for detecting edge pixels between the graphics shape and the surrounding video.
[0078] A G-criterion may be used to detect such boundaries (M. Mertens et al.: A robust nonlinear segment-edge finder, 1997 IEEE Workshop on Nonlinear Signal and Image Processing).
[0079] The G reference may be used to define characteristics (e.g., the color of a more planar natural object and a contrasting bright high-chroma color) as needed and calculate the amount of their co-occurrence in two regions. Since counting is involved, the shape of the region can also be adjusted as needed.
[0080] For example, two chromas of pixel color are used to define a first property. P1 = 1000*Cb + Cr
[0081] This property can be transformed by a further function, such as a deviation from a property based on a locally determined average chromaticity represented by, for example, the following formula. Delta_P = P1_pixel - P1_determined P2 = Function(Delta_P1)
[0082] This function classifies, for example, as 0 when the Delta_value is below a first threshold, as 1 - 9 for intermediate delta values, and as 10 when the delta values are sufficiently different (i.e., abs(1000*Cb - 1000*Cb_reference)>1000*Threshold1 or abs(Cr - Cr_reference)>Threshold2).
[0083] Next, for example, two regions are defined that move across the image until they are located on the left and right of the horizontal boundary of the ticker tape, such as two adjacent rectangles. The amount of movement may depend on the amount of coincidence.
[0084] The G_criterion(G) is the sum of the absolute values of the differences between the number of pixels having any value P2_i within rectangle 1 (e.g., P2_0 means the red component of the pixels being counted = 0, and P2_1 means red = 10) and the number of pixels having the same value P2_i within rectangle 2, for all possible different values of P2 in P2. Finally, this sum of absolute differences is divided by a normalization factor, which is usually the number of pixels in both rectangles (or the number of pixels adjusted by area if the areas of the regions are different). When comparing two test rectangles of the same size at adjacent positions within the image, the above formula becomes as follows. G = sum_i{abs[count(P2_i)_right_rectangle - count(P2_i)_left_rectangle]} / (2*L*W) [Equation 3]
[0085] Here, L and W are the lengths and widths of the two rectangles.
[0086] When the detector is within the graphics, almost the same color exists in both rectangles, and it is detected that all are zero, i.e., there are no edges. When one rectangle is on the graphics (e.g., saturated yellow) and the other is on the video (e.g., low-chroma green (excess of small green components)), the moving average from the green side (e.g., on the green graphics) becomes green, and the P2 value of the upper rectangle becomes almost zero. And compared to that continuous green color, the P2 values of the other sampling rectangles are all, for example, about 10 lower. In that case, in the left quadrilateral, there are L*W different pixels where almost all values are zero, and in the right quadrilateral, there are L*W pixels having P2 characteristics (e.g., having a value of 10 determined as an input to the G criterion). The different characteristic colors do not match on the left side, i.e., when corresponding to an edge, the G criterion approximates to a value of 1. The advantage of the G criterion is that any characteristic, such as a texture-based indicator, can be added for comparison. Instead of using other more classical edge detectors, a set of candidate points located at the edges of the graphics area and the start position of the natural video area may be found. Edge detectors often have the characteristic of being noisy, i.e., there may be both gaps and false edge pixels. Therefore, the shape analyzer 622 may include a preprocessing circuit for identifying the connected shapes in the HDR image from the detected edge pixels. Various techniques are known to those skilled in image analysis, for example, the Hough transform may be used to detect lines, circles may be collated, or splines or snakes may be used. By the analysis, if it is revealed that four lines (or one line and the boundary of the image) form a rectangle, and in particular, have specific characteristics such as having (exactly or approximately) the same width as the image and being at the bottom, the inner color is likely to be graphics pixels (e.g., the ticker tape of a news program). Therefore, the entire rectangle can be added to the second segmentation map SEGMAP_2. By determining symmetry based on the detected boundaries of what is expected to be graphics, it can be verified that the primary graphics element actually exists.In a simpler embodiment, for example, one may focus only on a few simple geometric elements, such as a rectangle detected at the bottom of the image (which is often sufficient for an initial consideration of the graphics range R_gr). And then, for example, only if such an element is not found, the algorithm can further search for more complex graphics objects, such as a star, a small area that jumps in, stays for a few seconds over several consecutive video shots, and then disappears again (i.e., the time-varying behavior of the graphics is also considered). Alternatively, conversely, the graphics region / object identified as a candidate may remain particularly invariant over several different shots of the movie, etc. The gradient analyzer circuit 623 may further analyze the graphics situation having an internal gradient (usually a long-distance gradient that slowly changes over dozens to hundreds of pixels) and not composed of a fixed set of colors. That is, such gradients include, for example, those that are all yellow but with the saturation changing from left to right. It may be positively verified or negatively verified that the gradient is likely to be a gradient generated by the graphics. For example, a blue gradient at the top of the image may be judged probably empty and discarded. This may depend on further geometric properties of the gradient in some embodiments, such as the size of the gradient region, the steepness of the gradient, the amount of colors spanned by the gradient, especially whether the gradient ends at a complex lower boundary (e.g., a tree). In the case of a blue line that may distinguish water from air, if the position of this line is much lower than the upper boundary of the image, the pixels may be discarded.
[0087] The suspected image analysis circuit 630 may provide additional certainty regarding the firmly determined graphics pixels in the third segmentation map SEGMAP_3. Modern and more attractive graphics may include, for example, a shape having clouds (generated by graphics) in the lower banner. Since this may look almost like a natural image, it may cause confusion. Even if correctly identified as graphics, it may not provide important new insights regarding the pixel lumen of the graphics that can be identified from other parts of the graphics. If it is actually part of the graphics, it will have a color that has been adjusted anyway. This may substantially overlap with the graphics range R_gra obtained by other methods, and the obtained upper limit lumen Y_high may be somewhat higher, or Y_low may be somewhat lower, but the method is not that important. Such areas attached to the identified graphics area, for example, the area at the lower left corner of a rectangular banner, may also be discarded (or, in an advanced embodiment, retained if it is verified to be a graphics element associated with the rest of the banner, for example, if it has a related color, but these suspected graphics may be complex, in which case it may be better to discard them). By discarding, a more reliable segmentation map can be obtained in determining the appropriate graphics range R_gra. More advanced embodiments may apply a texture or pattern recognition algorithm to areas that are not special areas of interest, i.e., areas with a small number of colors and simple gradients (i.e., deviating, for example, by 20% lower saturation), often geometrically simple and symmetric areas such as rectangles or substantially circular shapes. For example, a measure of busyness indicating how frequently and how fast the pixel color changes (e.g., which colors change) may be calculated for each 10x10 pixel area.Calculating the angular spread of lines (e.g., the center of gravity of an object with little color variation) is another measure that helps distinguish between natural complex patterns (e.g., leaves) and simple geometric patterns that appear in graphics (also, text usually has only two or a few stroke directions). In a more simple embodiment, the gradient may be calculated. If there are short-distance gradients (i.e., large changes over a few pixels) rather than the long-distance gradients commonly seen in graphics (changing slowly over dozens to hundreds of pixels), for robustness reasons, such regions may be excluded from the map. Also, if the color of a large surrounding geometry region (e.g., a text box) of a graphics region is significantly different from the average color of the rest of the graphics region, even if it is a complementary color (i.e., a color that may have been intentionally selected for this graphics, e.g., blue to complement orange), the processing logic of circuit 630 may ultimately remove those pixels from the set that determines the graphics range (i.e., remove those pixels from SEGMAP_3).
[0088] Regardless of the number of pixel region analysis sub - circuits or processes (more or less than the three illustrated), the final processing is typically performed by the luma histogram analysis circuit 650. This circuit examines all the luma of the pixels identified as primary graphics in the third segmentation map SEGMAP_3 (or an equivalent segmentation map or mechanism). It outputs a lower luma limit Y_low_ima based on image analysis that is near or at the lowest luma in the histogram of the luma of the pixels identified as graphics in SEGMAP_3. If the luma histogram contains many dark colors, the output may be higher than the minimum value, which is set, for example, to be more than 10% of the highest luma to obtain a practical graphics range. Conversely, if only white graphics pixels are detected in SEGMAP_3, instead of the non - useful setting of setting the lower and upper limits to be the same, the luma histogram analysis circuit 650 can output the lower luma limit Y_low_ima based on image analysis again, such as 25% of the upper luma limit Y_high_ima based on image analysis (which may be, for example, the luma of white text as more important parameters are typically determined first). Also, Y_high_ima does not have to always be the same as the maximum luma found in SEGMAP_3. For example, if only yellow is detected as the brightest color and it is known that their luminance is typically 90% of white, the luminance of 110% of the maximum luma detected in the histogram can be set as the upper luma limit. Or, when a bright - colored logo co - exists with dark - white subtitles, a value corresponding to the white value suitable for the brightest primary graphics element may be set.
[0089] Returning to FIG. 5, there may also be a luma mapping function analysis unit 512 for analyzing similar regions of the mapping function (i.e., regions where a subset of luma is mapped to approximately the same corresponding output luma by all different optimal image-dependent luminance mapping functions) for a series of (continuous or adjacent) images of the video. An exemplary algorithm thereof is described in FIG. 7. That is, any device may have unit 511, unit 512, or both, and usually there is some circuitry that summarizes the detected graphics ranges, such as overlaps or junctions. The same applies to unit 513.
[0090] Various luma mapping functions (LMF(t)) valid at a particular time are input to the identity analysis circuit 701. One or more previous luma mapping functions LMF_p are retrieved from the memory 702. If functions of different shapes are determined (i.e., if the input LMF(t) is of a different shape than the stored LMF_p at least at some points), the old LMF_p can be replaced or supplemented when comparing two or more functions. The identity analysis circuit 701 is configured to first check whether there is an overall identity of the functions, not just for sub-regions of the input luma. Some HDR codecs send one function per image, but since those functions may all be of the same shape for all images within the same scene, they need to be discarded. This is indicated by the establishment unit 703 detecting the ID bool FID as "yes" or "1". Next, simply the next function is read until an actually new function is loaded (e.g., the function of the dragon is old and the function of the sea 302 illuminated by the sun in FIG. 3 is new).
[0091] The partial range identity circuit 710 calculates the output value difference DEL for almost all input luma values Y_in. It is expected to find at least one outer region with a difference (e.g., below the graphics range) and find an intermediate range R_id where the outputs are the same. The lower luma Yb and upper luma Yt of this range can be determined. These can be output as the final lower and upper luma of the specified graphics range R_gra, regardless of whether they are suitable for the mixing of secondary graphics. Usually, the graphics range determination circuit 720 can perform further analysis before outputting the lower luma Y_lowfu and upper luma Y_highfu determined by the function. This can be based on, for example, checking whether the upper luma is within a typical range, i.e., below the typical upper limit HighTyp from the (dotted line, thus optional) range supply circuit 740. The same can happen for the typical lower luma LowTyp. If there is a 5000 nit master HDR image, finding a graphics range around 5000 nit may not be an appropriate graphics range as it may be considered too bright by some viewers. Embodiments can propose an average functional value for Y_highfu if the analysis fails, e.g., if the determined Yt far exceeds HighTyp, but usually an error state ERR2 is thrown. Also, a curve without an intermediate range of the identifier (R_id) may be communicated, in which case an error state may also be thrown (the first error state ERR). Here, for simplicity of explanation, it is shown that the graphics range has the same mapping for both curves for all points within that range, but generally, within a certain tolerance, this range can be relaxed and judged to be similar (e.g., a maximum 10% luminance deviation, etc.).
[0092] Returning to FIG. 5, in such a situation, the integrated logic circuit 515 cannot use the input of the lower limit luma and the upper limit luma from the luma mapping function analysis unit 512. In other cases, the similarity between Y_highfu and Y_high_ima is checked, and if they show close values, either value, or for example, the average value may be used as the final upper limit luma value Y_high, etc. Also, the upper limit value Yhi based on the metadata provided by the metadata extractor 513 may be directly used (similarly for the lower limit luma Y_lo of the metadata if it exists; if it does not exist, it can be replaced with a percentage of Y_hi, such as 1 / 3 of the corresponding luminance L_hi_met). In any case, at least the upper limit luma Y_high and usually also the lower limit luma Y_low are transmitted to the graphics generation circuit 520, so that these lumas can be taken into account when selecting the color of the secondary graphics element (or when converting if it is necessary to mix pre-created graphics elements). As explained, a graphics luma Y_gra suitable for each pixel is determined so that it usually fits within the graphics range (R_gra) or does not go much above or below it. The actual value of Y_gra depends on what the color and luma of the pixel contain in the graphics pattern. For example, if the darkest color (Y_orig_graph_min) is mapped to Y_low and the brightest color (Y_orig_graph_max) of the existing or to-be-created graphics is mapped to Y_high, the color of the pixels between Y_orig_graph_min and Y_orig_graph_max can be mapped proportionally (or non-linearly) between Y_low and Y_high.
[0093] Finally, the image mixer 530 mixes the color of the graphics (using the graphics pixel luminance Lgra because mixing is often more refined in the linear domain, or using the graphics pixel luminance Y_gra if, for example, the mixing consists of simple pixel replacement) and the color of the video pixel for (substantially) all pixels in the image.
[0094] In some embodiments, such as those that can be mixed in the output domain, there may be a lumamapper (533) (some embodiments of the device may be mixable only in the output domain, or only in the input domain (possibly before the final lumamapping), or both, and may be switched as needed). Even when appropriate graphics are mixed in the input domain, a luminance mapping function, such as a luminance mapping function for dragons, is applied to the pre-mixed image of the secondary graphics. This function may be scaled to various different end-user displays having different end-user display maximum luminance functions. However, with the innovation of the present disclosure, the graphics are relatively stable and remain relatively stable even when luminance mapping is performed. However, as described above, it is also possible to mix in the output domain of any function (i.e., the vertical axis in FIG. 3). In that case, however, the lumamapper knows how to mix in the output graphics range if it knows the luminance mapping function being applied. The reader will still understand that this makes blending easier because it is a future dynamic mapping (a mapping that can significantly change luminance or relative luminance) that has not yet been applied. In a typical application, for example, the end-user's TV pre-adjusts the graphics and sends them to the output domain for the final blend (e.g., can be coded with PQ).
[0095] The final mixed output image Im_out, having the mixed luminance Lmix_fi of the pixels (or usually the lumas that code them), is supplied via the image or video signal output 599.
[0096] The algorithm components disclosed herein may be implemented in practice as hardware (e.g., as part of an application-specific IC) or as software (in whole or in part) executed on a special digital signal processor or a general-purpose processor, etc.
[0097] A person skilled in the art will understand from the present disclosure which elements are optional improvements and can be implemented in combination with other elements, how the (optional) steps of the method correspond to the respective means of the apparatus, and vice versa. The term "apparatus" in the present application is used in the broadest sense, that is, as a group of means enabling the realization of a specific purpose. Thus, for example, it may be an IC (a small circuit part thereof), a dedicated device (such as a device with a display), or a part of a network system. "Configuration" is also used in the broadest sense and may include, inter alia, a single apparatus, a part of an apparatus, a set of cooperating apparatuses (parts thereof), etc.
[0098] The description of a computer program product is to be understood to encompass any physical realization of a set of commands that, after a series of loading procedures (which may include intermediate conversion procedures such as conversion to an intermediate language or a final processor language), enable a general-purpose or special-purpose processor to input commands to the processor and execute any of the characteristic functions of the invention. In particular, a computer program product may be realized as data on a carrier such as a disk or tape, data existing in a memory, data moving through a wired or wireless network connection, or program code on paper. In addition to program code, characteristic data necessary for the program may also be embodied as a computer program product.
[0099] Some of the steps necessary for the operation of the method, such as data input and output steps, may not be described in the computer program product but may already exist in the functions of the processor.
[0100] Note that the above embodiments are illustrative and do not limit the present invention. Those skilled in the art can easily map the presented examples to other areas of the claims, but for the sake of brevity, not all of these options are discussed in detail. In addition to the combinations of elements of the present invention combined in the claims, other combinations of elements are possible. Any combination of elements can also be realized by a single dedicated element.
[0101] The reference signs in parentheses in the claims are not intended to limit the claims. The term "comprising" does not exclude the presence of elements or aspects other than those listed in the claims. A singular element does not exclude the presence of a plurality of such elements. Each unit of the device in the present disclosure may, in a particular embodiment, be formed by a circuit on an application-specific integrated circuit (e.g., a color processing pipeline that applies a technical color change to the input pixel color such as YCbCr), be executed on a CPU or GPU (e.g., of a mobile phone), or be a software-defined algorithm executed on an FPGA. Usually, the computing hardware, whether it is a general-purpose bit-processing computer or a specific digital processing unit, is connected to a memory unit under the control of an operation command, and the memory unit may be mounted on a specific circuit or be off-board connected via a digital bus or the like. Some parts of such a computer may be directly connected to a large device such as a display panel controller or a hard disk controller for long-term storage, or may be connected to a physical medium such as a Blu-ray disc or a USB stick. Some functions may be distributed among various devices via a network. For example, some calculations may be executed on a server in the cloud.
Claims
Claim 1 A method for determining a second luma of pixels of a secondary graphics image element to be mixed with at least one high-dynamic range input image in a circuit for processing digital images, the method comprising: receiving a high-dynamic range image signal including the at least one high-dynamic range input image, in a method comprising: analyzing the high-dynamic range image signal to determine a range of primary graphics luma of primary graphics elements of the at least one high-dynamic range input image, the range of primary graphics luma being a sub-range of a range of luminance of the high-dynamic range input image, and determining the range of primary graphics luma includes determining a lower luma and an upper luma that specify endpoints of the range of primary graphics luma; luminance mapping the secondary graphics element using at least the brightest subset of its pixel luma included in the range of the graphics luma; mixing the secondary graphics image element with the high-dynamic range input image. Claim 2 The method of claim 1, wherein analyzing the high-dynamic range image signal includes detecting the one or more primary graphics elements in the high-dynamic range input image, establishing the luma of pixels of the one or more primary graphics elements, and summarizing the luma as a lower luma and an upper luma, the lower luma being lower than all or most of the luma of pixels of the one or more primary graphics elements, and the upper luma being higher than all or most of the luma of pixels of the one or more primary graphics elements. Claim 3 The method of claim 1, wherein analyzing the high-dynamic range image signal includes reading values of the lower luma and the upper luma written as metadata in the high-dynamic range image signal from the metadata in the high-dynamic range image signal. Claim 4 The analysis of the high dynamic range image signal includes reading two or more luma mapping functions associated with temporally continuous images from the metadata in the high dynamic range image signal, and establishing the range of the graphics luma as a range that satisfies a first condition regarding the two or more mapping functions. The two or more luma mapping functions map an input luma within the range of the graphics to a corresponding output luma, and the first condition is that for each input luma, the corresponding two or more output luma obtained by applying each mapping function from the two or more luma mapping functions to the input luma are substantially the same. The method according to claim 1.
5. Satisfy the verification of a second condition that the range of the graphics luma is included within the secondary range of an appropriate graphics luma, and the maximum value of the secondary range must be lower than a predetermined percentage of the maximum value of the range of the mixture of the image and the graphics. The method according to claim 4.
6. The method includes the step of luma mapping the pixel luma of the at least one high dynamic range input image to a corresponding output luma of at least one corresponding output image. The mixture is performed in the at least one corresponding output image by mapping the pixel luma using at least one luma mapping function obtained from the metadata of the high dynamic range image signal. The method according to any one of claims 1 to 5.
7. Including determining the secondary graphics image elements using a set of colors across a scale of colors of different brightnesses, a part of the colors of different brightnesses appears brighter than average to the human eye, a part appears darker than average, and at least the colors brighter than average are mixed into the at least one high dynamic range input image using the luma within the range of the graphics luma. The method according to any one of claims 1 to 6.
8. The mixture consists of pixel replacement or blending at a percentage less than 50% of the pixel brightness of the at least one high dynamic range input image, and the percentage is preferably 30% or less. The method according to any one of claims 1 to 7.
9. The determination of the set of colors involves selecting a limited set of relatively bright and darker-than-average colors, where the darker-than-average colors have a luminance higher than a secondary lower luminance, which is a certain percentage of the lower luminance, preferably higher than 70% of the lower luminance, according to the method of claim 7.
10. An apparatus for determining a second luminance of pixels of a secondary graphics image element and mixing the secondary graphics image element with at least one high-dynamic-range input image including a primary graphics element, the apparatus comprising: an input unit that receives a high-dynamic-range image signal including the at least one high-dynamic-range input image; an image signal analysis circuit that analyzes the high-dynamic-range image signal to determine a range that designates the luminance of pixels representing primary graphics and determines a range of graphics luminance of the at least one high-dynamic-range input image, wherein the range of graphics luminance is a partial range of the luminance range of the high-dynamic-range input image, and determining the range of graphics luminance includes determining a lower luminance and an upper luminance that designate the endpoints of the range of primary graphics luminance; a graphics generation circuit that generates or reads the secondary graphics element and luminance-maps the secondary graphics element using at least the brightest subset of the pixel luminances of the secondary graphics element included in the range of graphics luminance; an image mixer that mixes the secondary graphics image element with the high-dynamic-range input image to generate pixels having a mixed luminance; and an output unit that outputs at least one mixed image including the pixels having the mixed luminance.
11. The image signal analysis circuit includes an image graphics analysis circuit, and the image graphics analysis circuit: detects the one or more primary graphics elements within the high-dynamic-range input image; establishes the luminance of the pixels of the one or more primary graphics elements; Performing the summarization of the said luma into a lower luma and an upper luma, wherein the lower luma is lower than all or most of the luma of the pixels of the one or more primary graphics elements, and the upper luma is higher than all or most of the luma of the pixels of the one or more primary graphics elements, the apparatus according to claim 10.
12. The apparatus according to claim 10, wherein the image signal analysis circuit includes a metadata extraction circuit that reads values of the lower luma and the upper luma written in the metadata of the high dynamic range image signal from the metadata of the high dynamic range image signal.
13. The image signal analysis circuit includes a luma mapping function analysis unit, and the luma mapping function analysis unit reads two or more luma mapping functions of temporally consecutive images from the metadata in the high dynamic range image signal, and establishes the range of the graphics luma as a range that satisfies a first condition regarding the two or more mapping functions. The two or more luma mapping functions map an input luma within the range of the graphics to a corresponding output luma, and the first condition is that for each input luma, the corresponding two or more output luma obtained by applying each mapping function from the two or more luma mapping functions to the input luma are substantially the same. The apparatus according to claim 10.
14. The image mixer includes a luma mapper that maps the luma of the pixels of the at least one high dynamic range input image to a corresponding output luma of at least one corresponding output image. The at least one corresponding output image has a dynamic range different from the dynamic range of the at least one high dynamic range input image, for example, has a maximum luminance. The luma mapper further performs the mixing in the at least one corresponding output image according to at least one luma mapping function obtained from the metadata of the high dynamic range image signal. The apparatus according to claim 10.
15. The graphics generation circuit designs the secondary graphics image elements using a set of colors over a scale of colors of different brightnesses, a part of the colors of different brightnesses appears brighter than average to the human eye, a part appears darker than average, and at least the colors brighter than the average are mixed into the at least one high dynamic range input image using the luma within the range of the graphics luma, the apparatus according to claim 10.
16. The image mixer performs mixing by replacing the video pixels of the at least one high dynamic range input image with secondary graphics element pixels or blending at a percentage less than 50%, preferably 30% or less, of the pixel luminance of the at least one high dynamic range input image, the apparatus according to claim 10.
Citation Information
Patent Citations
Method and device for improved HDR image encoding and decoding
JP2018110403A
Graphics Blending for High Dynamic Range Video
US20160080716A1
Transitioning between video priority and graphics priority
US20180018932A1
Graphics-safe HDR image luminance re-grading
US20200193935A1