Estimating metadata for images without metadata or with metadata in unusable form

By iteratively updating the metadata set and optimization algorithms, the problem of lack or unusable metadata in images is solved, and high-quality mapping from high-dynamic range images to standard dynamic ranges is achieved to meet the creative intention of the display device.

CN120303938APending Publication Date: 2025-07-11DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380079631.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-16
Filing Date
2023-09-12
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the lack of or unusable metadata of the images results in a degradation of image processing and display quality, especially in the process of dynamic range conversion, inability to effectively match the creative intention.

Method used

By iteratively updating the metadata set, using optimization algorithms such as particle swarm optimization or Powell's method, the useable metadata of images generated based on the cost function minimization algorithm is realized to achieve the mapping of high-dynamic range images to standard dynamic range.

Benefits of technology

The quality and display effect of image processing are improved, so that high dynamic range images can be accurately mapped to standard dynamic range, satisfying creative intentions and adapting to different display devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303938A_ABST
    Figure CN120303938A_ABST
Patent Text Reader

Abstract

A method and apparatus for estimating metadata for an image without metadata or with metadata in an unusable form. In accordance with an example embodiment, a method of estimating metadata includes accessing a first image and a second image of a scene, the first image and the second image each having a first dynamic range (DR) and a different second DR. The method further includes generating a third image of the scene having a second DR by applying a mapping function to the first image, the mapping function configured using the applicable metadata set; generating a sequence of updated metadata sets by iteratively updating the applicable metadata sets based on a cost function that quantifies a difference between the second image and the third image; and calculating a value of the cost function to select an output metadata set from the sequence, the output metadata set having the estimated metadata of the second image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Provisional Application Serial No. 63 / 425,814, filed on November 16, 2022, which is hereby incorporated by reference in its entirety. Technical Field

[0003] Various example embodiments relate to image processing operations and, more specifically but not exclusively, to determining parameters for mapping image and video signals from a first dynamic range to a different second dynamic range. Background Art

[0004] This section presents aspects that may help in better understanding the present disclosure. Accordingly, statements in this section should be read from this perspective and should not be construed as admitting what is prior art or what is not prior art.

[0005] As used herein, the term "metadata" refers to any auxiliary information that is transmitted as part of an encoded bitstream and aids a decoder in rendering the corresponding image(s). For television broadcasting and video streaming, video metadata can be used to provide auxiliary information about a particular video and audio stream or file. Metadata can be directly embedded in the video or included as a separate file within a container (such as MP4 or MKV). Metadata can include information about the entire video stream or file or about a particular video frame. Metadata created by cameras, encoders, and other video processing elements can include, but are not limited to: timestamps, video resolution, digital film grain parameters, color space or gamut information, reference display parameters, primary display parameters, auxiliary signal parameters, file size, closed captions, audio language, ad insertion points, color space, error messages, and so on.

[0006] In some cases, for example, for legacy or older video and image content, the corresponding metadata is missing, non - existent, or only available in a format that is unusable or incompatible. Summary of the Invention

[0007] Various embodiments of methods and apparatuses for estimating metadata for images that either have no metadata or have metadata in an unusable form are disclosed. Various examples provide techniques for automatically generating usable metadata for such images based on iterative updates of a candidate image, the iterative updates being designed to minimize a cost function configured to quantify a relevant difference between the candidate image and a reference image. In some embodiments, the metadata is created using an optimization algorithm configured to use a per-pixel color error representation format specified in Recommendation ITU-R BT.2124. In various embodiments, the optimization algorithm may be selected from various exploration-exploitation or exploitation-based optimization algorithms.

[0008] According to an example embodiment, there is provided an image processing apparatus for estimating metadata, the apparatus comprising: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus to at least: access a first image of a scene and a second image of the scene, the first image having a first dynamic range (DR) and the second image having a second DR less than the first DR; generate a third image of the scene having the second DR by applying a mapping function to the first image, the mapping function being configured using an applicable metadata set; generate a sequence of updated metadata sets by iteratively updating the applicable metadata set based on a cost function configured to quantify a difference between the second image and the third image; and compute a value of the cost function to select an output metadata set from the sequence, the output metadata set having estimated metadata for the second image.

[0009] According to another example embodiment, there is provided an image processing method for estimating metadata, the method comprising: accessing, using an electronic processor, a first image of a scene and a second image of the scene, the first image having a first DR and the second image having a second DR less than the first DR; generating, using the electronic processor, a third image of the scene having the second DR by applying a mapping function to the first image, the mapping function being configured using an applicable metadata set; generating, using the electronic processor, a sequence of updated metadata sets by iteratively updating the applicable metadata set based on a cost function configured to quantify a difference between the second image and the third image; and computing, using the electronic processor, a value of the cost function to select an output metadata set from the sequence, the output metadata set having estimated metadata for the second image.

[0010] According to yet another example embodiment, a non-transitory machine-readable medium is provided, on which program code is encoded, wherein, when the program code is executed by a machine, the machine performs operations including the following: accessing, by an electronic processor, a first image of a scene and a second image of the scene, the first image having a first DR and the second image having a second DR that is less than the first DR; generating, by the electronic processor, a third image of the scene having the second DR by applying a mapping function to the first image, the mapping function being configured using an applicable metadata set; generating, by the electronic processor, a sequence of updated metadata sets by iteratively updating the applicable metadata set based on a cost function that quantifies a difference between the second image and the third image; and calculating, by the electronic processor, a value of the cost function to select an output metadata set from the sequence, the output metadata set having estimated metadata of the second image. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] By way of example, various other aspects, features, and advantages of the disclosed embodiments will become more fully apparent from the following detailed description and the accompanying drawings, in which:

[0012] Figure 1 is a block diagram illustrating a process flow for generating metadata according to various examples.

[0013] Figure 2 is a block diagram illustrating a metadata estimator employed in the process flow of Figure 1 according to various examples.

[0014] Figure 3 is a flowchart illustrating a method for generating metadata that can be used in the process flow of Figure 1 according to various examples.

[0015] Figure 4 is a block diagram illustrating a computing device according to various examples. DETAILED DESCRIPTION

[0016] As used herein, the term "dynamic range (DR)" can relate to the ability of the human visual system (HVS) to perceive the range of intensities (e.g., luminance, brightness) in an image, e.g., from the darkest black (shadow) to the brightest white (highlight). In this sense, DR relates to "scene-referred" intensities. DR can also relate to the ability of a display device to adequately or approximately render a range of intensities of a particular breadth. In this sense, DR relates to "display-referred" intensities. Unless a particular sense is specifically specified as having a particular meaning at any point in the description herein, it should be inferred that the term can be used in either sense, e.g., interchangeably.

[0017] As used herein, the term "high dynamic range (HDR)" relates to a DR breadth spanning 14 to 15 or more orders of magnitude of the HVS. In fact, relative to HDR, the DR of the wide range of intensities that humans can simultaneously perceive may be slightly truncated. As used herein, the term "enhanced dynamic range (EDR)" or "visual dynamic range (VDR)" can be related, either individually or interchangeably, to the DR that can be perceived by the human visual system, including eye movement, within a scene or image, thereby allowing some light adaptation changes over the scene or image. In this document, EDR can relate to a DR spanning 5 to 6 orders of magnitude. Although it may be slightly narrower than HDR relative to a reference real scene, EDR represents a wide DR breadth and can sometimes also be referred to as HDR.

[0018] In fact, an image includes one or more color components of a color space (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented with a precision of n bits per pixel (e.g., n = 8). Using non-linear luminance encoding (e.g., gamma encoding), an image where n ≤ 8 (e.g., a 24-bit color JPEG image) can be considered an image of standard dynamic range (SDR), while an image where n > 8 can be considered an EDR image.

[0019] The reference electro-optical transfer function (EOTF) of a given display characterizes the relationship between the color values (e.g., luminance) of an input video signal and the output screen color values (e.g., screen luminance) produced by the display. For example, ITU-R BT.1886, “Reference electro-optical transfer function for flat panel displays used in HDTV studio production” (March 2011) defines the reference EOTF for flat panel displays, the content of which is incorporated herein by reference in its entirety. Given a video stream, information about its EOTF can be embedded in the bitstream as (image) metadata.

[0020] As used herein, the term “PQ” refers to perceptual quantizer luminance. The HVS responds in a highly non-linear manner to increasing light levels. The ability of a human to observe a stimulus is affected by factors such as the luminance of the stimulus, the size of the stimulus, the spatial frequency that makes up the stimulus, and the light luminance level to which the eyes are adapted at the particular moment of viewing the stimulus. In some cases, the PQ function can map linear input gray levels to output gray levels that better match the contrast sensitivity threshold in the human visual system. An example of such a PQ mapping function is described in SMPTE ST 2084:2014, “High Dynamic Range EOTF of Mastering Reference Displays” (hereinafter “SMPTE”), the content of which is incorporated herein by reference in its entirety.

[0021] Many consumer displays can support a luminance of 100 to 300 cd / m 2 ² or nits. Many consumer HDTVs range from 300 to 500 nits, with new models reaching approximately 1000 nits. Such traditional displays represent lower dynamic range (LDR) displays, some of which are referred to as SDR displays. Traditional SDR video is a video technology that represents light intensity based on the brightness, contrast, and color characteristics and limitations of cathode ray tube (CRT) displays. Traditional SDR video typically represents images with a maximum luminance of approximately 100 nits, a black level of approximately 0.1 nit, and colors within the ITU 709 / sRGB color gamut.

[0022] The following description provides non-limiting examples of metadata that can be used in the various embodiments disclosed herein. In some examples, the metadata can be classified into several different sets, often referred to as metadata levels. The various embodiments can rely on all or only some of the metadata levels. In other words, in some examples, additional metadata can be generated and added to the image stream after the image processing disclosed herein is complete, or previously available metadata (if any) can be combined with the newly generated metadata. Examples of additional metadata that can be used in at least some embodiments are described in U.S. Patent Nos. 9,961,237, 10,540,920, and 10,600,166, all of which are incorporated herein by reference in their entireties.

[0023] Level 1 or L1 is the first set of metadata that can be created by performing pixel-level analysis on an image. L1 metadata includes the following values: (i) the lowest black level in the image, represented as minimum (or min); (ii) the average light luminance level on the image, represented as average (or avg, or mid); and (iii) the highest light luminance level in the image, represented as maximum (or max). Typically, L1 metadata is created for each image, and it can be assumed that the metadata is unique for each image (e.g., video frame) in a timeline or a content segment such as a movie, a TV series, or a documentary. However, in some examples, for instance, when a colorist copies the L1 metadata from one image to one or more other images in a timeline, multiple images can have the same metadata. This copying is sometimes done to match and apply the same mapping to similar shots of a scene. There are also other scenarios where multiple images have the same metadata. Such scenarios are known to those of ordinary skill in the relevant art.

[0024] In some examples, the L1 minimum represents the minimum value of the PQ-encoded min(RGB) value of the corresponding part of the video content (e.g., video frame or image), considering only the active area (e.g., by excluding gray bars or black bars, video black borders, etc.), where min(RGB) represents the minimum of the color component values {R, G, B} of a pixel. L1-mid and L1-max are calculated in a similar manner. In a specific example, L1-mid can represent the average of the PQ-encoded max(RGB) values of an image, and L1-max can represent the maximum of the PQ-encoded max(RGB) values of an image, where max(RGB) represents the maximum of the color component values {R, G, B} of a pixel. In some embodiments, the L1 metadata can be normalized to be within the range [0, 1].

[0025] In some examples, when it is necessary to create a dynamic or animated trim within a shot due to a color grading of the shot or a transition in the light / color composition, metadata is generated for each frame to create a smooth transition from one state of the image to another. In such examples, the per-frame metadata on each animated or dynamic frame can include L1 metadata, as well as level 2 (L2), level 3 (L3), and / or level 8 (L8) metadata - which are generally referred to as trim items and depend on the trim parameters that change over the extent of the frame. The trimming process provides the colorist with the option to examine the mapping generated by the L1 metadata and make changes or adjustments to obtain different results that match the creative intent. In some examples, a set of trim controls provided on a color correction or mastering system can be used to make changes to the metadata. In various examples, the trim controls produce corrected metadata and / or new metadata that modify the mapping, and the colorist can use any combination of the available controls to produce the desired result. While the trim controls are generally designed to mimic the look and feel of color correction tools / controls familiar to the colorist, it is important to note that the trim controls are essentially metadata modifier controls that generally do not perform any color correction or change the HDR master color grading. Adjustment of the trim controls generally produces new metadata, thereby producing a change in the mapping observed on an output display (e.g., a target display). The new metadata can be exported, for example, as an XML file.

[0026] In some examples, metadata for various trim levels is generated using some or all of the following controls. Lift, gamma, and gain are trim controls used to modify the shadows, midtones, and highlights of an image. In operation, these three controls adjust the tone mapping curve while essentially mimicking the response of traditional (non-metadata-based) lift controls, gamma controls, and gain controls. In other words, the lift, gamma, and gain trim controls only mimic the effects of traditional lift controls, gamma controls, and gain controls but have different functionality compared to traditional lift controls, gamma controls, and gain controls. Tone detail is a trim control that restores sharpness in the highlight regions of the mapped image. Tone detail works well in SDR by restoring some sharpness and detail of the highlights that may be lost when mapping from HDR to SDR. Chroma weight is a trim control that helps maintain the color saturation in the upper midtones and highlight regions, especially when mapping from HDR to SDR. This trim control is generally used to reduce the luminance of highly saturated colors, thereby increasing the detail in these regions. The range of chroma weight is from the minimum luminance with maximum saturation at one end of the control range to the maximum luminance with minimum saturation at the other end of the control range. Saturation gain is a trim control that enables the colorist to adjust the overall saturation of the mapped image. Saturation gain generally affects all the colors in the image.

[0027] In some examples, some or all of the following additional trimming controls are used. Midtone offset is a useful trimming control for matching the overall exposure of the mapped SDR signal to the HDR master or SDR reference. The midtone offset is used as an offset to the L1 midpoint and adjusts the midtones of the image without affecting the blacks and highlights. Changes made using the midtone offset will be recorded as part of the L3 metadata for each shot or frame of the item. Mid-contrast bias is a trimming control that compresses or stretches the image near the midtone region and can increase or decrease the contrast of the midtones of the mapped image. Mid-contrast bias is typically used with lift and / or gain to produce the desired overall result. Highlight clip is a trimming control that allows the colorist to set the level of detail in the highlights by retaining or clipping the highlights as needed. For example, when the mapped image shows unwanted detail, the highlights can be clipped. The resulting clip can extend into the upper midtones and can trigger some compensation using gamma or gain adjustments. For example, highlight clip can be useful when attempting to match the mapped SDR to an existing SDR reference (e.g., as described in some of the examples below).

[0028] In some examples, L8 metadata for each shot or frame of the item is used to record further trimming controls (which are referred to as secondary trimming controls). For example, the color saturation trimming control allows the colorist to adjust the saturation of the mapped image individually on red, yellow, green, cyan, blue, and magenta, or collectively on all colors when all colors are linked together. The color hue trimming control allows the colorist to offset the hue of the mapped image individually on red, yellow, green, cyan, blue, and magenta. These controls are useful when attempting to adapt / convert a larger color gamut to a smaller color gamut. Adjustments made to the mapping using the secondary trimming controls are typically recorded as L8 metadata in an XML file.

[0029] In various examples, the 100 nits (SDR) target is the lowest target for (the) mapping. Some studios only request the HDR master as the primary deliverable for their content and do not request a separate SDR version. In this case, the SDR version can be derived from the HDR master. Therefore, it becomes the responsibility of the entity to ensure that the resulting SDR matches the creative intent. A checking and trimming process under the 100 nits target can be used to ensure that the resulting SDR meets the creative intent and expectations. Some studios may also request additional trimming, for example, under a 600 nits PQ target. When performing the target trimming process for multiple targets, it is generally recommended to start with the lowest DR target and then proceed to higher DR targets.

[0030] Figure 1is a block diagram illustrating a process flow (100) for generating metadata according to various examples. For illustrative purposes, the process flow (100) is described below with reference to HDR and SDR examples. However, the various embodiments are not limited thereto. In various additional examples, two relevant DRs corresponding to the process flow (100) are typically a first DR and a different second DR that is less than the first DR. In Figure 1 the context of, HDR and SDR are examples of the first DR and the second DR, respectively.

[0031] The input to the process flow (100) includes an SDR image (110) and an HDR image (120). In a representative example, the SDR image (110) is generated by curating the HDR image (120) via a separate workflow ( Figure 1 not shown in). In various examples, such a separate workflow either does not generate metadata or generates metadata in an unusable form. Previously generated metadata, if any, may be unusable, for example, because the inherent structure of the metadata depends on parameters that are incompatible with the image curation tools currently available to the colorist responsible for processing the images (110, 120).

[0032] The process flow (100) employs a metadata estimator (130) to generate an SDR image (140) and metadata (150) based on the input images (110, 120). The SDR image (140) is an approximation of the SDR image (110) created by curating the HDR image (120) using an image processing tool that is compatible with the image curation tools available to the colorist responsible for processing the images (110, 120). In a representative example, the metadata (150) includes at least some of the above-described L2, L3, and L8 metadata corresponding to the SDR image (140) as well as L1 metadata. The metadata estimator (130) is configured to perform an iterative process aimed at generating the SDR image (140) such that the correlation difference between the SDR image (140) and the SDR image (110) is small based on one or more image comparison metrics employed by the metadata estimator (130). Thus, the metadata (150) can also be used as metadata corresponding to the SDR image (110).

[0033] In various examples, the metadata estimator (130) is configured based on multiple configuration and / or control inputs (128). The inputs (128) include one or more of the following: (i) an identification of the metadata level to be used for the metadata (150); (ii) an identification of the type of optimizer to be used for running the iterative process described above; (iii) optimization initialization parameters; (iv) an identification of one or more metrics (or objective functions) to be used for comparing relevant SDR images; and (v) an identification of the file format for which the metadata (150) is to be generated. In various examples, the metadata estimator (130) may be configured (based on the configuration / control inputs (128)) to process still images or sequences of video frames. In the case of video, a corresponding objective function specified via the inputs (128) may be selected so as to provide, for example, temporal smoothness of the trimming over the frame sequence in addition to meeting relevant intra-frame trimming objectives. Various embodiments, examples, and features of the metadata estimator (130) are described in more detail below.

[0034] Figure 2 is a block diagram illustrating a metadata estimator (130) according to various examples. The metadata estimator (130) includes an optimizer circuit or module (240). In operation, the optimizer circuit (240) generates metadata (260) based on the SDR image (110), the SDR image (220), and the cost function (250). A control signal (238) applied to the optimizer circuit (240) identifies the metadata level to be used for the metadata (260). In some examples, the control signal (238) is one of the configuration / control inputs (128). The metadata estimator (130) performs iterative calculations on the SDR image (220) and the metadata (260). The metadata (260) calculated in a previous iteration is applied via a feedback path (272) to the content mapping circuit or module (210) for the next iteration. When the iteration is stopped based on an applicable stop criterion, the SDR image (220) and the metadata (260) are output from the metadata estimator (130) as the SDR image (140) and the metadata (150), respectively.

[0035] The content mapping circuit (210) operates to map the HDR image (120) to the SDR image (220) based on applicable metadata. For the first iteration of the metadata estimator (130) intended to generate the metadata (150), the applicable metadata is provided via the control signal (208). In some examples, the control signal (208) is one of the configuration / control inputs (128), for example, an input signal configured to provide the optimization initialization parameters described above. For any of the next iterations of the metadata estimator (130) intended to generate the metadata (150), the applicable metadata is the metadata (260) provided via the feedback path (272).

[0036] In each iteration, the metadata estimator (130) performs calculations aimed at generating metadata (260) based on a comparison of SDR images (110, 220) quantified using a cost function (250). Several non-limiting examples of the cost function (250) are described in more detail below. The optimizer circuit (240) determines whether to stop or continue the iteration by: (i) calculating the value of the cost function (250) for the current pair of SDR images (110, 220), and (ii) comparing the calculated value of the cost function (250) with a fixed threshold. The fixed threshold is typically a configuration parameter of the corresponding optimization algorithm. When the calculated value of the cost function is greater than the fixed threshold, the optimizer circuit (240) advances the processing in the metadata estimator (130) to the next iteration. When the calculated value of the cost function is equal to or less than the fixed threshold, the optimizer circuit (240) stops the iteration.

[0037] In mathematical terms, the optimization problem numerically solved by the metadata estimator (130) can be formulated using Equation (1):

[0038]

[0039] where p represents the metadata; CM represents the content mapping function of the content mapping circuit (210); HDR(r, g, b) represents the HDR image (120) in the RGB color space; and SDR ref (r, g, b) represents the SDR image (110) in the RGB color space. In an example implementation of the metadata estimator (130), the above optimization problem is solved by iteratively finding an approximate minimum of the function Metric in the d-dimensional space of the relevant metadata parameters.

[0040] In some examples, the cost function (250) is implemented based on the function ΔE ITP, which is a per-pixel color error representation format specified in Recommendation ITU-R BT.2124, which is incorporated herein by reference in its entirety. The function ΔE ITP measures the distance between two pixels in the ICtCp color space. ICtCp is a color representation format specified in Recommendation ITU-R BT.2100, which is incorporated herein by reference in its entirety. Equation (2) provides an example mathematical expression of the function ΔE ITP:

[0041]

[0042] where the parameters I, T, P are expressed in terms of the coordinates of the ICtCp color space as follows:

[0043] I = I(3)

[0044] T = 0.5 * C T (4)

[0045] P = C P (5)

[0046] The subscripts "1" and "2" of the parameters I, T, and P refer to the first pixel and the second pixel of the pixel pair being compared, respectively. When applying the function ΔEITP to the SDR images (110, 220) in the optimizer circuit (240), the subscript "1" indicates the pixel of the SDR image (110), and the subscript "2" indicates the corresponding pixel of the SDR image (220). In this particular context, the term "corresponding" means that the first pixel and the second pixel have the same position within the pixel frame, and this position is generally the same for the images (110, 120, 140, 220).

[0047] It is apparent from the above description that the function ΔEITP is only a local, pixel-specific metric within the pixel frame. In contrast, the cost function (250) provides a metric for the entire pixel frame in the sense of equation (1). In various examples, the cost function (250) is calculated using the values of the function ΔEITP for multiple pixels. In one particular example, the cost function (250) is the average of the values of the function ΔEITP obtained over the pixel frame. In another particular example, the cost function (250) is the maximum value of the function ΔEITP in the pixel frame. In yet another particular example, the cost function (250) is a weighted sum of the average and the maximum value. As will be apparent to those of ordinary skill in the art from the above description, other embodiments of the cost function (250) based on the function ΔEITP or other suitable metrics for quantifying the difference between the SDR images (110, 220) are possible.

[0048] In various examples, the optimizer circuit (240) can be programmed to employ any suitable cost function optimization algorithm aimed at finding the optimal value of the metadata parameter p by locating the global minimum of the cost function (250). A variety of such algorithms (including but not limited to algorithms based on evaluating the Hessian, gradient, or only function values) are known to those of ordinary skill in the relevant art. In one particular non-limiting example, the optimizer circuit (240) is programmed to employ particle swarm optimization (PSO).

[0049] PSO is a computational method that optimizes the problem formulated by Equation (1) by iteratively attempting to improve candidate solutions based on a cost function (250). PSO solves the problem by having a swarm of candidate solutions, called particles, and moving these particles in the search space according to their positions and velocities. The movement of each particle is influenced by its local best position and is also guided towards the best-known positions in the search space, which are updated as other particles find better positions. This process gradually moves the swarm towards the best solution in the search space. PSO is metaheuristic because it makes few or no assumptions about the problem being optimized and can search a very large candidate solution space. Additionally, PSO does not rely on gradients, which means that, unlike some other optimization methods such as gradient descent and quasi-Newton methods, PSO does not require the optimization problem to be differentiable. Advantageously, PSO is suitable for efficient parallel computing implementations.

[0050] PSO is an example of an exploration-exploitation optimization algorithm. In various additional implementations of the optimizer circuit (240), other exploration-exploitation optimization algorithms can be used similarly. Generally, various optimization algorithms suitable for programming the optimizer circuit (240) can have different ratios between exploration and exploitation. Briefly defined, exploration is the ability of an algorithm to search those regions in the search space that have not been searched or visited. However, these un-searched regions may or may not produce better solutions. Therefore, exploration alone does not necessarily produce the best solution. In contrast, exploitation is the ability of an optimization algorithm to improve the best solution found by the optimization algorithm so far by searching a relatively small region around the solution.

[0051] In some additional examples, exploitation optimization algorithms can also be used to program the optimizer circuit (240). An example exploitation optimization algorithm suitable for programming the optimizer circuit (240) is the Powell method. The Powell method relies on the maximum gradient technique, which starts from an initial guess and moves towards the minimum in the search space by finding a good direction to move and calculating the actual distance to travel for each iteration. The corresponding algorithm iterates until no significant improvement is achieved by further iteration. The Powell method can be useful, for example, for finding the local minimum of a continuous but complex cost function, including non-differentiable functions.

[0052] Figure 3FIG. is a flowchart illustrating a method (300) for generating metadata (150) according to various examples. For illustrative purposes and without any implied limitation, method (300) is described below with reference to the PSO and Powell algorithms. In some examples, method (300) is implemented using a metadata estimator (130) as described below. Based on the provided description, a person of ordinary skill in the relevant art will be able to make and use additional implementations of method (300) without any undue experimentation, including implementations based on other exploration-exploitation and exploitation-based optimization algorithms.

[0053] Method (300) includes: at block (302), receiving an SDR image (110) and an HDR image (120). Method (300) further includes: at block (304), selecting a cost function (250). The cost function (250) can be selected, for example, from a plurality of available choices based on a specific objective (e.g., creative intent) that triggers the processing of the images (110, 120) in the metadata estimator (130). The selection of the cost function (250) can also depend on the specific optimization algorithm executed as part of method (300). For example, the above PSO algorithm and Powell algorithm can use their respective different cost functions (250).

[0054] Method (300) further includes: at block (306), initializing a content mapping function and an optimization algorithm. The content mapping function is implemented using the content mapping circuit (210) as described above and is initialized using a control signal (208). The optimization algorithm is run by the optimizer circuit (240) as described above and is initialized using a control signal (238).

[0055] Method (300) further includes: at block (308), calculating an SDR image (220) by applying the content mapping function to the HDR image (120) with applicable metadata. For the initial (first) iteration, the applicable metadata is provided via the control signal (208). For any subsequent iteration, the applicable metadata is the metadata (260) provided via the feedback path (272).

[0056] Method (300) further includes: at block (310), updating the metadata (260) by running the optimization algorithm using the optimizer circuit (240). In the initial iteration, the metadata (260) is regenerated. In any subsequent iteration, the metadata (260) is updated by the optimization algorithm based on the SDR image (220) calculated at block (308) and the calculation of the cost function (250) associated with that optimization algorithm. For example, for the PSO algorithm, the SDR image (220) is calculated at block (308) using the current candidate metadata set having the minimum value of the cost function (250) among a plurality of candidate metadata sets.

[0057] For the PSO algorithm, the operations performed in block (310) include computing a cost function (250) for each of a plurality of candidate metadata sets. In some examples, about 50 different candidate metadata sets may be used in each iteration. The operations performed in block (310) also include: changing and / or updating each of the plurality of candidate metadata sets based on the direction toward the respective weighted average of the personal best and the swarm best in the search space. The swarm best is the position in the search space of the current candidate metadata set having the minimum value of the cost function (250) among a plurality (e.g., 50) of candidate metadata sets. The personal best is determined based on the update history and is the position in the search space of the candidate metadata set for which the respective candidate metadata set has the personal minimum value of the cost function (250). The coefficients used to compute the weighted average are parameters of the PSO algorithm. In different embodiments, the computation of the candidate metadata sets in each iteration may be parallel or sequential.

[0058] For the Powell algorithm, there is one candidate metadata set in each iteration. The operations performed in block (310) include computing the cost function (250) of the current candidate metadata set. The operations performed in block (310) also include changing and / or updating the candidate metadata set based on the gradient direction of the cost function (250) (in the search space) or some approximation of that gradient direction.

[0059] The method (300) further includes: in decision block (312), determining whether an iteration stop criterion is satisfied. For the PSO algorithm, the stop criterion includes determining whether all of the plurality of candidate metadata sets are within a fixed distance of each other in the search space, e.g., within a multi-dimensional sphere of a fixed radius. The fixed distance (or radius) is a configuration parameter of the PSO algorithm. For the Powell algorithm, the stop criterion includes comparing the cost function value to a fixed threshold. The fixed threshold is a configuration parameter of the Powell algorithm. When the iteration stop criterion is not satisfied (\"no\" at decision block (312)), the processing of the method (300) in the metadata estimator (130) loops back to block (308). When the iteration stop criterion is satisfied (\"yes\" at decision block (312)), the processing of the method (300) in the metadata estimator (130) is directed to block (314).

[0060] The method (300) further includes: in block (314), outputting the last computed SDR image (220) and the best metadata (260) as the SDR image (140) and the metadata (150), respectively. When the output in block (314) is completed, the processing of the method (300) in the metadata estimator (130) terminates.

[0061] Figure 4FIG. is a block diagram illustrating a computing device (400) according to various examples. The device (400) may be used to implement, for example, a process flow (100). The device (400) includes an input / output (I / O) device (410), an image processing engine (IPE, 420), and a memory (430). The I / O device (410) may be used to enable the device (400) to receive input images (110, 120) and configuration / control inputs (128), as well as output images (140) and metadata (150). The I / O device (410) may also be used to connect the device (400) to a display.

[0062] The memory (430) may have a buffer for receiving image data and other related input data. The data may be in the form of, for example, image files, data packets, and XML files. Once the data is received, the memory (430) may provide a portion of the data to the IPE (420), for example, for performing a method (300). The IPE (420) includes a processor (422) and a memory (424). The memory (424) may store program code therein, which when executed by the processor (422) enables the IPE (820) to perform image processing, including but not limited to image processing according to some process flow (100) and method (300). Once the IPE (420) generates an image (140) and metadata (150) by executing the corresponding portions of the code, the IPE (420) operates to output the image and the metadata. The IPE (420) may perform rendering processing on various images (110, 120, 140, 220) and provide corresponding visual images for viewing on a display. The visual images may be in the form of, for example, a suitable image file output through the I / O device (410).

[0063] According to the example embodiments disclosed above, for example, in the Summary of the Invention section and / or with reference to Figures 1 to 4Any one or any combination of some or all of the following provides an image processing apparatus for estimating metadata, the apparatus comprising: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to cause the apparatus, together with the at least one processor, to at least: access a first image of a scene and a second image of the scene, the first image having a first DR and the second image having a second DR less than the first DR; generate a third image of the scene having the second DR by applying a mapping function to the first image, the mapping function being configured using an applicable metadata set; generate a sequence of updated metadata sets by iteratively updating the applicable metadata set based on a cost function that quantifies a difference between the second image and the third image; and calculate a value of the cost function to select an output metadata set from the sequence, the output metadata set having estimated metadata of the second image.

[0064] In some embodiments of the above apparatus, when the apparatus accesses the first image and the second image, there is no metadata associated with the second image.

[0065] In some embodiments of any of the above apparatuses, the first DR is a high DR; and wherein the second DR is a standard DR.

[0066] In some embodiments of any of the above apparatuses, the output metadata set includes level 1 metadata and another level metadata.

[0067] In some embodiments of any of the above apparatuses, for an initial iteration, the applicable metadata set is an initialized metadata set; and wherein, for any subsequent iteration, the applicable metadata set is the updated metadata set generated in the immediately preceding iteration.

[0068] In some embodiments of any of the above apparatuses, the iteratively updating includes running an optimization algorithm using a processor, the optimization algorithm being designed to find a minimum value of the cost function.

[0069] In some embodiments of any of the above apparatuses, the optimization algorithm includes a particle swarm optimization algorithm or a Powell-type optimization algorithm.

[0070] In some embodiments of any of the above apparatuses, the at least one memory and the program code are further configured to: cause the apparatus, using at least one processor, to calculate the cost function using a ΔE ITP function applicable to a pixel pair, one pixel of the pair being from the second image and the other pixel of the pair being from the third image.

[0071] In some embodiments of any one of the above devices, the value of the cost function is determined by finding the maximum value of the ΔE ITP function over the pixel frames corresponding to the second and third images.

[0072] In some embodiments of any one of the above devices, the value of the cost function is determined by calculating the average value of the ΔE ITP function over the pixel frames corresponding to the second and third images.

[0073] According to another exemplary embodiment disclosed above, for example, in the Summary of the Invention section and / or any one or any combination of some or all of the references Figures 1 to 4 there is provided an image processing method for estimating metadata, the method comprising: accessing, by an electronic processor, a first image of a scene and a second image of the scene, the first image having a first DR and the second image having a second DR less than the first DR; generating, by the electronic processor, a third image of the scene having the second DR by applying a mapping function to the first image, the mapping function being configured using an applicable metadata set; iteratively updating, by the electronic processor, the applicable metadata set to generate a sequence of updated metadata sets based on a cost function that quantifies the difference between the second image and the third image; and calculating, by the electronic processor, the value of the cost function to select an output metadata set from the sequence, the output metadata set having the estimated metadata of the second image.

[0074] In some embodiments of the above method, there is no metadata associated with the second image in the access.

[0075] In some embodiments of any one of the above methods, the first DR is a high DR; and wherein the second DR is a standard DR.

[0076] In some embodiments of any one of the above methods, the output metadata set includes level 1 metadata and another level of metadata.

[0077] In some embodiments of any one of the above methods, for an initial iteration, the applicable metadata set is an initialized metadata set; and wherein, for any subsequent iteration, the applicable metadata set is the updated metadata set generated in the immediately preceding iteration.

[0078] In some embodiments of any one of the above methods, the iteratively updating includes running, by the electronic processor, an optimization algorithm that is designed to find the minimum value of the cost function.

[0079] In some embodiments of any one of the above methods, the optimization algorithm includes a particle swarm optimization algorithm or a Powell-type optimization algorithm.

[0080] In some embodiments of any of the above methods, the method further comprises: using an electronic processor to calculate the cost function using a ΔE ITP function applicable to pixel pairs, one pixel in the pair being from a second image and the other pixel in the pair being from a third image.

[0081] In some embodiments of any of the above methods, the value of the cost function is determined by finding the maximum value of the ΔE ITP function over the pixel frames corresponding to the second and third images, or by calculating the average value of the ΔE ITP function over the pixel frames corresponding to the second and third images.

[0082] According to another example embodiment disclosed above, for example in the Summary section and / or with reference to Figures 1 to 4 any one or any combination of some or all of the above, there is provided a non-transitory machine-readable medium having program code encoded thereon, wherein when the program code is executed by a machine, the machine performs operations comprising: accessing, using an electronic processor, a first image of a scene and a second image of the scene, the first image having a first DR and the second image having a second DR less than the first DR; generating, using the electronic processor, a third image of the scene having the second DR by applying a mapping function to the first image, the mapping function being configured using an applicable metadata set; generating, using the electronic processor, a sequence of updated metadata sets by iteratively updating the applicable metadata set based on a cost function that quantifies the difference between the second image and the third image; and calculating, using the electronic processor, the value of the cost function to select an output metadata set from the sequence, the output metadata set having estimated metadata of the second image.

[0083] Regarding the processes, systems, methods, heuristics, etc. described herein, it should be understood that although the steps of these processes etc. have been described as being performed in a particular ordered sequence, these processes may be practiced using the described steps in an order different from that described herein. Further, it should be understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the process descriptions herein are provided for the purpose of illustrating certain embodiments and should in no way be construed as limiting the claims.

[0084] Accordingly, it should be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided will be apparent to those of ordinary skill in the art upon reading the above description. The scope should not be determined with reference to the above description, but should be determined with reference to the appended claims and the full scope of equivalents to which those claims are entitled. It is expected and intended that the technologies discussed herein will be developed in the future, and the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that this application is capable of modification and change.

[0085] All terms used in the claims are intended to be given the broadest reasonable interpretation and their ordinary meaning as understood by those who are aware of the technologies described herein, unless an explicit contrary indication appears herein. In particular, the use of singular articles such as "a", "the", and "said" should be understood to recite one or more of the indicated elements unless the claims recite an explicit contrary limitation.

[0086] A summary of the disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. This summary is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing detailed description, it can be seen that various features are grouped together in various embodiments for the purpose of presenting the disclosure as a unified whole. The methods of the disclosure should not be construed as reflecting an intention to claim embodiments that incorporate more features than are expressly recited in each claim. Rather, as reflected in the appended claims, the inventive subject matter lies in less than all of the features of a single disclosed embodiment. Accordingly, the appended claims are hereby incorporated into the detailed description, with each claim standing on its own as a separately claimed subject matter.

[0087] Although this disclosure includes references to illustrative embodiments, the specification is not to be construed as limiting. Various modifications of the described embodiments and other embodiments within the scope of the disclosure that are apparent to those skilled in the art to which the disclosure pertains are considered to be within the principles and scope of the disclosure, for example, as expressed in the claims.

[0088] Some embodiments may be implemented as circuit-based processes, including possible implementations on a single integrated circuit.

[0089] Some embodiments may be embodied in the form of methods and apparatuses for practicing these methods. Some embodiments may also be embodied in the form of program code recorded in a tangible medium, such as a magnetic recording medium, an optical recording medium, a solid-state memory, a floppy disk, a CD-ROM, a hard disk drive, or any other non-transitory machine-readable storage medium, wherein when the program code is loaded into and executed by a machine (such as a computer, etc.), the machine becomes an apparatus for practicing the claimed invention(s). Some embodiments may also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium (including being loaded into and / or executed by a machine), wherein when the program code is loaded into and executed by a machine (such as a computer or a processor, etc.), the machine becomes a device for practicing the claimed invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates similarly to a specific logic circuit.

[0090] Unless otherwise expressly stated, each numerical value and range should be interpreted as approximate as if the value or range were preceded by the word "about" or "approximately".

[0091] The use of figure numbers and / or reference numerals in the claims is intended to identify one or more possible embodiments of the claimed subject matter to facilitate the interpretation of the claims. Such use should not be construed as necessarily limiting the scope of these claims to the embodiments shown in the corresponding figures.

[0092] Although the elements in a method claim (if any) are recited in a specific order with corresponding labels, these elements are not necessarily intended to be limited to being implemented in a specific order unless the claim recites otherwise to imply a specific order for implementing some or all of these elements.

[0093] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present disclosure. In this specification, the phrase "in one embodiment" does not necessarily refer to the same embodiment everywhere and is not necessarily a separate or alternative embodiment mutually exclusive of other embodiments. The same applies to the term "embodiment".

[0094] Unless otherwise specified herein, the use of ordinal adjectives "first", "second", "third", etc. to refer to objects among a plurality of similar objects only indicates different instances of such similar objects being referred to and is not intended to imply that the similar objects so referred to must be in a corresponding order or sequence in time, in space, in ranking, or in any other way.

[0095] Unless otherwise specified herein, in addition to its plain meaning, the conjunction "if" may also or alternatively be construed to mean "when" or "upon" or "in response to determining" or "in response to detecting", which interpretation may depend on the corresponding specific context. For example, the phrase "if determining..." or "if detecting [stated condition]" may be construed to mean "after determining..." or "in response to determining..." or "after detecting [stated condition or event]" or "in response to detecting [stated condition or event]".

[0096] Also for the purposes of this description, the terms "coupled", "coupled to", "coupled with", "connected", "connected to", or "connected with" refer to any manner known in the art or later developed that permits energy to be transferred between two or more elements, and the insertion of one or more additional elements is contemplated, although this is not required. In contrast, the terms "directly coupled", "directly connected", etc. imply the absence of such additional elements.

[0097] As used herein with respect to elements and standards, the term "compatible" means that the element communicates with other elements in a manner specified in whole or in part by the standard and will be recognized by other elements as being sufficient to be able to communicate with other elements in the manner specified by the standard. Compatible elements need not operate internally in the manner specified by the standard.

[0098] The functionality of the various elements shown in the figures, including any functional blocks labeled "processor" and / or "controller", can be provided by using dedicated hardware as well as hardware capable of executing software associated with appropriate software. When provided by a processor, the functionality can be provided by a single dedicated processor, by a single shared processor, or by multiple individual processors, some of which may be shared. Additionally, the explicit use of the term "processor" or "controller" should not be construed to exclusively refer to hardware capable of executing software and may implicitly include, but is not limited to, digital signal processor (DSP) hardware, network processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), read only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage devices. Other conventional and / or custom hardware may also be included. Similarly, any switches shown in the figures are merely conceptual. Their functionality can be performed by the operation of program logic, by dedicated logic, by the interaction of program control and dedicated logic, or even manually, with the specific technique being selectable by the implementer, as more specifically understood from the context.

[0099] As used in this application, the terms "circuit" and "circuitry" can refer to one or more or all of the following: (a) a pure hardware circuit implementation (such as in a pure analog circuit and / or digital circuit); (b) a combination of hardware circuit and software, such as (if applicable): (i) a combination of analog hardware circuit and / or digital hardware circuit with software / firmware, and (ii) any part of a hardware processor with software (including (a) digital signal processor(s)), software, and (a) memory(s), which together enable a device such as a mobile phone or a server to perform various functions); and (c) (a) hardware circuit(s) and / or (a) processor(s), such as (a) microprocessor(s) or a part of (a) microprocessor(s), which require software (e.g., firmware) to operate, but the software may be absent when not needed for operation. This definition of circuit applies to all uses of the term in this application (including all claims). As a further example, as used in this application, the term "circuit" also covers an implementation of only one hardware circuit or one processor (or processors) or a part of a hardware circuit or processor and its accompanying software and / or firmware. The term "circuit" also covers, for example, a baseband integrated circuit or a processor integrated circuit for a mobile device, or a similar integrated circuit in a server, a cellular network device, or other computing or networking device when applicable to a particular claim element.

[0100] Those of ordinary skill in the art will recognize that any block diagrams herein represent a conceptual view of illustrative circuits embodying the principles of the present disclosure. Similarly, it will be recognized that any flowchart, flow diagram, state transition diagram, pseudocode, etc. represent various processes that can be substantially represented in a computer-readable medium and thus executed by a computer or a processor, whether or not the computer or processor is explicitly shown.

[0101] The "Summary of the Invention" in this specification is intended to introduce some example embodiments, and additional embodiments are described in the "Detailed Description" and / or with reference to one or more of the drawings. The "Summary of the Invention" is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

1. An image processing method for estimating metadata, the method comprising: Accessing, by an electronic processor, a first image of a scene and a second image of the scene, the first image having a first dynamic range (DR) and the second image having a second DR less than the first DR; Generating, by the electronic processor, a third image of the scene having the second DR by applying a mapping function to the first image, the mapping function being configured using an applicable metadata set; Generating, by the electronic processor, a sequence of updated metadata sets by iteratively updating the applicable metadata set based on a cost function that quantifies a difference between the second image and the third image; And Calculating, by the electronic processor, a value of the cost function to select an output metadata set from the sequence, the output metadata set having estimated metadata of the second image.

2. The method according to claim 1, wherein There is no metadata associated with the second image in the accessing.

3. The method according to claim 1 or 2, Among them, The first DR is a high DR; and Wherein the second DR is a standard DR.

4. The method according to any one of claims 1 to 3, wherein, The output metadata set includes level 1 metadata and another level of metadata.

5. The method according to any one of claims 1 to 4, Among them, For an initial iteration, the applicable metadata set is an initialized metadata set; and Wherein, for any subsequent iteration, the applicable metadata set is the updated metadata set generated in the immediately preceding iteration.

6. The method according to any one of claims 1 to 5, wherein, The iteratively updating includes running, by the electronic processor, an optimization algorithm that is designed to find a minimum value of the cost function.

7. The method according to claim 6, wherein The optimization algorithm includes a particle swarm optimization algorithm or a Powell-type optimization algorithm.

8. The method according to any one of claims 1 to 7, further comprising calculating, by the electronic processor, the cost function using a ΔE ITP function applicable to a pixel pair, one pixel of the pair being from the second image and the other pixel of the pair being from the third image.

9. The method according to claim 8, wherein, Determining the value of the cost function by finding a maximum value of the ΔE ITP function on a pixel frame corresponding to the second image and the third image, or by calculating an average value of the ΔE ITP function on a pixel frame corresponding to the second image and the third image.

10. A non-transitory machine-readable medium having program code encoded thereon, wherein, When the program code is executed by a machine, the machine performs operations including the method according to claim 1.

11. An image processing apparatus for estimating metadata, the apparatus comprising: At least one processor; And At least one memory including program code; Wherein the at least one memory and the program code are configured to cause the apparatus to perform at least the following operations using the at least one processor: Accessing a first image of a scene and a second image of the scene, the first image having a first dynamic range (DR) and the second image having a second DR less than the first DR; Generating a third image of the scene having the second DR by applying a mapping function to the first image, the mapping function being configured using an applicable metadata set; Generating a sequence of updated metadata sets by iteratively updating the applicable metadata set based on a cost function that quantifies the difference between the second image and the third image; and Calculating a value of the cost function to select an output metadata set from the sequence, the output metadata set having estimated metadata of the second image.

12. The apparatus according to claim 11, wherein, When the apparatus accesses the first image and the second image, there is no metadata associated with the second image.

13. The apparatus according to claim 11 or 12, Among them, The first DR is a high DR; and wherein the second DR is a standard DR.

14. The apparatus according to any one of claims 11 to 13, wherein The output metadata set includes level 1 metadata and another level of metadata.

15. The apparatus according to any one of claims 11 to 14, Among them, For an initial iteration, the applicable metadata set is an initialization metadata set; and wherein, for any subsequent iteration, the applicable metadata set is the updated metadata set generated in the immediately preceding iteration.

16. The device according to any one of claims 11 to 15, wherein, The iteratively updating includes running an optimization algorithm by the processor, the optimization algorithm being aimed at finding a minimum value of the cost function.

17. The apparatus according to claim 16, wherein, The optimization algorithm includes a particle swarm optimization algorithm or a Powell-type optimization algorithm.

18. The apparatus according to any one of claims 11 to 17, wherein The at least one memory and the program code are further configured to: cause the apparatus to calculate the cost function using a ΔEITP function applicable to a pixel pair by the at least one processor, one pixel in the pair being from the second image and the other pixel in the pair being from the third image.

19. The device according to claim 18, wherein, Determining the value of the cost function by finding a maximum value of the ΔE ITP function on a pixel frame corresponding to the second image and the third image.

20. The device according to claim 18, wherein, Determining the value of the cost function by calculating an average value of the ΔE ITP function on a pixel frame corresponding to the second image and the third image.

Citation Information

Patent Citations

  • Display management for high dynamic range video

    US10540920B2

  • Tone curve mapping for high dynamic range images

    US10600166B2

  • Display management for high dynamic range video

    US9961237B2