Efficient user-defined SDR to HDR conversion using model templates

Through machine learning, the brightness Gaussian process regression model and chromaticity dictionary are generated, combined with user input, and efficient conversion from SDR image to HDR image is achieved, solving the problems of low SDR to HDR conversion efficiency and poor user preference adaptability in the prior art, and generating a user-defined HDR appearance.

CN114223015BActive Publication Date: 2025-08-26DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080057574.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-15
Filing Date
2020-08-12
Publication Date
2025-08-26
Estimated Expiration
2040-08-12

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently convert standard dynamic range (SDR) images into high dynamic range (HDR) images, especially without relying on artificial color adjustments, and cannot adapt to user preferences of different display devices.

Method used

The brightness Gaussian process regression (GPR) model and chroma dictionary generated by machine learning are used to combine the image metadata input by users to generate user-defined compiler metadata, thereby achieving efficient conversion from SDR images to HDR images.

Benefits of technology

Generating user-defined HDR appearance on different display devices is realized, improving the efficiency and quality of image conversion, and reducing the dependence of manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114223015B_ABST
    Figure CN114223015B_ABST
Patent Text Reader

Abstract

A reverse shaping metadata prediction model is trained using training SDR images and corresponding training HDR images. Content creation user input is received, defining a user-adjusted HDR look for the corresponding training HDR images. A content creation user-specific modified reverse shaping metadata prediction model is generated based on the trained prediction model and the content creation user input. The content creation user-specific modified prediction model is used to predict operational parameter values ​​for a content creation user-specific reverse shaping mapping to reverse shape the SDR images into at least one mapped HDR image having the content creation user-adjusted HDR look.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 887,123, filed on August 15, 2019, and European Patent Application No. 19191921.6, filed on August 15, 2019, which are hereby incorporated by reference in their entireties. Technical Field

[0003] The present disclosure relates generally to images and more particularly to user-defined SDR to HDR conversion utilizing model templates. Background Art

[0004] As used herein, the term "dynamic range" (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., brightness, luminance) in an image (e.g., from darkest black (dark) to brightest white (highlight)). In this sense, DR is related to "scene-related" intensities. DR may also relate to the ability of a display device to adequately or approximately render a specific breadth of intensity range. In this sense, DR is related to "display-related" intensities. Unless a specific meaning is expressly specified to have a specific meaning at any point in the description herein, it should be inferred that the term can be used in either sense, e.g., interchangeably.

[0005] As used herein, the term high dynamic range (HDR) refers to a DR breadth of approximately 14-15 or more orders of magnitude across the human visual system (HVS). In practice, the DR over which humans can simultaneously perceive a wide breadth of intensity ranges may be slightly truncated relative to HDR. As used herein, the terms enhanced dynamic range (EDR) or visual dynamic range (VDR) may refer individually or interchangeably to the DR that can be perceived within a scene or image by the human visual system (HVS), including eye movements, thereby allowing for some light adaptation changes across the scene or image. As used herein, EDR may refer to a DR spanning 5 to 6 orders of magnitude. Therefore, although it may be slightly narrower relative to real scene-related HDR, EDR represents a wide DR breadth and may also be referred to as HDR.

[0006] In practice, an image includes one or more color components of a color space (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented by n bits of precision per pixel (e.g., n=8). Using nonlinear luminance coding (e.g., gamma coding), images where n≤8 (e.g., color 24-bit JPEG images) are considered to be images with a standard dynamic range, while images where n>8 can be considered to be images with an enhanced dynamic range.

[0007] The reference electro-optical transfer function (EOTF) for a given display characterizes the relationship between the color values ​​of the input video signal (e.g., luminance) and the output screen color values ​​produced by the display (e.g., screen brightness). For example, ITU Rec. ITU-R BT.1886, "Reference electro-optical transfer function for flat panel displays used in HDTV studio production" (March 2011), defines a reference EOTF for flat panel displays, which is incorporated herein by reference in its entirety. Given a video stream, information about its EOTF can be embedded in the bitstream as (image) metadata. The term "metadata" here refers to any auxiliary information that is transmitted as part of the coded bitstream and that helps a decoder render the decoded image. Such metadata may include, but is not limited to, color space or color gamut information, reference display parameters, and auxiliary signal parameters, as described herein.

[0008] The term "PQ" as used herein refers to perceived luminance magnitude quantization. The human visual system reacts to increasing light levels in a very nonlinear manner. The ability of humans to see a stimulus is affected by the brightness of the stimulus, the size of the stimulus, the spatial frequencies that make up the stimulus, and the brightness level to which the eye has adapted at the specific moment of viewing the stimulus. In some embodiments, the perceptual quantizer function maps a linear input grayscale to an output grayscale that better matches the contrast sensitivity threshold in the human visual system. An example PQ mapping function is described in SMPTE ST 2084:2014 "High Dynamic Range EOTF of Mastering Reference Displays" (hereinafter "SMPTE"), which is incorporated herein by reference in its entirety, wherein, given a fixed stimulus size, for each luminance level (e.g., stimulus level, etc.), the minimum visible contrast step size at that luminance level is selected based on the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model).

[0009] Supports 200 to 1000 cd / m 2A display with a brightness of 1,000 nits or nits represents a lower dynamic range (LDR) relative to EDR (or HDR), also known as standard dynamic range (SDR). EDR content can be displayed on an EDR display that supports a higher dynamic range (e.g., from 1,000 nits to 5,000 nits or more). Such a display can be defined using an alternative EOTF that supports high brightness capabilities (e.g., 0 to 10,000 nits or more). Examples of such EOTFs are defined in SMPTE 2084 and Rec. ITU-R BT.2100, "Image parameter values ​​for high dynamic range television for use in production and international programme exchange", (06 / 2017). As the inventors herein appreciate, there is a need for improved techniques for compiling video content data that can be used to support the display capabilities of a variety of SDR and HDR display devices.

[0010] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, problems identified with respect to one or more approaches should not be assumed to have been recognized in any prior art based on this section. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Embodiments of the invention are illustrated by way of example, and not by way of limitation, in the accompanying drawings in which like references indicate similar elements and in which:

[0012] Figure 1 Describes an example process of a video delivery pipeline;

[0013] Figure 2A and Figure 2B An example graphical user interface (GUI) display showing global and local modifications to a Gaussian process regression (GPR) model for brightness prediction;

[0014] Figure 3A An example distribution of mean predicted or estimated HDR chroma codeword values ​​is shown; Figure 3B shows an example angular distribution of clusters;

[0015] Figure 4A and Figure 4B An example process flow is shown; and

[0016] Figure 5A simplified block diagram of an example hardware platform is shown upon which the computers or computing devices described herein may be implemented. DETAILED DESCRIPTION

[0017] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are not described in detail in order to avoid unnecessarily obscuring, obscuring, or confusing the present disclosure.

[0018] Overview

[0019] This paper describes efficient user-defined SDR to HDR conversion using model templates. The technology described here uses efficient user-defined themes to generate user-defined compiler metadata that enables a receiving device to generate a user-defined mapped HDR image with a user-defined HDR appearance or look from an SDR image. The user-defined themes can be implemented on top of or as a starting point for dynamic SDR+model templates. The model templates include one or more machine learning (ML) generated luminance Gaussian process regression (GPR) models and one or more ML generated chrominance dictionaries that have been previously trained using a training dataset.

[0020] Example embodiments described herein relate to image metadata generation / optimization via machine learning and user input. A model template comprising a reverse shaping metadata prediction model is accessed. The reverse shaping metadata prediction model is trained using a plurality of training image feature vectors from a plurality of training standard dynamic range (SDR) images in a plurality of training image pairs and ground truth values ​​derived from a plurality of corresponding training high dynamic range (HDR) images in the plurality of training image pairs. Each training image pair in the plurality of training image pairs comprises a training SDR image from a plurality of training SDR images and a corresponding training HDR image from the plurality of corresponding training HDR images. The training SDR image and the corresponding training HDR image in each such training image pair depict the same visual content but have different luminance dynamic ranges. Content creation user input is received, the content creation user input defining one or more content creation user-adjusted HDR looks for the plurality of corresponding training HDR images. Based on the model template and the content creation user input, a content creation user-specific modified reverse shaping metadata prediction model is generated. using a content creation user specific modified reverse shaping metadata prediction model to predict operational parameter values ​​of a content creation user specific reverse shaping map for reverse shaping the SDR image into at least one of one or more content creation user adjusted HDR looks mapping the HDR image

[0021] Example embodiments described herein relate to image metadata generation / optimization using machine learning and user input. A standard dynamic range (SDR) image to be inverse-shaped into a corresponding mapped high dynamic range (HDR) image is decoded from a video signal. Compiler metadata is decoded from the video signal, the compiler metadata being used to derive one or more operational parameter values ​​for the content user-specific inverse shaping mapping. One or more content creation user-specific modified inverse shaping metadata prediction models are used to predict one or more operational parameter values ​​for the content user-specific inverse shaping mapping. One or more content creation user-specific modified inverse shaping metadata prediction models are generated based on a model template and content creation user input. The model template includes an inverse shaping metadata prediction model, wherein the inverse shaping metadata prediction model is trained using a plurality of training image feature vectors from a plurality of training SDR images in a plurality of training image pairs and ground truth values ​​derived from a plurality of corresponding training HDR images in the plurality of training image pairs. Each training image pair in the plurality of training image pairs includes a training SDR image from the plurality of training SDR images and a corresponding training HDR image from the plurality of corresponding training HDR images. The training SDR image and the corresponding training HDR image in each such training image pair depict the same visual content but have different luminance dynamic ranges. The content creation user input modifies a plurality of corresponding training HDR images into one or more content creation user-adjusted HDR looks. The SDR image is reverse-shaped using one or more operational parameter values ​​of the content user-specific reverse shaping mapping to a mapped HDR image of at least one of the one or more content creation user-adjusted HDR looks. A display image derived from the mapped HDR image is rendered by a display device.

[0022] Example Video Delivery Processing Pipeline

[0023] Figure 1 An example process of a video delivery pipeline (100) is depicted, showing various stages from video capture / generation to HDR or SDR display. Example HDR displays may include, but are not limited to, image displays operated in conjunction with televisions, mobile devices, home theaters, and the like. Example SDR displays may include, but are not limited to, SDR televisions, mobile devices, home theater displays, head-mounted display devices, wearable display devices, and the like.

[0024] Video frames (102) are captured or generated using an image generation block (105). The video frames (102) may be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation, etc.) to provide video data (107). Additionally, optionally or alternatively, the video frames (102) may be captured on film by a film camera. The film is converted to a digital format to provide video data (107). In some embodiments, the video data (107) may be edited or converted into an image sequence (e.g., automatically without human input, manually, automatically with human input, etc.) before being passed to the next processing stage / stage in the video delivery pipeline (100).

[0025] The video data (107) may include SDR content (e.g., SDR+ content, etc.), and image metadata that may be used by a receiving device downstream of the video delivery pipeline (100) to perform image processing operations on a decoded version of the SDR video content. Exemplary SDR video content may include, but is not necessarily limited to, SDR+ video content, SDR images, SDR movie releases, SDR+ images, SDR media programs, etc.

[0026] As used herein, the term "SDR+" refers to a combination of SDR image data and metadata that, when combined together, allow corresponding high dynamic range (HDR) image data to be generated. The SDR+ image metadata may include compiler data (e.g., user adjustments from a model template, etc.) to generate an inverse shaping mapping (e.g., an inverse shaping function / curve or set of polynomials, multivariate multiple regression (MMR) coefficients, etc.) that, when applied to an input SDR image, generates a corresponding HDR image of a user-defined HDR look or appearance. SDR+ images allow backward compatibility with legacy SDR displays, which may ignore the SDR+ image metadata and display only the SDR image.

[0027] Image metadata transmitted to a receiving device along with SDR video content may include compiler metadata generated (e.g., automatically, in real time, in an offline process, etc.) according to the techniques described herein. In some embodiments, video data (107) is provided to a processor for compiler metadata generation (115). The compiler metadata generation (115) may automatically generate the compiler metadata with little or no human interaction. The receiving device may use the automatically generated compiler metadata to perform an inverse reshaping operation to generate a corresponding high dynamic range (HDR) image from the SDR image in the video data (107).

[0028] Compiler metadata generation (115) can be used to provide one or more valuable services to make video content available for various display devices. One of the valuable services provided by compiler metadata generation (115) is the generation of HDR images from SDR images as described above in operational scenarios where an HDR image of the video content depicted in the SDR image is not available, but an SDR image depicting the video content is available. Thus, in these operational scenarios where an SDR image is available, the techniques described herein can be used to generate or compile HDR video content for an HDR display.

[0029] Another valuable service provided by compiler metadata generation (115) is the generation of HDR video content that is optimized (e.g., fully, partially, etc.) for HDR displays without relying on some or all of the manual work of a colorist, so-called "color adjustment" or "color grading."

[0030] In some operational scenarios, the encoding block (120) receives video data (107), automatically generated compiler metadata (177), and other image metadata; and encodes the video data (107) into an encoded bitstream (122) using the automatically generated compiler metadata (177), other image metadata, and the like. Example encoded bitstreams may include, but are not necessarily limited to, single-layer video signals, and the like. In some embodiments, the encoding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-ray, and other transport formats, to generate the encoded bitstream (122).

[0031] The encoded bitstream (122) is then transmitted downstream to a receiver, such as a decoding and playback device, a media source device, a media streaming client device, a television (e.g., a smart TV, etc.), a set-top box, a movie theater, etc. In the downstream device, the encoded bitstream (122) is decoded by a decoding block (130) to generate a decoded image 182, which may be similar to or identical to an image (e.g., an SDR image, an HDR image, etc.) represented in the video data (107) subject to quantization errors resulting from the compression performed by the encoding block (120) and the decompression performed by the decoding block (130).

[0032] In a non-limiting example, the video signal represented in the coded bitstream (122) can be a backward-compatible SDR video signal (e.g., an SDR+ video signal, etc.). Here, "backward-compatible video signal" refers to a video signal that carries an SDR image that is optimized for an SDR display (e.g., retains a particular artistic intent, etc.).

[0033] In some embodiments, the coded bitstream (122) output by the encoding block (120) may represent an output SDR video signal (e.g., an SDR+ video signal, etc.) embedded with image metadata, including but not limited to inverse tone mapping metadata, automatically generated compiler metadata (177), display management (DM) metadata, etc. The automatically generated compiler metadata (177) specifies an inverse shaping map that a downstream decoder may use to perform inverse shaping on an SDR image (e.g., an SDR+ image, etc.) decoded from the coded bitstream (122) to generate an inverse shaped image for rendering on an HDR (e.g., target, reference, etc.). In some embodiments, the inverse shaped image may be generated from the decoded SDR image using one or more SDR to HDR conversion tools that implement the inverse shaping map (or inverse tone mapping) specified in the automatically generated compiler metadata (177).

[0034] As used herein, inverse shaping refers to an image processing operation that converts a requantized image back to the original EOTF domain (e.g., gamma, PQ, hybrid log-gamma, or HLG, etc.) for further downstream processing (e.g., display management). Examples of inverse shaping operations are described in U.S. Provisional Patent Application No. 62 / 136,402, filed on March 20, 2015 (also published as U.S. Patent Application Publication No. 2018 / 0020224 on January 18, 2018), and PCT Application No. PCT / US2019 / 031620, filed on May 9, 2019, the entire contents of which are incorporated herein by reference as if fully set forth herein.

[0035] Furthermore, optionally or alternatively, a downstream decoder may use the DM metadata in the image metadata to perform display management operations on the inversely reshaped image to generate a display image (e.g., an HDR display image, etc.) optimized for rendering on an HDR reference display device or other display device such as a non-reference HDR display device.

[0036] In an operating scenario where the receiver works with (or is attached to) an SDR display (140) that supports standard dynamic range or relatively narrow dynamic range, the receiver may render the decoded SDR image on the target display (140) directly or indirectly.

[0037] In an operating scenario where the receiver operates with (or is attached to) an HDR display (140-1) that supports a high dynamic range (e.g., 400 nits, 1000 nits, 4000 nits, 10,000 nits, or more), the receiver may extract compiler metadata (e.g., user-adjusted from a model template, etc.) from the coded bitstream (122) (e.g., a metadata container therein, etc.), and use the compiler metadata to compile an HDR image (132) of a user-defined HDR appearance or look, which may be a reverse-shaped image generated by reverse-shaping an SDR image based on the compiler metadata. Furthermore, the receiver may extract DM metadata from the coded bitstream (122), apply a DM operation (135) to the HDR image (132) based on the DM metadata to generate a display image (137) optimized for rendering on an HDR (e.g., non-reference, etc.) display device (140-1), and render the display image (137) on the HDR display device (140-1).

[0038] Model templates, user-adjusted and modified templates

[0039] Single-layer inverse display management (SLiDM) or SDR+ can be used to enhance SDR content for rendering on HDR display devices. The luma and chroma channels (or color space components) of an SDR image can be separately mapped using image metadata (e.g., compiler metadata) to generate corresponding luma and chroma channels of a (mapped) HDR image.

[0040] The techniques described herein employ an efficient user-defined theme to generate user-defined compiler metadata that enables a receiving device to generate a user-defined mapped HDR image with a user-defined HDR look or appearance from an SDR image. The user-defined theme can be used in applications such as Figure 1 The dynamic SDR+ model template (142) is implemented based on or as a starting point. The model template (142) includes one or more machine learning (ML) generated luminance Gaussian process regression (GPR) models and one or more machine learning generated chrominance dictionaries, which are previously trained using a training data set. Exemplary ML generation of luminance GPR models and chrominance dictionaries is described in U.S. Provisional Patent Application No. 62 / 781,185, filed on December 18, 2018, the entire contents of which are incorporated by reference as if fully set forth herein.

[0041] The training dataset includes pairs of training SDR images and corresponding HDR images. Each training image pair described herein includes a training SDR image from the plurality of training SDR images and a corresponding HDR image from the plurality of corresponding training HDR images. The corresponding HDR image can be an HDR image derived from the SDR image in the same pair by professional color grading, manual color grading, or the like.

[0042] Training image features (e.g., content-related features, pixel value-related features, etc.) are extracted from a plurality of training SDR images. These image features are used to train an ML-generated luminance GPR model and an ML-generated chrominance dictionary for inclusion in a model template (142) accessible to a content creation user. The machine-learned optimal operating parameters of the ML prediction model / algorithm / method, the ML-generated luminance GPR model, and the ML-generated chrominance dictionary, as trained by the training image features, in the model template (142) can be persistently stored in a cache / memory, one or more cloud-based servers, etc., and made available (e.g., via a web portal, etc.) to a content creation user who wishes to create a relatively high-quality (e.g., professional-quality, near-professional-quality, non-trained, etc.) HDR image of a corresponding user-defined appearance or appearance from an (e.g., non-trained, user-owned, user-sourced, etc.) SDR image.

[0043] More specifically, a content creation user (e.g., a paying user, a subscriber, an authorized user, a designated user under a valid license, etc.) is allowed to access and modify a previously machine-trained model template (142) (e.g., a copy thereof, etc.) through user adjustments (144) to generate a user-defined theme (e.g., a user-defined HDR look or appearance, etc.) for SDR to HDR conversion, such as Figure 1 The correction template (146) includes one or more user-updated luminance GPR models and one or more user-updated chrominance dictionaries, which can be used to generate compiled metadata for reverse-shaping an SDR image into an HDR image according to a user-defined theme for SDR to HDR conversion.

[0044] The user adjustments (144) may be determined based on user input by a content creation user via one or more user interfaces rendered to the user by the system as described herein. In some operational scenarios, to create the correction templates (146), an intuitive solution is to keep all of the training SDR images unchanged and allow the user to adjust the (HDR) look or appearance of the training HDR images corresponding to the training SDR images. The training HDR images adjusted by the user may then be used as a new training dataset in combination with the unchanged training SDR images to generate a user-updated luminance GPR model and chrominance dictionary. However, for most end users, who may be amateurs and / or lack experience / education in color grading, adjusting all of the training HDR images in a training dataset containing a relatively large number of training images would be a very time-consuming task. Furthermore, retraining all of the parameters from the new training dataset including the modified training HDR images would also consume a relatively large amount of computing power.

[0045] Under the techniques described herein, correction templates (146) can be derived in a relatively simple and efficient manner. End users (e.g., content creation users, etc.) are allowed to perform (e.g., only perform, etc.) a relatively minimal amount of user interaction or manipulation, while still enabling a relatively maximum amount of appearance adjustments based on the creative intent of these end users.

[0046] These techniques can be used to calculate model parameters in a correction template (146) to achieve a user-defined HDR look in a relatively easy, simple, and efficient manner. The calculation of the model parameters can be based on user adjustments (144) that specify user-defined preferences for visual characteristics such as brightness, saturation, and hue in the HDR image.

[0047] Luma and chroma user adjustments may be handled differently. In some operational scenarios, for luminance, a simple least squares solution may be formulated or computed to generate a user-updated GPR model, thereby avoiding re-running the entire machine learning algorithm with the user-updated training HDR images. For chroma, a combined set of input training features in vector or matrix form computed from all image clusters in the training dataset may be used to generate a user-updated chroma dictionary via a simple matrix multiplication, avoiding performing a full complex numerical optimization. In some operational scenarios, the image clusters described herein may be generated based on an automatic clustering algorithm using similar (e.g., SDR, etc.) image features and / or characteristics and / or themes and / or events, etc.

[0048] The user can adjust the desired HDR look in global settings or local settings. Global settings can be applied (e.g., universally, etc.) to all GPR models and / or colorimetric dictionaries for all image clusters. Local settings can be applied (e.g., differently, etc.) to different image clusters to result in different user-adjusted modifications for different image clusters.

[0049] Luminance GPR model adjustment and retraining

[0050] Gaussian process regression (GPR) can be used to derive optimized operating parameters for an ML-generated luminance GPR model. The GPR model, representing a portion of a model template (142), can be trained (e.g., pre-trained, advanced, etc.) using a training dataset, as described herein. The trained ML-generated luminance GPR model (for simplicity, a pre-trained GPR model) can then be reused to compute new GPR parameters based on a modified desired HDR look derived from user input.

[0051] The image features extracted from the j-th training SDR image (or frame) among the plurality of training SDR images in the training dataset are indicated as a feature (e.g., columnar, etc.) vector x j The feature matrix X may be formed by a plurality of feature vectors including image features of a plurality of training SDR images.

[0052] The GPR model will be based on the j-th eigenvector x j The corresponding target value that is estimated or predicted (e.g., the target HDR codeword value to be mapped or reverse-shaped from a particular SDR codeword in the reverse shaping curve, etc.) is indicated as y j . The target vector y can be formed by multiple target values ​​to be estimated or predicted by the GPR model based on multiple feature vectors including image features extracted from multiple training SDR images.

[0053] The feature matrix X or the feature vectors therein are used as inputs in a GPR process that implements Gaussian process regression to derive optimized operating parameters of the GPR model, while the target vector Y or the target values ​​therein are used as responses (e.g., targets, references, etc.) in the same GPR process.

[0054] The Gaussian process (GP) used in the GPR process is a collection of random variables, where any finite number of random variables have a jointly Gaussian distribution. The GP is fully specified by its mean function, denoted m(x), and its covariance function, denoted r(x,x'). The mean function m(x) and covariance function r(x,x') of the true process f(x) are defined as follows:

[0055] m(x)=E[f(x)] (1-1)

[0056] r(x,x')=E[(f(x)-m(x))(f(x')-m(x'))] (1-2)

[0057] GP can be expressed or represented as follows:

[0058] f(x)~GP(m(x),r(x,x')) (2)

[0059] Where “~” indicates that the true process f(x) is distributed according to the GP and is characterized by the mean function of the GP denoted as m(x) and the covariance function of the GP denoted as r(x,x′).

[0060] Let f p =f(x p ) is consistent with the expected situation (x p ,y p ) corresponds to a random variable, where x p denotes the pth feature vector comprising image features extracted from the pth training SDR image, and y p represents the p-th target value to be estimated or predicted by the given GPR model.

[0061] Under the consistency / marginalization requirements of GP, if (y1,y2)~N(μ,Σ), then (y1)~N(μ1,Σ 11 ), where Σ 11 is the correlation submatrix of Σ. In other words, examination of the larger set of variables does not change or alter (eg, within the larger set of variables, etc.) the distribution of the smaller set of variables.

[0062] The GPR process can be based on a selected covariance function (or kernel) Example covariance functions may include, but are not necessarily limited to, the following rational quadratic (RQ) function:

[0063]

[0064] The hyperparameters (σ f ,α,l) can be found through the GPR optimization process as follows.

[0065] The covariance matrix is ​​constructed based on the RQ function as follows:

[0066]

[0067] For the case of noise-free data, {(x p ,f p)|p=1,...,F}, where F represents the total number of images in the multiple training SDR images, the joint distribution of the training output associated with the training dataset (denoted as f) and the test output associated with the given test dataset (denoted as f*) can be expressed as follows:

[0068]

[0069] The joint Gaussian prior distribution of observations or outputs associated with a given test data can be expressed as follows:

[0070] f * |X * ,X,f~N(R(X * ,X)R(X,X) -1 f,R(X * ,X * )-R(X * ,X)R(X,X) -1 R(X,X * )) (6)

[0072] For noise In the noisy data case, the joint distribution of the training output associated with the training dataset (denoted as y) and the test output f* associated with a given test dataset can be expressed as follows:

[0073]

[0074] The predicted output values ​​of the GPR process are as follows:

[0075]

[0076] in

[0077]

[0078] The prediction vector in the above formula (9) can be calculated relatively efficiently as follows:

[0079]

[0080] w=L T \(L\y) (12)

[0081]

[0082] Where cholesky(...) represents the Cholesky decomposition of the matrix in brackets (...); the operator "\" represents the left matrix division operation.

[0083] In fact, for the case of noisy data, the covariance matrix in Equation (4) can be directly calculated based on the data collected from the training dataset (e.g., without estimates, etc.).

[0084] Denote the qth element in w as w q Given a feature vector A new input (e.g., non-training, test, etc.) SDR image representing the extracted image features, and the predicted values ​​from the GPR model for the new input SDR image It can be given as follows:

[0085]

[0086] A GPR model may be characterized by some or all of the following parameters:

[0087] The kernel hyperparameters θ = {σ f ,α,l}

[0088] {x q}: Feature vector (F vector, each with K dimensions)

[0089] {w q}:Weight factor (F factor)

[0090] Hyperparameters (σ f ,α,l) represent some or all of the decisive parameters of the GPR model performance. Hyperparameters (σ f The optimal operating value of ,α,l) can be obtained or solved by maximizing the logarithm of the marginal likelihood as follows:

[0091] p(y|X)=∫p(y|f,X)p(f|X)df (15)

[0092] For the noise-free data case, the logarithm of the marginal likelihood can be given as:

[0093]

[0094] For the case of noisy data, the logarithm of the marginal likelihood can be given as follows:

[0095]

[0096] An example optimal solution or optimized value for each hyperparameter can be obtained by solving the partial derivatives of the marginal likelihood as follows:

[0097]

[0098] Retrain with updated target values

[0099] The model templates described in the article (e.g., Figure 1 142, etc.) may include multiple ML-generated luma GPR models for predicting or generating a luma reverse shaping curve (e.g., a luma reverse shaping function, a reverse shaping lookup table or BLUT, etc.). Each of the multiple ML-generated luma GPR models may be operated using an optimized operating value, which is generated from a training data set as described herein by the aforementioned operations represented by equations (1) to (18) above. Each such ML-generated luma GPR model may be used to predict or estimate an HDR mapping codeword (or value) mapped from a corresponding SDR codeword (or value) in a plurality of SDR codewords in an SDR codeword space. The HDR mapping codewords predicted or estimated by the multiple ML-generated luma GPR models and their corresponding SDR codewords may be used to construct a luma reverse shaping curve. Example generation of a luma reverse shaping function based on an ML-generated luma GPR model is also described in the aforementioned U.S. Provisional Patent Application No. 62 / 781,185.

[0100] User-defined themes (e.g., corresponding to a particular user-adjusted HDR look or appearance, etc.) can be based on user adjustments made to the model template (142) (e.g., Figure 1 144, etc.) to generate. In the user-defined theme, the training SDR images in the training dataset remain unchanged. Therefore, the input feature matrix X (or {x q} feature vector) remains unchanged. However, the desired target value y (e.g., the target HDR codeword to be inverse-shaped from a given SDR codeword in the inverse shaping function, etc.) is changed to y based on a user adjustment (144) made (or perceived to be made) to a training HDR image in the training dataset that corresponds to a training SDR image in the same training dataset.

[0101] Given a training SDR image that is unchanged under user adjustment (144) and a corresponding HDR image that is now changed (or perceived to be changed) under user adjustment (144), some or all of the following GPR operation parameters may be recomputed to reflect the changes in the HDR image, e.g., the kernel's hyperparameters θ = {σ f ,α,l} and {w q}: Weighting factor (F factor).

[0102] In some operational scenarios, all of the aforementioned GPR operating parameters can be directly retrained or recalculated by rerunning the GPR process as previously described based on a new training dataset comprising unchanged training SDR images and modified training HDR images. However, as previously described, the training process represented by the complete GPR process will take a relatively long time to complete and will also consume significant computational and other resources.

[0103] In some operational scenarios, by keeping the hyperparameter θ constant and only updating the weighting factor {w q}, a faster and more efficient solution or process can be used to update the GPR operating parameters.

[0104] Since the feature matrix X does not change, the hyperparameter θ (or its component σ f ,α,l) remains unchanged. Therefore, the covariance matrix R(X,X) also remains unchanged. In addition, the L matrix also remains the same as shown in the above expression (11).

[0105] Then, the new weighting factors corresponding to the user-defined topics can be obtained as a simple least squares solution as follows:

[0106]

[0107] in Indicates that the weight factor will be updated by The updated GPR model predicts or estimates the new target value.

[0108] The predicted value of a new input (e.g., non-training, testing, etc.) SDR image by the updated GPR model can be obtained by simply plugging in the new weighting factors as follows:

[0109]

[0110] or

[0111]

[0112] In other words, in some operational scenarios, rather than re-running the complete GPR (machine learning) process, the weighting factors {w q}, thus significantly reducing resource usage and generating correction templates (e.g. Figure 1 146, etc.) time.

[0113] Dictionary-color adjustment and retraining

[0114] Multivariate multiple regression (MMR) can be used to derive optimized operating parameters for the chromaticity dictionary. An example of an MMR model can be found in U.S. Patent No. 8,811,490, "Multiple color channel multiple regression predictor," the entire contents of which are incorporated herein by reference. The chromaticity dictionary representing a portion of the model template (142) can be trained (e.g., in advance, in advance, etc.) with a training data set (e.g., the same training data set used to train the GPR model, etc.), as described herein. The trained chromaticity dictionary (for simplicity, a pre-trained chromaticity dictionary) can then be reused to calculate new MMR operating parameters based on the modified desired HDR appearance derived from the user input.

[0115] A plurality of training SDR images and a plurality of corresponding training HDR images can be divided into a plurality of image clusters based on image features (e.g., brightness, color, resolution, etc.), feature recognition, image-related attributes (e.g., subject, time, event, etc.), etc., by automatic clustering techniques. In some operating scenarios, the training SDR images and their corresponding training HDR images (with or without user adjustment) can be automatically clustered into a plurality of image clusters based on feature vectors extracted from the training SDR images. The image clusters (or corresponding feature vector clusters) can be characterized by cluster centers in the image clusters that represent the centroids of the feature vectors extracted from the training SDR images. The aforementioned U.S. Provisional Patent Application No. 62 / 781,185 describes example automatic clustering associated with images in the training dataset.

[0116] Let triple and denote the normalized Y, C0, and C1 values ​​of the i-th pixel in the j-th training SDR image (or frame) and the j-th training HDR image (or frame), respectively.

[0117] The Y, C0 and C1 codeword ranges in the SDR codeword space can be divided into Q y , and As a result, a three-dimensional (3D) table is constructed for the jth training SDR image. --have Bin or entry. The 3D table Each bin in stores a (3-element) vector comprising three vector components, each of which is first initialized to zero. After all vectors in all bins (or all entries) of the 3D table are initialized to [0, 0, 0], each SDR pixel in the j-th training SDR image can be processed to determine the corresponding bin or bin index to which the SDR pixel belongs or is associated More specifically, the bin association between each such SDR pixel and its corresponding bin or bin index can be found as follows:

[0118]

[0119] in Represents the floor operator (removes any fractional values ​​from the value enclosed by the floor operator).

[0120] It should be noted that the bin association of SDR pixels and the bin association of HDR pixels corresponding to SDR pixels are dominated by SDR pixels only.

[0121] Once the bin association indicated in equation (22) above is determined, the Y, C0, and C1 codeword values ​​of the SDR pixel are added to the 3D table In other words, or where the bin vector accumulation is mapped to the t-th bin or the Y, C0, and C1 codeword values ​​associated with the t-th bin (e.g., for all SDR pixels in the j-th training SDR image), as shown below:

[0122]

[0123] For 3D tables The pixel in the j-th SDR frame of the t-th bin.

[0124] In addition, a 3D histogram Π can be constructed for the j-th training SDR image j , each bin in the 3D histogram stores the total number of SDR pixels in the jth SDR image that are mapped to the tth bin, as shown below:

[0125] Π j (t) = ∑Ι(i∈t) (24)

[0126] For the pixel in the jth SDR frame, Ι(·) in the above equation (24) represents the unit function.

[0127] Similarly, a second 3D table can be constructed in the HDR domain Second 3D table (or bins / entries therein) aggregate the Y, C0, and C1 codeword values ​​(e.g., for all HDR pixels in the j-th training HDR image) such that the collected SDR pixels map to the t-th bin as follows:

[0128]

[0129] For the HDR pixel in the j-th training HDR frame of the t-th bin.

[0130] For each image cluster c, let Φc is a set of training SDR and HDR images that are mapped or clustered into image clusters. Cluster-specific 3D tables (e.g., and ) and 3D histograms (e.g., Π c etc.) can be determined as follows:

[0131]

[0132] where p represents the pth SDR or HDR image belonging to the image cluster, or p∈Φ c .

[0133] 3D table and The non-zero entries in can be obtained by dividing the 3D histogram π c The pixel count (or total number) in the indicated bin is averaged. The vector components of the (3-element) vector in are normalized to the normalized value range [0, 1]. The averaging operation can be expressed as follows:

[0134]

[0135] with the same bin index (determined by the underlying SDR codeword value) arrive All mappings can be used to construct a 3D mapping table (3DMT), which converts the 3D SDR table Each bin in is mapped to a 3D HDR table The corresponding bin in .

[0136] These 3D tables can be used to construct a mapping matrix A for a specific image cluster c and B c , as shown below.

[0137] set up is a 3D vector with The average SDR codeword value in the t-th bin of set up is the second 3D vector, having the The average HDR codeword value in the t-th bin of To predict HDR chroma codeword values ​​by MMR for an image cluster, the following vector (with an R vector component) may be constructed for the t-th bin based on the SDR image data for the t-th bin in the training SDR images in the image cluster:

[0138]

[0139] Corresponding MMR coefficients of C0 and C1 channels in HDR domain and It can be expressed as follows:

[0140]

[0141] For an MMR process with second-order MMR coefficients (e.g., R=15), the expected (or predicted) HDR chromaticity values ​​are and It can be obtained as follows:

[0142]

[0143] Let W c It is a 3D table The count (or total number) of non-zero bins in the vector of expected HDR chroma values. and the combined matrix G of SDR values c It can be constructed as follows:

[0144]

[0145] Similarly, the vector of true HDR values ​​determined from the HDR image data in the clustered training HDR images is It can be constructed as follows:

[0146]

[0147] Therefore, the expected (or predicted) HDR value can be obtained by the following MMR process:

[0148]

[0149] MMR coefficient and The optimal value of can be obtained or solved by formulating the optimization problem as minimizing the total approximation error over all bins as follows:

[0150] For channel c0:

[0151] For channel c1:

[0152] This optimization problem can be solved using a linear least squares solver as follows:

[0153]

[0154] In formula (35), let A c =G c T G c , and These mapping matrices A c , and The colorimetric dictionary of each cluster can be calculated separately and formed together with its cluster centroid. As a result, the colorimetric dictionary of all clusters in the plurality of image clusters in the training dataset may include storing the following (e.g., core, primary, used to derive all other quantities, etc.) components:

[0155] A for each cluster c matrix

[0156] Each cluster and matrix

[0157] The cluster centroid Ψ of each cluster c (·)

[0158] The total number of clusters C

[0159] Retrain with updated target values

[0160] The model templates described in the article (e.g., Figure 1 142, etc.) may include multiple ML-generated chromaticity dictionaries for multiple image clusters, their respective cluster centroids including groups of image pairs of training SDR images and corresponding training HDR images. The ML-generated chromaticity dictionary can be used to predict or generate HDR chromaticity codewords for a mapped or inversely shaped HDR image from SDR luminance and chromaticity codewords for an input SDR image. Each of the multiple ML-generated chromaticity dictionaries for corresponding image clusters in the multiple image clusters may include optimized A and B matrices trained using the aforementioned operations represented by equations (19) to (35) above using training SDR images and training HDR images belonging to the corresponding image cluster. Example generation of chromaticity dictionaries is also described in the aforementioned U.S. Provisional Patent Application No. 62 / 781,185.

[0161] A user-defined theme (e.g., corresponding to a particular user-adjusted HDR look or appearance, etc.) can be based on user adjustments made to the model template (142) (e.g., Figure 1 144, etc.). In the user-defined theme, the training SDR images in the training dataset remain unchanged. Therefore, the bins in the aforementioned 3D tables and histograms with values ​​and pixel or codeword counts derived from the SDR image data in the training SDR images do not change. For example, the 3D SDR table remains constant as the user adjusts (144). In addition, all images are clustered automatically using the feature vectors extracted from the training SDR images to obtain the cluster centroid Ψ c (·) remains unchanged. However, 3D HDR table Modified to It is filled with HDR codeword values ​​modified from (e.g., hypothetical, actual, etc.) Values ​​collected or derived, or values ​​collected or derived (e.g., directly, indirectly, etc.) from modified bin values ​​that depend on an HDR image modified according to a user-adjusted HDR look (e.g., hypothetical, actual, etc.).

[0162] After user adjustment (144), the corresponding MMR coefficients of the C0 and C1 channels in the HDR domain and It can be expressed as follows:

[0163]

[0164] For an MMR process with second-order MMR coefficients (e.g., R=15), the expected (or predicted) HDR chrominance values ​​after user adjustment (144) are and It can be obtained as follows:

[0165]

[0166] Note that g t,c and G c are unchanged since they are derived from the SDR image data. After user adjustment (144), the expected (or predicted) HDR chromaticity values ​​for the user-adjusted HDR appearance are and It can be expressed as follows:

[0167]

[0168] in and is the following modified vector for each cluster

[0169]

[0170] Similarly, a vector of ground-truth HDR values ​​determined from HDR image data in a cluster of (e.g., actual, hypothetical, etc.) modified training HDR images according to the user-adjusted HDR appearance is It can be constructed as follows:

[0171]

[0172] MMR coefficient and The optimal value of can be obtained or solved by formulating the optimization problem to minimize the total approximation error over all bins as follows:

[0173] For channel c0:

[0174] For channel c1:

[0175] This optimization problem can be solved using a linear least squares solver as follows:

[0176]

[0177] Among them D c =((G c ) T G c ) -1 (G c ) T .

[0178] For each cluster, the aforementioned direct retraining / recomputation method may require the 3D SDR table to be stored in the memory space or data memory and retrieved from the stored 3D SDR table Export D c In addition, optionally or alternatively, the aforementioned direct retraining / recalculation method may require direct storage of D for each cluster c For D c The size of the data memory may be the total number of non-empty bins multiplied by the total number of non-empty bins, for example 10,000x10,000 per cluster, which corresponds to a relatively large storage consumption. In addition, this direct retraining / recalculation method may also require storing the 3DHDR table for each cluster. So that users can use 3D HDR table Modified to correct 3D HDR table

[0179] In some operational scenarios, the chromaticity dictionary can be retrained using a combined SDR set, as described below. A combined SDR set for all image clusters can be obtained or generated by collecting SDR bin information for all training SDR images in all image clusters, rather than having a separate SDR set for each cluster. More specifically, the combined SDR set includes (i) a combined 3D SDR table Ω s (t), which is cumulatively mapped to the combined 3D SDR table Ω s(i) the Y, C0, and C1 values ​​of all SDR pixels in all training SDR images in the (e.g., entire, etc.) training dataset for all corresponding bins (e.g., the t-th bin, etc.) in the combined 3D histogram Π(t), and (ii) the pixel counts of all SDR pixels in all training SDR images in the (e.g., entire, etc.) training dataset for all corresponding bins (e.g., the t-th bin, etc.) in the combined 3D histogram Π(t) cumulatively mapped to the pixel counts of all SDR pixels in all training SDR images in the (e.g., entire, etc.) training dataset for all corresponding bins (e.g., the t-th bin, etc.) in the combined 3D histogram Π(t), as follows:

[0180]

[0181] Combined 3D SDR table Ω s The bin values ​​in (t) can be normalized to the normalized value range [0, 1] as follows:

[0182] Ω s (t)=Ω s (t) / Π(t) (44)

[0183] set up is Ω s To predict or estimate the HDR chroma codeword value, we can first get the 3D average SDR vector from the t-th bin of (t). Construct the following vector:

[0184]

[0185] The total number of non-empty bins is denoted as W. The SDR value G can be constructed from the vector (t=0, 1, ... (W-1)) of all bins in expression (45) above c The merge matrix is ​​as follows:

[0186]

[0187] Note that the vector g of all clusters indicated in expression (45) t Different from the vector g of the c-th cluster indicated in expression (28) t,c , because the vector g t is generated from the combined SDR bins of all clusters rather than from a specific cluster. As a result, the merge matrix G of all clusters as indicated in expression (46) is also different from the merge matrix G of a specific cluster as indicated in expression (31-2) c .

[0188] Under this combined SDR ensemble approach, the expected (or predicted) HDR values ​​associated with the training HDR images in each cluster before user adjustment (144) and It can be obtained based on the combined SDR set through the MMR process as follows:

[0189]

[0190] Additionally, the expected (or predicted) HDR values ​​associated with the (e.g., hypothetical, actual, etc.) modified training HDR images in each cluster after user adjustment (144) according to the user-adjusted HDR appearance and can be expressed as follows:

[0191]

[0192] in and is the modification vector for each cluster.

[0193]

[0194] Under this combined SDR ensemble approach, the MMR coefficient of each cluster can be obtained as follows:

[0195]

[0196] As mentioned above, the calculation of ((G)) in expression (50) T G) -1 (G) T The operations involved are the same for each cluster. T G) -1 (G) T Obtaining the MMR coefficients becomes a simple matrix multiplication as follows:

[0197]

[0198] A video content creation / production support system (e.g., a cloud-based server) that supports the retraining process performed by a user-operated video production system may store the following information as part of the model template (142):

[0199] Ω s (t) (size is about ~1000×3 matrices)

[0200] Each cluster

[0201] At runtime, the user-operated video production system used by the user to generate an SDR+ encoded bitstream having the user-desired HDR appearance may access information of the model template (142) stored by the video content creation / production support system and use the accessed information to initially construct the following matrix, for example during startup time of the user-operated video production system, as shown below:

[0202] G(from)Ω s (t))

[0203] D=((G) T G) -1 (G) T

[0204]

[0205] For each cluster, a user-operated video production system may interact with the user through a user interface and generate a modified vector for each cluster based on the user input provided by the user according to the user's desired HDR appearance. As shown in expression (48) above, the correction vector can be used to update the following vector representing the expected HDR codeword value according to the user's desired HDR look.

[0206] A user-updated chromaticity dictionary based on the user's desired HDR look can be obtained to include, for each cluster, a cluster-common matrix A and two cluster-specific matrices and As shown below:

[0207] A=G T G (52-1)

[0208]

[0209] Additionally, optionally or alternatively, the user-updated chromaticity dictionary also includes the cluster centroid Ψ of each cluster c (·) and the total number of clusters C.

[0210] The MMR coefficient can be calculated as follows:

[0211]

[0212] Global and local updates of brightness GPR models

[0213] It should be noted that updating (or generating a revised) luminance GPR model as described herein can be achieved by completely direct (or iterative) retraining using a new reference or new target, which is a training HDR image updated (or modified) according to a user-defined HDR look. However, such retraining may take a relatively long time to complete and consume a relatively large amount of computational and other resources.

[0214] As previously mentioned, the techniques described herein can be used to relatively efficiently obtain a revised (or user-updated) luminance GPR model for inclusion in a revised template (e.g., Figure 1146, used to generate compiler metadata or reverse shaping metadata, etc.).

[0215] In some operational scenarios, a modified luminance GPR model in a modified template (146) to be used for generating compiler metadata according to a user-updated HDR look may be derived based on global changes made by the user that affect all training HDR images in the training dataset or / and local instances. In some operational scenarios, a modified luminance GPR model in a modified template (146) to be used for generating compiler metadata according to a user-updated HDR look may be derived based on local or individual changes made by the user to different subsets or / and local instances of all training HDR images in the training dataset. Furthermore, optionally or alternatively, in some operational scenarios, a combination of global and / or local changes to the luminance GPR model may be implemented.

[0216] Global Adjustment

[0217] As a starting point, L ML-generated luminance GPR models (where L is an integer greater than one (1)) can be used to derive (e.g., reference, etc.) ML-generated luminance reverse shaping curves (e.g., BLUTs, sets of polynomials, etc.) that are suitable for reverse shaping all training SDR images in the training dataset into mapped HDR images that approximate all corresponding training HDR images in the training dataset. During the model template training phase, the training HDR images can be (but are not necessarily limited to) professionally color-graded to serve as references, targets, and / or starting points from which users can create their own user-adjusted HDR looks. The user can be allowed to access and modify the ML-generated luminance GPR models used to generate the ML-generated luminance reverse shaping curves, and generate modified luminance reverse shaping curves based on the modified luminance GPR models.

[0218] Figure 2A Interaction with the user to perform some or all user adjustments (e.g., Figure 1 An example graphical user interface (GUI) display (e.g., a web page, etc.) of a system (e.g., a cloud-based portal serving a web page, etc.) of 144, etc., which updates L ML-generated luminance GPR models to L revised (or user-adjusted) luminance GPR models according to the user-adjusted HDR appearance.

[0219] The user's desired global HDR luminance appearance can be achieved by repeatedly, iteratively, progressively, etc., adjusting the L ML-generated luminance GPR models to L modified (or user-adjusted) luminance GPR models until the user completes and / or saves all user adjustments to the L ML-generated luminance GPR models (144). The L modified (or user-adjusted) luminance GPR models can be used to generate a modified (or user-adjusted) inverse shaping curve. The modified inverse shaping curve globally inversely shapes all training SDR images into corresponding training HDR images that are modified (or deemed to be modified) according to the user-adjusted HDR appearance.

[0220] Each of the L (eg, ML-generated, user-adjusted, etc.) GPR models controls a corresponding sampling point of a plurality of different sampling points on an inverse shaping curve (eg, tone curve, inverse tone map, etc.).

[0221] Figure 2A The GUI page of the embodiment includes a brightness band adjustment section 202, in which a plurality of user control components 204-1, 204-2, ... 204-L in the form of a plurality of vertical sliders are presented and can be operated by a user through user input (e.g., click, key, touch screen action, etc.) to adjust a plurality of mapped HDR values ​​in a plurality of sampling points on the inverse shaping curve. Each of the plurality of vertical sliders (e.g., 204-1, 204-2, ... 204-L, etc.) allows the user to adjust the brightness band by a value indicated as δ l A positive or negative (numerical) increment (or an increase or decrease in the mapped HDR value at the corresponding sampling point) controls or adjusts the corresponding mapped HDR value among the multiple mapped HDR values ​​for the multiple sampling points on the inverse shaping curve, where i ranges from 0 to (L-1). Figure 2A The GUI page also includes a user-defined HDR look preview portion having a plurality of display areas (e.g., 208-1 to 208-3, etc.). The plurality of display areas can be used to display a plurality (e.g., three, etc.) of different types of mapped HDR images, such as dark, mid-tone, and / or bright scene mapped HDR images. These mapped HDR images represent the previous looks of the (current) user-defined HDR look and allow the user to have immediate visual feedback on how different types of mapped HDR images are derived by inverse shaping the different types of SDR images based on the (current) corrected inverse shaping curve, the inverse shaping curve including a plurality of sampling points with the (current) user-adjusted mapped HDR values.

[0222] Furthermore, optionally or alternatively, instead of or in addition to the preview of the user-defined HDR look, a modified (eg, average-adjusted, etc.) inverse shaping curve may be displayed.

[0223] In some operating scenarios, user adjustments {δ l} imposes one or more constraints to help ensure that the modified reverse shaping curve is a monotonic function, such as a non-decreasing function. For example, the user can adjust {δ l Implement a simple constraint to make it a non-decreasing sequence as follows:

[0224] δ min ≤δ0≤δ1≤…≤δ L-1 ≤δ max (54)

[0225] where δ min and δ min Indicates minimum and maximum (eg, normalized, etc.) value adjustments, such as -0.2 and +0.2, -0.4 and +0.4, etc.

[0226] Furthermore, optionally or alternatively, a constraint may be imposed on the last or final adjusted sampling point (after consecutively adding all previous δ l After that), the mapped HDR value of the last or final adjusted sampling point is made to be within a specific HDR codeword value range, which is defined by SMPTE 2084 or ITU-R Rec.BT.2100, for example.

[0227] Category Adjustment

[0228] A user may want to create different HDR looks for images of different categories. Example categories of images with different HDR looks may include, but are not limited to, dark scenes, mid-tone scenes, bright scenes, scenes with faces present, scenes without faces present, landscape scenes, automatically generated image clusters, etc.

[0229] A simple solution is to create different adjustment sliders for each category. Figure 2B As shown, the desired category-specific HDR appearance of images of category 0 may be adjusted by a set of bars in the corresponding luminance band adjustment portion 202-1, while the desired category-specific HDR appearance of images of category 1 may be adjusted by a set of bars in the corresponding luminance band adjustment portion 202-2.

[0230] Assume Γ d is the d-th (e.g., mutually exclusive, etc.) subset of training SDR and HDR image pairs from a plurality (D) of subsets of training SDR and HDR image pairs from the (original) training dataset (which has F training SDR and HDR image pairs), where d is an integer between 1 and D. For each training SDR image (e.g., the j-th SDR image), the feature vector x jcan be extracted from each such training SDR image and used to classify each training SDR and HDR image pair containing each such SDR image into a different one of the D subsets.

[0231] By way of example and not limitation, it is indicated that The average brightness value (e.g., average picture level or APL, etc.) of can be calculated for each training SDR image and used as an indicator to classify the corresponding training SDR and HDR image pair containing each such training SDR image into the aforementioned D subsets denoted as Γ d The d-th subset of d It can be represented by {λ d-1} and {λ d}Partition boundaries are delineated as follows:

[0232]

[0233] For each subset (or image category) indexed by d, the lth sample point estimated or predicted by the lth GPR model can be adjusted by the user δ d,l The mapped (or target) HDR value of the lth sampling point can be derived as follows:

[0234]

[0235] The overall mapped (or target) HDR value vector can then be constructed as follows:

[0236]

[0237] Note that the user's adjustment {δ d,l} may still be constrained to be non-decreasing, as follows:

[0238] δ min ≤δ d,0 ≤δ d,1 ≤…≤δ d,L-1 ≤δ max (58)

[0239] However, scaling each class of an image by a constant for each sampling point may lead to visually perceptible issues, such as transition issues of images with different visual characteristics and features but still within the same boundaries of the class partition or between consecutive classes of an image.

[0240] In some operational scenarios, rather than applying the same inverse shaping curve to all images of a given class defined by the class split boundary, images from adjacent classes can be used to implement a soft transition to mitigate or resolve these issues in the correction model used to generate the inverse shaping metadata.

[0241] For example, user-adjusted interpolation can be implemented based on the feature vector. Consider the dth subset (or image category), for the image with the boundary between the two categories The average brightness value within The first distance (denoted as ω) from the center of the left class to the jth image (of the extracted features) d-1 ) and the second distance to the center of the right category (ω d ) can be calculated as follows:

[0242]

[0243] The new user adjustment value (δ j,d,l ) can be to two adjacent centers ω d-1 and ω d A weighted version of the distance is as follows:

[0244]

[0245] In some operational scenarios, a GUI display page may be provided to allow the user to control some or all of the following quantities: Partition Boundary {λ d}、{δ d,l}, and / or do not allow any partition boundaries for interpolation-based user-adjusted categories; instead, the user can only generate (non-interpolation-based) user-adjusted {δ d,l}.

[0246] User adjustment of chroma

[0247] It should be noted that, as in the case of the luma GPR model, updating (or generating a revised) chroma dictionary as described herein can be achieved by completely direct (or iterative) retraining using a new reference or new target, which is a training HDR image that has been updated (or modified) according to the user-defined HDR look. However, such retraining may take a relatively long time to complete, require a relatively large amount of data to implement such retraining, and consume a relatively large amount of computational and other resources.

[0248] As previously mentioned, the techniques described herein can be used to relatively efficiently obtain a revised (or user-updated) chromaticity dictionary for inclusion in a revised template (e.g., Figure 1 146, used to generate compiler metadata or reverse shaping metadata, etc.).

[0249] Example user adjustments to the chromaticity dictionary include, but are not necessarily limited to, global user adjustments that apply to all image clusters in the training dataset. These global user adjustments can include a saturation adjustment that allows the user to globally adjust color saturation across different brightness ranges, and a hue adjustment that allows the user to globally adjust hue across different brightness ranges for all image clusters in the training set.

[0250] In some operational scenarios, to perform global user adjustment using saturation adjustment, a modified target chrominance value vector for each image cluster (e.g., c-th cluster, etc.), or an expected (or predicted) HDR value associated with the modified training HDR images in each cluster (e.g., hypothetical, actual, etc.) after user adjustment according to the user-adjusted HDR appearance, can be constructed using addition as follows:

[0251]

[0252] In some operational scenarios, to perform global user adjustment with saturation adjustment, a modified target chrominance value vector for each image cluster (e.g., c-th cluster, etc.), or an expected (or predicted) HDR value associated with the modified training HDR images in each cluster (e.g., hypothetical, actual, etc.) after user adjustment according to the user-adjusted HDR appearance, can be constructed using a luminance modulation function as follows:

[0253]

[0254] The brightness modulation function f c0 () and f c1 () is a scaling function based on (or dependent on) brightness. By setting f c0 () = 1 and f c1 ()=1 can simplify the above expression (62) to the above expression (61). In various embodiments, The brightness modulation function f c0 () and f c1 () can be determined based on user input, determined based on heuristics, determined based on empirical studies of training data, etc. In addition, optionally or alternatively, the brightness modulation function f c0 () and f c1 () can be represented as a lookup table.

[0255] In some operational scenarios, in order to perform global user adjustments using hue adjustments, a simple solution can be achieved by rotating, as shown below:

[0256]

[0257] where θ t,c Indicates HDR brightness The brightness correction function is as follows:

[0258]

[0259] Brightness modulation function is a scaling function based on (or dependent on) brightness. In various embodiments, Brightness modulation function It can be determined based on user input, determined based on heuristics, determined based on empirical studies of training data, etc. In addition, optionally or alternatively, the brightness modulation function can be represented as a lookup table.

[0260] Local Adjustment

[0261] Similar to local user adjustment of luminance, local user adjustment of chrominance can be performed on image categories. In some operational scenarios, image categories associated with local user adjustment of chrominance can be divided based on clustered SDR Cb / Cr values.

[0262] As mentioned previously, the predicted or estimated HDR chroma codeword values ​​in each cluster can be calculated based on the combined SDR set as follows:

[0263]

[0264] where G is computed using the combined SDR set and applied to each cluster.

[0265] The mean predicted or estimated HDR chroma codeword value for each channel in each cluster can be given as follows:

[0266]

[0267] where i is the training image pair index of all training image pairs in each such cluster of training image pairs in the training dataset.

[0268] In some operating scenarios, such as Figure 3A As shown, the HDR chroma codeword value can be predicted or estimated using the mean value Distribution For example, based on the mean predicted or estimated HDR chroma codeword values Relative to Figure 3A The angle of the horizontal axis, the distribution can be grouped or divided into Mean predicted or estimated HDR chroma codeword value Distribution Each of the multiple areas can be adjusted independently.

[0269] For example, you can use Figure 3A The lines with the Cartesian origin (0, 0) are used to group or divide the g-th region as follows:

[0270] y-λ g x (67)

[0271] Where x and y represent the mean predicted or estimated HDR chrominance codeword values ​​respectively. These values ​​are close to the slope λ g The line, or the line with slope λ g The lines represent the same g-th region.

[0272] Alternatively, in a polar coordinate system, the angle of the mean predicted or estimated HDR chroma codeword value can be calculated for each cluster as follows:

[0273]

[0274] Figure 3B An example distribution of angles of clusters is shown, where the angles can be calculated using expression (68) above.

[0275] exist Figure 3A and Figure 3B The system described herein can interact with the user and allow the user to provide input to select the partition angle {θ g} and / or the total number of regions, where g is in 0, ..., G-1. Saturation and hue can be adjusted individually for each region.

[0276] Example Processing Flow

[0277] Figure 4AAn exemplary processing flow according to an embodiment of the present invention is shown. In some embodiments, one or more computing devices or components (e.g., an encoding device / module, a transcoding device / module, a decoding device / module, an inverse tone mapping device / module, a tone mapping device / module, a media device / module, a prediction model and feature selection system, an inverse mapping generation and application system, etc.) can perform the processing flow. In box 402, the image processing system accesses a model template including an inverse shaping metadata prediction model. The inverse shaping metadata prediction model is trained with a plurality of training image feature vectors from a plurality of training standard dynamic range (SDR) images in a plurality of training image pairs and true values ​​derived from a plurality of corresponding training high dynamic range (HDR) images in the plurality of training image pairs. Each training image pair in the plurality of training image pairs includes a training SDR image in the plurality of training SDR images and a corresponding training HDR image in the plurality of corresponding training HDR images. The training SDR image and the corresponding training HDR image in each such training image pair depict the same visual content but have different luminance dynamic ranges.

[0278] At block 404 , the image processing system receives content creation user input defining one or more content creation user adjusted HDR looks for a plurality of corresponding training HDR images.

[0279] In block 406 , the image processing system generates a content creation user-specific modified inverse shaping metadata prediction model based on the model template and the content creation user input.

[0280] In block 408, the image processing system uses the content creation user specific modified reverse shaping metadata prediction model to predict operational parameter values ​​of a content creation user specific reverse shaping mapping for reverse shaping the SDR image into a mapped HDR image of at least one of the one or more content creation user adjusted HDR looks.

[0281] In one embodiment, the one or more inverse shaping metadata prediction models include multiple Gaussian process regression (GPR) models for predicting a luma inverse shaping map to inversely shape an input luma SDR codeword into a mapped luma HDR codeword. A content creation user input modifies multiple sample points of the luma inverse shaping map.

[0282] In one embodiment, the plurality of sample points modified by the content creation user input are constrained to maintain the luminance inverse shaping map as a monotonically increasing function.

[0283] In one embodiment, a plurality of image pairs are categorized into a plurality of image categories; and the content creation user input modifies the luminance inverse shaping map differently for at least two image categories of the plurality of image categories.

[0284] In one embodiment, the content creation user input modifies a luminance inverse shaping map applied to all image pairs in the plurality of image pairs.

[0285] In one embodiment, one or more inverse shaping metadata prediction models include a set of multivariate multiple regression (MMR) mapping matrices for generating MMR coefficients for generating mapped chroma HDR codewords from input SDR codewords; wherein content creation user input utilizes multiplication operations to modify an appropriate subset of MMR mapping matrices in the set of MMR mapping matrices; and the remaining MMR mapping matrices in the set of MMR mapping matrices are exempted from modification by content creation user input.

[0286] In one embodiment, a plurality of image pairs are categorized into a plurality of image categories; and for at least two of the plurality of image categories, the content creation user input differently modifies an appropriate subset of the MMR mapping matrix.

[0287] In one embodiment, multiple image classes are classified based on multiple regions, each region including a different set of mean predicted Cb values ​​and mean predicted Cr values.

[0288] In one embodiment, multiple image categories are classified based on multiple different angular sub-ranges formed by different combinations of mean predicted Cb values ​​and mean predicted Cr values.

[0289] In one embodiment, the content creation user input modifies an appropriate set of MMR mapping matrices applicable to all image pairs in the plurality of image pairs.

[0290] In one embodiment, the image processing system is further configured to: encode one or more of the values ​​of the operation parameters of the inverse shaping mapping for inversely shaping the SDR image into the mapped HDR image together with the SDR image into a video signal as image metadata, wherein the video signal causes one or more receiving devices to render a display image derived from the mapped HDR image using one or more display devices.

[0291] In one embodiment, the reverse shaping metadata prediction model in the model template includes a set of hyperparameter values ​​and a set of weight factor values; by changing the set of weight factor values ​​while keeping the set of hyperparameter values ​​unchanged, content is exported from the reverse shaping metadata prediction model in the model template to create a user-specific modified reverse shaping metadata prediction model.

[0292] Figure 4BAn exemplary process flow according to an embodiment of the present invention is shown. In some embodiments, one or more computing devices or components (e.g., an encoding device / module, a transcoding device / module, a decoding device / module, an inverse tone mapping device / module, a tone mapping device / module, a media device / module, a prediction model and feature selection system, an inverse mapping generation and application system, etc.) may execute the process flow. In block 452, a video decoding system decodes a standard dynamic range (SDR) image from a video signal. The SDR image is to be inversely reshaped into a corresponding mapped high dynamic range (HDR) image.

[0293] In block 454 , the video decoding system decodes compiler metadata from the video signal, the compiler metadata being used to derive one or more operational parameter values ​​of the content user-specific inverse shaping map.

[0294] One or more operational parameter values ​​of a content user-specific reverse shaping mapping are predicted by creating one or more content user-specific modified reverse shaping metadata prediction models.

[0295] One or more content creation user-specific modified reverse shaping metadata prediction models are generated based on the model template and the content creation user input.

[0296] The model template includes an inverse shaping metadata prediction model. The inverse shaping metadata prediction model is trained using a plurality of training image feature vectors from a plurality of training SDR images in a plurality of training image pairs and ground truth values ​​derived from a plurality of corresponding training HDR images in the plurality of training image pairs, wherein each training image pair in the plurality of training image pairs includes a training SDR image in the plurality of training SDR images and a corresponding training HDR image in the plurality of corresponding training HDR images, the training SDR image and the corresponding training HDR image in each such training image pair depicting the same visual content but having different luminance dynamic ranges.

[0297] The content creation user input modifies the plurality of corresponding training HDR images into one or more content creation user adjusted HDR looks.

[0298] In block 456 , the video decoding system uses the one or more operational parameter values ​​of the content user-specific inverse shaping mapping to inversely reshape the SDR image into a mapped HDR image of at least one of the one or more content-creation user-adjusted HDR looks;

[0299] In block 458 , the video decoding system causes rendering by a display device of a display image derived from the mapped HDR image.

[0300] In one embodiment, a computing device such as a display device, mobile device, set-top box, multimedia device, etc. is configured to perform any of the aforementioned methods. In one embodiment, an apparatus includes a processor and is configured to perform any of the aforementioned methods. In one embodiment, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, result in the performance of any of the aforementioned methods.

[0301] In one embodiment, a computing device includes one or more processors and one or more storage media storing a set of instructions that, when executed by the one or more processors, result in performance of any of the aforementioned methods.

[0302] Note that although separate embodiments are discussed herein, any combination of embodiments and / or portions of embodiments discussed herein may be combined to form further embodiments.

[0303] Example Computer System Implementation

[0304] Embodiments of the present invention may be implemented using a computer system, a system configured in electronic circuits and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA) or another configurable or programmable logic device (PLD), a discrete-time or digital signal processor (DSP), an application-specific IC (ASIC), and / or an apparatus comprising one or more such systems, devices, or components. The computer and / or IC may execute, control, or perform instructions related to adaptive perceptual quantization of images with enhanced dynamic range, such as described herein. The computer and / or integrated circuit may calculate any of the various parameters or values ​​associated with the adaptive perceptual quantization process described herein. The image and video embodiments may be implemented using hardware, software, firmware, and various combinations thereof.

[0305] Certain implementations of the present invention include a computer processor that executes software instructions that cause the processor to perform the method of the present disclosure. For example, one or more processors in a display, an encoder, a set-top box, a transcoder, etc. can implement the method related to adaptive perceptual quantization of HDR images as described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention can also be provided in the form of a program product. The program product may include any non-transitory medium that carries a set of computer-readable signals, the set of computer-readable signals including instructions that, when executed by a data processor, cause the data processor to perform the method of the embodiment of the present invention. The program product according to an embodiment of the present invention may be in any of a variety of forms. The program product may include, for example, a magnetic data storage medium including a floppy disk, a physical medium such as a hard disk drive, an optical data storage medium including a CD ROM, a DVD, an electronic data storage medium including a ROM, a flash memory, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0306] When reference is made above to a component (e.g., a software module, a processor, a component, a device, a circuit, etc.), unless otherwise specified, reference to the component (including reference to "means") should be interpreted as including any component that performs the function of the component as an equivalent of the component (e.g., functional equivalent), including components that are not structurally equivalent to the disclosed structure for performing the function in the exemplary embodiments shown in the present invention.

[0307] According to one embodiment, the technology described herein is implemented by one or more special-purpose computing devices. Special-purpose computing devices can be hard-wired to perform these technologies, or can include digital electronic devices that are permanently programmed to perform these technologies, such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs), or can include one or more general-purpose hardware processors that are programmed to perform these technologies according to program instructions in firmware, memory, other storage, or a combination thereof. Such special-purpose computing devices can also be combined with custom hard-wired logic, ASICs, or field programmable gate arrays to implement these technologies through custom programming. Special-purpose computing devices can be desktop computer systems, portable computer systems, handheld devices, network devices, or any other device that combines hard-wiring and / or program logic to implement these technologies.

[0308] For example, Figure 5 5 is a block diagram illustrating a computer system 500 on which embodiments of the present invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled to bus 502 for processing information. Hardware processor 504 may be, for example, a general-purpose microprocessor.

[0309] The computer system 500 also includes a main memory 506, such as a random access memory or other dynamic storage device, coupled to the bus 502 for storing information and instructions to be executed by the processor 504. The main memory 506 may also be used for storing temporary variables or other intermediate information during execution of instructions by the processor 504. These instructions, when stored in a non-transitory storage medium accessible to the processor 504, render the computer system 500 as a special-purpose machine customized to perform the operations specified in the instructions.

[0310] Computer system 500 also includes a read-only memory 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, such as a magnetic or optical disk, is provided and coupled to bus 502 for storing information and instructions.

[0311] The computer system 500 may be coupled to a display 512, such as a liquid crystal display, via the bus 502 for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to the bus 502 for communicating information and command selections to the processor 504. Another type of user input device is a cursor controller 516, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 504 and for controlling cursor movement on the display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane.

[0312] Computer system 500 can implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, enables computer system 500 or programs computer system 500 to function as a special-purpose machine. According to one embodiment, the techniques described herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0313] The term "storage media" as used herein refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, ROM and EPROM, flash EPROM, non-volatile random access memory, any other memory chip or cartridge.

[0314] Storage media are distinct from transmission media, but can be used in conjunction with them. Transmission media participate in the transmission of information between storage media. Examples of transmission media include coaxial cables, copper wire, and optical fiber, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or optical waves, such as those generated during radio wave and infrared data communications.

[0315] When one or more sequences of one or more instructions are transmitted to the processor 504 for execution, various forms of media may be involved. For example, the instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the computer system 500 may receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on the bus 502. The bus 502 transmits the data to the main memory 506, from which the processor 504 retrieves and executes the instructions. The instructions received by the main memory 506 may optionally be stored on the storage device 510 before or after execution by the processor 504.

[0316] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides two-way data communication coupled to network link 520, which is connected to local network 522. For example, communication interface 518 can be an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 can be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links can also be implemented. In any such implementation, communication interface 518 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0317] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 can provide a connection to a host computer 524 or to data equipment operated by an Internet service provider (ISP) 526 through a local network 522. ISP 526, in turn, provides data communication services through a global packet data communication network 528, now commonly referred to as the "Internet." Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 520 and through communication interface 518 are example forms of transmission media that carry the digital data to and from computer system 500.

[0318] Computer system 500 can send messages and receive data, including program code, through the network, network link 520, and communication interface 518. In the Internet example, server 530 can transmit the requested code for an application through Internet 528, ISP 526, local network 522, and communication interface 518.

[0319] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510 or other non-volatile storage for later execution.

[0320] Equivalents, extensions, alternatives and others

[0321] In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details, which may vary from implementation to implementation. Thus, the sole and exclusive indicator of the embodiments for which protection is claimed, and which the applicants intend to claim, is the set of claims issuing from this application, and the specific form in which such claims issue, including any subsequent correction. Any express definitions herein for terms contained in a claim shall govern the meaning of the terms used in the claims. Accordingly, any limitation, element, attribute, feature, advantage, or property that is not expressly recited in a claim should not limit the scope of such claim in any way. The specification and drawings are, therefore, to be regarded in an illustrative rather than a restrictive manner.

[0322] Examples of Examples

[0323] The present invention may be embodied in any form described herein, including but not limited to the following Enumerated Example Embodiment (EEE), which describes the structure, features, and functions of some portions of embodiments of the present invention.

[0324] EEE 1. A method comprising:

[0325] accessing a model template comprising an inverse shaping metadata prediction model, wherein the inverse shaping metadata prediction model is trained using a plurality of training image feature vectors from a plurality of training standard dynamic range (SDR) images in a plurality of training image pairs and ground truth values ​​derived from a plurality of corresponding training high dynamic range (HDR) images in the plurality of training image pairs, wherein each training image pair in the plurality of training image pairs comprises a training SDR image in the plurality of training SDR images and a corresponding training HDR image in the plurality of corresponding training HDR images, wherein the training SDR image and the corresponding training HDR image in each such training image pair depict the same visual content but have different luminance dynamic ranges;

[0326] receiving content creation user input defining one or more content creation user adjusted HDR looks for the plurality of corresponding training HDR images;

[0327] Generate a content creation user-specific modified reverse shaping metadata prediction model based on the model template and content creation user input;

[0328] The content creation user specific modified reverse shaping metadata prediction model is used to predict operational parameter values ​​of a content creation user specific reverse shaping mapping for reverse shaping the SDR image into a mapped HDR image of at least one of the one or more content creation user adjusted HDR looks.

[0329] EEE 2. The method according to EEE 1, wherein the one or more reverse shaping metadata prediction models include multiple Gaussian process regression (GPR) models for predicting a luma reverse shaping mapping for reverse shaping an input luma SDR codeword into a mapped luma HDR codeword; wherein the content creation user input modifies multiple sample points of the luma reverse shaping mapping.

[0330] EEE 3. The method according to EEE 2, wherein the plurality of sample points modified by the content creation user input are constrained to maintain the luminance inverse shaping map as a monotonically increasing function. ...

[0331] EEE 4. The method according to any one of EEE 2, wherein the plurality of image pairs are classified into a plurality of image categories; wherein, for at least two image categories of the plurality of image categories, the content creation user input modifies the luminance inverse shaping map differently. EEE 4.

[0332] EEE 5. The method according to EEE 2, wherein the content creation user input modifies a luminance inverse shaping map applied to all image pairs in the plurality of image pairs. ...

[0333] EEE 6. A method according to EEE 1, wherein the one or more inverse shaping metadata prediction models include a set of multivariate multiple regression (MMR) mapping matrices for generating MMR coefficients for generating mapped chroma HDR codewords from input SDR codewords; wherein the content creation user input utilizes a multiplication operation to modify an appropriate subset of the MMR mapping matrices in the set of MMR mapping matrices; and wherein the remaining MMR mapping matrices in the set of MMR mapping matrices are exempt from modification by the content creation user input.

[0334] EEE 7. The method of EEE 6, wherein the plurality of image pairs are classified into a plurality of image categories; wherein the content creation user input modifies the appropriate subset of the MMR mapping matrix differently for at least two of the plurality of image categories. EEE 7. The method of claim 6, wherein the plurality of image pairs are classified into a plurality of image categories; and wherein the content creation user input modifies the appropriate subset of the MMR mapping matrix differently for at least two of the plurality of image categories.

[0335] EEE 8. The method according to EEE 7, wherein the plurality of image classes are classified based on a plurality of regions, each of the plurality of regions comprising a different set of mean predicted Cb values ​​and mean predicted Cr values. ...

[0336] EEE 9. The method according to EEE 7, wherein the plurality of image categories are classified based on a plurality of different angular sub-ranges formed by different combinations of mean predicted Cb values ​​and mean predicted Cr values. ...

[0337] EEE 10. The method of EEE 6, wherein the content creation user input modifies the appropriate subset of MMR mapping matrices applied to all image pairs in the plurality of image pairs. EEE 11. The method of claim 6, wherein the content creation user input modifies the appropriate subset of MMR mapping matrices applied to all image pairs in the plurality of image pairs.

[0338] EEE 11. The method according to EEE 1 also includes: encoding one or more of the operating parameter values ​​of the reverse shaping mapping for reverse shaping the SDR image into the mapped HDR image together with the SDR image into a video signal as image metadata, wherein the video signal enables one or more receiving devices to render a display image derived from the mapped HDR image using one or more display devices.

[0339] EEE 12. The method according to EEE 1, wherein the reverse shaping metadata prediction model in the model template includes a set of hyperparameter values ​​and a set of weight factor values; wherein, by changing the set of weight factor values ​​while keeping the set of hyperparameter values ​​unchanged, a user-specific modified reverse shaping metadata prediction model is created by deriving content from the reverse shaping metadata prediction model in the model template.

[0340] EEE 13. A method comprising:

[0341] decoding a standard dynamic range (SDR) image from a video signal to be inversely reshaped into a corresponding mapped high dynamic range (HDR) image;

[0342] decoding compiler metadata from the video signal, the compiler metadata for deriving one or more operating parameter values ​​of a content user-specific inverse shaping map;

[0343] wherein one or more operational parameter values ​​of a content user-specific reverse shaping mapping are predicted by one or more content creation user-specific modified reverse shaping metadata prediction models;

[0344] wherein one or more content creation user-specific modified reverse shaping metadata prediction models are generated based on the model template and the content creation user input;

[0345] wherein the model template comprises an inverse shaping metadata prediction model, wherein the inverse shaping metadata prediction model is trained using a plurality of training image feature vectors from a plurality of training SDR images in a plurality of training image pairs and ground truth values ​​derived from a plurality of corresponding training HDR images in the plurality of training image pairs, wherein each training image pair in the plurality of training image pairs comprises a training SDR image in the plurality of training SDR images and a corresponding training HDR image in the plurality of corresponding training HDR images, wherein the training SDR image and the corresponding training HDR image in each such training image pair depict the same visual content but have different luminance dynamic ranges;

[0346] wherein the content creation user input modifies the plurality of corresponding training HDR images into one or more content creation user-adjusted HDR looks;

[0347] inversely reshape the SDR image into a mapped HDR image of at least one of the one or more content-creation user-adjusted HDR looks using one or more operational parameter values ​​of the content user-specific inverse shaping mapping;

[0348] A display image derived from the mapped HDR image is caused to be rendered by a display device.

[0349] EEE 14. A computer system configured to perform the method according to any one of EEE 1-13.

[0350] EEE 15. An apparatus comprising a processor and configured to perform the method according to any one of EEE 1-13.

[0351] EEE 16. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon for executing the method according to any one of EEE 1-13.

Claims

1. An image processing method, comprising: accessing a model template comprising an inverse shaping metadata prediction model, wherein the inverse shaping metadata prediction model is trained using a plurality of training image feature vectors from a plurality of training standard dynamic range (SDR) images in a plurality of training image pairs and ground truth values ​​derived from a plurality of corresponding training high dynamic range (HDR) images in the plurality of training image pairs, wherein each training image pair in the plurality of training image pairs comprises a training SDR image in the plurality of training SDR images and a corresponding training HDR image in the plurality of corresponding training HDR images, wherein the training SDR image and the corresponding training HDR image in each such training image pair depict the same visual content but have different luminance dynamic ranges; receiving content creation user input defining one or more content creation user-adjusted HDR looks for the plurality of corresponding training HDR images; Generate a content creation user-specific modified reverse shaping metadata prediction model based on the model template and content creation user input; predicting operational parameter values ​​of a content creation user specific reverse shaping mapping for reverse shaping the SDR image into a mapped HDR image of at least one of the one or more content creation user adjusted HDR looks using a content creation user specific modified reverse shaping metadata prediction model, wherein a plurality of image pairs are categorized into a plurality of image categories; and wherein, for at least two image categories of the plurality of image categories, the content creation user input modifies the luminance inverse shaping map differently.

2. The method of claim 1 , wherein the one or more inverse shaping metadata prediction models comprise a plurality of Gaussian process regression (GPR) models for predicting a luma inverse shaping map for inversely shaping an input luma SDR codeword into a mapped luma HDR codeword; and wherein a content creation user input modifies a plurality of sample points of the luma inverse shaping map.

3. The method according to claim 2, wherein: The plurality of sample points modified by the content creation user input are constrained to maintain the luminance inverse shaping map as a monotonically increasing function.

4. The method according to claim 1, wherein The content creation user input modifies a luminance inverse shaping map applied to all image pairs in the plurality of image pairs.

5. The method according to claim 1, wherein The one or more inverse shaping metadata prediction models include a set of multivariate multiple regression MMR mapping matrices for generating MMR coefficients for generating mapped chroma HDR codewords from input SDR codewords; wherein the content creation user input utilizes a multiplication operation to modify an appropriate subset of the MMR mapping matrices in the set of MMR mapping matrices; and wherein the remaining MMR mapping matrices in the set of MMR mapping matrices are exempted from modification by the content creation user input.

6. The method of claim 5, wherein the plurality of image pairs are classified into a plurality of image categories; wherein, The content creation user input modifies the appropriate subset of the MMR mapping matrix differently for at least two image categories of the plurality of image categories.

7. The method of claim 5, wherein the plurality of image classes are classified based on a plurality of regions, each of the plurality of regions comprising a different set of mean predicted Cb values ​​and mean predicted Cr values.

8. The method of claim 5, wherein the plurality of image categories are classified based on a plurality of different angular sub-ranges formed by different combinations of mean predicted Cb values ​​and mean predicted Cr values.

9. The method according to claim 5, wherein: The content creation user input modification applies to the appropriate subset of MMR mapping matrices for all image pairs in the plurality of image pairs.

10. The method according to any one of claims 1-9 further comprises encoding one or more of the values ​​of the operating parameters of the inverse reshaping mapping for inverse reshaping the SDR image into the mapped HDR image together with the SDR image into a video signal as image metadata, wherein the video signal enables one or more receiving devices to render a display image derived from the mapped HDR image using one or more display devices.

11. The method according to any one of claims 1 to 9, wherein The reverse shaping metadata prediction model in the model template includes a set of hyperparameter values ​​and a set of weight factor values; wherein, by changing the set of weight factor values ​​while keeping the hyperparameter value set unchanged, content is exported from the reverse shaping metadata prediction model in the model template to create a user-specific modified reverse shaping metadata prediction model.

12. An image processing method comprising: decoding a standard dynamic range (SDR) image from the video signal to be inversely reshaped into a corresponding mapped high dynamic range (HDR) image; decoding compiler metadata from the video signal, the compiler metadata for deriving one or more operating parameter values ​​of a content user-specific inverse shaping map; wherein one or more operational parameter values ​​of a content user-specific reverse shaping mapping are predicted by one or more content creation user-specific modified reverse shaping metadata prediction models; wherein one or more content creation user-specific modified reverse shaping metadata prediction models are generated based on the model template and the content creation user input; wherein the model template comprises an inverse shaping metadata prediction model, wherein the inverse shaping metadata prediction model is trained using a plurality of training image feature vectors from a plurality of training SDR images in a plurality of training image pairs and ground truth values ​​derived from a plurality of corresponding training HDR images in the plurality of training image pairs, wherein each training image pair in the plurality of training image pairs comprises a training SDR image in the plurality of training SDR images and a corresponding training HDR image in the plurality of corresponding training HDR images, wherein the training SDR image and the corresponding training HDR image in each such training image pair depict the same visual content but have different luminance dynamic ranges; wherein the content creation user input modifies the plurality of corresponding training HDR images into one or more content creation user-adjusted HDR looks; inversely reshape the SDR image into a mapped HDR image of at least one of the one or more content-creation user-adjusted HDR looks using one or more operational parameter values ​​of the content user-specific inverse shaping mapping; causing a display image derived from the mapped HDR image to be rendered by a display device, wherein a plurality of image pairs are categorized into a plurality of image categories; and wherein, for at least two image categories of the plurality of image categories, the content creation user input modifies the luminance inverse shaping map differently.

13. An image processing device comprising a processor and configured to execute the method according to any one of claims 1 to 12.

14. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-12.

15. An image processing device comprising: processor; as well as A storage medium having computer-executable instructions stored thereon, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 12.

16. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 12.

17. An image processing apparatus comprising means for performing the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Signal reshaping approximation

    US20180020224A1

  • Multiple color channel multiple regression predictor

    US8811490B2

  • Screen-adaptive decoding of high dynamic range video

    US20180115777A1