Image prediction for hdr imaging in open-loop codecs

By generating noise data of noise intensity in an open-loop codec, combining it with HDR and SDR images to generate an enhanced input dataset, constructing a prediction model and minimizing the error, the problem of low image prediction efficiency in open-loop codecs is solved, and high-quality image reconstruction on HDR display devices is achieved efficiently.

CN116157824BActive Publication Date: 2026-05-12DOLBY LABORATORIES LICENSING CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2021-06-21
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In the existing technology, open-loop codecs lack efficient image prediction technology in high dynamic range (HDR) imaging, resulting in low coding efficiency and the inability to reconstruct high-quality HDR images on traditional SDR display devices.

Method used

By generating noise data with noise intensity, combining it with HDR and SDR images to generate an enhanced input dataset, constructing a prediction model and minimizing the error to generate prediction model parameters, compressing the image and generating a bitstream, image reconstruction is achieved.

Benefits of technology

It improves the encoding efficiency of open-loop codecs, enabling the reconstruction of high-quality HDR images on HDR display devices while reducing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116157824B_ABST
    Figure CN116157824B_ABST
Patent Text Reader

Abstract

A prediction model to predict an HDR image from a compressed representation of an input SDR image is generated in the following manner, given input HDR and SDR images representing the same scene: a) generating noise data based on at least a characteristic of the HDR image; b) generating a noisy SDR image by adding the noise data to the SDR image; c) generating an augmented HDR dataset and an augmented SDR dataset by using the input HDR and SDR images and the noisy SDR image; d) generating the prediction model to predict the augmented HDR dataset based on the augmented SDR dataset; and e) solving the prediction model according to a minimization error criterion to generate a set of prediction parameters to be transmitted to a decoder with the compressed representation of the input SDR image to reconstruct an approximation of the input HDR image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to European Patent Application No. 20182014.9 and U.S. Provisional Application No. 63 / 043,198, both filed on June 24, 2020, each of which is incorporated herein by reference in its entirety. Technical Field

[0003] This invention generally relates to images. More specifically, embodiments of the invention relate to image prediction for high dynamic range (HDR) imaging in open-loop codecs. Background Technology

[0004] As used herein, the term 'dynamic range (DR)' can refer to the ability of the human visual system (HVS) to perceive a range of intensity (e.g., luminance, luma) in an image, such as from the darkest gray (black) to the brightest white (highlight). In this sense, DR relates to the intensity of a 'scene-referred' image. DR can also refer to the ability of a display device to fully or approximately render a specific breadth of intensity. In this sense, DR relates to the intensity of a 'display-referred' image. Unless a particular meaning is explicitly specified to have a specific connotation at any point in the description herein, it should be inferred that the terms can be used interchangeably in either sense.

[0005] As used in this article, the term "high dynamic range (HDR)" refers to a DR width spanning 14 to 15 orders of magnitude across the human visual system (HVS). In reality, the DR, which humans can perceive simultaneously across a wide range of intensity, may be slightly truncated relative to HDR.

[0006] In fact, an image comprises one or more color components (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented by a pixel. n Bit precision representation (e.g., n = 8). Use linear or gamma brightness encoding, where n Images with a dynamic range ≤ 8 (e.g., a color 24-bit JPEG image) are considered to have a standard dynamic range, where... n Images with a dynamic range greater than 8 can be considered enhanced or high dynamic range images. HDR images can also be stored and distributed using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR document format developed by Industrial Light and Magic.

[0007] Most consumer desktop monitors currently support 200 to 300 cd / m³. 2 Or nits of brightness. Most consumer HDTVs range from 300 to 500 nits, with newer models reaching 1000 nits (cd / m²). 2 Therefore, such conventional displays represent the lower dynamic range (LDR) associated with HDR, also known as standard dynamic range (SDR). As the availability of HDR content increases due to advancements in both capture devices (e.g., cameras) and HDR displays (e.g., Dolby Laboratories' PRM-4200 professional reference monitor), HDR content can be color-graded and displayed on HDR displays that support higher dynamic ranges (e.g., from 1,000 nits to 5,000 nits or higher).

[0008] As used herein, the terms "reshaping" or "remapping" refer to the process of mapping a digital image from its original bit depth and original codeword distribution or representation (e.g., gamma, PQ, or HLG) to images with the same or different bit depths and different codeword distributions or representations, either sample-to-sample or codeword-to-codeword. Reshaping allows for improved compressibility or image quality at a fixed bit rate. For example, without limitation, forward reshaping can be applied to 10-bit or 12-bit PQ-coded HDR video to improve coding efficiency in 10-bit video coding architectures. In a receiver, after the received signal is decompressed (with or without reshaping), the receiver can apply an inverse (or backward) reshaping function to restore the signal to its original codeword distribution and / or achieve a higher dynamic range.

[0009] In HDR coding, image prediction (or shaping) allows the reconstruction of an HDR image using a baseline Standard Dynamic Range (SDR) image and a set of prediction coefficients representing a backward shaping function. Conventional devices can simply decode SDR images; however, HDR displays can reconstruct HDR images by applying a backward shaping function to the SDR image. In video coding, this image prediction can be used to improve coding efficiency while maintaining backward compatibility. Such a system can be called “closed-loop” when the encoder includes a decoding path and the prediction coefficients are derived based on the raw and decoded SDR and HDR data, or “open-loop” when there is no such decoding loop and the prediction coefficients are derived only from pairs of raw data. Here, as the inventors understand, there is a need for improved techniques for efficient image prediction in open-loop codecs.

[0010] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise instructed, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise instructed, any issues concerning one or more methods should not be considered to be in any prior art based on this section. Summary of the Invention

[0011] The example embodiments described herein relate to image prediction techniques. In these embodiments, in an apparatus including one or more processors, the processor receives input reference image pairs representing the same scene in both high dynamic range (HDR) and standard dynamic range (SDR) dimensions. Processor:

[0012] Noise data with noise intensity is generated based at least on features of HDR images;

[0013] A noisy input dataset is generated by adding noise data to an SDR image;

[0014] Generate the first enhanced input dataset based on HDR images;

[0015] Combine SDR images and noisy input datasets to generate a second enhanced input dataset;

[0016] Generate a predictive model to predict the first augmented input dataset based on the second augmented input dataset;

[0017] The prediction model is solved according to the error minimization criterion to generate a set of prediction model parameters;

[0018] Compress the second input image to generate a compressed bitstream; and

[0019] Generate an output bitstream that includes compressed bitstream and prediction model parameters. Attached Figure Description

[0020] Embodiments of the invention are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings:

[0021] Figure 1A An example single-layer decoder for HDR data using image prediction based on existing technology is depicted;

[0022] Figure 1B An example HDR open-loop encoder using image prediction based on existing technology is depicted;

[0023] Figure 1C An example HDR closed-loop encoder using image prediction based on existing technology is depicted;

[0024] Figure 1D An example HDR open-loop encoder using image prediction is depicted according to an embodiment of the present invention;

[0025] Figure 2 An example process for designing an enhanced data predictor according to embodiments of the present invention is described; and

[0026] Figure 3 An example process for designing an enhanced data predictor using 3DMT data representation according to an embodiment of the present invention is described. Detailed Implementation

[0027] This document describes an image prediction technique for efficiently encoding images in an open-loop codec. In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent that the invention can be practiced without these specific details. In other instances, well-known structures and devices are not described in detail to avoid unnecessarily obscuring, obscuring, or confusing the invention.

[0028] Example HDR encoding system

[0029] Figure 1A The illustration shows an example single-layer decoder architecture using image prediction, which can be implemented using one or more computational processors in a downstream video decoder. Figure 1B The illustration shows an example "open-loop" encoder architecture, which can also be implemented using one or more computing processors in one or more upstream video encoders. Figure 1C The diagram illustrates an example "closed-loop" encoder architecture.

[0030] Within this framework, given reference HDR content (120), the corresponding SDR content (125) (i.e., representing the same image as the HDR content, but color-graded and represented within the standard dynamic range) is encoded and transmitted in a single layer of the encoded video signal (144) by an upstream encoding device implementing the encoder-side codec architecture. The SDR content (144) is received and decoded in a single layer of the video signal by a downstream decoding device. Predictive metadata (e.g., backward shaping parameters) (152) is also encoded and transmitted in the video signal along with the SDR content, enabling the HDR display device to reconstruct the HDR content based on the SDR content (144) and the received metadata (152).

[0031] exist Figure 1B and Figure 1CIn one embodiment, given input HDR data (120), SDR data (125) can be generated from the HDR data by tone mapping, forward shaping, manual (during color grading), or a combination of techniques known in the art. In another embodiment, given reference SDR data (125), HDR data (120) can be generated from the SDR data by inverse tone mapping, backward shaping, manual (during color grading), or a combination of techniques known in the art. Compression block 140 (e.g., an encoder implemented according to any known video coding algorithm such as AVC, HEVC, AV1, etc.) compresses / encodes the SDR image (125) into a single layer 144 of the encoded bitstream.

[0032] The metadata (152) generated by unit 150 can be multiplexed as part of the video signal 144, for example, as a supplementary enhancement information (SEI) message. Therefore, metadata (152) can be generated or pre-generated on the encoder side to take advantage of the powerful computing resources and offline encoding processes available on the encoder side (including but not limited to content-adaptive multi-rounds, advance operation, inverse luma mapping, inverse chroma mapping, CDF-based histogram approximation and / or passing, etc.).

[0033] Figure 1B and Figure 1C The encoder architecture can be used to avoid directly encoding the input HDR image (120) into an encoded / compressed HDR image in the video signal; instead, the metadata (152) in the video signal can be used to enable downstream decoding devices to reconstruct the (encoded in the video signal) SDR image (125) into a reconstructed HDR image (167) that is the same as or close to / best approximated to the reference HDR image (120).

[0034] In some embodiments, such as Figure 1A As illustrated, the decoder side of the codec framework receives a video bitstream (144) with a compressed SDR image and metadata (152) with prediction parameters generated by the encoder as input. Decompression block 160 decompresses / decodes the compressed video data in a single layer (144) of the video signal into a decoded SDR image (162). Decompression 160 typically corresponds to the reverse process of compression 140. The decoded SDR image (162) may be identical to the SDR image (125), depending on the quantization errors in the compression block (140) and decompression block (160), which may have been optimized for SDR display devices. The decoded SDR image (162) may be output as part of the output SDR video signal (e.g., via an HDMI interface, via a video link, etc.) for rendering on an SDR display device.

[0035] Furthermore, prediction block 165 (also referred to as a “synthesizer”) applies metadata (152) from the input bitstream to the decompressed data (162) to generate a reconstructed HDR image (167). In some embodiments, the reconstructed image represents a production-quality or near-production-quality HDR image that is the same as or close to / best approximated by a reference HDR image (120). The reconstructed image (167) can be output in the output HDR video signal (e.g., via an HDMI interface, via a video link, etc.) for rendering on an HDR display device.

[0036] In some embodiments, as part of an HDR image rendering operation that renders a backward-shaped image (167) on an HDR display device, display management operations specific to the HDR display device may be performed on the reconstructed image (167).

[0037] Figure 1B An "open-loop" encoding architecture is described, in which metadata 152 is generated by unit 150 using only the input HDR and SDR images. Figure 1C A “closed-loop” coding architecture including an additional decompression block (160) is described. The closed-loop design uses an additional video decompression step 160, mimicking the operation of a decoder. This provides a more accurate data description to generate (e.g., in block 150) prediction parameters; however, it requires additional decoding steps. This works well when generating bitstreams at a single bitrate or profile, but becomes computationally more expensive when the server needs to generate streams at multiple bitrates (often referred to as a “bitrate ladder”). Therefore, as the inventors understand, it is beneficial to improve the open-loop architecture to provide performance as good or better than the closed-loop system while reducing computational complexity.

[0038] Example systems for improving predictions in open-loop systems

[0039] Single-channel predictor

[0040] Consider the input data { } and the observed output data { } of the correct, among which, i = 0, 1,…, P -1, where the output is generated as

[0041] (1)

[0042] in, Indicates having parameters of A polynomial model of the order "ground truth" is used, and without loss of generality, This indicates that it has zero mean and zero variance. Additive white Gaussian noise, denoted as .make

[0043] (2)

[0044] Let represent the vector of coefficients in the model, and let

[0045] (3)

[0046] A vector representing the output data of the observation.

[0047] In traditional predictive modeling, given P A set of baseline true data

[0048] {( , )}, hoping to use a new polynomial model The predictive model is constructed using the following formula.

[0049] (4)

[0050] Among them, polynomial coefficients The vector representation is

[0051] (5)

[0052] Given representation

[0053] and (6)

[0054] Equation (4) can be expressed as:

[0055] (7)

[0056] Given equation (7), an optimal set of polynomial coefficients can be defined to minimize the error between the observed and predicted data.

[0057] ,

[0058] Under minimum mean square error (MSE) optimization, the optimal solution is given by the following equation.

[0059] (8)

[0060] As long as the predictor has access to the original The model of Equation (7) works well for the data. Consider this scenario as an approximation of a closed-loop architecture, where the decompressor 160 provides a very accurate copy of the SDR data that the decoder will see. But what if such data is unavailable? In this embodiment, to better account for availability... To address uncertainties in the data (e.g., in open-loop architectures) and build more robust predictors, it is suggested to generate and use Gaussian white noise (e.g., Add to original input The generated set of duplicate input data .therefore,

[0061] , i = 0, 1,…, P -1. (9)

[0062] Figure 1D An example of an open-loop architecture supporting the proposed enhanced data prediction model, according to an embodiment, is depicted. Figure 1B compared to, Figure 1D The architecture also includes a noise insertion module that generates noisy SDR and / or HDR data. The original and noisy SDR and HDR data are then combined to form enhanced SDR and HDR data, which are fed into unit 170 to solve for the prediction parameters of the enhanced data prediction model. The enhanced input dataset can represent a combination of the input image and the noisy input dataset.

[0063] In this embodiment, the observation data of this new enhanced data prediction model Considered to be with same,

[0064] , i = 0, 1,…, P -1. (10)

[0065] In another embodiment, noise may be incorporated when modeling the observation data; however, experimental results show that modeling noise in the observation data does not provide a significant improvement. Therefore, without loss of generality, such noise will not be considered in the following discussion to simplify prediction modeling.

[0066] Given two pairs of training data, for example {( , )}and{( , A new order (e.g., = Polynomial model It can be represented as:

[0067] (11)

[0068] Similarly, the matrix / vector representation of the input and output data is given by the following equation:

[0069] (12)

[0070] ,and (13)

[0071] By combining new and old datasets, combined (or enhanced) datasets can be created.

[0072] ,and (14)

[0073] Furthermore, the enhanced data prediction model can be represented as

[0074] (15)

[0075] Solve This can be described as an optimization problem:

[0076] ,

[0077] The optimal solution (under MSE) is given by the following equation.

[0078] (16)

[0079] Figure 2 An example process for constructing an enhanced data predictor according to an embodiment is described. Figure 2 As described, the input to this process is a pair of input data and observable data, for example, ( , Yes. In step 205, noisy (or perturbative) input data is generated by adding noise to the original input data. (For example, see equation (9)). From the predictor's perspective, at the output of step 215, there is now a set of enhanced input data (e.g., This includes both raw input data and noisy input data. In step 210, noisy (or perturbed) observable data may optionally be generated based on the input observable data (e.g., + ), or for From the predictor's perspective, after step 220, there is now a set of enhanced observables and noisy observables (e.g., or Finally, in step 225, the coefficients of the enhanced data prediction model are solved (see, for example, equation (16)).

[0080] Augmented data prediction using multi-channel models

[0081] The preceding discussion used a relatively simple single-channel prediction model. In this section, the method is extended to multi-channel regression models, such as, but not limited to, those described in references [1] and [2]. As an example, without loss of generality, a detailed method will be described for an embodiment using a multi-channel multiple regression (MMR) predictor (reference [1]); however, those skilled in the art should be able to extend the method to other models, such as tensor product B-spline (TPB) models (reference [2]).

[0082] Consider a video sequence, where the first... t Samples in a frame (e.g., SDR images) are represented as ( , , ), i = 0,1,…, P -1, where each pixel has three color components. y , c 1 and c 2. For example (YCbCr, RGB, ICTCb, etc.). For example, an SDR image (125) can represent image data with 100 nits and an R709 color gamut, while the corresponding HDR image (120) can represent image data with 4,000 nits and a P3 color gamut. Using the MMR model, let the output... Represented as the following combination (where, ch (Indicates y, c0, or c1)

[0083] (17a)

[0084] For example, in an embodiment, a second-order vector with cross product MMR representation is used. It can be represented by 15 values.

[0085] (17b)

[0086] .

[0087] In equations (17a-17b), in some embodiments, certain terms can be removed to reduce computational load. For example, only one of the chromaticity components can be used in the model, or some higher-order cross components can be eliminated entirely. Not limiting, alternative linear or nonlinear predictors can also be employed.

[0088] make

[0089] (18)

[0090] Therefore, observable data (e.g., HDR images) can be represented as

[0091] (19)

[0092] Furthermore, the entire baseline truth model can be expressed as

[0093] (20)

[0094] in,

[0095] (twenty one)

[0096] Indicates additive noise, for example, .

[0097] Note: Using Gaussian white noise can be viewed as modeling quantization noise in an open-loop problem using worst-case noise. Those skilled in the art will understand that alternative models known in the art (e.g., Laplace, Cauchy models, etc.) can be used to model such noise.

[0098] Given a matrix form MMR model

[0099] ,(twenty two)

[0100] The parameters of a traditional predictor can be computed again using a minimization problem, for example...

[0101] ,

[0102] The optimal solution (under MSE) is given by the following equation.

[0103] ,(twenty three)

[0104] in,

[0105] and .(twenty four)

[0106] According to Figure 2 The method described in [the document] designs an enhanced data predictor. Similar to the single-channel case (see step 205), given input {( , , )|i = 0, 1,…, P-1}, by adding noise (e.g., a distribution of Gaussian noise) generates new noise or perturbation sets {( , , )|i = 0, 1,…, P-1}, for example,

[0107] , i = 0, 1,…, P-1(25)

[0108] make

[0109] ,

[0110] as well as

[0111] (26)

[0112] If the observed data Maintain and If the same (e.g., skip step 210), then

[0113] , i = 0, 1,…, P-1

[0114] as well as

[0115] (27)

[0116] In steps 215 and 220, the old and new datasets are combined to obtain...

[0117] as well as 。 (28)

[0118] Finally, in step 225, the optimization problem is...

[0119]

[0120] The solution can be obtained using the least squares minimization method.

[0121] 。 (29)

[0122] In another embodiment, the data can be augmented using additional perturbation input and / or output datasets (e.g., by using different noise variances for each perturbation set). For example, several sets can be created. and ( k = 0, 1,…, K -1), generating (e.g., in steps 215 and 220) a combined dataset:

[0123] as well as (30)

[0124] The solution to the prediction model is still given by equation (29).

[0125] Considerations for Noise Intensity Selection

[0126] A key part of augmented data prediction models is generating perturbed (or noisy) data by adding noise to the original input data. The question then arises: how much noise should be added? Intuitively, in video coding, a higher bitrate results in lower quantization noise; therefore, at least one parameter influencing the amount of noise added could be the target bitrate of the compressed bitstream.

[0127] As used herein, the term "within range" refers to the pixel range of the original test or training data to be used in the prediction model (e.g., [a, b]). As used herein, the term "beyond the lower bound of the range" refers to values ​​below the minimum range used in the prediction model (e.g., ...). a The pixel values ​​of the image. For example, these might be images with very low black values. As used in this paper, the term "out of range" means values ​​above the maximum range used in the prediction model (e.g., ...). b (pixel values). For example, these might be images with very high highlight values.

[0128] Experimental results show that for any out-of-range data, the augmented data predictor always performs better as the noise variance increases; however, for data within the range, the augmentation only improves when the standard deviation of the added noise is below a certain "optimal" value (denoted as...). When the noise variance is such that the augmented data predictor performs better, the data predictor will improve. Therefore, this optimal noise variance can be expressed as:

[0129] (31)

[0130] in Indicates the use of standard deviation The average distortion measure of Gaussian white noise-enhanced input data, and This represents the average distortion predicted using traditional prediction models, for example,

[0131] .

[0132] These observations indicate that another parameter affecting noise intensity is the dynamic range of the output (e.g., HDR) data, particularly the dynamic range of the chroma color components in the HDR input. Experimental data also show that... PThe larger the value, the more robust the augmented data model; however, in practice, due to the high computational cost, it's almost impossible to operate directly on all pixel values. Instead, subsampled images or "averaged" pixel values ​​can be used. For example, the input signal codeword can be divided into segments with equal intervals. w b (For example, for 16-bit input data,) w b = 65,536 / M )of M Non-overlapping bins (e.g., M = 16, 32, or 64), to cover the entire normalized dynamic range (e.g., (0, 1]). Then, instead of operating on pixel values, it can be operated on the average pixel values ​​within each such bin. The number of HDR bins (also known as a 3D mapping table (3DMT)) is represented as... P t In the embodiments, the noise intensity can be derived based on the following heuristic.

[0133] (32)

[0134] Wherein, given

[0135] ,

[0136] ,

[0137] but

[0138] (33)

[0139] Indicates the effective dynamic range of the observed data. Indicates the maximum noise intensity (e.g., = 0.08), It is a parameter that controls the expansion based on the input data count (e.g., = 3,000), and It is a parameter that controls the expansion based on the range of observed data (e.g., when the bit depth = 16 bits). =7,000). The model provides a slower decay as the input increases.

[0140] In another embodiment, an alternative approach is to provide faster decay with higher-order terms within the exponential function:

[0141] (34)

[0142] in, >1.

[0143] In the embodiments, a bit rate-related multiplier factor can be added to equations (32) and (34), for example:

[0144] (35)

[0145] in, The parameters controlling the spread are based on the average bit rate used to generate noise (e.g., = 2 Mb / s). For example, at high bit rates (e.g., 5.2 Mb / s or higher), the noise intensity may be almost zero. In the embodiment, in equation (35), the value of α in each exponential factor can have different values ​​(e.g., each α can be composed of different values ​​(e.g., ... , and )replace).

[0146] In one embodiment, an optimized noise intensity can be generated for each target bit rate, thereby generating a dedicated set of prediction parameters for each bit rate. In another embodiment, the service provider may wish to use one set (or just a few sets). For example, for a set of optimized MMR parameters, the worst-case scenario (e.g., minimum resolution at the lowest bit rate) can be used to add noise. In such a scenario, the bit rate-related exponent term in equation (35) can be considered as being absorbable. The fixed value (for example, see equation (34)).

[0147] Given a heuristic noise model (see Equation (35)), Figure 3 Depicting open-loop 3DMT architectures (e.g., such as...) Figure 1D An example process for enhancing data prediction (described). Referring to the HDR input... t The first frame i The color component values ​​of a pixel are represented as The corresponding SDR pixel value is represented as The minimum and maximum values ​​in each color channel are represented in the SDR image as ( , ), and is represented as ( ) in HDR images. , ).

[0148] like Figure 3 The process described, which begins in step 305 with the construction of a 3DMT representation (see also references [3-4]), can be summarized as follows:

[0149] a) In each channel, use a fixed number of bins for each component. To quantize the dynamic range of an SDR image. This partitioning can be used to cover the minimum / maximum (max) in each dimension. , ) range of unified partition boundaries to calculate ( 3D histogram. The quantization interval in each channel is given by the following formula:

[0150] .

[0151] Represent the 3D histogram box as ,in, .therefore, Total includes ( ) bins, such that each 3D bin is indexed by bin index q = ( The bin index is specified to represent the number of pixels with these 3-channel quantization values. To simplify the notation, the 3D bin index {q} can be vectorized to the 1-D index {q}.

[0152] .

[0153] b) Calculate the sum of each color component in the HDR for each 3D box. Let , and These are the luminance and chrominance values ​​mapped in the HDR image domain, such that each of these bins contains all HDR luminance and two chrominance values ​​(respectively, respectively). C 0 and C 1) Sum of pixel values, where the corresponding pixel value is located in the bin. Assuming there are P pixels, the operation can be summarized in pseudocode as follows:

[0154]

[0155] c) Locate the 3D histogram bin with a non-zero pixel count. In other words, collect all non-zero entries to set... Calculate HDR ( , , ) and SDR ( , , The average value of ).

[0156] Will The number of elements in is expressed as P t .make,

[0157] (36a)

[0158] as well as

[0159] (36b)

[0160] Then, for The elements in the set have mapping pairs { }and{ }

[0161] In step 310, the noise intensity can be calculated as follows: given P t The number of 3DMT boxes, in the embodiment, can be used to calculate the chromaticity range. R t As the average of the dynamic ranges in the two color channels:

[0162] (37)

[0163] Then, the noise intensity can be calculated according to equation (34) or (35).

[0164] In another embodiment, the noise standard deviation can be calculated separately for luminance and each color component, at the cost of increased complexity. Alternatively, it can be calculated using the maximum or minimum value of two chromaticity ranges instead of the average chromaticity range. However, overall, experimental results aimed at improving colorimetric quality indicate that, as described in the calculations... To produce satisfactory results with reasonable complexity and cost.

[0165] In step 315, without loss of generality, it is assumed that the MMR prediction model, given equation (36b), and the SDR input dataset can be expressed as follows:

[0166] .

[0167] Collect all P t 1 entry, obtained

[0168] .

[0169] Similarly, the vector form of 3DMT HDR chroma values ​​can be represented as:

[0170] and

[0171] In step 320, noise is added to each 3DMT entry.

[0172] ,

[0173] The obtained input 3DMT data is noisy, given by the following equation.

[0174] ,

[0175] In this case, the noise in each channel has the same distribution, for example,

[0176] .

[0177] In step 325, the augmented input 3DMT dataset is generated as follows: Noisy The extended form of the input MMR is represented as

[0178]

[0179] Then, for

[0180] ,

[0181] ,and ,

[0182] The augmented dataset is given by the following formula.

[0183] as well as .

[0184] In this embodiment, in step 330, the new prediction model can be described as follows:

[0185] (38)

[0186] The optimal solution (under the MSE criterion) is given by (references [3-4]).

[0187] (39)

[0188] References

[0189] Each of these references is incorporated into this paper in its entirety by way of citation.

[0190] 1. GM. Su et al., Multiple color channel multiple regression predictor [Multi-color channel multiple regression predictor] “, U.S. Patent 8,811,490.

[0191] 2. GM Su and others, Tensor-product B-spline predictor ", U.S. Provisional Patent Application No. 62 / 908,770, filed on October 1, 2019.

[0192] 3. N. J. Gadgil and GM. Su, “ Linear encoder for image / video processing Linear encoders for image / video processingThe PCT application filed on February 28, 2019, with the number PCT / US 2019 / 020115, was published as WO 2019 / 169174.

[0193] 4.Q. Song et al., " High-fidelity full reference and high-efficiency reduced reference encoding in end-to-end single-layer backward compatible Encoding pipeline [High-fidelity full-reference and efficient reduced-reference encoding in an end-to-end single-layer backward-compatible encoding pipeline] code] The WIPO PCT publication is WO2019 / 217751, dated November 14, 2019.

[0194] Example computer system implementation

[0195] Embodiments of the present invention may be implemented using computer systems, systems configured with electronic circuits and components, integrated circuit (IC) devices (such as microcontrollers, field-programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs)), and / or means including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or implement instructions related to image prediction techniques, as described herein. The computer and / or IC may calculate any of the various parameters or values ​​associated with the generation of the image prediction techniques described herein. Image and video dynamic range extension embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0196] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the method of the present invention. For example, one or more processors, such as those in a display, encoder, set-top box, transcoder, etc., can implement the method for the image prediction technique described above by executing software instructions in a processor-accessible program memory. The present invention can also be provided in the form of a program product. The program product may include any non-transitory and tangible medium carrying a set of computer-readable signals, including instructions that, when executed by a data processor, cause the data processor to perform the method of the present invention. The program product according to the present invention can take any of a variety of non-transitory and tangible forms. The program product may include, for example, physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0197] In the case of the components mentioned above (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise stated, references to such components (including references to “devices”) should be interpreted as including equivalents (e.g., functionally equivalents) of any component that performs the function of the described component, including components that are not structurally equivalent to components that perform the functions in the illustrated exemplary embodiments of the invention.

[0198] Equivalents, extensions, alternatives and miscellaneous

[0199] Therefore, embodiments relating to image prediction technology have been described. In the foregoing description, embodiments of the invention have been described with reference to numerous specific details that may vary depending on the implementation. Therefore, the sole and exclusive indication of the invention and the applicant's inventive intent is the set of claims issued in specific form according to this application, wherein such claims include any subsequent amendments. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, characteristics, advantages, or attributes not expressly referenced in the claims should not in any way limit the scope of such claims. Therefore, this specification and drawings should be viewed in an illustrative rather than restrictive sense.

[0200] Various aspects of the invention can be understood from the following enumerated example embodiments (EEE):

[0201] 1. A method for generating predictive coefficients using a processor, the method comprising:

[0202] Access a first input image (120) in a first dynamic range and a second input image (125) in a second dynamic range, wherein the first input image and the second input image represent the same scene;

[0203] Noise data with noise intensity is generated based at least on the features of the first input image;

[0204] A noise input dataset is generated by adding the noise data to the second input image;

[0205] Generate a first enhanced input dataset based on the first input image;

[0206] The second input image and the noisy input dataset are combined to generate a second enhanced input dataset;

[0207] Generate a prediction model to predict the first augmented input dataset based on the second augmented input dataset;

[0208] The prediction model is solved according to the error minimization criterion to generate a set of prediction model parameters;

[0209] Compress the second input image to generate a compressed bitstream; and

[0210] Generate an output bitstream that includes the compressed bitstream and the prediction model parameters.

[0211] 2. The method as described in EEE 1, further comprising, in the decoder:

[0212] Receive the output bitstream including the compressed bitstream and the prediction model parameters;

[0213] Decode the output bitstream to generate a first output image in the second dynamic range; and

[0214] The prediction model parameters are applied to the first output image to generate a second output image within the first dynamic range.

[0215] 3. The method as described in EEE 1 or EEE 2, wherein the first dynamic range includes a high dynamic range and the second dynamic range includes a standard dynamic range.

[0216] 4. The method of any one of EEE 1 to 3, wherein generating the noise data comprises:

[0217] Calculate statistical data based on the pixel values ​​of the first input image;

[0218] Calculate the noise standard deviation based on the statistical data; and

[0219] Noise samples of the noise data are generated using a Gaussian distribution with zero mean and the noise standard deviation.

[0220] 5. The method as described in EEE 4, wherein the calculation of the noise standard deviation is further based on the target bit rate for generating the compressed bitstream and / or features of the second input image.

[0221] 6. The method as described in EEE 4 or EEE 5, wherein calculating the statistical data includes calculating one or more of the following: the total number of pixel values ​​in the first input image, the range of pixel values ​​in the luminance component of the first input image, the range of pixel values ​​in the chrominance component of the first input image, or the number of bins representing groups representing the average pixel values ​​of the first input image.

[0222] 7. The method of any one of EEE 1 to 6, wherein the prediction model comprises a single-channel predictor and a multi-channel multiple regression (MMR) predictor.

[0223] 8. The method of any one of EEE 1 to 7, wherein solving the prediction model includes minimizing an error metric between the output of the prediction model and the first input image.

[0224] 9. The method as described in EE 8, wherein generating the set of predictive model parameters includes calculating

[0225] ,

[0226] in, The vector representation of the parameters of the prediction model. This represents the first augmented input dataset, and This represents a matrix based on the second enhanced input dataset.

[0227] 10. The method as described in EE 9, wherein, for color components ch ,

[0228] and ,

[0229] in, This represents the pixel values ​​of the first augmented input dataset. Including the pixel values ​​of the first input image, and Including the pixel values ​​of the first input image, wherein Alternatively, it may include pixel values ​​of the first input image with added noise.

[0230] 11. The method of any one of EEE 1 to 10, further comprising:

[0231] The first modified dataset is generated based on the modified representation of the first input image;

[0232] The second modified dataset is generated based on the modified representation of the second input image;

[0233] The noise input dataset is generated by adding the noise data to the second modified dataset;

[0234] The first enhanced input dataset is generated based on the first modified dataset; and

[0235] The second modified dataset and the noisy input dataset are combined to generate the second enhanced input dataset.

[0236] 12. The method as described in EEE 11, wherein the first modified dataset includes a subsampled version of the first input image or a three-dimensional table mapping (3DMT) representation of the first input image.

[0237] 13. The method as described in EEE 11 or EEE 12, wherein the second modified dataset includes a subsampled version of the second input image or a three-dimensional table mapping (3DMT) representation of the second input image.

[0238] 14. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for executing, with one or more processors, the method according to any one of EEE 1 to 13.

[0239] 15. An apparatus comprising a processor and configured to perform any of the methods described in EEE 1 to 13.

Claims

1. A method for generating predictive coefficients using a processor, the method comprising: Access a first input image (120) in a first dynamic range and a second input image (125) in a second dynamic range, wherein the first input image and the second input image represent the same scene; Calculate the dynamic range of one or more chromaticity color components of the first input image; Based on the calculated dynamic range of the one or more chromaticity color components of the first input image, noise data with noise intensity is generated. A noise input dataset is generated by adding the noise data to the second input image; Generate a first enhanced input dataset based on the first input image; The second input image and the noisy input dataset are combined to generate a second augmented input dataset as training data; Generate a prediction model to predict the first augmented input dataset based on the second augmented input dataset; The prediction model is solved according to the error minimization criterion to generate a set of prediction model parameters; Compress the second input image to generate a compressed bitstream; and Generate an output bitstream that includes the compressed bitstream and the prediction model parameters.

2. The method of claim 1, further comprising, in the decoder: Receive the output bitstream including the compressed bitstream and the prediction model parameters; Decode the output bitstream to generate a first output image in the second dynamic range; as well as The prediction model parameters are applied to the first output image to generate a second output image within the first dynamic range.

3. The method as described in claim 1 or 2, wherein, The first dynamic range includes a high dynamic range, and the second dynamic range includes a standard dynamic range.

4. The method as described in claim 1 or 2, wherein, The generation of the noise data includes: Calculate statistical data based on the pixel values ​​of the first input image; Calculate the noise standard deviation based on the statistical data; and Noise samples of the noise data are generated using a Gaussian distribution with zero mean and the noise standard deviation.

5. The method of claim 4, wherein, The noise standard deviation is calculated further based on the target bit rate used to generate the compressed bit stream and / or the characteristics of the second input image.

6. The method of claim 4, wherein, Calculating the statistical data includes calculating one or more of the following: the total number of pixel values ​​in the first input image, the range of pixel values ​​in the luminance component of the first input image, the range of pixel values ​​in the chrominance component of the first input image, or the number of bins representing the average pixel values ​​of the first input image.

7. The method as described in claim 1 or 2, wherein, The prediction model includes a single-channel predictor and a multi-channel multiple regression (MMR) predictor.

8. The method as claimed in claim 1 or 2, wherein, Solving the prediction model involves minimizing the error metric between the output of the prediction model and the first input image.

9. The method of claim 8, wherein, Generating this set of predictive model parameters includes calculating , in, The vector representation of the parameters of the prediction model. This represents the first augmented input dataset, and This represents a matrix based on the second enhanced input dataset.

10. The method of claim 9, wherein, For color components ch , and , in, This represents the pixel values ​​of the first augmented input dataset. Including the pixel values ​​of the first input image, and Including the pixel values ​​of the first input image, wherein Alternatively, it may include pixel values ​​of the first input image with added noise.

11. The method of claim 1 or 2, further comprising: The first modified dataset is generated based on the modified representation of the first input image; The second modified dataset is generated based on the modified representation of the second input image; The noise input dataset is generated by adding the noise data to the second modified dataset; The first enhanced input dataset is generated based on the first modified dataset; as well as The second modified dataset and the noisy input dataset are combined to generate the second enhanced input dataset.

12. The method of claim 11, wherein, The first modified dataset includes a subsampled version of the first input image or a three-dimensional table mapping (3DMT) representation of the first input image.

13. The method of claim 11, wherein, The second modified dataset includes a subsampled version of the second input image or a three-dimensional table mapping (3DMT) representation of the second input image.

14. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for performing the method according to any one of claims 1 to 13 using one or more processors.

15. An apparatus for generating prediction coefficients, comprising a processor and configured to perform any one of the methods described in claims 1 to 13.