Neural network for dynamic range conversion and display management of images

CN117716385BActive Publication Date: 2026-09-15DOLBY LABORATORIES LICENSING CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202280052320.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-29
Filing Date
2022-07-22
Publication Date
2026-09-15
Estimated Expiration
2042-07-22

Smart Images

  • Figure CN117716385B_ABST
    Figure CN117716385B_ABST
Patent Text Reader

Abstract

Methods and systems for dynamic range conversion and display mapping of standard dynamic range (SDR) images onto high dynamic range (HDR) displays are described. Given an SDR input image, a processor generates an intensity (luminance) image and optionally a base layer image and a detail layer image. A first neural network uses the intensity image to predict statistics of the SDR image in a higher dynamic range. These predicted statistics are used with original image statistics of the input image to derive an optimal tone mapping curve to map the input SDR image onto an HDR display. Optionally, using the intensity image and the detail layer image, a second neural network can generate a residual detail layer image in the higher dynamic range to augment the tone mapping of the base layer image into the higher dynamic range.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 226,847, filed July 29, 2021, and European Patent Application Serial No. 21188516.5, filed July 29, 2021, each of which is incorporated herein by reference in its entirety. Technical Field

[0003] This invention generally relates to images. More specifically, embodiments of the invention relate to dynamic range conversion and display management from standard dynamic range (SDR) images to high dynamic range (HDR) displays. Background Technology

[0004] As used herein, the term "dynamic range (DR)" can refer to the ability of the human visual system (HVS) to perceive a range of intensity (e.g., luminance, brightness) in an image, such as from the darkest gray (black) to the brightest white (highlight). In this sense, DR relates to the intensity "scene-referred". DR can also refer to the ability of a display device to fully or approximately render a specific breadth of intensity range. In this sense, DR relates to the intensity "display-referred". Unless a particular meaning is explicitly specified to have a specific connotation at any point in the description herein, it should be inferred that the term can be used interchangeably in either sense, for example.

[0005] As used herein, the term “high dynamic range (HDR)” refers to a DR width of approximately 14 to 15 or more orders of magnitude across the human visual system (HVS). In practice, the DR, which refers to the broad width of the intensity range that humans can simultaneously perceive relative to HDR, may be slightly truncated. As used herein, the terms “enhanced dynamic range (EDR)” or “visual dynamic range (VDR)” can be associated, individually or interchangeably, with this DR: a DR that can be perceived within a scene or image by the human visual system (HVS) including eye movements, thus allowing for some light-adaptive variations in the scene or image.

[0006] In practice, an image comprises one or more color components (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented by n bits per pixel with precision (e.g., n = 8). For example, using gamma luminance encoding / decoding, images where n ≤ 8 (e.g., color 24-bit JPEG images) are considered standard dynamic range images, while images where n ≥ 10 can be considered enhanced dynamic range images. EDR and HDR images can also be stored and distributed using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR file format developed by Industrial Light and Magic.

[0007] As used herein, the term "metadata" refers to any auxiliary information transmitted as part of the encoded bitstream and used by the decoder to render the decoded image. Such metadata may include, but is not limited to, minimum, average, and maximum luminance values, color space or gamut information, reference display parameters, and auxiliary signal parameters in an image as described herein.

[0008] Most consumer desktop monitors currently support 200 to 300 cd / m³. 2 Or nits of brightness. Most consumer HDTVs range from 300 to 500 nits, with newer models reaching 1000 nits (cd / m²). 2 Therefore, such conventional displays represent a lower dynamic range (LDR) relative to HDR or EDR, also known as standard dynamic range (SDR). As the availability of HDR content has increased due to advancements in both capture devices (e.g., cameras) and HDR displays (e.g., the Dolby Laboratories PRM-4200 professional reference monitor), HDR content can be color-graded and displayed on HDR displays that support higher dynamic ranges (e.g., from 1,000 nits to 5,000 nits or higher). Generally, and without limitation, the methods disclosed herein relate to any dynamic range above SDR.

[0009] As used herein, the term "display management" refers to the process performed on a receiver to render an image for a target display. For example, and without limitation, such a process may include tone mapping, gamut mapping, color management, frame rate conversion, etc.

[0010] The creation and playback of High Dynamic Range (HDR) content is becoming increasingly common because HDR technology offers more realistic and lifelike images than earlier formats. However, traditional content may only be available in Standard Dynamic Range (SDR), and broadcast infrastructure may not allow the transmission of metadata to convert such content into a format suitable for fully utilizing the capabilities of HDR displays. To improve existing display solutions, as understood herein by the inventors, improved techniques have been developed for the upconversion and display management of SDR images to HDR displays.

[0011] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise instructed, no method described in this section should be assumed to be prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise instructed, any issues concerning one or more methods should not be assumed to be found in any prior art based on this section. Attached Figure Description

[0012] Embodiments of the invention are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements in the drawings, and in the drawings:

[0013] Figure 1 An example process for a video transmission pipeline is described;

[0014] Figure 2A A dynamic range upconversion and display management pipeline with a single neural network processing unit is described according to a first exemplary embodiment of the present invention;

[0015] Figure 2B A dynamic range upconversion and display management pipeline with two neural network processing units is described according to a second exemplary embodiment of the present invention;

[0016] Figure 3 An example neural network architecture for predicting luminance metadata is described according to an exemplary embodiment of the present invention;

[0017] Figure 4A A processing pipeline is depicted in a Residual Network (ResNet) block used in a neural network to predict residual images of detail layers, according to an exemplary embodiment of the present invention; and

[0018] Figure 4B A processing pipeline for predicting detail layer residual images in a neural network according to an exemplary embodiment of the present invention is described. Detailed Implementation

[0019] This document describes a method for dynamic range conversion and display management of SDR images to HDR displays. In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent that the invention can be practiced without these specific details. In other instances, well-known structures and devices have not been described in detail to avoid unnecessarily obscuring, obscuring, or confusing the invention. Summary of the Invention

[0020] The example embodiments described herein relate to a method for dynamic range conversion and display management of an SDR image on an HDR display. In an embodiment, a processor receives an input image (202) with a first dynamic range and a first spatial resolution. The processor generates an intensity image (207) based on the input image, applies the intensity image to a first neural network (210) to generate predictive statistics when the intensity image is mapped in a second dynamic range higher than the first dynamic range, generates (215) a tone mapping curve based on the intensity image statistics and the predictive statistics, and applies the tone mapping curve to display the input image on a display having a target dynamic range different from the second dynamic range. In an embodiment, a method for dynamic range conversion and display mapping includes: accessing an input image with a first dynamic range and a first spatial resolution; generating an intensity image based on the input image; applying the intensity image to a first neural network to generate predictive statistics when the intensity image is mapped in a second dynamic range higher than the first dynamic range; and generating a tone mapping curve based on the intensity image statistics and the predictive statistics. In an embodiment, the method may include generating an output image based on the input image and the tone mapping curve for display on a display having a target dynamic range. The target dynamic range may be different from the second dynamic range. In an embodiment, the statistical data of the intensity image includes the intensity value of the intensity image in a first dynamic range, and the prediction statistical data includes the predicted intensity value in a second dynamic range (e.g., the corresponding predicted intensity value in the second dynamic range).

[0021] In one embodiment, a first neural network can be trained on a pair of reference images in a second dynamic range (e.g., HDR) and a pair of reference images in a first dynamic range (e.g., SDR). Reference images in the first dynamic range can be generated by mapping each reference image in the second dynamic range to the first dynamic range using a tone mapping operation. The first neural network can be trained on the pair of reference images in the second and first dynamic ranges to learn a relationship (e.g., minimizing error) between predicted statistics of the reference images in the first dynamic range (e.g., predicted statistics of the corresponding intensity image generated based on the reference images in the first dynamic range) and statistics of the reference images in the second dynamic range. In one embodiment, this can be accomplished by iteratively computing predicted statistics of the reference images in the first dynamic range for the first and second dynamic range reference image pairs using the first neural network, and backpropagating the error between the predicted statistics and the statistics of the corresponding reference images in the second dynamic range to the first neural network. Training can terminate when the error between the reference statistics and the predicted statistics is within a small threshold or reaches a non-decreasing steady-state period. Alternatively, in another embodiment, after generating prediction statistics for a reference image in a first dynamic range using a first neural network, the prediction statistics can be applied to upconvert the reference image in the first dynamic range to its corresponding prediction image in a second dynamic range. The corresponding reference image in the second dynamic range can be compared with the prediction image in the second dynamic range, and the error between the prediction image and the reference image in the second dynamic range can be backpropagated to the first neural network.

[0022] In one embodiment, the method may further include: generating a base layer image and a detail layer image based on an intensity image; and applying a tone mapping curve to the base layer image to generate a tone-mapped base layer image in a second dynamic range. In another embodiment, the method may further include: applying the intensity image and the detail layer image to a second neural network to generate a residual layer image in the second dynamic range; adding the residual layer image to the detail layer image to generate a second detail layer image; and adding the second detail layer image to the tone-mapped base layer image to generate an output image in the second dynamic range. In one embodiment, the second neural network can be trained on a pair of reference images in the second dynamic range (e.g., HDR) and a pair of reference images in the first dynamic range (e.g., SDR). A reference image in the first dynamic range can be generated by mapping each reference image in the second dynamic range to the first dynamic range using a tone mapping operation. The second neural network can be trained on the pair of reference images in the second and first dynamic ranges to learn the relationship between a predicted image in the second dynamic range and its corresponding reference image in the second dynamic range (e.g., minimizing error). Each pair of reference images in the second dynamic range and the first dynamic range can be processed by a second neural network, wherein the error between the reference image in the second dynamic range and the corresponding predicted image in the second dynamic range is backpropagated into the second neural network. The predicted image for each pair can be generated by: applying an intensity image generated based on the reference image in the first dynamic range and the corresponding detail layer image to the second neural network to generate a residual layer image in the second dynamic range; adding the residual layer image to the detail layer image to generate a second detail layer image; and adding the second detail layer image to a tone-mapped base layer image generated by applying a tone mapping curve to the intensity image to generate a predicted output image in the second dynamic range.

[0023] SDR to HDR image mapping and display management

[0024] Video encoding / decoding pipeline

[0025] Figure 1 An example process of a conventional video transmission pipeline (100) is depicted, illustrating the various stages from video capture to video content display. An image generation block (105) is used to capture or generate a sequence of video frames (102). The video frames (102) can be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) can be captured on film by a film camera. The film is converted to a digital format to provide video data (107). In the production stage (110), the video data (107) is edited to provide a video production stream (112).

[0026] The video data from the production stream (112) is then provided to the processor at block (115) for post-production editing. Post-production editing at block (115) may include adjusting or modifying the color or brightness of specific areas of the image to enhance image quality or achieve a specific look for the image according to the video creator's creative intent. This is sometimes referred to as "color timing" or "color grading". Other edits (e.g., scene selection and sorting, image cropping, adding computer-generated visual effects, etc.) may be performed at block (115) to produce a final version (117) of the work for distribution. During post-production editing (115), the video image is viewed on a reference monitor (125).

[0027] After post-production (115), the video data of the final work (117) can be transferred to the encoding block (120) for downstream transmission to decoding and playback devices such as televisions, set-top boxes, and cinemas. In some embodiments, the encoding block (120) may include audio encoders and video encoders such as those defined by ATSC, DVB, DVD, Blu-ray, and other transmission formats to generate an encoded bitstream (122). In the receiver, the encoded bitstream (122) is decoded by the decoding unit (130) to generate a decoded signal (132) that is the same as or nearly identical to the signal (117). The receiver may be attached to a target display (140) that may have characteristics completely different from those of the reference display (125). In this case, the display management block (135) may be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display mapping signal (137). Examples of the display management process are described in references [1] and [2] without limitation.

[0028] SDR to HDR Dynamic Range Conversion Pipeline

[0029] In traditional display mapping, mapping algorithms apply sigmoid-like functions (e.g., see references [3] and [4]) to map the input dynamic range to the target display's dynamic range. Such mapping functions can be represented as piecewise linear or nonlinear polynomials, characterized by anchor points, pivots, and other polynomial parameters generated using characteristics of the input source and the target display. For example, in references [3-4], the mapping function uses anchor points based on the luminance characteristics of the input image and the display (e.g., minimum, medium (average), and maximum luminance). However, other mapping functions can use different statistics, such as block-level or whole-image luminance variance or luminance standard deviation. For SDR images, the process can also be aided by additional metadata, which is either transmitted as part of the transmitted video or computed by the decoder or the display. For example, when a content provider has both an SDR and an HDR version of the source content, the source can use both versions to generate metadata (e.g., piecewise linear approximations of forward or backward shaping functions) to assist the decoder in converting the incoming SDR image into an HDR image. However, in many broadcast scenarios, limitations in HDR content availability, transmitters, communication media, and / or receivers may prevent the generation or transmission of such metadata, thus hindering the most efficient use of HDR displays.

[0030] Figure 2A A dynamic range upconversion and display management pipeline (200A) according to an example embodiment is depicted. Figure 2A As depicted herein, the input video (202) may include video received from a video decoder and / or video received from a graphics processing unit (e.g., from a set-top box), and / or other video inputs (e.g., from an HDMI port in a camera, television, or set-top box, a graphics processing unit (GPU), etc.). Without limitation, the input video 202 may be characterized as "SDR" video to be upconverted to "HDR" video for display on an HDR display.

[0031] In one embodiment, process 200A includes a neural network (NN) (210) for generating a set of predicted HDR statistics (or metadata) to facilitate the generation of an optimized SDR-to-HDR mapping. Due to computational limitations, in this embodiment, a preprocessing unit (205) may precede the NN unit (210) to convert the input image 202 into an image suitable in terms of color format and resolution. The mapping unit (215) uses the output of the NN unit 210 to generate an optimized mapping curve, which, together with the original input (202), is fed to the display mapping unit (220) to generate a mapping output 222. Details of each component are described below.

[0032] Neural Network Input Generation

[0033] In block 205, the input image is converted into a format suitable for processing by NN unit 210. In an embodiment, this process includes two steps: a) extracting the intensity or luminance of the input signal, and b) adjusting its resolution. For example, to extract intensity, the input RGB image can be converted to a luminance-chrominance color format, such as YCbCr, ICtCp, etc., using color conversion techniques known in the art (e.g., ITU-R Rec.BT 2100). In an alternative embodiment, intensity can be characterized as the maximum value per pixel of its R, G, and B components. If the source image has already been represented as a single-channel intensity image, the intensity extraction step can be bypassed. In some embodiments, pixel values ​​can also be normalized to [0, 1] according to a predefined standard dynamic range (e.g., between 0.005 and 100 nits) to facilitate the calculation of image statistics.

[0034] Global metadata generation neural networks (210) typically operate on a fixed image size, but the input image size can vary based on the source content (e.g., 480p, 720p, 1080i, etc.). In an embodiment, unit 205 can resample the image size to a size suitable for training and operating the NN metadata generator (e.g., 960×540). For example, a resolution of 960×540 has been found to offer a good trade-off between the complexity and resolution of state-of-the-art neural networks.

[0035] In one embodiment, if the input image is larger than the supported resolution of the neural network (NN), the input image is downsampled twice repeatedly until both the width and height are less than or equal to the desired resolution. As an example, and not a limitation, the downsampling operation can be performed using 4-tap separable horizontal and vertical low-pass filters (e.g., [1 3 3 1] / 8), followed by dropping every other pixel on both the horizontal and vertical dimensions. The width and height are then padded symmetrically along all four sides with padding values ​​to obtain the desired image size (e.g., 960 × 540). In other embodiments, the neural network can be trained for different image sizes, and the resampling step can be adjusted accordingly.

[0036] Neural network used to generate estimated HDR statistics

[0037] The predictive HDR statistics neural network (210) takes a single channel (its luminance) of an SDR image (202) as input and predicts statistics for the corresponding HDR image as needed to generate SDR-to-HDR mapping curves (e.g., minimum, average, and maximum luminance values). In some embodiments, the predicted HDR metadata (212) may be temporally filtered to ensure temporal consistency between images in a video scene. These values ​​may also be adjusted for inconsistent results to ensure they can be used for mapping, for example by clamping the results between 0 and 1, or by ensuring the monotonicity of the resulting image statistics.

[0038] In this embodiment, neural network 210 is defined as a set of 4-dimensional convolutions, each followed by a constant bias value added to all results. In some layers, negative values ​​are clamped to 0 after convolution. Convolutions are defined by their pixel size (M×N), the number of image channels they operate on (C), and the number of such kernels (K) in the filter bank. In this sense, each convolution can be described by a filter bank of size M×N×C×K. As an example, a filter bank of size 3×3×1×2 consists of two convolution kernels, each operating on one channel and having a size of 3 pixels × 3 pixels.

[0039] Some filter banks can also have a stride, meaning some of the convolution results are discarded. A stride of 1 means that each input pixel produces one output pixel. A stride of 2 means that only every other pixel in each dimension produces an output, and so on. Therefore, a filter bank with a stride of 2 will produce an output with (M / 2) × (N / 2) pixels, where M × N is the input image size. All inputs except for the fully connected kernel inputs are padded so that setting the stride to 1 will produce an output with the same number of pixels as the input. The output of each convolutional group is fed as input to the next convolutional layer.

[0040] like Figure 3 As depicted in the embodiment, the neural network (210) consists of four such convolutional layers:

[0041] • The first filter bank (305) has a size of 3×3×1×4, a step size of 2 and 4 biases, followed by the activation function of the first rectified linear unit (ReLU).

[0042] • The second filter bank (310) is 3×3×1×8 in size, with a step size of 2 and 8 biases, followed by a second ReLU.

[0043] • The third filter bank (315) is 7×7×2×16 in size, with a step size of 5 and 16 biases, followed by a third ReLU.

[0044] • The fourth filter bank (320), which is 48×27×16×3 in size, is fully connected and has 3 biases and one 1×3 output (212) that represents the estimated minimum, medium and maximum brightness levels of the HDR image corresponding to the SDR input.

[0045] In this embodiment, the NN (210) is trained on pairs of HDR and SDR images. For example, a large set of HDR images is mapped to corresponding SDR images using tone mapping operations, as described in references [1] and [2]. This process includes analyzing reference HDR metadata (e.g., minimum, intermediate, and maximum luminance values) from the HDR source images used during the tone mapping process. The goal of the network is to learn the relationship between the metadata from the estimated HDR images and the metadata from the reference HDR images. In one embodiment, this is accomplished by iteratively computing the predicted HDR metadata using a neural network architecture and minimizing the error between the predicted and reference HDR metadata by propagating the error back to the network weights. Training terminates when the error between the reference and predicted metadata is within a small threshold or reaches a non-decreasing steady-state period.

[0046] Alternatively, in another embodiment, after generating the predicted HDR metadata, the predicted metadata is applied to convert the input SDR image to its corresponding HDR image. The source HDR image and the predicted HDR image are compared, and the error is backpropagated to the network. It has been observed that training based on the error between the original image and the predicted image produces a neural network with better performance than training based on the error between the original metadata and the predicted metadata.

[0047] Given the predicted HDR metadata (212), step 215 generates the optimal mapping curve to be used by the display mapping process (220). Note that neural network 210 does not generate a mapping for a specific display. Therefore, the output of this SDR to HDR mapping may exceed the capabilities of the target display, thus requiring a second HDR (predicted image) to HDR (display) mapping that takes into account the characteristics of the target display. This second HDR to HDR mapping can be skipped if the generated HDR data is simply stored offline or transmitted for display by another downstream device.

[0048] For example, in one embodiment, the predicted HDR metadata (212) is processed to generate a “forward mapping” curve to map an HDR image with the predicted HDR metadata to the SDR signal range (see reference [3] or reference [4]). In an additional step, the forward mapping curve can be reversed to generate an “inverse mapping” curve that converts the SDR signal range of the source image to the HDR signal range of the predicted HDR image. The inverse mapping curve is then further adjusted based on the characteristics of the target display (such as its minimum and maximum brightness) or other parameters (such as desired contrast or ambient light) to map the dynamic range of the predicted HDR image. Finally, in step 220, the mapping process is displayed to generate a final HDR image (222) for the target display using the input SDR image (202) and the mapping curve (217) obtained in step 215 (see, for example, references [1-2]).

[0049] Local tone mapping adaptive

[0050] Since the generated mapping curve (217) is applied to the entire image (202), the upconversion process 200A can be considered a global dynamic range mapping process. As described in more detail in reference [2], the display mapping process 220 can be further improved by taking into account the local contrast and detail information of the input image. For example, as described in the appendix, the input image can be divided into two layers using downsampling and upsampling / filtering processes: a filtered base layer image and a detail layer image. By applying the tone mapping curve (217) to the filtered base layer and then adding the detail layer back to the result, the original contrast of the image can be preserved globally and locally. This can be referred to as “detail preservation” or “precise rendering”.

[0051] Therefore, display mapping can be performed as a multi-stage operation:

[0052] a) Generate a base layer (BL) image to guide SDR to HDR mapping;

[0053] b) Perform tone mapping on the base layer image;

[0054] c) Add the detail layer image to the tone-mapped base layer image.

[0055] In reference [2], the generated base layer (BL) represents a spatially blurred, edge-preserving version of the original image. That is, it preserves the important edges but blurs the finer details. More specifically, generating a BL image may include:

[0056] • Use the intensity of the original image to create an image pyramid with layers of lower resolution, and save each layer;

[0057] • Starting with the lowest resolution layer, upsample to higher layers to generate the base layer. Examples of generating base and detail layer images can be found in reference [2] or in the appendix of this specification.

[0058] Figure 2B An example embodiment of the inverse mapping and display management process (200B) using a second neural network (230) that leverages a pyramid representation and accurate rendering of the input image is depicted. Figure 2B As depicted, process 200B includes a new block (225) that provides the intensity (I) of the original image and generates the base layer (I). BL (BL) image and detail layer image (I) DL (DL). In this embodiment, the pixels (x, y) of the detail layer image are generated as

[0059] I DL (x, y) = I(x, y) - I BL (x, y)*dg, (1)

[0060] Where dg represents the detail gain scaler in [0, 1].

[0061] The predictive HDR detail neural network (230) takes two channels as input: the detail layer (DL) of the SDR image and the intensity (I) channel of the source SDR image. The predictive HDR detail neural network generates a single-channel predictive detail layer (PDL) image with the same resolution as the detail layer image, containing residual values ​​to be added to the detail layer image. In this embodiment, the detail layer residuals stretch the local contrast of the output image to increase its perceptual contrast and dynamic range. By utilizing both the detail layer input and the input image, the neural network can predict contrast stretching not only based on the content of the detail layer but also based on the content of the source image. To some extent, this provides the network with the possibility of correcting any problems that might arise from decomposing fixed-precision rendering into a base image and a detail image.

[0062] Given a block 225 in 200B that has already generated a brightness image I, block 205 in 200A can be simplified by performing appropriate downsampling only on I to achieve an appropriate input resolution (e.g., 960×540) for the neural network 210.

[0063] The neural network 230 consists of convolutional layers and residual network (ResNet) layers. For example... Figure 4AAs depicted in the embodiment, each ResNet block (410) comprises two convolutional layers (405a, 405b) with ReLU units, wherein the input (402) of each ResNet unit is added to the output of the second convolutional layer (405b) to generate a ResNet output (407). In the embodiment, each convolutional layer (405) has a 3×3×32×32 filter bank with no bias and a stride of 1.

[0064] like Figure 4B The neural network 230 for predicting HDR detail layers, as depicted in the diagram, consists of an input convolution (420), followed by five ResNet blocks (410) (each in...). Figure 4A The network consists of a final ReLU and an output convolution (430), which is described in the diagram. The network output forms an M×N image with the same size as the input detail layer image. This output is then added to the input detail layer image to form the final detail layer image. In some embodiments, a subsampled version can be used to reduce complexity instead of using a full-resolution (M×N) input image (I, DL). The output residual image (PDL) can then be upscaled to full resolution.

[0065] In one embodiment, convolutional network 420 has: input M×N×2; filter bank: 3×3×2×32, stride 1, no bias; and output M×N×32. Similarly, convolutional network 430 has: input M×N×32; filter bank: 3×3×32×1, stride 1, no bias; and output M×N×1.

[0066] The network can be trained on HDR and SDR image pairs. In an embodiment, a large set of HDR images is mapped to SDR using a tone mapping operation, as described in reference [2]. The pair is then processed by an HDR detail layer prediction NN, where the error signal between the reference image and the predicted HDR image is propagated to the weights of the neural network. Training terminates when the error falls below a threshold or when a non-decreasing stationary period is reached. During the training of NN 230, the dg scaler in equation (1) can be set to 1.

[0067] In some embodiments, the basic layer I can be used directly. BL Or it can be used in conjunction with the input intensity image I, such as

[0068] I B =α*I BL +(1-α)*I, where α is a scaler in [0,1]. When α=0, tone mapping is equivalent to traditional global tone mapping (see procedure 200A). When α=1, tone mapping is performed only on the base layer image.

[0069] Given IDL ,,image ID An optional scaler β in [0, 1] on L can be used to adjust the sharpening of the tone-mapped output to generate the final tone-mapped image.

[0070] I′=I′ BL +I DL *β, (2)

[0071] Among them, I′ BL Indicate I BL (or I) B The tone-mapped version of ). When using the prediction neural network 230, then

[0072] I′=I′ BL +(I DL +PDL)*β。 (3)

[0073] In an alternative implementation, process 200B can be simplified by bypassing (removing) the neural network (NN) used for HDR detail layer prediction (230) and by using only the original detail layer (DL). Therefore, given a pyramid representation of the input SDR image, process 200B can be adjusted as follows:

[0074] • In block 225, the intensity of the input image is divided into a base layer and a detail layer;

[0075] • As previously mentioned, the output (212) of the NN used for HDR metadata generation is used to generate the mapping curve 217;

[0076] • Use mapping curves to generate optimized mappings of only the base layers of the input image;

[0077] • Add the original detail layer (DL) to the optimization map to generate the final HDR image (e.g., see Equation (2)).

[0078] References

[0079] Each of the references listed in this article is cited in its entirety.

[0080] 1. U.S. Patent 9,961,237, "Display management for high dynamic range video", R. Atkins.

[0081] 2. PCT application PCT / US2020 / 028552, filed on April 16, 2020, WIPO Publication No. WO / 2020 / 219341, “Display management for high dynamic range images”, R. Atkins et al.

[0082] 3. U.S. Patent 8,593,480, “Method and apparatus for image data transformation”, A. Ballestad and A. Kostin.

[0083] 4. U.S. Patent 10,600,166, "Tone curve mapping for high dynamic range images", J.A. Pytlarz and R. Atkins.

[0084] Example computer system implementation

[0085] Embodiments of the present invention may be implemented using computer systems, systems configured with electronic circuits and components, integrated circuit (IC) devices (such as microcontrollers, field-programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs)), and / or means including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or implement instructions relating to image transformation, as described herein. The computer and / or IC may calculate any of the various parameters or values ​​relating to the image upconversion and display mapping process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0086] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors, such as those in a display, encoder, set-top box, transcoder, etc., can implement the methods related to image upconversion and display mapping as described above by executing software instructions in a processor-accessible program memory. The present invention can also be provided in the form of a program product. The program product may include any tangible and non-transitory medium carrying a set of computer-readable signals, including instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. The program product according to the present invention can take any of a variety of tangible forms. The program product may include, for example, physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash memory RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0087] In the case of the components mentioned above (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise specified, references to such components (including references to “modules”) should be interpreted as including any component that performs the function of the described component as an equivalent of that component (e.g., functionally equivalent), including components that are structurally different from those that perform the functions of the disclosed structures in the illustrated exemplary embodiments of the invention.

[0088] Equivalents, extensions, alternatives and miscellaneous

[0089] Therefore, example embodiments involving image dynamic range conversion and display mapping are described. In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the sole and exclusive indication of the invention, and of the invention as desired by the applicant, is the set of claims published in a particular form according to this application, wherein such claim publication includes any subsequent amendments. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, characteristics, features, advantages, or attributes not expressly recited in the claims shall not in any way limit the scope of such claims. Therefore, this specification and drawings should be viewed in an illustrative rather than restrictive sense.

[0090] appendix

[0091] Example process for generating base layer and detail layer images from an input image.

[0092] Pyramid downsampling

[0093] Given an intensity image of a source image, a base layer (BL) image can be constructed by combining downsampling and upsampling operations on the intensity image. In some embodiments, during downsampling, layers of the pyramid can be skipped to reduce memory bandwidth. For example, for a 4K input image, the first layer (e.g., 2K resolution) can be skipped. Then, during upsampling, the quarter-resolution image is simply doubled twice. For example, in an embodiment, given a 4K input, the pyramid can generate the following layers: 1024×576, 512×288, 256×144, 128×72, 64×36, 32×18, and 16×9. Similarly, for an 8K input image, half-resolution and quarter-resolution layers can be skipped. This ensures that subsequent layers of the pyramid will have the same size regardless of the input image size.

[0094] Although the pyramid is described using subsampling factors of 2, 4, 8, etc., other subsampling factors can be used without loss of generality.

[0095] As an example, when creating a pyramid, the k-th row of the nth pyramid level (e.g., n = 2 to 7) is generated by appropriately filtering the rows 2*k and 2*k-1 of the previous level. In an embodiment, this filtering is performed using a separable low-pass 2×2 filter (e.g., with filter coefficients [1 1] / 2) or a separable 4×4 low-pass filter (e.g., with filter coefficients [13 3 1] / 8). The 4×4 filter allows for better alignment between pyramid levels but requires additional row buffers. In another embodiment, different filters can be applied in the horizontal and vertical directions, for example, a 4-tap horizontal filter and a 2-tap vertical filter, and vice versa.

[0096] Before calculating the first level of the pyramid (e.g., 1024×576), the input image can be filled as follows:

[0097] • Ensure that all spatial dimensions of the pyramid, from smallest to largest, are divisible by 2.

[0098] • Copy the boundary pixels to account for the specified region of interest (ROI).

[0099] • Replicate boundary pixels to accommodate input images of various sizes or aspect ratios.

[0100] Sampling on the pyramid

[0101] In upsampling, the processor receives downsampled pyramid data and reconstructs the original image at its original resolution using an edge-aware upsampling filter at each layer. Upsampling is first performed on the smallest level of the pyramid, then on the other levels, upsampling continues until the resolution of the highest pyramid level is reached.

[0102] Let the pyramid image at layer i be denoted as P(i). Starting from the lowest resolution level (e.g., i = 7), the lowest resolution pyramid image (e.g., P(7)) is fed to an edge-preserving filter, which generates two coefficient “images” denoted as Ima(7) and Imb(7) (defined below). Next, both Ima and Imb are upsampled by a factor of two to generate upsampled coefficient images ImaU(7) and ImbU(7).

[0103] In the next layer, i=6, the P(6) layer of the pyramid is combined with the upsampled coefficient images ImaU(7) and ImbU(7) to generate an image.

[0104] S(6) = ImaU(7)*P(6) + ImbU(7), (4) This image, along with image P(6), is fed to an edge upsampling filter to generate the coefficient "images" Ima(6) and Imb(6). Next, Ima(6) and Imb(6) are upsampled by a factor of two to generate the upsampled coefficient images ImaU(6) and ImbU(6). The same process continues for other pyramid layers. Generally, for i = 7, 6, 5, ..., 2,

[0105] S(i-1)=ImaU(i)*P(i-1)+ImbU(i), (5)

[0106] The operation "*" that multiplies the coefficient images corresponds to multiplying their corresponding pixels pixel by pixel. For example, at pixel position (m, n), for pyramid level i of size W(i) × H(i),

[0107] S(i-1) m,n =ImaU(i) m,n *P(i-1) m,n +InbU(i) m,n (6)

[0108] m=1, 2, ..., W(i-1) and n=1, 2, ..., H(i-1).

[0109] After processing the second level of the pyramid (i=2), given S(1) and P(1), the edge filter will generate two parametric images Ima(1) and Imb(1). To generate a 4K image, Ima(1) and Imb(1) can be magnified by a factor of 2. To generate an 8K image, Ima(1) and Imb(1) can be magnified by a factor of 4. Combined with the intensity image (I) of the input video, the two magnified coefficient images (ImaU(1) and ImbU(1)) can be used to generate the base layer image, such as

[0110] BL=I BL=ImaU(1)*I+ImbU(1). (7)

[0111] In summary, given a pyramid with N layers of images (e.g., P(1) to P(N)), generating coefficient images Ima(1) and Imb(91) involves:

[0112] Ima(N) and Imb(N) are generated using an edge filter and an N-layer pyramid image P(N) with the lowest spatial resolution;

[0113] ImaU(N) and ImbU(N) are generated by scaling up Ima(N) and Imb(N) to match the spatial resolution of layer N-1;

[0114] For (i = N-1 to 2){

[0115] S(i)=ImaU(i+1)*P(i)+InbU(i+1)

[0116] Ima(i) and Imb(i) are generated using edge filters, S(i), and P(i).

[0117] }

[0118] Calculate S(1) = ImaU(2) * P(1) + ImbU(2); and

[0119] Ima(1) and Imb(1) are generated using edge filters, S(1) and P(1).

[0120] In the embodiments, S(i), P(i), P 2 Each of the inputs (i) and P(i)*S(i) is convolved horizontally and vertically using a 3×3 separable low-pass filter (e.g., H =

[121] / 4). Their corresponding outputs can be represented as Sout, Pout, P2out, and PSout. Although these signals are specific to each layer, the index i is not used for simplicity. Thus,

[0121]

Claims

1. A method for dynamic range conversion and display mapping, the method comprising: Access the input image with a first dynamic range and a first spatial resolution; Generate an intensity image based on the input image; The intensity image is applied to a first neural network to generate predictive statistics when the intensity image is mapped in a second dynamic range that is higher than the first dynamic range; as well as A tone mapping curve is generated based on the statistical data of the intensity image and the predicted statistical data. Generate a base layer image and a detail layer image based on the intensity image; The tone mapping curve is applied to the base layer image to generate a tone-mapped base layer image in the second dynamic range. The intensity image and the detail layer image are applied to the second neural network to generate the residual layer image in the second dynamic range; The residual layer image is added to the detail layer image to generate a second detail layer image; as well as The second detail layer image is added to the tone-mapped base layer image to generate the output image in the second dynamic range.

2. The method of claim 1, further comprising generating a mapped output image based on the input image and the tone mapping curve for display on a display having a target dynamic range.

3. The method of claim 1 or 2, further comprising applying the tone mapping curve to display the input image on a display having a target dynamic range different from the second dynamic range.

4. The method as described in claim 1 or 2, wherein, The statistical data of the intensity image includes the intensity value of the intensity image in the first dynamic range, and the prediction statistical data includes the predicted intensity value in the second dynamic range.

5. The method as described in claim 1 or 2, wherein, The statistical data of the intensity image includes the minimum intensity value, average intensity value, and maximum intensity value of the intensity image in the first dynamic range, and the prediction statistical data includes the predicted minimum intensity value, predicted average intensity value, and predicted maximum intensity value in the second dynamic range.

6. The method as described in claim 1 or 2, wherein, The first neural network comprises four layers, wherein, The first layer includes a first filter bank with a size of 3 × 3 × 1 × 4, a stride of 2 and 4 biases, followed by a first rectified linear unit ReLU activation function; The second layer includes a second filter bank with a size of 3 × 3 × 1 × 8, a stride of 2 and 8 biases, followed by a second ReLU; The third layer includes a third filter bank with a size of 7 × 7 × 2 × 16, a stride of 5, and 16 biases, followed by a third ReLU; and The fourth layer includes a fourth filter bank of size 48 × 27 × 16 × 3, which is fully connected and has 3 biases and a 1 × 3 output, which represents the prediction statistics when the intensity image is mapped in the second dynamic range.

7. The method of claim 1, wherein, The base layer image represents a spatially blurred, edge-preserved version of the intensity image, and generating the detail layer image includes calculations: , Among them, at pixel position ( x , y ) place, This represents the detail layer image. This represents the base layer image. I Represents the intensity image, and dg This represents the detail gain scaler in [0, 1].

8. The method of claim 1, further comprising: The detail layer image is added to the tone-mapped base layer image to generate the output image in the second dynamic range.

9. The method of claim 1, wherein, The second neural network includes: The input is a convolutional network, followed by five ResNet residual network blocks, then a final ReLU and an output convolutional network.

10. The method of claim 9, wherein, The input convolutional network includes: M × N × 2 inputs; Filter bank: 3 × 3 × 2 × 32, step size 1, no bias; and M × N × 32 output; and The output convolutional network includes: M × N × 32 inputs; Filter bank: 3 × 3 × 32 × 1, step size 1, no bias; and Output M × N × 1, where M and N are integers.

11. The method of claim 9, wherein, The residual network block includes: A first ReLU, followed by a first convolutional layer, followed by a second ReLU, followed by a second convolutional layer, followed by an adder. The adder is used to add the input of the ResNet network block to the output of the second convolutional layer to generate a ResNet output. Each of the first and second convolutional layers has a 3 × 3 × 32 × 32 filter bank with no bias and a stride of 1.

12. The method of claim 8, further comprising: The application displays a mapping process to map the output image to a display in a target dynamic range that is different from the second dynamic range.

13. The method of claim 1 or 2, further comprising reducing the first spatial resolution of the intensity image before applying the intensity image to the first neural network.

14. An apparatus comprising a processor and configured to perform the method as claimed in any one of claims 1 to 13.

15. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon for performing the method according to any one of claims 1 to 13 using one or more processors.

16. A computer program product comprising a computer program including computer executable instructions for performing the method according to any one of claims 1 to 13 using one or more processors.

Citation Information

Patent Citations

  • Tone curve mapping for high dynamic range images

    US10600166B2

  • Method and apparatus for image data transformation

    US8593480B1

  • Display management for high dynamic range video

    US9961237B2

  • Display management for high dynamic range images

    WO2020219341A1

  • Improved inverse tone mapping method and corresponding device

    EP3503019A1