Trim Path and Metadata Prediction in Video Sequences Using Neural Networks

A neural network-based method predicts trim-path metadata for HDR content, addressing the challenge of automatically generating metadata for converting HDR to SDR, and achieving improved display management and visual quality.

JP2025516767AActive Publication Date: 2025-05-30DOLBY LABORATORIES LICENSING CORP
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2024568247
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-01
Filing Date
2023-05-15
Publication Date
2025-05-30
Estimated Expiration
2043-05-15

AI Technical Summary

Technical Problem

Existing methods for converting high-dynamic range (HDR) content to standard dynamic range (SDR) content for legacy displays lack efficient techniques for automatically generating trim-path metadata, which are essential for optimal display management.

Method used

A neural network-based architecture is employed to predict trim-path metadata by extracting image features from HDR video sequences and mapping them to output trim-path metadata values, enabling automatic generation of metadata for tone mapping adjustments.

Benefits of technology

This approach allows for accurate and efficient generation of trim-path metadata, ensuring optimal tone mapping and display management for HDR content on SDR displays, thereby enhancing the visual quality and compatibility of HDR content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516767000006
    Figure 2025516767000006
  • Figure 2025516767000007
    Figure 2025516767000007
  • Figure 2025516767000008
    Figure 2025516767000008
Patent Text Reader

Abstract

A method and system for generating trim path metadata for high dynamic range (HDR) video are described. The trim path prediction pipeline includes a feature extraction network followed by a fully connected network that maps the extracted features to trim path values. In a first architecture, the feature extraction network is based on four cascaded convolutional networks. In a second architecture, the feature extraction network is based on a modified MobileNetV3 neural network. In either architecture, the fully connected network is formed by a set of three linear networks, each set being customized to best match its corresponding feature extraction network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims the benefit of priority based on U.S. Provisional Patent Application No. 63 / 342,306, filed on May 16, 2022, and European Patent Application No. 22 182 506.0, filed on Jul. 1, 2022, each of which is hereby incorporated by reference in its entirety.

[0002] [Technical Field] The present invention generally relates to images. More particularly, embodiments of the present invention relate to techniques for predicting trim - path metadata in a video sequence using a neural network.

Background Art

[0003] As used herein, the term "dynamic range" (DR) may relate to, for example, the ability of the human visual system (HVS) to perceive the range of intensities (e.g., luminance, luma) in an image from the darkest gray (black) to the brightest white (highlight). In this sense, DR relates to "scene - reference" intensity. DR may also relate to the ability of a display device to properly or approximately render a particular width of intensity range. In this sense, DR relates to "display - reference" intensity. At any point in the description herein, unless a particular meaning is explicitly specified to have a particular significance, it should be presumed that the term may be used in either sense, e.g., interchangeably.

[0004] As used herein, the term high dynamic range (HDR) relates to a DR range of approximately 14 to 15 orders of magnitude of the human visual system (HVS). In practice, the DR that a human can simultaneously perceive over a wide range in intensity can be somewhat truncated with respect to HDR. As used herein, the terms extended dynamic range (EDR) or visual dynamic range (VDR) may relate to the DR perceivable in a scene or image by the human visual system (HVS), including eye movements, which allows for some light adaptation changes across the scene or image, either individually or interchangeably with each other.

[0005] In practice, an image contains one or more color components (e.g., luma Y, and chroma Cb and Cr), and each color component is represented with a precision of n bits per pixel (e.g., n = 8). For example, using gamma luminance encoding, an image with n ≤ 8 (e.g., a 24-bit color JPEG image) is considered a standard dynamic range image, while an image with n ≥ 10 can be considered an extended dynamic range image. EDR and HDR images can also be stored and distributed using a high-precision (e.g., 16-bit) floating-point format such as the OpenEXR file format developed by Industrial Light and Magic.

[0006] Most consumer desktop displays currently support a luminance of 200 - 300 cd / m 2 or nits. Most consumer HDTVs are in the range of 300 - 500 nits, and new models are 1000 nits (cd / m 2) is reached. Thus, such conventional displays represent a lower dynamic range (LDR), also referred to as standard dynamic range (SDR), as opposed to HDR or EDR. As the availability of HDR content increases with the advancement of both capture devices (e.g., cameras) and HDR displays (e.g., Dolby Laboratories' PRM-4200 professional reference monitor), HDR content can be color-graded and displayed on an HDR display that supports a higher dynamic range (e.g., from 1,000 nits to 5,000 nits or more). Generally speaking, without limitation, the methods of the present disclosure relate to any dynamic range higher than SDR.

[0007] As used herein, the term "display management" refers to the processes performed on a receiver to render a picture for a target display. For example, without limitation, such processes may include tone mapping, gamut mapping, color management, frame rate conversion, and the like.

[0008] As used herein, the term "trim-pass" refers to the video post-production process in which a colorist or creative responsible for the content reviews the master grade of the content shot-by-shot and adjusts the lift, gamma, gain primary colors and / or other color parameters to create the desired color or effect. The parameters associated with this process (e.g., lift, gain, and gamma values) may be embedded as trim-pass metadata or "trims" within the video content to be used later as part of the display management process.

[0009] The creation and playback of high-dynamic range (HDR) content are becoming increasingly popular because HDR technology provides more realistic and vivid images than previous formats. However, when converting HDR content to SDR content for legacy displays, the broadcast infrastructure may not support the generation and transmission of custom trims. To improve existing encoding methods, as recognized herein by the inventors, improved techniques for automatically generating trim path metadata have been developed.

[0010] Patent Document 1 discloses a method for generating metadata used by a video decoder to display video content encoded by a video encoder, the method comprising accessing a target tone mapping curve, accessing a decoder tone curve corresponding to a tone curve used by the video decoder to tone map the video content, generating a plurality of parameters of a trim path function used by the video decoder for application after applying the decoder tone curve to the video content, the parameters of the trim path function being generated to approximate the target tone curve in combination with the trim path function and the decoder tone curve, and generating metadata used by the video decoder including the plurality of parameters of the trim path function.

[0011] Patent Document 2 discloses a method for automatic display management generation for gaming or SDR+ content. One or more specific image data feature types used when evaluating different candidate image data feature types and training a prediction model to optimize one or more image metadata parameters are identified. A plurality of image data features of one or more selected image data feature types are extracted from one or more images. The plurality of image data features of one or more selected image data feature types are grouped into a plurality of significant image data features. The total number of image data features among the plurality of significant image data features is not greater than the total number of image data features among the plurality of image data features of one or more selected image data feature types. The plurality of significant image data features are applied to train a prediction model for optimizing one or more image metadata parameters.

[0012] The approach described in this column is an approach that could be pursued, but is not necessarily an approach that has been previously considered or pursued. Therefore, unless otherwise indicated, none of the approaches described in this column should be assumed to be eligible as prior art merely by virtue of being included in this column. Similarly, the problems identified with respect to one or more approaches should not be assumed to be recognized in any prior art based on this column, unless otherwise indicated.

Prior Art Documents

Patent Documents

[0013]

Patent Document 1

Patent Document 2

Patent Document 3

[0014] [Non-Patent Document 1] “Module 2.8 - The Dolby Vision metadata Trim Pass, “ https: / / learning.dolby.com / hc / en-us / articles / 360056574431-Module-2-8-The-Dolby-Vision-Metadata-Trim-Pass-, downloaded May 3, 2022. (Reference [3]) [Non-Patent Document 2] A. Howard, et al., “Searching for MobileNetV3, " Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019, also arXiv:1905.02244v5, 20 Nov. 2019. (Reference [4]) [Non-Patent Document 3] S. Ioffe, and C. Szegedy. "Batch normalization: Accelerating deep network training by reducing internal covariate shift." International conference on machine learning. PMLR, 2015, also arXiv:1502.03167v3, 2 Mar. 2015. (Reference [5])

Non-Patent Document 4

Summary of the Invention

Means for Solving the Problems

[0015] The present invention is defined by the independent claims. The dependent claims relate to optional features of some embodiments of the present invention.

Brief Description of the Drawings

[0016]

Figure 1A

Figure 1B

Figure 2

Figure 3

Modes for Carrying Out the Invention

[0017] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0018] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings in which like reference symbols refer to similar elements and in which:

[0019] A method for trim-path metadata prediction for video is described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail in order to avoid unnecessarily obscuring, obscuring, or confusing the present invention.

[0020] [overview] Exemplary embodiments described herein relate to a method for generating trim path metadata for a video sequence. In one embodiment, a processor receives pictures in the video sequence. A feature extraction neural network extracts image features from the images and then passes them to a fully connected neural network that maps the image features to output trim path metadata values ​​for the input images.

[0021] [TrimPath Metadata Prediction Pipeline] In conventional display mapping (DM), the mapping algorithm applies a function similar to sigmoid (see, e.g., Reference [1] (Patent Document 3) and Reference [2] (Patent Document 4)) to map the input dynamic range to the dynamic range of the target display. Such a mapping function can be represented as a piecewise linear or non-linear polynomial characterized by anchor points, pivots, and other polynomial parameters generated using the characteristics of the input source and the target display. For example, in Reference [1] (Patent Document 3) and Reference [2] (Patent Document 4), the mapping function uses anchor points based on the luminance characteristics (e.g., minimum, intermediate (average), and maximum luminance) of the input image and the display. However, other mapping functions may use different statistical data such as luminance variance or luminance standard deviation values at the block level or for the entire image. In the case of an SDR image, the process may be assisted by additional metadata transmitted as part of the video being sent or calculated by the decoder or the display. For example, if the content provider has both an SDR version and an HDR version of the source content, the source can generate metadata (such as a piecewise approximation of the reshaping function in the forward or reverse direction) using both versions, which can assist the decoder in converting the input SDR image to an HDR image.

[0022] As used herein, the term "L1 metadata" refers to the minimum, intermediate, and maximum luminance values associated with an input frame or image. The L1 metadata may be calculated by converting the RGB data to a luma-chroma format (e.g., YCbCr) and then calculating the minimum, intermediate (average), and maximum values in the Y plane, or they may be calculated directly in the RGB space. For example, in one embodiment, L1Min represents the minimum value of the PQ-encoded min(RGB) value of the image when considering the active region (e.g., by excluding gray or black bars, letterbox bars, etc.). min(RGB) represents the minimum value of the color component values {R, G, B} of a pixel. The values of L1Mid and L1Max can also be calculated in the same manner by replacing the min() function with the average() and max() functions. For example, L1Mid represents the average of the PQ-encoded max(RGB) values of the image, and L1Max represents the maximum value of the PQ-encoded max(RGB) values of the image. In some embodiments, the L1 metadata can be normalized to be in [0,1].

[0023] Considering the L1Min value, L1Mid value, and L1Max value of the original HDR metadata, as well as the maximum (peak) and minimum (black) luminance of the target display indicated by Tmax and Tmin (see References [1] (Patent Document 3) and [2] (Patent Document 4)), an intensity tone-mapping mapping curve can be generated to map the intensity of the input image to the dynamic range of the target display. This can be considered an ideal single-stage tone-mapping curve to be matched by using the reconstructed metadata.

[0024] As described above, the term "trim" refers to tone curve adjustments performed by a colorist to improve the tone mapping operation. Trim is typically applied to the SDR range (e.g., a maximum luminance of 100 nits, a minimum luminance of 0.005 nits). These values are linearly interpolated to the target luminance range depending only on the maximum luminance. These values modify the default tone curve and exist for each trim.

[0025] Information about the trim may be part of the HDR metadata and can be used to adjust the original tone mapping curve (see References [1] (Patent Document 3), References [2] (Patent Document 4), and References [3] (Non-Patent Document 1)). For example, in Dolby Vision, the trim can be passed as Level 2 (L2) or Level 8 (L8) metadata that includes slope variables, offset variables, and power variables (collectively referred to as SOP parameters) representing gain and gamma values to adjust pixel values. For example, when the slope, offset, and power are within [-0.5, 0.5], given lift, gain, and gamma, it is shown as follows.

Equation

[0026] In certain content creation scenarios, it may not be possible to employ a full-range color grading tool to derive SDR content from HDR content. The embodiments described herein propose using a neural network-based architecture to automatically generate such trim path metadata. The trim path metadata is configured to effect an adjustment of the tone mapping curve applied to the input picture when displayed on the target display. The examples presented herein mainly explain how to predict slope, offset, and power values, although a similar architecture may be applied to directly predict lift, gain, or gamma, or other trims such as those described in reference [3] (Non-Patent Document 1).

[0027] FIG. 1A shows an exemplary trim prediction pipeline (100) for an HDR image (102) according to one embodiment. As shown in FIG. 1A, pipeline 100 includes the following modules or components. · HDR input (102) · Neural network (105) for feature extraction · A fully-connected neural network for mapping the extracted features to trim path metadata (110) · Trim path metadata output (112) (e.g., predicted SOP data).

[0028] The network (100) takes in a given frame (102) as input and passes it to a convolutional neural network (105) for feature extraction to identify high-level features of the image. The high-level features are then passed to a fully-connected network (110) that provides a mapping between the features and the trim metadata (112). Each of the high-level features corresponds to an image feature type in a set of image feature types. Each image feature type is represented by a plurality of related image features extracted from a corresponding set of training images and is used to train the convolutional neural network (105) for feature extraction. This network can be used as a stand-alone component or integrated into a video processing pipeline. To maintain the temporal consistency of the trim path metadata, the network can also take in a plurality of frames as input.

[0029] In one embodiment, the pipeline is formed from a neural network (NN) block trained to function with HDR images encoded using perceptual quantization (PQ) in the ICtCp color space as defined in Rec. BT. 2390, “High dynamic range television for production and international programme exchange.”

[0030] FIG. 1B shows an exemplary pipeline (130) for training the network 100. Compared to the system (100), this system also includes · a training input (132) of the true trim path metadata, and · an error / loss calculation module (120) that calculates an error function (e.g., mean squared error (MSE) or mean absolute error (MAE)) between the training data (132) and the predicted values (112), and · a backpropagation path (122) for training the networks 105 and 110, and includes.

[0031] [Trim Path Prediction Architecture] Two different trim path prediction architectures (100) are designed. These provide a balance between accuracy and speed. FIG. 2 shows a first architecture according to one embodiment.

[0032] In one embodiment, the neural network can be defined as a set of 4D convolutions, each followed by the addition of a constant bias value to all results. In some layers, the convolution is followed by clamping negative values to 0. The convolution is defined by their size in pixel units (M×N), the number of image channels they operate on (C), and the number of such kernels in the filter bank (K). In that sense, each convolution can be described by the size of the filter bank M×N×C×K. As an example, a filter bank of size 3×3×1×2 consists of two convolution kernels, each operating on one channel and having a size of 3 pixels × 3 pixels. The input and output sizes are shown as height × width × channels.

[0033] Some filter banks may also have a stride, i.e., it means that some of the results of the convolution are discarded. A stride of 1 means that each input pixel generates an output pixel. A stride of 2 means that only every other pixel in each dimension generates an output. Thus, a filter bank with a stride of 2 generates an output of (M / 2)×(N / 2) pixels, where M×N is the input image size. By setting the stride to 1, all inputs except the input to the fully connected kernel are padded so that an output having the same number of pixels as the input is generated. The output of each convolution bank is fed as input to the next convolution layer.

[0034] As shown in FIG. 2, in the first architecture, the feature extraction network (105) includes four convolution networks configured as follows. ·CONV1: Input 540×960×1, 3×3×1×4, stride = 2, output 270×480×4, bias, rectified linear unit (ReLU) activation. ·CONV2: 3×3×4×8, stride = 2, bias, ReLU activation, output: 135×240×8. ·CONV 3: 7×7×8×16, stride = 5, bias, ReLU activation, output: 27×48×16. ·CONV4: 27×48×16×3 (fully connected), stride = 1, bias, output: 1×1×3

[0035] Following the feature extraction network, the fully connected network (110) includes three linear networks configured as follows. ·Linear 1: Input: 1×3, output 1×6, Batch Norm, ReLU, DropOut ·Linear 2: Input: 1×6, output 1×6, Batch Norm, ReLU, DropOut ·Linear 3: Input: 1×6, output 1×3

[0036] DropOut refers to using dropout regularization, and Batch Norm refers to applying batch normalization (Reference [5] (Non-Patent Document 3)) to improve the training speed.

[0037] Figure 3 shows an example of a second architecture. As shown in Figure 3, in the second architecture, the feature extraction network (105) includes a modified version of the MobileNet_V3 network that was first described in reference [4] (non-patent document 2) for images with an aspect ratio of 1:1, but is here extended to operate on images of any aspect ratio. Table 1 (based on Table 2 of reference [4] (non-patent document 2)) provides additional details of the operation in the modified MobileNetV3 architecture for an example of a 540×960×3 input. Compared with reference [4] (non-patent document 2), after the second conv2d, 1×1 stage, the original pool, 7×7 stage is replaced by an average pool, 1×1 stage, generating a final output (e.g., 1×576), and two subsequent conv2d, 1×1, non-batch normalization (NBN) stages are removed.

[0038] In Table 1, "bneck" indicates bottleneck and inverted residual blocks, Exp size indicates the expansion block size, SE indicates whether squeeze and excite is enabled, NL indicates the non-linear activation function, HS indicates the hard-swish activation function, RE indicates the ReLU activation function, and s indicates the stride. The output of a neural network stage matches the input of the subsequent neural network stage.

[0039]

Table 1

[0040] Following the network 105, the fully-connected network (110) includes three linear networks configured as follows.

[0041] · Linear 1: Input: 1×576, Output: 1×128, Batch Norm, ReLU, DropOut · Linear 2: Input: 1×128, Output: 1×64, Batch Norm, ReLU, DropOut · Linear 3: Input: 1×64, Output: 1×3 As described in Reference [4] (Non-Patent Document 2), MobileNetV3 is tuned for the CPU of mobile phones for object detection and semantic segmentation (or high-density pixel prediction). This is · Depthwise convolution filters, and · Bottlenecks and inverted residual blocks (bneck), and · Squeeze and excitation (SE) blocks (Reference [6] (Non-Patent Document 4)), and · Hard-Swish (HS) activation function, that is,

Equation

[0042] [Fully Connected Network] The output of the feature extraction module (105) is fed into a fully connected network to obtain the predicted trim value. Based on the architecture of the feature extraction module, the fully connected network may have different input sizes and thus may require different output sizes. This module learns the mapping between the high-level features extracted from the image and its associated slope, offset, and power. As described above, using the same network design, lift, gain, and gamma, or any other parameters related to trim path metadata can be directly calculated.

[0043] In terms of complexity, the first architecture is simpler and less computationally intensive than the MobileNet-based architecture, but the MobileNet-based architecture is more accurate.

[0044] The second architecture can take three input channels (e.g., RGB or ICtCp) directly, and by enabling three input channels instead of just one, it is easy to modify the first architecture to take chroma into account as well.

[0045] [Network Training] As shown in Figure 1B, during training, either L1 loss (e.g., MAE) or L2 loss (e.g., MSE) can be applied during error calculation (120). This loss is calculated between the predicted trim path metadata value (112) and the ground truth trim path metadata value (132). Alternatively, a tone curve can be used as the loss function. Essentially, the trim path metadata value can be used together with other metadata values (such as L1) to define the tone curve corresponding to that image. Then, either L1 loss or L2 loss between the tone curves obtained using the predicted trim path metadata and the ground truth trim path metadata can be minimized.

[0046] For example, let true f(i) represent the tone curve generated using the L1 metadata for the input (102) and the training trim path metadata (132), and let f_pred(i) represent the tone curve generated using the same L1 metadata and the predicted trim path metadata (112). Then, during training, [Number] is defined as, where |i| represents the cardinality of the i values for which the MSE is calculated, and it may be desired to minimize the mean squared error. Alternatively, [Number] as shown, the L1 metric may be applied.

[0047] [References] Each of References [1] - [6] is hereby incorporated by reference in its entirety.

[0048] [Exemplary Computer System Implementations] Embodiments of the present invention may be implemented using a computer system, a system composed of electronic circuits and components, a microcontroller, a field programmable gate array (FPGA), or other integrated circuit (IC) devices such as configurable or programmable logic devices (PLDs), discrete time or digital signal processors (DSPs), application specific ICs (ASICs), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC can perform, control, or execute instructions related to image conversion, such as those described herein. The computer and / or IC can calculate any of a variety of parameters or values related to the trim path metadata prediction process described herein. Image and video embodiments can be implemented in hardware, software, firmware, and various combinations thereof.

[0049] Certain embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. may implement a method related to the trim path metadata prediction process as described above by executing software instructions in a program memory accessible to the processor. The present invention may be provided in the form of a program product. The program product can include any tangible non-transitory medium that carries a set of computer-readable signals that, when executed by a data processor, include instructions that cause the data processor to perform the methods of the present invention. The program product according to the present invention may be in any of a variety of tangible forms. The program product may include, for example, physical media such as floppy (registered trademark) disks, magnetic data storage media including hard disk drives, optical data storage media including CD ROMs, DVDs, ROMs, electronic data storage media including flash RAM, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0050] When components (e.g., software modules, processors, assemblies, devices, circuits, etc.) are described above, unless otherwise indicated, any reference to such components (including reference to "means") is to be construed as including any component that performs the function of the described component (e.g., is functionally equivalent) that is not structurally equivalent to the disclosed structure that performs the function in the illustrated exemplary embodiments of the present invention as an equivalent of such component.

[0051] [Equivalents, Extensions, Variations, and Others] Exemplary embodiments relating to a trim path metadata prediction process have been described above. In the above specification, embodiments of the present invention have been described with respect to many specific details that may vary from implementation to implementation. Therefore, the only and exclusive indicator of what the invention is and what the applicant intends to be the invention is the scope of the claims issued from this application, and the specific form in which such claims are issued includes any subsequent amendments. Any definitions explicitly set forth in this specification for terms included in such claims shall define the meaning of such terms as used in the claims. Accordingly, no limitations, elements, characteristics, features, advantages, or attributes not explicitly recited in the claims should limit the scope of such claims in any way. Thus, the specification and drawings should be considered in an illustrative rather than a limiting sense.

Claims

1. A method for generating trim path metadata for pictures in a video sequence, wherein the trim path metadata is configured to adjust a tone mapping curve applied to an input picture when displayed on a target display, the method comprising: receiving the input picture; providing a feature extraction network, the feature extraction network including a convolutional neural network for feature extraction trained to identify high-level image features of the input picture; applying the feature extraction network to the input picture to generate the image features; providing a fully connected network, the fully connected network including a plurality of cascaded linear neural networks trained to map the image features to output trim path metadata values for the input picture; applying the fully connected network to the image features to map the image features to the output trim path metadata values for the input picture; A method comprising the above.

2. The method according to claim 1, wherein the input picture is a high dynamic range (HDR) picture coded using PQ coding in the ICtCp color space.

3. The method according to claim 1, wherein the feature extraction network includes four cascaded convolutional networks.

4. The method according to claim 3, wherein the fully connected network includes three cascaded linear networks.

5. The method according to claim 1, wherein the feature extraction network includes a modified MobileNetV3 neural network that accepts inputs having a non-square aspect ratio.

6. The method according to claim 5, wherein the fully connected network includes three cascaded linear networks.

7. receiving input training trim path parameters corresponding to the input picture; applying an error loss unit to generate an error metric based on the input training trim path parameters and the output trim path metadata; training the feature extraction network and the fully connected network by minimizing the error metric; The method according to claim 1, further comprising the above. Claim 8 The method of claim 7, wherein calculating the error metric comprises calculating a minimum absolute error or a mean squared error between the input training trim path parameters and the output trim path metadata. Claim 9 Calculating the error metric comprises generating a first tone mapping function based at least on the input training trim path parameters; generating a second tone mapping function based at least on the output trim path metadata; calculating the minimum absolute error or the mean squared error between a value of the first tone mapping function and a value of the second tone mapping function; and the method of claim 7. Claim 10 An apparatus comprising a processor and configured to perform any one of the methods of claims 1 to 9. Claim 11 A non-transitory computer-readable storage medium storing computer-executable instructions for executing a method with one or more processors in accordance with any one of claims 1 to 9.

Citation Information

Patent Citations

  • Tone mapping processing, HDR video conversion method by automatic adjustment and update of tone mapping parameter, and device of the same

    JP2020017079A

  • Neural network circuit apparatus, neural network processing method and neural network execution program

    JP2020119462A

  • Tone curve optimization method and related video encoder and video decoder

    JP2020533841A

  • Situation identification device, situation learning device, and program

    JP2021135619A

  • HDR Image Representation Using Neural Network Mapping

    JP2021521517A