Trim Path Metadata Prediction in Video Sequences Using Neural Networks

A neural network-based approach automatically generates trim path metadata for HDR to SDR conversion, addressing inefficiencies in existing methods and ensuring effective tone mapping on target displays.

JP7819367B2Active Publication Date: 2026-02-24DOLBY LABORATORIES LICENSING CORP
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2024568247
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-07-01
Filing Date
2023-05-15
Publication Date
2026-02-24
Estimated Expiration
2043-05-15

AI Technical Summary

Technical Problem

Existing methods for generating trim path metadata to convert HDR content to SDR content are inefficient and lack automation, particularly in scenarios where full range of color grading tools is not available.

Method used

A neural network-based architecture is employed to automatically generate trim path metadata by extracting image features and mapping them to output values, using fully connected neural networks for tone mapping adjustments.

Benefits of technology

Enables efficient and accurate generation of trim path metadata, ensuring high-quality tone mapping on target displays without requiring full-range color grading tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007819367000006
    Figure 0007819367000006
  • Figure 0007819367000007
    Figure 0007819367000007
  • Figure 0007819367000008
    Figure 0007819367000008
Patent Text Reader

Abstract

A method and system for generating trim path metadata for high dynamic range (HDR) video are described. The trim path prediction pipeline includes a feature extraction network followed by a fully connected network that maps the extracted features to trim path values. In a first architecture, the feature extraction network is based on four cascaded convolutional networks. In a second architecture, the feature extraction network is based on a modified MobileNetV3 neural network. In either architecture, the fully connected network is formed by a set of three linear networks, each set being customized to best match its corresponding feature extraction network.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 342,306, filed May 16, 2022, and European Patent Application No. 22 182 506.0, filed July 1, 2022, each of which is incorporated by reference in its entirety.

[0002] [Technical field] The present invention relates generally to images. More particularly, embodiments of the present invention relate to techniques for predicting trim path metadata in video sequences using neural networks. [Background technology]

[0003] As used herein, the term "dynamic range" (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) in an image, e.g., from darkest gray (black) to brightest white (highlight). In this sense, DR relates to "scene-referred" intensity. DR may also relate to the ability of a display device to properly or approximately render an intensity range of a particular width. In this sense, DR relates to "display-referred" intensity. At any point in the description herein, unless it is explicitly specified that a particular meaning has a particular significance, it should be inferred that the terms may be used in either sense, e.g., interchangeably.

[0004] As used herein, the term high dynamic range (HDR) refers to a DR width that spans approximately 14 to 15 orders of magnitude of the human visual system (HVS). In practice, the DR that humans can simultaneously perceive across a wide range of intensity ranges may be somewhat truncated relative to HDR. As used herein, the terms extended dynamic range (EDR) and visual dynamic range (VDR) may individually or interchangeably refer to the DR perceivable in a scene or image by the human visual system (HVS), including eye movements, that allows for some light adaptation changes across a scene or image.

[0005] In practice, an image includes one or more color components (e.g., luma Y, and chroma Cb and Cr), each represented with n bits of precision per pixel (e.g., n=8). For example, using gamma luminance encoding, an image with n≦8 (e.g., a color 24-bit JPEG image) would be considered a standard dynamic range image, while an image with n≧10 would be considered an extended dynamic range image. EDR and HDR images can also be stored and distributed using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR file format developed by Industrial Light and Magic.

[0006] Most consumer desktop displays currently have a brightness of 200-300 cd / m 2 or nits (cd / m). Most consumer HDTVs are in the 300-500 nits range, with newer models supporting 1000 nits (cd / m). 2). Thus, such conventional displays represent a lower dynamic range (LDR), also referred to as standard dynamic range (SDR), as opposed to HDR or EDR. As HDR content becomes more available due to advances in both capture equipment (e.g., cameras) and HDR displays (e.g., Dolby Laboratories' PRM-4200 Professional Reference Monitor), HDR content can be color graded and displayed on HDR displays that support a higher dynamic range (e.g., 1,000 nits to 5,000 nits or more). Generally, but without limitation, the methods of the present disclosure relate to any dynamic range higher than SDR.

[0007] As used herein, the term "display management" refers to processes that occur on a receiver to render a picture for a target display. For example, but not limited to, such processes may include tone mapping, color gamut mapping, color management, frame rate conversion, etc.

[0008] As used herein, the term "trim-pass" refers to a video post-production process in which a colorist or creative responsible for content runs through a master grade of content on a shot-by-shot basis and adjusts lift, gamma, gain primaries, and / or other color parameters to create the desired color or effect. Parameters related to this process (e.g., lift, gain, and gamma values) may be embedded as trim-pass metadata or "trims" within the video content for subsequent use as part of the display management process.

[0009] The creation and playback of high dynamic range (HDR) content is currently becoming popular because HDR technology provides more realistic and lifelike images than previous formats. However, when converting HDR content to SDR content for legacy displays, broadcast infrastructure may not support the generation and transmission of custom trims. To improve upon existing encoding methods, improved techniques for automatically generating trim path metadata are being developed, as recognized by the inventors herein.

[0010] Patent document 1 discloses a method for generating metadata to be used by a video decoder to display video content encoded by a video encoder, the method including: accessing a target tone mapping curve; accessing a decoder tone curve that corresponds to the tone curve used by the video decoder to tone map the video content; generating a plurality of parameters of a trim path function to be used by the video decoder to apply after applying the decoder tone curve to the video content, wherein the parameters of the trim path function are generated such that a combination of the trim path function and the decoder tone curve approximates the target tone curve; and generating metadata to be used by the video decoder including the plurality of parameters of the trim path function.

[0011] Patent Literature 2 discloses a method for automatic display management generation for gaming or SDR+ content. Different candidate image data feature types are evaluated and one or more specific image data feature types are identified for use in training a predictive model for optimizing one or more image metadata parameters. A plurality of image data features of one or more selected image data feature types are extracted from one or more images. The plurality of image data features of the one or more selected image data feature types are aggregated into a plurality of significant image data features. The total number of image data features among the plurality of significant image data features is not greater than the total number of image data features among the plurality of image data features of the one or more selected image data feature types. The plurality of significant image data features are applied to train a predictive model for optimizing one or more image metadata parameters.

[0012] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Likewise, it should not be assumed that problems identified with one or more approaches have been recognized in any prior art based on this section, unless otherwise indicated. [Prior art documents] [Patent documents]

[0013] [Patent Document 1] U.S. Patent Application Publication No. 2021 / 076042 [Patent Document 2] U.S. Patent Application Publication No. 2021 / 350512 [Patent Document 3] Specifications of US Patent No. 8,593,480 (A. Ballestad and A. Kostin, US Patent 8,593,480, “Method and apparatus for image data transformation”, Reference [1])

Patent document 4

Non-licensed literature

[0014]

Non-licensed literature 1

Non-licensed Document 2

[0015] The invention is defined by the independent claims. The dependent claims relate to optional features of some embodiments of the invention. [Brief explanation of the drawings]

[0016] [Figure 1A] FIG. 1 illustrates an exemplary process for predicting trim path metadata for HDR video, in accordance with an exemplary embodiment of the present invention. [Figure 1B] FIG. 1B illustrates an exemplary process for training the neural network shown in FIG. 1A. [Figure 2] FIG. 1 illustrates a first example of a neural network architecture for generating trim path metadata according to an exemplary embodiment of the present invention. [Figure 3] FIG. 10 illustrates a second example of a neural network architecture for generating trim path metadata according to another exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0017] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0018] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings in which like reference symbols refer to similar elements and in which:

[0019] A method for trim-path metadata prediction for video is described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail in order to avoid unnecessarily obscuring, obscuring, or obfuscating the present invention.

[0020] [overview]

[0003] Exemplary embodiments described herein relate to a method for generating trim path metadata for a video sequence. In one embodiment, a processor receives pictures in the video sequence. A feature extraction neural network extracts image features from the images and then passes them to a fully connected neural network that maps the image features to output trim path metadata values ​​for the input images.

[0021] [TrimPath Metadata Prediction Pipeline] In traditional display mapping (DM), a mapping algorithm applies a sigmoid-like function (see, for example, Reference [1] and Reference [2]) to map the input dynamic range to the dynamic range of the target display. Such a mapping function can be expressed as a piecewise linear or nonlinear polynomial characterized by anchor points, pivots, and other polynomial parameters generated using the characteristics of the input source and the target display. For example, in Reference [1] and Reference [2], the mapping function uses anchor points based on the luminance characteristics of the input image and the display (e.g., minimum, mean (average), and maximum luminance). However, other mapping functions may use different statistics, such as luminance variance or luminance standard deviation values ​​at the block level or for the entire image. In the case of SDR images, the process may be aided by additional metadata, either transmitted as part of the transmitted video or calculated by the decoder or display. For example, if a content provider has both SDR and HDR versions of source content, the source may use both versions to generate metadata (such as a piecewise approximation of a forward or reverse reshaping function) to assist the decoder in converting an input SDR image into an HDR image.

[0022] As used herein, the term “L1 metadata” refers to the minimum, midpoint, and maximum luminance values ​​associated with an input frame or image. L1 metadata may be calculated by converting RGB data to luma-chroma format (e.g., YCbCr) and then calculating the minimum, midpoint (average), and maximum values ​​in the Y plane, or they may be calculated directly in RGB space. For example, in one embodiment, L1Min refers to the minimum PQ-encoded min(RGB) value of an image when considering active areas (e.g., by excluding gray or black bars, letterbox bars, etc.). min(RGB) refers to the minimum of a pixel's color component values ​​{R, G, B}. The L1Mid and L1Max values ​​may also be calculated in the same manner by replacing the min() function with the average() function and max() function. For example, L1Mid refers to the average PQ-encoded max(RGB) value of an image, and L1Max refers to the maximum PQ-encoded max(RGB) value of an image. In some embodiments, L1 metadata may be normalized to be in [0, 1].

[0023] Given the L1Min, L1Mid, and L1Max values ​​of the original HDR metadata and the maximum (peak) and minimum (black) luminance of the target display, denoted as Tmax and Tmin, an intensity tone-mapping curve can be generated that maps the intensity of the input image to the dynamic range of the target display, as described in References [1] and [2]. This can be considered an ideal single-stage tone-mapping curve to be matched by using the reconstructed metadata.

[0024] As mentioned above, the term "trim" refers to tone curve adjustments made by colorists to improve tone mapping operations. Trims are typically applied to the SDR range (e.g., maximum luminance of 100 nits, minimum luminance of 0.005 nits). These values ​​are linearly interpolated to the target luminance range, depending only on the maximum luminance. These values ​​modify the default tone curve and are present for each trim.

[0025] Information about the trim may be part of the HDR metadata and can be used to adjust the original tone mapping curve (see Reference [1] (Patent Document 3), Reference [2] (Patent Document 4), and Reference [3] (Non-Patent Document 1)). For example, in Dolby Vision, the trim can be passed as Level 2 (L2) or Level 8 (L8) metadata, which includes slope, offset, and power variables (collectively referred to as SOP parameters) that represent gain and gamma values ​​to adjust pixel values. For example, if the slope, offset, and power are within [-0.5, 0.5], then given the lift, gain, and gamma, it can be shown as follows:

number

[0026] In certain content creation scenarios, it may not be possible to employ a full range of color grading tools to derive SDR content from HDR content. The embodiments described herein propose using a neural network-based architecture to automatically generate such trim path metadata. The trim path metadata is configured to adjust the tone mapping curve applied to the input picture when it is displayed on a target display. While the examples presented herein primarily describe how to predict slope, offset, and power values, similar architectures can be applied to directly predict lift, gain, or gamma, or other trims such as those described in Reference [3].

[0027] 1A illustrates an exemplary trim prediction pipeline (100) for an HDR image (102) according to one embodiment. As shown in FIG. 1A, the pipeline 100 includes the following modules or components: HDR input (102) Neural Networks for Feature Extraction (105) A fully connected neural network for mapping extracted features to trim path metadata (110). · Trim path · Metadata output (112) (e.g. predicted SOP data).

[0028] The network (100) takes a given frame (102) as input and passes it to a convolutional neural network (105) for feature extraction, which identifies high-level features of the image. The high-level features are then passed to a fully connected network (110), which provides a mapping between the features and trim metadata (112). Each high-level feature corresponds to an image feature type from a set of image feature types. Each image feature type is represented by multiple associated image features extracted from a corresponding set of training images, which are used to train the convolutional neural network (105) for feature extraction. This network can be used as a standalone component or integrated into a video processing pipeline. To maintain temporal consistency of the trim path metadata, the network can also take multiple frames as input.

[0029] In one embodiment, the pipeline is formed from a neural network (NN) block trained to work with HDR images coded using perceptual quantization (PQ) in the ICtCp color space as defined in Rec. BT. 2390, “High dynamic range television for production and international program exchange.”

[0030] FIG. 1B shows an exemplary pipeline (130) for training network 100. In comparison to system (100), this system also: · True trim path · metadata training input (132) and an error / loss calculation module (120) that calculates an error function (e.g., mean squared error (MSE) or mean absolute error (MAE)) between the training data (132) and the predictions (112); A backpropagation path (122) for training the networks 105 and 110; Includes.

[0031] [Trimpath prediction architecture] Two different trim path prediction architectures (100) have been designed, which offer a balance between accuracy and speed. Figure 2 shows the first architecture according to one embodiment.

[0032] In one embodiment, a neural network can be defined as a set of four-dimensional convolutions, each of which is followed by the addition of a constant bias value to all results. In some layers, convolutions are followed by clamping negative values ​​to zero. Convolutions are defined by their size in pixels (M × N), how many image channels they operate on (C), and how many such kernels there are in the filter bank (K). In that sense, each convolution can be described by the filter bank's size M × N × C × K. As an example, a filter bank of size 3 × 3 × 1 × 2 consists of two convolution kernels, each of which operates on one channel and has a size of 3 pixels × 3 pixels. The input and output sizes are shown as height × width × channels.

[0033] Some filter banks may also have a stride, meaning that some results of the convolution are discarded. A stride of 1 means that every input pixel produces an output pixel. A stride of 2 means that only every other pixel in each dimension produces an output. Thus, a filter bank with a stride of 2 produces an output of (M / 2) × (N / 2) pixels, where M × N is the input image size. With padding = 1, all inputs except the input to the fully connected kernel are padded so that setting the stride to 1 produces an output with the same number of pixels as the input. The output of each convolution bank serves as the input to the next convolution layer.

[0034] As shown in FIG. 2, in the first architecture, the feature extraction network (105) includes four convolutional networks configured as follows: CONV1: Input 540x960x1, 3x3x1x4, stride=2, output 270x480x4, bias, rectified linear unit (ReLU) activation. ·CONV2: 3x3x4x8, stride=2, bias, activated ReLU, output: 135x240x8. CONV 3: 7x7x8x16, stride=5, bias, activated ReLU, output: 27x48x16. CONV4: 27×48×16×3 (fully connected), stride=1, bias, output: 1×1×3

[0035] Following the feature extraction network, the fully connected network (110) includes three linear networks configured as follows: Linear 1: Input: 1x3, Output: 1x6, Batch Norm, ReLU, DropOut Linear 2: Input: 1x6, Output: 1x6, Batch Norm, ReLU, DropOut Linear 3: Input: 1x6, Output: 1x3

[0036] DropOut refers to the use of dropout regularization, and Batch Norm refers to the application of batch normalization (Reference [5] (Non-Patent Document 3)) to increase the training speed.

[0037] Figure 3 shows an example of the second architecture. As shown in Figure 3, in the second architecture, the feature extraction network (105) comprises a modified version of the MobileNet_V3 network, which was first described in Reference [4] for images with a 1:1 aspect ratio, but is here extended to work for images of any aspect ratio. Table 1 (based on Table 2 in Reference [4]) provides additional details of the operation in the modified MobileNetV3 architecture for an example 540x960x3 input. Compared to Reference [4], after the second conv2d, 1x1 stage, the original pool, 7x7 stage is replaced by an average pool, 1x1 stage to produce the final output (e.g., 1x576), and the two subsequent conv2d, 1x1, non-batch normalized (NBN) stages are removed.

[0038] In Table 1, "bneck" indicates the bottleneck and inverted residual block, Exp size indicates the expansion block size, SE indicates whether squeeze and excite are enabled, NL indicates a nonlinear activation function, HS indicates a hard-swish activation function, RE indicates a ReLU activation function, and s indicates the stride. The output of a neural network stage matches the input of the subsequent neural network stage.

[0039] [Table 1]

[0040] Following network 105, the fully connected network (110) includes three linear networks configured as follows:

[0041] Linear 1: Input: 1x576, Output: 1x128, Batch Norm, ReLU, DropOut Linear 2: Input: 1x128, Output: 1x64, Batch Norm, ReLU, DropOut Linear 3: Input: 1x64, Output: 1x3 As described in Reference [4], MobileNetV3 is tuned for mobile phone CPUs for object detection and semantic segmentation (or dense pixel prediction). Depth-wise convolution filter, Bottleneck and inverted residual blocks (bneck), Squeeze and Excitation (SE) blocks (reference [6]), Hard-Swish (HS) activation function, i.e.,

number

[0042] [Fully connected network] The output of the feature extraction module (105) is fed into a fully connected network to obtain predicted trim values. Depending on the architecture of the feature extraction module, the fully connected network may have different input sizes and therefore require different output sizes. This module learns a mapping between high-level features extracted from the image and their associated slopes, offsets, and powers. As mentioned above, the same network design can be used to directly calculate lift, gain, and gamma, or any other parameters related to trim path metadata.

[0043] In terms of complexity, the first architecture is simpler and less computationally intensive than the MobileNet-based architecture, but the MobileNet-based architecture is more accurate.

[0044] The second architecture can directly take three input channels (e.g., RGB or ICtCp), but it is easy to modify the first architecture to also take chroma into account by allowing three input channels instead of just one.

[0045] [Network training] As shown in FIG. 1B, during training, either an L1 loss (e.g., MAE) or an L2 loss (e.g., MSE) may be applied during error calculation (120). This loss is calculated between the predicted trim path metadata values ​​(112) and the ground truth trim path metadata values ​​(132). Alternatively, a tone curve may be used as the loss function. Essentially, the trim path metadata values ​​may be used along with other metadata values ​​(e.g., L1) to define a tone curve corresponding to the image. Either the L1 loss or the L2 loss between the tone curve obtained using the predicted trim path metadata and the ground truth trim path metadata may then be minimized.

[0046] For example, f true Let (i) denote the tone curve generated using the L1 metadata for the input (102) and the training trim path metadata (132), and let fpred(i) denote the tone curve generated using the same L1 metadata and the predicted trim path metadata (112). Then, during training,

number

number

[0047] [References] Each of references [1]-[6] is incorporated by reference in its entirety.

[0048] Exemplary Computer System Implementation Embodiments of the present invention may be implemented using computer systems, systems comprised of electronic circuits and components, integrated circuit (IC) devices such as microcontrollers, field programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs), and / or apparatuses including one or more of such systems, devices, or components. The computers and / or ICs may perform, control, or execute instructions related to image transformations such as those described herein. The computers and / or ICs may calculate any of the various parameters or values ​​related to the trim-path metadata prediction process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0049] Certain embodiments of the present invention include a computer processor executing software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. may implement methods related to the trim path metadata prediction process described above by executing software instructions in a program memory accessible to the processor. The present invention may also be provided in the form of a program product. A program product may include any tangible, non-transitory medium bearing a collection of computer-readable signals that include instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. A program product according to the present invention may be in any of a wide variety of tangible forms. A program product may include physical media such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, and electronic data storage media including flash RAM. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0050] Where a component (e.g., a software module, processor, assembly, device, circuit, etc.) is described above, unless otherwise indicated, any reference to that component (including reference to "means") should be interpreted as including any component that performs the function of the described component (e.g., is functionally equivalent) as an equivalent of that component, including components that are not structurally equivalent to the disclosed structures that perform that function in the illustrated exemplary embodiments of the invention.

[0051] [Equivalents, extensions, variations, etc.] Exemplary embodiments of the TrimPath metadata prediction process have been described above. In the foregoing specification, embodiments of the present invention are described with reference to many specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indication of what the invention is and what the applicant intends to be the invention is the claims issuing from this application, including the specific form in which such claims are issued, including any subsequent amendments. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, no limitations, elements, properties, features, advantages, or attributes not expressly recited in a claim should in any way limit the scope of such claims. Accordingly, the specification and drawings should be regarded in an illustrative and not restrictive sense.

Claims

1. 1. A method of generating trim path metadata for a picture in a video sequence, the trim path metadata configured to adjust a tone mapping curve to be applied to an input picture when displayed on a target display, the method comprising: receiving the input picture; providing a feature extraction network, the feature extraction network including a convolutional neural network for feature extraction trained to identify high-level image features of the input picture; applying the feature extraction network to the input picture to generate the high-level image features; providing a fully-connected network, the fully-connected network including a plurality of cascaded linear neural networks trained to map the high-level image features to output trim-path metadata values ​​for the input picture; applying the fully connected network to the high-level image features to map the high-level image features to the output trim-path metadata values ​​for the input picture; Including, 10. A method wherein the feature extraction network comprises four cascaded convolutional networks.

2. The method of claim 1 , wherein the input picture is a high dynamic range (HDR) picture coded using PQ coding in ICtCp color space.

3. The method of claim 1 , wherein the fully connected network comprises three cascaded linear networks.

4. A method of generating trim path metadata for a picture in a video sequence, the trim path metadata configured to adjust a tone mapping curve that is applied to the input picture when displayed on a target display, the method comprising: receiving the input picture; providing a feature extraction network, the feature extraction network including a convolutional neural network for feature extraction trained to identify high-level image features of the input picture; applying the feature extraction network to the input picture to generate the high-level image features; providing a fully-connected network, the fully-connected network including a plurality of cascaded linear neural networks trained to map the high-level image features to output trim-path metadata values ​​for the input picture; applying the fully connected network to the high-level image features to map the high-level image features to the output trim-path metadata values ​​for the input picture; Including, The method, wherein the feature extraction network comprises a modified MobileNetV3 neural network that accepts inputs having non-square aspect ratios.

5. The method of claim 4 , wherein the fully connected network comprises three cascaded linear networks.

6. A method of generating trim path metadata for a picture in a video sequence, the trim path metadata configured to adjust a tone mapping curve that is applied to the input picture when displayed on a target display, the method comprising: receiving the input picture; providing a feature extraction network, the feature extraction network including a convolutional neural network for feature extraction trained to identify high-level image features of the input picture; applying the feature extraction network to the input picture to generate the high-level image features; providing a fully-connected network, the fully-connected network including a plurality of cascaded linear neural networks trained to map the high-level image features to output trim-path metadata values ​​for the input picture; applying the fully connected network to the high-level image features to map the high-level image features to the output trim-path metadata values ​​for the input picture; Including, receiving input training trim path parameters corresponding to the input picture; applying an error loss unit to generate an error metric based on the input training trimpath parameters and the output trimpath metadata values; training the feature extraction network and the fully connected network by minimizing the error metric; The method further comprises:

7. 7. The method of claim 6, wherein calculating the error metric comprises calculating a minimum absolute error or a mean squared error between the input training trim path parameters and the output trim path metadata values.

8. Calculating the error metric comprises: generating a first tone mapping function based on at least the input training trim path parameters; generating a second tone mapping function based on at least the output trim path metadata values; calculating a minimum absolute error or mean square error between the values ​​of the first tone mapping function and the values ​​of the second tone mapping function; The method of claim 6, comprising:

9. Apparatus comprising a processor and configured to perform any one of the methods according to claims 1 to 8.

10. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for carrying out a method on one or more processors according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Tone mapping processing, HDR video conversion method by automatic adjustment and update of tone mapping parameter, and device of the same

    JP2020017079A

  • Neural network circuit apparatus, neural network processing method and neural network execution program

    JP2020119462A

  • Tone curve optimization method and related video encoder and video decoder

    JP2020533841A

  • Situation identification device, situation learning device, and program

    JP2021135619A

  • HDR Image Representation Using Neural Network Mapping

    JP2021521517A