Trim-path correction for cloud-based encoding of HDR video

The method addresses temporal inconsistencies and overhead in cloud-based HDR video coding by using scene-based trim path correction and iterative segmentation, improving video quality and efficiency in multiprocessing environments.

JP7775296B2Active Publication Date: 2025-11-25DOLBY LABORATORIES LICENSING CORP
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
JP2023517839
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-08
Filing Date
2021-09-17
Publication Date
2025-11-25
Estimated Expiration
2041-09-17

Smart Images

  • Figure 0007775296000096
    Figure 0007775296000096
  • Figure 0007775296000097
    Figure 0007775296000097
  • Figure 0007775296000098
    Figure 0007775296000098
Patent Text Reader

Abstract

In a cloud-based system for encoding high dynamic range (HDR) video, each node receives video segments and bumper frames. To derive a scene-based forward reshaping function that minimizes the amount of reshaping-related metadata when encoding the video segments, each segment is subdivided into a primary scene and a secondary scene. If the parent scene of a secondary scene is processed by two or more neighboring nodes, the initial forward reshaping function and trim-path compensation parameters are adjusted using a reference tone mapping function and updated scene-based trim-path compensation parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 080,255, filed September 18, 2020; European Patent Application No. 20196876.5, filed September 18, 2020; U.S. Provisional Patent Application No. 63 / 089,154, filed October 8, 2020; and European Patent Application No. 20200781.1, filed October 8, 2020, which are hereby incorporated by reference.

[0002] technology The present disclosure relates generally to images. More particularly, some embodiments of the present invention relate to trim path correction for HDR video processing in cloud-based encoding architectures. [Background technology]

[0003] As used herein, the term "dynamic range" (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) in an image, e.g., from darkest gray (black) to brightest white (highlight). In this sense, DR relates to "scene-referred" intensity. DR may also relate to the ability of a display device to adequately or approximately render an intensity range of a particular width. In this sense, DR relates to "display-referred" intensity. Unless at any point in the description herein a particular meaning is explicitly designated as having particular importance, it should be presumed that the term can be used in either sense, e.g., interchangeably.

[0004] As used herein, the term high dynamic range (HDR) refers to a DR width spanning 14 to 15 orders of magnitude of the human visual system (HVS). In practice, the DR that humans can simultaneously perceive across a wide range of intensities may be somewhat truncated compared to HDR. As used herein, the terms visual dynamic range (VDR) or enhanced dynamic range (EDR), individually or interchangeably, may refer to the DR perceivable by the human visual system (HVS) within a scene or image, taking into account some light adaptation changes across the scene or image, and including eye movement. As used herein, VDR may refer to a DR spanning 5 to 6 orders of magnitude. Thus, although perhaps somewhat narrower than true scene-based HDR, VDR or EDR still represent a wide DR width and may be referred to as HDR.

[0005] In practice, an image contains one or more color components (e.g., luma Y and chroma Cb and Cr), each represented with n bits of precision per pixel (e.g., n = 8). For example, using gamma-luminance encoding, n <= 8 (e.g., a color 24-bit JPEG image) may be considered a standard dynamic range image, while n >= 10 may be considered an enhanced dynamic range image. HDR images may be stored and distributed using high-precision (e.g., 16-bit) floating-point formats such as the OpenEXR file format developed by Industrial Light and Magic.

[0006] Currently, most consumer desktop displays have a brightness of 200 to 300 cd / m 2 or nits luminance. Most consumer HDTVs are in the 300 to 500 nits range, with newer models supporting 1000 nits (cd / m 2). Such conventional displays are representative of low dynamic range (LDR), also known as standard dynamic range (SDR) as opposed to HDR. As HDR content becomes more available due to advances in both complementary equipment (e.g., cameras) and HDR displays (e.g., the PRM-4200 Professional Reference Monitor from Dolby Laboratories), HDR content may be color graded and displayed on HDR displays that support a higher dynamic range (e.g., 1000 nits to 5000 nits or higher).

[0007] As used herein, the term “forward reshaping” refers to the process of sample-to-sample or codeword-to-codeword mapping of a digital image from an original bit depth and original codeword distribution or representation (e.g., gamma, PQ, HLG, etc.) to the same or a different bit depth and a different codeword distribution or representation. The reshaping allows for improved compressibility or improved image quality at a fixed bit rate. For example, but not by way of limitation, reshaping may be applied to 10-bit or 12-bit PQ-coded HDR video to improve coding efficiency in a 10-bit video coding architecture. At a receiver, after decompressing the received signal (which may or may not have been reshaped), the receiver can apply an “inverse (or backward) reshaping function” to restore the signal to the original codeword distribution and / or achieve a higher dynamic range.

[0008] In many video distribution scenarios, HDR video may be encoded in multiprocessor environments, typically referred to as "cloud computing servers." In such environments, the trade-off between ease of computing, workload balancing among computing nodes, and video quality may force reshaping-related metadata to be updated frame-by-frame, which may result in unacceptable overhead, especially when transmitting video at low bitrates. Splitting a scene across multiple computing nodes may result in temporal inconsistencies in subscenes, which may affect how trim path data affects the output video. As understood by the inventors herein, improved techniques for trim path correction in cloud-based environments are desirable. The approaches described in this section could be pursued, but are not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise noted, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Likewise, unless otherwise noted, it should not be assumed that problems identified with one or more approaches have been recognized by the prior art based on this section. Each of the following references is incorporated by reference in its entirety. [Prior art documents] [Patent documents]

[0009] [Patent Document 1] U.S. Patent No. 10,575,028, H. Kadu et al., "Coding of high-dynamic range video using segment-based reshaping" [Patent Document 2] U.S. Patent No. 8,811,490, G.M. Su et al., "Multiple color channel multiple regression predictor" [Patent Document 3] International Publication No. 2019 / 217751, Q. Song et al., PCT Patent Application No. PCT / US2019 / 031620, "High-fidelity full reference and high-efficiency reduced reference encoding in end-to-end single-layer backward compatible encoding pipeline," filed May 9, 2019 [Patent Document 4] U.S. Patent No. 10,264,287, B. Wen et al., "Inverse luma / chroma mappings with histogram transfer and approximation" [Patent Document 5] U.S. Patent No. 10,397,576, H. Kadu and GM. Su, "Reshaping curve optimization in HDR coding" [Patent Document 6] U.S. Provisional Patent Application No. 63 / 049,673, G.M. Su et al., “Workload allocation and processing in cloud-based coding of HDR video,” filed July 9, 2020, also filed July 9, 2021 as PCT / US2021 / 040967. [Patent Document 7] U.S. Provisional Patent Application No. 63 / 080255, H. Kadu et al., "Recursive segment to scene segmentation for cloud-based coding of HDR video," filed September 18, 2020 [Patent Document 8] U.S. Patent No. 8,593,480, A. Ballestad and A. Kostin, “Method and apparatus for image data transformation” [Brief explanation of the drawings]

[0010] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference numerals refer to similar elements and in which:

[0011] [Figure 1A] 1 shows an exemplary single-layer encoder for HDR data using a reshaping function according to the prior art.

[0012] [Figure 1B] 1B shows an exemplary HDR decoder corresponding to the encoder of FIG. 1A according to the prior art.

[0013] [Figure 2] 1 illustrates an exemplary architecture and processing pipeline for cloud-based encoding of HDR video, according to one embodiment.

[0014] [Figure 3A] This example shows how to split a video input into segments and assign bumper frames to three nodes.

[0015] [Figure 3B] This example shows how to merge scene cuts to generate a list of primary scenes.

[0016] [Figure 3C] An example of a primary scene split across two computing nodes is shown.

[0017] [Figure 3D] 10 illustrates an example of a trim path window used to derive updated trim path correction parameters, according to an embodiment. [Figure 3E] 10 illustrates an example of a trim path window used to derive updated trim path correction parameters, according to an embodiment.

[0018] [Figure 4A] 1 illustrates an example of an iterative segment-to-scene segmentation process, according to an embodiment.

[0019] [Figure 4B] 1 illustrates an exemplary process for generating updated scene-based trim path correction parameters, according to an embodiment.

[0020] [Figure 4C] 1 illustrates an exemplary process for generating a trim-corrected forward reshape function, according to an embodiment.

[0021] [Figure 6A] 10 illustrates an example of trim path correction for a forward reshape function, according to one embodiment of the present invention. [Figure 6B] 10 illustrates an example of trim path correction for a forward reshape function, according to one embodiment of the present invention.

[0022] [Figure 5] 1 illustrates an exemplary encoder for scene-based encoding using reshaping, in accordance with one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0023] This application describes a method for trim path correction in cloud-based video coding of HDR video. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent that the present invention may be practiced without these specific details. In other instances, well-known structures and devices have not been described in exhaustive detail in order to avoid unnecessarily obscuring, obscuring, or obfuscating the present invention.

[0024] overview The exemplary embodiments described herein relate to cloud-based reshaping and coding of HDR images. In one embodiment, in a cloud-based system for encoding HDR video, a current node receives a first video sequence containing high dynamic range video frames. One or more processors within the node then: generating (455) for each video frame in the first video sequence a first forward reshaping function and frame-based trim path correction parameters, the forward reshaping function mapping frame pixels from a high dynamic range to a second dynamic range lower than the high dynamic range; For a subscene of the first video sequence that is part of the parent scene processed by the current computing node and one neighboring computing node, determining (457) a trim path parameter window including frames of the first video sequence and bumper frames, the bumper frames including video frames to be processed by the neighboring computing nodes; calculating (459) scene-based trim path correction parameters based on the frame-based trim path correction parameters of the trim path parameter window frame; generating a scene-based forward reshaping function based on the first forward reshaping function of frames in the subscene; and generating an output forward reshaping function based on the scene-based trim path correction parameters and the scene-based forward reshaping function for the subscene.

[0025] Exemplary HDR Encoding System 1A and 1B illustrate an exemplary single-layer backward-compatible codec framework with image reshaping according to the prior art. More specifically, FIG. 1A illustrates an exemplary encoder architecture that may be implemented in one or more computing processors in an upstream video encoder. FIG. 1B illustrates an exemplary decoder architecture that may also be implemented in one or more computing processors in one or more downstream video decoders.

[0026] Under this framework, given reference HDR content (120) and corresponding reference SDR content (125) (i.e., content representing the same image as the HDR content but color-graded and presented in standard dynamic range), reshaped HDR content (134) is encoded and transmitted as SDR content in a single layer of a coded video signal (144) by an upstream encoding device implementing an encoder architecture. The reshaped and then compressed SDR content is received and decoded in a single layer of the video signal by a downstream decoding device implementing a decoder architecture. Backward reshaping metadata (152) is also encoded and transmitted in the video signal along with the reshaped content, allowing an HDR display device to reconstruct the HDR content based on the (reshaped) SDR content and the backward reshaping metadata. Without loss of generality, in some embodiments, such as in non-backward compatible systems, the reshaped SDR content may not be viewable by itself but must be viewed in combination with a backward reshaping function that produces viewable SDR or HDR content. In other embodiments that support backward compatibility, legacy SDR decoders can still play received SDR content without the use of the backward reshaping function.

[0027] As shown in FIG. 1A , step 130 generates a forward reshaping function given an HDR image (120), an SDR image (125), and a target dynamic range. Given the generated forward reshaping function, a forward reshaping mapping step (132) is applied to the HDR image (120) to generate a reshaped SDR base layer (134). A compression block (142) (e.g., an encoder implemented according to any known video encoding algorithm, such as AVC, HEVC, AV1, etc.) compresses / encodes the SDR image (134) in a single layer (144) of the video signal. Furthermore, a backward reshaping function generator (150) can generate a backward reshaping function, which can be transmitted to a decoder as metadata (152). In some embodiments, the metadata (152) can represent the forward reshaping function (130), leaving it up to the decoder to generate the backward reshaping function (not shown).

[0028] Examples of backward reshaping metadata representing / specifying optimal backward reshaping functions may include, but are not necessarily limited to, any of the following: an inverse tone mapping function, an inverse luma mapping function, an inverse chroma mapping function, a look-up table (LUT), a polynomial, inverse display management coefficients / parameters, etc. In various embodiments, the luma backward reshaping function and the chroma backward reshaping function may be derived / optimized jointly or separately, and may be derived using a variety of techniques, such as, for example, without limitation, as described later in this disclosure.

[0029] The backward reshaping metadata (152) generated by the backward reshaping function generator (150) based on the reshaped SDR image (134) and the target HDR image (120) may be multiplexed as part of the video signal 144, for example as supplemental enhancement information (SEI) messaging.

[0030] In some embodiments, the backward reshaping metadata (152) is carried in the video signal as part of the overall image metadata, which is carried in the video signal separately from the single layer in which the SDR image is encoded in the video signal. For example, the backward reshaping metadata (152) may be encoded in a component stream in the encoded bitstream, which may or may not be separate from the single layer (of the encoded bitstream) in which the SDR image (134) is encoded.

[0031] In this way, the backward reshaping metadata (152) can be generated or pre-generated at the encoder side to take advantage of powerful computational resources and offline encoding flows available at the encoder side, including, but not limited to, content adaptive multiple passes, look-ahead operations, inverse luma mapping, inverse chroma mapping, CDF-based histogram approximation and / or transfer, etc.

[0032] The encoder architecture of FIG. 1A can be used to avoid directly encoding the target HDR image (120) into the coded / compressed HDR image in the video signal; instead, backward reshaping metadata (152) in the video signal can be used to enable a downstream decoding device to backward reshape the SDR image (134) (encoded in the video signal) into a reconstructed image that is identical to or a good / optimal approximation of the reference HDR image (120).

[0033] In some embodiments, as shown in FIG. 1B , a video signal encoded with a reshaped SDR image in a single layer (144) and backward reshaping metadata (152) as part of the overall image metadata are received as input at the decoder side of the codec framework. A decompression block (154) decompresses / decodes the compressed video data in the single layer (144) of the video signal into a decoded SDR image (156). Decompression 154 typically corresponds to the inverse of compression 142. The decoded SDR image (156) may be identical to the SDR image (134) subject to quantization errors in the compression block (142) and decompression block (154), which may be optimized for an SDR display device. In a backward-compatible system, the decoded SDR image (156) may be output in an output SDR video signal (e.g., through an HDMI interface, through a video link, etc.) for rendering on an SDR display device.

[0034] Optionally, alternatively, or additionally, in the same or another embodiment, the backward reshaping block 158 extracts backward (or forward) reshaping metadata (152) from the input video signal, constructs a backward reshaping function based on the reshaping metadata (152), and performs a backward reshaping operation on the decoded SDR image (156) based on the optimal backward reshaping function to generate a backward reshaped image (160) (or a reconstructed HDR image). In some embodiments, the backward reshaped image represents a production-quality or near-production-quality HDR image that is identical to or closely / optimally approximates the reference HDR image (120). The backward reshaped image (160) may be output in an output HDR video signal (e.g., through an HDMI interface, through a video link, etc.) to be rendered on an HDR display device.

[0035] In some embodiments, display management operations specific to an HDR display device may be performed on the rear reshaped image (160) as part of an HDR image rendering operation that renders the rear reshaped image (160) on an HDR display device.

[0036] Cloud-based coding Existing reshaping techniques may be frame-based, i.e., new reshaping metadata is sent for each new frame, or scene-based, i.e., new reshaping metadata is sent for each new scene. As used herein, the term "scene," with respect to a video sequence (a sequence of frames / images), may refer to a series of consecutive frames in the video sequence that share similar luminance, color, and dynamic range characteristics. Scene-based methods work well in video workflow pipelines where complete scenes are accessible. However, it is not uncommon for content providers to use cloud-based multiprocessing, where after dividing a video stream into segments, each segment is processed independently by a single computational node in the cloud. As used herein, the term "segment" refers to a series of consecutive frames in a video sequence. A segment may be part of a scene or may contain one or more scenes. Thus, the processing of a scene may be divided across multiple processors.

[0037] As discussed in U.S. Patent No. 6,119,149, in certain cloud-based applications, under certain quality constraints, segment-based processing may require generating reshaping metadata for each frame, resulting in undesirable overhead. This can be problematic in very low bitrate (e.g., less than 1 Mbit / s) applications. U.S. Patent No. 6,213,149, proposed a solution to this problem using a two-stage architecture including: a) an orderer stage implemented on a single computing node that assigns scenes to segments; and b) an encoding stage in which each node in the cloud encodes a sequence of segments. After the scenes are segmented, the proposed scene-to-segment allocation process involves one or more iterative steps using an initial random assignment of scenes to nodes, followed by a refined assignment based on optimizing the allocation cost across all nodes. In such an implementation, the total length of the video processed at each node may vary across all nodes.

[0038] The embodiment discussed in US Patent Application Publication No. 2009 / 0129997 provided an alternative solution: after a sequence is divided into segments such that each segment is processed by a separate node, each segment is subdivided into sub-segments (or scenes) at each node in a way that minimizes the need to update each sub-segment's corresponding reshaping function, thus minimizing the overhead required to transmit reshaping-related metadata.

[0039] In some embodiments, a reference SDR signal (e.g., 125) based on the HDR signal (120) may not have the desired "look." In this case, a colorist may adjust "trim parameters" (commonly referred to as lift, gain, and gamma (LGG)) to achieve the desired effect. This process is sometimes called a "trim pass," and its primary goal is to preserve the director's intent and look. As used herein, the term "trim pass" refers to the post-production process of producing a sequence for display with a target dynamic range typically lower than the master's dynamic range, and may include trimming, rotating, cropping, flipping, and adjusting the brightness, color, and saturation of the video sequence. For example, given a film mastered at 4000 nits, a colorist may produce "trims" at 400 nits, 500 nits, and 1000 nits. While such a trim pass may preserve the artistic intent in the SDR signal, it may also introduce objectionable clipping artifacts in the shadows (low-intensity areas) or highlights (high-intensity areas) after reshaping. Such artifacts may be further exaggerated by the video-to-node segmentation process and / or the video encoding process (142). While one could attempt to reduce such artifacts by modifying the LGG "trim" data itself, studios do not allow data changes after the director's approval, and such modifications would require an additional review process. Thus, as understood by the inventors, it would be beneficial to be able to reduce the artifacts introduced by the HDR-to-SDR mapping process. Such a methodology is now described.

[0040] As first disclosed in U.S. Patent Application Publication No. 2009 / 0129990, Figure 2 shows an exemplary architecture and processing pipeline for cloud-based encoding of HDR video according to one embodiment. Given a video source (202) for content distribution, typically called a mezzanine file, and a collection of worker nodes, each node (e.g., nodes 205-N) fetches (e.g., from an XML file) video pictures (or frames) and corresponding video metadata (207) to be processed as follows:

[0041] In the preprocessing stage 210, the mezzanine input is divided into segments, and each segment is assigned to a different computing node (e.g., node 205-N). These segments are mutually exclusive, meaning they have no frames in common. Each node also obtains a certain number of frames before the first frame in the segment and a certain number of frames after the last frame in the segment. These overlapping frames, called bumper frames, are used solely to maintain temporal consistency with the previous and next nodes, respectively. Bumper frames are not encoded by the node. Without loss of generality, in one embodiment, these video segments may all be of equal, fixed length, except perhaps for the segment assigned to the last node. As an example, Figure 3A shows a sample of distributing the mezzanine (305) into three segments (307-1, 307-2, 307-3) along with a bumper frame (e.g., 309) and assigning these frames to different nodes. Without limitation, for a 30 second long segment and a 2 second long bumper section on each side, example embodiments may include the following arrangement: 1800 frame segment section and 120 frame bumper section, 60fps 1500 frame segment section and 100 frame bumper section, 50fps 720 frame segment section and 48 frame bumper section, 24fps

[0042] After the preprocessing stage 210 is complete, each node has access to the frame and a two-pass approach is followed. In pass 1 (stages 215, 220), a list of scenes in the segment (222) is generated. Scene cuts extracted from the XML file (209) and those generated using the automatic scene cut detector (215) are combined in stage 220 to obtain a first list of primary scenes. Primary scenes belonging to a parent scene coded across multiple nodes may be subdivided into secondary scenes. To maintain temporal consistency in scenes distributed across multiple nodes, bumper frames and a novel recursive scene segmentation algorithm are also provided. The segmentation generates additional scenes, which are added to the first list of scenes to obtain a second list. This second list of scenes (222) is passed to pass 2. To reduce the computational overhead in pass 2, auxiliary data (212) needed to correct the trim pass data is also passed to pass 2. Pass 2 (stages 225, 230, 235) uses the scene list received from pass 1 to perform forward and backward reshaping for each scene in the segment. In the forward reshaping stage (225), trim pass data correction is performed. Forward reshaping (225) using a scene-based forward reshaping function (227) produces a reshaped SDR segment (229), and the backward reshaping unit (235) produces metadata parameters used by the decoder to reconstruct the HDR input. The reshaped SDR input (229) is compressed (230), and the compressed video data and reshaped metadata are combined together to produce a compressed bitstream (240).

[0043] For simplicity, let L denote the number of frames in a segment, and B denote the number of frames in each bumper section. The ith frame in the mezzanine is denoted by fi In one embodiment, the first node is a node that is located in the segment portion between frames f0 and f L-1 This node has no left bumper and its right bumper is in the frame range f L ~f L+B-1 The segment part of node N spans frame f (N-1)L ~f NL-1 where f (N-1)L-B ~f i(N-1)L-1 is the left bumper and f NL ~f NL+B-1 is the right bumper. The last node does not have a right bumper section and may have fewer than L frames in the segment.

[0044] Given a node N, node N-1 is the left / previous neighbor and node N+1 is the right / next or subsequent neighbor. A reference to a node that is the left / previous node with respect to N refers to all nodes from 0 to N-1. Similarly, a reference to a node that is the right / next with respect to N refers to all nodes from N+1 to the last node. We will now explain the two paths mentioned above in more detail.

[0045] Pass 1: Generate a scene from segments The main purpose of this pass is to generate a list of scenes within the segment assigned to the node. The process begins by detecting scene cuts in all frames assigned to the node, including the node segment and both bumper sections. Only the scene cuts within the segment are ultimately used for scene-based encoding by pass 2. However, scenes within the bumper sections are still useful for maintaining temporal coherence with neighboring nodes.

[0046] Colorist-specified scene cuts (209) are read from an XML file (207). An automatic scene cut detector (215) may identify possible scene cut locations. These scene cuts from the colorist and automatic detector are merged to obtain a first list of scenes, known as primary scenes. Primary scenes on segment boundaries are split using bumper frames and a novel scene splitting technique. Splitting primary scenes on segment boundaries creates additional scenes, called secondary scenes or subscenes. Secondary scenes are added to the first list of scenes to obtain a second list. This list is then used by pass 2 for scene-based encoding. Apart from the list of scenes, pass 2 may also require auxiliary data (212) for forward reshaping of secondary scenes. Each stage is described in more detail below.

[0047] Colorists and professional color graders typically process each scene as a unit. To achieve their goals (e.g., proper color grading, inserting fade-ins and fade-outs, etc.), they must manually detect scene cuts in a sequence. This information is stored in an XML file and can be used for other purposes. Every node reads only the significant scene cuts for that segment from the XML file. These scene cuts can be within the segment section or the bumper section.

[0048] XML scene cuts are defined by colorists, but are not completely accurate. For grading purposes, colorists sometimes introduce scene cuts in the middle of a dissolve scene or at the beginning of a fade-in or fade-out portion of a scene. These scene cuts, when taken into account in the reshaping phase, can cause flashing in the reconstructed HDR video and should generally be avoided. For this reason, an automatic scene-cut detector (Auto-SCD) 215 is also employed in some embodiments.

[0049] An automatic scene cut detector, or Auto-SCD, uses changes in luminance levels in different sections of consecutive video pictures to detect scene changes. Any scene cut detector known in the art can be used as the automatic detector. In one embodiment, such an automatic detector is unaware of parts of the video that dissolve, fade in, or fade out, yet is able to correctly detect all true scene cuts.

[0050] A potential problem with the automatic detector is false positives. There may be lighting changes in a scene due to camera pans, motion, occlusions, etc., and these lighting changes may also be detected as scene cuts by Auto-SCD. To discard these false positives, in one embodiment, the scene cuts from the XML file and the scene cuts from Auto-SCD are merged together in step 220. Those skilled in the art will understand that if no scene cuts are defined in the XML file, the output of the automatic scene detector may simply be used. Similarly, other embodiments may strictly rely on the scene cuts defined in the XML file. Alternatively, more than two scene cut detectors may be used, each detecting a different attribute of interest, and then a primary scene may be defined based on the combination of all of their results (e.g., their intersection or other set operation combination, e.g., their union, intersection, etc.).

[0051] Ψ XML N Let Ψ be the set of frame indices representing the scene start frames at node N as reported in the XML file. Similarly, Ψ Auto-SCD N Let N be the set of frame indices representing the scene start frames at node N as reported by Auto-SCD. In one embodiment, merging scene cuts from these two sets is equivalent to taking the intersection of these two sets. Ψ1N =Ψ XML N ∩Ψ Auto-SCD N (1) where Ψ1 N = indicates the first list of scene cuts (or scenes) at node N. These scenes are also called primary scenes. Figure 3B shows an example scenario. In this example, the XML file reports three scene cuts. Auto-SCD also reports three scene cuts; two are identical to the XML scene cuts, but the third is in a different location. Because only two are common among the six reported scene cuts, the node segment is divided into only three primary scenes (310-1, 310-2, 310-3) according to those two common scene cuts. In some embodiments, scene cuts in XML and Auto-SCD may be recognized as the same even if they are reported on different frames, as long as the scene cut indexes between the two lists differ within a given small tolerance (±n frames, e.g., n in [0, 6]).

[0052] As shown in Figure 3B, primary scene 2 (310-2) resides entirely in node N. Therefore, it can be processed entirely by node N in pass 2. Conversely, primary scenes 1 (310-1) and 3 (310-3) lie on a segment boundary. Their parent scenes are distributed across multiple nodes and are processed independently by those nodes. To ensure a consistent appearance in boundary frames encoded by different nodes, primary scenes 1 and 3 require some special treatment. We next consider several alternative scenarios.

[0053] Consider a simple scenario where P is a parent scene distributed across two nodes N and N+1, as shown in Figure 3C. Nodes N and N+1 have access to only a portion of the parent scene. We assume that these nodes process and encode their own portions of the parent scene (without bumpers). Since the reshaping parameters are computed for different sets of frames, we can use the last frame in node N's segment, i.e., f (N+1)L-1 and the first frame in the segment of node N+1, i.e., f (N+1)L The reshaped SDR and reconstructed HDR for a scene may appear visually different. Such visual differences typically manifest as flickering, flashing, or sudden brightness changes. This problem is called temporal inconsistency across nodes. Part of the reason for the inconsistency in the above scenario is the lack of a common frame when calculating the reshaping parameters. As shown in Figure 3C, including bumper frames when generating these statistics provides a smoother transition across nodes. However, because bumper sections can be relatively short compared to the parent scene, they may not be long enough to guarantee temporal consistency. In one embodiment, to address these issues, the portions of the parent scene at nodes N and N+1 are partitioned into secondary scenes or subscenes. Even if the reshaping statistics within a scene change significantly from one node to the next, these statistics do not change significantly from frame to frame. Secondary scenes use only statistics within a small neighborhood to evaluate the reshaping parameters. Therefore, these reshaping parameters do not change significantly from one subscene to the next. Thus, the partitioning achieves temporal consistency. Note that nearby subscenes can also be previous / next nodes.

[0054] It should be noted that splitting introduces additional scenes and therefore increases the metadata bitrate. The challenge is to achieve temporal coherence using a minimum number of splits to keep the metadata bitrate low. Bumper frames play an important role in achieving good visual quality while reducing the number of splits. 1. Bumper frames help the splitting algorithm mimic the splits that the previous / next node passes through. The valuable insight gained by mimicking splits on other nodes helps minimize the number of splits. 2. By using bumper frames to calculate the reshaping parameters, smoother transitions can be achieved at segment boundaries. The scene segmentation algorithm originally described in U.S. Patent No. 6,239,699 is described in the following subsections, but those skilled in the art will understand that the trim-path correction proposed here is independent of the scene segmentation algorithm used and is applicable regardless of the method for deriving the primary and secondary scenes. The discussion begins with segmenting parent scenes without considering multi-node assignments, and then the method is extended to scene segmentation for parent scenes distributed across two or more neighboring nodes.

[0055] Consider a parent scene P with M frames (M>1) ranging from the Qth index frame to the Q+M-1 frame in the mezzanine. FIG. 4A shows an exemplary process (400) for dividing a primary scene into subscenes, according to one embodiment. The goal of this process is to divide the primary scene into "temporally stable" or "temporally consistent" subscenes. All frames within each subscene are reshaped using the same scene-based reshaping function. Temporal stability thus allows for a reduction in the amount of reshaping metadata while maintaining video quality at a given bitrate.

[0056] The process 400 begins with an initialization stage 410, where, given input HDR and SDR frames (405) for a primary scene P, the HDR and SDR histograms h v and h s and individual forward reshaping functions (FLUTs)

number

number

number

[0057] The segmentation methods described herein are agnostic to how the frame-based reshaping functions are generated, and thus, in some embodiments, such reshaping functions may be generated directly from available HDR video using any of the known reshaping techniques, without relying on the availability of corresponding SDR video.

[0058] Scene FLUT

number

number

number

[0059] The scene FLUT and the generated histogram are used to determine the "DC" value χ for every frame in the scene P. jIf the height and width of the frame are H and W respectively, then its DC value is

number

[0060] In one embodiment, the DC difference of each frame from the previous frame is:

number

number

[0061] The maximum absolute value of the element-wise difference between each frame's FLUT and the previous frame's FLUT is also stored during the initialization stage and used as an additional set of thresholds for detecting smoothness violations.

number

[0062] Secondary Scene Cut C g is the sorted set of subscenes Ω P where g is the index into the collection. Frame index Q+M acts as an end-of-list marker and is not used as a scene cut. In one embodiment, the secondary scene cuts at initialization are:

number

[0063] In one embodiment, a violation subscene set Υ is used to store subscenes that violate the smoothness criterion. To start splitting a parent scene P, at initialization, Υ={P}. Only scenes or subscenes in the violation set will be split later. In summary, the initialization step in step 410:

number

[0064] In step 415, a violation set Υ and a sorted set of secondary scene cuts Ω are P Given as input, a new round of subscene partitioning begins: iterate through all subscenes in the violation set Y and determine how to partition them.

[0065] P g The frame range [C g ,C g+1 -1]. For the purpose of the division, the subscene FLUT

number

number

number

[0066] After splitting, subscene P g is split into two sub-scenes or secondary scenes, and the new split index is inserted into the secondary set at the correct position.

number

[0067] In step 420, the updated set Ω P A new subscene FLUT is calculated for every secondary scene in the set Ω. P But from C0 to C G It contains G+1 secondary scene cuts up to

number

number

[0068] Frame Range [C g ,C g+1 -1] spanning subscene P g Consider P for g∈[0,G-1]. g Regarding subscene FLUT, that is

number

number

number

[0069] In the current round of the segmentation process, let the DC values ​​be defined by λ. These DC values ​​are later used in step 425 to find threshold violations at sub-scene boundaries.

number

[0070] In step 425, temporal stability violations at the boundaries between sub-scenes are detected. g-1 ,C g -1} in the secondary scene P g-1 and {C g ,C g+1 -1} in the secondary scene P g About C g If any of those checks fail, then the subscene P g-1 and P g Both of the subscenes P are moved to the violation set Υ. g and P g+1 Regarding C g+1 A bounds check needs to be computed in The first frame of the segment is C0(Q) and the last frame of the segment is Q+M-1=C G In addition to -1, each sub-scene boundary C g The same checks apply.

[0071] Using equation (15), Ω P After iterating through all subscenes in, the updated DC value (λ j) is available for all frames in the primary scene P. These values ​​are used to perform the bounds violation check in steps 425 and 430. Cg is the index C g and a frame with index C g This is the difference of the DC value of the previous frame with -1.

number

[0072] Violation Check #1:

number

number

[0073] Violation Check #2:

number

number

number

[0074] Violation Check #3:

number

number

number

[0075] All violation checks are at subscene boundaries. If there is a violation, both subscenes are placed in a violation set. This ends the current segmentation round. In step 430, if the updated violation set is not empty, control returns to step 415 with the updated sets Ω and Y for the next segmentation round. Otherwise, if there are no boundary violations and the violation set is empty, the process ends and step 440 outputs the final secondary set of subscenes. In one embodiment, in step 425, if a secondary scene in Y is only one frame long, it cannot be further segmented and can be removed from the Y set. Alternatively, such single-frame scenes can be ignored in step 415.

[0076] In practice, a parent scene is split only if it is processed across two or more nodes. For example, a node may look for scene cuts in the left and right bumper sections. If no such scene cuts are detected, it can infer that the beginning or end of the segment is also processed by a neighboring node, and therefore one or more primary scenes need to be split.

[0077] Consider the scenario shown in Figure 3C, where a parent scene P is processed by two nodes. Each portion of the parent scene within a node is the primary scene for that node. One approach is to segment them independently using the segmentation approach described above. In that case, due to missing statistics, the scene cuts in the overlapping regions at nodes N and N+1 may not coincide with each other. The proposed segmentation algorithm performs much better in resolving temporal inconsistencies if there are good estimates of secondary scene boundaries within nearby subscenes on neighboring nodes.

[0078] In one embodiment, for the example of Figure 3C, two new synchronized sub-scene cuts (320) in the two primary scenes are introduced, one at each node. These synchronized cuts split the primary scene into two parts: 1. The first portion (e.g., between the first cut and the first synchronized cut at node N) is visible to the current node but not to the other node. As shown in FIG. 3C, in one embodiment, the first synchronized cut on node N is at position C L-1 - B, where B indicates the number of buffered frames, and C L-1 indicates the last frame of the segment. 2. The second portion (e.g., the tail bumper frames for node N and the early bumper frames for node N+1) is "visible" to both nodes. As shown in Figure 3C, in one embodiment, the second synchronization cut on node N+1 may be at position C0+B, where C0 denotes the first frame of the segment.

[0079] In one embodiment, these initial synchronous segmentations may be performed as part of stage 410, and the segmentation algorithm 400 can be applied to these primary scenes. The only minor change is in the initialization stage 410, where the set Ω for each node is P contains one additional synchronized scene cut (320). After that, no further splitting needs to be done, so we can jump directly to stage 420 after initialization. The algorithm then proceeds as normal.

[0080] Or, the original Omega P Given a set, if we detect that the primary scene is not entirely at the current node, this synchronization refinement is performed by using the rules mentioned above (e.g., for node N, if the primary scene does not end at node N, then the position C L-1 -B) to add a scene cut.

[0081] With these initial synchronization divisions, the sub-scene cuts computed by isolated nodes N and N+1 are expected to be reasonably aligned with each other. For node N, Ψ1 NLet denote the first list of scenes obtained after merging XML scene cuts and Auto-SCD scene cuts, as shown in Figure 3B. These scenes are called primary scenes. Scenes that lie on segment boundaries are split into secondary scenes or sub-scenes. Secondary scenes or sub-scenes generate additional scene cuts that are appended to the first list. Apart from these secondary scene cuts, the first frame f of the segment NL is also a scene cut. Since a node can only start encoding from the beginning of a segment, that frame is treated as a scene cut. Similarly, the last frame of a segment, f (N+1)L-1 is the end of the last scene in the segment. For example, consider node N with the following initial assignments: primary scenes 1, 2, and 3. Primary scene 1 may be subdivided into secondary scenes A, B, and C, primary scene 2 may remain unchanged, and primary scene 3 may be split into secondary scenes D, E, and F. Ψ l N and Ψ r N Let Ψ2 denote the set of scene cuts near the left and right segment boundaries for node N, respectively. Then, let Ψ2 denote the set of scene cuts near the left and right segment boundaries for node N, respectively. N can be expressed mathematically as follows:

number

[0082] For scenes longer than the segment length, there may be a single set of secondary scene cuts, rather than separate sets for left or right. k is a list Ψ2 N Let represent the starting frame index for the kth scene in the list. If there are K scenes in the list, then the elements in the list are

number

[0083] A second list of scenes Ψ2 N contains details about the primary and secondary scenes in the segment. The primary scene does not require additional data from pass 1. However, the secondary scene requires the following auxiliary data from pass 1: 1. The number of overlapping frames between the left and right sides for each secondary scene. 2. Trim path correction data, which may include one or more of the following: a. Average HDR luma value for each subscene b. Average SDR luma values ​​for each subscene c. Merging points in low intensity regions (dark areas) for each subscene d. Merging points in high intensity regions (highlights) for each subscene Collectively, the trim path correction data may be referred to as trim path parameters, trim parameters, trim path correction parameters, or trim correction parameters.

[0084] The term "merge point" refers to a luminance value in the SDR or HDR domain that serves as a boundary point for modifying the original forward reshaping function (FLUT) to compensate for potential effects of trim pass operations. For example,

number

[0085] The proposed architecture has two main types of scenes: primary and secondary scenes. Pass 2 processes every scene and generates the same set of metadata parameters for every single frame in that scene. In the forward phase of Pass 2, reshaping parameters are calculated from all frames within the statistics collection window of that scene. For the primary scene, the statistics collection window includes all frames within the primary scene. Frames outside the primary scene are not referenced. Conversely, for a secondary scene, the statistics collection window includes all frames in that secondary scene plus some overlapping frames from the previous or next subscene. These additional frames are called overlapping frames.

[0086] In principle, a primary scene has no overlap with any neighboring scenes, and a secondary scene is only allowed to have overlap with neighboring secondary scenes. That is, overlapping frames for a subscene can never come from neighboring primary scenes. The overlap parameter θ (see Eq. (13)) is set by the user and has a default value of 1. The backward phase in pass 2 does not use such overlap for primary or secondary scenes.

[0087] As mentioned above, colorists often introduce trim in the reference SDR (125) to meet the director's intent. Sometimes, trimming leads to clipping of highlights and / or crushing of low-intensity values. Reconstructing HDR from trim-affected SDR introduces undesirable artifacts in the reconstructed HDR domain. To reduce these artifacts, as discussed in U.S. Patent No. 5,999,499 and later sections of this specification, a trim correction algorithm requires (i) a reference display mapping (DM) curve calculated using the minimum, maximum, and average HDR and SDR luma values ​​in the content, and (ii) merge points in the low- and / or high-intensity regions.

[0088] As used herein, the term "display mapping" (DM) curve refers to a function that maps pixel values ​​in an image frame (at a first dynamic range) to pixel values ​​on a target display (at a second, different dynamic range). For example, without limitation, as discussed in U.S. Patent Application Publication No. 2009 / 0129990, given minimum, maximum, and average luminance values ​​in a video frame and corresponding minimum, average, and maximum luminance values ​​on a target display, a tone mapping curve can be determined that maps input HDR values ​​to corresponding HDR or SDR values. In an embodiment, such a tone mapping or display mapping (DM) curve may be used to better determine the range of HDR values ​​over which the original forward reshaping function needs to be adjusted so that trim pass artifacts are reduced.

[0089] For trimming in secondary scenes, the minimum and maximum HDR and SDR luma values ​​used in the reference DM curve are fixed globally for the entire video sequence. This means that there is no need to calculate the minimum and maximum HDR or SDR luma values ​​for each secondary scene. We define the average HDR and SDR luma values ​​as v avg and s avgLet the merge points in the low-intensity and high-intensity SDR regions be s l m and s h m As an example, consider the scenario in Figure 3D where segment P is processed by two nodes. After the insertion of the two synchronized scene cuts, node N processes subscenes A, B, and C, and node N+1 processes subscenes D, E, and F. Nodes N and N+1 have access to the same set of frames in the trim parameter window L (345) used to collect the trim parameters. For each frame index j in L, the luma mean value v avg Y,j and two SDR merge points s l m,j , s h m,j (Later, scene-based HDR merging points v l m , v h m These averages are then averaged to obtain the output trim parameter set for L, as shown below.

number

number

[0090] In one embodiment, the trim parameter window may be defined as follows: If you share the parent scene with subsequent neighboring nodes: Frame C sr =C L-1 -B to frame C sr Frames up to the end of the bumper frames at the end of the node segment, including B, where B represents the number of bumper frames and C L-1 represents the last frame of the segment. If you share the parent scene with the previous neighbor: Frame C from the beginning of the bumper frames before the start of the node segment sl =C0+B, but C sl where C0 represents the first frame of the segment.

[0091] For longer parent scenes that span multiple nodes, the trim parameters are copied or interpolated. Figure 3E shows an example of a long parent scene being processed by three computing nodes. avg Y,L1 , s l m,L1 , s h m,L1 Let v be the set of trim parameters in L1 (350). Every subscene in node N-1 gets a copy of these parameters. Let v be the trim parameters in L2 (355). avg Y,L2 , s l m,L2 , s h m,L2 Then all subscenes in node N+1 use the L2 set of values ​​for trim correction.

[0092] For node N, the trim value for any subscene is calculated as follows: Let ρ represent any of the three parameters in the trim set. Let a0 represent the first frame (351) after the end of the starting trim parameter window 350 at node N. Frame a0 and all frames before a0 use trim set L1. a n Let (353) denote the first frame of the tail trim window 355 at node N. Frame a n and a n All frames starting from a use trim set L2. All other frames between these two trim parameter windows (i.e., frame index j∈[a0,a n -1]), for example, for a frame in subscene A, the trim parameters are determined by linear interpolation from trim sets L1 and L2 to anchor frames a0 and a n For example, in one embodiment, for the jth frame, ρ j The trim parameters may be calculated as follows:

number

number

[0093] ​The same formula applies for any parameter in a trim set. Trim parameter interpolation across subscenes is calculated using the trim parameters in L1 (i.e., ρ) computed at both ends of the node. L1 ) and the trim parameter in L2 (i.e., ρ L2 ) without interpolation, different trim parameters at the edges of the segments would cause the front reshape parameters to change significantly from one subscene to the next, which could cause flickering or sudden brightness changes.

[0094] Although equation (24) applies a simple linear interpolation, alternative interpolation methods known in the art may be used, such as quadratic, bilinear, polynomial, spline, etc.

[0095] FIG. 4B outlines an exemplary process (450) for generating updated scene-based correction trim path parameters, according to one embodiment. As shown in FIG. 4B, the process begins in step 455 by accessing the original trim path correction parameters (e.g., average SDR and HDR values ​​and merge points for each frame). If the primary scene is entirely performed at the current node, no action is required, and the process proceeds to forward reshaping (225), where the forward reshape function may be trim path corrected as described below (see process 470). If the node detects that the scene is shared with a single neighboring node (e.g., a predecessor or successor node), in step 457, the node uses its synchronization scene cut to define a trim path parameter window (e.g., 345) and generate a new set L of trim path correction parameters (see, e.g., equations (22) and (23)). The new set is used to generate scene-based forward reshape functions (225) for all secondary scenes at the current node (step 459).

[0096] If the node detects that one or more scenes are shared with both its neighboring nodes (e.g., predecessor and successor nodes), then in step 462, the node uses the two synchronized scene cuts to define two trim path parameter windows (e.g., 350 and 355) and determine two new sets of updated trim path correction parameters (L1 and L2) (see, e.g., (22) and (23)). Using L1 and L2, in step 464, the node generates updated trim path parameters for all frames between the two trim path parameter windows by interpolating values ​​between L1 and L2 according to their distance. Then (step 466), the following assignments are made: For all secondary scenes prior to the first sync frame, use the L1 set of trim correction parameters. For all secondary scenes after the second sync frame, use the L2 set of trim correction parameters. For each secondary scene between the first and second synchronized frames, a new scene-based trim correction set (e.g., ρ) is calculated by averaging the corresponding interpolated correction trim parameters for all frames in that secondary scene. A ) (see, for example, equation (25)).

[0097] Pass 2: Scene-based encoding As shown in Figure 2, every node is assigned a segment to be encoded into the bitstream. Pass 1, described above, generates a list of primary and secondary scenes that minimizes the bitrate of reshaping-related metadata while maintaining temporal consistency across neighboring nodes. Pass 2 uses this list to encode the scenes within the segment and generate a reshaped SDR-encoded bitstream. As shown in Figure 3A, bumper frames (309) are not encoded and are used to collect statistics in the forward reshaping pass to maintain temporal stability of secondary scenes.

[0098] As shown in Figure 2, pass 2 includes stages 225, 230, and 235. An alternative view of the same processing pipeline, at the scene or subscene level, is also shown in Figure 5. For forward reshaping, the primary and secondary scenes are processed in a similar manner, with one major difference: the primary scene has no overlap in forward reshaping, while the secondary scene has some overlap with neighboring subscenes. For backward reshaping, the process is exactly the same for the primary and secondary scenes. There is no overlap in the backward reshaping phase. A combined bitstream (240) consisting of reshaping metadata and the compressed base layer is generated as output. Each block is described in more detail below.

[0099] Given a segment to scene list (222), Figure 5 shows an example architecture for scene-based encoding at each node in the cloud. Without limitation, block 225 of Figure 2 can be split as shown using block 505 and block 132. Let S be the starting frame index for the kth scene. k Therefore, given a scene k, a node is a node in frame S k , S k +1, S k +2, …, and S k+1 -1 must be processed. The reference HDR frame (504) and the corresponding SDR frame (502) for the scene may be stored in corresponding SDR and HDR scene buffers (not shown). As discussed above, bumper frames are only used to generate statistical data about the secondary scene and are ignored when processing the primary scene.

[0100] From Figure 5, in step 505, the input SDR and HDR frames are used to generate a scene-based forward reshaping function. The parameters of such a function are used for the entire scene (rather than updated frame by frame), thus reducing the overhead for metadata 152. Next, in step 132, the forward reshaping is applied to the HDR scene (504) to generate a reshaped base layer 229, which is encoded by the compression unit (230) to generate the coded bitstream 144. Finally, in step 235, the reshaped SDR data 229 and the original HDR data (504) are used to generate parameters 152 for a backward reshaping function, which are to be transmitted together to a downstream decoder. These steps are described in more detail below. Without limitation, the steps are described in the context of what is called a three-dimensional mapping table (3DMT) representation. Here, to simplify operation, each frame is represented as a three-dimensional mapping table, each color component (e.g., Y, Cb, or Cr) is subdivided into "bins," and instead of using explicit pixel values ​​to represent the image, pixel averages within each bin are used. Details of the 3DMT formulation can be found in U.S. Patent No. 5,649,999.

[0101] The scene-based generation of the forward reshaping function (505) consists of two levels of operation. First, statistics are collected for each frame. For example, for luma, SDR(h j s (b)) and HDR(h j v (b)) Compute histograms for both frames and store them in the frame buffer for the jth frame, where b is the bin index. After generating the 3DMT representation for each frame, generate an "a / B" matrix representation, which can be expressed as follows:

number

[0102] Given the statistics of each frame in the current scene, a scene-level algorithm can be applied to calculate the optimal forward reshaping coefficients. For example, for luma, SDR(h s (b)) and HDR data (h v The scene-based histogram for (b) can be generated by summing or averaging the frame-based histograms. For example, in one embodiment:

number

[0103] Having both scene-level histograms, in one embodiment, by way of example and without limitation, cumulative density function (CDF) matching (Patent Documents 4-5) can be applied to generate a forward mapping function (FLUT) from HDR to SDR. For example,

number

[0104] While frame-based FLUTs are typically designed based on the minimum, average, and maximum HDR and SDR luminance values ​​of a frame, reference DM curves are constructed slightly differently. For primary scenes, the reference DM curve uses the minimum and maximum HDR and SDR luminance values ​​for all frames in the scene. For secondary scenes, the reference DM curve uses global minimum and maximum values. These global minimum and global maximum HDR and SDR values ​​may define a wider dynamic range than the corresponding frame-based values. In some embodiments, these values ​​may be preselected to represent the normal range of values ​​allowed during data transmission. For example, in some embodiments, they may be defined based on the SMPTE range, so for 16-bit HDR data and 8-bit SDR data, the DM curve maps HDR codewords in [4096,60160] to SDR codewords in [16,235]. Alternatively, the entire range of possible values ​​may be applied. This is for example by mapping the full 16-bit HDR range [0,65535] to the full 10-bit SDR range [0,1023].

[0105] For primary scenes, SDR and HDR scene histograms are used to measure clipping-related distortions within the scene. For example, if the value of the normalized SDR histogram (see, e.g., Eq. (55)) is below a certain threshold (e.g., 0.20 for 8-bit SDR data and 0.05 for 10-bit SDR data) and / or there is no "variance peak" (see Eqs. (56)-(58)), then no trim correction is required. That is,

number

number

number

number

number

number

number

number

[0106] This bounded differential DM curve is denoted as a range-constrained differential DM curve. Accumulating the elements of this range-constrained differential DM curve (corresponding to a simple integral operation) gives the range-constrained DM curve:

number

[0107] For low intensity trim correction, the original FLUT and DM curves are min Y In high intensity trim correction, on the other hand, the curves are combined from the high merge point onwards. The merge point can be defined as a FLUT index value that marks the beginning / end of the FLUT and DM merging process, depending on the area of ​​the trim correction. There can be one merge point in the low intensity region and / or one merge point for the high intensity region.

[0108] As described in more detail in U.S. Patent No. 5,649,999 and in a later section (see "Calculating Merge Points for Trim-Path Correction"), we start with initial estimates of merge points for the SDR codeword range. We then calculate these initial estimates as s for the low-intensity and high-intensity regions, respectively. l m,1 and s h m,1 FLUT curve

number

number

number

number

[0109] During merging, to avoid clipping, the differential FLUT curve is replaced by a range-constrained differential DM curve in low or high intensity regions. To keep the SDR-mapped range of the resulting hybrid curve the same as that of the original FLUT curve, the range-constrained differential DM curve should be appropriately scaled as shown. Low intensity trim correction:

number

number

number

number

number

[0110] An example of the process is shown in FIG. 6A. Plots 610 and 605 show an example of the original FLUT and the reference DM curve, respectively. Plot 615 provides an example of a scaled and range-constrained DM curve. HDR values ​​620 show an example highlight merge point used for trim-path correction. Next, to generate a trim-path corrected FLUT, the mapping function in the original FLUT (610) after the merge point (620) is replaced based on the corresponding portion of the scaled DM curve. A similar approach can be applied to replace the original FLUT below the merge point for dark areas (not shown).

[0111] For the secondary scene, the process starts with the updated trim parameters previously generated (see process 450). As discussed previously (see process 450 in Figure 4B), a new set of trim correction parameters is determined for all subscenes using only the common frames in the trim parameter window of neighboring nodes. This set of trim correction parameters is read from the auxiliary data (212) and used during trim correction for that subscene.

[0112] The SDR merge points read from the datablock for the current subscene are now s for the low intensity range and s for the high intensity range, respectively. l m and s h m Let the HDR equivalent of these merge points v l m and v h m can be determined using the inverse mapping. These merge points can then be used to construct the hybrid FLUT and DM curves.

number

[0113] In one embodiment, the reference DM curve is calculated over the maximum possible range [0,2 BH-1], HDR min and HDR max is generated based on the global minimum and maximum possible HDR luma codeword values ​​(e.g., as defined by a constrained SMPTE range), denoted as B H is the bit depth of the HDR codeword, where the average value v avg Y is obtained from the trim correction parameter set for that subscene.

number

[0114] For temporal stability, it is recommended to use similar DM curves for all subscenes. Choosing fixed min / max values ​​allows the DM curve to be consistent throughout the sequence, while also covering the entire valid codeword range. These min / max values ​​are configurable and should be set to the lowest / highest possible codeword values ​​possible in the sequence.

[0115] Next, the differential DM curve

number

number

number

[0116] For trim correction in low intensity regions, the DM curve is scaled similarly to the legacy method. In one embodiment, scaling is preferred in low intensity regions to avoid introducing a constant offset in the FLUT. In high intensity regions, the FLUT curve is replaced by the original DM curve without scaling. Low intensity trim correction:

number

number

number

[0117] Finally, the updated or trimmed differential FLUT curve

number

number

[0118] An example based on the curve in Fig. 6A and high-intensity-only merge points is shown in Fig. 6B. The reason for using the original DM curve rather than a scaled version of it is to avoid modifying the FLUT when there are small changes in the SDR range, as well as to maintain better temporal consistency between successive sub-scenes that may be processed by different nodes.

[0119] 4C provides an example process (470) of these stages according to one embodiment. As shown in FIG. 4C, in stage 472, the process receives the original FLUT to be trim-corrected, a set of original merge points used for trim correction (e.g., see equation (23) or equation (25)), and scene-based minimum, average, and maximum HDR luminance values ​​(e.g., v min Y , v avg Y , v max Y ) next, the node identifies whether the scene to be processed is primary or secondary. For primary scenes, if trim-path correction is required, a reference DM curve is generated in step 474 (see equation (29)). Using the differential FLUT and DM curve (step 476), the node generates a differential DM curve and a corresponding range-constrained DM curve in step 478 (see equations (30)-(32)), such that the mapped SDR range of the range-constrained DM curve is the same as that of the original FLUT. Given the range-constrained DM curve, the node generates a second, more accurate set of merge points in step 480 (see equation (35)). The original differential FLUT curve is then edited in step 482 using the new merge points, using segments from the range-constrained differential DM curve generated in step 478 and scaled as needed to match the original SDR range of the FLUT curve (see equations (36)-(38)). Finally, in step 485, a trim-corrected FLUT is generated from the edited difference FLUT.

[0120] For secondary scenes, in step 490, a reference DM is generated using the scene-based average luminance value but without the global minimum and maximum HDR luminance values ​​(e.g., HDR min , v avg Y , HDR max). As in the primary scene case, in step 492, a difference FLUT and a reference DM curve are constructed. Unlike in the primary scene case, the original merge points are not refined. Step 494 is similar to step 482; using the original merge points, the original difference FLUT curve is edited, this time using segments from the difference DM curve generated in step 492, but with a slightly different treatment of the low and high merge points. For low intensity merge points, the difference reference DM curve is scaled to match the SDR dynamic range (equation (43)). For high intensity merge points, the difference reference DM curve is used without additional scaling (see equation (44)). Finally, in step 485, a trim-corrected FLUT is generated from the edited difference FLUT.

[0121] For chroma (e.g., ch=Cb or ch=Cr), we again average the a / B frame-based representation in (26) to obtain

number

number

number

[0122] The generation of the scene-based backward reshaping function (152) also involves both frame-level and scene-level operations. Since the luma mapping function is a single-channel predictor, we simply invert the forward reshaping function to obtain the backward reshaping function. For chroma, we use the reshaped SDR data (229) and the original HDR data (504) to form a 3DMT representation, and then generate a new frame-based a / B representation.

number

[0123] At the scene level, for luma, the histogram-weighted BLUT construction of U.S. Patent No. 5,999,499 may be applied to generate a backward luma reshaping function. For chroma, the frame-based a / B representations can again be averaged to compute a scene-based a / B representation.

number

number

number

[0124] Trim-related clipping may exist in low- or high-intensity regions. To correct these clipping artifacts, as discussed previously, a trim-affected forward reshaping curve (FLUT) is merged with a trim-free display mapping (DM) curve to construct a hybrid curve that avoids clipping. Two key parameters for generating this hybrid curve are the merge points for the FLUT and DM curves. First, we define these merge points (e.g., s ), one for the low-intensity region and one for the high-intensity region. l m and s h m ) is calculated in the SDR codeword range. Then, as discussed above and in U.S. Pat. No. 5,519,239, an equivalent merge point (e.g., v l m and v h m) are derived using one or more transforms. These HDR merge points are ultimately used to construct the hybrid (trim-corrected) FLUT. This section describes the merge points s in the SDR codeword range. l m and s h m 1 shows an exemplary method for calculating

[0125] The SDR and HDR histograms and original FLUT required for trim path correction may be frame-based, scene-based, or window-based (e.g., calculated within a narrow window of the frame), but the merge points are evaluated in the same way.

number

[0126] The first step is h sm s The SDR luma histogram h is calculated using a moving average filtered version of the SDR histogram, expressed as s In one embodiment, element-by-element differences between the original and smoothed SDR histograms may be used to estimate the locations of these peaks.

number

number

[0127] Normalized values ​​higher than a threshold (e.g., 0.20 for 8-bit SDR, 0.05 for 10-bit SDR)

number

number

[0128] Due to the HDR-to-SDR mapping, multiple HDR codewords map to a single SDR codeword because the bit depth of HDR codewords is higher than that of SDR. Clipping exacerbates this many-to-one mapping because in the clipped domain, many more codewords map to the same SDR codeword. For any SDR codeword c, we can find the HDR codewords assigned to it and calculate the variance of those codewords.

number

number

[0129] Cross-domain variance σ c 2 is evaluated at all SDR codeword values ​​and normalized by dividing by the sum of all SDR codeword variances,

number

number

number

[0130] The peaks in the SDR histogram and the cross-domain scatter array are collected into an ensemble and separated into low-intensity and high-intensity groups. j represents any peak in the ensemble, and τ loc j <2 Bs-1 If so, it is a low intensity peak, otherwise the peak belongs to the high intensity region. For the low intensity region, the merge point s l m,j is estimated for the jth peak as follows:

number

[0131] Then, the low merge point s in the SDR range l m is 2 Bs-1 This is the maximum of all these merge points bounded by

number

number

[0132] Exemplary Computer System Implementation Embodiments of the present invention may be implemented using computer systems, systems comprised of electronic circuits and components, integrated circuit (IC) devices, such as microcontrollers, field programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or apparatuses including one or more of such systems, devices, or components. The computers and / or ICs may execute, control, or perform instructions related to trim-path correction in cloud-based video encoding of HDR video as described herein. The computers and / or ICs may calculate any of a variety of parameters or values ​​related to trim-path correction in cloud-based video encoding of HDR video as described herein. Image and video dynamic range enhancement embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0133] Certain implementations of the present invention include computer processors executing software instructions that cause the processors to perform the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. can implement the method for trim path correction in cloud-based video encoding of HDR video described above by executing software instructions in program memory accessible to the processor. The present invention may also be provided in the form of a program product. A program product may include any non-transitory, tangible medium bearing a set of computer-readable signals containing instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. A program product according to the present invention may be in any of a wide variety of non-transitory, tangible forms. A program product may include physical media, such as magnetic data storage media including floppy diskettes and hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, and electronic data storage media including flash RAM. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0134] Where a component (e.g., a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, reference to that component (including reference to "means") should be interpreted as including any component that performs the function of the described component (e.g., is functionally equivalent) as an equivalent of that component, including components that are not structurally equivalent to the disclosed structures that perform that function in the illustrated exemplary embodiments of the present invention.

[0135] Equivalents, Extensions, Substitutes and Others Thus, exemplary embodiments relating to trim-path correction and node-based processing in cloud-based video encoding of HDR video have been described. In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details that may vary depending on the implementation. Accordingly, the sole and exclusive indication of what is and what is intended by the applicant to be the invention is the set of claims issuing from this application in the particular form in which the claims are allowed, including any subsequent amendments. Definitions, if any, expressly set forth herein for terms contained in such claims will control the meaning of such terms as used in such claims. Accordingly, no limitations, elements, properties, features, advantages, or attributes not expressly recited in a claim should in any way limit the scope of such claims. Accordingly, the specification and drawings are to be regarded in an illustrative and not restrictive sense.

[0136] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE). [EEE1] 1. A method for trim path correction in a computing node for a video segment to be encoded, the method comprising: receiving, at a current computing node, a first video sequence including high dynamic range video frames; generating (455) for each video frame in the first video sequence a first forward reshaping function and frame-based trim path correction parameters, the forward reshaping function mapping frame pixels from the high dynamic range to a second dynamic range lower than the high dynamic range; For a subscene in the first video sequence that is part of a parent scene processed by the current computing node and one neighboring computing node, determining (457) a trim path parameter window including frames and bumper frames in the first video sequence, the bumper frames including video frames processed by the neighboring computing nodes; calculating (459) scene-based trim path correction parameters based on the frame-based trim path correction parameters for frames in the trim path parameter window; generating a scene-based forward reshaping function based on the first forward reshaping function of frames in the subscene; generating, for the subscene, an output forward reshape function based on the scene-based trim path correction parameters and the scene-based forward reshape function; method. [EEE2] For said subscene in said first video sequence: applying the output forward reshaping function to frames in the subscene to generate an output video subscene in the second dynamic range; and compressing the output video sub-scene to generate an encoded sub-scene in the second dynamic range. Method described in EEE1. [EEE3] For a high dynamic range (HDR) frame in the first video sequence, frame-based trim-path correction parameters are: the average luminance value of said HDR frame; a first merging point in a low luminance region of a standard dynamic range (SDR) frame depicting the same scene as the HDR frame but in the second dynamic range; and a second merge point in a high luminance region of the SDR frame; The method of any one of claims 1 to 3, including one or more of the following: [EEE4] wherein the scene-based trim path correction parameters are calculated by:

number

number

number

Claims

1. 1. A method for trim path correction for a video segment to be encoded in a computing node, the method comprising: receiving, at a current computing node, a first video sequence including high dynamic range video frames; generating (455) for each video frame in the first video sequence a first forward reshaping function and frame-based trim path correction parameters, the forward reshaping function mapping frame pixels from the high dynamic range to a second dynamic range lower than the high dynamic range; For a sub-scene in the first video sequence that is part of a parent scene processed by the current computing node and one neighboring computing node, determining (457) a trim path parameter window including frames and bumper frames in the first video sequence, the bumper frames including video frames processed by the neighboring computing nodes; calculating (459) scene-based trim path correction parameters based on the frame-based trim path correction parameters of frames in the trim path parameter window; generating a scene-based forward reshaping function based on the first forward reshaping function of frames in the subscene; generating, for the subscene, an output forward reshape function based on the scene-based trim path correction parameters and the scene-based forward reshape function; method.

2. 1. A method for trim path correction for a video segment to be encoded in a computing node, the method comprising: receiving, at a current computing node, a first video sequence that is part of the video segment, the first video sequence including high dynamic range video frames; generating (455) for each video frame in the first video sequence a first forward reshaping function and frame-based trim path compensation parameters, the forward reshaping function mapping frame pixels from the high dynamic range to a second dynamic range lower than the high dynamic range, the trim path compensation parameters representing parameters for modifying given trim path parameters for one or more of trimming, rotating, cropping, flipping, and adjusting brightness, color, or saturation of the first video sequence, and the frame-based trim path compensation parameters being generated for each frame based on the corresponding frame for which the parameters are generated; For a sub-scene in the first video sequence that is part of a parent scene processed by a current computing node and one neighboring computing node, the sub-scene and the first video sequence currently being processed by computing code: determining (457) a trim path parameter window including frames and bumper frames in the first video sequence, the bumper frames including video frames processed by the neighboring computing nodes; calculating (459) scene-based trim path correction parameters based on the frame-based trim path correction parameters of frames in the trim path parameter window; generating a scene-based forward reshaping function based on the first forward reshaping function of frames in the subscene; generating, for the subscene, an output forward reshape function based on the scene-based trim path correction parameters and the scene-based forward reshape function; method.

3. For the subscene in the first video sequence: applying the output forward reshaping function to frames in the subscene to generate an output video subscene in the second dynamic range; and compressing the output video sub-scene to generate an encoded sub-scene in the second dynamic range. The method according to claim 1 or 2.

4. For a high dynamic range (HDR) frame in the first video sequence, frame-based trim path correction parameters are: an average luminance value of said HDR frame; a first merging point in a low luminance region of a standard dynamic range (SDR) frame depicting the same scene as the HDR frame but in the second dynamic range; and a second merge point in a high luminance region of the SDR frame; 4. The method of claim 1, further comprising one or more of:

5. Calculating the scene-based trim path correction parameters includes: [Number 93] where B represents the number of bumper frames, L represents the set of frames in the trim path parameter window, and v avg j represents the average luminance value of the jth HDR frame, and s l m,j represents the first merging point of the j-th HDR frame, and s h m,j represents the second merging point of the jth HDR frame, and v avg L , s l m,L and s h m,L represents the scene-based trim path correction parameters, The method of claim 4.

6. The method of claim 5 , further comprising assigning the scene-based trim path correction parameters to all sub-scenes being processed on a current computing node.

7. Generating an output forward reshaping function for an HDR frame in the subscene comprises: generating a reference tone mapping function that maps pixel values ​​from the high dynamic range to pixel values ​​in the second dynamic range based on at least an average luminance value and global minimum and maximum HDR and SDR luminance values ​​in the sub-scene; calculating a high strength merge point based on the second merge point; Starting from the high-strength merge point, by replacing the scene-based forward reshaping function based on a corresponding section in the reference tone mapping function; generating the output forward reshaping function; 7. The method according to any one of claims 4 to 6.

8. Generating the output forward reshaping function: generating a differential tone mapping function based on the reference tone mapping function; generating a differential scene-based forward reshaping function based on the scene-based forward reshaping function; generating an updated differential scene-based forward reshaping function by replacing the differential scene-based forward reshaping function with a corresponding portion of the differential tone mapping function, starting from the high strength merge point; generating the output forward reshaping function by integrating the updated differential scene-based forward reshaping function; The method of claim 7.

9. Given the average HDR luminance value and the average SDR luminance value in the sub-scene, the reference tone mapping function is: the average HDR luminance value to the average SDR luminance value; a maximum global luminance in the high dynamic range to a maximum global luminance value in the second dynamic range; mapping a minimum global luminance value in the high dynamic range to a minimum global luminance value in the second dynamic range; 9. The method according to claim 7 or 8.

10. For a sub-scene in the first video sequence that is part of a parent scene processed by the current computing node and one neighboring computing node: If the neighboring computing node is ahead of the current computing node, The trim path parameter window is set to the first frame (C 0 ) all (B) bumper frames in front of frame C s = C 0 Including up to +B, If the neighboring computing node is behind the current computing node, the trim path parameter window includes Cs=CL-1-B and all subsequent frames, including all bumper frames after the last frame of the first video sequence; 10. The method according to any one of claims 1 to 9.

11. For a sub-scene in the first video sequence that is part of a parent scene being processed by the current computing node, a nearby predecessor node, and a nearby successor node, further: determining (462) a first trim path parameter window that includes a frame in the first video sequence and a bumper frame shared with the neighboring predecessor node; determining a second trim path parameter window that includes frames in the first video sequence and bumper frames shared with the proximate successor node; calculating a first set of scene-based trim path correction parameters based on the frame-based trim path correction parameters of frames in the first trim path parameter window; calculating a second set of scene-based trim path correction parameters based on the frame-based trim path correction parameters of frames in the second trim path parameter window; calculating interpolated frame-based trim path correction parameters based on the first and second sets of scene-based trim path correction parameters; assigning the first set of scene-based trim path correction parameters to the subscene if the subscene is within the first trim path parameter window; assigning the second set of scene-based trim path correction parameters to the subscene if the subscene is within the second trim path parameter window; and if the subscene is between the first and second trim path parameter windows, generating scene-based trim path correction parameters based on the interpolated frame-based trim path correction parameters for frames within the subscene, where L represents the number of frames in the video segment.

11. The method according to any one of claims 1 to 10.

12. Calculating the interpolated frame-based trim path correction parameters includes: [Number 94] Calculating where ρ L1 is a trim path correction parameter in the first set of scene-based trim path correction parameters, and ρ L2 are trim path correction parameters in the second set of scene-based trim path correction parameters, and ρ x is the last frame a of the first trim path parameter window 0 and the first frame a of the second trim path parameter window n represents the interpolated frame-based trim path correction parameters for the x-th frame between The method of claim 11.

13. A scene-based trim path correction parameter (ρ) is calculated based on the interpolated scene-based trim path correction parameter. A ) is generated by [Number 95] , including calculating a s represents the start frame of the sub-scene, and a e represents the start frame of the next subscene, The method of claim 12.

14. A non-transitory computer readable storage medium storing computer executable instructions for executing the method of any one of claims 1 to 13 on one or more processors.

15. Apparatus comprising a processor and configured to perform the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Segment-Based Reconstruction for High Dynamic Range Video Coding

    JP2020507958A

  • Integrated image reconstruction and video coding

    JP2020526942A

  • US.63/080255

  • US10,264,287、B

  • US10,397,576、H