Recursive segment-to-scene segmentation for cloud-based encoding of HDR video
The recursive segment-to-scene segmentation method in cloud-based HDR video coding minimizes metadata overhead and ensures quality by subdividing segments and using bumper frames for temporal consistency, addressing inefficiencies in existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2021-09-17
- Publication Date
- 2026-05-25
AI Technical Summary
Existing cloud-based video coding systems for HDR video face challenges in managing resizing-related metadata overhead, especially at low bitrates, due to frame-based or scene-based reshaping techniques that require frequent updates, leading to inefficient processing and quality degradation.
A recursive segment-to-scene segmentation method is employed in cloud-based coding, where video segments are subdivided into subsegments to minimize the need for reshaping function updates, using bumper frames for temporal consistency and automatic scene detection to optimize metadata transmission.
This approach reduces metadata overhead and maintains video quality by ensuring temporal consistency across nodes, even at low bitrates, thereby enhancing the efficiency and effectiveness of HDR video processing in cloud environments.
Smart Images

Figure 0007864698000074 
Figure 0007864698000075 
Figure 0007864698000076
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 080,255, filed on 18 September 2020, and European Patent Application No. 20196876.5, filed on 18 September 2020, which are incorporated herein by reference.
[0002] technology This disclosure generally relates to images. More specifically, one embodiment of the present invention relates to recursive video segment-to-scene segmentation for processing HDR video in a cloud-based coding architecture. [Background technology]
[0003] As used in this application, the term “dynamic range” (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensity in an image (e.g., luminance, luma), from the darkest gray (black) to the brightest white (highlight). In this sense, DR relates to “scene-referred” intensity. DR may also relate to the ability of a display device to render a range of intensity of a particular width sufficiently or approximately. In this sense, DR relates to “display-referred” intensity. Unless explicitly stated at any point in this description that a particular meaning has particular importance, the term should be presumed to be interchangeable in either sense.
[0004] As used here, the term High Dynamic Range (HDR) refers to a DR width spanning 14 to 15 orders of magnitude in the human visual system (HVS). In practice, a DR that humans can simultaneously perceive across a wide range of intensity may be somewhat truncated compared to HDR. As used here, the terms Visual Dynamic Range (VDR) or Enhanced Dynamic Range (EDR) may refer to the DR perceptible by the human visual system (HVS), including eye movements within a scene or image, taking into account some light adaptation changes across the scene or image as a whole, individually or interchangeably. As used here, VDR may refer to a DR spanning 5 to 6 orders of magnitude. Thus, although perhaps somewhat narrower than true scene-based HDR, VDR or EDR still represents a wide DR width and is sometimes referred to as HDR.
[0005] In practice, an image contains one or more color components (e.g., lumens Y and chromins Cb and Cr), and each color component is represented with n bits per pixel (e.g., n=8) of precision. For example, using gamma-luminance coding, n ≤ 8 (e.g., a 24-bit color JPEG image) may be considered a standard dynamic range image, and n ≥ 10 may be considered an enhanced dynamic range image. HDR images may be stored and delivered using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR file format developed by Industrial Light and Magic.
[0006] Currently, most consumer desktop displays have a brightness of 200 to 300 cd / m². 2 Alternatively, it supports luminance in nits. Most consumer HDTVs are in the 300 to 500 nits range, and newer models are 1000 nits (cd / m²). 2) reaches. Such conventional displays are typical of low dynamic range (LDR), also known as standard dynamic range (SDR) in contrast to HDR. Due to advancements in both supplemental equipment (e.g., cameras) and HDR displays (e.g., the PRM-4200 professional reference monitor from Dolby Laboratories), as the availability of HDR content increases, HDR content may be color graded and displayed on HDR displays that support a higher dynamic range (e.g., 1000 nits to 5000 nits or more).
[0007] Where used herein, the term “forward reshaping” refers to the process of mapping a digital image sample-to-sample or codeword-to-codeword from its original bit depth and original codeword distribution or representation (e.g., Gamma, PQ, HLG, etc.) to the same or a different bit depth and a different codeword distribution or representation. Reshaping allows for improved compressibility or improved image quality at a fixed bitrate. For example, reshaping may be applied to 10-bit or 12-bit PQ-encoded HDR video to improve coding efficiency in a 10-bit video coding architecture, but is not limited to this. At the receiver, after decompressing the received signal (which may or may not be reshaped), the receiver may apply a “reverse (or backward) reshaping function” to restore the signal to its original codeword distribution and / or achieve a higher dynamic range.
[0008] In many video distribution scenarios, HDR video is encoded in a multiprocessor environment, typically referred to as a “cloud computing server.” In such environments, trade-offs between computing ease, workload balancing across computing nodes, and video quality may force resizing-related metadata to be updated frame by frame, which can result in unacceptable overhead, especially when transmitting video at low bitrates. As understood by the inventors hereby, there is a need for improved techniques for segment-to-scene segmentation to minimize the overhead of resizing-related metadata in cloud-based environments. The approaches described in this section could have been pursued, but are not necessarily previously conceived or pursued. Therefore, unless otherwise noted, none of the approaches described in this section should be assumed to qualify as prior art simply because they are included in this section. Similarly, unless otherwise noted, no problem identified with respect to one or more approaches should be assumed to have been recognized in any prior art based on this section. Each of the following references is referenced in its entirety. [Prior art documents] [Patent Documents]
[0009] [Patent Document 1] U.S. Patent No. 10,575,028, H. Kadu et al., "Coding of high-dynamic range video using segment-based reshaping" [Patent Document 2] U.S. Patent No. 8,811,490, GM. Su et al., "Multiple color channel multiple regression predictor" [Patent Document 3] International Publication No. 2019 / 217751, Q. Song et al., PCT Patent Application No. PCT / US2019 / 031620, "High-fidelity full reference and high-efficiency reduced reference encoding in end-to-end single-layer backward compatible encoding pipeline," filed May 9, 2019. [Patent Document 4] U.S. Patent No. 10,264,287, B. Wen et al., "Inverse luma / chroma mappings with histogram transfer and approximation" [Patent Document 5] U.S. Patent No. 10,397,576, H. Kadu and GM. Su, "Reshaping curve optimization in HDR coding" [Patent Document 6] U.S. Provisional Patent Application No. 63 / 049,673, GM. Su et al., “Workload allocation and processing in cloud-based coding of HDR video,” filed July 9, 2020. [Brief explanation of the drawing]
[0010] Embodiments of the present invention are shown in the accompanying drawings as examples, not limiting, and similar reference numerals refer to similar elements.
[0011] [Figure 1A] This shows an example of a single-layer encoder for HDR data using a reshaping function based on conventional technology.
[0012] [Figure 1B] This shows an exemplary HDR decoder corresponding to the encoder in Figure 1A, using conventional technology.
[0013] [Figure 2] This document illustrates an exemplary architecture and processing pipeline for cloud-based encoding of HDR video according to one embodiment.
[0014] [Figure 3A] This example shows how to split a video input into segments and assign bumper frames to three nodes.
[0015] [Figure 3B] This example shows how to merge scene cuts to generate a list of primary scenes.
[0016] [Figure 3C] This shows an example of a primary scene split across two computing nodes.
[0017] [Figure 3D] An example of a statistical window used to derive a scene-based forward reshaping function according to one embodiment is shown.
[0018] [Figure 4] This illustrates an example of a sequential, iterative segment-to-scene segmentation process according to one embodiment.
[0019] [Figure 5] An exemplary encoder for scene-based encoding using reshaping, according to one embodiment of the present invention, is shown. [Modes for carrying out the invention]
[0020] This application describes a method for scene segmentation and node-based processing in cloud-based video coding of HDR video. The following description includes numerous specific details to provide a full understanding of the invention. However, it will be apparent that the invention may be carried out without these specific details. On the other hand, well-known structures and apparatus are not described in comprehensive detail to avoid unnecessarily obscuring, obscuring, or making the invention difficult to understand.
[0021] overview The exemplary embodiments described herein relate to cloud-based reshaping and coding of HDR images. In one embodiment, in a cloud-based system for encoding HDR video, the current node receives a first video sequence containing high dynamic range video frames. Then, one or more processors within the node: For each video frame in a first video sequence, the process involves generating a frame-based forward reshaping function, wherein the forward reshaping function maps the frame pixels from a high dynamic range to a second dynamic range lower than the high dynamic range; The first step is to generate a set of primary scenes for the first video sequence; A step of generating a second set of scenes for a first video sequence based on a set of primary scenes, secondary scenes derived from one or more primary scenes, and the frame-based forward reshaping function; The steps include: generating a scene-based forward reshaping function based on a second set of scenes; The process involves applying the aforementioned scene-based forward reshaping function to a first video sequence to generate a second dynamic range output video sequence; The steps include compressing the output video sequence to generate a second dynamic range encoded bitstream, where, given a primary scene, a list of secondary scenes for that primary scene is generated: The first step is to initialize the set of secondary scenes and the set of violation scenes based on the set of primary scenes; A step of generating one or more sets of smoothing thresholds based on the frame-based forward reshaping function; Until boundary violations cease: Divide each scene in the collection of violation scenes into two new subscenes; Using the empty set, generate an updated set of violation scenes; An updated set of secondary scenes is generated by adding the aforementioned new subscene to the aforementioned set of secondary scenes; Using one or more sets of smoothing thresholds, perform one or more boundary violation checks in the set of secondary scenes; If there is at least one boundary violation between two subscenes in the set of secondary scenes, add the two subscenes to the set of violating scenes, and continue subdividing the primary scene using the updated set of violating scenes and the updated set of secondary scenes; Otherwise, the process repeats, signaling that there are no boundary violations and outputting a list of secondary scenes.
[0022] Exemplary HDR coding system Figures 1A and 1B illustrate exemplary single-layer, backward-compatible codec frameworks using image reshaping based on prior art. More specifically, Figure 1A shows an exemplary encoder architecture that may be implemented in one or more computing processors within an upstream video encoder. Figure 1B shows an exemplary decoder architecture that may also be implemented in one or more computing processors within one or more downstream video decoders.
[0023] Under this framework, given reference HDR content (120) and corresponding reference SDR content (125) (i.e., content representing the same image as the HDR content but color-graded and represented in standard dynamic range), the reshaped HDR content (134) is encoded and transmitted as SDR content in a single layer of the encoded video signal (144) by an upstream encoding device implementing an encoder architecture. The received SDR content is received and decoded in a single layer of the video signal by a downstream decoding device implementing a decoder architecture. Back-reshaping metadata (152) is also encoded and transmitted in the video signal together with the reshaped content, so that an HDR display device can reconstruct the HDR content based on the (reshaped) SDR content and the back-reshaping metadata. Without loss of generality, in some embodiments, such as in non-backward compatible systems, the reshaped SDR content may not be viewable on its own but must be viewed in combination with a back-reshaping function that generates viewable SDR or HDR content. In other embodiments that support backward compatibility, legacy SDR decoders can also play back received SDR content without using a backward reshaping function.
[0024] As shown in Figure 1A, given an HDR image (120), an SDR image (125), and a target dynamic range, step 130 generates a forward reshaping function. Given the generated forward reshaping function, a forward reshaping mapping step (132) is applied to the HDR image (120) to generate a reshaped SDR base layer (134). A compression block (142) (an encoder implemented according to any known video coding algorithm such as AVC, HEVC, AV1, etc.) compresses / encodes the SDR image (134) in a single layer (144) of the video signal. Furthermore, a backward reshaping function generator (150) can generate a backward reshaping function that can be transmitted to the decoder as metadata (152). In some embodiments, the metadata (152) can represent the forward reshaping function (130), and thus it is left to the decoder to generate the backward reshaping function (not shown).
[0025] Examples of back-reshaping metadata that represent / specify the optimal back-reshaping function are not limited to these, but may include any of the following: inverse tone mapping functions, inverse lumar mapping functions, inverse chroma mapping functions, lookup tables (LUTs), polynomials, inverse display control coefficients / parameters, etc. In various embodiments, the lumar back-reshaping function and the chroma back-reshaping function may be derived / optimized congruently or separately, and may be derived using, for example without limitation, a variety of techniques as described below in this disclosure.
[0026] The back-reshaping metadata (152) generated by the back-reshaping function generator (150) based on the reshaped SDR image (134) and the target HDR image (120) may be multiplexed as part of the video signal 144, for example, as supplemental enhancement information (SEI) messaging.
[0027] In some embodiments, the back-reshaping metadata (152) is carried in the video signal as part of the overall image metadata, and the metadata is carried in the video signal separately from the single layer on which the SDR image is encoded. For example, the back-reshaping metadata (152) may be encoded in a component stream in the encoded bitstream, and the component stream may or may not be separate from the single layer (of the encoded bitstream) on which the SDR image (134) is encoded.
[0028] Thus, the back-reformatted metadata (152) can be generated or pre-generated on the encoder side to take advantage of the powerful computing resources and offline encoding flows available on the encoder side, including, but not limited to, content adaptive multiple passes, lookahead operations, inverse lumer mapping, inverse chroma mapping, CDF-based histogram approximation and / or transfer.
[0029] The encoder architecture in Figure 1A can be used to avoid directly encoding the target HDR image (120) into the encoded / compressed HDR image in the video signal; instead, by using back-reshaping metadata (152) in the video signal, a downstream decoder can back-reshape the SDR image (134) (encoded in the video signal) into a reconstructed image that is identical to, or closely approximates / optimally approximates, the reference HDR image (120).
[0030] In some embodiments, as shown in Figure 1B, a video signal encoded with a reshaped SDR image in a single layer (144) and backward reshaped metadata (152) as part of the overall image metadata are received as input on the decoder side of the codec framework. The decompression block (154) decompresses / decodes the compressed video data in a single layer (144) of the video signal into a decoded SDR image (156). Decompression 154 typically corresponds to the inverse of compression 142. The decoded SDR image (156) may be identical to the SDR image (134) under quantization errors in the compression block (142) and decompression block (154), which may be optimized for an SDR display device. In backward-compatible systems, the decoded SDR image (156) may be output in the output SDR video signal (e.g., via an HDMI® interface, via a video link, etc.) to be rendered on an SDR display device.
[0031] Optionally, alternatively, or additionally, in the same or another embodiment, the back-reshaping block 158 extracts back-(or forward) reshaping metadata (152) from the input video signal, constructs a back-reshaping function based on the reshaping metadata (152), and performs a back-reshaping operation on the decoded SDR image (156) based on the optimal back-reshaping function to produce a back-reshaped image (160) (or reconstructed HDR image). In some embodiments, the back-reshaped image represents a production-quality or near-production-quality HDR image that is identical to or closely approximates / optimally approximates the reference HDR image (120). The back-reshaped image (160) may be output in the output HDR video signal (e.g., via an HDMI® interface, via a video link, etc.) for rendering on an HDR display device.
[0032] In some embodiments, display management operations specific to the HDR display device may be performed on the back-reshaped image (160) as part of an HDR image rendering operation that renders the back-reshaped image (160) on the HDR display device.
[0033] Cloud-based coding Existing reshaping techniques may be frame-based, meaning that new reshaping metadata is sent for each new frame, or scene-based, meaning that new reshaping metadata is sent for each new scene. As used herein, the term “scene” in relation to a video sequence (sequence of frames / images) may refer to a series of consecutive frames within a video sequence that share similar luminance, color, and dynamic range characteristics. Scene-based methods work well in video workflow pipelines that have access to the entire scene. However, it is not uncommon for content providers to use cloud-based multiprocessing, in which case the video stream is divided into segments, and each segment is processed independently by a single compute node in the cloud. As used herein, the term “segment” refers to a series of consecutive frames in a video sequence. A segment may be part of a scene, or it may contain one or more scenes. Thus, the processing of a scene may be divided across multiple processors.
[0034] As discussed in Patent Document 1, in certain cloud-based applications, under certain quality constraints, segment-based processing may require the generation of re-formatted metadata for each frame, resulting in undesirable overhead. This can be problematic in applications with very low bitrates (e.g., less than 1 Mbit / s). Patent Document 6 proposes a solution to this problem using a two-stage architecture including a) an ordering stage implemented on a single computing node that assigns scenes to segments, and b) an encoding stage in which each node in the cloud encodes a sequence of segments. After the scenes are segmented, the proposed scene-to-segment assignment process includes one or more iterative steps using an initial random assignment of scenes to the nodes, followed by a refined assignment based on optimizing the assignment cost across all nodes. In such an implementation, the total length of the video processed at each node may vary across all nodes.
[0035] The embodiments presented herein provide an alternative solution. After the sequence is divided into segments so that each segment is processed by a separate node, each node subdivides each segment into subsegments (or scenes) such that the need to update the corresponding reshaping function for each subsegment is minimized, thus minimizing the overhead required to transmit reshaping-related metadata.
[0036] Figure 2 shows an exemplary architecture and processing pipeline for cloud-based encoding of HDR video according to one embodiment. Given a video source (202) for content delivery, typically called a mezzanine file, and a set of working nodes, each node (e.g., nodes 205-N) fetches the video pictures (or frames) and corresponding video metadata (207) to be processed (e.g., from an XML file) as follows:
[0037] In preprocessing step 210, the mezzanine input is divided into segments, each segment assigned to a different computing node (e.g., nodes 205-N). These segments are mutually exclusive, meaning they have no common frames. Each node also obtains a certain number of frames preceding the first frame in its segment and several frames following the last frame in its segment. These preceding and succeeding overlapping frames are called bumper frames and are used solely to maintain temporal consistency between the preceding and succeeding nodes, respectively. Bumper frames are not encoded by the node. Without loss of generality, in one embodiment, these video segments may all be of equal fixed length, except perhaps the segment assigned to the last node. As an example, Figure 3A shows a sample where a mezzanine (305) is distributed into three segments (307-1, 307-2, 307-3) along with bumper frames (e.g., 309), and these frames are assigned to different nodes. Without limitation, for a segment 30 seconds long and a bumper section 2 seconds long on each side, exemplary embodiments may include the following arrangement: • Segment section with 1800 frames and bumper section with 120 frames, 60fps • Segment section with 1500 frames and bumper section with 100 frames, 50fps • Segment section with 720 frames and bumper section with 48 frames, 24fps
[0038] After the preprocessing stage 210 is complete, each node gains access to the frame, and the two-pass approach continues. · In Path 1 (Stages 215, 220), a list of scenes within a segment (222) is generated. The scene cuts (209) extracted from the XML file and the scene cuts generated using the automatic scene cut detector (215) are combined at Stage 220 to obtain a first list of primary scenes. The primary scenes belonging to a parent scene encoded across multiple nodes may be subdivided into secondary scenes. To maintain temporal consistency in scenes distributed across multiple nodes, bumper frames and a new recursive scene segmentation algorithm are also provided. The segmentation generates additional scenes. These are added to the first list of scenes to obtain a second list. This second list of scenes (222) is passed to Path 2. · Path 2 (Stages 225, 230, 235) uses the list of scenes received from Path 1 to perform forward and backward reshaping for each scene within the segment. Forward reshaping (225) using the scene-based forward reshaping function (227) generates a reshaped SDR segment (229), and the backward reshaping unit (235) generates metadata parameters used by the decoder to reconstruct the HDR input. The reshaped SDR input (229) is compressed (230), and the compressed video data and the reshaped metadata are combined together to generate a compressed bitstream (240).
[0039] For the sake of simplicity of discussion, let L represent the number of frames within a segment and B represent the number of frames within each bumper section. Let the i-th frame in the mezzanine be represented as f i In one embodiment, the first node encodes frames f0 to f L-1 in a segment portion. This node has no left bumper, and its right bumper spans frames f L to f L+B-1 . The segment portion of node N processes frames f (N-1)L to f NL-1 . Here, f (N-1)L-B to f i(N-1)L-1 is the left bumper, and f NL to fNL+B-1 This is the right bumper. The last node does not have a right bumper section and may have fewer than L frames in the segment portion.
[0040] Given a node N, node N-1 is the left / previous neighbor node, and node N+1 is the right / next neighbor node. A reference to a node that is the left / previous node for N represents all nodes from 0 to N-1. Similarly, a reference to a node that is the right / next node for N represents all nodes from N+1 to the last node. Now, let's elaborate on the two paths mentioned above.
[0041] Path 1: Scene generation from segments The primary purpose of this pass is to generate a list of scenes within the segments assigned to the node. This process begins by detecting scene cuts within all frames assigned to the node, including the node segment and both bumper sections. Only the scene cuts within the segments are ultimately used for scene-based encoding by pass 2. However, scenes within the bumper sections are still useful for maintaining temporal consistency with neighboring nodes.
[0042] The scene cuts (209) specified by the colorist are read from the XML file (207). An automatic scene cut detector (215) may identify possible scene cut locations. These scene cuts from the colorist and the automatic detector are merged to obtain a first list of scenes known as primary scenes. Primary scenes on segment boundaries are divided using bumper frames and a novel scene division technique. Dividing primary scenes on segment boundaries creates additional scenes called secondary scenes or subscenes. The secondary scenes are added to the first list of scenes to obtain a second list. This list is then used by pass 2 for scene-based encoding. Apart from the list of scenes, pass 2 may also require auxiliary data (212) for forward reshaping of secondary scenes. Details of each step are described below.
[0043] Colorists and professional color graders typically treat each scene as a single unit. To achieve their goals (e.g., proper color grading, insertion of fade-ins and fade-outs), they must manually identify scene cuts within the sequence. This information is stored in an XML file and can be used for other purposes. Every node reads only significant scene cuts for its segment from the XML file. These scene cuts can be located within segment sections or bumper sections.
[0044] XML scene cuts are defined by the colorist, but are not always perfectly precise. For grading purposes, colorists sometimes introduce scene cuts in the middle of dissolve scenes or at the beginning of fade-in or fade-out portions of scenes. These scene cuts, when taken into account during the reshaping phase, can cause flickering in the reconstructed HDR video and should generally be avoided. For this reason, in some embodiments, an automatic scene-cut detector (Auto-SCD) 215 is also employed.
[0045] An automatic scene cut detector, or Auto-SCD, detects scene changes using changes in luminance levels in different sections of a sequence of video pictures. Any scene cut detector known in the art can be used as the automatic detector. In some embodiments, such an automatic detector is unaware that parts of the video dissolve, fade in, or fade out, yet can still correctly detect all true scene cuts.
[0046] A potential problem with the automatic detector is false positives. There may be changes in brightness within the scene due to camera panning, movement, occlusion, etc., and these brightness changes may also be detected as scene cuts by Auto-SCD. To discard these false positives, in one embodiment, scene cuts from the XML file and scene cuts from Auto-SCD are merged together in step 220. Those skilled in the art will understand that if no scene cuts are defined in the XML file, the output of the automatic scene detector may simply be used. Similarly, in other embodiments, it may strictly rely on scene cuts defined in the XML file. Alternatively, two or more scene cut detectors may be used, each detecting different attributes of interest, and then a primary scene may be defined based on a combination of all their results (e.g., their intersection or other set operations, e.g., their union, intersection, etc.).
[0047] Ψ XML N Let Ψ be a set of frame indices representing the scene start frame at node N, as reported in the XML file. Similarly, Ψ Auto-SCD N Let be a set of frame indices representing the scene start frame at node N, as reported by Auto-SCD. In one embodiment, merging scene cuts from these two sets is equivalent to taking the intersection of these two sets. Ψ1N =Ψ XML N ∩Ψ Auto-SCD N (1) Here, Ψ1 N = indicates the first list of scene cuts (or scenes) at node N. These scenes are also called primary scenes. Figure 3B shows an exemplary scenario. In this example, the XML file reports three scene cuts. Auto-SCD also reports three scene cuts; two are the same as the XML scene cuts, but the third is in a different location. Since only two of the six reported scene cuts are common, the node segment is divided into only three primary scenes (310-1, 310-2, 310-3) according to those two common scene cuts. In some embodiments, the scene cuts from XML and Auto-SCD may be recognized as the same, even if they are reported on different frames, as long as the scene cut indices between the two lists differ within a given small tolerance (±n frames, e.g., n is in [0,6]).
[0048] As shown in Figure 3B, primary scene 2 (310-2) is entirely at node N. Therefore, it can be processed entirely by node N in pass 2. Conversely, primary scenes 1 (310-1) and 3 (310-3) are on the segment boundary. Their parent scenes are distributed across multiple nodes and are processed independently by those nodes. To ensure a consistent appearance in boundary frames encoded by different nodes, primary scenes 1 and 3 require some special handling. Next, we consider several alternative scenarios.
[0049] As shown in Figure 3C, consider a simple scenario where P is a parent scene distributed across two nodes N and N+1. Nodes N and N+1 can only access parts of the parent scene. Assume these nodes process and encode their respective parts of the parent scene (without bumpers). Since the reshaping parameters are calculated for different sets of frames, the last frame in node N's segment, i.e., f (N+1)L-1 and the first frame in the segment of node N+1, i.e., f (N+1)L Reshaped SDR and reconstructed HDR images may appear visually different. Such visual differences typically manifest as flickering, blinking, or sudden changes in brightness. This problem is called temporal inconsistency across nodes. Part of the reason for the inconsistency in the above scenario is the lack of a common frame when calculating the reshaping parameters. Including bumper frames when generating these statistics, as shown in Figure 3C, provides smoother transitions across nodes. However, bumper sections may be relatively short compared to the parent scene and may not be long enough to guarantee temporal consistency. In some embodiments, to solve these problems, portions of the parent scene at nodes N and N+1 are split into secondary scenes or subscenes. Even if the reshaping statistics within a scene change significantly from one node to the next, those statistics do not change significantly from frame to frame. The secondary scenes use only statistics within a small neighborhood to evaluate the reshaping parameters. Therefore, these reshaping parameters do not change much from one subscene to the next. Thus, the splitting achieves temporal consistency. Note that nearby subscenes can also exist in the previous / next node.
[0050] It should be noted that splitting creates additional scenes, thus increasing the metadata bitrate. The challenge is to achieve temporal consistency using a minimum number of splits while keeping the metadata bitrate low. Bumper frames play a crucial role in achieving good visual quality while reducing the number of splits. 1. Bumper frames help the partitioning algorithm mimic the partitions that the previous / next node will pass through. The valuable insights gained by mimicking partitions on other nodes help minimize the number of partitions. 2. By using bumper frames to calculate reshaping parameters, smoother transitions can be achieved at segment boundaries. The scene partitioning algorithm is described in the following subsections. The discussion begins with partitioning a parent scene without considering multi-node assignments, and then extends this method to scene partitioning for parent scenes distributed across two or more neighboring nodes.
[0051] Consider a parent scene P having M frames (M>1) in the range from the Q-th index frame to Q+M-1 frames in a mezzanine. Figure 4 shows an exemplary process (400) for splitting a primary scene into subscenes according to one embodiment. The objective of this process is to split the primary scene into subscenes that are "temporally stable" or "temporally consistent". All frames within each subscene are reshaped using the same scene-based reshaping function. Thus, temporal stability allows for a reduction in the amount of reshaping metadata while maintaining video quality at a given bitrate.
[0052] Process 400 begins with initialization stage 410. Here, given input HDR and SDR frames (405) for the primary scene P, HDR and SDR histograms h v and h s and individual forward reshaping functions (FLUT)
number
number
number
[0053] The segmentation methods described herein do not concern how the frame-based reshaping function is generated. Therefore, in some embodiments, such a reshaping function may be generated directly from the available HDR video using any known reshaping technique, without depending on the availability of the corresponding SDR video.
[0054] Scene FLUT
number
number
number
[0055] The scene FLUT and the generated histogram show the "DC" value χ for each frame within scene P. jIt is used to predict. If the height and width of the frame are H and W respectively, its DC value
number
[0056] In one embodiment, the DC difference between each frame and the previous frame is,
number
number
[0057] The maximum absolute value of the element-wise difference between the FLUT of each frame and the FLUT of the previous frame is also stored during the initialization phase and used as an additional set of thresholds for detecting smoothness violations. Here, α and β are configurable parameters, with typical values of 2.0 and 0.60, respectively.
number
[0058] Secondary Scene Cut C g However, the sorted subscene set Ω P It is collected in the following location, where g is the index within the set. The frame index Q+M acts as the list end marker and is not used as a scene cut. In one embodiment, the secondary scene cut at initialization is as follows:
number
[0059] In one embodiment, a set of violating subscenes Υ is used to store subscenes that violate the smoothness criterion. To begin the splitting of the parent scene P, Υ = {P} is set during initialization. Only scenes or subscenes in the violating set are split later. In summary, the initialization stage in step 410 is:
number
[0060] In stage 415, the set of violations Υ and the sorted set of secondary scene cuts Ω are obtained. P Given as input, a new round of subscene splitting is initiated. It iterates sequentially through all subscenes in the violation set Υ to determine how to split them.
[0061] P g frame range [C g ,C g+1 This will be a subscene within the violation set spanning [-1]. For division, use subscene FLUT.
number
number
number
[0062] After splitting, subscene P g It is divided into two subscenes or secondary scenes, and the new division index is inserted into the secondary set at the correct position.
number
[0063] In stage 420, the updated set Ω P For all secondary scenes within the set Ω, a new subscene FLUT is calculated. P However, as shown in the following equation, C0 to C G It includes G+1 secondary scene cuts up to that point.
number
number
[0064] Frame range [C g ,C g+1 Subscene P spanning -1] g Consider P for g∈[0,G-1]. g Regarding the subscene FLUT, that is
number
number
number
[0065] In the current round of the partitioning process, the DC value is defined by λ. These DC values will later be used in stage 425 to find threshold violations at subscene boundaries.
number
[0066] In stage 425, a temporal stability violation is detected at the boundary between subscenes. For example, {C g-1 ,C g Secondary scene P in -1} g-1 and {C g ,C g+1 Secondary scene P in -1} g Regarding C g Boundary checks need to be calculated in this area. If any of those checks fail, subscene P g-1 and P g Both are moved to the infringement set Υ. Subscene P g and P g+1 Regarding C g+1 Boundary checks need to be calculated in this case. The first frame of the segment C0(Q) and the last frame of the segment Q+M-1=C G Aside from -1, each subscene boundary C g The same check applies in this case.
[0067] Using equation (15), Ω P After iterating through all subscenes sequentially, the updated DC value (λ j) becomes available for all frames in the primary scene P. These values are used to perform boundary violation checks in stages 425 and 430. DC difference Δ Cg is index C g A frame with index C g This is the difference in DC value from the previous frame, which has a value of -1.
number
[0068] Violation Check #1:
number
number
[0069] Violation check #2:
number
number
number
[0070] Violation Check #3:
number
number
number
[0071] All violation checks are performed at subscene boundaries. If a violation exists, both subscenes are placed in the violation set. This completes the current splitting round. In step 430, if the updated violation set is not empty, control returns to step 415 with the updated sets ΩP and Υ for the next splitting round. If, instead, there are no boundary violations and the violation set is empty, the process ends and step 440 outputs the final quadratic set of subscenes. In one embodiment, in step 425, if a quadratic scene in Υ is only one frame long, it can be removed from the Υ set because it cannot be split further. Alternatively, such a single-frame scene can be ignored in step 415.
[0072] In practice, a parent scene is only split if it is processed across two or more nodes. For example, a node may look for scene cuts in its left and right bumper sections. If no such scene cuts are found, it can be inferred that the beginning or end of that segment is also processed by neighboring nodes, and thus one or more primary scenes need to be split.
[0073] Consider the scenario shown in Figure 3C, where a parent scene P is processed by two nodes. Each part of the parent scene within a node is a primary scene for that node. One approach is to partition them independently using the partitioning approach described above. In this case, due to missing statistics, scene cuts in overlapping regions at nodes N and N+1 may not coincide with each other. The proposed partitioning algorithm works much better in resolving temporal inconsistencies if there is a good estimation of the secondary scene boundaries within neighboring subscenes on neighboring nodes.
[0074] In one embodiment, for the example in Figure 3C, two new synchronized subscene cuts (320) are introduced in the two primary scenes, one for each node. These synchronized cuts divide the primary scene into two parts: 1. The first portion (for example, between the first cut and the first sync cut at node N) is visible to the current node but not to the other node. As shown in Figure 3C, in one embodiment, the first sync cut on node N is at position C L-1 -B may also be, where B indicates the number of bumper frames, C L-1 This indicates the last frame of the segment. 2. The second portion (for example, the trailing bumper frames for node N and the initial bumper frames for node N+1) is "visible" to both nodes. As shown in Figure 3C, in one embodiment, the second synchronization cut on node N+1 may be at position C0+B, where C0 represents the first frame of the segment.
[0075] In one embodiment, these initial synchronous partitions may be performed as part of stage 410, and the partition algorithm 400 can be applied to these primary scenes. The only minor change is in the initialization stage 410, with a set Ω for each node. P This includes one additional sync scene cut (320). Since no further splitting is needed, we can jump directly to stage 420 after initialization. The algorithm then proceeds as usual.
[0076] Or, the original Omega P Given a set, if this sync subdivision detects that the primary scene is not entirely at the current node, it uses the aforementioned rule instead of equation (9) (for example, for node N, if the primary scene does not end at node N, then position C L-1 This may be done using -Add a scene cut to B.
[0077] Using these initial synchronization partitions, the subscene cuts calculated by isolated nodes N and N+1 are expected to be aligned to a reasonable degree. For node N, Ψ1 NLet it represent the first list of scenes obtained after merging the XML scene cut and the Auto - SCD scene cut shown in Figure 3B. These scenes are called primary scenes. Scenes on the segment boundary are split into secondary scenes or sub - scenes. The secondary scenes or sub - scenes generate additional scene cuts that are appended to the first list. Apart from these secondary scene cuts, the first frame f NL of the segment is also a scene cut. Since the node can only start encoding from the beginning of the segment, that frame is treated as a scene cut. Similarly, the last frame f (N+1)L-1 of the segment is the end of the last scene in the segment. For example, consider a node N with the following initial assignment: primary scenes 1, 2, and 3. Primary scene 1 is subdivided into secondary scenes A, B, and C, primary scene 2 remains unchanged, and primary scene 3 may be split into secondary scenes D, E, and F. Ψ l <...>and Ψ r <...>respectively show the sets of scene cuts near the left and right segment boundaries for node N. Then, the second list of scenes Ψ2 N can be mathematically expressed as follows.
Equation
[0078] For scenes longer than the segment length, instead of separate left or right sets, there may be a single set of secondary scene cuts. S k represents the start - frame index for the k - th scene in the list Ψ2 N . Assuming there are K scenes in the list, the elements in the list can be represented by the following formula. Here, S K indicates a dummy scene cut immediately after the last frame of the segment. It is only used as a marker at the end of the list.
Equation
[0079] The second list Ψ2 N of scenes has details about the primary and secondary scenes within the segment. The primary scenes do not require additional data from path 1 while the secondary scenes require the next auxiliary data from path 1. 1. The number of overlapping frames on the left and right for each secondary scene. 2. Trim - path correction data.
[0080] As used herein, "trim - pass" data or metadata refers to the "trim" data generated by the colorist during color grading to meet the director's intent. Sometimes, trimming leads to clipping of highlights and / or crushing of low - intensity values. When reconstructing HDR from an SDR affected by trimming, unwanted artifacts are introduced into the reconstructed HDR video. To reduce these artifacts, as discussed in Patent Document 5, the trim correction algorithm may require some supplementary data. The trim - path correction process may be part of the node - based processing, the details of which are beyond the scope of the present invention and are not discussed here.
[0081] [[ID=• For the primary scene, the statistics collection window includes all frames within the primary scene. Frames outside the primary scene are not referenced. Conversely, for secondary scenes, the statistics collection window includes all frames within that secondary scene, plus some overlapping frames from the previous or next subscene. These additional frames are called overlapping frames.
[0082] As a general rule, primary scenes do not overlap with any neighboring scenes, and secondary scenes are only allowed to overlap with neighboring secondary scenes. In other words, overlapping frames for subscenes can never come from neighboring primary scenes. The overlap parameter θ (see equation (13)) is set by the user, and its default value is 1. The backward phase in pass 2 does not use such overlaps for primary or secondary scenes.
[0083] To explain in detail the number of overlapping frames on the left and right, please refer to Figure 3D, which shows an exemplary embodiment with subscenes A through H. For subscene A, the forward reshaping statistics window shows no additional frames on the left and θ additional frames on the right (part of subscene B). Subscene H has additional frames only on the left (part of subscene G). Subscenes B, E, F, and G have additional frames on both the left and right. For subscenes C and D, the number of overlapping frames on the left and right is calculated in a slightly different way. Subscene C uses θ additional frames from subscene B on the left. On the right, all frames up to the next scene cut are used. In this example, since there are no scene cuts on the right, all bumper frames are taken. The scene cuts at the beginning of the segment are ignored when calculating the number of overlapping frames. The dotted box (330) on subscene C indicates the frames within subscene C and the overlapping frames on the right. For subscene D, there are θ additional frames on the right. On the left, all frames up to the previous scene split are taken. The scene divisions indicated by solid vertical lines mark the start of segments, and overlapping frames are ignored in the calculation. The dotted box (340) on subscene D indicates the frames within the subscene and the overlapping frames to the left.
[0084] The reason for ignoring the start of segment scene cuts is to ensure that the forward reshaping statistics windows (e.g., 330 and 340) for C and D are the same. Even if C and D are on different nodes, the same forward reshaping parameters are calculated for C and D. This helps achieve a consistent appearance in neighboring subscenes that span nodes. Synchronized scene cuts play a crucial role in aligning all scene cuts on nodes N and N+1, resulting in C and D having the same statistics window.
[0085] Path 2: Scene-based encoding As shown in Figure 2, each node is assigned a segment to be encoded into a bitstream. Pass 1, described above, generates a list of primary and secondary scenes that minimize the bitrate of reshaping-related metadata while maintaining temporal consistency across neighboring nodes. Pass 2 uses this list to encode the scenes within the segments and generate a reshaped SDR encoded bitstream. As shown in Figure 3A, bumper frames (309) are not encoded and are used to collect statistics in the forward reshaping pass to maintain the temporal stability of the secondary scenes.
[0086] As shown in Figure 2, pass 2 includes stages 225, 230, and 235. An alternative view of the same processing pipeline at the scene or subscene level is also shown in Figure 5. For forward reshaping, primary and secondary scenes are processed in a similar manner, with one major difference: primary scenes have no overlap in forward reshaping, while secondary scenes have some overlap with neighboring subscenes. For backward reshaping, the process is exactly the same for primary and secondary scenes. There is no overlap in the backward reshaping phase. A composite bitstream (240) consisting of reshaping metadata and a compressed base layer is generated as output. Details of each block are described below.
[0087] Given a segment-to-scene list (222), Figure 5 shows an exemplary architecture for scene-based encoding at each node in the cloud. Without limitation, block 225 in Figure 2 can be divided as shown using blocks 505 and 132. The starting frame index for the k-th scene is S k We will write it this way. Therefore, given scene k, the node is frame S k S k +1, S k +2, ..., and S k+1-1 needs to be processed. The reference HDR frame (504) and corresponding SDR frame (502) for the scene may be stored in the corresponding SDR and HDR scene buffers (not shown). As discussed earlier, bumper frames are used only to generate statistical data for secondary scenes and are ignored when processing primary scenes.
[0088] From Figure 5, in step 505, the input SDR and HDR frames are used to generate a scene-based forward reshaping function. The parameters of such a function are used for the entire scene (rather than being updated per frame), thus reducing the overhead for metadata 152. Next, in step 132, the forward reshaping is applied to the HDR scene (504) to generate a reshaped base layer 229, which is encoded by a compression unit (230) to produce an encoded bitstream 144. Finally, in step 235, the reshaped SDR data 229 and the original HDR data (504) are used to generate parameters 152 for a backward reshaping function, which should be sent together to a downstream decoder. These steps are described in more detail below. Without limitation, the steps are described in the context of what is called a three-dimensional mapping table (3DMT) representation. Here, to simplify the operation, each frame is represented as a 3D mapping table, where each color component (e.g., Y, Cb, or Cr) is subdivided into "bins," and instead of using explicit pixel values to represent the image, the pixel average within each bin is used. Details of the 3DMT formulation can be found in Patent Document 3.
[0089] The scene-based generation of the forward reshaping function (505) consists of two levels of operation. First, statistics are collected for each frame. For example, for rumor, SDR(h j s (b)) and HDR(h j v(b) Calculate histograms for both frames and store them in the frame buffer for the j-th frame, where b is the bin index. After generating the 3DMT representation for each frame, generate the "a / B" matrix representation which is expressed as follows:
number
[0090] Given statistics for each frame in the current scene, a scene-level algorithm can be applied to calculate the optimal forward reshaping coefficient. For example, for rumor, SDR(h s (b)) and HDR data (h v A scene-based histogram for (b)) can be generated by summing or averaging the frame-based histograms. For example, in one embodiment,
number
[0091] Having histograms at both scene levels, cumulative density function (CDF) matching (Patent Documents 4-5) can be applied to generate a forward mapping function (FLUT) from HDR to SDR. For example,
number
number
number
number
[0092] The generation of the scene-based back-reshaping function (152) also involves both frame-level and scene-level operations. Since the chroma mapping function is a single-channel predictor, the back-reshaping function can be obtained simply by inverting the forward-reshaping function. For chroma, the reshaped SDR data (229) and the original HDR data (504) are used to form a 3DMT representation, and a new frame-based a / B representation is generated.
number
[0093] At the scene level, for rumors, a backward rumor reshaping function may be generated by applying the histogram-weighted BLUT construction described in Patent Document 3. For chromas, a scene-based a / B representation can be calculated by averaging frame-based a / B representations.
number
number
number
[0094] Exemplary Receiving System Implementation Embodiments of the present invention may be implemented using computer systems, systems comprising electronic circuits and components, integrated circuit (IC) devices such as microcontrollers, field-programmable gate arrays (FPGAs), or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or apparatus comprising one or more such systems, devices, or components. Computers and / or ICs can execute, control, or run instructions relating to segment-to-scene segmentation and node-based processing in cloud-based video coding of HDR video, as described herein. Computers and / or ICs can calculate any of the various parameters or values relating to scene segmentation and node-based processing in cloud-based video coding of HDR video, as described herein. Embodiments of image and video dynamic range enhancement may be implemented in hardware, software, firmware, and various combinations thereof.
[0095] Certain implementations of the present invention include a computer processor that executes software instructions causing the processor to perform the method of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc., can implement the method for scene segmentation and node-based processing in cloud-based video coding of HDR video as described above by executing software instructions in program memory accessible to the processor. The present invention may be provided in the form of a program product. The program product may include any non-temporary and tangible medium that, when executed by a data processor, carries a set of computer-readable signals containing instructions causing the data processor to perform the method of the present invention. The program product according to the present invention may be any of a wide variety of non-temporary and tangible forms. The program product may include physical media such as magnetic data storage media including floppy diskettes and hard disk drives, optical data storage media including CD-ROMs and DVDs, and electronic data storage media including ROMs and flash RAM. The computer-readable signals on the program product may optionally be compressed or encrypted.
[0096] Where components (e.g., software modules, processors, assemblies, devices, circuits, etc.) are referred to above, unless otherwise indicated, references to such components (including references to “means”) should be interpreted as including any components that perform the function of the described component (e.g., functionally equivalent), including components that are not structurally equivalent to the disclosed structure performing the function of the described component in the illustrated embodiments of the present invention.
[0097] Equivalents, extensions, substitutes, and others In this manner, exemplary embodiments relating to scene segmentation and node-based processing in cloud-based video encoding of HDR video are described. In the foregoing specification, embodiments of the invention are described with reference to numerous specific details that may vary by implementation. Therefore, the sole and exclusive indicator of what constitutes the invention and what the applicant intends to constitute the invention is the set of claims issued from this application in the specific form in which the claims are granted, including any subsequent amendments. If there are definitions of terms contained herein that are expressly provided, those definitions govern the meaning of such terms as used in those claims. Thus, no limitations, elements, characteristics, features, advantages or attributes not expressly provided in the claims should in any way limit the scope of such claims. Accordingly, the specification and drawings should be taken in an illustrative rather than restrictive sense.
[0098] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE). [EEE1] A method for segmenting a video segment into scenes using a processor, the method being: Currently, the computing node receives a first video sequence containing high dynamic range video frames; For each video frame in the first video sequence, a step is to generate a frame-based forward reshaping function, wherein the forward reshaping function maps the frame pixels from the high dynamic range to a second dynamic range lower than the high dynamic range; The steps include generating a set of primary scenes for the first video sequence; A step of generating a second set of scenes for the first video sequence based on the set of primary scenes, secondary scenes derived from one or more primary scenes, and the frame-based forward reshaping function; The steps include: generating a scene-based forward reshaping function based on the aforementioned second set of scenes; The steps include: applying the scene-based forward reshaping function to the first video sequence to generate the second dynamic range output video sequence; The steps include compressing the output video sequence to generate the second dynamic range encoded bitstream, and given a primary scene, generating a list of secondary scenes for that primary scene: Based on the aforementioned set of primary scenes, the steps include initializing the set of secondary scenes and the set of violation scenes; A step of generating one or more sets of smoothing thresholds based on the frame-based forward reshaping function; Until boundary violations cease: Divide each scene in the aforementioned set of violation scenes into two new subscenes; Using the empty set, generate an updated set of violation scenes; An updated set of secondary scenes is generated by adding the aforementioned new subscene to the aforementioned set of secondary scenes; Using one or more sets of smoothing thresholds, perform one or more boundary violation checks in the set of secondary scenes; If there is at least one boundary violation between two subscenes in the set of secondary scenes, add the two subscenes to the set of violating scenes and continue subdividing the primary scene using the updated set of violating scenes and the updated set of secondary scenes; Otherwise, the process includes repeatedly signaling that there is no boundary violation and outputting the aforementioned list of secondary scenes, method. [EEE2] To generate the aforementioned set of primary scenes: The steps include: accessing a first set of scene cuts from an XML file associated with the first video sequence; A step of generating a second set of scene cuts for the first video sequence using an automatic scene change detector; A step of generating a final set of scene cuts based on the intersection of the first set of scene cuts and the second set of scene cuts; The step of generating the set of primary scenes using the final set of scene cuts, Methods described in EEE1. [EEE3] The method according to EEE1 or 2, wherein the primary scene is divided into secondary scenes only if the primary scene belongs to a parent scene having a picture frame encoded across the current computing node and neighboring computing nodes to the current computing node. [EEE4] Scene P in the set of violation scenes mentioned above g Given, the scene is at frame position C s It is divided into, Scene P g However, if it includes a primary scene which is part of a parent scene that has frames processed in nodes prior to the current node, C s =C0+B Here, C0 represents the first frame of the first video sequence, and B represents the number of bumper frames shared by the two adjacent nodes; Otherwise, Scene P g However, if it includes a primary scene which is part of a parent scene that has frames processed in nodes later than the current node, C s =C L-1 -B And here, C L-1 This indicates the last frame in the first video sequence; Otherwise, Scene P gIf it includes a secondary scene,
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
Claims
1. A method for segmenting video segments into scenes using a cloud-based system for encoding high dynamic range video, the method being: The cloud-based system's current computing nodes receive a first video sequence containing high dynamic range video frames; For each video frame in the first video sequence, the step of generating a frame-based forward reshaping function that maps the video frame from the high dynamic range to a second dynamic range lower than the high dynamic range; The steps include: generating a set of primary scenes for the first video sequence using a set of scene cuts for the first video sequence; A step of generating a second set of scenes for the first video sequence based on the set of primary scenes, wherein primary scenes belonging to a parent scene having video frames encoded across the current computing node and the neighboring computing nodes of the cloud-based system are divided into secondary scenes; For each scene in the second set of scenes, the step of generating a scene-based forward reshaping function that maps the video frames within that scene from the high dynamic range to the second dynamic range; The steps include: applying the scene-based forward reshaping function to the video frames in the first video sequence to generate an output video sequence including the video frames of the second dynamic range; The step includes compressing the output video sequence to generate an encoded bitstream, method.
2. Given a primary scene, generating a list of secondary scenes for that primary scene is: Based on the aforementioned set of primary scenes, the steps include initializing the set of secondary scenes and the set of violation scenes; The steps include generating one or more sets of smoothing thresholds based on the frame-based forward reshaping function; Until boundary violations cease: Divide each scene in the aforementioned set of violation scenes into two new subscenes; Using the empty set, generate an updated set of violation scenes; An updated set of secondary scenes is generated by adding the aforementioned new subscene to the aforementioned set of secondary scenes; Using one or more sets of smoothing thresholds, perform one or more boundary violation checks in the set of secondary scenes; If there is at least one boundary violation between two subscenes in the set of secondary scenes, add the two subscenes to the set of violating scenes and continue to subdivide the primary scene using the updated set of violating scenes and the updated set of secondary scenes; Otherwise, the process includes repeatedly signaling that there is no boundary violation and outputting the aforementioned list of secondary scenes, The method according to claim 1.
3. To generate the set of primary scenes using the set of scene cuts for the first video sequence: The steps include: accessing a first set of scene cuts for the first video sequence from a file; The steps include: generating a second set of scene cuts for the first video sequence using an automatic scene change detector; A step of generating a final set of scene cuts based on the intersection of the first set of scene cuts and the second set of scene cuts; The process includes the step of generating the set of primary scenes using the final set of scene cuts, The method according to claim 1 or 2.
4. Scene P in the set of violation scenes mentioned above g Given, the scene is at frame position C s It is divided into, Scene P g However, if it includes a primary scene which is part of a parent scene that has frames processed in the computing nodes of the cloud-based system, prior to the current computing node, C s =C 0 +B And here, C 0 represents the first frame of the first video sequence, and B represents the number of bumper frames shared by two adjacent computing nodes of the cloud-based system; Otherwise, Scene P g However, if it includes a primary scene which is part of a parent scene that has frames processed in a computing node of the cloud-based system, which is currently behind the computing node, C s =C L-1 -B And here, C L-1 indicates the last frame in the first video sequence; Otherwise, Scene P g If it includes a secondary scene, [Number 59] And here, [Number 60] This represents a frame-based forward reshaping function for frame j in the first video sequence, as a function of the input codeword b. [Number 61] Scene P g This represents the average of the frame-based forward reshaping function for the frames related to the specified frame. The method according to claim 2.
5. Generating one or more sets of smoothing thresholds is the smoothing threshold for each frame j in the first video sequence. [Number 62] This includes calculating the first set of, [Number 63] And here, [Number 64] And here, [Number 65] This represents a frame-based forward reshaping function for frame j in the first video sequence as a function of the input codeword b, and h j v (b) represents the histogram of the j-th frame in the first video sequence, where H and W represent the width and height values for the frame in the first video sequence. The method according to claim 2.
6. Second set of smoothness thresholds [Number 66] This further includes the step of calculating, [Number 67] The method according to claim 5, wherein α and β are constants.
7. Frame C g-1 Secondary scene P that begins from g-1 and frame C g Secondary scene P that begins from g Regarding this, one or more boundary violation checks can be performed between those two scenes: [Number 68] Test whether the following is true, and if it is true, declare a boundary violation, where ω is a constant, D Cg =λ Cg -l Cg-1 And, Scene P g Regarding frame j within the frame, [Number 69] And, [Number 70] This is the second scene P g And representing the average of the frame-based forward reshaping function in neighboring quadratic scenes, The method according to claim 6. [Request Item 8] [Number 71] And θ is an integer constant representing the frame overlap between the two subscenes, C 0 and C L-1 represents the first and last frames in the first video sequence. The method according to claim 7. [Request Item 9] [Number 72] This further includes testing whether the condition is true, and if true, declaring a boundary violation. The method according to claim 7 or 8, wherein, for a real number x, sign(x) returns 0 if x = 0, 1 if x > 0, and -1 if x < 0. [Request Item 10] [Number 73] The method according to any one of claims 7 to 9, further comprising testing whether the following is true and declaring a boundary violation if it is true.
11. For each scene in the aforementioned second set of scenes, a scene-based forward reshaping function can be generated: If a scene in the second set of scenes is a primary scene, a scene-based forward reshaping function for that scene is generated based solely on statistical data generated from the frames within the scene; otherwise, If a scene in the second set of scenes is a quadratic scene, the function includes generating a scene-based forward reshaping function for that scene based on statistics from frames in the scene and frames from neighboring quadratic scenes. The method according to any one of claims 1 to 10.
12. The steps include: generating a scene-based backward reshaping function based on the output video sequence, the first video sequence, and the scene-based forward reshaping function; The steps include: generating metadata based on the parameters of the aforementioned scene-based backward reshaping function; The step further includes outputting an output bitstream that includes the encoded bitstream and the metadata, The method according to claim 11.
13. A non-temporary computer-readable storage medium storing computer-executable instructions for executing the method according to any one of claims 1 to 12 on one or more processors.
14. An apparatus having a processor and configured to perform the method according to any one of claims 1 to 12.