Workload distribution and processing in cloud-based encoding of HDR video

By adopting a scenario-based workload distribution and shaping method in a cloud computing environment, the problems of workload distribution and shaping metadata overhead in HDR video encoding in a cloud computing environment are solved, and high-quality video encoding and uniform load distribution are achieved.

CN115968547BActive Publication Date: 2025-09-09DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180048767.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-09
Filing Date
2021-07-08
Publication Date
2025-09-09
Estimated Expiration
2041-07-08

AI Technical Summary

Technical Problem

When encoding HDR video in cloud computing environments, existing technologies have difficulty in achieving a balance between workload distribution and shaping-related metadata overhead, resulting in unacceptable overhead, especially at low bitrate transmission.

Method used

A scene-based workload distribution method is adopted to segment the video into scenes and generate scene-to-segment allocation through the scheduler node. The output bitstream is generated using scene-based forward and backward shaping functions to reduce the transmission of shaping metadata.

Benefits of technology

It effectively reduces the data rate of the shaping metadata and improves the quality of the encoded video. At the same time, it achieves a uniform distribution of workload among computing nodes and reduces the overhead in low-bitrate transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115968547B_ABST
    Figure CN115968547B_ABST
Patent Text Reader

Abstract

In a cloud-based system for encoding high dynamic range (HDR) video, compute nodes are assigned to become scheduler nodes, which segment the input video into scenes and generate scene-to-segment assignments for use by other compute nodes. The scene-to-segment assignment process consists of one or more iterations with an initial random assignment of scenes to compute nodes, followed by a refined assignment based on optimizing the allocation cost across all compute nodes. Methods for generating scene-based forward and backward shaping functions are also investigated to optimize video encoding and improve the coding efficiency of metadata related to the shaping.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Patent Application No. 63 / 049,673, filed on July 9, 2020, and European Patent Application EP20184883.5, each of which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure relates generally to images and, more particularly, to workload distribution and processing in cloud-based encoding of high dynamic range (HDR) video. Background Art

[0004] As used herein, the term 'dynamic range' (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, brightness) in an image, for example, from the darkest gray (black) to the brightest white (highlight). In this sense, DR is related to 'scene-referenced' intensities. DR may also relate to the ability of a display device to adequately or approximately render a range of intensities of a particular width. In this sense, DR is related to 'display-referenced' intensities. Unless expressly stated anywhere in this specification that a particular meaning has a particular meaning, it should be inferred that the term may be used in either sense, e.g., interchangeably.

[0005] As used herein, the term high dynamic range (HDR) refers to a DR width that spans 14-15 orders of magnitude of the human visual system (HVS). In practice, the DR of the extended width of the intensity range that humans can perceive simultaneously may be somewhat truncated relative to HDR. As used herein, the terms visual dynamic range (VDR) or enhanced dynamic range (EDR) may refer individually or interchangeably to the DR within a scene or image that can be perceived by the human visual system (HVS) including eye movement, thereby allowing for some light adaptation variations across the scene or image. As used herein, VDR may refer to a DR that spans 5 to 6 orders of magnitude. Therefore, although HDR, VDR, or EDR relative to a true scene reference may be somewhat narrow, it still represents a wider DR width and may also be referred to as HDR.

[0006] In practice, an image includes one or more color components (e.g., luma Y and chroma Cb and Cr), where each color component is represented by n bits of precision per pixel (e.g., n=8). For example, using gamma luma encoding where n≤8 (e.g., color 24-bit JPEG images) is considered an image with standard dynamic range, while images with n≥10 can be considered an image with enhanced dynamic range. HDR images can also be stored and distributed using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR file format developed by Industrial Light and Magic.

[0007] Most consumer desktop monitors currently support 200 to 300 cd / m 2 Most consumer HDTVs range from 300 to 500 nits, with newer models reaching 1000 nits (cd / m 2 Such conventional displays therefore represent a lower dynamic range (LDR) relative to HDR, also known as standard dynamic range (SDR). Due to the increased availability of HDR content, thanks to advances in capture devices (e.g., cameras) and HDR displays (e.g., the PRM-4200 Professional Reference Monitor from Dolby Laboratories), HDR content can now be color graded and displayed on HDR displays that support a higher dynamic range (e.g., from 1000 nits to 5000 nits or more).

[0008] As used herein, the term "forward reshaping" refers to the process of mapping a digital image from its original bit depth and original codeword distribution or representation (e.g., gamma, PQ, HLG, etc.) to a sample-to-sample or codeword-to-codeword image of the same or different bit depth and different codeword distribution or representation. Reshaping allows for improved compressibility or improved image quality at a fixed bit rate. As an example and not limitation, reshaping can be applied to 10-bit or 12-bit PQ encoded HDR video to improve coding efficiency in a 10-bit video coding architecture. In a receiver, after decompressing the received signal (which may or may not be shaped), the receiver can apply a "reverse (or backward) shaping function" to restore the signal to its original codeword distribution and / or achieve a higher dynamic range.

[0009] In many video distribution scenarios, HDR video may be encoded in a multi-processor environment, often referred to as a "cloud computing server." In such an environment, trade-offs between computational simplicity, workload balancing among computing nodes, and video quality may force encoding-related metadata to be updated on a frame-by-frame basis, which can result in unacceptable overhead, particularly when transmitting video at low bitrates. As appreciated by the inventors herein, improved techniques for workload distribution and node-based processing are desirable to improve the quality of encoded video in cloud-based environments while minimizing the overhead of encoding-related metadata.

[0010] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any approach described in this section qualifies as prior art solely by virtue of its inclusion in this section. Similarly, unless otherwise indicated, it should not be assumed that the problems identified for one or more approaches have been recognized in any prior art based on this section. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Embodiments of the invention are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals refer to similar elements, and in which:

[0012] Figure 1A depicts an example single-layer encoder for HDR data using a shaping function according to the prior art;

[0013] Figure 1B Depicts a method according to the prior art corresponding to Figure 1A An example HDR decoder for the encoder;

[0014] Figure 2 Depicting an example architecture for cloud-based encoding of HDR video according to an embodiment;

[0015] Figure 3A depicts an example process of scene-to-fragment dispatching according to an embodiment;

[0016] Figure 3B depicts an example of an improved dispatching process within a scene-to-fragment dispatching process according to an embodiment; and

[0017] Figure 4 An example encoder for scene-based encoding using shaping according to an embodiment of the present invention is depicted. DETAILED DESCRIPTION

[0018] This document describes methods for workload distribution and node-based processing in cloud-based video encoding of HDR video. In the following description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the present invention. However, it should be understood that the present invention can be practiced without these specific details. In other cases, well-known structures and devices are not described in exhaustive detail to avoid unnecessarily obscuring, obscuring, or obfuscating the present invention.

[0019] summary

[0020] Example embodiments described herein relate to cloud-based reshaping and encoding for HDR images. In one embodiment, in a cloud-based system for encoding HDR video, a node is arranged as a scheduler node that segments an input video into scenes and generates scene-to-segment assignments for use by other compute nodes. A processor in the scheduler node receives a sequence of scenes, where each scene includes one or more video frames, and then the processor:

[0021] receiving a sequence of scenes, wherein each scene comprises one or more video frames; and

[0022] Performing one or more dispatch iterations to generate an optimal output dispatch, wherein performing the one or more dispatch iterations comprises:

[0023] For an iteration in the one or more dispatch iterations:

[0024] generating an initial random assignment of the scene sequence to M computing nodes based on a random seed selection for the assignment iteration (305), where M>1;

[0025] Performing an improved allocation step (310) based on the initial random allocation to generate an improved allocation of the scene sequence to the M computing nodes and an improved allocation cost; and

[0026] The optimal dispatch cost and the optimal output dispatch are updated based on the improved dispatch and the improved dispatch cost (315).

[0027] In another embodiment, for a node among the M computing nodes, a processor in the node accesses a scene assigned to the node according to scene-to-fragment dispatching, the scene including a high dynamic range (HDR) frame sequence and a corresponding standard dynamic range (SDR) frame sequence, and uses a scene-based forward shaping function and a scene-based backward shaping function to generate an output bitstream and corresponding shaping metadata.

[0028] Example HDR encoding system

[0029] Figure 1A and Figure 1BAn example single-layer backward-compatible codec framework using image reshaping is shown according to the prior art. More specifically, Figure 1A An example encoder architecture is shown that may be implemented with one or more computing processors in an upstream video encoder. Figure 1B An example decoder architecture is shown that may also be implemented with one or more computing processors in one or more downstream video decoders.

[0030] In this framework, given reference HDR content (120) and corresponding reference SDR content (i.e., content representing the same image as the HDR content, but color-graded and represented in a standard dynamic range), the shaped HDR content (134) is encoded by an upstream encoding device implementing the encoder architecture and transmitted as SDR content in a single layer of an encoded video signal (144). The received SDR content is received and decoded by a downstream decoding device implementing the decoder architecture in a single layer of a video signal. Backward shaping metadata (152) is also encoded and transmitted in the video signal along with the shaped content, so that an HDR display device can reconstruct the HDR content based on the (shaped) SDR content and the backward shaping metadata. Without loss of generality, in some embodiments, as in non-backward compatible systems, the shaped SDR content may not be viewable by itself, but must be viewed in conjunction with a backward shaping function that will generate viewable SDR or HDR content. In other embodiments that support backward compatibility, a legacy SDR decoder can still play back the received SDR content without employing the backward shaping function.

[0031] like Figure 1A As shown, given an HDR image (120) and a target dynamic range, after generating a forward shaping function (132) in step 130; given the forward shaping function, a forward shaping mapping step (132) is applied to the HDR image (120) to generate a shaped SDR base layer (134). A compression block (142) (e.g., an encoder implemented according to any known video coding algorithm such as AVC, HEVC, AV1, etc.) compresses / encodes the SDR image (134) in a single layer (144) of the video signal. In addition, a backward shaping function generator (150) can generate a backward shaping function, which can be sent to a decoder as metadata (152). In some embodiments, the metadata (152) can represent the forward shaping function (130), so that the backward shaping function will be generated by the decoder (not shown).

[0032] Examples of backward shaping metadata representing / specifying an optimal backward shaping function may include, but are not necessarily limited to, any of the following: an inverse tone mapping function, an inverse luma mapping function, an inverse chroma mapping function, a lookup table (LUT), a polynomial, inverse display management coefficients / parameters, etc. In various embodiments, the luma backward shaping function and the chroma backward shaping function may be derived / optimized jointly or separately, and may be derived using various techniques, such as, but not limited to, those described later in this disclosure.

[0033] The backward shaping metadata generated by the backward shaping function generator based on the shaped SDR image and the target HDR image may be multiplexed as part of the video signal 144, for example, delivered as a Supplemental Enhancement Information (SEI) message.

[0034] In some embodiments, the backward shaping metadata (152) is carried in the video signal as part of the overall image metadata, separately from the single layer of the video signal that is encoded with the SDR image. For example, the backward shaping metadata (152) may be encoded in a component stream in the coded bitstream that may or may not be separate from the single layer of the coded bitstream that is encoded with the SDR image (134).

[0035] Therefore, the backward shaping metadata (152) can be generated or pre-generated on the encoder side to take advantage of the powerful computing resources and offline encoding streams available on the encoder side (including but not limited to content-adaptive multi-pass, look-ahead operations, inverse luma mapping, inverse chroma mapping, CDF-based histogram approximation and / or transmission, etc.).

[0036] Figure 1A The encoder architecture may be used to avoid encoding the target HDR image (120) directly as an encoded / compressed HDR image in the video signal; instead, backward shaping metadata (152) in the video signal may be used to enable a downstream decoding device to backward shape the SDR image (134) (which is encoded in the video signal) into a reconstructed image that is identical to or closely / best approximates the reference HDR image (120).

[0037] In some embodiments, as Figure 1BAs shown, a video signal encoded with a shaped SDR image in a single layer and backward shaping metadata as part of the overall image metadata are received as input at the decoder side of the codec framework. A decompression block (154) decompresses / decodes the compressed video data in the single layer (144) of the video signal into a decoded SDR image (156). Decompression 154 typically corresponds to the inverse of compression 142. The decoded SDR image (156) may be identical to the SDR image (134), subject to quantization errors in the compression block (142) and the decompression block (154), which may be optimized for an SDR display device. In a backward compatible system, the decoded SDR image (156) may be output in an output SDR video signal (e.g., via an HDMI interface, via a video link, etc.) for presentation on an SDR display device.

[0038] Optionally, alternatively, or additionally, in the same or another embodiment, a backward shaping block 158 extracts backward (or forward) shaping metadata (152) from the input video signal, constructs a backward shaping function based on the shaping metadata (152), and performs a backward shaping operation on the decoded SDR image (156) based on the optimal backward shaping function to generate a backward shaped image (160) (or a reconstructed HDR image). In some embodiments, the backward shaped image represents an HDR image of the same or closely / best approximation of the production quality or near production quality as the reference HDR image (120). The backward shaped image (160) can be output in an output HDR video signal (e.g., via an HDMI interface, via a video link, etc.) for presentation on an HDR display device.

[0039] In some embodiments, as part of an HDR image rendering operation to render the back-shaped image (160) on an HDR display device, display management operations specific to the HDR display device may be performed on the back-shaped image (160).

[0040] Cloud-based coding

[0041] Existing shaping techniques can be frame-based, i.e., new shaping metadata is transmitted with each new frame, or scene-based, i.e., new shaping metadata is transmitted with each new scene. As used herein, the term "scene" of a video sequence (frame / image sequence) can refer to a series of consecutive frames in the video sequence that share similar brightness, color, and dynamic range characteristics. Scene-based approaches work well in video workflow pipelines that have access to the entire scene; however, it is not uncommon for content providers to use cloud-based multiprocessing, where after dividing the video stream into segments, each segment is processed independently by a single compute node in the cloud. As used herein, the term "segment" refers to a series of consecutive frames in a video sequence. A segment can be part of a scene or can include one or more scenes. Therefore, the processing of a scene can be split across multiple processors.

[0042] As discussed in Reference [1], in some cloud-based applications, under certain quality constraints, segment-based processing may require the generation of shaping metadata on a frame-by-frame basis, resulting in undesirable overhead. This can be a problem in very low bitrate applications (e.g., below 1 Mbit / s). Figure 2 Depicted is an example architecture of a novel, scenario-based distributed architecture that allows reducing the data rate of shaping metadata without compromising the quality of the decoded video.

[0043] like Figure 2 As shown, the proposed architecture comprises two stages: a) a scheduler stage (205), typically but not limitedly implemented on a single compute node that distributes scenes into segments; and b) an encoding stage (210), where each node in the cloud encodes a sequence of segments.

[0044] Given a video source (202) for content distribution (often referred to as a mezzanine file), the first stage node obtains video metadata (e.g., from an XML file) and: (a) in step 215, it determines scene boundaries; and (b) in step 220, it determines a scene-to-segment dispatch list for each worker node. The main goal of scene boundary determination is to ensure that there are no significant brightness or color changes during normal playback within a scene, including fade-ins, fade-outs, and dissolves. (Dissolve in video editing refers to a smooth transition from one image to another. A dissolve between a blank (or black) image to another image is also called a fade-in or fade-out.) The goal of the scene-to-segment dispatch unit (220) is to ensure that a scene is not split and encoded in two different compute nodes, which may result in abrupt changes near segment boundaries. In addition, the dispatch tasks should strive for an even workload across all compute nodes (210).

[0045] In the second phase (210), each compute node receives its own scene-to-segment list (S2S list) (230) from phase one and its own partial mezzanine (225) of the corresponding segment from the input video (202). Each node encodes the assigned segment in parallel and outputs a coded bitstream. The details of each processing task are discussed below.

[0046] Scheduler node

[0047] There are two main scenarios of interest, depending on the workload distribution requirements within each node. In one embodiment, segments can have non-uniform lengths, allowing for non-uniform workloads across different nodes. This is tailored for scenario-based solutions, where a scenario cannot be split to be encoded in more than one node. In another embodiment, segments have a fixed length, enforcing a uniform workload across all nodes. With the proposed scheduler and worker node model, the proposed architecture can handle both scenarios.

[0048] In a scenario where the segment lengths are uneven, each worker node may receive a different workload for processing. To enable scene-based encoding, the scheduler node reads the XML file (extracted from the mezzanine) and determines the scene boundaries, in particular how to split or merge frames in fade-ins / fades-outs, and dissolve scenes into new scene cut boundaries. In these newly defined scenes, the scheduler determines which scene should be encoded by which node. The output of this process will be a scene-to-segment (S2S) list (230). The main goal is to distribute the number of frames in each node as evenly as possible. In one embodiment, without limitation, the metric for measuring uniformity in this stage is the standard deviation of the number of frames distributed in each node. A lower standard deviation means a more even workload in each node. In one embodiment, the S2S list can be derived as the output of an optimization problem for the best uniform load across all nodes under the constraint of uninterrupted scene processing.

[0049] Scene cuts can be defined in an XML file in the video source (202), but typically such metadata defines color grading boundaries. For example, to facilitate color grading, a scene cut marker can be inserted during a dissolve scene so that the display management process does not distort the colors during playback. However, such XML data does not take into account that the baseline data has been shaped, and shaping may affect the final appearance within the dissolve. In one embodiment, to avoid such problems, the dissolve can be divided into multiple single frames per scene to allow for slow transitions along the time domain. Note that this approach will increase the bitrate of the metadata related to shaping during these special transition effects. The same technique can also be applied to fade-in and fade-out transitions.

[0050] When the XML file is not available, the scheduler will need to identify scene cuts itself using any known scene cut detection technique known in the art. For example, in one embodiment, the brightness change along the time domain can be measured and checked to see if the change has a constant rate. Once a scene cut is detected, the entire scene can be split into individual frames, each representing a separate "scene".

[0051] In addition to the above methods, soft transitions near scene cut boundaries can be considered to avoid false scene cut boundaries. For example, for detected scene cuts, a small number of single-frame "scenes" can be added before the scene cut and a small number of single-frame "scenes" can be added after the scene cut. This approach will improve the bitrate of scene-based metadata.

[0052] Given the scene boundary decision (215), the scene to segment unit (220) decides which scene should be included in which segment. This assignment will produce a scene to segment assignment list (S2S list) (230). The scheduler node will output an S2S list for each worker node.

[0053] Consider a video sequence with a total of J frames, grouped into K scenes. Let the corresponding starting frame index of the kth scene be denoted as S k , and denote the number of frames of the kth scene as D k , where k = 0, 1, ..., K-1. Therefore:

[0054] D k =S k+1 -S k (1)

[0055]

[0056] The number of working nodes is expressed as M. In order to assign scenarios to each node, in one embodiment, the following rules can be implemented:

[0057] A scene cannot be split into two or more smaller sub-scenes to be processed in more than one node. In other words, a complete scene must be processed within one node to maintain temporal stability and compression efficiency of shaping-related metadata.

[0058] Nodes should not process scenes that are not sequential in time. For example, you would not want node n to process scenes 3, 6, and 7 because scenes 3 and 6 are not sequential, and at some point scenes 4 and 5 would need to be inserted between scenes 3 and 6. Processing non-sequential scenes would require a post-processing step to reassemble all scenes sequentially, thus requiring additional post-processing.

[0059] The set of scenarios assigned to node m is represented as Φ m, where m = 0, 1, ..., M - 1. Following the aforementioned rules, a first scenario index can be defined within Φ m as φ m , where φ m has a value range between 0 and K - 1. In one embodiment, to simplify implementation, a monotonically increasing rule can be implemented, i.e.:

[0060] φ m < φ n when m < n

[0061] In one embodiment, φ0 = 0, and thus, the first scenario is always assigned to the first segment. When the number of scenarios K is greater than the number of nodes M, φ m must be unique, i.e., its value cannot be the same in any other node. This is to ensure that no node has a zero workload. When K < M, a simple solution is to assign one scenario to each node; and leave the remaining nodes without scenarios.

[0062] Figure 3A Illustrates an example process (300) of scenario to segment assignment according to an embodiment. As Figure 3A shown, the process starts with an initial random assignment (305). In this step, the list of K scenarios is randomly split into M segments. This initial list will be further adjusted using an iterative algorithm. Step 305 initializes two sets:

[0063] · The candidate set (Ω (t) ) is the original list of scenario indices at the end of the t-th iteration (for t > 0); and

[0064] · The selected set (Ψ (t) ) is the list of assigned scenario indices at the end of the t-th iteration.

[0065] At the start of the operation, i.e., t = 0, Ω (0) includes all scenarios except the first scenario (i.e., Ω (0) = {φ m |m = 1, ..., K - 1}) and Ψ (0) contains only the first scenario φ0. In the t-th iteration (t > 0), an element is randomly selected from Ω (t-1) , removed from Ω (t-1) , and the selected element is placed in the set Ψ (t) . This process is repeated M - 1 times until Ψ (t) contains M elements sorted in ascending order. The sorted Ψ (t) will be the output of this stage. Table 1 expresses this process in pseudocode.

[0066] Table 1: Initialization steps in scene-to-segment assignment

[0067]

[0068] As an example, consider a list of 10 scenes to be distributed among 3 nodes, each scene having a variable number of frames as shown below

[0069] Scene index (k) 0 1 2 3 4 5 6 7 8 9 <![CDATA[Number of frames (D k )]]> 3 9 4 7 2 5 4 8 3 10

[0070] Let the output of step 305 be Ψ (2) ={0,3,8}, then after this step, the scene is assigned to the nodes (or fragments) as follows:

[0071] Node 0: Scene 0-2

[0072] Node 1: Scenarios 3 to 7

[0073] Node 2: Scenes 8-9

[0074] In step 310, the initial random assignment (Ψ (M-1) ),like Figure 3B As shown. Figure 3B As shown, there are two iteration steps: a) one at the node level (for all nodes), and b) one at the total dispatch cost level (until convergence). Starting from step 345, the total dispatch cost is initialized to a large value that can approximate the maximum possible dispatch cost (e.g., ). Next, the algorithm iterates for each node m, where m=0, 1, 2, ...M-1. When iterating over each node, the workload of each node is checked. Three possible scenarios can be used to adjust the workload of each node (350):

[0075] (A) Remove its last scene and move it to the next node (this does not apply to node M-1)

[0076] (B) Add one more scene from the previous node (this does not apply to node 0)

[0077] (C) Maintain the current allocation.

[0078] Among these three options, in step 355, the cost associated with the dispatch (e.g., the standard deviation of the number of frames in each segment) is measured. A lower cost means a more uniform workload, and this is more preferred. Therefore, for each node, in step 360, the setting that produces the lowest dispatch cost is selected. After all nodes have been processed, in step 362, the lowest dispatch cost (e.g., )) and the existing total dispatch costs (e.g. ) is compared. If the lowest dispatch cost is deemed lower than the total dispatch cost, the value of the total dispatch cost is updated with the lowest dispatch cost, and the process returns to step 350. Otherwise, if there is no cost improvement, or the improvement is deemed too small, then in step 365, the improved dispatch phase 310 terminates by outputting the last scene-to-segment dispatch, referred to as the improved S2S dispatch and the improved dispatch cost, and its corresponding cost (i.e., the last value of the total dispatch cost).

[0079] In another embodiment, instead of starting the node iteration (e.g., steps 350, 355, and 360) at node 0 and moving forward, the node iteration may also start at node M-1 and move backward. Alternatively, a bidirectional iteration may be attempted between all nodes, and the workload with the lowest cost may be selected between the two.

[0080] After stage 310, given the improved scene-to-segment assignment, a new optimal overall dispatch cost (and associated S2S assignment) may be calculated in step 315. In one embodiment, to avoid poor random initialization steps 305 that may result in suboptimal assignments, steps 305-315 are repeated L times for L different random initialization steps 305 (e.g., by using different random seed generators), each yielding an overall dispatch cost (l), l = 1, 2, ... L (e.g., Then, in step 315, the algorithm with the best overall cost (e.g., the smallest standard deviation) is selected. The experimental results show that L = 100 combined with the improved allocation step (310) produces satisfactory results, and a larger value of L does not significantly improve the overall S2S allocation strategy.

[0081] Therefore, when l = 1, Simply express the first improvement dispatch cost (i.e., ),in represents the optimal overall dispatch cost. In subsequent iterations, if then ignore this iteration, otherwise update the optimal dispatch cost (e.g., ), and the corresponding workload of this iteration is considered the best scene-to-fragment dispatch.

[0082] Step 320 checks whether all L iterations are completed, if so, then in step 325 the best scene-to-segment assignment is output, ie the one with the best cost among all L iterations, otherwise the process is repeated with another initial random assignment (305).

[0083] To facilitate discussion, add one more variable To indicate the end of the video sequence. Set, m = 0, 1, ..., M, the number of frames in each node at the tth iteration can be calculated as

[0084] For m=0,1,…,M-1. (3) In one embodiment, workload uniformity or dispatch cost can be defined as The standard deviation of

[0085]

[0086]

[0087] The smaller the value of , the more evenly the workload is distributed to each node. Table 2 lists the pseudo code for this improved dispatch phase.

[0088] Table 2: Example code for the improved dispatch phase in scene-to-fragment assignment

[0089]

[0090]

[0091] While in one embodiment and not limitation, the standard deviation of the frames being used provides a good cost metric for improving the dispatch phase, alternative cost metrics may also be applied, such as:

[0092] The workload range is measured by subtracting the minimum total number of dispatched frames from the node from the maximum total number of dispatched frames in the node

[0093] The average workload of all frames in each node (e.g., the )

[0094] · The average distance of each node's workload from the overall average

[0095] Returning to our example, Tables 3 and 4 describe the scene-to-segment assignments and the corresponding S2S parameters after the random initialization phase. The total cost of measuring the standard deviation between the values ​​in can be calculated as

[0096] Table 3: Example S2S allocation after initialization

[0097]

[0098] Table 4: Example allocation parameters after the initialization phase

[0099]

[0100]

[0101] Now consider the example of improving the dispatch (310). In the first iteration of this phase, where t=0, for the first node, m=0, three different strategies are tried and the standard deviation of each case is measured. The results are shown in Tables 5 and 6. As shown in Table 5, under Option A, node 0 was assigned only scenarios 0 and 1, with a cost of 10.12. Under Option B, node 0 was assigned scenarios 0-3, with a cost of 5.04. Under Option C (which was unchanged from before), the cost remained unchanged (6.81). Therefore, Option B was selected as the optimal strategy to continue improving scenario assignments at subsequent nodes, where the same process would be repeated.

[0102] Table 5: Example Improved Dispatch for Node 0

[0103]

[0104] Table 6: S2S parameters for options A, B, and C for node 0

[0105] <![CDATA[φ0 (0) ]]> <![CDATA[φ1 (0) ]]> <![CDATA[φ2 (0) ]]> <![CDATA[f0 (0) ]]> <![CDATA[f1 (0) ]]> <![CDATA[f2 (0) ]]> <![CDATA[σ f (0) ]]> A 0 2 8 12 30 13 10.12 B 0 4 8 23 19 13 5.04 C 0 3 8 16 26 13 6.81

[0106] In this example, at the end of t=0, the optimal S2S is still the cost shown in Table 6 One of 5.04; namely:

[0107] Node 0: Scenes 0-3

[0108] Node 1: Scenarios 4-7

[0109] Node 2: Scenes 8-9

[0110] Next, when t = 1, steps 350, 355, and 360 are repeated. In this example, when t = 1, the overall cost has not improved, so the process will terminate.

[0111] In some embodiments, it may be preferred that all segments have the same number of frames. In this case, the number of frames for the first M-1 nodes can be assigned as

[0112]

[0113] The remaining frames will be dispatched to the last node (node ​​M-1)

[0114]

[0115] Scenario-based coding

[0116] Given a scene to segment assignment (230), Figure 4 An example architecture for scene-based encoding on each node in the cloud is depicted (210). Recall that the starting frame index of the k-th scene is denoted as S k .. Therefore, given a scene k, the node needs to process frame S k ,S k +1,S k +2,…, and S k+1 - 1. The reference HDR frame (404) and corresponding SDR frame (402) of the scene may be stored in corresponding SDR and HDR scene buffers (not shown).

[0117] according to Figure 4 In step 405, a scene-based forward shaping function is generated using the input SDR and HDR frames. The parameters of such a function will be used for the entire scene (rather than updated on a frame-by-frame basis), thereby reducing the overhead of metadata 152. Next, in step 132, forward shaping is applied to the HDR scene (404) to generate a shaped base layer 407, which will be encoded by the compression unit (142) to generate the coded bitstream 144. Finally, in step 410, the shaped SDR data 407 and the original HDR data (404) are used to generate parameters 152 for the backward shaping function to be sent together to the downstream decoder. These steps will be described in more detail below. Without limitation, the steps are described in the context of a so-called three-dimensional mapping table (3DMT) representation, where, to simplify operations, each frame is represented as a three-dimensional mapping table where each color component (e.g., Y, Cb, or Cr) is subdivided into "bins" and, rather than using explicit pixel values ​​to represent the image, the average of the pixels within each bin is used. Details of the 3DMT formula can be found in reference [3].

[0118] The scene-based generation of the forward shaping function (405) involves two stages of operation. First, statistics of each frame are collected. For example, for brightness, the SDR is calculated. and HDR The histograms of both frames are generated and stored in the frame buffer of the jth frame, where b is the interval index. After generating the 3DMT representation for each frame, the "a / B" matrix representation is generated, which is expressed as:

[0119]

[0120]

[0121] Where ch refers to the luminance or chrominance channel (for example, Y, Cb, or Cr), represents the transposed matrix of the parametric model based on the reference HDR scene data and the forward shaping function, and A vector representing a parametric model based on the SDR scene data and a forward shaping function.

[0122] Given the statistics of each frame in the current scene, a scene-level algorithm can be applied to calculate the optimal forward shaping coefficients. For example, for brightness, the SDR (B s (b)) and HDR data (B v (b) Generate a scene-based histogram. For example, in one embodiment,

[0123]

[0124]

[0125] With two scene-level histograms, cumulative density function (CDF) matching (references [4-5]) can be applied to generate a forward mapping function (FLUT) from HDR to SDR, e.g.,

[0126]

[0127] For chrominance (e.g., ch=Cb or ch=Cr), the frame-based representations of a / B in equation (7) can again be averaged to generate a scene-based a / B matrix representation given by

[0128]

[0129]

[0130] The parameters of the multi-color, multivariate regression (MMR) model for the shaping function are as follows (references [2-3])

[0131] m F,ch =(B F ) -1 a F,ch (11)

[0132] The shaped SDR signal (407) can then be generated as:

[0133]

[0134] Generating the scene-based backward shaping function (410) also involves both frame-level and scene-level operations. Since the luma mapping function is a single-channel predictor, the forward shaping function can be simply restored to obtain the backward shaping function. For chrominance, the shaped SDR data (407) and the original HDR data (404) are used to form a 3DMT representation, and the new frame-based a / B representation is calculated as follows:

[0135]

[0136]

[0137] At the scene level, for luminance, the histogram-weighted BLUT construction in reference [3] can be applied to generate the backward luminance shaping function. For chrominance, the frame-based a / B representation can be similarly averaged to calculate the scene-based a / B representation.

[0138]

[0139]

[0140] The MMR model solution of the backward shaping mapping function is given by

[0141] m B,ch =(B B ) -1 a B,ch (15)

[0142] Then, the reconstructed HDR signal (160) can be generated as:

[0143]

[0144] References

[0145] Each of these references is incorporated herein by reference in its entirety.

[0146] 1. H. Kadu et al., “Coding of high-dynamic range video using segment-based reshaping,” U.S. Patent 10,575,028.

[0147] 2. GM. Su et al., “Multiple color channel multiple regression predictor,” U.S. Patent 8,811,490.

[0148] 3. Q. Song et al., PCT patent application serial number PCT / US2019 / 031620, “High-fidelity full reference and high-efficiency reduced reference encoding in end-to-end single-layer backward compatible encoding pipeline,” filed May 9, 2019, published as WO 2019 / 217751.

[0149] 4. B. Wen et al., “Inverse luma / chroma mappings with histogram transfer and approximation,” U.S. Patent 10,264,287.

[0150] 5. H. Kadu and GM. Su, “Reshaping curve optimization in HDR coding,” U.S. Patent 10,397,576.

[0151] Example Computer System Implementation

[0152] Embodiments of the present invention may be implemented using a computer system, a system configured in electronic circuits and components, an integrated circuit (IC) device (such as a microcontroller, a field programmable gate array (FPGA) or another configurable or programmable logic device (PLD)), a discrete-time or digital signal processor (DSP), an application-specific IC (ASIC), and / or an apparatus comprising one or more such systems, devices, or components. The computer and / or IC may implement, control, or execute instructions related to workload distribution and node-based processing in cloud-based video encoding of HDR video, such as those described herein. The computer and / or IC may calculate any of the various parameters or values ​​related to workload distribution and node-based processing in cloud-based video encoding of HDR video described herein. The image and video dynamic range extension embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0153] Certain implementations of the present invention include a computer processor executing software instructions that cause the processor to perform the method of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. can implement the method of workload distribution and node-based processing in cloud-based video encoding of HDR video as described above by executing software instructions in a program memory accessible to the processor. The present invention can also be provided in the form of a program product. The program product can include any non-transient and tangible medium carrying a set of computer-readable signals that, when executed by a data processor, cause the data processor to perform the method of the present invention. The program product according to the present invention can be in any of a variety of non-transient and tangible forms. The program product can include, for example, physical media, such as magnetic data storage media including floppy disks and hard drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAM, etc. The computer-readable signals on the program product can optionally be compressed or encrypted.

[0154] Where a component (e.g., a software module, processor, accessory, device, circuit, etc.) is mentioned above, unless otherwise specified, reference to the component (including reference to "means") should be interpreted to include any component that performs the function of the described component (e.g., functionally equivalent) as an equivalent form of the component, including components that are not structurally equivalent to the disclosed structure that performs the function in the example embodiments shown in this invention.

[0155] Equivalent forms, extended forms, alternative forms and others

[0156] The Enumerated Example Embodiments (EEE) of the present invention are defined as follows, but are not limited thereto:

[0157] EEE1. A method for assigning a scene sequence to segments for encoding by one or more computing nodes, the method comprising:

[0158] receiving a sequence of scenes, wherein each scene comprises one or more video frames; and

[0159] Performing one or more dispatch iterations to generate an optimal output dispatch, wherein performing the one or more dispatch iterations comprises:

[0160] For an iteration in the one or more dispatch iterations:

[0161] generating an initial random assignment of the scene sequence to M computing nodes based on a random seed selection for the assignment iteration (305), where M>1;

[0162] Performing an improved assignment step (310) based on the initial random assignment to generate an improved assignment of the scene sequence to the M computing nodes and an improved assignment cost; and

[0163] The optimal dispatch cost and the optimal output dispatch are updated based on the improved dispatch and the improved dispatch cost (315).

[0164] EEE2. The method of EEE1, wherein performing the improved dispatching step (310) comprises:

[0165] Initialize the total dispatch cost with a first value;

[0166] For each computing node, setting a node workload according to the initial random assignment of the scene sequence to the M computing nodes; and

[0167] Repeat the following steps until convergence:

[0168] For each node sequentially, starting with the first node:

[0169] removing a scenario from the node workload and assigning it to the workload of its next available node, and calculating a first cost metric for the M compute nodes;

[0170] adding a scenario to the node workload taken from the workload of an available node thereon, and calculating a second cost metric for the M compute nodes;

[0171] Keeping the node workload unchanged, and calculating a third cost metric for the M computing nodes; and

[0172] generating (360) an updated node workload based on a minimum of the first cost metric, the second cost metric, and the third cost metric;

[0173] calculating an iterative dispatch cost based on the updated node workload; and

[0174] If the total dispatch cost is less than the iterative dispatch cost, then: signal convergence, output the updated node workload as the improved dispatch, and output the total dispatch cost as the improved dispatch cost, otherwise: continue by replacing the total dispatch cost with the iterative dispatch cost.

[0175] EEE3. The method of EEE2, wherein the first value comprises an estimate of a maximum possible standard deviation of the total number of frames dispatched to each node. EEE4.

[0176] EEE4. The method of any one of EEE1-EEE3, wherein generating the initial random assignment comprises:

[0177] generating a candidate set having scene indices ranging from 1 to K-1, where K represents the total number of scenes in the scene sequence to be assigned to the M computing nodes;

[0178] Generate a dispatch set whose first element is 0;

[0179] updating the dispatch set according to a random selection selected using the random seed to generate an updated dispatch set;

[0180] sorting the updated dispatch set in ascending order to generate a sorted dispatch set;

[0181] Generating the initial random assignment according to the sorted set of assignments, wherein updating the set of assignments comprises:

[0182] For t=1 to M-1:

[0183] Select a random integer p between 0 and Kt-1;

[0184] identifying the pth element in the candidate set and appending it to the dispatch set;

[0185] Remove the pth element from the candidate set; and

[0186] The candidate set is sorted in ascending order.

[0187] EEE5. The method of EEE4, wherein generating the initial random assignment according to the sorted assignment set comprises:

[0188] All scenes with indices between values ​​equal to or greater than the mth element in the sorted dispatch set but less than the m+1th element in the sorted dispatch set are dispatched to node m.

[0189] EEE6. The method of any one of EEE1-EEE5, wherein calculating the cost metric for all computing nodes based on the scenario-to-node assignment of each computing node comprises:

[0190] For each computing node, calculating a total number of frames assigned to the computing node based on the scene-to-node assignment; and

[0191] Calculate the standard deviation of the total number of frames dispatched to each compute node.

[0192] EEE7. The method of any one of EEE1-EEE6, wherein removing a scenario from the node workload and assigning it to the workload of its next available node comprises:

[0193] The last scenario assigned to the node is identified and assigned as the first scenario to the workload of the next available node.

[0194] EEE8. The method of any one of EEE1-EEE7, wherein adding a scenario to a node workload taken from the workload of its previous available node comprises:

[0195] The last scenario assigned to the last available node is identified and assigned to the node workload as the first scenario.

[0196] EEE9. The method of any one of EEE1-EEE8, wherein updating the optimal dispatch cost and the optimal output dispatch comprises:

[0197] For a first dispatch iteration, setting the improved dispatch as the best output dispatch and setting the improved dispatch cost to the best dispatch cost; and

[0198] For subsequent dispatch iterations, the improved dispatch cost is compared to the optimal dispatch cost; and if the optimal dispatch cost is greater than the improved dispatch cost, the improved dispatch is selected as the optimal output dispatch and the improved dispatch cost is selected as the optimal dispatch cost.

[0199] EEE10. The method as described in any one of EEE1-EEE9, further comprising:

[0200] For a node among the M computing nodes:

[0201] accessing a sequence of high dynamic range (HDR) frames and a corresponding sequence of standard dynamic range frames (SDR) for a scene assigned to the node according to the optimal output assignment of the scene sequence to the node; and

[0202] An output bitstream is generated for the scene assigned to the node.

[0203] EEE11. The method of EEE10, wherein generating the output bit stream further comprises:

[0204] generating a scene-based forward shaping function based on the HDR frame sequence and the SDR frame sequence;

[0205] mapping the HDR frame sequence to a shaped SDR frame sequence based on the scene-based forward shaping function;

[0206] generating a coded bitstream by compressing the shaped SDR frame sequence;

[0207] generating a scene-based backward shaping function based on the shaped SDR frame sequence, the HDR frame sequence, and the scene-based forward shaping function;

[0208] generating metadata based on parameters of the scene-based backward shaping function; and

[0209] The output bitstream includes the encoded bitstream and the metadata.

[0210] EEE12. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions being used to execute the method described in any one of EEE1-EEE11 using one or more processors.

[0211] EEE13. A device comprising a processor and configured to perform any method described in EEE1-EEE11.

[0212] Thus described are example embodiments of workload distribution and node-based processing in cloud-based video encoding involving HDR video. In the foregoing specification, embodiments of the invention have been described with reference to many specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indicator of what is the invention, and what the applicants intend the invention to be, is the set of claims issuing from this application, in the specific form in which such claims issue, including any subsequent corrections. Any definitions of terms contained in such claims expressly set forth herein shall govern the meaning of such terms as used in the claims. Accordingly, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should in any way limit the scope of such claim. Accordingly, the specification and drawings should be regarded in an illustrative sense, and not in a restrictive sense.

Claims

1. A method for distributing scenes of a sequence of scenes to be encoded by a plurality of computing nodes, the method comprising: receiving a sequence of scenes, wherein each scene comprises one or more video frames; as well as performing one or more dispatch iterations to generate an optimal output dispatch, wherein each scene in the sequence of scenes is scheduled for encoding at a particular compute node among the M compute nodes, where M>1, wherein performing the one or more dispatch iterations comprises: For an iteration in the one or more dispatch iterations: For each computing node, generating an initial random assignment of one or more scenes in the scene sequence to the corresponding computing node based on a random seed selection for the assignment iteration (305); For each computing node, performing a refined assignment step (310) based on the initial random assignment to generate a refined assignment of one or more scenes in the scene sequence to the corresponding computing node and a refined assignment cost reflecting a workload imposed on the corresponding computing node when encoding the assigned scenes; and The uniformity of workload distribution on the computing nodes is maximized by minimizing an improved dispatch cost of the M computing nodes, wherein the dispatch cost includes a standard deviation of the number of video frames scheduled for encoding at the corresponding computing node, and updating an optimal dispatch cost and an optimal output dispatch based on the improved dispatch (315).

2. The method of claim 1, wherein performing the improving dispatching step (310) comprises: Initialize the total dispatch cost with a first value; For each computing node m, setting a node workload according to the initial random assignment of the scene sequence to the M computing nodes; and Repeat the following until convergence: Sequentially for each compute node m, starting from compute node m = 0 until reaching compute node m = M-1: Only for compute node m<M-1, remove the scenario from the node workload and assign it to the workload of compute node m+1, and calculate a first cost metric for the M compute nodes; Only for compute node m>0, adding a scenario to the node workload taken from the workload of compute node m-1, and calculating a second cost metric for said M compute nodes; Keeping the node workload unchanged, and calculating a third cost metric for the M computing nodes; as well as generating (360) an updated node workload based on a minimum of the first cost metric, the second cost metric, and the third cost metric; calculating an iterative dispatch cost based on the updated node workload; as well as If the total dispatch cost is less than the iterative dispatch cost, then: signal convergence, output the updated node workload as the improved dispatch, and output the total dispatch cost as the improved dispatch cost, otherwise: continue by replacing the total dispatch cost with the iterative dispatch cost.

3. The method of any one of claims 1-2, wherein generating the initial random assignment comprises: generating a candidate set having scene indices ranging from 1 to K-1, where K represents the total number of scenes in the scene sequence to be assigned to the M computing nodes; Generate a dispatch set whose first element is 0; updating the dispatch set according to a random selection using the random seed selection to generate an updated dispatch set; sorting the updated dispatch set in ascending order to generate a sorted dispatch set; as well as Generating the initial random assignment according to the sorted set of assignments, wherein updating the set of assignments comprises: For t = 1 to M-1: Select a random integer p between 0 and Kt-1; identifying the pth element in the candidate set and appending it to the dispatch set; Remove the pth element from the candidate set; and The candidate set is sorted in ascending order.

4. The method of claim 3 , wherein generating the initial random assignment from the sorted set of assignments comprises: All scenarios with indices between values ​​equal to or greater than the mth element in the sorted dispatch set but less than the m+1th element in the sorted dispatch set are dispatched to compute node m.

5. The method of any one of claims 1-2, wherein calculating the cost metric for all computing nodes based on the scenario-to-node assignment of each computing node comprises: For each computing node, calculating a total number of frames assigned to the computing node based on the scene-to-node assignment; as well as Calculate the standard deviation of the total number of frames dispatched to each compute node.

6. The method of any one of claims 1-2, wherein removing a scenario from the node workload and assigning it to the workload of computing node m+1 comprises: The last scene scheduled for encoding at compute node m is identified and assigned as the first scene to be scheduled for encoding at compute node m+1.

7. The method of any one of claims 1-2, wherein adding a scenario to the node workload taken from the workload of computing node m-1 comprises: The last scene scheduled for encoding at compute node m-1 is identified and assigned as the first scene to be scheduled for encoding at compute node m.

8. The method of any one of claims 1-2, wherein updating the optimal dispatch cost and the optimal output dispatch comprises: For a first dispatch iteration, setting the improved dispatch as the optimal output dispatch and setting the improved dispatch cost to the optimal dispatch cost; as well as For subsequent dispatch iterations, the improved dispatch cost is compared to the optimal dispatch cost; and if the optimal dispatch cost is greater than the improved dispatch cost, the improved dispatch is selected as the optimal output dispatch and the improved dispatch cost is selected as the optimal dispatch cost.

9. The method according to any one of claims 1 to 2, further comprising: For a computing node among the M computing nodes: accessing a sequence of high dynamic range (HDR) frames and a corresponding sequence of standard dynamic range (SDR) frames of a scene assigned to the compute node based on an optimal output assignment of the scene sequence to the compute node; as well as An output bitstream is generated for the scenario assigned to the compute node.

10. The method of claim 9, wherein generating an output bitstream further comprises: generating a scene-based forward shaping function based on the HDR frame sequence and the SDR frame sequence; mapping the HDR frame sequence to a shaped SDR frame sequence based on the scene-based forward shaping function; generating a coded bitstream by compressing the shaped SDR frame sequence; generating a scene-based backward shaping function based on the shaped SDR frame sequence, the HDR frame sequence, and the scene-based forward shaping function; generating metadata based on parameters of the scene-based backward shaping function; as well as An output bitstream including the encoded bitstream and the metadata is output.

11. A computer-readable storage medium storing computer-executable instructions for executing the method according to any one of claims 1 to 10 using one or more processors.

12. An apparatus comprising a processor and configured to perform any one of the methods of claims 1-10.