Methods, apparatuses, and media to encode images into video signals or render images

By injecting luminance-related noise and performing spatial downsampling during high dynamic range image encoding, the stripe artifact problem was solved, achieving efficient video signal encoding and improved rendering quality.

CN116034394BActive Publication Date: 2026-05-29DOLBY LABORATORIES LICENSING CORP

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2021-08-05
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively mitigate striping artifacts in video signals, especially when encoding high dynamic range images into standard dynamic range signals, leading to a decline in rendering quality.

Method used

By generating a forward shaping map, a high dynamic range image is mapped to a low dynamic range image. After spatial downsampling, light-related noise is injected to generate an image with embedded noise, which is then encoded into a video signal. The receiving device renders the image as a low dynamic range image through a backward shaping map.

Benefits of technology

It effectively reduces stripe artifacts, improves the rendering quality of video on low dynamic range devices, and reduces computational costs and resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116034394B_ABST
    Figure CN116034394B_ABST
Patent Text Reader

Abstract

A forward reshaping mapping is generated to map a source image to a corresponding forward reshaped image of a lower dynamic range. The source image is spatially downsampled to generate an image, noise is injected into the resized image to generate an injected noise image. The forward reshaping mapping is applied to map the injected noise image to generate the lower dynamic range embedded noise image. The video signal is encoded with the embedded noise image and transmitted to a recipient device for the recipient device to render a display image generated from the embedded noise image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 061,937 and European Patent Application No. 20189859.0, both filed on August 6, 2020, each of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure generally relates to image processing operations, and more particularly to video codecs. Specifically, this disclosure relates to methods, image processing apparatus, and computer-readable storage media for encoding images into video signals or rendering images. Background Technology

[0004] As used herein, the term "dynamic range (DR)" can refer to the ability of the human visual system (HVS) to perceive a range of intensity (e.g., luminance, luma) in an image, such as from the darkest black (deep) to the brightest white (highlight). In this sense, DR relates to the intensity of a "scene-referred" image. DR can also refer to the ability of a display device to fully or approximately render a specific breadth of intensity. In this sense, DR relates to the intensity of a "display-referred" image. Unless a particular meaning is explicitly specified to have a specific connotation at any point in the description herein, it should be inferred that the terms can be used in either sense (e.g., interchangeably).

[0005] As used herein, the term "high dynamic range (HDR)" refers to a DR width of approximately 14 to 15 or more orders of magnitude across the human visual system (HVS). In practice, the DR, on which humans can simultaneously perceive a wide range of intensity, may be slightly truncated relative to HDR. As used herein, the terms "enhanced dynamic range (EDR)" or "visual dynamic range (VDR)" may be associated, individually or interchangeably, with this type of DR: a DR that can be perceived within a scene or image by the human visual system (HVS), including eye movements, allowing for some light-adaptive variations across that scene or image. As used herein, EDR may refer to a DR spanning 5 to 6 orders of magnitude. While perhaps slightly narrower than HDR relative to a reference real-world scene, EDR indicates a wide DR width and can also be referred to as HDR.

[0006] In practice, an image comprises one or more color components in a color space (e.g., luminance Y and chroma Cb and Cr), where each color component is represented by n bits per pixel with precision (e.g., n = 8). Using non-linear luminance coding (e.g., gamma coding), images where n ≤ 8 (e.g., a color 24-bit JPEG image) are considered images with standard dynamic range, while images where n > 8 can be considered images with enhanced dynamic range.

[0007] A reference electro-optical transfer function (EOTF) for a given display characterizes the relationship between the color values ​​(e.g., luminance) of the input video signal and the color values ​​(e.g., screen luminance) of the output screen generated by the display. For example, ITU-R BT.1886, “Reference electro-optical transfer function for flat panel displays used in HDTV studio production” (March 2011), defines a reference EOTF for flat panel displays, the contents of which are incorporated herein by reference in their entirety. Given a video stream, information about its EOTF can be embedded in the bitstream as (image) metadata. The term “metadata” in this document refers to any auxiliary information transmitted as part of the encoded bitstream and used to assist the decoder in rendering the decoded image. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters as described herein.

[0008] As used herein, the term "PQ" refers to Perceived Luminance Amplitude Quantization. The human visual system responds to increasing light levels in a significantly non-linear manner. The human ability to perceive stimuli is influenced by factors such as the luminance of the stimulus, the size of the stimulus, the spatial frequency constituting the stimulus, and the luminance level to which the eye has adapted at a particular moment of viewing the stimulus. In some embodiments, the perceived quantizer function maps linear input gray levels to output gray levels that better match the contrast sensitivity threshold in the human visual system. An example PQ mapping function is described in SMPTE ST 2084:2014 "High Dynamic Range EOTF of Mastering Reference Displays" (hereinafter "SMPTE"), which is incorporated herein by reference in its entirety, wherein, given a fixed stimulus size, for each luminance level (e.g., stimulus level, etc.), the minimum visible contrast step size at that luminance level is selected based on the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model).

[0009] Supports 200 to 1,000 cd / m³ 2 A display with a brightness of nits represents a lower dynamic range (LDR) associated with EDR (or HDR), also known as standard dynamic range (SDR). EDR content can be displayed on an EDR display that supports a higher dynamic range (e.g., from 1,000 nits to 5,000 nits or higher). Such a display can be defined using an alternative EOTF that supports high brightness capabilities (e.g., 0 to 10,000 nits or higher). Examples of such EOTFs are defined in SMPTE 2084 and Rec.ITU-RBT.2100, “Image parameter values ​​for high dynamic range television for use in production and international programme exchange” (06 / 2017). As the inventors understand herein, improved techniques are desired for synthesizing video content data that can be used to transmit media content to and support the display capabilities of various SDR and HDR display devices, including mobile devices.

[0010] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise instructed, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise instructed, any problem identified with respect to one or more methods should not be considered to have been identified in any prior art based on this section. Summary of the Invention

[0011] According to a first aspect of this disclosure, a method for encoding an image into a video signal is provided, the method comprising: at a first stage, performing the following operations: generating a forward shaping map to map a source image of a first dynamic range to a corresponding forward-shaped image of a second dynamic range lower than the first dynamic range; and at a second stage, performing the following operations: generating an image of a first dynamic range and a first spatial resolution by spatially downsampling the source image of the first dynamic range; calculating a light intensity-dependent noise intensity using the forward shaping map; injecting noise having a light intensity-dependent noise intensity into the image of the first dynamic range and the first spatial resolution to generate an image of injected noise of the first dynamic range and the first spatial resolution; applying the forward shaping map to map the image of injected noise of the first dynamic range and the first spatial resolution to generate an image of embedded noise of the second dynamic range and the first spatial resolution; and encoding the image of embedded noise of the second dynamic range and the first spatial resolution into a video signal.

[0012] According to a second aspect of this disclosure, a method for rendering an image is provided, the method comprising: receiving, by a receiving device, a video signal generated by a method for encoding an image into a video signal according to this disclosure, and image metadata including a backward-shaping map; decoding, by the receiving device, an image with embedded noise of a second dynamic range and a first spatial resolution from the video signal; applying, by the receiving device, a backward-shaping map to the image with embedded noise of the second dynamic range and the first spatial resolution to generate a backward-shaped image of the first dynamic range; and rendering, by the receiving device, a display image representing the backward-shaped image of the first dynamic range.

[0013] According to a third aspect of this disclosure, an image processing apparatus is provided, the image processing apparatus including a processor and configured to perform any one of the methods of this disclosure for encoding an image into a video signal or for rendering an image.

[0014] It also provides other aspects. Attached Figure Description

[0015] Embodiments of the invention are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings:

[0016] Figure 1 An example process of a video transmission pipeline is described;

[0017] Figures 2A to 2F The diagram illustrates an example system configuration for adaptive video streaming; Figure 2G The illustration shows an example of generating an input video clip from an input video signal; Figure 2H The diagram illustrates an example system configuration or architecture implemented by cluster nodes;

[0018] Figure 3A The illustration shows an example of a forward lookup table (FLUT) for luminance. Figure 3B The illustration shows an example of codeword addition in the codeword bin determined by FLUT; Figure 3C The illustration shows an example of smoothed noise intensity; Figure 3D The illustration shows an example frequency domain pattern or block used to generate a noisy image;

[0019] Figures 4A to 4C The example process flow is illustrated; and

[0020] Figure 5 A simplified block diagram of an example hardware platform is shown, on which a computer or computing device as described herein can be implemented. Detailed Implementation

[0021] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent that this disclosure may be practiced without these specific details. In other instances, well-known structures and devices have not been described in detail in order to avoid unnecessarily obscuring, obscuring, or confusing this disclosure.

[0022] Overview

[0023] Techniques as described herein can be implemented to support or provide an adaptive streaming framework in which striped video content can be adaptively streamed to end-user devices with relatively moderate decoding and / or display capabilities. The adaptive streaming framework can implement a multi-resolution / bitrate ladder based on video coding, wherein multiple spatial resolution video signals can be generated for adaptive streaming to the end-user device. Example end-user devices may include, but are not limited to, mobile devices with 8-bit or 10-bit video codecs (such as Advanced Video Coding (AVC), HEVC, AV1, etc.).

[0024] While 8-bit video signals are typically very susceptible to striping artifacts without the implementation of techniques as described herein, with techniques as described herein, noise such as film grain noise can be injected into the HDR video signal used to generate the video signal (including 8-bit video signals) at a brightness-dependent noise intensity to achieve effective false contouring or striping reduction.

[0025] Striped-light adaptive streaming can be implemented in various system configurations (or architectures). In one example, striped-light adaptive video streaming can be combined with segment / node-style parallel processing. In another example, striped-light adaptive video streaming can be combined with linear / real-time encoding, such as in broadcast applications.

[0026] A media streaming system or service (which may be referred to as a media streamer) can support adaptive streaming using multiple video signals at multiple spatial resolutions and / or multiple bit rates. Each of these multiple video signals depicts the same visual semantic content as the input source HDR video signal and can be generated by the media streamer from the input source HDR video signal.

[0027] A media streaming transmitter can dynamically adapt to the media streaming operation of streaming video content to a given media streaming client (including but not limited to mobile client devices) based on some or all of a variety of client-specific media streaming conditions / factors. These client-specific media streaming conditions / factors may include, but are not limited to, any of the following: network conditions or network bandwidth, available system resources, transmission latency and / or system latency, user-specific settings / selections, etc.

[0028] At a given time, a media streaming transmitter can stream one or more video segments or stream portions of a specifically selected video signal with a specific spatial resolution and / or a specific bit rate to a given media streaming client. The specific video signal may (e.g., dynamically, in real-time, during runtime, etc.) be selected from multiple video signals with multiple spatial resolutions and / or multiple bit rates depending on some or all of the real-time or non-real-time streaming conditions / factors specific to and / or common to a given media streaming client.

[0029] In some operational scenarios, such as the media streaming transmitters described herein, a system configuration comprising multiple instances of a complete media processing pipeline can be implemented. Each instance of the complete media processing pipeline can be used to generate a corresponding video signal from multiple video signals with different settings or combinations of supported spatial resolutions and / or supported bit rates. For example, the complete media processing pipeline or each instance thereof can generate forward and backward shaping maps between (input) source images and base (BL) images; inject noise into the source image to mitigate or prevent false contour artifacts; resize or spatially downsample the source image with injected noise; forward shape the resized source image with injected noise into a shaped BL image with embedded noise (e.g., striped masked, injected noise, etc.); encode / compress the BL image with injected noise into the corresponding video signal, and so on.

[0030] In some operational scenarios, media streaming transmitters, as described herein, can implement system configurations for multi-level encoding. By way of illustration and not limitation, media streaming transmitters can implement two-level system configurations (including, but not limited to, two-level single-layer backward-compatible (SLBC) encoder system configurations or architectures), which include a first level that performs the first part of the complete media processing pipeline as described above and a second level that performs the remainder of the complete media processing pipeline.

[0031] System configurations for multi-level coding can be designed to encode video content for adaptive streaming applications relatively efficiently, leveraging cloud computing resources to create multiple encoded bitstreams with various combinations of spatial resolution (or image resolution) and bitrate. The media streamer can implement a single instance of the first level and multiple instances of the second level to generate multiple video signals with different settings or combinations of spatial resolutions and / or supported bitrates in the bitrate ladder. In a two-level system configuration, forward-shaping data and film grain noise injection mechanisms can be used to encode and output multiple encoded bitstreams. Noise injection (including but not limited to film grain injection) can be used to mask or significantly reduce false contour artifacts, thus significantly improving the visual quality when rendering video content on end-user devices such as mobile phones.

[0032] The first part of the complete media processing pipeline is executed by a first encoding stage in a two-stage system configuration and includes: generating a forward shaping map and a backward shaping map between the (input) source image and the base layer (BL) image, forward passing the binary data obtained from the forward shaping map to the second stage, etc. The forward and backward shaping maps enable conversion between relatively high dynamic range (e.g., EDR, etc.) video content and relatively low dynamic range (e.g., SDR, etc.) video content. HDR video content and SDR video content can depict the same visual semantic content, despite having different dynamic ranges (e.g., different brightness ranges, etc.).

[0033] The second part of the complete media processing pipeline is executed by each of the multiple instances of the second coding level in the two-level system configuration, and includes: receiving forward binary data from the first level, resizing or spatially downsampling the source image, using the forward binary data to determine the intensity of light-dependent noise, injecting noise into the resized source image at the determined intensity of light-dependent noise to mitigate or prevent false contour artifacts, forward-shaping the resized source image with injected noise into a BL image with embedded noise, encoding / compressing the BL image with embedded noise into a corresponding video signal among multiple video signals of different settings or combinations of supported spatial resolutions and / or bit rates, etc.

[0034] In contrast to a full media processing pipeline, a two-level system configuration effectively implements a full / reduced media processing pipeline. The first part of a full media processing pipeline can be executed only once using either a two-level system configuration or the first level of a full / reduced media processing pipeline, while the second part of a full media processing pipeline can be executed or instantiated multiple times using either a two-level system configuration or the second level of a full / reduced media processing pipeline.

[0035] Therefore, in a two-stage system configuration, only the reduced pipeline (or the second stage reduced from the full media processing pipeline) is executed multiple times by the second stage (multiple instances) of the full / reduced media processing pipeline. However, each instance of the second stage of the full / reduced media processing pipeline can be combined with the first stage of the full / reduced media processing pipeline to provide functionality fully or completely equivalent to that of the full media processing pipeline (instances) for a corresponding video signal among multiple video signals. Additionally, optionally, or alternatively, multiple instances of the second stage can, but are not limited to, run independently and / or in parallel on different processing threads of one or more computing processors or on multiple different computing processors.

[0036] As a result, the cost of redundant media processing and computation can be significantly reduced in full / reduced media processing pipelines compared to deploying multiple instances of a full / reduced media processing pipeline, enabling adaptive streaming with multiple spatial resolutions and / or multiple bit rates.

[0037] The adaptive streaming described herein can be implemented using a variety of computing systems, including but not limited to any of the following: a single computing system, a combination of multiple computing systems, a geographically distributed computing system, a network of one or more computing systems, etc.

[0038] In some operational scenarios, a two-tier system configuration or a full / reduced media processing pipeline can be implemented by a cloud-based computing cluster comprising multiple cluster computing nodes, each of which can be a virtual computer launched via cloud computing services. An input source video stream can be used to generate multiple consecutive (e.g., partially overlapping, etc.) input video segments. Each of the multiple input video segments generated from the input source video stream can be assigned to a specific cluster computing node in the cloud computing cluster to produce a corresponding encoded bitstream portion or output video segment that supports different settings or combinations of spatial resolution and / or bitrate for adaptive streaming.

[0039] The adaptive streaming video signal being streamed to the receiving media streaming client may include a sequence of consecutive coded bitstream portions or a sequence of consecutive output video segments. The coded bitstream portions or output video segments in the consecutive coded bitstream portions or the sequence of consecutive output video segments (covering multiple time segments / intervals collectively representing the entire duration of the media program covered by the consecutive coded bitstream portions or the sequence of consecutive output video segments) may be selected from different video signals, specifically from different settings or combinations of spatial resolution and / or bitrate, depending on real-time or non-real-time streaming conditions / factors specific to or common to the receiving media streaming client.

[0040] Video signals can be transmitted, streamed, and / or transmitted directly or indirectly to a receiving media streaming client. For example, if a noisy BL image is matched with the applicable display capabilities of the receiving media streaming client, the noisy BL image decoded from the video signal can be directly rendered by the receiving media streaming client.

[0041] Additionally, optionally, or alternatively, the video signal may further carry image metadata, including but not limited to some or all of backshaping maps (or compositor metadata), display management (DM) metadata, etc. The receiving media streaming client may use the image metadata or backshaping maps received along with the video signal or portions thereof to synthesize images with higher dynamic range, wider color gamut, higher spatial resolution, etc., matching the applicable display capabilities of the receiving media streaming client. These synthesized (or reconstructed) images may be rendered by the receiving media streaming client, rather than by the BL image decoded from the video signal.

[0042] The example embodiments described herein relate to encoding video data. A forward shaping map is generated to map a source image of a first dynamic range to a corresponding forward-shaped image of a second dynamic range lower than the first dynamic range. Noise is injected into the images of the first dynamic range and a first spatial resolution to generate an image with injected noise of the first dynamic range and the first spatial resolution. The images of the first dynamic range and the first spatial resolution are generated by spatially downsampling the source image of the first dynamic range. The forward shaping map is applied to map the image with injected noise of the first dynamic range and the first spatial resolution to generate an image with embedded noise of the second dynamic range and the first spatial resolution. A video signal encoded with the image with embedded noise of the second dynamic range and the first spatial resolution is transmitted to a receiving device for the receiving device to render a display image generated from the image with embedded noise.

[0043] The example embodiments described herein relate to decoding video data. A video signal is received, generated by an upstream encoder and encoded with an image containing embedded noise at a second dynamic range and a first spatial resolution. The second dynamic range is lower than the first dynamic range. The image containing embedded noise at the second dynamic range and the first spatial resolution has been generated by applying a forward shaping map to an image with injected noise at the first dynamic range and the first spatial resolution via the upstream encoder. The image with injected noise at the first dynamic range and the first spatial resolution has been generated by injecting noise into the image at the first dynamic range and the first spatial resolution via the upstream encoder. The image at the first dynamic range and the first spatial resolution is generated by spatially downsampling a source image at the first dynamic range. A display image is generated from the image containing embedded noise at the second dynamic range and the first spatial resolution. The display image is rendered on an image display.

[0044] Example video transmission and processing pipeline

[0045] Figure 1 An example process of a video transmission pipeline (100) is depicted, illustrating the various stages from video capture to video content display. An image generation block (105) is used to capture or generate a sequence of video frames (102). The video frames (102) can be captured digitally (e.g., by a digital camera, etc.) or generated by a computer (e.g., using computer animation, etc.) to provide video data (107). Additionally, optionally, or alternatively, the video frames (102) can be captured on film by a film camera. The film can be converted to a digital format to provide video data (107). At the production stage (110), the video data (107) is edited to provide a video production stream (112).

[0046] The video data from the production stream (112) is then provided to the processor for post-production editing (115). Post-production editing (115) may include adjusting or modifying the color or brightness in specific areas of the image to enhance image quality or achieve a specific look for the image according to the creative intent of the video creator. This is sometimes referred to as "colortiming" or "color grading". Other editing may be performed at post-production editing (115) (e.g., scene selection and sorting, manual and / or automatic scene cut information generation, image cropping, adding computer-generated visual effects, etc.) to generate post-production versions of HDR images and content-mapped versions of SDR images.

[0047] The post-production version of the HDR image and the content-mapped version of the SDR image depict the same set of visual scenes or semantic content. The content-mapped version of the SDR image can be obtained from the post-production version of the HDR image through a combination of manual, automated content mapping and / or color grading, or manual and automated image processing operations. In some operational scenarios, during post-production editing (115), one or both of the post-production version of the HDR image and the content-mapped version of the SDR image are viewed and color-graded, for example, by a colorist on HDR and SDR reference monitors that respectively support (e.g., guided rendering of HDR and SDR images).

[0048] As an example and not a limitation, the HDR image (117-1) may represent a post-production version of an HDR image, and the SDR image (117) may represent a content-mapped version of an SDR image. The encoding block (120) receives the HDR image (117-1) and the SDR image (117) from the post-production editor (115) and forward-shapes the HDR image (117-1) into a (forward-shaped) SDR image. The forward-shaped SDR image may approximate the SDR image (117) from an automatic or manual content-mapping (and / or color grading) operation.

[0049] The coding block (120) can implement some or all of the stripe reduction and adaptive streaming operations as described herein to generate multiple target versions of the stripe-reduced forward-shaping SDR image for various different combinations of spatial resolution and / or bit rate.

[0050] In some operational scenarios, each of some or all of a plurality of target versions of a strip-reduced forward-shaping SDR image can be compressed / encoded into a coded bitstream (122) by a coding block (120) in a linear video coding mode. The coded bitstream (122) includes the SDR image in the target version (e.g., the forward-shaping strip-reduced SDR image, etc.). Additionally, optionally, or alternatively, the coded bitstream (122) may include image metadata (e.g., backward-shaping metadata, etc.), which includes operational parameters to be used by the receiving device of the coded bitstream (122) to reconstruct an HDR image from the forward-shaping strip-reduced SDR image in the target version.

[0051] In some operational scenarios, each of some or all of the multiple target versions of a strip-reduced forward-shaping SDR image can be compressed / encoded by a coding block (120) in a segmented video coding mode into a sequence of consecutive video segments (122-1). Each video segment in the sequence of consecutive video segments (122-1) constituting some or all of the multiple target versions of the strip-reduced forward-shaping SDR image can be an independently accessible video streaming file (or a set of video streaming files including a main video file and zero or more accompanying files), providing video content within time sub-intervals (e.g., 10 seconds, 20 seconds, etc.) within the entire time interval covered by the target version. Additionally, optionally, or alternatively, the video segments may include image metadata (e.g., backward-shaping metadata, etc.) including operational parameters to be used by the receiving device of the video segment to reconstruct an HDR image from the forward-shaping strip-reduced SDR image encoded in the video segment.

[0052] The coding block (120) can be implemented at least in part with an audio encoder and a video encoder, such as those defined by ATSC, DVB, DVD, Blu-ray and other transport formats, to generate some or all of a plurality of target versions of a strip-reduced forward-shaped SDR image and to encode each of the plurality of target versions of the strip-reduced forward-shaped SDR image into a corresponding encoded bitstream (e.g., 122 etc.) and / or a corresponding sequence of consecutive video segments (e.g., 122-1 etc.).

[0053] In some operational scenarios, encoded bitstreams (e.g., 122, etc.) or sequences of video segments (e.g., 122-1, etc.) can represent backward-compatible video signals (e.g., 8-bit SDR video signals, 10-bit SDR video signals, etc.) for various SDR display devices (e.g., SDR displays, etc.). In a non-limiting example, a video signal encoded with an SDR image reduced by forward-shaping striping can be a single-layer backward-compatible video signal. Here, "single-layer backward-compatible video signal" can refer to a video signal carrying an SDR image optimized or color-graded for SDR displays in a single signal layer. An example of single-layer video encoding operation is described in U.S. Patent Application Publication No. 2019 / 0110054, "Encoding and decoding reversible production-quality single-layer video signals," the entire contents of which are incorporated herein by reference as fully set forth herein.

[0054] Some or all of the operational parameters in the image metadata provided along with the forward-strip-reduced SDR image encoded in the video signal can be decoded by the receiving device of the video signal and used in image processing operations (e.g., prediction operations, backward-stripping operations, inverse tone mapping operations, etc.) to generate a reconstructed image with a higher dynamic range than that represented by the forward-strip-reduced SDR image.

[0055] In some operational scenarios, the decoded image represents an SDR image that has undergone forward shaping and stripe reduction by an upstream video encoder (e.g., with coding blocks (120), etc.), which is generated by forward shaping a post-processed HDR image (e.g., possibly spatially downsampled, etc.) of a post-processed version of the HDR image (117-1) to approximate a post-processed SDR image (e.g., possibly spatially downsampled, etc.) of a content-mapped version of the SDR image (117). The reconstructed image generated from the decoded image using operational parameters in the image metadata transmitted in the video signal represents an HDR image that approximates (e.g., possibly spatially downsampled, etc.) the post-processed HDR image in the post-processed version of the HDR image (117-1) on the encoder side.

[0056] Example reshaping operations are described in U.S. Patent 10,080,026, "Signal reshaping approximation," to which the entire contents, as fully set forth herein, are incorporated by reference.

[0057] Additionally, optionally, or alternatively, the video signal is encoded with additional image metadata, including but not limited to display management (DM) metadata, which a downstream decoder can use to perform display management operations on the decoded or back-shaped image to generate a display image optimized for rendering on a target display.

[0058] The video signal, in the form of an encoded bitstream (e.g., 122, etc.) or a sequence of video segments (e.g., 122-1, etc.), is then transmitted downstream to a receiver such as a mobile device, decoding and playback device, media source device, media streaming client device, television (e.g., smart TV, etc.), set-top box, cinema, etc. At the receiver (or downstream device), the decoding block (130) decodes the video signal to generate a decoded image 182, which may be the same as the image encoded into the video signal by the encoding block (120) (e.g., a forward-shaping stripe-reduced SDR image, etc.), and is subject to quantization errors generated in the compression performed by the encoding block (120) and the decompression performed by the decoding block (130).

[0059] In an operational scenario where the receiver operates together with (or is attached to or operatively linked to) a target display 140 that supports rendering the decoded image (182), the decoding block (130) can decode the image (182) from the encoded bitstream (122) (e.g., a single layer in the encoded bitstream, etc.) and use the decoded image (182) (e.g., a forward-shaped SDR image, etc.) directly or indirectly for rendering on the target display (140).

[0060] In some operating scenarios, the target display (140) has similar characteristics to the SDR reference display (125), and the decoded image (182) is a forward-shaping stripe-reduced SDR image that can be directly viewed on the target display (140).

[0061] In some embodiments, the receiver operates in conjunction with (or is attached to or operatively linked to) a target display having display capabilities different from those of a reference display to which the decoded image (182) is optimized. Some or all of the operating parameters in the image metadata (or synthesizer metadata) can be used to synthesize or reconstruct an image from the decoded image (182) optimized for the target display.

[0062] For example, the receiver can operate with an HDR target display 140-1 that supports a higher dynamic range (e.g., 100 nits, 200 nits, 300 nits, 500 nits, 1,000 nits, 4,000 nits, 10,000 nits or more, etc.) compared to the decoded image (182). The receiver can extract image metadata from the video signal (e.g., the metadata containers therein) and use the operating parameters in the image metadata (or synthesizer metadata) to synthesize or reconstruct image 132-1 from the decoded image (182) (such as a forward-shaping stripe-reduced SDR image).

[0063] In some operational scenarios, the reconstructed image (132-1) refers to a reconstructed HDR image optimized for viewing on an HDR display that is the same as or comparable to the HDR target display used by the receiver (e.g., a reference display, etc.). The receiver can directly use the reconstructed image (132-1) for rendering on the HDR target display.

[0064] In some operational scenarios, the reconstructed image (132-1) refers to a reconstructed HDR image optimized for viewing on an HDR display (e.g., a reference display, etc.) that is different from the HDR target display (140-1) operated in conjunction with the receiver. A display management block (e.g., 135-1, etc.) (which may be located in the receiver, in the HDR target display (140-1), or in a separate device) further adjusts the reconstructed image (132-1) according to the characteristics of the HDR target display (140-1) by generating a display mapping signal (137-1) suitable for the characteristics of the HDR target display (140-1). The display image or the adjusted reconstructed image can be rendered on the HDR target display (140-1).

[0065] Adaptive video streaming with stripe reduction

[0066] The techniques described herein can be used, for example, to support adaptive streaming video images under various combinations of spatial resolution and bit rate in cloud computing environments. Simultaneously, relatively efficient false contour (or striping) mitigation is implemented to mask false contours or striping artifacts in the video image.

[0067] These technologies enable media streamers to maintain relatively high (e.g., as good as possible) stripe mitigation or masking capabilities in the presence of noise or film grain injection, and reduce computational costs and disk space usage for constructing bitrate ladders that generate video images with various combinations of spatial resolution and bitrate.

[0068] Adaptive video streaming can be implemented to generate, encode, and / or stream target video content with different spatial resolutions and bit rates, thereby adapting to time-varying or dynamically changing network conditions / bandwidths, and providing relatively smooth video playback of the streamed video content under these different network conditions / bandwidths.

[0069] As described herein, bitrate tiering can be implemented in media streaming providers to generate some or all of the target video content with different combinations of spatial resolution and / or bitrate. A media streaming provider may be referred to, but is not limited to, a media streaming server / service, a video streaming server / service, a media / video content provider, a media or video encoder, a media broadcasting system, an upstream device, etc. In some operational scenarios, media streaming providers can be deployed or accessed in a cloud computing environment. The adaptive streaming architecture described herein can be implemented with relatively high efficiency and relatively low (e.g., cloud-based, etc.) compute resource utilization, thereby significantly reducing the ongoing operational costs of streaming media / video content to end-user devices. Example compute resource utilization may include, but is not limited to, any utilization related to (e.g., cloud-based, etc.) CPU time, disk space, etc.

[0070] In some operational scenarios, a separate, complete media processing pipeline can be deployed for each setting or combination of spatial resolution and / or bitrate in the bitrate hierarchy. This can easily incur significant computational costs due to relatively high CPU utilization, relatively high disk space utilization, and so on.

[0071] The media streaming transmitters described herein implement multi-level coding, such as a two-level video coding pipeline, and utilize this multi-level coding to perform cost-effective computation to improve video quality by striping the video signal through bitrate ladders supporting multiple spatial resolutions and / or multiple bitrates. In some operational scenarios, the media streaming transmitter can be implemented using computing resources provided or leased by shared or shared cloud-based systems or services in a cloud computing environment.

[0072] For illustrative purposes only, a media streamer receives a mezzanine video content item from an input HDR video source, comprising an input or source HDR image with a relatively high input dynamic range. The mezzanine video content item can be, but is not limited to, TV programs, movies, video recordings of events, etc. In some operational scenarios, the input HDR video source can be implemented or include Figure 1 The post-production block (115) in the end-to-end video transmission pipeline is systematically generated and provided to the media streaming receiver. The media streaming receiver (or the video coding system implemented or operated with the media streaming receiver) can... Figure 1 Encoding blocks (120) are implemented in the end-to-end video transmission pipeline to encode or generate multiple (e.g., positive integers M greater than one (1) etc.) video signals under different combinations of spatial resolution and / or different bit rates, such as encoded bit streams or output video segment sequences.

[0073] A media streamer can select (e.g., time-varying, fluctuating, etc.) video frames or video images of the best possible quality that available network bandwidth / conditions can support from some or all of the bitstreams or output video segments to be streamed in real time to an end-user device for (e.g., real-time, near real-time, etc.) playback or image rendering.

[0074] The encoded bitstream or output video clip described herein may include a base layer encoded with a relatively low dynamic range forward-shaping (SDR) image (e.g., 8-bit, image data, etc.), which is generated by the coding block (120) through forward-shaping of a source HDR image (e.g., a spatially downsampled film grain-injected version thereof).

[0075] In some operational scenarios, target stripe (or pseudo-contour) reduced video content (e.g., a target version of an SDR image, etc.) can be streamed from a media streamer to an end-user device, including but not limited to mobile devices with relatively small display sizes and / or relatively dim displays operated with (e.g., only) 8-bit video decompression or decoding modules. Without implementing the techniques described herein, 8-bit video systems may lack sufficient image processing capabilities to avoid or improve pseudo-contour artifacts in image rendering operations.

[0076] In some operating scenarios, film grain noise can be injected into the forward shaping path on the encoder side. Additionally, the inverse tone mapping curve can be adjusted in the corresponding backward shaping path to reduce or mitigate banding artifacts. These operations performed in both the forward and backward shaping paths can incur relatively high computational costs. Furthermore, these operations may sacrifice specular contrast to reduce banding artifacts. Therefore, these operations may be relatively effective in operating scenarios where images are rendered using large, bright displays (rather than mobile devices with limited display capabilities). Even for large, bright displays, in many cases, due to the large display size and high visibility of false outlines or banding artifacts on bright displays, they cannot be adequately or completely removed without implementing other methods as described herein.

[0077] Given that mobile device displays are much smaller and darker than typical non-mobile device displays such as TVs, mobile devices tend to generate generally fewer false contours or banding / compression artifacts compared to the larger and brighter non-mobile device displays. Additionally, optionally, or alternatively, mobile devices tend to pay less attention to injected film grain noise compared to non-mobile device displays.

[0078] Techniques as described herein can be implemented to utilize these (e.g., unique, dissimilar, etc.) characteristics of mobile device displays. Under these techniques, relatively efficient and effective stripe reduction methods can be implemented to simply modulate the film grain intensity according to or covariantly with the slope of the forward shaping function (e.g., brightness, etc.) until a relatively strong film grain intensity is achieved.

[0079] In order to achieve relatively high-quality stripe masking or mitigation of image content encoded in a base layer (e.g., 8-bit, etc.), film grain parameters such as Discrete Cosine Transform (DCT) block size, DCT frequency, minimum / maximum noise intensity, etc., can be tuned by the coding block (120) for different settings or combinations of spatial resolution and bit rate in the bit rate ladder.

[0080] Complete media processing pipeline

[0081] Figure 2A The illustration depicts an example system configuration implementing a bitrate ladder for adaptive video content streaming using multiple instances of a complete media processing pipeline. In this system configuration, a complete encoding instance is created for each setting or combination of spatial resolution and bitrate in the bitrate ladder to achieve relatively high performance. For example, M complete encoding instances (or M instances of the complete media processing pipeline) can be created to generate bitstreams or sequences of video clips for M spatial resolution and bitrate settings or combinations in the bitrate ladder on the server side (e.g., encoder side, adaptive streaming service side, etc.).

[0082] Each of the M complete encoding instances represents a complete pipeline, wherein: spatial downsampling is performed on (e.g., each, etc.) the input HDR image to generate a corresponding resized image (denoted as "HDR1", ... "HDR M"); content mapping (denoted as "CM") is performed on the resized HDR image to generate a corresponding SDR image, the SDR image depicting the same visual semantic content as the HDR image, but with a reduced dynamic range and possibly reduced spatial resolution; the resized HDR image and SDR image can be used to generate a forward shaping map (denoted as "compute forward function coefficients"); the forward shaping map is used to forward shape the resized HDR image into a shaped SDR image; the resized HDR image and shaped SDR image are used to generate a backward shaping map, the backward shaping map being part of the image metadata (denoted as "Rpu 1", ... "Rpu One of the “M” is provided to the receiving device and used by the receiving device to perform backward shaping or reconstruction on the HDR image, which is approximately the resized HDR image; the resized HDR image is injected with noise (referred to as “film grain injection”); the resized HDR image with injected noise is forward shaped (referred to as “perform forward shaping”) to generate an SDR image with embedded noise; the SDR image with embedded noise is encoded (referred to as “video compression”) into a video signal or output video segment sequence in the BL image data layer (referred to as “BL 1”, … “BL M”); and so on.

[0083] While this implementation of bitrate laddering can generate relatively high-quality bitstreams or video clips, its computational efficiency can be relatively low because the encoded bitstream or video clip sequence for each setup or combination in the bitrate ladder is generated by running multiple complete independent encoder-side processing pipelines, such as... Figure 2A As shown in the diagram.

[0084] In this system configuration, image metadata, such as the backward-shaping metadata in each different setting or combination of spatial resolution and bitrate in the bitrate ladder, is different and generated separately relative to other settings or combinations in the bitrate ladder. This is because the forward-shaping SDR images in different settings or combinations have different spatial dimensions through the spatial downsampling process and different injected film grain. Additionally, optionally, or alternatively, in operational scenarios where multi-level lossy video compression / encoding is used to achieve relatively high video quality and to avoid the expensive writing of image data to the BL (Browser Blade), each encoding instance may have to be run twice, resulting in relatively high cost-inefficiency.

[0085] Downsampling after the complete media processing pipeline

[0086] Figure 2B and Figure 2C The diagram illustrates an example two-stage system configuration that implements a bitrate ladder for adaptive video content streaming using downsampling after the full media processing pipeline. In this system configuration, to achieve relatively high computational efficiency, such as... Figure 2B As illustrated, both the first and second stages (which constitute the complete encoding processing pipeline) are performed at the highest spatial resolution (the same or equivalent spatial resolution as the source HDR image) to obtain the highest spatial resolution uncompressed BL image (e.g., SDR image, etc.) and the corresponding image metadata of the BL image.

[0087] In such Figure 2B In the illustrated complete pipeline, (e.g., each) input HDR image is content-mapped (referred to as "CM") to generate a corresponding SDR image, which depicts the same visual semantic content as the HDR image but has a reduced dynamic range and potentially reduced spatial resolution; the HDR image and SDR image can be used to generate a forward-shaping map (referred to as "compute forward function coefficients"); the forward-shaping map is used to forward-shape the HDR image into a shaped SDR image; the HDR image and the shaped SDR image are used to generate a backward-shaping map, which can be provided to the receiving device as image metadata (referred to as "Rpu") and used by the receiving device to backward-shape or reconstruct an HDR image that approximates the HDR image; the HDR image is injected with noise (referred to as "film grain injection"); the noise-injected HDR image is forward-shaped (referred to as "perform forward shaping") to generate a noise-embedded SDR image or a maximum resolution uncompressed BL image; and so on.

[0088] Then, in Figure 2C In the second stage illustrated, for each setting or combination of spatial resolution and bit rate in the bit rate ladder, the BL image with the highest spatial resolution is spatially downsampled to different spatial resolutions (denoted as "resized BL 1", ... "resized BL M"). Then, depending on the bit rate setting or combination, the downsampled BL images with different spatial resolutions are compressed (denoted as "video compression") or encoded into different video signals (denoted as "BL1", ... "BL M") or different sequences of output video segments.

[0089] Therefore, in such Figure 2B and Figure 2C In the illustrated system configuration, for each combination of spatial resolution and bitrate other than the highest quality combination in the bitrate ladder, spatial downsampling and video compression operations can be performed to obtain the corresponding bitstream or video segment sequence. In this configuration or architecture, from... Figure 2BThe image metadata generated by the complete pipeline can be reused in each spatially resized (or spatially downsampled) bitstream or video clip sequence.

[0090] Despite significant computational savings, the system configuration may or may not allow optimization of film grain parameters for different spatial resolutions and bit rates, as the film grain parameters are identical across all different spatial resolutions and bit rates. Because the injected film grains in the injected image content are downsampled along with the (master) image content during spatial downsampling operations, mid-spatial-frequency film grains generated from the highest spatial resolution image can be low-pass filtered during these spatial downsampling operations. As a result, while the system configuration allows for the reuse of image metadata generated from the highest spatial resolution image in other settings or combinations of spatial resolution and bit rate within a bit rate ladder, the reduced sharpness of the film grains in the spatially downsampled images may make stripe reduction in these downsampled images either sufficiently or insufficiently effective.

[0091] This system configuration may also result in high disk space usage, because of factors such as... Figure 2B The first-level BL image shown in the diagram needs to be written to disk space, so that... Figure 2C Each of the multiple instances in the illustrated second level can receive a BL image as input for corresponding coding operations such as spatial downsampling and video compression. Alternatively, to avoid input from, for example, Figure 2B The first-level BL image, as illustrated, is written to disk space, and the operations in the first level can be repeated. In fact, as... Figure 2B and Figure 2C The system configuration or architecture illustrated can be transformed or degenerated into Figure 2A The system configuration includes multiple complete processing pipelines to generate various settings or combinations of spatial resolution and bitrate in a bitrate ladder. This problem can be exacerbated if multi-level video compression is to be supported. Multi-level video compression can be better performed by utilizing BL images available on disk space, rather than repeatedly running the first level.

[0092] Full / Reduced Media Processing Pipeline

[0093] Figure 2D and Figure 2E The illustration depicts an example system configuration or architecture with a two-stage full and reduced pipeline, featuring a single instance of the first stage and multiple instances of the second stage in the full / reduced media processing pipeline. This system configuration can be used to optimize film grain settings in each setting or combination of spatial resolution and bitrate in a bitrate ladder, while simultaneously saving computational costs and disk space.

[0094] In such Figure 2D The first stage illustrated involves performing content mapping (denoted as "CM") on (e.g., each, etc.) the input HDR image to generate a corresponding SDR image, which depicts the same visual semantic content as the HDR image but has a reduced dynamic range and potentially reduced spatial resolution; the HDR image and SDR image can be used to generate a forward shaping map (denoted as "compute forward function coefficients"); the forward shaping map is used to forward shape the HDR image into a shaped SDR image; the HDR image and the shaped SDR image are used to generate a backward shaping map, which can be provided as image metadata to the receiving device and used by the receiving device to backward shape or reconstruct an HDR image that approximates the HDR image; and so on.

[0095] In such Figure 2D The first stage, as illustrated, calculates and generates the operation parameters or coefficients for a specified forward shaping (or forward shaping function) and / or outputs them to a binary file (represented as "forward binary" or FB). Additionally, it calculates and generates the operation parameters or coefficients for a specified backward shaping (or backward shaping function) and / or outputs them as image metadata (represented as "Rpu" or "rpu") to be encoded or formatted into (multiple) bitstreams or (multiple) video segments according to the encoding syntax specification.

[0096] In such Figure 2D In the first stage illustrated, although operational parameters for forward shaping are calculated or generated, it is not necessary (e.g., in fact, etc.) to perform forward shaping based on these operational parameters to obtain the BL image for compression. Noise injection can be disabled in the first stage to avoid using noise-injected luminance codewords to calculate or generate operational parameters for forward shaping (or forward shaping functions).

[0097] In such Figure 2EIn each instance of the second level illustrated, the input HDR image is resized to each of the different settings or combinations of spatial resolutions and / or bitrates supported by the bitrate ladder (e.g., target, etc.) spatial resolutions (denoted as "resized HDR 1", ... "resized HDR M"). Using operating parameters or coefficients generated or output from the first level for forward shaping (or forward shaping function), film grain parameters can be adjusted, and film grain noise can be injected (denoted as "film grain injection") with the adjusted film grain parameters into each such spatial resolution size-resized HDR image ("resized HDR 1", ... "resized HDR M") in the bitrate ladder. Film grain parameter settings including a specific set of film grain parameters can be used to inject noise into HDR images as described herein; the noise-injected HDR images can be forward-shaped into SDR or BL images to be encoded into video signals supporting specific spatial resolutions and / or bitrates. Injecting noise into an HDR image is separate and independent from applying a forward shaping function to the HDR image to generate a shaped SDR or BL image; in some operational scenarios, information obtained from the forward shaping function can be used to calculate the brightness-related noise intensity. Additionally, optionally, or alternatively, noise injection parameters can be selected or chosen individually for each spatial resolution and / or bit rate in the second stage.

[0098] In such Figure 2E In the second stage illustrated, a forward shaping process (denoted as "perform forward shaping") can be performed on the resized, noise-injected HDR image. Figure 2D The illustrated forward shaping, specified by the operating parameters or coefficients calculated or generated from the first level, generates the corresponding BL image for compression / encoding (referred to as "video compression") into the corresponding video signal (referred to as "BL 1", ... "BL M").

[0099] In such Figure 2E The second stage, as illustrated, does not require calculation of the operational parameters used for forward and backward shaping, because these operational parameters are already present in... Figure 2D The first stage, as illustrated, is generated. Therefore, in the second stage, it is not necessary to generate the content-mapped SDR image for the purpose of generating operational parameters for forward and backward shaping, using the content-mapped SDR image as a reference SDR image. The corresponding content-mapped operations, which are potentially the most computationally intensive part of video coding, can be skipped in the second stage. As a result, film grain parameters can be optimally and relatively efficiently adjusted and applied in the second stage without incurring relatively high computational costs, such as... Figure 2E As shown in the diagram.

[0100] like Figure 2E The second level shown in the diagram is as follows: Figure 2A The diagram shows a scaled-down version of the complete coding pipeline. (See also:) Figure 2D and Figure 2E The system configuration or architecture illustrated is well-suited for implementing a two-stage video compression pipeline. This is particularly relevant in scenarios where multi-stage encoding is used to encode bitstreams or video segments, such as... Figure 2E The illustrated reduced pipeline can be run multiple times. In, for example... Figure 2E The BL images generated in the second stage, as illustrated, do not need to be written to disk space. Instead, these BL images (e.g., in BL YUV file format, etc.) can be directly output or fed into the input storage (e.g., random access memory, main memory, cache memory, etc.) of subsequent video compression modules that perform video compression of the BL images, thereby significantly reducing disk space costs in cloud computing environments.

[0101] In some operational scenarios, such as Figure 2D The first level shown in the diagram specifies the operation parameters for forward integer shaping (or forward integer function) (as in...). Figure 2D The first level (generated) is output as or written from Figure 2D Level 1 to Figure 2E The second-level forward binary file. The forward binary file may include downsampled luma or chroma shaped data. For example, in a FLUT that maps higher-bit-depth luma or chroma codewords to lower-bit-depth luma or chroma codewords, the higher-bit-depth luma or chroma codewords can be downsampled from the higher bit depth to an intermediate bit depth below the higher bit depth but above the lower bit depth to generate a downsampled FLUT that maps intermediate-bit-depth luma or chroma codewords to lower-bit-depth luma or chroma codewords. As a result, in a multi-level coding system configuration as described herein, a smaller data-size FLUT can be generated and passed from the first level to the second level. Additionally, optionally, or alternatively, the forward binary file may contain luma or chroma shaped data that is not downsampled. For example, shaped data in a multivariate regression (MMR) representation may not be downsampled.

[0102] In multi-level encoding, forward binary data is passed from level 1 to level 2.

[0103] For illustrative purposes only, a forward binary file includes a file header and forward shaping information for each frame (or each image) of multiple consecutive frames in frame order.

[0104] The file header includes a set of header parameters that can be used by a second level to perform (e.g., correct, etc.) read operations, for example, with forward-shaping information from each frame of the forward binary file.

[0105] In some operational scenarios, the set of header parameters includes some or all of the following: (forward binary) version information; the number of frames / images that provide forward shaping information in the forward binary; HDR bit depth; SDR or BL bit depth; the total number of entries in the forward lookup table (forward LUT or FLUT) specifying the luminance forward shaping for each frame; specifying the highest order of the MMR coefficients for the chrominance forward shaping for each frame; etc.

[0106] Each frame of forward-shaping information includes a luminance one-dimensional (1D) LUT (or FLUT) for forward-shaping HDR luminance channel codewords (e.g., 12-bit precision, etc.) into forward-shaped SDR luminance channel codewords (e.g., 8-bit precision, etc.) and MMR coefficients for mapping HDR luminance and chrominance codewords to codewords belonging to each forward-shaped SDR chrominance channel.

[0107] Given an HDR bit depth, the size of the forward shaping information per frame depends on either the SDR or BL bit depth. For an 8-bit BL bit depth, a 1D-LUT (or FLUT) can use one byte per entry. Therefore, the total number of entries in a 1D-LUT (or FLUT) is 2^3. 12 *1 = 4096 bytes, for example, for 12-bit precision video / image data.

[0108] Given that the highest order of MMR coefficients is 3, the floating-point precision (4 bytes) MMR coefficients used to generate forward-integer chromaticity (Cb and Cr) channel codewords can be 2 (channels) * 2^2 (up to order 3 MMR coefficients) * 4 (floating-point precision bytes) = 176 bytes.

[0109] Therefore, the forward shaping information used to forward shape an HDR image into an SDR image consumes less than 5KB per frame, which is relatively small compared to image data (e.g., luminance and chrominance codewords for a given spatial resolution).

[0110] In some operational scenarios, such as the source (or original) HDR image described in this article, the image may be 16-bit HDR bit depth.

[0111] Let F(.) be derived from... Figure 2D The first stage of generation or prediction of the raw 16-bit luminance FLUT is illustrated to forward-shape the 16-bit HDR luminance codeword into an SDR luminance codeword.

[0112] Let F'(.) be a subsampled luminance FLUT, used to forward-shape subsampled HDR luminance codewords (e.g., more quantized than 16-bit HDR luminance codewords, or with a smaller bit depth than 16-bit HDR luminance codewords, etc.) into SDR luminance codewords. The total number of entries in F'(.) (or the size of the subsampled luminance FLUT) is N.F , where N F This represents the total number of downsampled or double-sampled HDR luma codewords. In some operational scenarios where the pre-subsampled HDR luma codeword is a 16-bit codeword space, N represents the total number of downsampled or double-sampled HDR luma codewords within the downsampled / double-sampled HDR luma codewords. F ≤2 16 For example, 4096 (corresponding to a 12-bit downsampled HDR bit depth).

[0113] Let ε be the step size (or "stride"). Now, each entry has an index u = 0, 1, ... N F The FLUT entries for quadratic sampling of -1 can be obtained as follows:

[0114]

[0115] Where F′(u) represents the FLUT entry with index u that is sampled twice; Indicates the entry index is The original FLUT entry.

[0116] As an example rather than a constraint, given N F =4096, the subsampled FLUT entries with entry index u=2028 can be obtained as follows: the subsampled FLUT entries corresponding to the original FLUT entries with the entry index u=4096. Therefore, F′(2028) = F(32456). This second-sampled FLUT(F'(.)) can be written into the forward binary file as the luma forward shaping information of the applicable or corresponding frame (e.g., using f as the frame index).

[0117] Table 1 below illustrates an example program for writing each frame of luminance forward shaping data for the image / frame covered by the forward binary file.

[0118] Table 1

[0119]

[0120] As an example, and not a limitation, chroma forward shaping (e.g., mapping HDR luminance and chroma codewords to forward-shaped SDR chroma codewords) can be used in... Figure 2DThe first stage is performed by (a) MMR coefficients or (b) (e.g., singlepiece, etc.) polynomials. In various embodiments, each frame can be mapped with a per-frame chroma forward-shaping data including (e.g., individual, different, etc.) numbers of MMR coefficients to generate forward-shaping chroma (e.g., Cb and Cr channels, etc.) codewords for that frame.

[0121] In various operational scenarios, chromaticity shaping data, as described in this paper, can take on variable or fixed sizes. Chromaticity shaping data in MMR representations can include MMR coefficients of any given order. Chromaticity shaping data in polynomial representations can include zero values ​​for some polynomial coefficients at certain polynomial locations.

[0122] In some operational scenarios, the size of each frame of chroma forward-shaped data written / read from the forward binary file can remain the same for each frame. This allows relatively large amounts of chroma forward-shaped data (e.g., for multiple frames) to be written / read at once and to be relatively easy to correctly divide into multiple fixed blocks of chroma forward-shaped data per frame, thereby improving data access / update speed and efficiency, which is particularly useful in cloud-based storage or computing environments.

[0123] To ensure that the size of the chroma-shaping data per frame is fixed or constant, the chroma-shaping data per frame of a given type (e.g., any type) (whether (a) or (b) or another type) can be converted or transformed into fixed-order MMR coefficients (or MMR coefficients up to the highest MMR order). For example, the global MMR order Λ can be specified in the header of the forward binary for the chroma-shaping data of all images / frames covered by the forward binary. fix The global MMR order (which is the same for all images / frames covered in the forward binary) indicates how many MMR coefficients should be signaled for each frame. This may mean that, for some frames / images covered in the forward binary, if the chroma forward shaping data for each frame of those frames / images uses a second-highest MMR order lower than the global MMR order indicated in the header of the forward binary, then one or more of the highest MMR order coefficients in the forward binary can be set to zero (0).

[0124] As mentioned earlier, in some operational scenarios, it is used in Figure 2D The (a) MMR coefficients calculated in the first level are used to specify or define the chroma forward shaping data for each frame. This will be achieved through... Figure 2D The first-level calculation of the m-th C (Cb or Cr; alternatively, it can be represented as c) channel MMR coefficient is expressed as... The calculated MMR order is represented as λ. C .

[0125] Table 2 below illustrates an example procedure for writing each frame of chroma forward shaping data (in the form of MMR coefficients) for the images / frames covered by the forward binary file.

[0126] Table 2

[0127]

[0128] In some operational scenarios, it is used in Figure 2D The (b) polynomial (e.g., a single-segment or 1-segment 2nd order polynomial) computed in the first level is used to specify or define the chroma forward shaping data for each frame. Here, the polynomial coefficients of the polynomial as described herein can be placed in the corresponding or corresponding MMR positions.

[0129] An MMR matrix, including MMR coefficients, can be formed using the i-th normalized HDR pixel value from the luminance channel Y and the chrominance channels Cb and Cr. The i-th normalized HDR pixel value comprises the normalized HDR pixel values ​​(e.g., between [0,1)) from the HDR image / frame to be forward-shaped. The normalized HDR pixel values ​​from the chrominance channels Cb and Cr can be represented as... and The i-th normalized HDR pixel value also includes the normalized HDR pixel value in the luminance channel value at the i-th pixel in the downsampled HDR image after noise injection (if enabled) (e.g., normalized to the range [0,1), etc.). The normalized HDR pixel value in the luminance channel Y can be represented as In some operational scenarios, luminance downsampling can be performed to match the size of the luminance and chrominance planes of a YUV 420 image (or an HDR image / frame in a 420 color space subsampled format).

[0130] The MMR vector of the i-th chromaticity pixel (denoted as V) i HDR pixel values ​​can be specified or defined as follows:

[0131]

[0132] The k-th index entry is For example

[0133] A second-order Cb polynomial can be specified or defined as follows:

[0134]

[0135] The polynomial coefficients α of the second-order Cb polynomial in expression (3) above CbThey can be placed at their respective MMR positions or indices: 0, 2, and 9 in the above expression (2), where zero (0) is used as the starting position index in expression (2).

[0136] Similarly, the multiple polynomial coefficients α of the second-order Cr polynomial Cr (Not shown, but similar to the above expression (3)) can be placed at their respective MMR positions or indices: 1, 3, 10 in the above expression (2).

[0137] The polynomial coefficients of the chroma channel c are expressed as α. C Table 3 below illustrates an example procedure for writing each frame of chroma forward shaping data (stored as polynomial coefficients in the corresponding or transformed MMR position / index) for the images / frames covered by the forward binary file.

[0138] Table 3

[0139]

[0140] At the second level, HDR image data can be forward-shaped using forward-shaping data signaled in the forward binary file, and corresponding BL or SDR image data with multiple spatial resolutions can be generated and encoded.

[0141] Noise intensity adjustment

[0142] Noise such as film grain noise can be injected into the original or resized HDR image to prevent or reduce false contours or stripe artifacts in the rendered image obtained directly or indirectly from the HDR image. The original or resized image includes the original HDR image, the source HDR image, the input HDR image, the resized HDR image, the spatially downsampled HDR image, etc.

[0143] Noise can be injected into different original or resized HDR images randomly or non-repeatingly selected from multiple film grain images in a noisy image library. For each frame to be injected with noise (e.g., current, HDR, source HDR, etc.), a film grain image can be randomly or non-repeatingly selected from multiple film grain images in the noisy image library. In some operational scenarios, noisy images, such as film grain images, can be indexed with corresponding index values. A non-repeating pseudo-random number generator can be used to generate index values ​​that avoid repetition in consecutively generated index values. As a result, two consecutive images are injected with two different noisy images or film grain images. Additionally, optionally, or alternatively, noise such as film grain (noise) in the selected noise or film grain images as described herein can be scaled, adjusted, and / or modulated with brightness-related noise intensity and then added to the brightness channel of the original or resized HDR image, which in some operational scenarios can be represented as an HDR YUV image.

[0144] As an example, and not a limitation, the bit depth of an HDR image is n. v =16, and its available luminance codeword range is [0, 65535]. Noisy images (e.g., patterns, film grains, etc.) in the noisy image library can include noise values ​​normalized on the scale of [-1, 1].

[0145] In some operational scenarios, constant scaling (or a constant scaling factor) can be applied to noise values ​​in noisy images. As a result, the noise intensity is constant across different luminance subranges or bins within the range of available luminance codewords. On the one hand, many images or video clips may still exhibit striping artifacts when the scaling factor used to scale the noise intensity across these different luminance bins is relatively low. On the other hand, when the scaling factor used to scale the noise intensity across different luminance bins increases or becomes relatively high, some areas or portions of the image or video clip (especially highlighted areas) may exhibit excessive noise that can be visually perceptible and unpleasant to the viewer.

[0146] In some operational scenarios (e.g., HDR to SDR), forward shaping functions (e.g., represented by per-frame luminance forward-shaped data in the forward binary file as described herein) can be used to calculate the noise intensity at each HDR codeword within the available luminance codeword range. Forward shaping functions can be used to control the allocation of codewords across various luminance intensity subranges or bins.

[0147] For example, in operational scenarios where HDR luminance codeword subranges or bins are mapped to relatively few SDR codewords (e.g., consecutive HDR codewords), the distances between the mapped SDR codewords of the subranges or bins (e.g., average, etc.) are relatively large or relatively high. This can cause or result in relatively large or relatively high striping in the image constructed by backward shaping the forward-shaped SDR or BL image in the received video signal (e.g., reconstructed HDR image, etc.). Therefore, for those HDR codeword subranges or bins, relatively high noise intensity can be applied to prevent, mask, or reduce striping artifacts.

[0148] On the other hand, in scenarios where HDR bins are mapped to a relatively large number of SDR codewords, those bins may contain fewer visual stripes in the backward-shaped HDR image constructed by backward-shaping the forward-shaped SDR or BL image in the received video signal. For those luminance bins, the lower noise intensity is sufficient to mask any false contours or stripe artifacts, while simultaneously preventing unwanted and annoying high noise.

[0149] In these operating scenarios, the noise intensity in an HDR luma subrange or bin may be inversely correlated with the total number of mapped SDR codewords assigned to the HDR luma subrange or bin.

[0150] In various embodiments, noise injection as described herein can be injected into a raw HDR image, such as a 16-bit source HDR image or a 12-bit source HDR image. Similarly, noise injection as described herein can be injected into a downsampled HDR image (such as a 12-bit downsampled HDR image generated by downsampling (which is bit-depth downsampling rather than spatial downsampling), or a 16-bit raw or source HDR image.

[0151] Noise injection

[0152] Figure 4A The diagram illustrates the method for injecting film grain noise into n v Example processing flow in a bit-high HDR image. Let F(.) be the expression for processing n... v Bit-level (e.g., 16-bit, 12-bit, etc.) HDR luminance codewords are mapped to n s FLUT of a bit (e.g., 8-bit, 10-bit, etc.) SDR codeword. Let v be an HDR codeword; therefore, F(v) becomes its mapped SDR codeword. Let the entire HDR range be divided into N... B Individual bins. The total number of HDR codewords in each HDR luminance bin. Then, the HDR luminance bin b contains the HDR luminance codeword v = [(b*π), (b+1)*π)-1], where the bin index b = 0, 1, 2, ..., N. B -1.

[0153] Box 402 involves finding the corresponding per-binary codeword increment in each HDR luminance (or luminance) bin. The FLUT can be constructed as a monotonically non-decreasing function, or F(v2) ≥ F(v1) if v2 > v1. In some operational scenarios, each HDR luminance bin can be mapped to at least one SDR luminance codeword. The total number (more than 1) of additional SDR luminance codewords assigned to each HDR luminance bin (bin index b) can be represented as φ. b In other words, φ b Indicates the amount of SDR codeword increase in the HDR brightness bin b.

[0154] In a scenario where a 16-bit HDR image is mapped or forward-shaped to an 8-bit SDR image, the total number of SDR luminance codewords is 256. If the entire range of available 16-bit HDR luminance codewords is divided into 64 subranges or bins, each HDR luminance bin has 65536 / 64 = 1024 HDR codewords mapped or forward-shaped to a subset of the entire range of available 8-bit SDR luminance codewords (e.g., the total SDR luminance codeword range of 0-255 codewords, etc.).

[0155] For example, an HDR luminance bin with bin index b = 5 contains v = [1024*5, (1024*6-1)] HDR luminance codewords. Assume this HDR luminance bin is mapped to the SDR codeword F(v) = [30, 35]. Then, for an HDR luminance bin with bin index b = 5, the codeword increases by φ5 = 35 - 30 = 5.

[0156] Box 404 includes normalization of the codeword increment for each HDR luminance bin. Let φ max This is the maximum per-bin codeword increment across all HDR luminance bins. The maximum per-bin codeword increment can be used to normalize all per-bin codeword increments, for example, on a scale of [0,1]. The normalized per-bin codeword increment for HDR luminance with bin index b is... It can be given as follows:

[0157]

[0158] Box 406 includes assigning noise intensity based on the increase in codewords per bin in the HDR luminance bin.

[0159] Let ψ min The minimum noise level of the HDR image to be injected into the luminance channel, and let ψ maxThis refers to the maximum noise intensity. In various embodiments, these two noise intensities (or intensity values) can be automatically set by the system described herein and / or can be specified as user input from a designated user. The minimum and maximum noise intensities can be set to optimal or selected values ​​(e.g., optimal values ​​determined empirically or procedurally), which may vary or depend on some or all of the following: spatial resolution, target bit rate, the video compression codec used, and corresponding parameters, etc.

[0160] The normalized codeword increment per bin can be used to obtain the minimum noise intensity ψ configured for the HDR luminance codeword bin. min With maximum noise intensity ψ max The corresponding noise intensity per compartment between them. Let ψ v This is the noise intensity per cell for the HDR luminance codeword v in the HDR luminance codeword cell. The cell index b of the HDR luminance codeword cell. v It can be calculated as in, This is a floor operation. The noise intensity ψ per cell of the HDR luminance codeword v. v (between [ψ] min , ψ max (between) can utilize based on a warehouse index b v Increased normalized codewords in HDR luminance codeword bin (For example, The scaling factor (etc.) is calculated as follows:

[0161]

[0162] Figure 3A The illustration shows an example of luminance FLUT. Figure 3B The illustration shows an example of codeword increments per warehouse determined by FLUT.

[0163] Table 4 below illustrates an example process for calculating the per-bin noise intensity of each HDR codeword based on the per-bin codeword increment determined according to FLUT.

[0164] Table 4

[0165]

[0166] The noise intensity per cell calculated using cell-by-cell measurements can be smoothed using an adaptive smoothing filter with a kernel length of Θ codewords. As an example, and not a limitation, for 16-bit HDR luminance codewords, the kernel length Θ of the adaptive smoothing filter can be set to 2049.

[0167] set up It is the filtered noise intensity (per codeword) of the HDR luminance codeword v. Figure 3C The illustration shows an example (per codeword) noise intensity generated by smoothing the noise intensity per bin (e.g., in...). Figure 3B (Medium). As an example and not a limitation, the configured minimum and maximum noise intensities [ψ] are... min , ψ max [800, 1200]. The pre-filtering curve represents the noise intensity per codeword calculated using expression (5). The post-filtering curve represents the noise intensity per codeword generated by applying a smoothing filter to the noise intensity per codeword. Figure 3C The X-axis represents the entire HDR luminance codeword range of 0-65535, which is processed or divided into 64 bins.

[0168] Table 5 below illustrates an example program for calculating the noise intensity per codeword for each HDR codeword based on smoothing filtering.

[0169] Table 5

[0170]

[0171] Noise intensity varying with light intensity level

[0172] The techniques used to calculate noise intensity can be applied to both the original image and downsampled images (e.g., bit depth downsampled, spatial downsampled, etc.).

[0173] These technologies can be used in, for example Figure 2E The illustrated second level is used to determine the per-codeword noise intensity of each HDR luminance codeword, which can be in the original HDR image or in a downsampled HDR image obtained by downsampling (e.g., bit depth, spatial resolution, etc.) a source HDR image. As an example and not a limitation, noise intensity calculation can be performed on a downsampled HDR image with a bit depth of 12 bits. The luminance FLUT of the downsampled HDR image contains the entire range of luminance codewords with only 4096 entries, fewer than the 65536 entries of a 16-bit (e.g., original, source, etc.) HDR image. In some operational scenarios, the luminance FLUT F'(.) of the downsampled HDR image can be obtained by subsampling the original 65536 entries of the FLUT F(.) of the 16-bit (e.g., original, source, etc.) HDR image, as shown in expression (1).

[0174] Here, the forward-shaped SDR or BL luminance codeword (the HDR luminance codeword v in the 16-bit (e.g., original, source, etc.) HDR image is mapped to this codeword) can be specified or defined by the (lookup forward mapping) entry F'(floor(v / ε)) in the subsampled FLUT F'(.), where floor(v / ε) represents the entry index in the subsampled FLUT F'(.); ε represents "stride", where, Furthermore, floor(.) represents a floor operation that discards the decimal part of the argument and retains only the integer part (or number) of the argument.

[0175] In the example, N F =4096; ε = 65536 / 4096 = 16. To find the forward mapping (or forward-shaped SDR or BL luminance codeword) of the 16-bit HDR luminance codeword v = 32449, the entry index in the double-sampled FLUT F' is calculated as follows: floor(v / ε)) = floor(32456 / 16) = 2028.

[0176] Table 6 below illustrates example programs for determining the noise intensity per codeword based on the increment of the codeword per codeword in the double-sampled FLUT F' and for determining the noise intensity per codeword based on the application of a smoothing filter to the noise intensity per codeword.

[0177] Table 6

[0178]

[0179] In some operational scenarios, it is possible to Figure 2E Applications in the second level, such as Figure 4A The same process flow is illustrated, but per-codeword noise intensity is calculated using a double-sampled FLUT F' instead of the original FLUT F (which can be obtained from the forward binary file). Figure 2D The first level of input to Figure 2E The second stage injects noise, such as film grain noise, into the size-adjusted (or spatially downsampled) HDR image supported by each spatial resolution, as described in this paper's bitrate ladder.

[0180] A size-resized, noise-injected HDR image can be forward-shaped by forward-shaping operations based on luminance and chrominance forward-shaping data in the forward binary file (e.g., including luminance FLUT and chrominance MMR coefficients up to a fixed highest order) to generate a corresponding forward-shaped SDR or BL image that depicts the same visual semantic content as the HDR image.

[0181] The HDR luminance codeword of the i-th pixel in the resized, noise-injected HDR image is represented as: Table 7 below shows the corresponding n values ​​in the SDR or BL image used to calculate the corresponding forward-shaping injected noise for the HDR luminance codeword of the i-th pixel. s Forward-facing shaping SDR or BL brightness code Example program.

[0182] Table 7

[0183]

[0184] The HDR chroma (Cb / Cr) codeword of the i-th pixel in the resized, noise-injected HDR image is represented as: As previously shown, HDR luminance and chrominance codewords can be mapped to forward-shaped SDR or BL chrominance codewords for the Cb and Cr channels using fixed-order MMR coefficients signaled in the forward binary file. Table 8 below shows the corresponding forward-shaped SDR or BL chrominance codewords used to calculate the i-th pixel in the corresponding forward-shaped injected noise SDR or BL image. Example program.

[0185] Table 8

[0186]

[0187] Full and reduced cascaded pipelining architectures for real-time coding

[0188] like Figure 2D and Figure 2E The illustrated two-stage full / reduced processing pipeline can be used in, for example... Figure 2F The individual (combined) levels shown in the diagram are cascaded together. Figure 2F This single level can, but is not necessarily limited to, encoding bitstreams and / or video clips in real-time streaming / broadcast scenarios. Figure 2F In a single level, such as Figure 2D The output of the complete pipeline section (or first stage) illustrated, such as the forward binary (multiple forward binary files), can be directly fed into... Figure 2E This reduces the input to the streamlined section (or second stage) of the pipeline to minimize processing time latency. Additionally, Figure 2F A single level can be combined Figure 2E Multiple instances of the reduced pipeline section (or second stage) allow for parallel or simultaneous encoding of different combinations of spatial resolution and bit rate in the supported bit rate ladder.

[0189] Fragment encoding in cloud computing

[0190] like Figure 2D and Figure 2E The illustrated full / reduced pipelines can be cascaded together as follows: Figure 2F The illustrated single level generates a corresponding output bitstream portion (or output video segment) for each input video segment of a media program represented in an input or source HDR video signal, including an input or source HDR image, under different settings or combinations of spatial resolution and bitrate in the bitrate ladder. The corresponding output bitstream portion (or output video segment) and the input video segment depict the same visual semantic content, but with different combinations of spatial resolution and bitrate.

[0191] In some operational scenarios, implementation Figure 2D and Figure 2E Full / reduced level or Figure 2F A single, combined-level media streaming server / service dynamically or adaptively streams the generated output video clips to different receiver playback devices with different network conditions / bitrates and / or different display capabilities (e.g., spatial resolution and / or dynamic range and / or color gamut supported by the receiver playback device).

[0192] In some operational scenarios, implementation Figure 2D and Figure 2E Full / reduced level or Figure 2F A single, combined-level media streaming server dynamically or adaptively switches to a specific portion (or segment) of the generated output bitstream based on real-time network conditions and / or system resource utilization associated with the media streaming server / service and the receiving playback device. A specific portion (or segment) of output bitstream to be streamed from the media streaming server / service to the receiving playback device can be selected from among multiple corresponding portions (or segments) of output bitstream (or segments) as the best bitstream (or segment) supporting visual quality under real-time network conditions and / or system resource utilization.

[0193] Media streaming servers can be implemented in cloud-based media streaming scenarios using a cloud cluster comprising multiple computer nodes (e.g., virtual computers, computer instances launched using cloud computing services, etc.). A single computer node in the cluster can be assigned to process a corresponding input video segment from multiple input video segments generated from an input or source HDR video signal, to generate or encode (multiple) output bitstream portions (or output video segments) from the corresponding input video segments.

[0194] Figure 2GThe illustration shows an example of generating multiple input video segments from a sequence of consecutive input or source HDR images in an input HDR video signal. For different settings or combinations of spatial resolution and bitrate, the input video segments can be processed and encoded by cluster nodes in a computer cluster (e.g., cloud-based, etc.) into corresponding output bitstream portions or output video segments in the output SLBC video signal.

[0195] like Figure 2G As illustrated, an entire media program (or video clip) represented by a sequence of consecutive input or source HDR images can include a total of F N A series of images / frames. F in a media program (or video clip) N A continuous image / frame can be divided into multiple real segments. As used herein, the term "real segment" refers to a subset (or block of frames) of continuous input or source HDR images in a media program (or video clip), which do not overlap with each other and comprise mutually exclusive portions (or subsequences) of the (overall) sequence of continuous input or source HDR images. By way of example and not limitation, each of the multiple real segments divided from the sequence of continuous input or source HDR images in a media program (or video clip) comprises F images / frames.

[0196] By adding one or two adjacent portions of the overhead (bumper) frame to the corresponding ground truth segment, an input video segment as described herein can be generated or synthesized from the ground truth segment. One or more adjacent portions of the overhead (bumper) frame extend to one or more adjacent ground truth segments adjacent to, or (partially) overlap with, the corresponding ground truth segment.

[0197] More specifically, the first (or initial) input video segment (in Figure 2G The input video segments, denoted as "Seg 0" and consecutively as "Seg 1", "Seg 2", ..., "Seg(T-2)" and "Seg(T-1)" (which serve as edge input video segments), can be generated from the first (or starting) real segment by adding the trailing adjacent portion of the overhead (bumper) frame. The trailing adjacent portion of the overhead (bumper) frame added to the first input video segment extends to or (partially) overlaps with the second real segment immediately following the first real segment.

[0198] Second input video segment ( Figure 2GThe “Seg 1” (which serves as the internal input video segment) can be generated from the second real segment by adding a leading adjacent portion and a trailing adjacent portion of the overhead (bumper) frame. The leading adjacent portion and the trailing adjacent portion added to the second input video segment extend to, or (partially) overlap with, the first real segment immediately preceding the second real segment and the third real segment immediately following the second real segment.

[0199] The final (or ending) input video clip (in) Figure 2G The input video segment, denoted as "Seg(T-1)" (which serves as an edge input video segment), can be generated from the last (or end) real segment by adding a leading adjacent portion of an overhead (bumper) frame. The leading adjacent portion of the overhead (bumper) frame added to the last input video segment extends to or (partially) overlaps with the penultimate real segment immediately preceding the last real segment.

[0200] Therefore, a total of T input video clips can be generated from media programs (or video clips), where T = F. N / F. Each input video segment can be distributed, processed, or encoded into a corresponding output bitstream portion or output video segment by cluster nodes in a computer cluster (e.g., cloud-based, etc.) for different combinations of spatial resolution and bit rate.

[0201] Figure 2H The illustration shows an example system configuration or architecture implemented by a cluster node (e.g., a cloud-based virtual computer, etc.) denoted as "Node K" for processing input video clips as described herein. This is an example, and not a limitation, of what a cluster node can use or implement. Figure 2D and Figure 2E The diagram illustrates two levels for processing video clips. Additionally, optionally, or alternatively, as... Figure 2H The two levels shown in the diagram can be combined to form, as Figure 2F The illustration shows a single level.

[0202] At the first level (referred to as "Level 1"), the input video clips assigned to cluster nodes can be processed at the highest spatial resolution, which can be the same as the local spatial resolution of the input or source HDR image in the input video clip, to generate or produce forward binary files (referred to as ".bin files") and image metadata such as backward shaping metadata (referred to as ".rpu files").

[0203] In the second level (referred to as "Level 2"), each input or source HDR image (e.g., an HDR YUV image, etc.) is downsampled to obtain multiple (e.g., M, where M is an integer greater than one (1, etc.) downsampled HDR images at various spatial resolutions supported by the bitrate ladder described herein. These spatial resolutions are indicated as Res 1 (e.g., 720p, etc.), Res 2 (e.g., 432p, etc.) to Res M.

[0204] In the second stage, the luminance and chrominance forward mapping data and / or image metadata (“RPU”) from the forward binary file generated in the first stage can be used to process these downsampled HDR images at different spatial resolutions into corresponding output bitstream portions or corresponding output video clips at different spatial resolutions. For different resolutions, the same RPU and forward binary file can be used in the second stage. In some operational scenarios, given the same RPU and forward binary file, multiple threads and / or multiple processes can be used to generate the output bitstream portions or output video clips in parallel and / or independently.

[0205] In some operational scenarios, these output bitstream portions or output video segments can be combined, for example, at the central cluster node of a computer cluster (e.g., cloud-based) with other output bitstream portions or output video segments generated by other cluster nodes (e.g., other cloud-based virtual computers) to form multiple overall bitstreams or multiple video segment sequences.

[0206] In operational scenarios using linear coding mode, each of the combined multiple overall bitstreams may include a sequence of consecutively output SDR or BL image sequences at the corresponding spatial resolution and / or corresponding bit rate in the bitrate ladder supported by the computer cluster, and represent a complete version of a media program or video clip at the corresponding spatial resolution and / or corresponding bit rate.

[0207] In operational scenarios using segment coding mode, each of the combined multiple video segment sequences may include a sequence of consecutive video segments, each of which includes multiple consecutive output SDR or BL images at a corresponding spatial resolution and / or corresponding bit rate in the bit rate ladder supported by the computer cluster, and represents a complete version of a media program or video clip at the corresponding spatial resolution and / or corresponding bit rate.

[0208] Optimal film grain settings

[0209] As previously mentioned, noise injection, such as film grain noise injection, can be performed in a single-stage combined processing pipeline / architecture or in the second stage of a two-stage processing pipeline / architecture. The operational parameters for noise injection described herein can be specifically configured for one or more settings or combinations of different spatial resolutions and / or bit rates.

[0210] Film grain noise can be constructed using two-dimensional (2D) random fields in the spatial frequency domain. ω G ×ω G A film grain noise patch G(m, n) can be rendered in the following ways:

[0211] G(m,n)=p·iDCT(Q(x,y)) (6)

[0212] Where Q(x, y) represents ω G ×ω G DCT coefficients, where ω G It is a positive integer; p represents the noise standard deviation of film grain noise; iDCT(·) represents the inverse DCT operation or operator.

[0213] In some operational scenarios, the ω of the noise patch G ×ω G A subset of the AC coefficients of Q(x, y) in a set of DCT coefficients (where at least one coefficient in x or y is non-zero) can be set to Gaussian random numbers with a mean of 0 and a standard deviation of 1 (or p = 1). The ω of the noise patch G ×ω G All other coefficients in the DCT coefficients (including the DC coefficients where both x and y are equal to zero) are set to zero.

[0214] When a subset of frequency bands in multiple spatial frequencies, represented by G(m,n), has non-zero AC coefficients in the DCT domain, the noise patch in the pixel domain corresponding to G(m,n) in the DCT domain exhibits film grain noise. The lower the frequency band with the non-zero AC coefficient distribution, the larger the film grain appears. The DCT size (ω...) G ×ω G Control the maximum available film grain size.

[0215] Examples of noise injection operations include, but are not limited to, film grain noise injection in PCT application serial number PCT / US2019 / 054299, "Reducing Banding Artifacts in Backward-Compatible HDR Imaging", filed October 2, 2019, and published as WO 2020 / 072651, and in U.S. Provisional Patent Application Serial No. 62 / 950,466, "Noise synthesis for digital images", filed December 19, 2019, by H. Kadu et al., the entire contents of which are incorporated herein by reference as fully set forth herein.

[0216] Film grain parameters (such as one or more of the following: DCT size (ω) G ×ω G (e.g., frequency bands with non-zero AC coefficients) can be configured by the system and / or designated users.

[0217] For example, let f s and f e These are the start and end frequencies in each of the 2D directions (horizontal and vertical). These frequencies can be used to control or define the non-zero spatial frequencies f as follows: f s ≤f≤f e The start and end frequencies, which are operational parameters for film noise injection, can be adjusted based on the image spatial resolution and / or the bit rate (to be encoded). In some operational scenarios, the ratio of film grain size to image / frame size can remain approximately the same across all different image spatial resolutions, for example, with an error tolerance of 10%, 20%, 30%, or another percentile. Therefore, in these operational scenarios, smaller film grain can be used for smaller spatial resolutions.

[0218] The optimal film grain settings can be individually configured for film grain noise injection, especially for the corresponding settings or combinations of spatial resolution and / or bit rate among multiple settings or combinations of spatial resolution and / or bit rate.

[0219] Figure 3D The illustration shows an example 16×16 frequency domain pattern or block. The letter "n" at a frequency location subset indicates that the frequency location within that subset is set to a norm or a Gaussian random number. Other frequency locations outside the frequency location subset that do not contain the letter "n" are set to zero.

[0220] Inverse DCT (IDCT) can be applied to a 16×16 frequency domain pattern or block with a norm or Gaussian random number in a subset of frequency locations to generate a corresponding film grain patch in the pixel domain. This can be repeated to generate more than one film grain patch. A film grain noise image with a specific spatial resolution and a unit standard deviation can be generated by stitching together multiple non-overlapping film grain noise patches.

[0221] For 16-bit HDR video signals with first and second settings or combinations of spatial resolution and / or bit rate, the minimum noise intensity and maximum noise intensity can be set as follows: [ψ min , ψ max [500, 1000]. Since the HDR codeword range of a 16-bit HDR video signal is 0-65535, noise addition or injection of the HDR luminance (or brightness) codeword of pixels in the HDR input or source image can be set to a minimum noise intensity and a maximum noise intensity in the range of [500, 1000].

[0222] In another example, for the third setting or combination of an image spatial resolution of 768×432 and a frame rate of 24fps at a bitrate of 1Mbps, and the fourth setting or combination of an image spatial resolution of 480×360 and a frame rate of 24fps at a bitrate of 0.5Mbps, adding film grain noise may degrade compression performance. This is primarily because noise injected into the image at these settings may make it difficult for codecs such as standard 8-bit AVC compressors to represent the image efficiently. Therefore, as indicated in Table 4 above, images with these resolutions / bitrates in the third and fourth settings or combinations may not be injected with noise in some operating scenarios.

[0223] Example process flow

[0224] Figure 4B An example process flow according to an embodiment is illustrated. In some embodiments, one or more computing devices or components (e.g., encoding devices / modules, transcoding devices / modules, decoding devices / modules, inverse tone mapping devices / modules, tone mapping devices / modules, media devices / modules, inverse mapping generation and application systems, etc.) may perform this process flow. In block 422, the image processing system generates a forward-shaping map to map a source image of a first dynamic range to a corresponding forward-shaping image of a second dynamic range lower than the first dynamic range.

[0225] In block 424, the image processing system injects noise into images with a first dynamic range and a first spatial resolution to generate images with injected noise at the first dynamic range and the first spatial resolution. The images with the first dynamic range and the first spatial resolution are generated by spatially downsampling the source image of the first dynamic range.

[0226] In box 426, the image processing system applies a forward shaping map to map an image of injected noise with a first dynamic range and a first spatial resolution to generate an image of embedded noise with a second dynamic range and a first spatial resolution.

[0227] In block 428, the image processing system transmits a video signal encoded with an image containing embedded noise using a second dynamic range and a first spatial resolution to the receiving device for the receiving device to render a display image generated from the image containing embedded noise.

[0228] In this embodiment, the video signal represents a single-layer backward compatible signal.

[0229] In the embodiments, the first dynamic range is the high dynamic range; wherein, the second dynamic range is the standard dynamic range.

[0230] In one embodiment, the video signal is transmitted to the receiving device along with image metadata including a backward-shaping map; the displayed image represents a backward-shaping image of a first dynamic range, the backward-shaping image being generated by applying a backward-shaping map to an image with embedded noise of a second dynamic range and a first spatial resolution.

[0231] In the embodiments, the noise is one of the following: film grain noise or non-film grain noise.

[0232] In one embodiment, noise is injected into an image with a first dynamic range and a first spatial resolution, using a brightness-related noise intensity calculated using forward shaping mapping.

[0233] In an embodiment, noise is injected with one or more operating parameters, which are configured based on one or more of the following: a first spatial resolution, a target bit rate for transmitting the video signal to the receiving device, etc.

[0234] In an embodiment, the video signal includes one of the following: a coded bitstream encoded with a sequence of images with continuously embedded noise of a second dynamic range and a second spatial resolution, a sequence of consecutive video segments, each of the consecutive video segments in the sequence of consecutive video segments including a subsequence of images with continuously embedded noise of a second dynamic range and a second spatial resolution, etc.

[0235] In an embodiment, the image processing system is further configured to perform the following operations: injecting second noise into a second image with a first dynamic range and a second spatial resolution to generate a second image with injected noise of the first dynamic range and a second spatial resolution, the second image with the first dynamic range and the second spatial resolution being generated by spatially downsampling a source image of the first dynamic range; applying the same forward shaping mapping to map the second image with injected noise of the first dynamic range and the second spatial resolution to generate a second image with embedded noise of the second dynamic range and the second spatial resolution; and transmitting a second video signal encoded with the second image with embedded noise of the second dynamic range and the second spatial resolution to a second receiving device for the second receiving device to render a second display image generated from the image with embedded noise.

[0236] In the embodiments, the video signal and the second video signal are generated for different combinations of spatial resolution and bit rate in a bit rate ladder supported by the cloud-based media content system.

[0237] In an embodiment, a source image with a first dynamic range is provided in the input video signal; a plurality of input video segments are generated from the input video signal; each of the plurality of input video segments is assigned to a corresponding cluster node among a plurality of cluster nodes in a cloud-based computer cluster; the corresponding cluster node processes the input video segment into one of the following: a plurality of coded bitstream portions of different combinations of spatial resolution and bit rate, a plurality of output video segments of different combinations of spatial resolution and bit rate, etc.

[0238] Figure 4C An example process flow according to an embodiment is illustrated. In some embodiments, one or more computing devices or components (e.g., encoding devices / modules, transcoding devices / modules, decoding devices / modules, inverse tone mapping devices / modules, tone mapping devices / modules, media devices / modules, inverse mapping generation and application systems, etc.) may perform this process flow. In block 442, the image processing system receives a video signal generated by an upstream encoder and encoded with embedded noise using a second dynamic range and a first spatial resolution, the second dynamic range being lower than the first dynamic range.

[0239] The image with embedded noise in the second dynamic range and the first spatial resolution has been generated by applying forward shaping mapping to the image with injected noise in the first dynamic range and the first spatial resolution by the upstream encoder.

[0240] An image with injected noise of the first dynamic range and the first spatial resolution is generated by injecting noise into the image of the first dynamic range and the first spatial resolution through an upstream encoder; the image of the first dynamic range and the first spatial resolution is generated by spatially downsampling the source image of the first dynamic range.

[0241] In box 444, the image processing system generates a display image from an image with embedded noise of a second dynamic range and a first spatial resolution.

[0242] In box 446, the image processing system renders and displays an image on the image display.

[0243] In one embodiment, a video signal is received along with image metadata including a backward-shaping map; a display image is shown representing an image with injected noise at a first dynamic range and a first spatial resolution; and an image processing system is further configured to perform the following operations: applying the backward-shaping map to the image with the embedded noise at a second dynamic range and the first spatial resolution to generate the display image.

[0244] In embodiments, computing devices such as display devices, mobile devices, set-top boxes, and multimedia devices are configured to perform any of the methods described above. In embodiments, an apparatus includes a processor and is configured to perform any of the methods described above. In embodiments, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause any of the methods described above to be performed.

[0245] In one embodiment, a computing device includes one or more processors and one or more storage media storing an instruction set that, when executed by the one or more processors, causes any of the methods described above to be performed.

[0246] Note that although individual embodiments are discussed herein, any combination of the embodiments and / or some of the embodiments discussed herein can be combined to form further embodiments.

[0247] Example computer system implementation

[0248] Embodiments of the present invention may be implemented using computer systems, systems configured with electronic circuits and components, integrated circuit (IC) devices (such as microcontrollers, field-programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs)), and / or means including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or implement instructions relating to adaptive perceptual quantization of images with enhanced dynamic range, as described herein. The computer and / or IC may calculate any of the various parameters or values ​​relating to the adaptive perceptual quantization process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0249] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present disclosure. For example, one or more processors, such as those in a display, encoder, set-top box, transcoder, etc., can implement methods related to adaptive perceptual quantization of HDR images as described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. The program product may include any non-transitory medium carrying a set of computer-readable signals, including instructions that, when executed by a data processor, cause the data processor to perform the methods of embodiments of the present invention. Program products according to embodiments of the present invention can take any of a variety of forms. The program product may include, for example, physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0250] In the case of the components mentioned above (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise specified, references to said components (including references to “devices”) should be interpreted as including any component that performs the function of the described component as an equivalent of said component (e.g., functionally equivalent), including components that are structurally different from those that perform the functions in the illustrated exemplary embodiments of the invention.

[0251] According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to perform these techniques, or may include digital electronic devices persistently programmed to perform these techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hardwired and / or program logic to implement the techniques.

[0252] For example, Figure 5 This is a block diagram illustrating a computer system 500 on which embodiments of the present invention may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 to process information. The hardware processor 504 may be, for example, a general-purpose microprocessor.

[0253] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When stored in non-transitory storage media accessible to processor 504, such instructions enable computer system 500 to become a dedicated machine defined to perform the operations specified in the instructions.

[0254] Computer system 500 further includes read-only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions of processor 504. Storage device 510 (such as a magnetic disk or optical disk) is provided and coupled to bus 502 for storing information and instructions.

[0255] Computer system 500 can be coupled to display 512, such as an LCD, via bus 502 for displaying information to a computer user. Input device 514, including alphanumeric keys and other keys, is coupled to bus 502 for transmitting information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, trackball, or arrow keys, for transmitting directional information and command selections to processor 504 and for controlling cursor movement on display 512. Typically, this input device has two degrees of freedom on two axes (a first axis (e.g., x-axis) and a second axis (e.g., y-axis)), allowing the device to specify a position in a plane.

[0256] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic. These custom hardwired logics, one or more ASICs or FPGAs, firmware, and / or program logic, combined with the computer system, enable computer system 500 to be a dedicated machine or programmed to be a special-purpose machine. According to one embodiment, computer system 500 performs the techniques described herein in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium (such as storage device 510). Execution of the sequence of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0257] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, floppy hard disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, flash EPROMs, NVRAMs, any other memory chips or memory cartridges.

[0258] Storage media differ from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including conductors containing bus 502. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.

[0259] Various forms of media can involve loading one or more sequences of one or more instructions to processor 504 for execution. For example, instructions may initially be loaded onto a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 500 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 502. Bus 502 loads the data into main memory 506, from which processor 504 fetches and executes the instructions. Instructions received in main memory 506 may optionally be stored on storage device 510 before or after execution by processor 504.

[0260] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides bidirectional data communication coupled to network link 520, which connects to local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity with a corresponding type of telephone line. As another example, communication interface 518 may be a Local Area Network (LAN) card for providing data communication connectivity with a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.

[0261] Network link 520 typically provides data communication to other data devices via one or more networks. For example, network link 520 may provide a connection via local network 522 to host computer 524 or to data devices operated by Internet Service Provider (ISP) 526. ISP 526, in turn, provides data communication services via a global packet data communication network now commonly referred to as the “Internet” 528. Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals that carry streams of digital data. Signals through various networks, as well as signals on network link 520 and through communication interface 518 (which carries digital data to and from computer system 500), are example forms of transmission media.

[0262] Computer system 500 can send messages and receive data, including program code, through multiple networks, network links 520, and communication interfaces 518. In the Internet example, server 530 can transmit application request codes through the Internet 528, ISP 526, local network 522, and communication interface 518.

[0263] The received code may be executed by processor 504 and / or stored in storage device 510 or other non-volatile storage device for later execution upon receipt.

[0264] Equivalents, extensions, alternatives and miscellaneous

[0265] In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the sole and exclusive indication of the claimed embodiments of the invention, and of the applicant's opinion, is the set of claims issued in specific form according to this application, wherein such claims include any subsequent amendments. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, characteristics, advantages, or attributes not expressly referenced in the claims should not in any way limit the scope of such claims. Therefore, this specification and the drawings should be viewed in an illustrative rather than restrictive sense.

Claims

1. A method for encoding an image into a video signal, the method comprising: At the first level, perform the following operations: Generate a forward-shaping map to map a source image of a first dynamic range to a corresponding forward-shaping image of a second dynamic range that is lower than the first dynamic range; as well as At the second level, perform the following operations: An image with the first dynamic range and the first spatial resolution is generated by spatially downsampling the source image of the first dynamic range; The forward shaping mapping is used to calculate the noise intensity related to the first luminance. Noise with a noise intensity related to the first brightness is injected into the image with the first dynamic range and the first spatial resolution to generate an image with injected noise with the first dynamic range and the first spatial resolution. The forward shaping map is applied to map the injected noise image at the first dynamic range and the first spatial resolution to generate an embedded noise image at the second dynamic range and the first spatial resolution. as well as Encoding the image with embedded noise at the second dynamic range and the first spatial resolution into a video signal, and wherein the method further includes: In the second level, perform the following operations: The first dynamic range and a second spatial resolution different from the first spatial resolution are generated by spatially downsampling the source image of the first dynamic range; The forward shaping mapping is used to calculate the noise intensity related to the second luminance. Noise with a noise intensity related to the second brightness is injected into the image with the first dynamic range and the second spatial resolution to generate an image with injected noise with the first dynamic range and the second spatial resolution; The forward shaping map is applied to map the injected noise image at the first dynamic range and the second spatial resolution to generate an image with embedded noise at the second dynamic range and the second spatial resolution; and The image with embedded noise at the second dynamic range and the second spatial resolution is encoded into a second video signal.

2. The method as described in claim 1, wherein, Calculating the first luminance-related noise intensity using the forward shaping mapping includes: Divide the first dynamic range into several compartments; The corresponding increase in codewords per warehouse is determined from the forward shaping mapping; The increment of codewords per warehouse is normalized using the maximum increment of codewords per warehouse; and The noise intensity is increased for each warehouse codeword.

3. The method of claim 1 or 2, further comprising: In the first level, perform the following operations: Generate a forward binary file, wherein the forward binary file includes operation parameters or coefficients specifying the forward integer mapping; and The forward binary file is passed to the second level, and In the second level, perform the following operations: Receive the forward binary file from the first level; and The forward binary file is used to determine the noise intensity associated with the first brightness.

4. The method as described in claim 1 or 2, wherein, The video signal represents a single-layer backward compatible signal.

5. The method as described in claim 1 or 2, wherein, The first dynamic range is a high dynamic range, and the second dynamic range is a standard dynamic range.

6. The method of claim 1 or 2, further comprising: The video signal, along with image metadata including backward-shaping mapping, is transmitted to the receiving device. The receiving device decodes the embedded noise image from the video signal, which has the second dynamic range and the first spatial resolution. The receiving device applies the backward shaping mapping to the image with embedded noise of the second dynamic range and the first spatial resolution to generate a backward-shaped image of the first dynamic range; as well as The receiving device renders a display image representing the backward-shaped image of the first dynamic range.

7. The method as described in claim 1 or 2, wherein, The noise is one of the following: film grain noise or non-film grain noise.

8. The method as claimed in claim 1 or 2, wherein, The noise is injected with one or more operating parameters configured based on one or more of the following: the first spatial resolution or the target bit rate.

9. The method as claimed in claim 1 or 2, wherein, The video signal includes one of the following: a coded bitstream encoded with a sequence of images with continuously embedded noise of the second dynamic range and the first spatial resolution, or a sequence of consecutive video segments, each of the consecutive video segments comprising a subsequence of images with continuously embedded noise of the second dynamic range and the first spatial resolution.

10. The method of claim 1, wherein, The video signal and the second video signal are generated for different spatial resolutions and / or bit rates in a bit rate ladder supported by a cloud-based media content system.

11. The method as claimed in claim 1 or 2, wherein, The source image of the first dynamic range is provided in the input video signal; wherein a plurality of input video segments are generated from the input video signal; wherein each of the plurality of input video segments is assigned to a corresponding cluster node among a plurality of cluster nodes in a cloud-based computer cluster; wherein the corresponding cluster node processes the input video segment into one of the following: a plurality of coded bitstream portions with different spatial resolutions and / or bit rates, or a plurality of output video segments with different spatial resolutions and / or bit rates.

12. A method for rendering an image, the method comprising: The receiving device receives a video signal generated by the method according to any one of claims 1 to 11, as well as image metadata including backward shaping mapping; The receiving device decodes the embedded noise image from the video signal, which has the second dynamic range and the first spatial resolution. The receiving device applies the backward shaping mapping to the image with embedded noise of the second dynamic range and the first spatial resolution to generate a backward-shaped image of the first dynamic range; as well as The receiving device renders a display image representing the backward-shaped image of the first dynamic range.

13. An image processing apparatus, the image processing apparatus comprising a processor and configured to perform any one of the methods claimed in claims 1 to 12.

14. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon for performing the method using one or more processors according to any one of the methods of claims 1 to 12.