Method and module for video pipeline
By applying weighted and overpositive correction of residual data between the base decoding layer and the enhancement decoding layer of the video decoder, the hardware limitations of existing video decoders are addressed, enabling efficient decoding of high-quality UHD video, reducing power consumption, and simplifying platform upgrades.
Patent Information
- Application Number
- CN202380093032.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-12-22
- Publication Date
- 2025-11-14
AI Technical Summary
Existing video decoder hardware suffers from hardware limitations and inefficiency when implementing enhanced decoding, especially when processing UHD resolution in the video output path, making it difficult to efficiently decode high-quality video.
Enhanced decoding is achieved using existing video pipeline hardware by applying weighted and overpositive correction to residual data between the base decoding layer and the enhancement decoding layer of the video decoder. This includes using a hybrid unit for computation and controlling the operation of the enhancement decoder through a decoder integration layer.
It improves video decoding quality and efficiency without requiring hardware upgrades, enables real-time decoding of UHD video, reduces power consumption, and simplifies upgrade requirements for older platforms.
Smart Images

Figure CN120958813A_ABST
Abstract
Description
Technical Field
[0001] This application relates to decoders, and more particularly to enhanced decoders. Furthermore, this application relates to video pipelines. A video pipeline may include one or more of the following: one or more decoders (e.g., one or more base decoders and enhanced decoders), a video mixer, secure memory, non-secure memory, and secure and / or non-secure interfaces for connection to a display. Background Technology
[0002] In enhanced coding algorithms, such as MPEG-5 Part 2 Low Complexity Enhanced Video Coding (LCEVC), one or more layers of residual data can be used to improve the performance of the underlying coding algorithm.
[0003] In an enhanced encoder, residual data is calculated based on a comparison between the base-decoded video signal and the original input video signal. Generally, the base-decoded video signal is the result of encoding and then decoding the video signal according to the base codec. Each image element (such as a pixel) of the original input video signal can have a value lower than or higher than the corresponding image element in the base-decoded video signal. Typically, the difference will include a mixture of higher values (positive residual values) for some image elements and lower values (negative residual values) for others.
[0004] At the enhanced decoder, the residual data is used to recover the original input video signal based on the base-decoded video signal. This typically involves adding positive values of the residual data to the base-decoded video signal and subtracting negative values of the residual data from the base-decoded video signal. Summary of the Invention
[0005] The goal is to implement the enhanced decoder using existing video pipeline hardware as much as possible.
[0006] Video pipelines, and especially secure video pipelines such as those found in devices like televisions and set-top boxes, typically include fixed decoder hardware for one or more basic codecs. Video pipelines also typically include a mixing unit for blending two images, for example, to display a configuration menu on the displayed video. Video pipelines often further include numerical mapping functionality, for example, in the form of color correction or lookup tables.
[0007] This invention provides an enhanced decoding method that can utilize such characteristics and effectively recover the enhanced decoded video signal by applying residuals.
[0008] Furthermore, by utilizing existing hardware blocks in new ways, decoders can decode video at qualities higher than previously possible. For example, if a set-top box has a chip that previously could decode HD in real time, it may be possible to use that chip to decode UHD in real time without any upgrades to the hardware blocks and simply by using the hardware module in a new way.
[0009] According to a first aspect, a method for use in a video pipeline is described, the method comprising: obtaining a base-decoded video signal from a base decoding layer; obtaining one or more layers of positive residual data from an enhancement decoding layer, the one or more layers of residual data being generated based on a comparison of data derived from the base-decoded video signal and data derived from an original input video signal, wherein the base-decoded video signal includes values greater than and less than the original input video signal, and the positive residual data includes only values greater than or equal to zero; calculating a weighted sum of the base-decoded video signal and the one or more layers of positive residual data; and converting the weighted sum into an enhancement-decoded video signal by applying an over-positive correction.
[0010] Using the above method, most of it can be executed within the video output path of the secure decoder, and the method can be executed more efficiently, thereby achieving higher quality decoded video for the same processing and memory resources.
[0011] The alternative is similar to the first aspect, except that over-positive correction is applied to one or more layers of positive residual data instead of a weighted sum, and the resulting corrected residual data is used when calculating the weighted sum. In other words, a method for use in a video pipeline is described, comprising: obtaining a base-decoded video signal from a base decoding layer; obtaining one or more layers of positive residual data from an enhancement decoding layer, the residual data being generated based on a comparison of data derived from the base-decoded video signal and data derived from the original input video signal, wherein the base-decoded video signal includes values greater than and less than the original input video signal, and the positive residual data includes only values greater than or equal to zero; converting the one or more layers of positive residual data into one or more layers of corrected residual data by applying over-positive correction; and calculating a weighted sum of the base-decoded video signal and the one or more layers of corrected residual data.
[0012] Optionally, in the first or alternative aspect, the method includes using hybrid units to compute a weighted sum.
[0013] Optionally, in the first aspect or an alternative aspect, obtaining one or more layers of positive residual data includes: obtaining one or more layers of residual data from an enhanced decoding layer, the one or more layers of residual data being generated based on a comparison of data derived from the decoded video signal and data derived from the original input video signal; and processing the one or more layers of residual data to generate one or more layers of positive residual data.
[0014] Optionally, in the first aspect or an alternative aspect, processing one or more layers of residual data to generate one or more layers of positive residual data includes: cutting the one or more layers of positive residual data according to the maximum residual value and / or the minimum residual value.
[0015] Optionally, in the first or alternative aspect, the weight associated with each of the underlying decoded video signal and one or more layers of residual data is greater than or equal to zero.
[0016] Optionally, in the first aspect or an alternative aspect, the over-positive correction includes a predetermined gain and / or a predetermined offset.
[0017] Optionally, in the first or alternative aspect, overpositive correction includes using a lookup table to convert the weighted sum value into the enhanced decoded video signal value.
[0018] Optionally, in the first or alternative aspect, overpositive correction includes the application of color correction algorithms.
[0019] Optionally, in the first aspect or an alternative aspect, over-positive correction further includes cropping the enhanced decoded video signal based on the maximum and / or minimum decoded values.
[0020] According to the second aspect, a module for use in a video pipeline is described, wherein the module is configured to perform a method according to the first aspect or an alternative aspect.
[0021] According to another aspect, a bit stream is described, which includes data generated by the method according to the first aspect or data generated at an intermediate step of the method according to the first aspect.
[0022] In one aspect, a bitstream is described comprising a weighted sum of a base decoded video signal and one or more layers of positive residual data, the positive residual data being generated based on a comparison of data derived from the base decoded video signal and data derived from the original input video signal, wherein the base decoded video signal includes values greater than and less than the original input video signal, and the positive residual data includes only values greater than or equal to zero.
[0023] On the other hand, a bitstream is described that includes positive residual data generated based on a comparison of data derived from a base-decoded video signal and data derived from an original input video signal, wherein the base-decoded video signal includes values greater than and less than the original input video signal, and the positive residual data includes only values greater than or equal to zero.
[0024] In another aspect, a computer-readable storage device is described that stores a bit stream of one aspect of the preceding aspects. Attached Figure Description
[0025] Examples of systems and methods according to the invention will now be described with reference to the accompanying drawings, wherein:
[0026] Figure 1 A known high-level schematic diagram of the LCEVC decoding process is shown;
[0027] Figures 2a and 2b show schematic diagrams of a basic decoder and a decoder integration layer in a video pipeline, respectively.
[0028] Figure 3 A known high-level schematic of a video decoder chipset is shown;
[0029] Figure 4 A schematic diagram of a video pipeline according to an example of this disclosure is shown;
[0030] Figure 5 A schematic block diagram illustrating a video decoding method according to an example of this disclosure is shown;
[0031] Figure 6A A flowchart illustrating a method for generating positive residuals according to an example of this disclosure is provided.
[0032] Figure 6B A flowchart illustrating a method for generating an intermediate video signal according to an example of this disclosure is provided.
[0033] Figure 6C A flowchart illustrating a method for reconstructing the original input video signal according to an example of this disclosure;
[0034] Figure 7 A high-level schematic diagram of a video decoder chipset according to an example of this disclosure is shown; and
[0035] Figure 8 A block diagram illustrating the integration of an enhanced decoder according to an example of this disclosure is shown. Detailed Implementation
[0036] Hybrid backward-compatible coding techniques have previously been proposed in documents such as WO 2013 / 171173, WO 2014 / 170819, WO 2019 / 141987, and WO 2018 / 046940, the contents of which are incorporated herein by reference. Further examples of layer-based coding formats include ISO / IEC MPEG-5 Part 2 LCEVC (hereinafter referred to as “LCEVC”). LCEVC has been described in WO2020 / 188273A1, GB 2018723.3, WO 2020 / 188242, and related standard specifications, including the ISO / IEC DIS 23094-2 Low Complexity Enhanced Video Coding Draft, published at the MPEG 129 meeting held in Brussels from Monday, January 13, 2020 to Friday, January 17, 2020, the entire contents of which are incorporated herein by reference.
[0037] In these encoding formats, the signal is broken down into multiple data "tiers" (also called "hierarchical layers"), each corresponding to a "quality level," ranging from the highest tier to the lowest tier based on the original signal's sampling rate. The lowest tier typically represents a low-quality reproduction of the original signal, while the other tiers contain information about corrections applied to the reconstructed reproduction to produce the final output.
[0038] LCEVC employs this multi-layered approach, where any base codec (e.g., Advanced Video Coding (AVC, also known as H.264) or High Efficiency Video Coding (HEVC, also known as H.265)) can be enhanced via an additional low-bit-rate stream. LCEVC is defined by two component streams: the base stream, which is typically decoded by a hardware decoder; and the enhancement stream, which consists of one or more enhancement layers suitable for software processing implementations with sustainable power consumption.
[0039] In specific LCEVC examples of these layered formats, the process works by encoding a lower-resolution version of the source image using any existing codec (the base codec), and by using a different compression method (enhanced) to differentiate the reconstructed lower-resolution image from the source.
[0040] The remaining details that constitute the difference from the source are efficiently and quickly compressed using LCEVC, which employs specific tools designed to compress residual data. LCEVC enhances the compression of residual information across at least two layers: one layer at the base resolution to correct artifacts caused by the base encoding process; and another layer at the source resolution to add details to reconstruct the output frame. Between the two reconstructions, the image is upscaled using a canonical upsampler or a custom upsampler specified by the encoder in the bitstream. Furthermore, LCEVC performs nonlinear operations called residual prediction, which further refine the reconstruction process before residual addition, resulting in a low-complexity, intelligent content-adaptive (i.e., encoder-driven) upscaling.
[0041] Because LCEVC and similar encoding formats fully utilize existing decoders and are inherently backward compatible, efficient and effective integration with existing video coding implementations is possible without a complete redesign. Examples of known video decoding implementations include the software tool FFmpeg, used by the simple media player FFplay.
[0042] Furthermore, LCEVC is not limited to known codecs and can theoretically make full use of codecs that are yet to be developed. Therefore, any LCEVC implementation should be able to integrate with any known or undeveloped codecs implemented in hardware or software without introducing decoding complexity.
[0043] LCEVC is an enhanced codec, meaning it not only upsamples well but also encodes and compresses the residual information needed to achieve true fidelity to the source (by transforming, quantizing, and encoding it). LCEVC can also produce mathematically lossless reconstructions, meaning all information can be encoded and transmitted, perfectly reconstructing the image. LCEVC preserves the creator's intent, small text, logos, advertisements, and unpredictable high-resolution details.
[0044] As an example: –
[0045] -LCEVC can transmit 2160p 10-bit HDR video via an 8-bit AVC base encoder.
[0046] - When using the HEVC base encoder for 2160p streaming, LCEVC can deliver the same quality at a bit rate typically 33% lower than the original bit rate, i.e., reducing the typical bit rate of 20 Mbit / s (HEVC only) to 15 Mbit / s or lower (LCEVC on HEVC).
[0047] The many unique benefits of LCEVC can be summarized as follows. LCEVC…
[0048] - Quickly enhance the quality and cost-effectiveness of all codec workflows.
[0049] - Reduce the processing power requirements for serving a given resolution.
[0050] - It can be deployed via software, resulting in much lower power consumption.
[0051] -Simplify the transition from older generation codecs to newer generation codecs.
[0052] - Improve engagement by enhancing visual quality at a given positioning rate.
[0053] - Modifiable and backward compatible.
[0054] - Can be deployed at scale immediately via software updates.
[0055] - It features low battery consumption on user devices.
[0056] - Reduce the complexity of new codecs and make them easy to deploy.
[0057] Considering all of the above, LCEVC allows for interesting and highly cost-effective ways to achieve higher resolutions and frame rates on older devices / platforms without replacing the entire hardware, ignoring customers with older devices, or creating duplicate services for new devices. This approach to introducing higher-quality video services on older platforms simultaneously creates demand for devices with even better encoding performance. Furthermore, LCEVC not only eliminates the need for platform upgrades but also allows for the delivery of higher-resolution content over existing transport networks that may have limited bandwidth capabilities.
[0058] The LCEVC approach is a software-driven, codec-independent enhancer that leverages available hardware acceleration, as illustrated in the broader range of implementation options on the decoder side. While existing decoders are typically implemented in hardware at the bottom of the stack, LCEVC essentially allows for implementations at multiple levels: from scripts and applications to OS and driver levels, and all the way down to SoCs and ASICs. In other words, there are more than one solution for implementing LCEVC on the decoder side. Generally, the lower the implementation is in the stack, the more device-specific the approach becomes. No new hardware is required except for implementations at the ASIC level.
[0059] There are challenges in attempting to integrate LCEVC decoding into a video decoder pipeline chipset without redesigning the chipset. At least in the short term, it is desirable to implement LCEVC in a simple manner using existing architectures and designs. Specific implementation challenges exist related to secure decoding of protected (e.g., premium) content.
[0060] Generally, one place to perform operations for the LCEVC reconstruction phase (i.e., the combination of the decoded enhanced video and the residual of the base decoded video) is in the video output path. This is because the video output path is the safest, and also because such uses are memory-efficient, involving direct operations performed on secure memory.
[0061] However, such implementations in the video output path involve dealing with inherent hardware limitations. These limitations include, for example, low memory bandwidth and restrictions on the types of operations that can be performed. Components in the video output path (such as video shifters (alternatively called graphics feeders)) are specifically designed for and excel at functions such as overlay and color space conversion, but are limited to a wider range of applications.
[0062] Different blocks in the video output path have different limitations and trade-offs, and different blocks from different manufacturers have different functionalities. For example, a hardware upscaler designed for this specific purpose may have different trade-offs than a video shifter. Identifying how to implement LCEVC reconstruction within the video output path involves compromises. These challenges are exacerbated when processing operations at UHD resolution.
[0063] Alternative implementations that integrate LCEVC reconstruction into the decoder CPU may be unsafe because the CPU is not a protected pipeline, while implementations that integrate LCEVC into the video output path are potentially limited by the inherent hardware constraints of the path's blocks. Therefore, these implementations may be inefficient.
[0064] Seeking to address the limitations of video decoder chipsets and foster innovation in introducing and implementing enhanced decoders such as LCEVC into the broader video decoder ecosystem.
[0065] This disclosure describes implementations for integrating hybrid backward-compatible coding techniques with existing decoders, optionally via software updates. In a non-limiting example, this disclosure relates to implementations and integration of MPEG-5 Part 2 Low Complexity Enhanced Video Coding (LCEVC). LCEVC is a hybrid backward-compatible decoding technique that is a flexible, adaptable, efficient, and computationally inexpensive decoding format that combines different video decoding formats, underlying codecs (i.e., encoder-decoder pairs, such as AVC / H.264, HEVC / H.265, or any other current or future codecs, and non-standard algorithms such as VP9, AV1, etc.) with encoded data from one or more enhancement layers.
[0066] Exemplary hybrid backward-compatible decoding techniques use a downsampled source signal encoded using a base codec to form a base stream. Enhanced streams are formed using, for example, an encoded set of residuals from the base stream, corrected by increasing resolution or by increasing frame rate. Multiple levels of enhanced data can exist in the hierarchy. In some arrangements, the base stream can be decoded by a hardware decoder, while the enhanced streams can be adapted for processing using a software implementation. Therefore, a stream is viewed as a base stream and one or more enhanced streams, where typically two enhanced streams may exist but often only one is used. It is worth noting that typically the base stream can be decoded by a hardware decoder, while the enhanced streams can be adapted for software processing implementations with appropriate power consumption. Streams can also be viewed as layers.
[0067] Compared to the block-based approach used in the MPEG algorithm family, which encodes video frames in a hierarchical manner, hierarchical frame encoding involves generating residuals for the entire frame, and then generating residuals for reduced or extracted frames. In the examples described herein, residuals can be considered as errors or differences at a specific quality level or resolution.
[0068] For context purposes only, since the detailed structure of LCEVC is known and described in the approved draft standard specification, Figure 1 The logical flow explains how LCEVC operates on the decoding side, assuming H.264 as the underlying codec. Those skilled in the art will understand how the examples described herein also rely on references. Figure 1 The general description of LCEVC presented applies to other multi-layer coding schemes (e.g., those using a base layer and enhancement layers). [Go to...] Figure 1 The LCEVC decoder 10 operates at the individual video frame level. It takes as input the decoded low-resolution image from the base (H.264 or other) video decoder 11 and LCEVC enhancement data to produce a decoded full-resolution image ready to be displayed in the view. The LCEVC enhancement data is typically received in the Supplemental Enhancement Information (SEI) of the H.264 Network Abstraction Layer (NAL) or in the Additional Data Packet Identifier (PID), and is separated from the base-coded video by the demultiplexer 12. Therefore, the base video decoder 11 receives the demultiplexed encoded base stream, and the LCEVC decoder 10 receives the demultiplexed encoded enhancement stream, which is decoded by the LCEVC decoder 10 to generate a residual set for combination with the decoded low-resolution image from the base video decoder 11.
[0069] LCEVC can be quickly implemented in existing decoders with software updates and is essentially backward compatible, as devices that have not yet been updated to decode LCEVC can play video using the basic codec, which further simplifies deployment.
[0070] In this context, this paper proposes a decoder implementation to integrate decoding and rendering with existing systems and devices that perform the underlying decoding. This integration is easy to deploy. It also enables support for a wide range of encoding and player vendors and is easily updatable to support future systems. Specifically, embodiments of the invention relate to how to implement LCEVC in a manner that provides secure decoding of protected content.
[0071] The proposed decoder implementation is available through an optimized software library for decoding MPEG-5 LCEVC enhanced streams, providing a simple yet powerful control interface or API. This allows developers the flexibility to deploy LCEVC at any level, from low-level command-line tools to software stacks integrated with common open-source encoders and players. Specifically, embodiments of the invention generally relate to driver-level and system-on-chip (SoC)-level implementations.
[0072] The terms LCEVC and enhancement are used interchangeably in this document. For example, an enhancement layer may include one or more enhancement streams, i.e., residual data of LCEVC enhancement data.
[0073] Figure 2a illustrates the unmodified video pipeline 20. In this conceptual pipeline, acquired or received Network Abstraction Layer (NAL) units are input to a base decoder 22. The base decoder 22 can be a low-level media codec, accessed via mechanisms such as MediaCodec (e.g., found in the Android (RTM) operating system), VTDecompression Session (e.g., found in the iOS (RTM) operating system), or Media Foundation Transforms (MFT, e.g., found in the Windows (RTM) family of operating systems), depending on the operating system. The output of the pipeline is a surface 23 representing the decoded raw video signal (e.g., frames of this video signal, with the order of successful frames displayed in the video presentation).
[0074] Figure 2b conceptually illustrates the proposed video pipeline using the LCEVC decoder integration layer. Similar to the comparison video decoder pipeline in Figure 2a, NAL unit 24 is acquired or received and processed by LCEVC decoder 25 to provide surface 28 of reconstructed video data. By using LCEVC decoder 25, the quality of surface 28 can be higher than that of comparison surface 23 in Figure 2a, or the quality of surface 28 can be the same as that of comparison surface 23, but requires less processing and / or network resources.
[0075] In Figure 2b, the LCEVC decoder 25 is implemented in conjunction with the base decoder 26. The base decoder 26 can be provided by various providers and includes operating system functions as discussed above (e.g., using MediaCodec, VTDecompressionSession, or MFT interfaces or commands). The base decoder 26 can be hardware-accelerated, for example, using a dedicated processing chip to implement operations for a specific codec. The base decoder 26 can be the same base decoder shown as 22 in Figure 2a and used for other non-LCEVC video decoding; for example, it can include a pre-existing base decoder.
[0076] In Figure 2b, the LCEVC decoder 25 is implemented using a decoder integration layer (DIL) 27. The decoder integration layer 27 provides a control interface for the LCEVC decoder 25, allowing client applications to use the LCEVC decoder 25 in a similar manner to the base decoder 22 shown in Figure 2a, for example, as a complete solution from buffer to output. The decoder integration layer 27 controls the operation of decoder plug-in (DPI) 27a and enhancement decoder 27b to generate a decoded reconstruction of the original input video signal. In some variations, as shown in Figure 2b, the decoder integration layer may also control GPU functions 27c, such as GPU shaders, to reconstruct the original input video signal from the decoded base stream and the decoded enhancement stream.
[0077] NAL unit 24, comprising the encoded video signal along with associated enhancement data, may be provided in one or more input buffers. The input buffers may be fed to (or made available to) the base decoder 26 and the decoder integration layer 27, particularly the enhancement decoder controlled by the decoder integration layer 27. In some examples, the encoded video signal may comprise an encoded base stream and is received separately from an encoded enhancement stream comprising the enhancement data; in other preferred examples, the encoded video signal comprising the encoded base stream may be received together with the encoded enhancement stream, for example, as a single multiplexed encoded video stream. In the latter case, the same buffer may be fed to (or made available to) both the base decoder 26 and the decoder integration layer 27. In this case, the base decoder 26 may retrieve the encoded video signal comprising the encoded base stream and ignore any enhancement data in the NAL unit. For example, the enhancement data may be carried in an SEI message for the base stream of video data, which may be ignored by the base decoder 26 if the enhancement data is not suitable for processing custom SEI message data. In this case, the base decoder 26 can operate according to the base decoder 22 in Figure 2a, but in some cases, the base video stream may be at a resolution lower than that in the comparison case.
[0078] Upon receiving the encoded video signal, which includes the encoded underlying stream, the underlying decoder 26 is configured to decode the encoded video signal and output it as one or more underlying decoded frames. This output can then be received or accessed by the decoder integration layer 27 for enhancement. In one example set, the underlying decoded frames are passed as input to the decoder integration layer 27 in presentation order.
[0079] Decoder integration layer 27 extracts LCEVC enhancement data from the input buffer and decodes the enhancement data. Decoding of the enhancement data is performed by enhancement decoder 27b, which receives the enhancement data from the input buffer as an encoded enhancement signal and extracts the residual data by applying the enhancement decoding pipeline to one or more streams of encoded residual data. For example, enhancement decoder 27b may implement the LCEVC standard decoder as described in the LCEVC specification.
[0080] Decoder plugins provide functions at the decoder integration layer to control the base decoder. In some cases, decoder plugin 27a can handle the reception and / or access of base decoded video frames, and preferably applies LCEVC enhancements to these frames during playback. In other cases, the decoder plugin can arrange the output of the base decoder 26 to be accessible by the decoder integration layer 27, which is then arranged to control the addition of residual outputs from the enhancement decoder to generate output surface 28. Once integrated into the decoding device, LCEVC decoder 25 enables the decoding and playback of video encoded with LCEVC enhancements. The display of the decoded and reconstructed video signal can be supported by one or more GPU functions 27c, such as GPU shaders controlled by decoder integration layer 27.
[0081] Generally, the decoder integration layer 27 controls the operation of one or more decoder plug-ins and enhancement decoders to generate a decoded reconstruction of the original input video signal 28 using one or more layers of decoded video signal from the base coding layer (i.e., as implemented by the base decoder 26) and residual data from the enhancement coding layer (i.e., as implemented by the enhancement decoder). The decoder integration layer 27 provides a control interface to the video decoder 25, for example, to an application within a client device.
[0082] Depending on the configuration, the decoder integration layer can output the decoded data to surface 28 in different ways. For example, as a buffer, as an off-screen texture, or as the upper surface of the screen. The output format can be configured in the settings provided after creating an instance of the decoder integration layer 27, as explained further below.
[0083] In some implementations, if no enhancement data is found in the input buffer, for example, if NAL unit 24 does not contain enhancement data, the decoder integration layer 27 can back off to pass the video signal to the output at a lower resolution, i.e., the output of the base decoding layer implemented by base decoder 26. In this case, LCEVC decoder 25 can operate according to the video decoder pipeline 20 in FIG2a.
[0084] Decoder integration layer 27 can be used for both application integration and operating system integration, for example, for use by both client applications and the operating system. Decoder integration layer 27 can be used to control operating system functions, such as function calls to the hardware-accelerated underlying codec, without requiring the client application to be aware of these functions. In some cases, multiple decoder plugins can be provided, where each decoder plugin provides a wrapper for a different underlying codec. It is also possible for a common underlying codec to have multiple decoder plugins. This could be in cases where different implementations of the underlying codec exist, such as a GPU-accelerated version, a native hardware-accelerated version, and an open-source software version.
[0085] When viewing the schematic diagram of Figure 2b, the decoder plugin can be considered as integrated with the base decoder 26 or, alternatively, as a wrapper around the base decoder 26. Effectively, Figure 2b can be viewed as a visualization of a stack. The decoder integration layer 27 in Figure 2b conceptually includes functionality 27b for extracting augmented data from NAL units, functionality 27a for communicating with the decoder plugin and applying the augmented decoded data to the base decoded data, and one or more GPU functions 27c.
[0086] The collection of decoder plugins is configured to present a common interface (i.e., a common set of commands) to the decoder integration layer 27, allowing the decoder integration layer 27 to operate without knowing the specific commands or functionality of each underlying decoder. Plugins thus allow underlying codec-specific commands, such as MediaCodec, VTDecompression Session, or MFT, to map to a set of plugin commands accessible by the decoder integration layer 27 (e.g., multiple different decoding function calls can be mapped to a single common plugin "Decode(...)" function).
[0087] Since the decoder integration layer 27 effectively includes a 'residual engine,' i.e., a library of correction plane sets at different quality levels generated from LCEVC via the encoded NAL unit, this layer can behave as a complete decoder (i.e., the same as decoder 22) by controlling the underlying decoder.
[0088] For simplicity, the entity referred to herein will be called the client, but it should be understood that the client can be considered as any application layer or functional layer, and the decoder integration layer 27 can be easily and readily integrated into the software solution. The terms client, application layer, and user are used interchangeably herein.
[0089] In application integration, decoder integration layer 27 can be configured to be displayed directly on a screen surface of arbitrary size (typically different from the content resolution) provided by the client. For example, even if the underlying decoded video is standard definition (SD), decoder integration layer 27 can use enhancement data to display the surface at high definition (HD), ultra-high definition (UHD), or a custom resolution. Further details of non-standard methods for upscaling and post-processing of LCEVC decoded video streams that can be applied are found in PCT / GB2020 / 052420, which is incorporated herein by reference. Exemplary application integration includes, for example, the use of LCEVC decoder 25 by ExoPlayer (an application-level media player for Android) or VLCKit (a target C wrapper for the libVLC media framework). In these cases, VLCKit and / or ExoPlayer can be configured to decode LCEVC video streams "behind the scenes" using LCEVC decoder 25, wherein the computer program code for VLCKit and / or ExoPlayer functionality is configured to use and invoke commands provided by decoder integration layer 27, i.e., the control interface of LCEVC decoder 25. VLCKit integration can be used to provide LCEVC display on iOS devices, and ExoPlayer integration can be used to provide LCEVC display on Android devices.
[0090] In operating system integration, decoder integration layer 27 can be configured to decode the buffer or draw on an off-screen texture of the same size as the final resolution of the content. In this case, decoder integration layer 27 can be configured such that it does not handle the final display to a monitor, such as a display device. In these cases, the final display can be handled by the operating system, and therefore the operating system can use the control interface provided by decoder integration layer 27 to provide LCEVC decoding as part of an operating system call. In these cases, the operating system can implement additional operations related to LCEVC decoding, such as YUV to RGB conversion, and / or resizing the destination surface prior to final rendering on the display device. Examples of operating system integration include integration with an MFT decoder (or backend) for Microsoft Windows (RTM) operating systems or with an Open Media Acceleration (OpenMAX-OMX) decoder (or backend), OMX being a set of C-language-based programming interfaces (e.g., at the kernel level) for low-power and embedded systems, including smartphones, digital media players, game consoles, and set-top boxes.
[0091] These integration modes can be set by the client device or application.
[0092] The configuration and use of the decoder integration layer in Figure 2b allow LCEVC decoding and rendering to be integrated with many different types of existing legacy (i.e., basic) decoder implementations. For example, the configuration in Figure 2b can be viewed as a modification of the configuration in Figure 2a, as can be found on computing devices. Further examples of integration include LCEVC decoding libraries available within common video decoding tools such as FFmpeg and FFplay. For example, FFmpeg is often used as a basic video decoding tool within client applications. By configuring the decoder integration layer as a plugin or patch for FFmpeg, an LCEVC-enabled FFmpeg decoder can be provided, allowing client applications to use the known functionality of FFmpeg and FFplay to decode LCEVC (i.e., enhanced) video streams. For example, an LCEVC-enabled FFmpeg decoder can provide video decoding operations such as playback, decoding to YUV, and running metrics (e.g., Peak Signal-to-Noise Ratio (PSNR) or Video Multi-Method Evaluation Fusion (VMAF) metrics) without first decoding to YUV. This can be achieved through plugin or patch computer program code provided by the decoder integration layer for calling FFmpeg functions.
[0093] As described above, in order to integrate an LCEVC decoder (such as 25) into a client (i.e., an application or operating system), a decoder integration layer (such as 27) provides a control interface or API to receive instructions and configuration and exchange information.
[0094] Figure 3 A computing system 100a including a conventional video shifter 131a is shown. The computing system 100a is configured to decode video signals encoded using a single codec (e.g., VVC, AVC, or HEVC). In other words, the computing system 100a is not configured to decode video signals encoded using a layer-based codec (such as LCEVC). The computing system 100a further includes a receiving module 103a, a video decoding module 117a, an output module 131a, insecure memory 109a, secure memory 110a, and a CPU or GPU 113a. The computing system 100a is connected to a protected display (not shown).
[0095] The receiving module 103a is configured to: receive the encrypted stream 101a, separate the encrypted stream, and output decrypted secure content 107a (e.g., a decrypted encoded video signal encoded using a single codec) to a secure memory 110a. The receiving module 103a is also configured to output unprotected content 105a (such as audio or subtitles) to a non-secure memory 109a. The unprotected content can be processed 111a by the CPU or GPU 113a. The processed unprotected content is then output 115a to the video shifter 131a.
[0096] Video decoder 117a is configured to receive decrypted secure content (e.g., decrypted encoded video signal) from 119a and decode the decrypted secure content. The decoded secure content is sent to secure memory 110a from 121a and subsequently stored in secure memory 110a. The decoded secure content is output from secure memory 125a to video shifter 131a.
[0097] In other words, the video shifter 131a: reads decoded and decrypted secure content 125a from secure memory; reads insecure content 115a (e.g., subtitles) from insecure memory 109a; combines the decoded and decrypted secure content with the subtitles; and outputs the combined data 133a to the protected display.
[0098] The various components (i.e., modules and memories) are connected via multiple channels. A channel (also called a pipe) is a communication channel that allows data to flow between two components at either end of the channel. Generally, the channel connected to the secure memory 110c is a secure channel. The channel connected to the insecure memory 109c is an insecure channel.
[0099] Various examples of implementing LCEVC reconstruction on video decoders used, such as in set-top boxes, are discussed by referencing PCT / GB2022 / 051238, which is incorporated herein by reference in its entirety. The security-related portion of layer-based (e.g., LCEVC) decoder implementations resides in the processing steps where decoded enhancement layers are combined with decoded (and resolution-upgraded) base layers to create the final output sequence. Depending on which layer of the stack the layer-based (e.g., LCEVC) decoder is implementing, different approaches exist to establish secure and ECP-compliant content workflows.
[0100] For base decoders utilizing secure memory, PCT / GB2022 / 051238 discusses how to combine the output of a base decoder from secure memory with the output of an LCEVC decoder from general-purpose memory to assemble an enhanced output sequence. Two similar approaches are proposed: providing a secure decoder when implementing LCEVC at the driver-level implementation; or providing a secure decoder when implementing LCEVC at the system-on-chip (SoC) level. Which approach to use may depend on the capabilities of the chipset used in the corresponding decoding device.
[0101] Implementing LCEVC (or other layer-based codecs) at the device driver level utilizes hardware blocks or GPUs. Generally, once the base layer and (e.g., LCEVC) enhancement layers are separated, most of the decoding of the (e.g., LCEVC) enhancement layers can be performed in the CPU and therefore in general-purpose (insecure) memory. PCT / GB2022 / 051238 proposes using modules (e.g., secure hardware blocks or GPUs) to upsample the output of the base encoder using secure memory, combine the upsampled output with the prediction residual, and apply the decoded enhancement layer (e.g., LCEVC residual map) from general-purpose (insecure) memory. The output sequence (e.g., output plane) can then be sent to a protected display via an output module (e.g., a video shifter), which is part of the output video path in the decoder (i.e., in the chipset).
[0102] In short, the LCEVC reconstruction phase can be performed on the aspect of a computing system that has access to secure memory; that is, the steps of upsampling the underlying decoded video signal and combining it with one or more residual layers to create the reconstructed video. Examples include video output paths (such as video shifters), hardware blocks (such as hardware resolution upscalors), or the GPU of a computing system. Video shifters can also be called graphics feeders.
[0103] When the module implementing the LCEVC refactoring phase is a hardware block, it can be used to process data very efficiently (e.g., by maximizing page efficiency double data rate (DDR) memory).
[0104] However, not all devices have hardware blocks, and not all of these blocks have access to secure memory. In such cases, it may be preferable to have modular functionality within a GPU module (which many relevant devices have), providing a flexible approach and allowing implementation on many different devices, including telephones. By writing the modular functionality as a layer that runs on the GPU (e.g., using open GLES), the implementation can function on a variety of different GPUs (and therefore different devices), providing a single solution to a problem that can be implemented on many devices (i.e., providing secure video). This generally contrasts with SoC-level implementations, which typically use device (video shifter) architecture-specific implementations and therefore a unique solution for each video shifter to, for example, call the correct functions and connect them.
[0105] While the examples described herein are provided in the context of secure memory, such as the embodiments described in PCT / GB2022 / 051238, it should be understood that the principles presented herein are not limited thereto, but are provided only for the context. The advantages of the invention can be realized in video decoder embodiments that do not require protected pipelines, and protected pipelines are provided only for additional explanation.
[0106] When integrating LCEVC into existing video decoder architectures, the goal can be to do so in the simplest and most efficient way. While retrofitting LCEVC into existing set-top boxes is envisioned, integrating it into new chipsets is also advantageous. The expectation might be to integrate LCEVC without significant architectural changes, allowing chipset manufacturers to easily roll out fast and effortless LCEVC decoding without needing to modify their designs. Ease of integration is one of the many known advantages of LCEVC. However, implementing LCEVC in this way on existing chipset designs presents challenges.
[0107] As identified above, handling secure content is one such example. Another example of these integration challenges is the inherent hardware limitations of existing video decoder architectures.
[0108] Generally, the most appropriate place to perform the LCEVC reconstruction phase is probably in the video output path of the video decoder chipset. This addresses security requirements by storing the video in a protected pipeline, but it is also the most memory-efficient.
[0109] Examples of hardware limitations include resource issues in handling UHD, the inability to handle "signed" values (i.e., the hardware block may only process positive values), and / or the inability to perform subtraction operations. These three exemplary limitations introduce compromises due to the nature of the LCEVC reconstruction phase. That is, one or more residual layers output by the enhancement decoder typically include "signed" values (i.e., positive or negative values), and the underlying decoded video signal must be upsampled and combined with these signed values to reconstruct the original video.
[0110] In more specific examples of hardware limitations, some chipsets only allow addition when video is overlaid (because the blocks are being blended). Typically, for blending hardware, the hardware can be placed in addition mode rather than signed addition mode. LCEVC relies on both addition and subtraction.
[0111] In a specific example of hardware limitations, the set-top box might have limited memory bandwidth. Addition and subtraction of UHD values is 4x HD. For UHD, the underlying video is an HD image. Therefore, it might be possible to use hardware blocks to perform addition and subtraction of HD values, but when attempting these with UHD, the hardware blocks would either break or be extremely inefficient (due to insufficient memory bandwidth).
[0112] Therefore, in the example, a hardware block (such as a hardware resolution booster or other similar component) may be able to perform subtraction, but a video shifter cannot, and a video shifter may not be able to handle signed values.
[0113] Typically, the processor in a video pipeline cannot perform the necessary operations at UHD resolution, but may be able to perform certain operations on specific types of input.
[0114] Figure 4 The document presents an overview of the invention. The invention describes an implementation in which a video output path (“video pipeline”) is used for as many operations as possible, and a hardware block (CPU or GPU) is used for any remaining operations. The guiding principle of the implementation is primarily simplicity, and secondarily security, namely, the ability to decode secure content.
[0115] like Figure 4 As shown, the base decoder 401 is configured to receive a base-coded signal corresponding to the lower-quality video, and the enhancement decoder 402 is configured to receive an enhancement signal that can be used to enhance the lower-quality video to obtain a higher-quality video. The base-coded signal and the enhancement signal can be combined in the transmitted and received signals.
[0116] The enhancement decoder 402 (such as, for example, the LCEVC decoder) includes a residual generator 403. The residual generator is part of the enhancement operation and generates one or more layers of residual data R by decoding the enhanced signal. Each layer may correspond to a different quality level (e.g., a different resolution). The residual data is a set of signed values (i.e., positive and negative values) that generally correspond to the difference between the decoded version of the input video decoded using the underlying codec and the original input video signal.
[0117] This paper proposes module 404, which processes one or more layers of residual data to generate one or more layers of positive residual data R. + Positive residual data is a modified form of residual data that uses only positive values. The negative portions of the layers obtained by residual generator 403 are still included in the positive residual data, but are modified to have values greater than or equal to zero. In other words, "positive residual" can be equivalently called "modified residual" and has a similar meaning.
[0118] In some embodiments, the positive residual generator 404 generates positive residual data by applying a fixed offset to the entire layer of residual data obtained by the residual generator 403. For example, if the layer of the enhanced signal includes residuals with values between -X (e.g., -128) and +Y (e.g., 127), the positive residual generator 404 can increment all residual values by X to give a range of 0 to (X+Y) (e.g., 0 to 255).
[0119] Alternatively, the positive residual generator 404 can apply a dynamic offset to the residuals obtained by the residual generator 403. For example, the positive residual generator 404 can monitor the distribution of the residual values obtained by the residual generator 403 and identify offsets that will be sufficient to ensure that at least a certain percentage (e.g., 90% or 100%) of the modified residuals are greater than or equal to zero. The advantage of this is that the positive residual generator 404 can handle variable bit depths for enhancement layers.
[0120] The positive residual generator 404 can apply a weighting factor before or after applying the offset. For example, a 50% weighting can be applied to reduce the bit depth of the residual before the enhancement layer is combined with the underlying decoding layer, for example, to prevent integer overflow. The weighting factor can be between 0 and 1. Alternatively, the weighting factor can be greater than 1, for example, in the case where the received enhancement layer includes downsampled or additionally compressed residual values.
[0121] The positive residual generator 404 can apply shearing, which can be applied before or after applying the offset and / or the weighting factor. For example, the positive residual generator can be configured with an allowed range of positive residual values and can be configured to replace any individual residual value obtained by the residual generator 403 that falls outside the allowed range with the maximum or minimum residual value of the allowed range. In other words, the allowed range of positive residual values can be a range smaller than the range of residual values that can be obtained by the residual generator 403.
[0122] Positive residual generator 404 in Figure 4 The module shown is within the enhancement decoder 402. It should be understood that this module can be separate from the enhancement decoder and can receive the residual generated by the enhancement decoding process, or the module can be integrated into the enhancement decoder itself.
[0123] The base decoded video I0 is fed from the base decoder 401 to the upsampler 405. The upsampler 405 will then process the upsampled base decoded video I0... U The associated resolution is matched to the positive residual data R. + The resolution associated with each layer (or layer). (Of course, if the resolution of the base-decoded video already matches the resolution of the layers with positive residual data, upsampler 405 can be omitted.) Upsampler 405 can also apply a weighting factor to combine the base-decoded video with the positive residual data. For example, a 50% weighting can be applied to reduce the bit depth of the base-decoded video before combining the enhancement layer with the base-decoded layer.
[0124] In one case, weighting is applied to prevent integer overflow. In other words, weighting ensures that the sum of each image element (e.g., pixel) of the base-decoded video and its corresponding image element in the enhancement layer is no greater than the maximum supported value for the image elements of the base-decoded video.
[0125] The weighting factor used for the base-decoded video can be between 0 and 1. Alternatively, the weighting factor can be greater than 1.
[0126] The upsampled, fundamentally decoded video I0 generated by upsampler 405 U Compared with the positive residual data R generated by the positive residual generator 404 + Combining 406 to generate intermediate video signal I1 +The combination can be a simple sum of the residual value and the data value of the base decoded video. In this case, adder 406 can be used to implement the combination. Alternatively, the combination can be a weighted sum. As previously mentioned, the weights can be pre-applied by the positive residual generator 404 and / or upsampler 405. Alternatively, the weights can be applied at the stage of combination 406. In the above description, weighting is used to prevent overflow, but weighting can serve other purposes, such as supporting custom ranges of residual values or specific compression in the enhancement layer. Furthermore, in some cases, weighting may not be necessary (i.e., weight 1 can be used for both the residual value and the data value of the base decoded video).
[0127] Some hardware components (such as the mixing unit) are designed to combine values as required at stage 406, but are only capable of (or optimized for) combining positive values. Therefore, by generating positive residual data before combining the residual data with the base-decoded video, it becomes possible to implement combination with a wider range of hardware. This is particularly advantageous when there are constraints on the available hardware (such as in the video output path of a secure video decoder). As a specific example, a hardware constraint could be that the enhancement decoding must be backward compatible. An example of backward-compatible enhancement decoding is LCEVC. The advantage of backward compatibility is that existing video pipelines supporting the base decoder can be configured to also support enhancement decoding without replacing any hardware.
[0128] The intermediate video signal I1 generated by combining 406 + The original video signal, which is expected to differ from the encoded and transmitted signal, is modified before combination. Accordingly, an additional module 407 is provided to effectively reverse the effect of the positive residual generator 404 and output the enhanced decoded video signal data I1. This is described herein as "over-positive correction".
[0129] As previously mentioned, some hardware (especially in the context of secure video decoders) cannot perform a full range of mathematical operations. For example, in some embodiments, it is again necessary to avoid subtraction when implementing module 407. For example, module 407 may include a lookup table for converting the values of intermediate video signals into the values of enhanced decoded video signals. The conversion performed using the lookup table may be equivalent, for example, to applying a predetermined negative offset that is equal to and opposite to the offset applied by the positive residual generator 404.
[0130] Existing decoders often incorporate lookup table functionality into their video output paths, for example, to perform color correction. In such examples, a lookup table for "over-positive correction" can be combined with an existing lookup table to produce a combined lookup table, instead of a color correction lookup table, that can be installed in the video output path without increasing processing or memory requirements for that path. In decoders capable of receiving software updates, this lookup table can even be installed into the existing decoder.
[0131] Similarly, if color correction is performed by any other algorithm, a combined algorithm that performs both overpositive correction and color correction can be computed, and the combined algorithm can be installed in the decoder instead of the color correction algorithm.
[0132] Furthermore, in embodiments where the hardware supports subtraction, module 407 can directly apply a negative offset to the intermediate video signal without using a lookup table. For example, while the mixing unit used for combining 406 may not support subtraction, another hardware module available downstream of the mixing unit can support it.
[0133] Alternatively, the over-positive correction performed by module 407 may include applying a predetermined gain to the intermediate video signal. For example, each value of the intermediate video signal may be multiplied by a weight that depends on a weight pre-applied by the positive residual generator 404, the upsampler 405, and / or at a stage of combination 406.
[0134] Module 407 can also apply shearing, which can be applied before or after applying the offset and / or the weighting factor. For example, module 407 can be configured with a permissible range of enhanced decoded values and can be configured to replace any enhanced decoded value falling outside the permissible range with the maximum or minimum enhanced decoded value within the permissible range.
[0135] like Figure 4 As further illustrated, a secure decoder can typically have a secure section. The secure section is protected against actions such as unauthorized copying of the video signal, and the decoder is designed to prevent the decoded signal from being passed from the secure section to any non-secure (or "clear") section of the decoder. The non-secure section of the decoder can implement additional functions that do not require security. For example, in the case of video containing subtitles, the decoder can process the video signal within the secure section, while the subtitles are initially processed in the non-secure section and then securely combined with the video signal (e.g., using a mixing unit). The hardware in the secure section is typically heavily constrained, and therefore the non-secure section can be used more flexibly for some functions.
[0136] Enhanced video decoding may include decoding the base layer in the secure portion, decoding the residual layer in the non-secure portion, and generating a combined video signal in the secure portion.
[0137] The security component may advantageously include a base video decoder 401 and a video output path 408. The video output path 408 may include one or more hardware blocks, such as a mixing unit, CPU, GPU, etc., and the output path 408 may implement an upsampler 405, a combining module 406, and an over-positive correction module 407.
[0138] The non-secure portion can be configured to implement the enhancement decoder 402 using dedicated hardware and / or software configuration, the enhancement decoder including residual generator 403 and positive residual generator 404.
[0139] Figure 5 A schematic block diagram illustrating a video decoding method according to an example of this disclosure is shown. This is intended to provide feasible examples of methods according to this disclosure, summarized in the table below:
[0140] Original enhanced video signal value 14 Decoded base value 5A 64 Weighted base value 5B 32 Residual value 5C -50 Positive residual 5D 78 Weighted residual 5E 39 Intermediate video signal value 5F 71 Enhanced decoded video signal value 5H 14
[0141] exist Figure 5 In the example, the enhanced video signal comprises a sequence of values before encoding and transmission. The meaning of each value in the sequence can depend on how the video signal is described. For example, each value could be associated with a pixel in a frame of the video signal. Alternatively, each value could be associated with an attribute of a vector graphics element. Different values in the video signal can have different meanings—for example, the signal could include some values associated with multiple frames and other values that define a portion of the content of a single frame. The meaning of a “value” in the video signal is not constrained here. One of these values used in the feasible example is considered to be “14”.
[0142] After the enhanced video signal is received at its destination and the base layer of the enhanced video signal is decoded, the decoded base layer 5A has an exemplary value of "64".
[0143] The base video signal value 5A is converted into a weighted base value 5B. In this example, the base weight is 50%. Therefore, the base value "64" becomes the weighted base value "32". This weight can be applied by, for example, an upsampler 405 or a combiner 406.
[0144] After the enhancement layer of the enhanced video signal is decoded, the decoded enhancement layer includes a corresponding residual value 5C with a value of "-50". The decoded enhancement layer can be derived from... Figure 4 The residual generator 403 in the configuration is used for output.
[0145] The residual value 5C is converted to a positive residual value 5D. In this example, the conversion involves applying a fixed offset of 128 to each residual value. Therefore, the residual value "-50" becomes the positive residual value "78". This can be performed by the positive residual generator 404 as discussed above.
[0146] The positive residual value 5D is converted into a weighted residual value 5E. In this example, the residual weight is 50%. Therefore, the positive residual value "78" becomes the weighted residual value "39". This weight can be applied, for example, by the positive residual generator 404 or the combiner 406 discussed above.
[0147] The weighted base value 5B is combined with the weighted residual value 5E to give the intermediate video signal value 5F. In this example, the combination is a simple sum. Typically, the weights used for the simple sum are added together at 100%. The weighted base value "32" is combined with the weighted residual value "39" to give the intermediate video signal value "71". This can be performed by a combiner 406 (e.g., a mixing unit) as discussed above. For example, the mixing unit can be configured to perform the weighted sum combination stages 5B, 5E, and 5F simultaneously.
[0148] A correction of 5G is applied to the intermediate video signal value 5F to produce the enhanced decoded video signal 5H. The enhanced decoded video signal can then be output to a display for viewing. In this example, the correction of 5G includes a gain of 2 (inversely proportional to the previously mentioned weights), followed by an offset of -128 (matching the previously mentioned offset for the positive residual value 5D). Therefore, the intermediate video signal value "71" becomes "142", and then becomes "14", as the value of the enhanced decoded video signal. The value "14" is the same as the enhanced video signal before encoding and transmission.
[0149] Of course, the above is relative to Figure 5 The values discussed are merely examples, and different weights and offsets may be applied in other embodiments.
[0150] Figure 6A , Figure 6B and Figure 6C Each of the three exemplary stages representing the proposed concept is illustrated in a flowchart. As noted, each stage can be performed by the same or different modules of the video pipeline. For convenience, we will refer to these as generation, composition, and correction.
[0151] exist Figure 6A During the generation phase, the module receives one or more layers of residual data (step 601) and then processes the residual data to generate one or more layers of positive residuals (step 602). The positive residual data includes only values greater than or equal to zero. The generation phase can be implemented in the non-secure portion of the video pipeline, or alternatively, in the secure portion of the video pipeline.
[0152] In the Figure 6B At the assembly stage shown in the flowchart, the base-decoded video signal is received (step 611), for example, from a standardized base decoder. The base decoder, here referring to the decoder implementing the base codec (e.g., Advanced Video Coding (AVC, also known as H.264) or High Efficiency Video Coding (HEVC, also known as H.265)). If it is necessary to match the resolution of the residual data, the base-decoded video signal is upsampled or resolution-enhanced (step 612). The terms "upsampling" and "resolution enhancement" are used interchangeably herein. A positive residual is received (step 613) and combined with the resolution-enhanced base-decoded video signal (step 614). After combination, the assembly stage can generate or output an intermediate video signal from the combination of the positive residual and the (resolution-enhanced) base-decoded video signal (step 615).
[0153] Figure 6C The steps demonstrate the process of modifying intermediate decoded video signals to compensate for the original residuals, converting them into only positive values. The correction stage first receives the intermediate video signals after combination (step 621). Then, over-positive correction is applied (step 622), which may include using lookup tables, gain, offset, clipping, etc., as discussed above. The correction stage outputs or generates a reconstructed video signal (step 623). The final step may include storing the output plane and outputting it to an output module for transmission to a display.
[0154] Figure 7 A high-level schematic diagram of a video decoder chipset according to an example of this disclosure is shown.
[0155] Figure 7 The principles of this disclosure are illustrated in a video decoding computer system 100b, which includes general-purpose memory and secure memory. The computing system includes a receiving module 103b, a basic decoding module 117b, an output module 706b, an enhancement layer decoding module 113b, non-secure memory 109b, and secure memory 110b. The computing system is connected to a protected display (not shown).
[0156] The various components (i.e., modules and memories) are connected via multiple channels. A channel (also called a pipe) is a communication path that allows data to flow between two components at either end of the channel. Generally, the channel connected to the secure memory 110c is a secure channel. The channel connected to the insecure memory 109c is an insecure channel. For clarity, the channels are not explicitly shown in the figure; instead, the data flow between the various modules is illustrated.
[0157] Receiver module 103b is configured to receive video signal 101b as a single stream. The video signal includes an encrypted encoded reproduction of the base layer 107b and an encoded reproduction of the enhancement layer 105b. Receiver module 103b is configured to separate the video signal into an encrypted encoded reproduction of the base layer and an encoded reproduction of the enhancement layer. Receiver module 103b is configured to decrypt the encrypted encoded reproduction of the base layer. Receiver module 103b is configured to output the encoded reproduction of the enhancement layer 105b to the insecure memory 109b. Receiver module 103b is configured to output the decrypted encoded reproduction of the base layer 107b to the secure memory 110b.
[0158] The received encoded reproducible of the enhancement layer can be received by the receiving module 103b as an encrypted version of the encoded reproducible of the enhancement layer. In such an example, the receiving module 103b is configured to decrypt the encrypted version of the encoded reproducible of the enhancement layer before outputting the encoded reproducible of the enhancement layer to obtain the encoded reproducible of the enhancement layer 105b.
[0159] Insecure memory 109b is configured to receive (via insecure channel) and store the encoded reproduction of the video signal from enhancement layer 105b via receiving module 103b. Insecure memory 109b is configured to output the encoded reproduction of the enhancement layer to enhancement decoding module 113b, which is configured to generate a decoded reproduction of the enhancement layer by decoding the encoded reproduction. The decoded reproduction of the residual layer has a first resolution. Insecure memory 109b is configured to receive and store the decoded reproduction of the enhancement layer from insecure decoding module 113b.
[0160] The insecure memory 109b is configured to output the decoded reproduction of the enhancement layer to the enhancement decoding module 113b, which is configured to generate a positive residual layer at a first resolution. The insecure memory 109b is configured to receive and store the positive residual layer from the insecure decoding module 113b.
[0161] The insecure memory 109b can be further configured to output the decoded reproduction of the enhancement layer to the enhancement decoding module 113b, which is configured to identify correction requirements. The correction requirements provide a way to implement corrections for the secure portion of the decoder, which flexibly matches the processing of the enhancement layer in the insecure decoding module 113b. For example, the correction requirements may describe the processing that generates a positive residual layer (e.g., offset parameters or weights). The insecure memory 109b is configured to receive and store correction requirements from the insecure decoding module 113b.
[0162] The generation of the decoded reconstruction of the enhancement layer, the generation of the positive residual layer, and the identification of the correction requirements can be performed in multiple stages 702b, 704b, 706b, or a single stage 113b. In a single stage, the non-secure memory 109b outputs the encoded reconstruction of the enhancement layer 105b and stores the positive residual map and correction requirements.
[0163] Secure memory 110b is configured to receive the decrypted encoded replay of the base layer 107b of the video signal from receiving module 103b. Secure memory 110b is configured to output the decrypted encoded replay of the base layer 119b to base decoding module 117b. Secure memory 110b is configured to receive the decrypted decoded replay of the base layer 121b of the video signal generated by base decoding module 117b. Secure memory 110b is configured to store the decrypted decoded replay of base layer 121b.
[0164] Output module 708b has access to secure memory 110b and non-secure memory 109b.
[0165] Output module 708b is configured to read (via a secure channel) a base layer decrypted-decrypted reproduction of the video signal from secure memory 110b. The base layer decrypted-decrypted reproduction has a second resolution. In this illustrated embodiment, the second resolution is lower than the first resolution (however, this is not necessary; the second resolution can be the same as the first resolution, in which case upsampling of the base layer decrypted-decrypted reproduction is not required). Output module 708b is configured to generate an upsampled base layer decrypted-decrypted reproduction of the video signal by upsampling the base layer decrypted-decrypted reproduction, such that the upsampled base layer decrypted-decrypted reproduction has a first resolution.
[0166] Output module 708b is configured to read (e.g., via a non-secure channel) the decoded reproduction of the positive residual layer 712b of the video signal from non-secure memory 109b, the decoded reproduction Figure 7 The image is marked as an LCEVC positive residual map. Output module 708b is configured to apply the decoded and reconstructed positive residual layer 712b to the upsampled base layer to generate an intermediate video signal.
[0167] Output module 708b is further configured to convert the intermediate video signal into an enhanced decoded video signal by applying over-positive correction as discussed above. Output module 708b may be further configured to read (e.g., via an insecure channel) correction requirement 714b from insecure memory 109b, which correction requirement in... Figure 7The intermediate video signal is marked as an LCEVC correction requirement. In embodiments where a correction requirement is not used or when it is unavailable, output module 708b uses static parameters for correction. On the other hand, when a correction requirement is available, output module 708b can use dynamic parameters to control how the intermediate video signal is corrected.
[0168] Output module 708b is configured to output output plane 133b to a protected display (not shown) via a secure channel.
[0169] The output module may include a video shifter. The output module may further include a subtraction module, which may be, for example, a hardware scaling and compositing block, as typically found in a video decoder SoC or within a GPU operating in secure memory. Upsampling, generation of intermediate video signals, and generation of enhanced decoded video signals can be performed in multiple stages or a single stage 708b. Different stages can be performed by different hardware blocks, such as a mixing unit, a hardware scaling and compositing module, a hardware 2D processor, or a GPU.
[0170] The prediction residuals can be processed by the output module 131b, for example, using the prediction average based on lower resolution data, as described in WO 2013 / 171173 (incorporated by reference), and as can be applied as part of a modified upsampling procedure as described in WO / 2020 / 188242 (incorporated by reference) (such as in section 8.7.5 of the LCEVC standard). WO / 2020 / 188242 is specifically for section 8.7.5 of LCEVC because the prediction average is applied via a process referred to as “modified upsampling”. Generally, WO 2013 / 171173 describes calculating / reconstructing the predicted average at the pre-inverse transform stage (i.e., in the transformed coefficient space), but the modified upsampling in WO 2020 / 188242 moves the application of the predicted average modifier outside the pre-inverse transform stage and applies it during upsampling (in the post-inverse transform or reconstructed image space). This is possible because the transforms are (e.g., simple) linear operations, and their application can be moved within the processing pipeline. Therefore, output module 708b can be configured to: generate the predicted residual (consistent with the method described in WO 2020 / 188242); and apply the predicted residual (generated by modified upsampling) to the decrypted-decoded reconstructed image of the base layer (in addition to applying the decrypted reconstructed image of enhancement layer 115b) to generate the output plane. Generally, the output module 708b generates the prediction residual by determining the difference between: the average value of the upsampled and decrypted 2x2 block of the base layer; and the value of the corresponding pixel in the base layer (i.e., the unupsampled) decrypted and decoded block.
[0171] Figure 8A block diagram of the enhancement decoder is shown, incorporating the steps of the separation and subtraction stages described elsewhere in this disclosure, as well as the broader general steps of the enhancement decoder. As elsewhere described, the residuals can be generated in a separated form, rather than separated from a set of residuals created by the enhancement decoder.
[0172] The decoder receives the encoded base stream and one or more enhancement streams at decoder 800.
[0173] The encoded base stream is decoded at the base decoder 880 to produce a base reconstruction of the input signal received at the encoder. This base reconstruction can be used in practice to provide a visual reproduction of the signal at a lower quality level. However, the base reconstructed signal also provides a basis for a higher quality reproduction of the input signal.
[0174] Figure 8 Both sublayer 1 reconstruction and sublayer 2 reconstruction are shown. In the enhanced decoder shown, the sublayer 1 reconstruction is optional.
[0175] At sublayer 1, the decoded base stream is provided to the processing block for reconstructing the layer 1 video signal. The processing block also receives the encoded layer 1 stream and reverses any encoding, quantization, and transform applied by the encoder. The processing block includes an entropy decoding process 810-1, an inverse quantization process 820-1, and an inverse transform process 830-1. Optionally, only one or more of these steps may be performed depending on the operations performed at the corresponding block at the encoder. By performing these corresponding steps, the decoded layer 1 stream, including the first set of residuals, becomes available at the decoder 800.
[0176] The layer that processes the first set of residuals in 840-1 to generate positive residual data, for example, according to the reference above. Figure 4 or Figure 5 The technology described.
[0177] The first set of positive residuals is combined with the decoded base stream from the base decoder 880 (i.e., a summation operation 860-1 is performed on the decoded base stream and the first set of decoded residuals to generate the intermediate sublayer 1 video signal).
[0178] Overpositive correction 870-1 is applied to the intermediate sublayer 1 video signal to reconstruct a downsampled version of the input video.
[0179] Furthermore, and optionally in parallel, the encoded Level 2 stream is processed to produce another set of decoded residuals. Similar to the Level 1 processing block described above, the Level 2 processing block includes an entropy decoding process 810-2, an inverse quantization process 820-2, and an inverse transform process 830-2. These operations will correspond to the operations performed at the block in the encoder, and one or more of these steps may be omitted as needed.
[0180] The layer that processes the second set of residuals in 840-2 to generate positive residual data, for example, according to the reference above. Figure 4 or Figure 5 The technology described.
[0181] The decoded base stream is upsampled at upsampler 850-2 and summed with the positive residual at higher resolution at operation 860-2 to create the intermediate sublayer 2 video signal.
[0182] Overpositive correction 870-2 is applied to the intermediate sublayer 2 video signal to reconstruct the layer 2 version of the input video.
[0183] As described above, the enhancement stream may include two streams: a coding level 1 stream (first enhancement level) and a coding level 2 stream (second enhancement level). The coding level 1 stream provides a set of correction data, which can be combined with the decoded form of the base stream to generate a corrected image.
[0184] The architecture used to implement the above concepts may include two main components.
[0185] The first component can be a user-space application. Its purpose may be to parse the input transport stream (e.g., MPEG2), extract the underlying video and LCEVC stream (e.g., SEI NALU and dual-track multiplexing). The application's functions are: configuring the hardware underlying video decoder and passing the underlying video for decoding; decoding the LCEVC stream using DPI to create a positive residual plane; and sending the decoded underlying video and residual to the display driver.
[0186] The second component of this architecture can be the display driver. Its purpose is to modify the video device driver to perform resolution upscaling and compositing using a mixer and a set of hardware compositors. The mixer can be used to combine multiple video planes into a single output. The display driver's functions are: to upscale the underlying video decoding, then use the mixer (by adding to a pre-computed α) to composite it with a full-resolution positive residual, apply corrections using color management, and place a randomly generated dither mask on the on-screen display (OSD) plane; and to send the mixer's output to the display.
[0187] In the implementation, the base video and the enhanced video will be stored in a hardware-protected buffer throughout the process (i.e., a secure video path).
[0188] Implementation methods vary slightly across different SoC variants. Some SoC variants have additional features that allow for extra capabilities, such as enhanced negative residuals at higher resolutions, secondary resolution upscaling, color management, or image sharpening. Fundamentally, the architecture remains the same: a mixer is used for positive residuals; and corrections are applied at the post-mixing stage (e.g., using color management).
[0189] The desired method for enhancing a base video with LCEVC is as follows: performing a 2x2 resolution upscaling of the base video using specified scaler factors (kernels); adding a predicted average, i.e., the difference between pixel values in the base video and the average of four pixels in the corresponding 2x2 upscaled block; applying a signed offset plane to the result; and dithering the output by adding a plane of signed random values. Preferably, these steps are performed in hardware. For dithering, one implementation uses a blending method to add planes of positive random values—scaling—such that they do not exceed a specified dithering intensity.
[0190] Dithering can be applied at a lower resolution, which is then combined with the video signal to produce the final output. This method yields surprisingly good visual quality. In other words, dithering is applied at a separate plane and at a resolution lower than the output resolution. Furthermore, dithering can be applied to each of the YUV planes, whereas typically dithering can be applied to only one.
[0191] Dithered planes (such as 960x540 dithered planes) can be scaled at the post-mixing stage and applied to a scaled version of the LCEVC enhanced output itself, scaled for the display resolution. The video can then be output for display.
[0192] In other words, the LCEVC enhanced output (i.e., the output of the premixed and enhanced video data) can be scaled to the display resolution, such as 4:4:4, where luminance and chrominance have the same spatial resolution (other display resolutions, such as 4:2:2 or 4:2:0, are also contemplated). The dither plane can also be scaled to the display resolution, which is 4:2:2 in this example. The dither plane and the scaled enhanced video signal are then combined at the post-mixing stage to generate the video for display.
[0193] As noted above, applying dithering in this way—that is, outputting enhanced video and then applying dithering at the post-mixing stage—produces surprisingly good visual quality. Furthermore, arranging the video display path in this way allows for display at any resolution.
[0194] The dithered plane is input (i.e. applied) at a lower resolution before scaling.
[0195] In another example, the dither plane can be combined with the base decoded video signal at the premixing stage, and then combined with the positive residual.
[0196] In another example, the dither plane can be combined with the base decoded video simultaneously with the positive residual. This can be performed in the mixing unit, i.e., at the mixing stage.
[0197] Positive dithering can be applied at a lower resolution, upscaled, and then added to the final output. In other words, dithering can be applied on a separate plane at a resolution lower than the output resolution. Furthermore, dithering can be applied to every YUV plane, whereas typically dithering will be applied to only one.
[0198] Generally, any of the functionalities described in this text or illustrated in the diagrams may be implemented using software, firmware (e.g., a fixed logic circuit system), programmable or non-programmable hardware, or a combination of these embodiments. Generally, as used herein, the terms "component" or "function" refer to software, firmware, hardware, or a combination of these. For example, in the case of a software embodiment, the terms "component" or "function" may refer to program code that performs a specified task when executed on one or more processing devices. The illustrated separation of components and functions into distinct units may reflect any actual or conceptual physical grouping and allocation of such software and / or hardware and tasks.
Claims
1. A method for use in a video pipeline, the method comprising: Obtain the decoded video signal from the basic decoding layer; One or more layers of positive residual data are obtained from the enhancement decoding layer, the positive residual data being generated based on a comparison between data derived from the base decoded video signal and data derived from the original input video signal, wherein the base decoded video signal includes values greater than and less than the original input video signal, and the positive residual data includes only values greater than or equal to zero; and Calculate the weighted sum of the basic decoded video signal and the one or more layers of positive residual data; and The weighted sum is converted into an enhanced decoded video signal by applying an overpositive correction.
2. The method of claim 1, further comprising using a mixing unit to calculate the weighted sum.
3. The method according to any preceding claim, wherein obtaining the one or more layers of positive residual data comprises: One or more layers of residual data are obtained from the enhanced decoding layer, the one or more layers of residual data being generated based on a comparison between data derived from the decoded video signal and data derived from the original input video signal; and Process the one or more layers of residual data to generate the one or more layers of positive residual data.
4. The method of claim 3, wherein processing the one or more layers of residual data to generate the one or more layers of positive residual data comprises: The positive residual data of one or more layers are cut according to the maximum residual value and / or the minimum residual value.
5. The method according to any of the preceding claims, wherein the weight associated with each of the underlying decoded video signal and the one or more layers of residual data is greater than or equal to zero.
6. The method according to any of the preceding claims, wherein the over-positive correction includes a predetermined gain and / or a predetermined offset.
7. The method according to any of the preceding claims, wherein the over-positive correction comprises using a lookup table to convert the weighted sum value into an enhanced decoded video signal value.
8. The method according to any of the preceding claims, wherein the over-positive correction comprises applying a color correction algorithm.
9. The method of claim 8, wherein the overpositive correction comprises applying a combined algorithm that performs both overpositive correction and color correction.
10. The method according to any one of claims 6 to 9, wherein the over-positive correction further comprises cutting the enhanced decoded video signal according to the maximum decoded value and / or the minimum decoded value.
11. A module for use in a video pipeline, the module being configured to: Obtain the decoded video signal from the basic decoding layer; One or more layers of positive residual data are obtained from the enhancement decoding layer, the positive residual data being generated based on a comparison between data derived from the base decoded video signal and data derived from the original input video signal, wherein the base decoded video signal includes values greater than and less than the original input video signal, and the positive residual data includes only values greater than or equal to zero; and Calculate the weighted sum of the basic decoded video signal and the one or more layers of positive residual data; and The weighted sum is converted into an enhanced decoded video signal by applying an overpositive correction.
12. The module of claim 11, comprising a mixing unit, wherein the weighted sum is calculated by the mixing unit.
13. The module according to claim 11 or claim 12, wherein obtaining the one or more layers of positive residual data comprises: One or more layers of residual data are obtained from the enhanced decoding layer, the one or more layers of residual data being generated based on a comparison between data derived from the decoded video signal and data derived from the original input video signal; and Process the one or more layers of residual data to generate the one or more layers of positive residual data.
14. The module of claim 13, wherein processing the one or more layers of residual data to generate the one or more layers of positive residual data comprises: The positive residual data of one or more layers are cut according to the maximum residual value and / or the minimum residual value.
15. The module according to any one of claims 11 to 14, wherein the weight associated with each of the underlying decoded video signal and the one or more layers of residual data is greater than or equal to zero.
16. The module according to any one of claims 11 to 15, wherein the over-positive correction includes a predetermined gain and / or a predetermined offset.
17. The module according to any one of claims 11 to 16, wherein the over-positive correction comprises using a lookup table to convert the weighted sum value into an enhanced decoded video signal value.
18. The module according to any one of claims 11 to 17, wherein the over-positive correction comprises applying a color correction algorithm.
19. The module according to any one of claims 16 to 18, wherein the over-positive correction further comprises cutting the enhanced decoded video signal according to the maximum decoded value and / or the minimum decoded value.
20. A bitstream comprising a weighted sum of a basic decoded video signal and one or more layers of positive residual data. The one or more layers of positive residual data are generated based on a comparison between data derived from the base-decoded video signal and data derived from the original input video signal, wherein the base-decoded video signal includes values greater than and less than the original input video signal, and the positive residual data includes only values greater than or equal to zero.
21. A bitstream comprising positive residual data, the positive residual data being generated based on a comparison of data derived from a base-decoded video signal and data derived from an original input video signal, wherein the base-decoded video signal includes values greater than and less than the original input video signal, and the positive residual data includes only values greater than or equal to zero.
22. A computer-readable storage device storing a bit stream according to claim 20 or claim 21.
Citation Information
Patent Citations
Integrating a decoder for hierachical video coding
GB2597551A
Decomposition of residual data during signal encoding, decoding and reconstruction in a tiered hierarchy
WO2013171173A1
Hybrid backward-compatible signal encoding and decoding
WO2014170819A1
Video compression using differences between a higher and a lower layer
WO2018046940A1
Multi-codec processing and rate control
WO2019141987A1