Dynamic range adjustment parameter signaling and enabling variable bit depth support

By removing the parsing dependency between parameter sets in video coding technology, allowing independent parsing of dynamic range adjustment syntax elements, the video decoding latency problem is solved, enabling more efficient video data processing and support for different bit depths.

CN115362678BActive Publication Date: 2026-04-03QUALCOMM INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-09
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing video coding technologies, the parsing dependencies between parameter sets cause video decoders to wait for the sequence parameter sets to arrive, resulting in decoding delays.

Method used

By removing the parsing dependencies between parameter sets, the video decoder is allowed to parse the image parameter set and the adaptive parameter set before receiving the sequence parameter set, and to independently parse the dynamic range adjustment syntax elements, thereby reducing decoding latency.

Benefits of technology

It reduces video decoding latency, improves the efficiency and flexibility of video data processing, and supports video data compression with different bit depths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115362678B_ABST
    Figure CN115362678B_ABST
Patent Text Reader

Abstract

An example device for processing video data includes: a memory configured to store video data and one or more processors implemented in circuitry and coupled to the memory. The one or more processors are configured to parse a first set of parameters, which is signaled once in a bitstream of data for each set of encoded pictures. The one or more processors are configured to parse one or more Dynamic Range Adjustment (DRA) syntax elements in a second set of parameters, which is signaled in the bitstream and associated with at least one picture in the set of encoded pictures, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements of the first set of parameters; and to process the at least one picture based on the first and second set of parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Application No. 17 / 225,801, filed April 8, 2021, and U.S. Provisional Patent Application No. 63 / 008,533, filed April 10, 2020, the entire contents of which are incorporated herein by reference. U.S. Application No. 17 / 225,801, filed April 8, 2021, claims the benefit of U.S. Provisional Patent Application No. 63 / 008,533, filed April 10, 2020. Technical Field

[0002] This disclosure relates to video encoding and video decoding. Background Technology

[0003] Digital video functionality can be integrated into a wide variety of devices, including digital televisions, digital live broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite wireless phones, so-called "smartphones," video conferencing equipment, video streaming devices, and so on. Digital video devices implement video decoding technologies, such as those described in the standards defined below: MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Decoding (AVC), ITU-T H.265 / High-Efficiency Video Decoding (HEVC), and extensions to the above standards. By implementing these video decoding technologies, the aforementioned video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove inherent redundancy in video sequences. For block-based video decoding, video slices (e.g., video pictures or portions of video pictures) can be divided into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in the (I) slice of intra-frame decoding of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in the (P or B) slice of inter-frame decoding of a picture can use spatial prediction relative to reference samples in adjacent blocks within the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005] In general, this disclosure describes techniques for signaling and operations applied to video data to achieve more efficient compression of certain types of video data, such as high dynamic range (HDR) and wide color gamut (WCG) video data. More specifically, this disclosure describes techniques for implementing dynamic range adjustment decoding tools and removing parsing dependencies, enabling support for video data with different bit depths. By removing parsing dependencies, the techniques of this disclosure enable independent parsing of different parameter sets, thereby improving decoding latency.

[0006] In one example, a method includes: parsing a first set of parameters, which is signaled once per bitstream of encoded video data for each set of encoded images; parsing one or more dynamic range adjustment (DRA) syntax elements in a second set of parameters, which is signaled in the bitstream of encoded video data and associated with at least one image in that set of encoded images, wherein parsing the one or more DRA syntax elements does not depend on any syntax elements in the first set of parameters; and processing the at least one image based on the first and second set of parameters.

[0007] In another example, an apparatus for processing video data includes: a memory configured to store the video data; and one or more processors implemented in circuitry and coupled to the memory. The one or more processors are configured to: parse a first set of parameters, which is signaled once in the bitstream of encoded video data for each set of encoded pictures; parse one or more Dynamic Range Adjustment (DRA) syntax elements in a second set of parameters, which is signaled in the bitstream of encoded video data and associated with at least one picture in that set of encoded pictures, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements of the first set of parameters; and process the at least one picture based on the first and second set of parameters.

[0008] In another example, a non-temporary computer-readable storage medium is encoded with instructions. When executed, these instructions cause one or more processors to: parse a first set of parameters, which is signaled once in the bitstream of encoded video data for each set of encoded pictures; parse one or more Dynamic Range Adjustment (DRA) syntax elements in a second set of parameters, which is signaled in the bitstream of encoded video data and associated with at least one picture in that set of encoded pictures, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements of the first set of parameters; and process the at least one picture based on the first and second set of parameters.

[0009] In another example, an apparatus for processing video data includes: a component for parsing a first set of parameters, the first set of parameters being signaled once in the bitstream of encoded video data for each sequence of encoded pictures; a component for parsing one or more dynamic range adjustment (DRA) syntax elements in a second set of parameters, the second set of parameters being signaled in the bitstream of encoded video data and associated with at least one picture in the set of encoded pictures, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements of the first set of parameters; and a component for processing the at least one picture based on the first set of parameters and the second set of parameters.

[0010] Details of one or more examples are illustrated in the following figures and description. Other features, objects, and advantages will be apparent from the specification, figures, and claims. Attached Figure Description

[0011] Figure 1 This is a block diagram illustrating an example video encoding and decoding system that can perform the techniques disclosed herein.

[0012] Figure 2 This is a block diagram illustrating an example video encoder that can perform the techniques disclosed herein.

[0013] Figure 3 This is a block diagram illustrating an example video decoder that can perform the techniques disclosed herein.

[0014] Figure 4 This is a conceptual diagram illustrating human vision and display capabilities.

[0015] Figure 5 This is a conceptual diagram illustrating the color gamut.

[0016] Figure 6 This is a block diagram illustrating an example of HDR / WCG representation conversion.

[0017] Figure 7 This is a block diagram illustrating an example of inverse HDR / WCG conversion.

[0018] Figure 8 This is a conceptual diagram illustrating an example of a transfer function.

[0019] Figure 9 This is a conceptual diagram illustrating another example of a pass function.

[0020] Figure 10 This is a conceptual diagram illustrating an example of a luminance-driven chroma scaling (LCS) function.

[0021] Figure 11 This is a conceptual diagram illustrating an example table of functions that specify Qpc as qPi.

[0022] Figure 12 This is a conceptual diagram illustrating an example of an HDR buffer model.

[0023] Figure 13 This is a block diagram of a video encoder and video decoder system that includes a DRA unit.

[0024] Figure 14 This is a flowchart illustrating the dynamic range adjustment parameter parsing technique according to this disclosure.

[0025] Figure 15 This is a flowchart illustrating an example of video encoding.

[0026] Figure 16 This is a flowchart illustrating an example of video decoding. Detailed Implementation

[0027] In some draft video coding standards, resolution dependencies can exist between one parameter set and another. For example, the resolution of syntax elements in a Picture Parameter Set (PPS) or Adaptive Parameter Set (APS) may depend on syntax elements in a Sequence Parameter Set (SPS). This dependency is undesirable because parameter sets typically reside in different Network Abstraction Layer (NAL) units and can arrive at the video decoder at different times. Due to this dependency, if a PPS or APS applicable to a specific block of video data arrives at the video decoder before the SPS applicable to that specific block, the video decoder must wait for that SPS to arrive and resolve the SPS before resolving the PPS or APS. This results in decoding latency for the video decoder while waiting for the SPS to arrive.

[0028] According to the technique disclosed herein, dependencies between the PPS and SPS can be removed from the DRA syntax elements of the PPS and / or APS. The video decoder can receive the PPS and / or APS before receiving the SPS and can parse the PPS and / or APS without waiting for the SPS, because the DRA syntax elements of the PPS and / or APS do not depend on any syntax elements within the SPS. In this way, decoding latency is reduced compared to cases where one parameter set depends on syntax elements in another parameter set.

[0029] Figure 1 This is a block diagram illustrating an example video encoding and decoding system 100 that can perform the techniques of this disclosure. The techniques of this disclosure are generally aimed at decoding (encoding and / or decoding) video data. Typically, video data includes any data used for processing video. Thus, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.

[0030] like Figure 1 As shown, in this example, system 100 includes source device 102, which provides encoded video data to be decoded and displayed by target device 116. Specifically, source device 102 provides video data to target device 116 via computer-readable medium 110. Source device 102 and target device 116 can include any of a variety of devices, including desktop computers, laptops, tablets, set-top boxes, mobile phones such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and target device 116 are configured for wireless communication and are therefore referred to as wireless communication devices.

[0031] exist Figure 1 In the example, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Target device 116 includes an input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of target device 116 can be configured to apply techniques for implementing dynamic range adjustment decoding tools and removing parsing dependencies, and enabling support for video data with different bit depths. Thus, source device 102 represents an example of a video encoding device, while target device 116 represents an example of a video decoding device. In other examples, the source device and target device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, target device 116 may be connected to an external display device, rather than including an integrated display device.

[0032] Figure 1The system 100 shown is merely an example. In general, any digital video encoding and / or decoding device can implement techniques for implementing dynamic range adjustment decoding tools and removing parsing dependencies, enabling support for video data with different bit depths. Source device 102 and target device 116 are merely examples of these decoding devices, wherein source device 102 generates encoded video data for transmission to target device 116. This disclosure refers to a “decoding” device as a device that performs the decoding (encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices, specifically examples of a video encoder and a video decoder, respectively. In some examples, source device 102 and target device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and target device 116 includes both a video encoding component and a video decoding component. Therefore, system 100 can support one-way or two-way video transmission between video source device 102 and target device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0033] Typically, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a continuous series of pictures (also called “frames”) of the video data to video encoder 200. Video encoder 200 encodes the data for these pictures. Video source 104 of source device 102 may include, for example, a video capture device of a camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As an alternative, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured video data, pre-captured video data, or computer-generated video data. Video encoder 200 may rearrange the pictures from the receiving order (sometimes referred to as the “display order”) to a decoding order for decoding. Video encoder 200 may generate a bitstream that includes the encoded video data. Subsequently, source device 102 can output encoded video data to computer-readable medium 110 via output interface 108 for reception and / or retrieval via, for example, input interface 122 of target device 116.

[0034] The memory 106 of source device 102 and the memory 120 of target device 116 represent general-purpose memory. In some examples, memory 106 and memory 120 may store raw video data, such as raw video from video source 104 and raw, decoded video data from video decoder 300. Alternatively or additionally, memory 106 and memory 120 may store software instructions executable by, for example, video encoder 200 and video decoder 300, respectively. Although memory 106 and memory 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memory 106 and memory 120 may store, for example, encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of memory 106 and memory 120 may be allocated as one or more video buffers, for example, to store raw video data, decoded video data, and / or encoded video data.

[0035] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded video data from source device 102 to target device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to transmit encoded video data directly to target device 116 in real time, for example via a radio frequency network or a computer-based network. According to communication standards such as wireless communication protocols, output interface 108 can demodulate the transmitted signal including the encoded video data, and input interface 122 can demodulate the received transmitted signal. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, switch, base station, or any other equipment that facilitates communication from source device 102 to target device 116.

[0036] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, target device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 can include any of a variety of distributed or locally accessible data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.

[0037] In some examples, source device 102 may output encoded video data to file server 114 or another intermediate storage device that may store the encoded video data generated by source device 102. Target device 116 may access the video data stored on file server 114 in a streaming or download manner. File server 114 may be any type of server device capable of storing and sending encoded video data to target device 116. File server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network-attached storage (NAS) device. Target device 116 may access the encoded video data of file server 114 via any standard data connection, including an internet connection. This may include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or combinations thereof suitable for accessing the encoded video data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming protocol, a loading transport protocol, or a combination thereof.

[0038] Output interface 108 and input interface 122 can represent a wireless transmitter / receiver, a modem, a wired network component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 can be configured to transmit data, such as encoded video data, according to cellular communication standards, including, for example, 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, etc. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to transmit data, such as encoded video data, according to other wireless standards, including the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee), etc. TM Bluetooth TM Standards, etc. In some examples, source device 102 and / or target device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include an SoC device that performs functions distributed to video encoder 200 and / or output interface 108, while target device 116 may include an SoC device that performs functions distributed to video decoder 300 and / or input interface 122.

[0039] The technology disclosed herein can be applied to support video decoding in a variety of multimedia applications, including, for example, over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission, such as HTTP-based Dynamic Adaptive Streaming (DASH), digital video encoded to a data storage medium, decoding digital video stored on a data storage medium, or other applications.

[0040] The input interface 122 of the target device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, which is also used by the video decoder 300, such as syntax elements having values ​​describing the characteristics and / or processing of video blocks or other decoded units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays a decoded image of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as liquid crystal displays (LCDs), plasma displays, organic light-emitting diode (OLED) displays, or another type of display device.

[0041] Although Figure 1 As not shown, in some examples, the video encoder 200 and video decoder 300 may be integrated with the audio encoder and / or audio decoder, respectively, and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams that include both audio and video in a normal data stream. Where applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol or other protocols such as User Datagram Protocol (UDP).

[0042] Both the video encoder 200 and the video decoder 300 can be implemented as any of a variety of suitable encoder and / or decoder circuits, including, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination of the above circuits. When the above techniques are partially implemented using software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute these instructions in hardware using one or more processors to perform the techniques disclosed herein. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, and either the video encoder 200 or the video decoder 300 can be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or the video decoder 300 can include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.

[0043] The video encoder 200 and video decoder 300 can operate according to video decoding standards such as ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)) or its extensions (e.g., Multi-View and / or Scalable Video Coding Extensions). Alternatively, the video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards such as ITU-T H.266 (also known as Universal Video Coding (VVC)). The latest draft of the VVC standard (hereinafter referred to as "VVC Draft 8"), in the JVET-Q2001-vE chapter of the proposal submitted to the 17th meeting of the Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, held in Brussels, Belgium (BE) from January 7 to 17, 2020, describes the “Versatile Video Coding (Draft 8)” by Bross et al. (hereinafter referred to as “VVC Draft 8”). Alternatively, the video encoder 200 can operate according to MPEG-5 Enhanced Video Coding (EVC). However, the technology disclosed herein is not limited to any particular decoding standard.

[0044] Typically, video encoder 200 and video decoder 300 can perform block-based image decoding. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or used during encoding and / or decoding). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, video encoder 200 and video decoder 300 can decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, instead of decoding the red, green, and blue (RGB) data of the image samples, video encoder 200 and video decoder 300 can decode the luminance and chrominance components, where the chrominance components may include both red and blue chrominance components. In some examples, video encoder 200 converts received RGB format data to YUV representation before encoding, while video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and postprocessing units (not shown) can perform these conversions.

[0045] This disclosure can generally refer to the decoding (e.g., encoding and decoding) of an image, including the process of encoding or decoding data of that image. Similarly, this disclosure can refer to the decoding of blocks of an image, including the process of encoding or decoding data of the blocks, such as prediction and / or residual decoding. Encoded video bitstreams typically include a series of values ​​for syntax elements that represent decoding decisions (e.g., decoding modes) and how images are divided into blocks. Therefore, the description of decoding an image or block should generally be understood as the decoded values ​​of the syntax elements used to form that image or block.

[0046] HEVC defines various blocks, including decoding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (e.g., a video encoder 200) divides the decoding tree unit (CTU) into CUs according to a quadtree structure. That is, the video decoder divides the CTU and CU into four equal and non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and the CU of such leaf nodes can include one or more PUs and / or one or more TUs. The video decoder can further divide the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of the TU. In HEVC, PUs represent inter-frame prediction data, while TUs represent residual data. CUs resulting from intra-frame prediction include intra-frame prediction information, such as intra-frame mode indication.

[0047] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, the video decoder (e.g., video encoder 200) partitions the image into multiple decoding tree units (CTUs). Video encoder 200 can partition the CTUs according to a tree structure, including structures such as quadtree-binary tree (QTBT) or multi-type tree (MTT) structures. The QTBT structure eliminates the concept of multiple partitioning types, such as the distinction between CU, PU, ​​and TU in HEVC. The QTBT structure includes two levels: a first level obtained by partitioning according to a quadtree, and a second level obtained by partitioning according to a binary tree. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoding units (CUs).

[0048] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT), binary tree (BT), and one or more ternary tree (TT) partitions (also known as tritree (TT)). In a ternary or tritree partition, a block is divided into three sub-blocks. In some examples, a ternary or tritree partition divides a block into three sub-blocks without using a center point to partition the original block. Partition types in MTT (such as QT, BT, and TT) can be symmetric or asymmetric.

[0049] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components. In other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, for example, one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).

[0050] The video encoder 200 and video decoder 300 can be configured to use HEVC-based quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures. For illustrative purposes, the techniques of this disclosure have been described based on QTBT partitioning. However, it should be understood that the techniques of this disclosure can also be applied to video decoders configured to use quadtree partitioning or other types of partitioning.

[0051] In some examples, the CTU includes a decoded tree block (CTB) of luma samples, two corresponding CTBs of chroma samples of an image with three sample arrays, or a CTB of samples of a monochrome image or an image decoded using three independent color planes and a syntax structure for decoding the samples. For some value of N, the CTB can be an N×N sample block such that the partitioning of components to the CTB is a partition. A component is an array, or a single sample of one of three arrays (luma and two chroma) that make up an image with a color format of 4:2:0, 4:2:2, or 4:4:4, or an array or a single sample of an array that makes up an image in monochrome format. In some examples, for some values ​​of M and N, the decoded block is an M×N sample block such that the partitioning of CTB to the decoded block is a partition.

[0052] Individual blocks (e.g., CTUs or CUs) in an image can be grouped in several ways. As an example, a brick refers to a rectangular area within a row of CTUs in a specific tile of an image. A tile is a rectangular area composed of CTUs within a specific tile column and a specific tile row in an image. A tile column is a rectangular area composed of CTUs whose height is equal to the height of the image, and whose width (e.g., as in an image parameter set) is specified by the syntax element. A tile row is a rectangular area composed of CTUs whose height (e.g., as in an image parameter set) is specified by the syntax element, and whose width is equal to the width of the image.

[0053] In some examples, a tile can be divided into multiple bricks, each brick comprising one or more CTU rows within that tile. A tile that is not divided into multiple bricks can also be called a brick. However, a brick that is a proper subset of a tile cannot be called a tile.

[0054] The bricks in an image can also be arranged in slices. A slice can be an integer number of bricks in an image that can be exclusively contained in a single Network Abstraction Layer (NAL) unit. In some examples, a slice includes multiple complete tiles or a continuous sequence of complete bricks containing only one tile.

[0055] This disclosure uses "N×N" and "N multiplied by N" interchangeably to refer to the sample size of a block (e.g., a CU or other video block) in the vertical and horizontal dimensions, such as 16×16 samples or 16 by 16 samples. Typically, a 16×16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, the number of samples in the horizontal direction of a CU does not necessarily have to be the same as the number of samples in the vertical direction. For example, a CU may include N×M samples, where M is not necessarily equal to N.

[0056] Video encoder 200 encodes video data of a CU (Complex Unit) representing prediction and / or residual information, as well as other information. Prediction information indicates how the CU will be predicted to form a prediction block for that CU. Residual information typically represents the point-to-point difference between the prediction block and the CU samples before encoding.

[0057] To predict the Cubic Frame (CU), the video encoder 200 typically forms prediction blocks for the CU using either inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU based on data from previously decoded images, while intra-frame prediction typically refers to predicting the CU based on previously decoded data from the same image. To perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate prediction blocks. The video encoder 200 can typically perform motion search, for example, identifying a reference block that closely matches the CU based on the difference between the CU and a reference block. The video encoder 200 can calculate difference metrics using sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.

[0058] Some examples of VVC and EVC also provide affine motion compensation modes that can be considered as inter-frame prediction modes. In affine motion compensation modes, the video encoder 200 can determine two or more motion vectors representing non-translational motion, such as zooming in or out, rotation, perspective motion, or other irregular motion types.

[0059] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of VVC provide 67 intra-frame prediction modes, including various orientation modes as well as planar and DC modes. Examples of EVC can provide additional intra-frame prediction modes as well as planar and DC modes. Typically, the video encoder 200 selects an intra-frame prediction mode that describes the neighboring samples of the current block (e.g., a block of a CU) based on its prediction of the current block. Assuming the video encoder 200 decodes the CTU and CU in raster scan order (from left to right, from top to bottom), neighboring samples are typically located above, to the upper left, or to the left of the current block in the same image.

[0060] The video encoder 200 encodes data representing the prediction mode of the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used, as well as the motion information used for the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may use Advanced Motion Vector Prediction (AMVP) or merging modes to encode motion vectors. For affine motion compensation modes, the video encoder 200 may use similar modes to encode motion vectors.

[0061] Following a prediction, such as intra-frame or inter-frame prediction of a block, the video encoder 200 can compute residual data for that block. The residual data, such as a residual block, represents the sample-by-sample difference between that block and the predicted block formed using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to generate transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply a Discrete Cosine Transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 can apply a secondary transform after a primary transform, such as a Mode-Dependent Inseparable Secondary Transform (MDNSST), a Signal-Dependent Transform, a Karhunen-Loeve Transform (KLT), etc. The video encoder 200 generates transform coefficients after applying one or more transforms.

[0062] As described above, after performing any transformation that generates transform coefficients, the video encoder 200 can quantize these transform coefficients. Quantization typically refers to quantizing transform coefficients to reduce the amount of data used to represent them, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 can round an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, the video encoder 200 can right-shift the value to be quantized bit by bit to perform quantization.

[0063] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. This scan can be designed to place higher-energy (and therefore lower-frequency) transform coefficients at the beginning of the vector and lower-energy (and therefore higher-frequency) transform coefficients at the end. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to generate a serialized vector, and then entropy-encode the quantized transform coefficients of that vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy-encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic decoding (CABAC). The video encoder 200 can also entropy-encode the values ​​of syntax elements describing metadata associated with the encoded video data used by the video decoder 300 in decoding video data.

[0064] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. This context may involve, for example, whether the neighboring values ​​of the symbol are zero. Probability determination can be based on the context assigned to the symbol.

[0065] The video encoder 200 can further generate, for example, block-based syntax data, image-based syntax data, and sequence-based syntax data for the video decoder 300 in image headers, block headers, and slice headers, or generate other syntax data such as sequence parameter sets (SPS), picture parameter sets (PPS), video parameter sets (VPS), or adaptive parameter sets (APS). The video decoder 300 can also decode such syntax data to determine how to decode the corresponding video data.

[0066] In this way, the video encoder 200 can generate a bitstream. This bitstream includes encoded video data, such as syntax elements describing the division of images into blocks (e.g., CUs) and prediction and / or residual information for those blocks. Finally, the video decoder 300 can receive this bitstream and decode the encoded video data.

[0067] Typically, the video decoder 300 performs the reverse process of the video encoder 200 to decode the encoded video data of the bitstream. For example, although the CABAC encoding process is reversed compared to that of the video encoder 200, the video decoder 300 can use CABAC in a generally similar manner to decode the values ​​of the syntax elements of the bitstream. The syntax elements can define partitioning information for dividing an image into CTUs, and the partitioning of each CTU according to the corresponding partitioning structure (such as a QTBT structure), thereby defining the CU of that CTU. The syntax elements can also define prediction and residual information for blocks (e.g., CUs) of video data.

[0068] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inverse transform the quantized transform coefficients of a block to reconstruct the residual block of that block. The video decoder 300 uses a signaled prediction mode (intra-frame or inter-frame prediction) and associated prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for that block. The video decoder 300 can then combine this prediction block with the aforementioned residual block (on a sample-by-sample basis) to reconstruct the original block. The video decoder 300 can perform additional processing, such as a deblocking process, to reduce visual artifacts along the block boundaries.

[0069] As will be detailed below, this invention describes techniques for parsing syntax elements. In some examples, this disclosure describes techniques for parsing syntax elements related to DRA, wherein specific dependencies between syntax elements in different parameter sets are removed. In one example, the method includes: parsing an SPS applicable to a block of video data; parsing one or more DRA syntax elements applicable to the block in a second parameter set, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements in the SPS; and processing the block based on the SPS and the second parameter set.

[0070] According to the technology disclosed herein, the method includes: parsing a first parameter set, the first parameter set being signaled once in the bitstream of encoded video data for each set of encoded images; parsing one or more dynamic range adjustment (DRA) syntax elements in a second parameter set, the second parameter set being signaled in the bitstream of encoded video data and associated with at least one image in the set of encoded images, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements of the first parameter set; and processing the at least one image based on the first parameter set and the second parameter set.

[0071] According to the technology disclosed herein, the apparatus includes: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory. The one or more processors are configured to: parse a first set of parameters, the first set of parameters being signaled once in the bitstream of encoded video data for each set of encoded pictures; parse one or more dynamic range adjustment (DRA) syntax elements in a second set of parameters, the second set of parameters being signaled in the bitstream of encoded video data and associated with at least one picture in that set of encoded pictures, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements of the first set of parameters; and process the at least one picture based on the first set of parameters and the second set of parameters.

[0072] According to the technology disclosed herein, the apparatus includes: components for parsing a first parameter set, the first parameter set being signaled once in the bitstream of encoded video data for each set of encoded pictures; components for parsing one or more dynamic range adjustment (DRA) syntax elements in a second parameter set, the second parameter set being signaled in the bitstream of encoded video data and associated with at least one picture in the set of encoded pictures, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax element of the first parameter set; and components for processing the at least one picture based on the first parameter set and the second parameter set.

[0073] According to the technology of this disclosure, a non-temporary computer-readable storage medium is encoded with instructions. When executed, the instructions cause one or more processors to: parse a first parameter set, which is signaled once in the bitstream of encoded video data for each set of encoded pictures; parse one or more Dynamic Range Adjustment (DRA) syntax elements in a second parameter set, which is signaled in the bitstream of encoded video data and associated with at least one picture in that set of encoded pictures, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements of the first parameter set; and process the at least one picture based on the first and second parameter sets.

[0074] This disclosure can generally refer to "signaling" certain information, such as syntax elements. The term "signaling" can generally refer to the communication of the value of a syntax element and / or other data used for decoding encoded video data. That is, the video encoder 200 can signal the value of a syntax element in the bitstream. Typically, signaling refers to generating a value in the bitstream. As described above, the source device 102 can transmit the bitstream to the target device 116 substantially in real time, or it can be non-real-time, such as when storing syntax elements in storage device 112 for later retrieval by the target device 116.

[0075] Figure 2 This is a block diagram illustrating an example video encoder 200 that can perform the techniques disclosed herein. Provided for illustrative purposes. Figure 2 This should not be construed as limiting the techniques broadly illustrated and described in this disclosure. For illustrative purposes, this disclosure describes a video encoder 200 based on the techniques of VVC (ITU-T H.266 under development), HEVC (ITU-T H.265), and EVC. However, the techniques of this disclosure can be implemented by video encoding devices configured to other video decoding standards.

[0076] exist Figure 2 In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded image buffer (DPB) 218, and an entropy encoding unit 220. Any one or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy encoding unit 220 can be implemented in one or more processors or in processing circuitry. For example, units of the video encoder 200 can be implemented as one or more circuit or logic elements, as part of hardware circuitry, or as part of an FPGA, processor, or ASIC. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to implement these or other functions. In some examples, the video encoder 200 may also include DRA signaling techniques capable of performing the present disclosure and will be discussed later herein. Figure 13 The described positive DRA unit.

[0077] The video data storage 230 can store video data to be encoded by components of the video encoder 200. The video encoder 200 can receive video data stored in the video data storage 230 (e.g., from video source 104). Figure 1 The video data memory 230 and DPB 218 can act as a reference picture memory, storing reference video data used by the video encoder 200 to predict subsequent video data. The video data memory 230 and DPB 218 can be formed from any of a variety of storage devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The video data memory 230 and DPB 218 can be provided as the same storage device or as separate storage devices. In several examples, the video data memory 230 can be on-chip with other components of the video encoder 200, as shown above, or off-chip relative to those components.

[0078] In this disclosure, the description of video data memory 230 should not be construed as being limited to memory within video encoder 200 unless so specifically described, or to memory external to video encoder 200 unless so specifically described. More precisely, the description of video data memory 230 should be understood as a reference memory that stores video data received by video encoder 200 for encoding (e.g., video data for the current block to be encoded). Figure 1 The memory 106 can also provide temporary storage for the outputs from the various units of the video encoder 200.

[0079] right Figure 2 The various units are described below to aid in understanding the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination of both. Fixed-function circuits are circuits that provide specific functions and are preset for the operations that can be performed. Programmable circuits are circuits that can be programmed to perform various tasks and provide flexible functionality in the operations that can be performed. For example, a programmable circuit can execute software or firmware to cause the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., receiving or outputting parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of the above units may be different circuit blocks (fixed-function or programmable), while in other examples, one or more of the above units may be integrated circuits.

[0080] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the video encoder 200 is operated using software executed by programmable circuitry, memory 106 ( Figure 1 The instructions (e.g., object code) of the software received and executed by the video encoder 200 can be stored, or these instructions can be stored in another memory (not shown) in the video encoder 200.

[0081] The video data storage unit 230 is configured to store received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be the raw video data to be encoded.

[0082] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units for video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-block copying unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0083] The mode selection unit 202 typically coordinates multiple coding paths to test combinations of coding parameters and the resulting rate distortion (RDD) values ​​of those combinations. These coding parameters may include the partitioning from CTU to CU, the prediction mode used for the CU, the transformation type of the residual data used for the CU, and the quantization parameters of the residual data used for the CU. The mode selection unit 202 can ultimately select a combination of coding parameters that has a better RTD value than other tested combinations.

[0084] The video encoder 200 can divide images retrieved from the video data storage 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can divide the image's CTUs according to a tree structure, such as the QTBT structure described above or the quadtree structure of HEVC. As mentioned above, the video encoder 200 can form one or more CUs by dividing CTUs according to a tree structure. Such CUs can also be referred to as "video blocks" or "blocks".

[0085] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate predicted blocks for the current block (e.g., the current CU, or the overlapping portion of PU and TU in HEVC). For inter-frame prediction of the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference images (e.g., one or more pre-decoded images stored in DPB 218). Specifically, motion estimation unit 222 may calculate values ​​indicating how similar a potential reference block is to the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. Motion estimation unit 222 may typically use sample-by-sample differences between the current block and the considered reference blocks to perform these calculations. Motion estimation unit 222 may identify the reference block with the minimum value obtained from these calculations as the reference block that most closely matches the current block.

[0086] Motion estimation unit 222 can generate one or more motion vectors (MVs), each defining the position of a reference block in a reference image relative to the position of the current block in the current image. Motion estimation unit 222 can then provide these motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, it can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data for the reference blocks. As another example, if the motion vectors have fractional sample precision, motion compensation unit 224 can interpolate values ​​for the prediction blocks based on one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for the two reference blocks identified by the corresponding motion vectors and combine the retrieved data, for example, by sample-by-sample averaging or weighted averaging.

[0087] As another example, for intra-frame prediction or intra-frame prediction decoding, intra-frame prediction unit 226 can generate a prediction block based on samples adjacent to the current block. For example, in directional mode, intra-frame prediction unit 226 can typically perform a mathematical combination of the values ​​of adjacent samples and fill these calculated values ​​across the current block in a defined direction to generate a prediction block. As another example, in DC mode, intra-frame prediction unit 226 can calculate the average of the adjacent samples of the current block and generate a prediction block to include the obtained average for each sample of the prediction block.

[0088] Mode selection unit 202 provides a prediction block to residual generation unit 204. Residual generation unit 204 receives the original, uncoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for that current block. In some examples, residual generation unit 204 may also determine the difference between sample values ​​in the residual block to generate the residual block using Residual Differential Pulse Code Modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form residual generation unit 204.

[0089] In some examples, mode selection unit 202 divides a CU into PUs, each PU potentially associated with a luminance prediction unit and a corresponding chrominance prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As shown above, the size of a CU refers to the size of its luminance decoding block, and the size of a PU refers to the size of its luminance prediction unit. Assuming a specific CU has a size of 2N×2N, video encoder 200 can support PUs with sizes of 2N×2N or N×N for intra-frame prediction, and symmetrical PUs with sizes of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support PUs with sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for asymmetric partitioning for inter-frame prediction.

[0090] In some examples, the mode selection unit 202 no longer divides the CU into PUs; each CU can be associated with a luma decoding block and a corresponding chroma decoding block. As mentioned above, the size of the CU refers to the size of the luma decoding block of the CU. The video encoder 200 and the video decoder 300 can support CU sizes of 2N×2N, 2N×N, or N×2N.

[0091] For other video decoding techniques (such as intra-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding), as some examples, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples, such as palette mode decoding, mode selection unit 202 does not generate a prediction block, but instead generates syntax elements indicating how the block should be reconstructed based on the selected palette. In this mode, mode selection unit 202 can provide these syntax elements to entropy coding unit 220 for encoding.

[0092] As described above, the residual generation unit 204 receives the video data of the current block and the video data of the corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. The residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block to generate the residual block.

[0093] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply multiple transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply discrete cosine transform (DCT), direction transformation, Karhunen-Loeve transform (KLT), or conceptually similar transformations to the residual block. In some examples, transform processing unit 206 performs multiple transformations on the residual block, such as primary and secondary transformations, such as rotation transformations. In some examples, transform processing unit 206 does not perform any transformations on the residual block.

[0094] Quantization unit 208 can quantize the transform coefficients in the transform coefficient block to generate a quantized transform coefficient block. Quantization unit 208 can quantize the transform coefficients of the transform coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization causes information loss; therefore, the accuracy of the quantized transform coefficients will be lower than the accuracy of the original transform coefficients generated by transform processing unit 206.

[0095] The inverse quantization unit 210 and the inverse transform processing unit 212 can perform inverse quantization and inverse transform on the quantized transform coefficient block, respectively, to reconstruct the residual block based on the transform coefficient block. The reconstruction unit 214 generates a reconstruction block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although it may have some degree of distortion). For example, the reconstruction unit 214 can add the samples of the reconstructed residual block to the corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstruction block.

[0096] Filter unit 216 can perform one or more filtering operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce blocky artifacts along the boundaries of the CU. In some examples, the operation of filter unit 216 can be skipped.

[0097] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 can store the reconstructed blocks in the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 can store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve a reference image formed by the reconstructed (and possibly filtered) blocks from the DPB 218 to perform inter-frame prediction for blocks of subsequent encoded images. Additionally, the intra-frame prediction unit 226 can use the reconstructed blocks of the current image in the DPB 218 to perform intra-frame prediction for other blocks in that current image.

[0098] Typically, entropy coding unit 220 can entropy-encode syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy-encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy-encode predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, another example of video data, to generate entropy-encoded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, exponential-Golomb coding, or another entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in a bypass mode where syntax elements are not entropy-encoded. In some examples, the entropy coding unit 220 can encode the SPS, PPS, and / or APS, where each parameter set is independent of the others, so that the video decoder does not need to parse any syntax elements of the SPS before parsing the PPS and / or APS.

[0099] The video encoder 200 can output a bitstream containing the entropy-encoded syntax elements required to reconstruct slices or blocks of images. Specifically, the entropy coding unit 220 can output the bitstream.

[0100] The operations described above are relative to blocks. This description should be understood as operations applied to luma decoding blocks and / or chroma decoding blocks. As mentioned above, in some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the CU, respectively. In some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the PU.

[0101] In some examples, it is unnecessary to repeat the operations performed on the luma decoder for the chroma decoder block. As an example, the operations used to identify the motion vector (MV) and reference image of the luma decoder block do not need to be repeated for identifying the MV and reference image of the chroma block. More precisely, the MV of the luma decoder block can be scaled to determine the MV of the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma decoders.

[0102] Video encoder 200 represents an example of a device configured to encode video data. The device includes a memory configured to store video data and one or more processors implemented in circuitry, the processors being configured to: signal a sequence of sequences of encoded pictures once to a first set of parameters available in the bitstream of encoded video data; determine one or more DRA syntax elements applicable to at least one picture in the set of encoded pictures; and signal the second set of parameters in the bitstream of encoded video, wherein the video decoder's parsing of the one or more DRA syntax elements is independent of any syntax elements of the SPS.

[0103] Figure 3 This is a block diagram illustrating an example video decoder 300 that can perform the techniques disclosed herein. Provided for illustrative purposes. Figure 3 This should not be construed as limiting the techniques broadly illustrated and described in this disclosure. For illustrative purposes, this disclosure describes a video decoder 300 based on technologies such as VVC (ITU-T H.266 under development), HEVC (ITU-T H.265), and EVC. However, the techniques of this disclosure can be implemented by video decoding devices configured for other video decoding standards.

[0104] exist Figure 3In the example, the video decoder 300 includes a decoded image buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded image buffer (DPB) 314. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or processing circuitry. For example, units of the video decoder 300 can be implemented as one or more circuit or logic elements, as part of hardware circuitry or as part of an FPGA, processor, or ASIC. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these or other functions. In some examples, the video decoder 300 may also include components capable of performing the DRA parsing techniques of this disclosure and, later herein, relative to... Figure 13 The output DRA unit being described.

[0105] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-frame prediction unit 318. The prediction processing unit 304 may include additional units to perform predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0106] CPB memory 320 can store video data to be decoded by components of video decoder 300, such as encoded video bitstreams. The video data stored in CPB memory 320 can be, for example, from computer-readable medium 110 (…). Figure 1The CPB memory 320 may include a CPB for storing encoded video data (e.g., syntax elements) based on the encoded video bitstream. Furthermore, the CPB memory 320 may store video data other than the syntax elements of the decoded images, such as temporary data representing the output from multiple units of the video decoder 300. The DPB 314 typically stores decoded images. When decoding subsequent data and images from the encoded video bitstream, the video decoder 300 may output and / or use the aforementioned decoded images as reference video data. The CPB memory 320 and DPB 314 may be formed from any of a variety of memory devices, such as DRAM. These various memory devices include SDRAM, MRAM, RRAM, or other types of memory devices. The CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In several examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.

[0107] Alternatively or alternatively, in some examples, the video decoder 300 can be drawn from the memory 120 ( Figure 1 The decoded video data can be retrieved from the memory. In other words, memory 120 can store data together with CPB memory 320 as described above. Similarly, when some or all of the functions of video decoder 300 are implemented by software to be executed by the processing circuitry of video decoder 300, memory 120 can store instructions to be executed by video decoder 300.

[0108] right Figure 3 The various units shown are illustrated to aid in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination of both. Figure 2 Similarly, a fixed-function circuit is a circuit that provides a specific function and is pre-defined for the operations it can perform. A programmable circuit is a circuit that can be programmed to perform multiple tasks and provides flexible functionality within the operations it can perform. For example, a programmable circuit can execute software or firmware to make it operate in a manner defined by the instructions in the software or firmware. A fixed-function circuit can execute software instructions (e.g., receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more of the above units may be different circuit blocks (fixed-function or programmable), while in other examples, one or more of the above units may be integrated circuits.

[0109] The video decoder 300 may include an ALU, an EFU, digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executed on the programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0110] Entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to reconstruct the syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0111] Typically, the video decoder 300 reconstructs the image on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").

[0112] Entropy decoding unit 302 can entropy decode the syntax elements and transform information of the quantized transform coefficients that define a block of quantized transform coefficients. The transform information may be, for example, quantization parameters (QPs) and / or (multiple) transform mode indicators. In some examples, entropy decoding unit 302 can decode SPS, PPS, and / or APS, where each parameter set is independent of the others, so that video decoder 300 does not need to parse any syntax elements of the SPS before parsing the PPS and / or APS. For example, video decoder 300 can parse the SPS applicable to a block of video data. Video decoder 300 can independently parse a second parameter set (e.g., PPS or APS) applicable to the block, where parsing the second parameter set is independent of any syntax elements of the SPS. Video decoder 300 can process the block based on the SPS and the second parameter set. Parsing syntax elements or parameter sets may mean analyzing bits or bit strings to determine their values.

[0113] The inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization for application by the inverse quantization unit 306. The inverse quantization unit 306 can, for example, perform a bit-left shift operation to inverse quantize the quantized transform coefficients. The inverse quantization unit 306 can thus form a transform coefficient block including the transform coefficients.

[0114] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the transform coefficient block.

[0115] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements obtained by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is an inter-frame prediction, the motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements may indicate the reference picture from which the reference block is retrieved in the DPB 314 and a motion vector identifying the position of the reference block in the reference picture relative to the current block in the current picture. The motion compensation unit 316 can typically be configured for the motion compensation unit 224 ( Figure 2 The method described is basically similar to the one used for inter-frame prediction.

[0116] As another example, if the prediction information syntax element indicates that the current block is intra-predictive, then intra-predictive unit 318 can generate a prediction block according to the intra-predictive mode indicated by the prediction information syntax element. Similarly, intra-predictive unit 318 can typically be configured with the same syntax as intra-predictive unit 226. Figure 2 The intra-prediction process is performed in a manner substantially similar to that described above. The intra-prediction unit 318 can retrieve data from neighboring samples of the current block from the DPB 314.

[0117] Reconstruction unit 310 can use prediction blocks and residual blocks to reconstruct the current block. For example, reconstruction unit 310 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block.

[0118] Filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce blocky artifacts along the edges of the reconstructed block. The operation of filter unit 312 does not necessarily need to be performed in all examples.

[0119] The video decoder 300 can store reconstructed blocks in the DPB 314. For example, in an example where the operation of the filter unit 312 is not performed, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the operation of the filter unit 312 is performed, the filter 312 can store the filtered reconstructed blocks in the DPB 314. As described above, the DPB 314 can provide reference information to the prediction processing unit 304, such as samples of the current image for intra-frame prediction and pre-decoded images for subsequent motion compensation. Simultaneously, the video decoder 300 can output the decoded image (e.g., decoded video) from the DPB 314 for subsequent processing, such as... Figure 1 The display is presented on the display device 118.

[0120] In this manner, video decoder 300 represents an example of a video decoding device, which includes a memory configured to store video data, and one or more processors implemented in circuitry and coupled to the memory. The one or more processors are configured to: parse a first set of parameters, which is signaled once in the bitstream of encoded video data for each set of encoded pictures; parse one or more Dynamic Range Adjustment (DRA) syntax elements in a second set of parameters, which is signaled in the bitstream of encoded video data and associated with at least one picture in that set of encoded pictures, wherein the parsing of the one or more DRA syntax elements does not depend on any syntax elements of the first set of parameters; and process the at least one picture based on the first and second set of parameters.

[0121] Next-generation video applications can operate using video data representing captured scenes with High Dynamic Range (HDR) and Wide Color Gamut (WCG). The parameters used for dynamic range and color gamut are two separate properties of video content, and their specifications for digital television and multimedia services are defined by multiple international standards. For example, ITU-R Rec.709 defines parameters for High Definition Television (HDTV), such as Standard Dynamic Range (SDR) and Standard Color Gamut (SCG), while ITU-R Rec.2020 specifies parameters for Ultra High Definition Television (UHDTV), such as HDR and WCG. Other standards development organization (SDO) documents specify these properties in other systems; for example, the P3 color gamut is defined in SMPTE-231-2 from the Society of Motion Picture and Television Engineers (SMPTE), and some parameters for HDR are defined in SMPTE ST-2084. A brief introduction to the dynamic range and color gamut of video data is provided below.

[0122] Dynamic range is typically defined as the ratio between the minimum and maximum brightness of a video signal. Dynamic range is also measured in "f-stops," where one f-stop corresponds to twice the dynamic range of the signal. In the MPEG definition, HDR content is content characterized by brightness variations exceeding 16 f-stops. In some definitions, levels between 10 and 16 f-stops are considered intermediate dynamic range, but in others, they are considered HDR. Meanwhile, the human visual system (HVS) can perceive an even greater dynamic range, and the HVS includes an adaptive mechanism that narrows the range of motion, known as the synchronization range.

[0123] Video applications and services can be regulated by Rec.709 and offer SDR, typically supporting a luminance (or brightness) range of approximately 0.1 to 100 candela (cd) per m² (often referred to as "nits"), resulting in less than 10 f-stops. Next-generation video services are expected to offer a dynamic range of up to 16 f-stops. Some initial parameters are specified in SMPTE-2084 and Rec.2020.

[0124] Figure 4 These are conceptual diagrams illustrating the example dynamic range of SDR in HDTV, the expected HDR in UHDTV, and the HVS dynamic range. For example, Figure 400 shows the brightness of starlight, moonlight, indoor light, and sunlight in nits on a logarithmic scale. Figure 402, relative to Figure 400, shows the human visual range and the simultaneous HVS dynamic range (also known as steady-state dynamic range). Figure 404, relative to Figures 400 and 402, depicts the SDR used for display in HDTV. For example, the SDR used for the display may overlap with a portion of the brightness range from moonlight to indoor light, and may overlap with a portion of the HVS sync range. Figure 406, relative to Figures 400 through 404, depicts the expected HDR used for the display in UHDTV. It can be seen that this HDR is much larger than the SDR in Figure 404 and includes a larger portion of the HVS sync range than the SDR in Figure 404.

[0125] Now let's describe the color gamut. Figure 5 This is a conceptual diagram illustrating an example color gamut map. Beyond HDR, a more realistic aspect of video experience is the color dimension, which is typically defined by the color gamut. Figure 5 The examples show the visual representation of the SDR color gamut (triangle 500 based on the BT.709 red, green, and blue primary colors) and the wider color gamut used for UHDTV (triangle 502 based on the BT.2020 red, green, and blue primary colors). Figure 5 It also depicts the so-called spectral locus (defined by the tongue-shaped region 504), representing the boundaries of natural colors. For example... Figure 5As shown, the primary colors from BT.709 (Triangle 500) to BT.2020 (Triangle 502) are designed to provide approximately 70% or more of the color for UHDTV services. D65 specifies white for a given specification.

[0126] Table 1 below shows some examples of color gamut specifications.

[0127] Table 1 - Colorimetric parameters for a selected color space

[0128]

[0129] Now let's discuss the compression of HDR video data. HDR / WCG is typically acquired and stored (even with floating-point precision) for each component in a 4:4:4 chroma format and a very wide color space (e.g., XYZ). This representation targets high precision and is (almost) mathematically lossless. However, this format characterizes many redundancies and is not optimal for compression purposes. Lower-precision formats with an HVS-based assumption are often used in state-of-the-art video applications.

[0130] Figure 6 This is a block diagram illustrating an example format conversion technique. The video encoder 200 can perform a format conversion technique to transform linear RGB 510 into HDR data 518. These techniques may include, for example... Figure 6 The three main elements described are: 1) a non-linear transfer function (TF) for dynamic range compression 512; 2) color conversion to a more compact or robust color space 514; and 3) floating-point to integer representation conversion (quantization 516).

[0131] Figure 7 This is a block diagram illustrating an example of reverse format conversion technology. The video decoder 300 or another processor can perform this operation. Figure 7 The inverse transformation technology includes inverse quantization 522, inverse color conversion 524, and inverse TF 526 to inverse transform HDR data 520 into linear RGB 528.

[0132] An example transfer function (TF) is now described. The TF is applied to data to compress its dynamic range and make it possible to represent the data with a finite number of bits. This function is typically a one-dimensional (1D) nonlinear function that reflects the inverse of the electro-optical transfer function (EOTF) for end-user displays specified in Rec. 709 for SDR, or approximates HVS perception as a change in luminance as the PQ TF specified for HDR in SMPTE-2084. The inverse process of the OETF is mapping the decoding level back to the luminance of the EOTF (electro-optical transfer function). Figure 8 This is a conceptual diagram illustrating an example of EOTF. Figure 8 Chart 530 depicts the EOTF using linear brightness on the y-axis and decoding levels on the x-axis. Figure 8 An example EOTF is the HDR EOTF of the ST2084.

[0133] The SMPTE ST-2084 specification defines the EOTF application as follows. TF is applied to the normalized linear R, G, B values ​​to obtain the nonlinear representation R'G'B'. SMPTE ST-2084 defines normalization by NORM = 10000, which is associated with a peak luminance of 10000 nits (cd / m2).

[0134]

[0135] in

[0136] Where L is the normalized linear R value, G value, or B value.

[0137]

[0138]

[0139]

[0140]

[0141]

[0142] Figure 9 This is a conceptual diagram showing a visualization of the PQ TF (SMPTE ST-2084EOTF). Figure 9 The PQ TF shown in Figure 540 depicts the input values ​​(linear color values) normalized to the range 0…1 and the output values ​​(non-linear color values) normalized to the range 1. As can be seen from the curve, one percent of the dynamic range of the input signal (low illumination) is converted to 50% of the dynamic range of the output signal.

[0143] Typically, EOTF is defined as a function with floating-point precision, so if the inverse TF (so-called OETF) is applied, no error will be introduced into the signal having this non-linear relationship. The inverse TF (OETF) specified in SMPTE ST-2084 is defined as the inverse PQ function:

[0144]

[0145] in

[0146] Where N is the value of R', G', or B'.

[0147]

[0148]

[0149]

[0150]

[0151]

[0152] Utilizing floating-point precision, consecutive applications of EOTF and OETF provide error-free, perfect reconstructions. However, this representation is not optimal for streaming or broadcast services. A more compact representation of nonlinear R'G'B' data with fixed-bit precision is described in the following paragraphs.

[0153] It should be noted that EOTF and OETF are very effective research topics, and the TF used in some HDR video decoding systems may be different from SMPTE ST-2084.

[0154] Now let's discuss color transformation. Since RGB data is generated by image capture sensors, it is commonly used as input. However, this color space has high redundancy between components and is not optimal for compact representation. To achieve a more compact and robust representation, the RGB components are typically converted to a more uncorrelated color space, better suited for compression, such as YCbCr. This color space separates luminance in the form of lightness (Y) and color information (Cb and Cr) in different uncorrelated components.

[0155] For modern video decoding systems, the commonly used color space is YCbCr, as specified in ITU-R BT.709 or ITU-R BT.709. The YCbCr color space in the Broadcast Services (Television) BT.709 standard specifies the following conversion process from R'G'B' to Y'CbCr (non-constant luminance representation):

[0156]

[0157] The above can also be achieved using the following approximate transformation that avoids the division between Cb and Cr components:

[0158]

[0159]

[0160] The ITU-R BT.2020 standard specifies the following conversion process from R'G'B' to Y'CbCr (non-constant luminance representation):

[0161]

[0162] The above can also be achieved using the following approximate transformation that avoids the division between Cb and Cr components:

[0163]

[0164] It should be noted that both color spaces are kept normalized; therefore, for input values ​​normalized in the range 0...1, the resulting values ​​will be mapped to the range 0...1. Typically, color transformations implemented with floating-point precision provide perfect reconstruction, making the process lossless.

[0165] Now let's discuss quantization / fixed-point conversion. The aforementioned processing stages can be implemented using floating-point precision, thus it can be considered lossless. However, for most consumer electronics applications, this type of precision can be considered redundant and expensive. In some examples, input data in the target color space is converted to target bit-depth fixed-point precision. Specific studies have shown that 10-12 bit precision combined with PQ TF is sufficient to provide 16 f-stop HDR data, where distortion is below the Just-Noticeable Difference (JDD). The JDD is the amount of change that must be made to make a difference perceptible or detectable at least half the time. Data represented with 10-bit precision can be further decoded using most state-of-the-art video decoding solutions. This conversion process includes signal quantization, an element of lossy decoding, and a source of inaccuracies introduced into the converted data.

[0166] The following illustrates an example of such quantization applied to codewords in a target color space (e.g., YCbCr). The input value YCbCr, expressed in floating-point precision, is converted into a signal with a fixed bit depth BitDepthY for the Y value and a BitDepthC for the chromaticity values ​​(Cb, Cr), where D... Y' D represents the converted luminance signal. Cb and D Cr Represents the converted chroma signal

[0167]

[0168]

[0169] Where x represents the rounding operation formula in Equation 7 for a given luminance or chromaticity component.

[0170] Round(x)=Sign(x)*Floor(Abs(x)+0.5),

[0171] Sign(x)=-1 if x<0,0 if x=0,1 if x>0

[0172] Floor(x) is the largest integer less than or equal to x.

[0173] Abs(x)=x if x>=0,-x if x<0

[0174] Clip1 Y (x) = Clip3(0, (1< <BitDepth Y )-1,x)

[0175] Clip1 C (x) = Clip3(0, (1< <BitDepth C )-1,x)

[0176] Clip3(x,y,z) = x if z<x,y if z> y,z and others

[0177] Dynamic Range Adjustment (DRA) will now be discussed. DRA was originally proposed in SEI (D. Rusanovskyy, AK Ramasubramonian, D. Bugdayci, S. Lee, J. Sole, M. Karczewicz, VCEG document COM16-C 1027-E, September 2015) to enable high dynamic range video decoding with backward compatibility.

[0178] The authors suggest implementing the DRA as a piecewise linear function f(x), defined for a set of non-overlapping dynamic range partitions (ranges) {Ri} of the input value x, where i is the index of a range ranging from 0 to N-1 (inclusive), and N is the total number of ranges {Ri} used to define the DRA function. For example, the range of the DRA is defined by the minimum and maximum x values ​​belonging to the range Ri, e.g., [x i ,x i+1 -1], where x i and x i+1 Representing the range R respectively i and R i+1 The minimum value. Applied to the Y color component (luma) of video, the DRA function Sy is obtained by a scale S. y,i and a bias O y,i Defined, it is applied to every x∈[x i ,x i+1-1], therefore S y ={S y,i O y,i}

[0179] In this way, for any Ri, and for each x ∈ [x i ,x i+1 -1], the output value X is calculated as follows:

[0180] X = S y,i *(xO y,i (8)

[0181] For the inverse DRA mapping process of the luminance component Y implemented at the decoder (e.g., video decoder 300) or another processor, the DRA function Sy is determined by the scaling factor S. y,i The reciprocal and bias O y,i The value is defined and applied to each X∈[X] i ,X i+1 -1].

[0182] In this way, for any Ri, and for every X∈[X i ,X i+1 -1], the reconstructed value x is calculated as follows:

[0183] x = X / S y,i +O y,i (9)

[0184] The forward DRA mapping process for chromaticity components Cb and Cr is defined as follows. An example is given, where the term "u" represents a sample of the Cb color component belonging to the range Ri, u ∈ [u...]. i ,u i+1 -1], therefore S u ={S u,i O u,i}:

[0185] U = S u,i *(uO y,i )+Offset (10)

[0186] Where Offset equals 2 (bitdepth-1) This indicates the bias of the bipolar Cb and Cr signals.

[0187] The inverse DRA mapping process performed at the decoder for the chromaticity components Cb and Cr is defined as follows. An example is given, where the U term represents a sample of the remapped Cb color component belonging to the range Ri, U∈[U... i U i+1 -1]:

[0188] u = (U - Offset) / Su,i +O y,i (11)

[0189] Where Offset equals 2 (bitdepth-1) This indicates the bias of the bipolar Cb and Cr signals.

[0190] The topic of luminance-driven chroma scaling (LCS) will now be discussed. In JCTVC-W0101, HDRCE2: On CE2.a-1 LCS (AK Ramasubramonian, J. Sole, D. Rusanovskyy, D. Bugdayci, M. Karczewicz), a technique for adjusting chroma information (e.g., Cb and Cr) by utilizing luminance information associated with processed chroma samples is disclosed. Similarly, the DRA method for enabling high dynamic range video decoding with backward compatibility through dynamic range adjustment (SEI) (D. Rusanovskyy, AK Ramasubramonian, D. Bugdayci, S. Lee, J. Sole, M. Karczewicz, VCEG document COM16-C 1027-E, September 2015) proposes applying a scaling factor S of Cb to the chroma samples. u and Cr's S v,i However, the DRA function is not defined as in equations (3) and (4) for a range {R} that can be taken by chromaticity values ​​u or v. i Piecewise linear function S u ={S u,i O u,i The LCS method proposes using the luminance value Y to derive a scaling factor for the chrominance samples. In this way, the forward LCS mapping of the chrominance sample u (or v) is implemented as follows:

[0191] U = S u,i (Y)*(u-Offset)+Offset (12)

[0192] The inverse LCS process implemented on the decoder side is defined as follows:

[0193] u = (U - Offset) / S u,i (Y)+Offset (13)

[0194] More specifically, for a given pixel located at (x,y), the LCS function S is used to calculate the value of the pixel from its luminance value Y'(x,y). Cb (or S) Cr The derived factor scales the chromaticity samples Cb(x,y) or Cr(x,y).

[0195] At the forward LCS of the chroma sample (e.g., video encoder 200), the Cb (or Cr) value and its associated luminance value Y' are treated as the chroma scaling function S. Cb (or S) Cr The input is taken as Cb or Cr, and as shown in equation (14), Cb or Cr is converted to Cb' and Cr'. On the decoder side (e.g., video decoder 300), the inverse LCS is applied, and the reconstructed Cb' or Cr' is converted to Cb or Cr, as shown in equation (15).

[0196] Cb'(x,y)=S cb (Y'(x,y))*Cb(x,y),

[0197] Cr'(x, y) = S cr (Y'(x,y))*Cr(x,y) (14)

[0198]

[0199]

[0200] Figure 10 This is a conceptual diagram illustrating an example of the LCS function. Using the LCS function 550, the chromaticity component of pixels with smaller luminance values ​​is multiplied by a smaller scaling factor.

[0201] The relationship between DRA sample scaling and the QP of the video decoder will now be discussed. To adjust the compression ratio of the video encoder, block transform-based video decoding schemes (such as HEVC) utilize scalar quantizers applied to the block transform coefficients.

[0202] Xq = X / scalerQP

[0203] Where Xq is the quantized code value of the transform coefficient X generated by applying the scalerQP derived from the QP parameters. In most decoders, the quantized code value is approximated as an integer value (e.g., through rounding). In some decoders (e.g., video encoder 200 or video decoder 300), quantization can be a different function that depends not only on QP but also on other parameters of the decoder.

[0204] The scaler value, scalerQP, is controlled by the quantization parameter (QP), where the relationship between QP and the scalar quantizer is defined as follows, where k is a known constant:

[0205] scalerQP=k*2^(QP / 6) (16)

[0206] The relationship between the inverse function and the scalar quantizer applied to the transform coefficients and the QP of HEVC is defined as follows:

[0207] QP=ln(scalerQP / k)*6 / ln(2); (17)

[0208] Correspondingly, an additive change in the QP value, such as deltaQP, will result in a multiplicative change in the scalerQP value applied to the transformation coefficients.

[0209] DRA effectively applies the scaleDRA value to the pixel sample value, and taking into account the transformation properties, it can be combined with the scalerQP value as follows:

[0210] Xq = T(scaleDRA*x) / scaleQP

[0211] Here, Xq is the quantized transform coefficient generated by scaling the transform T of the scaled x-sample values ​​using scaleQP applied to the transform domain. Therefore, applying the multiplier scaleDRA in the pixel domain produces an effective change in the scaler quantizer scaleQP, which is then applied to the transform domain. This can be interpreted, in turn, as an additive change in the QP parameters applied to the currently processed data block:

[0212] dQP=log2(scaleDRA)*6; (18)

[0213] dQP is an approximate QP bias introduced by HEVC by deploying DRA on the input data.

[0214] We now discuss the dependency of the chroma QP on the luma QP value. Some state-of-the-art video decoding designs, such as HEVC and newer designs, can utilize a predefined dependency between the luma and chroma QP values ​​that are effectively applied to the block Cb processing the current decoding. This dependency can be used to achieve optimal bitrate allocation between the luma and chroma components.

[0215] Examples of this dependency are illustrated in Tables 8-10 of the HEVC specification, where the QP values ​​applied to decode chroma samples are derived from the QP values ​​used to decode luma samples. The chroma QP values ​​are derived from the relevant portion of the QP values ​​of the corresponding luma samples (the QP values ​​applied to the block / TU, corresponding to the luma sample to which the chroma QP value belongs). The chroma QP biases and Tables 8-10 of the HEVC specification are reproduced below:

[0216] When ChromaArrayType is not equal to 0, the following applies:

[0217] – Variable qP Cb and qP Cr The following was exported:

[0218] – If tu_residual_act_flag[xTbY][yTbY] equals 0, then the following applies:

[0219] qPi Cb =Clip3(-QpBdOffset) C ,57,Qp Y +pps_cb_qp_offset+slice_cb_qp_offset+CuQpOffset Cb (8-287)

[0220] qPi Cr =Clip3(-QpBdOffset) C ,57,Qp Y +pps_cr_qp_offset+slice_cr_qp_offset+CuQpOffset Cr (8-288)

[0221] – Otherwise (tu_residual_act_flag[xTbY][yTbY] equals 1), the following applies:

[0222] qPi Cb =Clip3(-QpBdOffsetC,57,QpY+PpsActQpOffsetCb+slice_act_cb_qp_offset+CuQpOffsetCb) (8-289)

[0223] qPi Cr =Clip3(-QpBdOffsetC,57,QpY+PpsActQpOffsetCr+slice_act_cr_qp_offset+CuQpOffsetCr) (8-290)

[0224] – If ChromaArrayType equals 1, then the variable qP Cb and qP Cr Based on the index qPi, they are respectively equal to qPi Cb and qPi Cr It is set to the value of Qpc as specified in Tables 8 to 10.

[0225] Otherwise, based on index qPi being equal to qPi respectively Cb and qPi Cr variable qP Cb and qP Cr It is set to equal Min(qPi,51).

[0226] – A colorimetric parameter used for the Cb and Cr components, Qp' Cb and Qp' Cr The following was exported:

[0227] Qp′ Cb =qP Cb +QpBdOffset C (8-291)

[0228] Qp′ Cr =qP Cr +QpBdOffset C (8-292)

[0229] Figure 11 This is a conceptual diagram illustrating an example table of functions where Qpc is specified as qPi. In some examples, Figure 11 The tables are tables 8 through 10 of the HEVC specification. Tables 8 through 10 (560) detail the functions for which Qpc is specified as qPi when ChromaArrayType is equal to 1.

[0230] The derivation of chroma scaling of DRA will now be discussed. In a video decoding system (e.g., video encoder 200 or video decoder 300) employing uniform scalar quantization in both the transform domain and the pixel domain scaled by DRA, the derivation of the proportional DRA value applied to the chroma sample (Sx) can depend on the following:

[0231] -S Y : The brightness ratio value of the associated brightness sample.

[0232] -S CX : The proportion derived from the color gamut of the content, where CX replaces Cb or Cr as needed.

[0233] -S corr : Correction scaling term, used to resolve mismatches in transform decoding and DRA scaling, for example, to compensate for the dependencies introduced by HEVC in Tables 8 to 10.

[0234] S X =fun(S) Y ,S CX ,S corr ).

[0235] An example is a separable function defined as follows: SX = f1(S Y )*f2(S CX )*f3(S CX ), where f1 is the first part of the separable function, f2 is the second part of the separable function, and f3 is the third part of the separable function.

[0236] The bumping process is now described. A decoded picture buffer (DPB) (e.g., DPB 218 or DPB 314) maintains a set of pictures / frames that can be used as references for inter-frame prediction in the decoding loop of a decoder (e.g., video encoder 200 or video decoder 300). Depending on the decoding state, one or more pictures can be output for use by or read by an external application. Depending on the decoding order, DPB size, or other conditions, pictures no longer used in the decoding loop and no longer used by external applications can be removed from the DPB 314 ( Figure 3 Remove or replace with a new reference image. Output the DPB 314 image and the DPB( Figure 3 The process of potentially removing images in HEVC is called the bumping process. An example of the bumping process defined for HEVC is shown below:

[0237] C.5.2.4 "Collision" Process

[0238] The "collision" process includes the following ordered steps:

[0239] 1. Select the image with the smallest PicOrderCntVal value from all images in DPB as the first image to be output and mark it as "needs to be output".

[0240] 2. Use the consistent cropping window specified for the image in the effective SPS to crop the image, output the cropped image, and mark the image as "not required for output".

[0241] 3. When the image storage cache, which includes cropped and output images, contains images marked as "not for reference", the image storage cache is cleared.

[0242] Note – For any two images (picA and picB) belonging to the same CVS [decoded video sequence] and output by the “collision process”, when picA is output earlier than picB, the value of PicOrderCntVal of picA is less than the value of PicOrderCntVal of picB.

[0243] The collision process using DRA will now be discussed. The DRA specification post-processing, in the form of a modified collision process, is adopted in the draft text of the EVC specification. (In the markup...) <change> and< / change> The text shown here is an excerpt from the EVC specification text containing a varying collision process.

[0244] Appendix C Assuming a Reference Decoder

[0245] The HRD [hypothetical reference decoder] includes a decoded picture buffer (CPB), a transient decoding process, a decoded picture buffer (DPB), an output DRA [which is an additional process related to some video decoder], and a crop, as shown in Figure C-2 [reproduced in this disclosure]. Figure 12 ]. Figure 12 This is a conceptual diagram illustrating an example of an HDR buffer model.

[0246] Sub-clause C.3 specifies the operation of DPB. Sub-clauses C.3.3 and C.5.2.4 specify the output DRA process and trimming.

[0247] C.3.3 Image Decoding and Output

[0248] Image n is decoded and its DPB output time t o,dpb (n) is derived from the following formula.

[0249] t o,dpb (n)=t r (n)+t c *dpb_output_delay(n) (C-12)

[0250] The output of the current image is specified as follows.

[0251] –If t o,dpb (n)=t r If (n), then the current image is output.

[0252] –otherwise(t) o,dpb (n)>t r (n)), the current image will be output later and stored in the DPB (as specified in sub-clause C.2.4), and at time t o,dpb (n) is output until time t is reached. o,dpb (n) The time before decoding or inference is equal to 1 if no_output_of_prior_pics_flag indicates that it should not be output.

[0253] <change> The output image should be exported by calling the DRA procedure specified in sub-clause 8.9.2 and cropped using the clipping rectangle specified for the sequence in SPS.< / change>

[0254] When image n is the image to be output and is not the last image in the output bitstream, t o,dpb The value of (n) is defined as:

[0255] Δt o,dpb (n)=t o,dpb (n n )-t o,dpb (n) (C-13)

[0256] Where n nIndicates the image that follows image n in the output order.

[0257] The decoded image is stored in the DPB.

[0258] C.5.2.4 "Collision" Process

[0259] The "collision" procedure is invoked in the following situations.

[0260] - As specified in sub-clause C.5.2.2, the current image is an IDR image, and no_output_of_prior_pics_flag is not equal to 1.

[0261] - As specified in the sub-clause, there is no empty image storage buffer (i.e., the DPB fullness is equal to the DPB size), and an empty image storage buffer is required to store decoded images.

[0262] The "collision" process includes the following ordered steps:

[0263] <change>

[0264] 4. Select the image with the smallest PicOrderCntVal value from all images in DPB as the first image to be output and mark it as "needs to be output".

[0265] The selected image comprises an array of pic_width_in_luma_samples multiplied by pic_height_in_luma_samples for luminance samples currPicL and two arrays of PicWidthInSamplesC multiplied by PicHeightInSamplesC for chrominance samples currPicCb and currPicCr. The sample arrays currPicL, currPicCb, and currPicCr correspond to the decoded sample array S. L S Cb and S Cr .

[0266] 5. When dra_table_present_flag equals 1, the DRA derivation procedure specified in Item 8.9 is invoked, taking the selected image as input and the output image as output; otherwise, the sample array of the output image is initialized by the sample array of the selected image.< / change>

[0267] 6. Use the consistent cropping window specified for the image in the effective SPS to crop the output image, output the cropped image, and mark the image as "not needed for output".

[0268] 7. When including being <change> Mapping< / change> When the image storage buffer of the cropped and output image contains images marked as "not for reference", the image storage buffer is cleared.

[0269] The APS signaling for DRA data will now be discussed. The EVC specification defines how DRA parameters are signaled in the Adaptive Parameter Set (APS). The syntax and semantics of DRA parameters are provided below:

[0270]

[0271]

[0272]

[0273] DRA Data Syntax

[0274]

[0275] A value of 1 for sps_dra_flag indicates that dynamic range adjustment mapped to the output samples is used. A value of 0 for sps_dra_flag indicates that dynamic range adjustment mapped to the output samples is not used.

[0276] A value of 1 for `pic_dra_enabled_present_flag` indicates that `pic_dra_enabled_flag` exists in PPS. A value of 0 for `pic_dra_enabled_present_flag` indicates that `pic_dra_enabled_flag` does not exist in PPS. When `pic_dra_enabled_flag` does not exist, it is inferred to be equal to 0.

[0277] `pic_dra_enabled_flag` equal to 1 specifies that DRA is enabled for all decoded images in the reference PPS. `pic_dra_enabled_flag` equal to 0 specifies that DRA is not enabled for all decoded images in the reference PPS. When `pic_dra_enabled_flag` does not exist, it is inferred to be equal to 0.

[0278] pic_dra_aps_id specifies the adaptation_parameter_set_id of the DRA APS that is enabled for the decoded image of the reference PPS.

[0279] The adaptation_parameter_set_id provides an identifier for APS to reference by other syntax elements.

[0280] aps_params_type specifies the type of APS parameter carried in the APS as specified in Table 6 [as shown in Table 2 in this disclosure].

[0281]

[0282] Table 2—APS Parameter Type Codes and APS Parameter Types

[0283] The value of dra_descriptor1 should be in the range of 0 to 15 (inclusive). In the current version of the specification, the value of the syntax element dra_descriptor1 is restricted to 4, and other values ​​are reserved for future use.

[0284] `dra_descriptor2` specifies the precision of the decimal part of the DRA scaling parameter signaling and reconstruction process. The value of `dra_descriptor2` should be in the range of 0 to 15 (inclusive). In the current version of the specification, the value of the syntax element `dra_descriptor2` is limited to 9; other values ​​are reserved for future use.

[0285] The variable numBitsDraScale is exported as follows:

[0286] numBitsDraScale=dra_descriptor1+dra_descriptor2

[0287] The value of dra_number_ranges_minus1 plus 1 specifies the number of ranges described by the signaling notification for the DRA table. The value of dra_number_ranges_minus1 should be in the range of 0 to 31 (inclusive).

[0288] A dra_equal_ranges_flag of 1 specifies that the DRA table is derived using ranges of equal size, where the size is specified by the syntax element dra_delta_range[0]. A dra_equal_ranges_flag of 0 specifies that the DRA table is derived using dra_number_ranges, where the size of each of these ranges is specified by the syntax element dra_delta_range[j].

[0289] dra_global_offset specifies the starting codeword position used to derive the DRA table, and initializes the variable inDraRange[0] as follows:

[0290] inDraRange[0]=dra_global_offset

[0291] The number of bits used for signaling notification dra_global_offset is BitDepth. Y Bit.

[0292] `dra_delta_range[j]` specifies the size of the j-th range in the codeword used to derive the DRA table. The value of `dra_delta_range[j]` should be between 1 and (1 < 1). <BitDepth Y The range of 1 (including endpoint values).

[0293] For j within the range of 1 to dra_number_ranges_minus1 (inclusive), the variable inDraRange[j] is derived as follows:

[0294] inDraRange[j]=inDraRange[j–1]+(dra_equal_ranges_flag==1)?

[0295] dra_delta_range[0]:dra_delta_range[j]

[0296] The bitstream consistency requirement is that inDraRange[j] should be in the range of 0 to (1 << BitDepth Y ) - 1.

[0297] dra_scale_value[j] specifies the DRA scaling value associated with the j-th range of the DRA table. The number of bits used to signal dra_scale_value[j] is equal to numBitsDraScale.

[0298] dra_cb_scale_value specifies the scaling value for the chroma samples of the Cb component used to derive the DRA table. The number of bits used to signal dra_cb_scale_value is equal to numBitsDraScale. In the current version of the specification, the value of the syntax element dra_cb_scale_value should be less than 4 << dra_descriptor2, and other values are reserved for future use.

[0299] dra_cr_scale_value specifies the scaling value for the chroma samples of the Cr component used to derive the DRA table. The number of bits used to signal dra_cr_scale_value is equal to numBitsDraScale bits. In the current version of the specification, the value of the syntax element dra_cb_scale_value should be less than 4 << dra_descriptor2, and other values are reserved for future use.

[0300] The values of dra_scale_value[j], dra_cb_scale_value, and dra_cr_scale_value should not be equal to 0.

[0301] dra_table_idx specifies the access entry of the ChromaQpTable used to derive the chroma scaling value. The value of dra_table_idx should be in the range of 0 to 57 (including the endpoint values).

[0302] In some examples of the MPEG EVC specification, the signaling of DRA parameters is based on the parsing dependencies of specific syntax elements in the SPS (Signal Set of Parameters) that are signaled in the SPS (Signal Set of Parameters) for specific syntax elements in the PPS (Picture Header) and APS (Adaptive Parameter Set). These dependencies are undesirable because the video decoder 300 has to wait for the SPS before parsing the PPS and / or APS, which introduces decoding latency. For example, the parsing of the syntax elements pic_dra_enabled_present_flag, pic_dra_enabled_flag, and pic_dra_aps_id may depend on the value of the syntax element sps_dra_flag in the SPS. For example, the parsing of the syntax elements dra_global_offset and dra_delta_range[j] may depend on the value of the syntax element bit_depth_luma_minus8 in the SPS, because bit_depth_luma_minus8 can be used to determine the bit depth of dra_global_offset and dra_delta_range[j].

[0303] The first example of parsing dependencies existing in PPS is shown in the syntax table below, where dependencies on syntax elements in SPS are marked with tags. <mark> and< / mark> between.

[0304]

[0305]

[0306] A second example of parsing dependencies present in the APS DRA is shown in the syntax table and semantics of the DRA APS:

[0307]

[0308] dra_global_offset specifies the starting codeword position used to derive the DRA table, and initializes the variable InDraRange[0] as follows:

[0309] InDraRange[0]=dra_global_offset (7-70)

[0310] `dra_delta_range[j]` specifies the size of the j-th range in the codeword used to derive the DRA table. The value of `dra_delta_range[j]` should be between 1 and (1 < 1). <BitDepth Y The range of 1 (including endpoint values).

[0311] For j within the range of 1 to dra_number_ranges_minus1+1 (inclusive), the variable InDraRange[j] is derived as follows:

[0312] InDraRange[j]=

[0313] InDraRange[j–1]+(dra_equal_ranges_flag==1)? dra_delta_range[0]:dra_delta_range[j-1] (7-71)

[0314] Bitstream consistency requirement InDraRange[j] should be between 0 and (1 < 0). <BitDepth Y Within the range of 1 (including endpoint values).

[0315] The number of bits used for signaling notifications dra_global_offset and dra_delta_range[j] is <mark>BitDepth Y < / mark> Bit.

[0316] The variable BitDepth Y The following is derived from the SPS syntax elements:

[0317] bit_depth_luma_minus8 specifies the luminance array BitDepth Y The range of bit depth and brightness quantization parameters of the sample, and the offset QpBdOffset. Y The values ​​are as follows:

[0318] BitDepth Y =8+bit_depth_luma_minus8 (7-3)

[0319] In the latter case of (APS), it is understood that bit depth dependency is advantageous for DRA designs that can accommodate multiple bit depths (e.g., 12-bit or 16-bit data). However, as mentioned above, having this dependency is undesirable.

[0320] One possible solution to this problem is to introduce an additional syntax element that can identify the number of bits in the syntax element used for signaling notification of the global DRA offset (e.g., dra_global_offset) and the bit depth indicating brightness (e.g., bit_depth_luma_minus8). However, this design is also undesirable because it actually increases the bit rate and complexity of the video encoder 200 and video decoder 300 by introducing additional syntax elements.

[0321] To address the aforementioned problems, the following changes to the design in EVC are disclosed based on the technology of this disclosure. In some examples, these techniques may also be applied to other video decoder designs. For example, the video encoder 200 and video decoder 300 may be configured to operate according to the following changes.

[0322] According to the technology disclosed herein, the dependency of DRA syntax elements in PPS and / or APS on SPS can be removed. In this way, the video decoder 300 can: parse the SPS applicable to a block of video data; parse one or more Dynamic Range Adjustment (DRA) syntax elements in a second parameter set applicable to that block, wherein the parsing of one or more DRA syntax elements does not depend on any syntax elements in the SPS; and process the block based on the SPS and the second parameter set. In some examples, the SPS and the second parameter set are in different NAL units.

[0323] 1) Remove dependencies in PPS:

[0324] In some examples, the second parameter set is the PPS. One or more DRA syntax elements of the second parameter set may include a first DRA syntax element (e.g., pic_dra_enabled_present_flag) indicating the presence of a second DRA syntax element (which itself indicates whether DRA is enabled) in the PPS, a second DRA syntax element (e.g., pic_dra_enabled_flag), or a third DRA syntax element (e.g., pic_dra_aps_id) indicating an identifier of the DRA adaptive parameter set. In some examples, the video decoder 300 may parse the PPS before parsing syntax elements in the SPS indicating whether DRA is used on output samples mapped to blocks. By making the DRA syntax elements in the PPS independent of the syntax elements in the SPS, the video decoder 300 can parse the PPS even if the video decoder 300 has not yet received the SPS or has not parsed the SPS.

[0325] To remove SPS syntax element dependencies in PPS within EVC, you can remove the following from the tag: <delete> and< / delete> The conditions between. Additionally, this can be achieved by removing the following from the markers shown below. <delete> and< / delete> The `pic_dra_enabled_present_flag` can be used to reduce the number of control syntax elements. For example, video encoder 200 may not signal `pic_dra_enabled_present_flag` and video decoder 300 may not be certain about the following deletion conditions.

[0326]

[0327] 2) Remove resolution dependencies in APS:

[0328] In some examples, the second parameter set is an adaptive parameter set. In some examples, one or more DRA syntax elements include at least one syntax element whose value indicates the starting codeword position for deriving the DRA table, or whose value indicates at least one syntax element in the codeword used to derive the DRA table. In some examples, the video decoder 300 may parse the second parameter set before parsing the syntax element in the SPS that indicates the luminance bit depth. In some examples, the first DRA syntax element of one or more DRA syntax elements has a predetermined fixed bit depth (e.g., 10 bits). In some examples, the first DRA syntax element is a syntax element whose value indicates the starting codeword position for deriving the DRA table. In some examples, the first DRA syntax element is a syntax element whose value indicates the range in the codeword used to derive the DRA table. By making the DRA syntax elements in the DRA APS have a fixed bit depth instead of making the bit depth of the DRA syntax elements based on the values ​​of the syntax elements in the SPS, the video decoder 300 can still parse the DRA APS even if the video decoder 300 has not yet received the SPS or has not parsed the SPS.

[0329] To eliminate parsing dependencies in DRA APS, the number of bits used for signaling notification syntax elements can be fixed, for example, 10 bits, and the video decoder 300 can determine the actual value of the syntax element based on the bit depth of the signaling notification determined during the DRA reconstruction process. Since the syntax element interpretation process is performed after parsing, no parsing dependency is introduced. For example, the video encoder 200 can use a fixed number of bits for signaling notification syntax elements, such as those marked below. The markings will be changed in the following text. <change> and< / change> between.

[0330] dra_global_offset specifies the starting codeword position used to derive the DRA table, that is, it defines the value of the variable InDraRange[0]. <change> The number of bits used for signaling notification dra_global_offset is 10 bits.< / change>

[0331] dra_delta_range[j] specifies the size of the j-th range in the codeword used to derive the DRA table, where j is in the range from 1 to dra_number_ranges_minus1+1. <change> The number of bits used for signaling notification dra_delta_range[j] is 10 bits.< / change> The value of dra_delta_range[j] should be between 1 and Min(1023, (1< <BitDepth Y Within the range of )-1) (including endpoint values).

[0332] The variable InDraRange[0] is initialized as follows:

[0333] <change>InDraRange[0]=dra_global_offset< <Max(0,BitDepth Y –10)< / change> (7-70)

[0334] For j within the range of 1 to dra_number_ranges_minus1+1 (inclusive), the variable InDraRange[j] is derived as follows:

[0335] deltaRange=(dra_equal_ranges_flag==1)? dra_delta_range[0]:dra_delta_range[j–1]

[0336] InDraRange[j]=InDraRange[j–1]+ <change>(deltaRange< <Max(0,BitDepth Y –10))< / change> (7-71)

[0337] Bitstream consistency requirement InDraRange[j] should be between 0 and (1 < 0). <BitDepth Y Within the range of 1 (including endpoint values).

[0338] Figure 13 This is a block diagram of a video encoder and video decoder system including a DRA unit. A video encoder, such as video encoder 200, may include a forward DRA unit 240 and a decoder core 242. In some examples, the decoder core 242 may include... Figure 2 The units depicted in the text can be referenced as above. Figure 2 The operation described above. The video encoder 200 can also determine multiple APS 244 and multiple PPS 246, which may include information from the forward DRA unit 240.

[0339] According to the technology disclosed herein, the video encoder 200 can signal the SPS (Syntax Parameter Set) applicable to a block of video data. The forward DRA unit 240 can determine one or more DRA syntax elements from a second parameter set applicable to that block, wherein the video decoder's parsing of the one or more DRA syntax elements is independent of any syntax elements in the SPS. The video encoder 200 can signal the second parameter set.

[0340] A video decoder, such as video decoder 300, may include a decoder core 340 and an output DRA unit 342. In some examples, the decoder core 340 may include... Figure 3 The units depicted in the text can be referenced as above. Figure 3 The operation described above. The video decoder 300 may also determine multiple APS 344 and multiple PPS 346 that may include information to be used by the output DRA unit 342.

[0341] According to the technology disclosed herein, video decoder 300 can parse a Sequence Parameter Set (SPS) applicable to a block of video data. Output DRA unit 342 can parse one or more Dynamic Range Adjustment (DRA) syntax elements in a second parameter set applicable to the block, wherein the parsing of one or more DRA syntax elements does not depend on any syntax elements of the SPS. Output DRA unit 342 can process the block based on the SPS and the second parameter set.

[0342] Figure 14 This is a flowchart illustrating the dynamic range adjustment parameter parsing technique according to this disclosure. The video decoder 300 can parse a first set of parameters that is signaled once per bitstream of encoded video data for each set of encoded picture sequences (580). For example, the video decoder 300 can receive a SPS in the bitstream and can parse the SPS to determine the values ​​of the syntax elements included within the SPS. The SPS can be applied to picture sequences because the values ​​of the syntax elements within the SPS can be used when decoding the picture sequence. In other words, the SPS can be applied to sequences of video data comprising multiple pictures.

[0343] The video decoder 300 can parse one or more Dynamic Range Adjustment (DRA) syntax elements in a second parameter set. The second parameter set is signaled in the bitstream of the encoded video data. The second parameter set is associated with at least one picture in the set of encoded pictures, wherein the parsing of one or more DRA syntax elements does not depend on any syntax elements in the first parameter set (582). In this way, the parsing of the second parameter set is independent of the applicable first parameter set. For example, the video decoder 300 can parse the second parameter set independently of the first parameter set, and may parse the second parameter set without depending on or relying on any syntax elements in the first parameter set.

[0344] The video decoder 300 can process the at least one image (584) based on a first parameter set and a second parameter set. For example, the video decoder 300 can utilize the values ​​of syntax elements within the first and second parameter sets when processing the at least one image.

[0345] In some examples, the first parameter set and the second parameter set reside in different Network Abstraction Layer (NAL) units. In some examples, the video decoder 300 may parse the second parameter set before parsing the first parameter set, as part of independently parsing the second parameter set. In some examples, the first parameter set is a Sequence Parameter Set (SPS) and the second parameter set is a Picture Parameter Set (PPS). In some examples, one or more DRA syntax elements include one or more of the following: i) a first DRA syntax element indicating the presence of the second DRA syntax element in the PPS, ii) a second DRA syntax element indicating whether DRA is enabled, or iii) a third DRA syntax element indicating an identifier for a third parameter set describing the DRA parameters.

[0346] In some examples, as part of parsing one or more DRA syntax elements, the video decoder 300 may parse the PPS before parsing the syntax element in the SPS that indicates whether dynamic range adjustment on the output samples mapped to the block is used. For example, the video decoder 300 may parse the PPS before parsing the sps_dra_flag.

[0347] In some examples, the first parameter set is a Sequence Parameter Set (SPS) and the second parameter set is an Adaptive Parameter Set (APS). In some examples, one or more DRA syntax elements include at least one of a first syntax element whose value indicates the starting codeword position for deriving the DRA table or a second syntax element whose value indicates the range in the codeword used to derive the DRA table. In some examples, as part of parsing one or more DRA syntax elements, the video decoder 300 may parse the APS before parsing the syntax element in the SPS that indicates the luma bit depth. In some examples, the DRA syntax elements in the APS have a predetermined fixed bit depth. In some examples, the DRA syntax elements include a syntax element whose value indicates the starting codeword position for deriving the DRA table. For example, a DRA syntax element may include dra_global_offset. In some examples, the DRA syntax elements include a syntax element whose value indicates the range in the codeword used to derive the DRA table. For example, a DRA syntax element may include dra_delta_range[j]. In some examples, the predetermined fixed bit depth is 10 bits.

[0348] Figure 15 This is a flowchart illustrating an example method for encoding the current block. The current block may include the current CU. Although referencing video encoder 200 ( Figure 1 and Figure 2 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 15 Similar to the method.

[0349] In this example, the video encoder 200 initially makes a prediction for the current block (350). For example, the video encoder 200 may form a prediction block for the current block. Next, the video encoder 200 may compute a residual block for the current block (352). To compute the residual block, the video encoder 200 may calculate the difference between the original uncoded block and the prediction block for the current block. Afterward, the video encoder 200 may transform the residual block and quantize the transform coefficients of the residual block (354). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (358). For example, the video encoder 200 may use CAVLC or CABAC to encode the transform coefficients. Then, the video encoder 200 may output the entropy-encoded data of the block (360).

[0350] Figure 16 This is a flowchart illustrating an example method for decoding the current block of video data. The current block may include the current CU. Although the reference video decoder 300 ( Figure 1 and Figure 3 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 16 A similar approach.

[0351] The video decoder 300 can receive entropy-coded data for the current block, such as entropy-coded prediction information and entropy-coded data for the transform coefficients of the residual block corresponding to the current block (370). The video decoder 300 can entropy decode the entropy-coded data to determine the prediction information for the current block and reproduce the transform coefficients of the residual block (372). The video decoder 300 can, for example, use an intra-frame or inter-frame prediction mode indicated by the prediction information for the current block to predict the current block (374) to compute a prediction block for the current block. Then, the video decoder 300 can perform an inverse scan on the reproduced transform coefficients (376) to create a block of quantized transform coefficients. Next, the video decoder 300 can inverse quantize the transform coefficients and apply the inverse transform to the transform coefficients to produce a residual block (378). The video decoder 300 can finally decode the current block by combining the prediction block and the residual block (380).

[0352] By removing the dependency of the second parameter set (e.g., PPS or APS) on the syntax elements of SPS, the technique disclosed herein facilitates the video decoder 300 in parsing the second parameter set without waiting for SPS in the received bitstream. Therefore, the technique disclosed herein can reduce decoding latency.

[0353] This disclosure contains the following terms.

[0354] Item 1A. A method for processing video data, the method comprising: suppressing a pic_dra_enabled_present_flag in a determined set of image parameters; and processing the video data according to the set of image parameters.

[0355] Clause 2A. A method for processing video data, the method comprising: fixing the bit depth of a syntax element in a dynamic range adjustment adaptive parameter set to a predetermined bit depth; signaling the syntax element with the predetermined bit depth; and processing the video data based on the syntax element.

[0356] Clause 3A. The method described in accordance with Clause 2A, wherein the syntax element is dra_global_offset or dra_delta_range[j].

[0357] Clause 4A. The method described in accordance with Clause 2A or Clause 3A, wherein the predetermined number of digits is 10.

[0358] Clause 5A. The method described in Clause 1A, wherein the processing includes decoding.

[0359] Clause 6A. The method according to any one of Clauses 1A-5A, wherein the processing includes encoding.

[0360] Clause 7A. An apparatus for processing video data, the apparatus comprising one or more components for performing the method described in any one of Clauses 1A-6A.

[0361] Clause 8A. The device pursuant to Clause 7A, wherein the one or more components include one or more processors implemented in a circuit.

[0362] Clause 9A. The device pursuant to any one of Clauses 7A and 8A further includes a memory for storing video data.

[0363] Clause 10A. The device pursuant to any one of Clauses 7A-9A further includes a display configured to display decoded video data.

[0364] Clause 11A. The device pursuant to any one of Clauses 7A-10A, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device, or set-top box.

[0365] Clause 12A. The device according to any one of Clauses 7A-11A, wherein the device includes a video decoder.

[0366] Clause 13A. The device according to any one of Clauses 7A-12A, wherein the device includes a video encoder.

[0367] Clause 14A. A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the method described in any one of Clauses 1A-6A.

[0368] Clause 15A. An apparatus for encoding video data, the apparatus comprising: components for suppressing the determination of a pic_dra_enabled_present_flag in a set of picture parameters; and components for processing the video data based on the set of picture parameters.

[0369] Clause 16A. An apparatus for encoding video data, the apparatus comprising: means for fixing the bit depth of syntax elements in a dynamic range adjustment adaptive parameter set to a predetermined bit depth; means for signaling the syntax elements with a predetermined bit depth; and means for processing the video data based on the syntax elements.

[0370] Clause 1B. A method for processing video data, the method comprising: parsing a first set of parameters, the first set of parameters being signaled once in a bitstream of encoded video data for each sequence of encoded pictures in a set; parsing one or more dynamic range adjustment (DRA) syntax elements in a second set of parameters, the second set of parameters being signaled in a bitstream of encoded video data and associated with at least one picture in the set of encoded pictures, wherein parsing of the one or more DRA syntax elements is independent of any syntax elements in the first set of parameters; and processing the at least one picture based on the first set of parameters and the second set of parameters.

[0371] Clause 2B. The method according to Clause 1B, wherein the first parameter set and the second parameter set are in different network abstraction layer units.

[0372] Clause 3B. The method described in Clause 1B or 2B, wherein parsing one or more DRA syntax elements includes parsing a second set of parameters before parsing a first set of parameters.

[0373] Clause 4B. The method according to any combination of Clauses 1B-3B, wherein the first parameter set is a sequence parameter set (SPS) and the second parameter set is a picture parameter set (PPS).

[0374] Clause 5B. The method described in Clause 4B, wherein one or more DRA syntax elements include one or more of the following syntax elements: i) a first DRA syntax element indicating whether a second DRA syntax element exists in the PPS, ii) a second DRA syntax element indicating whether DRA is enabled, or iii) a third DRA syntax element indicating an identifier of a third set of parameters describing DRA parameters.

[0375] Clause 6B. The method according to Clause 4B, wherein parsing one or more DRA syntax elements includes parsing the PPS before parsing the syntax element in the SPS that indicates whether the DRA mapped to the output sample of the at least one picture is used.

[0376] Clause 7B. The method according to any combination of Clauses 1B-3B, wherein the first parameter set is a sequence parameter set (SPS) and the second parameter set is an adaptive parameter set (APS).

[0377] Clause 8B. The method according to Clause 7B, wherein one or more DRA syntax elements include at least one of a first syntax element whose value indicates the starting codeword position for deriving the DRA table or a second syntax element whose value indicates the range of codewords for deriving the DRA table.

[0378] Clause 9B. The method described in accordance with Clause 7B or Clause 8B, wherein parsing one or more DRA syntax elements includes parsing the APS before parsing the syntax element in the SPS that indicates the luminance bit depth.

[0379] Clause 10B. The method according to any combination of Clauses 7B-9B, wherein the first DRA syntax element of one or more DRA syntax elements has a predetermined fixed bit depth.

[0380] Clause 11B. The method according to Clause 10B, wherein the first DRA syntax element includes a syntax element whose value indicates the starting codeword position for deriving the DRA table.

[0381] Clause 12B. The method according to Clause 10B, wherein the first DRA syntax element includes a syntax element in its value indicator codeword used to derive the range of the DRA table.

[0382] Clause 13B. The method according to any combination of Clauses 10B-12B, wherein the predetermined fixed bit depth is 10 bits.

[0383] Clause 14B. An apparatus for processing video data, the apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: parse a first set of parameters, the first set of parameters being signaled once in a bitstream of encoded video data for each sequence of a set of encoded pictures; parse one or more dynamic range adjustment (DRA) syntax elements in a second set of parameters, the second set of parameters being signaled in a bitstream of encoded video data and associated with at least one picture in the set of encoded pictures, wherein the parsing of the one or more DRA syntax elements is independent of any syntax element in the first set of parameters; and process the at least one picture based on the first set of parameters and the second set of parameters.

[0384] Clause 15B. The device as described in Clause 14B, wherein the first parameter set and the second parameter set are in different network abstraction layer units.

[0385] Clause 16B. The apparatus as described in Clause 14B or Clause 15B, wherein, as part of parsing one or more DRA syntax elements, one or more processors are configured to: parse a second set of parameters before parsing a first set of parameters.

[0386] Clause 17B. The device described under any combination of Clauses 14B-16B, wherein the first parameter set is a sequence parameter set (SPS) and the second parameter set is a picture parameter set (PPS).

[0387] Clause 18B. The device pursuant to Clause 17, wherein one or more DRA syntax elements include one or more of the following syntax elements: i) a first DRA syntax element indicating the presence of a second DRA syntax element in the PPS, ii) a second DRA syntax element indicating whether DRA is enabled, or iii) a third DRA syntax element indicating an identifier of a third set of parameters describing DRA parameters.

[0388] Clause 19B. The apparatus as described in Clause 17B, wherein, as part of parsing one or more DRA syntax elements, one or more processors are configured to: parse the PPS before parsing the syntax element in the SPS that indicates whether dynamic range adjustment mapped to the output sample of the at least one picture is used.

[0389] Clause 20B. The device described under any combination of Clauses 14B-16B, wherein the first set of parameters is a sequence parameter set (SPS) and the second set of parameters is an adaptive parameter set (APS).

[0390] Clause 21B. The device according to Clause 20B, wherein one or more DRA syntax elements include at least one of a first syntax element whose value indicates the starting codeword position for deriving the DRA table or a second syntax element whose value indicates the range of codewords for deriving the DRA table.

[0391] Clause 22B. The apparatus according to Clause 20B or Clause 21B, wherein, as part of parsing one or more DRA syntax elements, one or more processors are configured to: parse the APS before parsing the syntax element in the SPS that indicates the luminance bit depth.

[0392] Clause 23B. The device according to any combination of Clauses 20B-22B, wherein the first DRA syntax element of one or more DRA syntax elements has a predetermined fixed bit depth.

[0393] Clause 24B. The apparatus according to Clause 23B, wherein the first DRA syntax element includes a syntax element whose value indicates the starting codeword position for deriving the DRA table.

[0394] Clause 25B. The device as described in Clause 23B, wherein the first DRA syntax element includes a syntax element in its value indicator codeword used to derive the range of the DRA table.

[0395] Clause 26B. The device described under any combination of Clauses 23B-25B, wherein the predetermined fixed bit depth is 10 bits.

[0396] Clause 27B. A non-provisional computer-readable storage medium storing instructions, which, when executed, cause one or more processors to: parse a first set of parameters, the first set of parameters being signaled once in a bitstream of encoded video data for each set of encoded pictures; parse one or more dynamic range adjustment (DRA) syntax elements in a second set of parameters, the second set of parameters being signaled in a bitstream of encoded video data and associated with at least one picture in that set of encoded pictures, wherein the parsing of the one or more DRA syntax elements is independent of any syntax element in the first set of parameters; and process the at least one picture based on the first set of parameters and the second set of parameters.

[0397] Clause 28B. An apparatus for processing video data, the apparatus comprising: a component for parsing a first set of parameters, the first set of parameters being signaled once in a bitstream of encoded video data for each sequence of encoded pictures in a set; a component for parsing one or more dynamic range adjustment (DRA) syntax elements in a second set of parameters, the second set of parameters being signaled in a bitstream of encoded video data and associated with at least one picture in the set of encoded pictures, wherein the parsing of the one or more DRA syntax elements is independent of any syntax element in the first set of parameters; and a component for processing the at least one picture based on the first set of parameters and the second set of parameters.

[0398] It should be recognized that, depending on the examples, certain actions or events of any of the techniques described herein may be performed in a different order, or may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for practicing these techniques). Furthermore, in some examples, actions or events may be performed concurrently, for example, through multithreading, interrupt handling, or multiple processors, rather than sequentially.

[0399] In one or more examples, the functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium or a communication medium. A computer-readable storage medium corresponds to a tangible medium such as a data storage medium. A communication medium includes any medium that facilitates, for example, the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures to implement the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0400] By way of example and not limitation, the aforementioned computer-readable storage media may include: RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and is accessible to a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the aforementioned coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. The disks and optical discs used herein include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically copy data magnetically, while optical discs copy data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0401] Instructions can be executed by one or more processors, such as one or more signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuit" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combinational decoder. Similarly, the techniques can be fully implemented in one or more circuit or logic elements.

[0402] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Multiple components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Specifically, as described above, multiple units can be combined in a decoder hardware unit, or provided by a set of interoperable hardware units, including one or more processors as described above, in conjunction with appropriate software and / or firmware.

[0403] Several examples have been described. These and other examples are all within the scope of the appended claims.

Claims

1. A method for processing video data, the method comprising: The Sequence Parameter Set (SPS) is parsed, which is signaled once in the bitstream of the encoded video data for each set of encoded images; One or more Dynamic Range Adjustment (DRA) syntax elements in a Picture Parameter Set (PPS) are parsed, the PPS being signaled in the bitstream of the encoded video data and associated with at least one picture in the set of encoded pictures, wherein the parsing of the one or more DRA syntax elements in the PPS is independent of any syntax elements of the SPS, the one or more DRA syntax elements in the PPS including one or more of the following: i) a first DRA syntax element indicating whether DRA is enabled, or ii) a second DRA syntax element indicating an identifier of an Adaptive Parameter Set (APS); Parsing one or more DRA syntax elements in the APS, the APS being signaled in the bitstream of the encoded video data and associated with at least one of the set of encoded pictures, wherein parsing the one or more DRA syntax elements in the APS does not depend on any syntax elements of the SPS, the one or more DRA syntax elements in the APS including at least one of a first syntax element whose value indicates the starting codeword position for deriving the DRA table or a second syntax element whose value indicates the range of codewords for deriving the DRA table; and The at least one image is processed based on the SPS, the PPS, and the APS.

2. The method according to claim 1, wherein the SPS and the PPS are in separate network abstraction layer units.

3. The method of claim 1, wherein parsing the one or more DRA syntax elements in the PPS includes parsing the PPS before parsing the SPS.

4. The method of claim 1, wherein parsing the one or more DRA syntax elements in the PPS includes parsing the PPS before parsing the syntax elements in the SPS that indicate whether a DRA mapping is used on the output sample of the at least one image.

5. The method of claim 1, wherein parsing the one or more DRA syntax elements in the APS includes parsing the APS before parsing the syntax elements in the SPS that indicate the luminance bit depth.

6. The method of claim 1, wherein the first DRA syntax element of the one or more DRA syntax elements in the APS has a predetermined fixed bit depth.

7. The method of claim 6, wherein the first DRA syntax element in the APS includes a syntax element whose value indicates the starting codeword position for deriving the DRA table.

8. The method of claim 6, wherein the first DRA syntax element in the APS includes a syntax element in its value indicator codeword for deriving the range of the DRA table.

9. The method according to claim 6, wherein the predetermined fixed bit depth is 10 bits.

10. An apparatus for processing video data, the apparatus comprising: The memory is configured to store the video data; as well as One or more processors are implemented in the circuit and coupled to the memory, the one or more processors being configured to: The Sequence Parameter Set (SPS) is parsed, which is signaled once in the bitstream of the encoded video data for each set of encoded images; One or more Dynamic Range Adjustment (DRA) syntax elements in a Picture Parameter Set (PPS) are parsed, the PPS being signaled in the bitstream of the encoded video data and associated with at least one picture in the set of encoded pictures, wherein the parsing of the one or more DRA syntax elements in the PPS is independent of any syntax elements in the SPS, the one or more DRA syntax elements in the PPS including one or more of the following: i) a first DRA syntax element indicating whether DRA is enabled, or ii) a second DRA syntax element indicating an identifier of an Adaptive Parameter Set (APS); Parsing one or more DRA syntax elements in the APS, the APS being signaled in the bitstream of the encoded video data and associated with at least one of the set of encoded pictures, wherein parsing the one or more DRA syntax elements in the APS does not depend on any syntax elements of the SPS, the one or more DRA syntax elements in the APS including at least one of a first syntax element whose value indicates the starting codeword position for deriving the DRA table or a second syntax element whose value indicates the range of codewords for deriving the DRA table; and The at least one image is processed based on the SPS, the PPS, and the APS.

11. The device of claim 10, wherein the SPS and the PPS are in separate network abstraction layer units.

12. The device of claim 10, wherein, as part of parsing the one or more DRA syntax elements in the PPS, the one or more processors are configured to: The PPS is parsed before the SPS is parsed.

13. The device of claim 10, wherein, as part of parsing the one or more DRA syntax elements in the PPS, the one or more processors are configured to: The PPS is parsed before parsing the syntax element in the SPS that indicates whether the DRA mapping on the output sample of the at least one image is used.

14. The apparatus of claim 10, wherein, as part of parsing the one or more DRA syntax elements in the APS, the one or more processors are configured to: The APS is parsed before parsing the syntax element indicating the luminance bit depth in the SPS.

15. The device of claim 10, wherein the first DRA syntax element of the one or more DRA syntax elements in the APS has a predetermined fixed bit depth.

16. The apparatus of claim 15, wherein the first DRA syntax element in the APS includes a syntax element whose value indicates the starting codeword position for deriving the DRA table.

17. The device of claim 15, wherein the first DRA syntax element in the APS includes a syntax element in its value indicator codeword for deriving the range of the DRA table.

18. The device of claim 15, wherein the predetermined fixed bit depth is 10 bits.

19. A non-transitory computer-readable medium storing instructions, which, when executed, cause one or more processors to: The Sequence Parameter Set (SPS) is parsed, which is signaled once in the bitstream of the encoded video data for each set of encoded images; One or more Dynamic Range Adjustment (DRA) syntax elements in a Picture Parameter Set (PPS) are parsed, the PPS being signaled in the bitstream of the encoded video data and associated with at least one picture in the set of encoded pictures, wherein the parsing of the one or more DRA syntax elements in the PPS is independent of any syntax elements of the SPS, the one or more DRA syntax elements in the PPS including one or more of the following: i) a first DRA syntax element indicating whether DRA is enabled, or ii) a second DRA syntax element indicating an identifier of an Adaptive Parameter Set (APS); Parsing one or more DRA syntax elements in the APS, the APS being signaled in the bitstream of the encoded video data and associated with at least one of the set of encoded pictures, wherein parsing the one or more DRA syntax elements in the APS does not depend on any syntax elements of the SPS, the one or more DRA syntax elements in the APS including at least one of a first syntax element whose value indicates the starting codeword position for deriving the DRA table or a second syntax element whose value indicates the range of codewords for deriving the DRA table; and The at least one image is processed based on the SPS, the PPS, and the APS.

Citation Information

Patent Citations

  • Fixed point implementation of range adjustment of components in video coding

    CN108028936A

  • Method of signaling availability of profile_idc flag and decoding thereof

    JP2005341184A

  • Parameter set groups for coded video data

    US20130114694A1