Decoded Picture Buffer (DPB) parameter signaling notification for video decoding
By signaling the DPB parameters in the Sequence Parameter Set (SPS), the decoding problem of the video decoder when there is no signaling notification in the Video Parameter Set (VPS) is solved, and the correct reconstruction of video data is achieved.
Patent Information
- Application Number
- CN202180010298.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-27
- Filing Date
- 2021-01-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-01-28
AI Technical Summary
In some cases, if the video decoder fails to signal the decoded picture buffer (DPB) parameters in the video parameter set (VPS), the video decoder may fail to correctly reconstruct the video data.
The video decoder selectively signals DPB parameters in the Sequence Parameter Set (SPS), especially when the VPS is unavailable, by referencing the SPS as the only layer of the Output Layer Set (OLS) to achieve signaling notification of DPB parameters.
This ensures that the video decoder can correctly reconstruct the video data, avoiding decoding errors caused by a lack of signaling notification parameters.
Smart Images

Figure CN115004712B_ABST
Abstract
Description
[0001] This application claims the benefit of U.S. Application No. 17 / 159,508, filed January 27, 2021; U.S. Provisional Application No. 62 / 967,507, filed January 29, 2020; and U.S. Provisional Application No. 63 / 004,022, filed April 2, 2020, the entire contents of any one of which are incorporated herein by reference. U.S. Application No. 17 / 159,508 claims the benefit of U.S. Provisional Application No. 62 / 967,507, filed January 29, 2020; and U.S. Provisional Application No. 63 / 004,022, filed April 2, 2020. Technical Field
[0002] This disclosure relates to video encoding and video decoding. Background Technology
[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones, so-called "smartphones," video conferencing equipment, and video streaming devices. Digital video devices implement video decoding technologies, such as those described in the standards defined below: MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Decoding (AVC), ITU-T H.265 / High-Efficiency Video Decoding (HEVC), and extensions to these standards. By implementing such video decoding technologies, video devices can more efficiently send, receive, encode, decode, and / or store digital video information.
[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate inherent redundancy in video sequences. For block-based video decoding, video slices (e.g., video pictures or portions of video pictures) can be divided into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in intra-frame decoded (I) slices of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in inter-frame decoded (P or B) slices of a picture can use spatial prediction relative to reference samples in adjacent blocks within the same picture, or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention
[0005] This disclosure typically describes techniques for signalalling DPB parameters in video decoding. DPB parameters can be included in a DPB parameter set, and various aspects used to construct the DPB can be specified, such as the DPB size, maximum number of picture reorderings, and maximum latency. In some examples, the video decoder (e.g., a video encoder or video decoder) can signal the DPB parameters within a video parameter set (VPS). However, in certain situations (e.g., where only a single layer is decoded), the video decoder may not signal the VPS. When the VPS is not signaled, problems may arise when the video decoder attempts to reference the DPB parameters from the VPS.
[0006] According to one or more techniques of this disclosure, a video decoder can selectively signal DPB parameters outside the VPS. For example, if DPB parameters would be unavailable in the VPS, the video decoder can signal DPB parameters in the Sequence Parameter Set (SPS). As an example, the video decoder can signal DPB parameters in the SPS, where the SPS is referenced by a layer that is the only layer in the OLS (i.e., the OLS has only one layer, or only one layer is encoded in the bitstream). In this way, the techniques of this disclosure enable the video decoder to avoid attempting to reference parameters from parameter sets that have never been signaled.
[0007] In one example, the method includes decoding the sequence parameter set (SPS) of the current bitstream of video data from the decoded picture buffer (DPB) parameter syntax structure when the layer is referenced as the only layer of the output layer set (OLS); and reconstructing the video data represented by the current bitstream based on the DPB parameter syntax structure.
[0008] In another example, the device includes a memory configured to store at least a portion of a decoded video bitstream; and one or more processors implemented in the circuit and configured to: decode a decoded picture buffer (DPB) parameter syntax structure from the SPS when the sequence parameter set (SPS) of the decoded video bitstream is referenced by a layer that is the only layer of the output layer set (OLS); and reconstruct the video data represented by the current bitstream based on the DPB parameter syntax structure.
[0009] In another example, the device includes components for decoding the sequence parameter set (SPS) of the current bitstream of video data from the decoded picture buffer (DPB) parameter syntax structure when the layer is referenced as the only layer of the output layer set (OLS); and components for reconstructing the video data represented by the current bitstream based on the DPB parameter syntax structure.
[0010] In another example, a computer-readable storage medium stores instructions that, when executed, cause one or more processors to: decode the sequence parameter set (SPS) of the current bitstream of video data from the decoded picture buffer (DPB) parameter syntax structure when the layer is referenced as the only layer of the output layer set (OLS); and reconstruct the video data represented by the current bitstream based on the DPB parameter syntax structure.
[0011] Details of one or more examples are set forth in the accompanying drawings and description below. Other features, objectives, and advantages will become apparent from the description, drawings, and claims. Attached Figure Description
[0012] Figure 1 This is a block diagram illustrating an example video encoding and decoding system that can perform the techniques of this disclosure.
[0013] Figure 2A and Figure 2B This is a conceptual diagram illustrating an example quadtree binary tree (QTBT) structure and its corresponding decoding tree unit (CTU).
[0014] Figure 3 This is a block diagram illustrating an example video encoder that can perform the techniques of this disclosure.
[0015] Figure 4 This is a block diagram illustrating an example video decoder that can perform the techniques of this disclosure.
[0016] Figure 5 This is a flowchart illustrating an example method for encoding the current block according to one or more techniques of this disclosure.
[0017] Figure 6 This is a flowchart illustrating an example method for decoding the current block according to one or more techniques of this disclosure.
[0018] Figure 7 This is a flowchart illustrating an example method for signaling notification of a decoded picture buffer (DPB) parameter syntax structure according to one or more techniques of this disclosure. Detailed Implementation
[0019] Figure 1This is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. The techniques of this disclosure generally relate to decoding (encoding and / or decoding) video data. Typically, video data includes any data used for processing video. Therefore, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling notification data.
[0020] like Figure 1 As shown, in this example, system 100 includes a source device 102 that provides encoded video data for decoding and display by a target device 116. Specifically, source device 102 provides video data to target device 116 via computer-readable medium 110. Source device 102 and target device 116 can include any of a wide range of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, handheld telephone devices such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and target device 116 may be equipped for wireless communication and may therefore be referred to as wireless communication devices.
[0021] exist Figure 1 In the example, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Target device 116 includes an input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of target device 116 can be configured to apply techniques for decoding a decoded picture buffer (DPB) structure. Therefore, source device 102 represents an example of a video encoding device, while target device 116 represents an example of a video decoding device. In other examples, the source device and target device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, target device 116 may interface with an external display device, rather than including an integrated display device.
[0022] like Figure 1The system 100 shown is merely an example. Generally, any digital video encoding and / or decoding device can perform techniques for decoding the DPB structure. Source device 102 and target device 116 are merely examples of such decoding devices, where source device 102 generates decoded video data for transmission to target device 116. In this disclosure, "decoding device" refers to a device that performs the decoding (encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices, specifically, examples of a video encoder and a video decoder, respectively. In some examples, source device 102 and target device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and target device 116 includes video encoding and decoding components. Therefore, system 100 can support one-way or two-way video transmission between source device 102 and target device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0023] Typically, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a continuous series of pictures (also referred to as "frames") of video data to video encoder 200, which encodes the data for these pictures. Video source 104 of source device 102 may include a video capture device such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from the received order (sometimes referred to as "display order") into a decoding order for decoding. Video encoder 200 may generate a bitstream comprising encoded video data. The source device 102 can then output encoded video data to a computer-readable medium 110 via the output interface 108 for reception and / or retrieval, for example, by the input interface 122 of the target device 116.
[0024] The memory 106 of source device 102 and the memory 120 of target device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data, such as raw video from video source 104 and raw, decoded video data from video decoder 300. Additionally or alternatively, memories 106 and 120 may store, for example, software instructions executable by video encoder 200 and video decoder 300, respectively. Although memories 106 and 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store, for example, encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106 and 120 may be allocated as one or more video buffers, for example, to store raw, decoded, and / or encoded video data.
[0025] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded video data from source device 102 to target device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to transmit encoded video data directly to target device 116 in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, output interface 108 can modulate the transmitted signal including the encoded video data, while input interface 122 can demodulate the received transmitted signal. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, switch, base station, or any other equipment that may be useful in facilitating communication from source device 102 to target device 116.
[0026] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, target device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 can include any of a variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0027] In some examples, source device 102 may output encoded video data to file server 114 or to another intermediate storage device that may store the encoded video generated by source device 102. Target device 116 may access the stored video data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing and sending encoded video data to target device 116. File server 114 may represent (e.g., for a website) a web server, file transfer protocol (FTP) server, content delivery network device, or network attached storage (NAS) device. Target device 116 may access the encoded video data from file server 114 via any standard data connection including an Internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., digital subscriber line (DSL), cable modem, etc.), or a combination of both suitable for accessing encoded video data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.
[0028] Output interface 108 and input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data such as encoded video data according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, and 5G. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to operate according to specifications such as IEEE 802.11, IEEE 802.15 (e.g., ZigBee). TM ),Bluetooth TM Other wireless standards, such as those used for transmitting encoded video data, may be employed. In some examples, source device 102 and / or target device 116 may include corresponding system-on-chip (SoC) devices. For instance, source device 102 may include an SoC device for performing functions attributed to video encoder 200 and / or output interface 108, while target device 116 may include an SoC device for performing functions attributed to video decoder 300 and / or input interface 122.
[0029] The techniques disclosed herein can be applied to video decoding in a variety of multimedia applications that support any of the following: over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as HTTP-based Dynamic Adaptive Streaming (DASH), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0030] The input interface 122 of the target device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling notification information defined by the video encoder 200 and also used by the video decoder 300, such as syntax elements having values describing the characteristics and / or processing of video blocks or other decoded units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays a decoded image of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device.
[0031] Although not in Figure 1 As shown, but in some examples, the video encoder 200 and video decoder 300 may each be integrated with the audio encoder and / or audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams that include both audio and video in a common data stream. Where applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol or other protocols, such as User Datagram Protocol (UDP).
[0032] The video encoder 200 and video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute these instructions in hardware using one or more processors to perform the technology of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.
[0033] The video encoder 200 and video decoder 300 can operate according to video decoding standards such as ITU-T H.265, also known as High Efficiency Video Decoding (HEVC), or its extensions such as Multi-View and / or Scalable Video Decoding Extensions. Alternatively, the video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards such as the Joint Exploratory Test Model (JEM) or ITU-T H.266, also known as Multi-Functional Video Decoding (VVC). At the 17th meeting of the Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Brussels, BE, June 7-17, 2020, JVET-S2001-v11, Bross et al. described a draft VVC standard (hereinafter referred to as "VVC Draft 8") in "Multi-Functional Video Decoding (Draft 8)". However, the technology disclosed herein is not limited to any particular decoding standard.
[0034] Typically, video encoder 200 and video decoder 300 can perform block-based decoding of images. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during encoding and / or decoding). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, video encoder 200 and video decoder 300 can decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, video encoder 200 and video decoder 300 can decode luminance and chrominance components, rather than decoding red, green, and blue (RGB) data for samples of an image, where chrominance components may include both red and blue chrominance components. In some examples, video encoder 200 converts received RGB-formatted data to YUV representation before encoding, while video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and post-processing units (not shown) can perform these conversions.
[0035] This disclosure can generally relate to the decoding (e.g., encoding and decoding) of images, and is intended to include the process of encoding or decoding data of an image. Similarly, this disclosure can relate to the decoding of blocks of an image to include the process of encoding or decoding data for the blocks, such as prediction and / or residual decoding. Encoded video bitstreams typically include a series of values for syntax elements that represent decoding decisions (e.g., decoding modes) and the segmentation of images into blocks. Therefore, references to decoding images or blocks should generally be understood as decoding the values of syntax elements that form an image or block.
[0036] HEVC defines various blocks, including decoding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (such as a video encoder 200) partitions the decoding tree unit (CTU) into CUs according to a quadtree structure. That is, the video decoder partitions the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and the CU of such leaf nodes can include one or more PUs and / or one or more TUs. The video decoder can further partition the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a partition of the TU. In HEVC, the PU represents inter-frame prediction data, while the TU represents residual data. CUs with intra-frame prediction include intra-frame prediction information, such as intra-frame mode indication.
[0037] As another example, the video encoder 200 and video decoder 300 can be configured to operate according to JEM or VVC. According to JEM or VVC, the video decoder (such as the video encoder 200) segments the image into multiple decoding tree units (CTUs). The video encoder can segment the CTUs according to a tree structure such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple segmentation types, such as the distinction between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level of segmentation based on quadtree segmentation and a second level of segmentation based on binary tree segmentation. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to the decoding units (CUs).
[0038] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning. A ternary tree partition is a partition where a block is divided into three sub-blocks. In some examples, a ternary tree partition divides a block into three sub-blocks without partitioning the original block through a center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.
[0039] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).
[0040] The video encoder 200 and video decoder 300 can be configured to use quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structures according to HEVC. For illustrative purposes, the description of the techniques of this disclosure is presented with reference to QTBT segmentation. However, it should be understood that the techniques of this disclosure can also be applied to video decoders configured to use quadtree segmentation or other types of segmentation.
[0041] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a brick can refer to a rectangular area of a row of CTUs within a specific tile in an image. A tile can be a rectangular area of CTUs within a specific tile column and a specific tile row in an image. A tile column refers to a rectangular area of CTUs with a height equal to the height of the image and a width specified by a syntax element (e.g., such as in an image parameter set). A tile row refers to a rectangular area of CTUs with a height specified by a syntax element (e.g., such as in an image parameter set) and a width equal to the width of the image.
[0042] In some examples, a tile can be divided into multiple bricks, where each brick can include one or more CTU rows within the tile. A tile that is not divided into multiple bricks can also be referred to as a brick. However, bricks that are a true subset of a tile cannot be referred to as a tile.
[0043] Bricks in an image can also be arranged in slices. A slice can be an integer number of bricks from an image that can be exclusively contained within a single Network Abstraction Layer (NAL) unit. In some examples, a slice consists of a continuous sequence of multiple complete tiles or a single complete tile.
[0044] This disclosure uses "N×N" and "N multiplied by N" interchangeably to refer to the sample dimension of a block (such as a CU or other video block) in terms of both vertical and horizontal dimensions, for example, 16×16 samples or 16 by 16 samples. Typically, a 16×16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal and vertical directions. For example, a CU may include N x M samples, where M is not necessarily equal to N.
[0045] The video encoder 200 encodes video data for a control unit (CU), representing prediction and / or residual information, as well as other information. Prediction information indicates how the CU will be predicted to form a prediction block for the CU. Residual information typically represents the point-to-point difference between the samples of the CU before encoding and the prediction block.
[0046] To predict the Cubic Frame (CU), the video encoder 200 typically forms prediction blocks for the CU using either inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU from data of a previously decoded image, while intra-frame prediction generally refers to predicting the CU from previously decoded data of the same image. To perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate prediction blocks. The video encoder 200 can typically perform motion search to identify a reference block that closely matches the CU, for example, based on the difference between the CU and a reference block. The video encoder 200 can use sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to compute difference metrics to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.
[0047] Some examples of JEM and VVC also provide an affine motion compensation mode, which can be considered an inter-frame prediction mode. In the affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion such as zooming in or out, rotation, perspective motion, or other irregular motion types.
[0048] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of JEM and VVC provide sixty-seven intra-frame prediction modes, including various directional modes, as well as planar and DC modes. Generally, the video encoder 200 selects the following intra-frame prediction mode: describing the neighboring samples of the current block (e.g., a block of a CU) to predict the samples of the current block from the neighboring samples. Assuming the video encoder 200 decodes the CTU and CU in raster scan order (from left to right, from top to bottom), such samples can typically be located above, to the upper left, or to the left of the current block in the same picture as the current block.
[0049] The video encoder 200 encodes data representing the prediction mode used for the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data representing which of the various available inter-frame prediction modes is used, and the motion information used for the corresponding mode. For example, for unidirectional or bidirectional inter-frame prediction, the video encoder 200 may encode motion vectors using Advanced Motion Vector Prediction (AMVP) or merge modes. For affine motion compensation modes, the video encoder 200 may use similar modes to encode motion vectors.
[0050] Following the prediction of a block, such as after intra-frame or inter-frame prediction of the block, the video encoder 200 can compute residual data for the block. Residual data, such as a residual block, represents the sample-by-sample difference between the block and the predicted block formed using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to generate transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply a Discrete Cosine Transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 can apply secondary transforms after the primary transform, such as Mode-dependent Inseparable Subtransform (MDNSST), Signal-dependent Transform, Karl Henle-Loup Transform (KLT), etc. The video encoder 200 generates transform coefficients after applying one or more transforms.
[0051] As described above, after any transformation used to generate the transform coefficients, the video encoder 200 can perform quantization on the transform coefficients. Quantization generally refers to the process in which transform coefficients are quantized to reduce the amount of data used to represent them, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 can round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bit-by-bit right shift on the value to be quantized.
[0052] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix containing the quantized transform coefficients. The scan can be designed to place higher-energy (and therefore lower-frequency) transform coefficients before the vector and lower-energy (and therefore higher-frequency) transform coefficients after the vector. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to generate a serialized vector, and then entropy-encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy-encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic decoding (CABAC). The video encoder 200 can also entropy-encode values for syntax elements that describe metadata associated with the encoded video data and used by the video decoder 300 when decoding the video data.
[0053] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. The context may involve, for example, whether the neighboring values of a symbol are zero. Probability determination can be based on the context assigned to the symbols.
[0054] The video encoder 200 can further generate, for example, block-based syntax data, image-based syntax data, and sequence-based syntax data, or other syntax data such as sequence parameter sets (SPS), image parameter sets (PPS), or video parameter sets (VPS) from image headers, block headers, and slice headers, which are then passed to the video decoder 300. The video decoder 300 can similarly decode this syntax data to determine how to decode the corresponding video data.
[0055] In this manner, the video encoder 200 can generate a bitstream that includes encoded video data, such as syntax elements describing the segmentation of images into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.
[0056] Generally, the video decoder 300 performs the reverse process of the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC, although in a substantially similar manner to the CABAC encoding process of the video encoder 200, to decode values for syntax elements in the bitstream. Syntax elements can define image-to-CTU segmentation information, and the segmentation of each CTU according to a corresponding segmentation structure such as a QTBT structure, to define the CU of the CTU. Syntax elements can further define prediction and residual information for blocks (e.g., CUs) of video data.
[0057] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inverse transform the quantized transform coefficients of the block to reconstruct the residual block for the block. The video decoder 300 uses the prediction mode (intra-frame or inter-frame prediction) notified by signaling and associated prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for the block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reconstruct the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the block boundaries.
[0058] During the decoding process, the video encoder 200 and video decoder 300 can store the decoded video data in decoded picture buffers (DPBs). The structure of these DPBs can vary, and the video encoder 200 can determine the structure of the DPBs and signaling notifications indicating one or more syntax elements of the determined structure. In VVC draft standards (e.g., VVC draft 8), the video encoder 200 can signal notifications indicating the syntax elements of the DPB parameter structure that can be signaled in the Video Parameter Set (VPS). The video encoder 200 can also signal notifications the number of DPB structures (vps_num_dpb_params), but this signaling notification can be conditional on whether all layers are decoded independently (vps_all_independent_layers_flag).
[0059] In VVC Draft 8, the following syntax table is presented:
[0060]
[0061]
[0062]
[0063] The following semantics describe the syntactic elements from the above syntax table:
[0064] A `vps_all_independent_layers_flag` value of 1 indicates that all layers in CVS are decoded independently without using inter-layer prediction. A `vps_all_independent_layers_flag` value of 0 indicates that one or more layers in CVS can use inter-layer prediction. When it does not exist, the value of `vps_all_independent_layers_flag` is inferred to be 1.
[0065] `each_layer_is_an_ols_flag` equal to 1 indicates that each OLS contains only one layer, and each layer in the reference VPS's CVS is itself the single included layer and the only output layer in the OLS. `each_layer_is_an_ols_flag` equal to 0 means the OLS can contain more than one layer. If `vps_max_layers_minus1` equals 0, then the value of `each_layer_is_an_ols_flag` is inferred to be equal to 1. Otherwise, when `vps_all_independent_layers_flag` equals 0, then the value of `each_layer_is_an_ols_flag` is inferred to be equal to 0.
[0066] `vps_num_dpb_params` specifies the number of `dpb_parameters()` syntax structures in the VPS. The value of `vps_num_dpb_params` will be in the range of 0 to 16 (inclusive). When it does not exist, the value of `vps_num_dpb_params` is inferred to be equal to 0.
[0067] `ols_dpb_params_idx[i]` specifies the index of the list of `dpb_parameters()` syntax structures in the VPS, applied to the `i`th OLS, when `NumLayersInOls[i]` is greater than 1. When `ols_dpb_params_idx[i]` exists, its value will be in the range of 0 to `vps_num_dpb_params-1` (inclusive). When `ols_dpb_params_idx[i]` does not exist, its value is inferred to be 0.
[0068] When NumLayersInOls[i] equals 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0069] `vps_sublayer_dpb_params_present_flag` is used to control the presence of the syntax elements `max_dec_pic_buffering_minus1[]`, `max_num_reorder_pics[]`, and `max_latency_increase_plus1[]` in the `dpb_parameters()` syntax structure of the VPS. When these elements are not present, `vps_sub_dpb_params_info_present_flag` is inferred to be equal to 0.
[0070] A value of 1 for `vps_general_hrd_params_present_flag` indicates that the syntax structure `general_hrd_parameters()` and other HRD parameters exist in the VPS RBSP syntax structure. A value of 0 for `vps_general_hrd_params_present_flag` indicates that the syntax structure `general_hrd_parameters()` and other HRD parameters do not exist in the VPS RBSP syntax structure. When they do not exist, the value of `vps_general_hrd_params_present_flag` is inferred to be 0.
[0071] When NumLayersInOls[i] equals 1, the general_hrd_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0072] In VVC draft 8, the video decoder 200 can notify one or more syntax elements of the DPB structure in the Indicator Sequence Parameter Set (SPS) according to the following conditional signaling:
[0073]
[0074] `sps_ptl_dpb_hrd_params_present_flag` equal to 1 indicates that the `profile_tier_level()` and `dpb_parameters()` syntax structures exist in SPS, and the `general_hrd_parameters()` and `ols_hrd_parameters()` syntax structures can also exist in SPS. `sps_ptl_dpb_hrd_params_present_flag` equal to 0 indicates that none of these four syntax structures exist in SPS. The value of `sps_ptl_dpb_hrd_params_present_flag` should be equal to `vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]`.
[0075] The technology in VVC Draft 8 may have one or more drawbacks. For example, there may be situations where all layers are decoded independently, and more than one layer is included in the output layer set (OLS). In this case, the video encoder 200 may signal zero DPB structures in the VPS, but when the DPB structure is inferred to be equal to 0, it is later referenced by ols_dpb_params_idx. Additionally, DPB parameters from the VPS DPB structures are used to define, for example, the maximum DPB size in level limits. The following is an excerpt from VVC Draft 8 entitled "A.4.1 General tier and level limits":
[0076] Otherwise (NumLayersInOls[TargetOlsIdx] is greater than 1), PicWidthMaxInSamplesY is set to equal ols_dpb_pic_width[TargetOlsIdx], PicHeightMaxInSamplesY is set to equal ols_dpb_pic_height[TargetOlsIdx], PicSizeMaxInSamplesY is set to equal PicWidthMaxInSamplesY*PicHeightMaxInSamplesY, and the applicable dpb_parameters() syntax structure is identified by ols_dpb_params_idx[TargetOlsIdx] found in the VPS.
[0077] In this scenario, dpb_parameters() is not signaled, but it is referenced (e.g., used) and utilized. In such cases, the video decoder 300 might attempt to reference syntax elements that were not signaled, potentially leading to unpredictable and unintended operation of the video decoder 300.
[0078] In SPS, when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 1, it means that the layer is decoded independently and may require video encoder 200 signaling notification to the DPB structure. However, independently decoded layers can be included in OLS and dpb_parameters() can be used instead.
[0079] This invention proposes several techniques that can solve the aforementioned problems. As an example, the video encoder 200 can signal the DPB parameters for an Output Layer Set (OLS) having more than one layer, regardless of whether the layers included in the OLS are independent or all layers included in the OLS are independent. A layer can be considered an independent layer in which it can be decoded without any other layer information. As another example, if the SPS is referenced by a layer that is the only layer in the OLS (i.e., the OLS has only one layer, or only one layer is encoded in the bitstream), the video encoder can always signal the DPB parameters in the SPS (e.g., signaling notification may be required).
[0080] In VVC Draft 8, `vps_num_dpb_params` can be signaled to be equal to 0 (i.e., indicating that no `dpb_parameters()` exists in the VPS). However, when it does not exist, `vps_num_dpb_params` is inferred to be 0. Since a signaling notification of a 0 value is not required, this disclosure proposes that the syntax element specifying the number of `dpb_parameters()` syntax structures in the VPS (i.e., `vps_num_dpb_params` in VVC Draft 8) can be modified to specify the number of `dpb_parameters()` syntax structures in the VPS minus one (e.g., `vps_num_dpb_params` can be replaced by `vps_num_dpb_params_minus1`). If at least one `dpb_parameters()` structure exists, this syntax element (e.g., `vps_num_dpb_params_minus1`) can be signaled. In one example, the semantics of this syntax element can be represented as follows:
[0081] `vps_num_dpb_params_minus1` specifies the number of `dpb_parameters()` syntax structures in the VPS minus 1. The value of `vps_num_dpb_params_minus1` will be in the range of 0 to 15 (inclusive). When it does not exist, the value of `vps_num_dpb_params_minus1` is inferred to be equal to 0.
[0082] Then, the signaling notification of the ols_dpb_params_idx syntax element can be conditional on the value of the vps_num_dpb_params_minus1 syntax element. As an example, this constraint can be implemented as follows:
[0083] if(vps_num_dpb_params_minus1>0) ols_dpb_params_idx[i] ue(v)
[0084] Signaling notifications for the `vps_sublayer_dpb_params_present_flag` syntax element can also be modified to be conditional on the value of the `vps_num_dpb_params_minus1` syntax element. As an example, this constraint can be implemented as follows:
[0085] if(vps_num_dpb_params_minus1>0&&vps_max_sublayers_minus1>0) vps_sublayer_dpb_params_present_flag u(1)
[0086] In another example, the signaling notification of vps_sublayer_dpb_params_present_flag can be combined with vps_num_dpb_params, then the condition vps_num_dpb_params_minus1>0 can be omitted, and only the condition with more than one temporal layer (vps_max_sublayers_minus1>0) can be retained.
[0087] As another example, when all independent layers can be included in the OLS, the video encoder 200 can also signal the vps_num_dpb_params syntax element for this situation. According to one or more techniques of this disclosure, the video encoder 200 can conditionally signal the vps_num_dpb_params syntax element such that the vps_num_dpb_params syntax element is signaled when more than one layer exists in the VPS and when more than one layer is included in any OLS. As an example, this constraint can be implemented as follows:
[0088] if(vps_max_layers_minus1>0&&!each_layer_is_an_ols_flag) vps_num_dpb_params ue(v)
[0089] Alternatively, the video encoder 200 may constrain signaling notifications to the vps_num_dpb_params syntax element based on one of the aforementioned conditions, rather than both. For example, the video encoder 200 may constrain signaling notifications to the vps_num_dpb_params syntax element based on either the condition vps_max_layers_minus1>0 (more than one layer) or ! each_layer_is_an_ols_flag (more than one layer is included in the OLS).
[0090] In another alternative, the video encoder 200 can unconditionally signal the `vps_num_dpb_params` syntax element. In some examples, the semantics of the `vps_num_dpb_params` syntax element can be constrained so that the value of `vps_num_dpb_params` is equal to zero when only one layer exists in the CVS, or when all OLS contain only one layer. In some examples, the semantics of the `vps_num_dpb_params` syntax element can be modified so that the value of `vps_num_dpb_params` is greater than 0 if more than one layer exists in the CVS or if more than one layer is included in any OLS.
[0091] The video encoder 200 can conditionally signal hypothetical reference decoder (HRD) parameters (vps_general_hrd_params_present_flag) under the same conditions as the vps_num_dpb_params signaling notification (e.g., when more than one layer is included in the OLS (!each_layer_is_an_ols_flag)). In this case, the signaling notifications for vps_num_dpb_params and vps_general_hrd_params_present_flag can be combined under one condition to avoid checking the condition twice. In one example, this can be implemented as follows:
[0092]
[0093] In SPS, when a standalone layer is included in OLS, dpb_parameters() can be signaled but not used because dpb_parameters() in the VPS is utilized. According to one or more techniques of this disclosure, the semantics of the sps_ptl_dpb_hrd_params_present_flag syntax element can be modified so that signaling notification of dpb_parameters() is only required when the SPS is referenced by a single layer. In one example, it is modified as follows:
[0094] `sps_ptl_dpb_hrd_params_present_flag` equal to 1 indicates that the `profile_tier_level()` and `dpb_parameters()` syntax structures exist in the SPS, and the `general_hrd_parameters()` and `ols_hrd_parameters()` syntax structures may also exist in the SPS. `sps_ptl_dpb_hrd_params_present_flag` equal to 0 indicates that none of these four syntax structures exist in the SPS. When the SPS is referenced by a single layer included in the OLS, the value of `sps_ptl_dpb_hrd_params_present_flag` should be equal to 1.
[0095] The following example syntax and semantics illustrate one or more implementations of the techniques described above. Modifications relative to VVC Draft 8 are presented in italics.
[0096]
[0097]
[0098]
[0099]
[0100] The following example syntax and semantics illustrate one or more implementations of the techniques described above. Modifications relative to VVC Draft 8 are presented in italics.
[0101] `vps_num_dpb_params_minus1` specifies the number of `dpb_parameters()` syntax structures in the VPS minus 1. The value of `vps_num_dpb_params_minus1` will be in the range of 0 to 15 (inclusive). When it does not exist, the value of `vps_num_dpb_params_minus1` is inferred to be equal to 0.
[0102] `ols_dpb_params_idx[i]` specifies the list of `dpb_parameters()` syntax structures in the VPS, when `NumLayersInOls[i]` is greater than 1, and is the index of the `dpb_parameters()` syntax structure applied to the `i`th OLS. When it exists, the value of `ols_dpb_params_idx[i]` will be in the range from 0 to `vps_num_dpb_params_minus1` (inclusive). When `ols_dpb_params_idx[i]` does not exist, its value is inferred to be 0.
[0103] When NumLayersInOls[i] equals 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0104] According to the technology disclosed herein, video encoder 200 and / or video decoder 300 can decode a syntax element in the video parameter set (VPS) of the current bitstream of specified video data minus one of the number of decoded picture buffer (DPB) parameter syntax structures; in response to determining that the syntax element does not exist in the bitstream, infer that the number of DPB syntax structures in the VPS is zero; and reconstruct the video data represented by the current bitstream.
[0105] Although typically described with reference to the Decoded Picture Buffer (DPB) structure, the techniques disclosed herein are equally applicable to other syntax structures. As an example, the techniques disclosed herein can be applied to profiletier level (PTL) syntax structures. As another example, the techniques disclosed herein can be applied to hypothetical reference decoder (HRD) syntax structures.
[0106] In some examples, video decoders may need to utilize a consistent design in the signaling notifications of the PTL, HRD, and DPB structures. For instance, a video decoder may use a common flag in the Sequence Parameter Set (SPS) to indicate the presence or absence of PTL, DBP, and HRD parameters (e.g., sps_ptl_dpb_hrd_params_present_flag).
[0107] According to one or more techniques disclosed herein, a video decoder may signal public flags in a Video Parameter Set (VPS) to indicate the presence or absence of DBP and HRD parameters. The video decoder may also separately signal flags used to indicate the presence or absence of PTL parameters (e.g., because PTL parameters can be used for session negotiation purposes).
[0108] The following example syntax and semantics illustrate one or more implementations of the techniques described above. Modifications relative to VVC Draft 8 are presented in italics.
[0109]
[0110]
[0111]
[0112]
[0113] vps_num_dpb_params_minus1 specifies the number of dpb_parameters() syntax structures in the VPS minus 1. The value of vps_num_dpb_params_minus1 will be in the range of 0 to 15 (inclusive).
[0114] `ols_dpb_params_idx[i]` specifies the list of `dpb_parameters()` syntax structures in the VPS, when `NumLayersInOls[i]` is greater than 1, and is the index of the `dpb_parameters()` syntax structure applied to the `i`th OLS. When it exists, the value of `ols_dpb_params_idx[i]` will be in the range from 0 to `vps_num_dpb_params_minus1` (inclusive). When `ols_dpb_params_idx[i]` does not exist, its value is inferred to be 0.
[0115] When NumLayersInOls[i] equals 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0116] A value of 1 for `vps_dpb_hrd_params_present_flag` indicates that the syntax structures `dpb_parameters()`, `general_hrd_parameters()`, and other HRD parameters exist in the VPS RBSP syntax structure. A value of 0 for `vps_dpb_hrd_params_present_flag` indicates that the syntax structures `dpb_parameters()`, `general_hrd_parameters()`, and other HRD parameters do not exist in the VPS RBSP syntax structure. When `vps_dpb_hrd_params_present_flag` does not exist, its value is inferred to be 0.
[0117] When NumLayersInOls[i] equals 1, the dpb_parameters() and general_hrd_parameters() syntax structures applied to the i-th OLS exist in the SPS referenced by the layer in the i-th OLS.
[0118] If more than one layer is included in any OLS in the VPS, then vps_dpb_hrd_params_present_flag should be equal to 1.
[0119] In SPS semantics, the inference of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is eliminated because it is insufficient, since in the case of more than one layer located in OLS, dpb_parameters() and ols_hrd_parameters() are derived from VPS and will be signaled there.
[0120] `sps_ptl_dpb_hrd_params_present_flag` equal to 1 indicates that the `profile_tier_level()` and `dpb_parameters()` syntax structures exist in SPS, and the `general_hrd_parameters()` and `ols_hrd_parameters()` syntax structures may also exist in SPS. `sps_ptl_dpb_hrd_params_present_flag` equal to 0 indicates that none of these four syntax structures exist in SPS. The value of `sps_ptl_dpb_hrd_params_present_flag` should be equal to 1 when `sps_video_parameter_set_id` equals 0 or when only one layer is included in any OLS of the referenced VPS.
[0121] The following is a clean version of the example syntax and semantics above.
[0122] <Clean Version>
[0123]
[0124]
[0125]
[0126]
[0127]
[0128] vps_num_dpb_params_minus1 specifies the number of dpb_parameters() syntax structures in the VPS minus 1. The value of vps_num_dpb_params_minus1 will be in the range of 0 to 15 (inclusive).
[0129] `ols_dpb_params_idx[i]` specifies the list of `dpb_parameters()` syntax structures in the VPS, when `NumLayersInOls[i]` is greater than 1, and is the index of the `dpb_parameters()` syntax structure applied to the `i`th OLS. When it exists, the value of `ols_dpb_params_idx[i]` will be in the range from 0 to `vps_num_dpb_params_minus1` (inclusive). When `ols_dpb_params_idx[i]` does not exist, its value is inferred to be 0.
[0130] When NumLayersInOls[i] equals 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0131] A value of 1 for `vps_dpb_hrd_params_present_flag` indicates that the syntax structures `dpb_parameters()`, `general_hrd_parameters()`, and other HRD parameters exist in the VPS RBSP syntax structure. A value of 0 for `vps_dpb_hrd_params_present_flag` indicates that the syntax structures `dpb_parameters()`, `general_hrd_parameters()`, and other HRD parameters do not exist in the VPS RBSP syntax structure. When they do not exist, the value of `vps_dpb_hrd_params_present_flag` is inferred to be 0.
[0132] When NumLayersInOls[i] equals 1, the dpb_parameters() and general_hrd_parameters() syntax structures applied to the i-th OLS exist in the SPS referenced by the layer in the i-th OLS.
[0133] If more than one layer is included in any OLS in the VPS, then vps_dpb_hrd_params_present_flag should be equal to 1.
[0134] In the SPS semantics, the inference of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is eliminated because it is not sufficient since in the case where more than one layer is in the OLS, dpb_parameters() and ols_hrd_parameters() are derived from the VPS and will be signaled there.
[0135] A sps_ptl_dpb_hrd_params_present_flag equal to 1 specifies that the profile_tier_level() syntax structure and the dpb_parameters() syntax structure are present in the SPS, and also that the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure may be present in the SPS. A sps_ptl_dpb_hrd_params_present_flag equal to 0 specifies that none of these four syntax structures are present in the SPS. When sps_video_parameter_set_id is equal to 0 or only one layer is included in any OLS of the referenced VPS, the value of sps_ptl_dpb_hrd_params_present_flag shall be equal to 1.
[0136] < / Pure version>
[0137] In some examples, the video decoder may signal the presence or absence of the PTL, DPB, and HRD structures under a common gating flag in the SPS, while the PTL can always be signaled in the VPS (vps_num_ptls_minus1). The PTL signaled in the VPS can be used for session negotiation, but there may currently not be a mechanism to disable PTL signaling in the VPS, even in the case of having only a single layer in the OLS.
[0138] According to one or more techniques of the present disclosure, the video decoder may signal a common gating flag in the VPS to indicate the presence or absence of the PTL, DPB, and HRD syntax structures. For example, the video decoder may include the PTL together with the DPB and HRD under the common gating flag in the VPS.
[0139] The following example syntax and semantics may illustrate one or more implementations of the above techniques. Modifications relative to VVC draft 8 are presented in italics.
[0140]
[0141]
[0142]
[0143]
[0144]
[0145] `vps_num_dpb_params_minus1` specifies the number of `dpb_parameters()` syntax structures in the VPS minus 1. The value of `vps_num_dpb_params` will be in the range of 0 to 15 (inclusive).
[0146] `ols_dpb_params_idx[i]` specifies the list of `dpb_parameters()` syntax structures in the VPS, when `NumLayersInOls[i]` is greater than 1, and is the index of the `dpb_parameters()` syntax structure applied to the `i`th OLS. When it exists, the value of `ols_dpb_params_idx[i]` will be in the range from 0 to `vps_num_dpb_params_minus1` (inclusive). When `ols_dpb_params_idx[i]` does not exist, its value is inferred to be 0.
[0147] When NumLayersInOls[i] equals 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0148] A value of 1 for `vps_ptl_dpb_hrd_params_present_flag` indicates that the syntax structures `profile_tier_level()`, `dpb_parameters()`, `general_hrd_parameters()`, and other HRD parameters exist in the VPSRBSP syntax structure. A value of 0 for `vps_ptl_dpb_hrd_params_present_flag` indicates that these syntax structures do not exist in the VPS RBSP syntax structure. When they do not exist, the value of `vps_ptl_dpb_hrd_params_present_flag` is inferred to be 0.
[0149] When NumLayersInOls[i] equals 1, the general_hrd_parameters() and dpb_parameters() syntax structures applied to the i-th OLS exist in the SPS referenced by the layer in the i-th OLS.
[0150] If more than one layer is included in any OLS in the VPS, then vps_ptl_dpb_hrd_params_present_flag should be equal to 1.
[0151] In SPS semantics, the inference of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is eliminated because it is insufficient, since in the case of more than one layer located in OLS, dpb_parameters() and ols_hrd_parameters() are derived from VPS and will be signaled there.
[0152] `sps_ptl_dpb_hrd_params_present_flag` equal to 1 indicates that the `profile_tier_level()` and `dpb_parameters()` syntax structures exist in SPS, and the `general_hrd_parameters()` and `ols_hrd_parameters()` syntax structures may also exist in SPS. `sps_ptl_dpb_hrd_params_present_flag` equal to 0 indicates that none of these four syntax structures exist in SPS. The value of `sps_ptl_dpb_hrd_params_present_flag` should be equal to 1 when `sps_video_parameter_set_id` equals 0 or when only one layer is included in any OLS of the referenced VPS.
[0153] The following is a clean version of the example syntax and semantics above.
[0154] <Clean Version>
[0155]
[0156]
[0157]
[0158]
[0159] `vps_num_dpb_params_minus1` specifies the number of `dpb_parameters()` syntax structures in the VPS minus 1. The value of `vps_num_dpb_params` will be in the range of 0 to 15 (inclusive).
[0160] `ols_dpb_params_idx[i]` specifies the list of `dpb_parameters()` syntax structures in the VPS, when `NumLayersInOls[i]` is greater than 1, and is the index of the `dpb_parameters()` syntax structure applied to the `i`th OLS. When it exists, the value of `ols_dpb_params_idx[i]` will be in the range from 0 to `vps_num_dpb_params_minus1` (inclusive). When `ols_dpb_params_idx[i]` does not exist, its value is inferred to be 0.
[0161] When NumLayersInOls[i] equals 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0162] A value of 1 for `vps_ptl_dpb_hrd_params_present_flag` indicates that the syntax structures `profile_tier_level()`, `dpb_parameters()`, `general_hrd_parameters()`, and other HRD parameters exist in the VPSRBSP syntax structure. A value of 0 for `vps_ptl_dpb_hrd_params_present_flag` indicates that these syntax structures do not exist in the VPS RBSP syntax structure. When they do not exist, the value of `vps_ptl_dpb_hrd_params_present_flag` is inferred to be 0.
[0163] When NumLayersInOls[i] equals 1, the general_hrd_parameters() and dpb_parameters() syntax structures applied to the i-th OLS exist in the SPS referenced by the layer in the i-th OLS.
[0164] If more than one layer is included in any OLS in the VPS, then vps_ptl_dpb_hrd_params_present_flag should be equal to 1.
[0165] In the SPS semantics, the inference of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is eliminated because it is not sufficient, as in the case where more than one layer is in the OLS, dpb_parameters() and ols_hrd_parameters() are derived from the VPS and will be signaled there.
[0166] A sps_ptl_dpb_hrd_params_present_flag equal to 1 specifies that the profile_tier_level() syntax structure and the dpb_parameters() syntax structure are present in the SPS, and that the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure may also be present in the SPS. A sps_ptl_dpb_hrd_params_present_flag equal to 0 specifies that none of these four syntax structures are present in the SPS. When sps_video_parameter_set_id is equal to 0 or only one layer is included in any OLS of the referenced VPS, the value of sps_ptl_dpb_hrd_params_present_flag shall be equal to 1.
[0167] < / Pure version>
[0168] In the above example, for all cases, including the case where all layers are independent, it may be necessary for the video decoder to signal the presence or absence of the DPB, HRD structures, and optionally the PTL syntax structure in the VPS. However, there is a possibility that an independent layer is extracted into a single bitstream where the presence of the VPS may be optional.
[0169] According to one or more techniques of the present disclosure, the video decoder may signal the DPB, HRD, or PTL in the SPS or in any other parameter set for an independent layer, even in the case where more than one layer is included in the OLS. When the DPB, HRD, or PTL parameters are accessed. In the case where more than one layer is included in the OLS, the video decoder may determine whether the layers in the OLS are independent, and if the layers are independent, then the video decoder may access the DPB, HRD, or PTL from the SPS or any other parameter set referenced by the layer, rather than from the VPS. In some examples, the video decoder may access those parameters (e.g., DPB, HRD, or PTL) from the VPS for dependent layers. Similar methods may be used for any other parameters. In this way, the video decoder may enable VPS optional functionality.
[0170] This disclosure may generally refer to certain information such as “signaling”, including syntax elements. The term “signaling” generally refers to communication regarding the values of syntax elements and / or other data used for decoding encoded video data. That is, the video encoder 200 may signal values for syntax elements in the bitstream. Typically, signaling involves generating values in the bitstream. As described above, the source device 102 may transmit the bitstream to the target device 116 substantially in real time or non-real time, with non-real-time transmission occurring, such as when syntax elements are stored in storage device 112 for later retrieval by the target device 116.
[0171] Figure 2A and Figure 2B This is a conceptual diagram illustrating an example Quadtree Binary Tree (QTBT) structure 130 and its corresponding Decoding Tree Unit (CTU) 132. Solid lines represent quadtree partitions, and dashed lines indicate binary tree partitions. In each partition (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which partition type (i.e., horizontal or vertical) is used, where in this example, 0 indicates a horizontal partition and 1 indicates a vertical partition. For quadtree partitions, it is not necessary to indicate the partition type because the quadtree node divides the block horizontally and vertically into four equal-sized sub-blocks. Accordingly, the video encoder 200 can encode and the video decoder 300 can decode syntax elements (such as partition information) at the region tree level (i.e., solid lines) and the prediction tree level (i.e., dashed lines) of the QTBT structure 130. The video encoder 200 can encode and the video decoder 300 can decode video data such as prediction and transform data for the CU represented by the terminal leaf nodes of the QTBT structure 130.
[0172] Generally speaking, Figure 2B The CTU 132 can be associated with parameters that define the size of the block corresponding to the nodes of the first and second levels in the QTBT structure 130. These parameters may include the CTU size (the size of the CTU 132 in sample points), the minimum quadtree size (MinQTSize, representing the minimum allowed size of the leaf nodes of the quadtree), the maximum binary tree size (MaxBTSize, representing the maximum allowed size of the root node of the binary tree), the maximum binary tree depth (MaxBTDepth, representing the maximum allowed depth of the binary tree), and the minimum binary tree size (MinBTSize, representing the minimum allowed size of the leaf nodes of the binary tree).
[0173] The root node of the QTBT structure corresponding to CTU can have four child nodes at the first level of the QTBT structure, each of which can be partitioned according to a quadtree. That is, the nodes at the first level are either leaf nodes (with no child nodes) or have four child nodes. An example of QTBT structure 130 represents such nodes as including a parent node and child nodes with solid lines for branching. If the nodes at the first level are not larger than the maximum allowed binary tree root node size (MaxBTSize), these nodes can be further partitioned by the corresponding binary tree. The binary tree partitioning of a node can be iterated until the resulting nodes reach the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). An example of QTBT structure 130 represents such nodes as having dashed lines for branching. The binary tree leaf nodes are called decoding units (CUs), which are used for prediction (e.g., intra-image or inter-image prediction) and transformation without any further partitioning. As mentioned above, CUs can also be referred to as “video blocks” or “blocks”.
[0174] In one example of a QTBT partitioning structure, the CTU size is set to 128×128 (luminance samples and two corresponding 64×64 chrominance samples), MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. Quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If a leaf quadtree node is 128×128, it will not be further partitioned into binary trees because its size exceeds MaxBTSize (i.e., 64×64 in this example). Otherwise, the leaf quadtree node will be further partitioned into binary trees. Therefore, the quadtree leaf node is also the root node of the binary tree and has a binary tree depth of 0. When the depth of the binary tree reaches MaxBTDepth (4 in this example), further partitioning is not allowed. Similarly, when a binary tree node has a width equal to MinBTSize (4 in this example), further horizontal partitioning is not allowed. Likewise, a binary tree node with a height equal to MinBTSize means that further vertical partitioning is not allowed for that node. As mentioned above, the leaf nodes of the binary tree are called CUs and are further processed according to the prediction and transformation without further partitioning.
[0175] Figure 3 This is a block diagram illustrating an example video encoder 200 that can perform the techniques of this disclosure. Figure 3This disclosure is provided for illustrative purposes and should not be construed as limiting the techniques broadly illustrated and described herein. For illustrative purposes, this disclosure describes a video encoder 200 based on the following technologies: JEM, VVC (ITU-T H.266, under development), and HEVC (ITU-T H.265). However, the techniques of this disclosure can be implemented by video encoding devices configured to other video decoding standards.
[0176] exist Figure 3 In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy coding unit 220 can be implemented in one or more processors or processing circuits. For example, the units of the video encoder 200 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuits to perform these and other functions.
[0177] The video data storage device 230 can store video data encoded by components of the video encoder 200. The video encoder 200 can retrieve data from, for example, a video source 104 (…). Figure 1 The video data memory 230 receives video data stored in the video data memory 230. The DPB 218 can act as a reference picture memory storing reference video data for prediction of subsequent video data by the video encoder 200. The video data memory 230 and DPB 218 can be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 can be located on-chip with other components of the video encoder 200, as illustrated, or off-chip relative to those components.
[0178] In this disclosure, references to video data memory 230 should not be construed as being limited to memory internal to video encoder 200 unless so specifically described, or memory external to video encoder 200 unless so specifically described. Rather, references to video data memory 230 should be understood as reference memory storing video data received by video encoder 200 for encoding (e.g., video data for the current block to be encoded). Figure 1 The memory 106 can also provide temporary storage for the outputs from various units of the video encoder 200.
[0179] Figure 3 The various units are illustrated to aid in understanding the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides specific functionality and is pre-programmed to perform certain operations. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations it can perform. For example, a programmable circuit can execute software or firmware that causes it to operate in a manner defined by instructions from software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), while in other examples, one or more of the units may be integrated circuits.
[0180] The video encoder 200 may include an arithmetic logic unit (ALU), an elementary function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video encoder 200 is performed using software executed by programmable circuitry, memory 106 ( Figure 1 The video encoder 200 may store instructions (e.g., object code) of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.
[0181] The video data storage unit 230 is configured to store the received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be the raw video data to be encoded.
[0182] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units for performing video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-frame block copying unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, and so on.
[0183] The mode selection unit 202 typically coordinates multiple coding passes to test combinations of coding parameters and the resulting rate-distortion values for such combinations. Coding parameters may include the CTU-CU split, the prediction mode for the CU, the transformation type of the residual data for the CU, the quantization parameters of the residual data for the CU, etc. The mode selection unit 202 can ultimately select a combination of coding parameters that has a better rate-distortion value than other tested combinations.
[0184] The video encoder 200 can segment images retrieved from the video data storage 230 into a series of CTUs, and encapsulate one or more CTUs within a slice. The mode selection unit 202 can segment the CTUs of the image according to a tree structure such as the QTBT structure or quadtree structure of HEVC described above. As mentioned above, the video encoder 200 can form one or more CUs by segmenting CTUs according to a tree structure. Such CUs can also generally be referred to as "video blocks" or "blocks".
[0185] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate prediction blocks for the current block (e.g., the current CU, or the overlapping portion of PU and TU in HEVC). For inter-frame prediction of the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference images (e.g., one or more previously decoded images stored in DPB 218). Specifically, motion estimation unit 222 may calculate values indicating how similar a potential reference block is to the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. Motion estimation unit 222 may typically perform these calculations using the sample-by-sample difference between the current block and the considered reference blocks. Motion estimation unit 222 may identify reference blocks with the lowest values obtained from these calculations, which indicate the reference block that most closely matches the current block.
[0186] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the position of a reference block in a reference image relative to the position of a current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data for reference blocks. As another example, if the motion vectors have fractional sample precision, motion compensation unit 224 can interpolate the values for the prediction blocks according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by corresponding motion vectors and combine the retrieved data, for example, by sample-by-sample averaging or weighted averaging.
[0187] As another example, for intra-prediction or intra-prediction decoding, intra-prediction unit 226 can generate a prediction block from samples adjacent to the current block. For example, in directional mode, intra-prediction unit 226 can typically mathematically combine the values of adjacent samples and fill these calculated values in a defined direction across the current block to produce a prediction block. As another example, in DC mode, intra-prediction unit 226 can calculate the average of adjacent samples for the current block and generate a prediction block such that this resulting average is included for each sample in the prediction block.
[0188] Mode selection unit 202 provides the prediction block to residual generation unit 204. Residual generation unit 204 receives the original uncoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, residual generation unit 204 may also determine the difference between sample values in the residual block to generate the residual block using residual differential pulse decode-modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form residual generation unit 204.
[0189] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As mentioned above, the size of a CU can refer to the size of its luma decoding block, while the size of a PU can refer to the size of its luma prediction unit. Assuming a specific CU has a size of 2Nx2N, video encoder 200 can support PUs of 2Nx2N or NxN size for intra-frame prediction, and symmetric PUs of 2Nx2N, 2NxN, Nx2N, NxN, or similar sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric segmentation of PUs of 2NxnU, 2NxnD, nLx2N, and nRx2N size for inter-frame prediction.
[0190] In the example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luminance decoding block and a corresponding chrominance decoding block. As mentioned above, the size of the CU can refer to the size of the luminance decoding block of the CU. The video encoder 200 and the video decoder 300 can support CUs with sizes of 2Nx2N, 2NxN, or Nx2N.
[0191] For other video decoding techniques such as intra-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding, as examples, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples, such as palette mode decoding, mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block is reconstructed based on the selected palette. In such modes, mode selection unit 202 may provide these syntax elements to entropy coding unit 220 for encoding.
[0192] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.
[0193] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a discrete cosine transform (DCT), direction transform, Karl Henroff transform (KLT), or a conceptually similar transform to the residual block. In some examples, transform processing unit 206 may perform multiple transformations on the residual block, such as primary and secondary transformations, such as rotation transformations. In some examples, transform processing unit 206 does not apply any transformations to the residual block.
[0194] Quantization unit 208 can quantize the transform coefficients in a transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 can quantize the transform coefficients of the transform coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and therefore, the quantized transform coefficients may have lower precision than the original transform coefficients generated by transform processing unit 206.
[0195] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform to the quantized transform coefficient block, respectively, to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although it may have some degree of distortion). For example, the reconstruction unit 214 can add the samples of the reconstructed residual block to the corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.
[0196] Filter unit 216 can perform one or more filter operations on the reconstructed block. For example, filter unit 216 can perform deblocking to reduce blockiness artifacts along the edges of the CU. In some examples, the operations of filter unit 216 can be skipped.
[0197] The video encoder 200 stores reconstructed blocks in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 can store the reconstructed blocks in the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 can store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve a reference picture formed by the reconstructed (and possibly filtered) blocks from the DPB 218 to perform inter-frame prediction of blocks in subsequent encoded pictures. Furthermore, the intra-frame prediction unit 226 can use the reconstructed blocks of the current picture in the DPB 218 to perform intra-frame prediction of other blocks in the current picture.
[0198] Generally, entropy coding unit 220 can entropy-encode syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy-encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy-encode predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, another example of video data, to generate entropy-encoded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, exponential-Golomb coding, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in a bypass mode in which syntax elements are not entropy-encoded.
[0199] The video encoder 200 can output a bitstream containing entropy-coded syntax elements of the blocks needed to reconstruct slices or images. Specifically, the entropy coding unit 220 can output a bitstream.
[0200] The operations described above are relative to blocks. Such descriptions should be understood as operations applied to luma decoding blocks and / or chroma decoding blocks. As mentioned above, in some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the CU. In some instances, the luma decoding block and chroma decoding block are the luma and chroma components of the PU.
[0201] In some examples, the operations performed relative to the luma decoding block do not need to be repeated for the chroma decoding block. As an example, the operations used to identify the motion vector (MV) and reference image for the luma decoding block do not need to be repeated for identifying the MV and reference image for the chroma decoding block. Specifically, the MV for the luma decoding block can be scaled to determine the MV for the chroma block, while the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma decoding blocks.
[0202] Video encoder 200 represents an example of a device configured to encode video data, including a memory configured to store the video data, and one or more processing units implemented in circuitry and configured to: decode a syntax element in the Video Parameter Set (VPS) of the current bitstream of the specified video data by subtracting one from the number of Decoded Picture Buffer (DPB) parameter syntax structures; infer that the number of DPB syntax structures in the VPS is zero in response to determining that the syntax element does not exist in the bitstream; and reconstruct the video data represented by the current bitstream. Video encoder 200 may configure the structure of DPB 218 based on one or more structures described by the DPB syntax structures in the VPS.
[0203] Figure 4 This is a block diagram illustrating an example video decoder 300 that can perform the techniques of this disclosure. Figure 4 This disclosure is provided for illustrative purposes and is not intended to limit the techniques broadly exemplified and described herein. For illustrative purposes, this disclosure describes a video decoder 300 based on the following technologies: JEM, VVC (ITU-T H.266, under development), and HEVC (ITU-T H.265). However, the techniques of this disclosure can be implemented by video decoding devices configured to other video decoding standards.
[0204] exist Figure 4In the example, the video decoder 300 includes a decoded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or processing circuits. For example, units of the video decoder 300 can be implemented as one or more circuit or logic elements, as part of hardware circuitry or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuits to perform these and other functions.
[0205] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-prediction unit 318. The prediction processing unit 304 may include additional units for performing predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copying unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, and so on. In other examples, the video decoder 300 may include more, fewer, or different functional components.
[0206] CPB memory 320 can store video data, such as encoded video bitstreams, that will be decoded by components of video decoder 300. The video data stored in CPB memory 320 can be, for example, from computer-readable medium 110 (…). Figure 1 The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Similarly, the CPB memory 320 may store video data other than the syntax elements of the decoded picture, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores decoded pictures that the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and DPB 314 may be formed of any of a variety of memory devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of memory devices. The CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be located on-chip with other components of the video decoder 300, or off-chip relative to those components.
[0207] Additionally or alternatively, in some examples, the video decoder 300 can be drawn from the memory 120 ( Figure 1 Retrieving decoded video data. That is, memory 120 can store data as discussed above with CPB memory 320. Similarly, when some or all of the functionality of video decoder 300 is implemented in software for execution by the processing circuitry of video decoder 300, memory 120 can store instructions for execution by video decoder 300.
[0208] Figure 4 The various units shown are illustrated to aid in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Figure 3 Similarly, a fixed-function circuit refers to a circuit that provides a specific function and is pre-programmed to perform certain operations. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations it can perform. For example, a programmable circuit can execute software or firmware that causes it to operate in a manner defined by instructions from the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more of the units may be different circuit blocks (fixed-function or programmable), while in other examples, one or more of the units may be integrated circuits.
[0209] The video decoder 300 may include an ALU, an EFU, digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executed on the programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.
[0210] Entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to reproduce the syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.
[0211] Typically, the video decoder 300 reconstructs the image on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed, i.e., the block being decoded, can be called the "current block").
[0212] Entropy decoding unit 302 can entropy decode the syntax elements defining the quantized transform coefficients of the quantized transform coefficient block, as well as transform information such as quantization parameters (QP) and / or (one or more) transform mode indications. Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the quantization level, and similarly, determine the inverse quantization level to be applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bit-by-bit left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 can thus form a transform coefficient block including the transform coefficients.
[0213] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply the inverse DCT, inverse integer transform, inverse Karl Henle-Loughlin transform (KLT), inverse rotation transform, inverse direction transform, or another inverse transform to the transform coefficient block.
[0214] Furthermore, prediction processing unit 304 generates prediction blocks based on prediction information syntax elements entropy decoded by entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is predicted inter-frame, motion compensation unit 316 can generate prediction blocks. In this case, the prediction information syntax elements may indicate a reference picture from which the reference block is retrieved in DPB 314, and a motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. Motion compensation unit 316 can typically be used in conjunction with motion compensation unit 224 ( Figure 3 The method described is basically similar to the method used to perform the inter-frame prediction process.
[0215] As another example, if the prediction information syntax element indicates that the current block is intra-predictable, then intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, intra-prediction unit 318 can typically be used in conjunction with intra-prediction unit 226 ( Figure 3 The intra-prediction process is performed in a similar manner. The intra-prediction unit 318 can retrieve data from the neighboring samples of the current block from the DPB 314.
[0216] Reconstruction unit 310 can use prediction blocks and residual blocks to reconstruct the current block. For example, reconstruction unit 310 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block.
[0217] Filter unit 312 can perform one or more filter operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce block artifacts along the edges of the reconstructed block. The operation of filter unit 312 does not need to be performed in all examples.
[0218] The video decoder 300 can store reconstructed blocks in the DPB 314. For example, in an example where the filter unit 312 is not operated, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the filter unit 312 is operated, the filter unit 312 can store the filtered reconstructed blocks in the DPB 314. As described above, the DPB 314 can provide reference information, such as samples of the current image for intra-frame prediction and samples of the previously decoded image for subsequent motion compensation, to the prediction processing unit 304. Moreover, the video decoder 300 can output decoded images (e.g., decoded video) from the DPB 314 for use in, for example... Figure 1 The subsequent presentation on display devices such as display device 118.
[0219] In this manner, video decoder 300 represents an example of a video decoding device, including a memory configured to store video data, and one or more processing units implemented in circuitry and configured to: decode a syntax element in the Video Parameter Set (VPS) of the current bitstream of specified video data minus one of the number of Decoded Picture Buffer (DPB) parameter syntax structures; infer that the number of DPB syntax structures in the VPS is zero in response to determining that the syntax element does not exist in the bitstream; and reconstruct the video data represented by the current bitstream. Video decoder 300 can configure the structure of DPB 314 based on one or more structures described by the DPB syntax structures in the VPS.
[0220] Figure 5 This is a flowchart illustrating an example method for encoding a current block according to one or more techniques of this disclosure. The current block may include the current CU. Although relative to video encoder 200 ( Figure 1 and Figure 3 This is described in detail, but it should be understood that other devices can be configured to perform similar actions. Figure 5 Similar to the method.
[0221] In this example, the video encoder 200 initially predicts the current block (350). For example, the video encoder 200 may form a predicted block for the current block. The video encoder 200 may then compute a residual block for the current block (352). To compute the residual block, the video encoder 200 may compute the difference between the original uncoded block for the current block and the predicted block. The video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (354). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (358). For example, the video encoder 200 may encode the transform coefficients using CAVLC or CABAC. The video encoder 200 may then output the entropy-encoded data of the block (360). The video encoder 200 may further selectively encode syntax elements in the Video Parameter Set (VPS) of the current bitstream of the specified video data, minus one number of Decoded Picture Buffer (DPB) parameter syntax structures.
[0222] Figure 6 This is a flowchart illustrating an example method for decoding the current block of video data. The current block may include the current CU. Although relative to video decoder 300 ( Figure 1 and Figure 4 This is described in detail, but it should be understood that other devices can be configured to perform similar actions. Figure 6 Similar to the method.
[0223] The video decoder 300 can receive entropy-coded data for the current block, such as entropy-coded prediction information and entropy-coded data for the coefficients of the residual block corresponding to the current block (370). The video decoder 300 can entropy decode the entropy-coded data to determine the prediction information for the current block and reproduce the coefficients of the residual block (372). The video decoder 300 can predict the current block, for example, using an intra-frame prediction or inter-frame prediction mode indicated by the prediction information for the current block (374), to compute a prediction block for the current block. The video decoder 300 can then inverse scan the reproduced coefficients (376) to create a block of quantized transform coefficients. The video decoder 300 can then inverse quantize and inverse transform the transform coefficients to produce a residual block (378). The video decoder 300 can finally decode the current block by combining the prediction block and the residual block (380). The video decoder 300 can further decode a syntax element in the video parameter set (VPS) of the current bitstream of the specified video data, minus one number of decoded picture buffer (DPB) parameter syntax structures; in response to determining that the syntax element does not exist in the bitstream, infer that the number of DPB syntax structures in the VPS is zero; and reconstruct the video data represented by the current bitstream.
[0224] Figure 7 This is a flowchart illustrating an example method for signaling notification of DPB parameters according to one or more aspects of this disclosure. Although relative to the video decoder 300 ( Figure 1 and Figure 4 This is described in detail, but it should be understood that other devices can be configured to perform similar actions. Figure 7 (For example, the method of video encoder 200) is similar to that of video encoder 200.
[0225] like Figure 7 As shown, the video decoder 300 can determine whether the decoded picture buffer (DPB) parameter syntax structure exists in the sequence parameter set (SPS) of the current bitstream (702). For example, the entropy decoding unit 302 of the video decoder 300 can decode the syntax element indicating whether the specified DPB parameter syntax structure exists in the SPS. As an example, the entropy decoding unit 302 can decode the sps_ptl_dpb_hrd_params_present_flag syntax element. When the sps_ptl_dpb_hrd_params_present_flag syntax element is equal to 1, the DPB parameter syntax structure may exist in the SPS. Alternatively, when the sps_ptl_dpb_hrd_params_present_flag syntax element is equal to 0, the SPS may not include the DPB parameter syntax structure. According to one or more aspects of this disclosure, when the SPS is referenced by a layer that is the only layer in the Output Layer Set (OLS) (i.e., where sps_video_parameter_set_id equals 0), the sps_ptl_dpb_hrd_params_present_flag syntax element can always be 1 (i.e., the SPS contains a DPB parameter syntax structure). Additionally or alternatively, when sps_video_parameter_set_id equals 0 or only one layer is included in any OLS of the referenced VPS, the value of sps_ptl_dpb_hrd_params_present_flag should be equal to 1.
[0226] If the DPB parameter syntax structure exists in the SPS (the "yes" branch of 702), the video decoder 300 can decode the DPB parameter syntax structure from the SPS (704). For example, the entropy decoding unit 302 can decode the DPB parameter syntax structure by decoding at least the syntax structure that includes syntax elements that provide information about the DPB size, the maximum number of picture reorderings, and the maximum latency for one or more OLS.
[0227] If the DPB parameter syntax structure is not present in the SPS (the "No" branch of 702), the video decoder 300 can decode from the Video Parameter Set (VPS) a syntax element specifying whether each OLS contains only one layer or is allowed to contain multiple layers (706), and determine whether each OLS is allowed to contain multiple layers based on that syntax element (708). For example, the entropy decoding unit 302 can decode the each_layer_is_an_ols_flag syntax element from the VPS. If the each_layer_is_an_ols_flag syntax element is equal to 1, the video decoder 300 can determine that each OLS contains only one layer. If the each_layer_is_an_ols_flag syntax element is equal to 0, the video decoder 300 can determine that each OLS is allowed to contain multiple layers.
[0228] If each OLS is allowed to contain multiple layers (the "yes" branch of 708), the video decoder 300 can decode syntax elements (710) from the VPS that are one less than the number of DPB parameter syntax structures in the specified VPS. For example, the entropy decoding unit 302 can decode the vps_num_dpb_params_minus1 syntax element.
[0229] The video decoder 300 can reconstruct the video data represented by the current bitstream based on one or more DPB parameter syntax structures (712). For example, the video decoder 300 can configure one or more aspects of the DPB 314 (e.g., size, maximum number of image reorderings, maximum latency). As described above, the DPB 314 can store images, which the video decoder 300 can output and / or use as reference video data when decoding subsequent data or images of the encoded video bitstream.
[0230] In cases where each OLS is not allowed to contain multiple layers (i.e., each OLS contains only one layer) (the "No" branch of 708), the video decoder 300 can infer that the VPS contains zero DPB syntax structures (714). For example, when the each_layer_is_an_ols_flag syntax element is equal to 1, the decoded bitstream may not include the vps_num_dpb_params_minus1 syntax element. By signaling the number of DPB parameter syntax structures in the VPS to a number minus one, the technique of this disclosure enables the video decoder to avoid having to explicitly signal that the VPS includes zero DPB parameter syntax structures. Thus, the technique of this disclosure reduces the number of bits used for signaling the number of DPB parameter syntax structures included in the VPS, which improves decoding efficiency.
[0231] In some examples, video decoder 300 can constrain the decoding of one or more other syntax elements based on a value specifying whether each OLS contains only one layer or is allowed to contain multiple layers. For example, in response to a syntax element indicating that each OLS is allowed to contain multiple layers, video decoder 300 can decode from a VPS a syntax element specifying whether the VPS includes a hypothetical reference decoder (HRD) parameter syntax structure. As an example, entropy decoding unit 302 can decode hrd_params_present_flag from a VPS.
[0232] The following numbered clauses may illustrate one or more aspects of this disclosure:
[0233] Clause 1. A method for decoding video data, the method comprising: decoding a syntax element in a video parameter set (VPS) of a current bitstream of specified video data, the number of decoded picture buffer (DPB) parameter syntax structures minus 1; in response to determining that the syntax element does not exist in the bitstream, inferring that the number of DPB syntax structures in the VPS is zero; and reconstructing the video data represented by the current bitstream.
[0234] Clause 2. The method as described in Clause 1, wherein decoding a syntax element comprises selectively decoding the syntax element based on the value of the syntax element which specifies the number of layers contained in each Output Layer Set (OLS).
[0235] Clause 3. As described in Clause 2, wherein the syntax element specifying the number of DPB parameter syntax structures in the VPS minus 1 includes the vps_num_dpb_params_minus1 syntax element.
[0236] Clause 4. The method as described in Clause 3, wherein the syntax element specifying the number of layers contained in each OLS includes the each_layer_is_an_ols_flag syntax element.
[0237] Clause 5. A method for decoding video data, the method comprising: decoding in a video parameter set (VPS) of a current bitstream of the video data a syntax element that commonly indicates the presence of a decoded picture buffer (DPB) parameter syntax structure and a hypothetical reference decoder (HRD) parameter syntax structure in the VPS; in response to the syntax element indicating the presence of the DPB parameter syntax structure and the HRD parameter syntax structure in the VPS, decoding the DPB parameter syntax structure and the HRD parameter syntax structure from the VPS; and reconstructing the video data represented by the current bitstream based on the DPB parameter syntax structure and the HRD parameter syntax structure.
[0238] Clause 6. The method as described in Clause 5, wherein decoding a syntax element comprises selectively decoding the syntax element based on the value of the syntax element which specifies the number of layers contained in each Output Layer Set (OLS).
[0239] Clause 7. As described in Clause 6, wherein the syntax elements that commonly indicate the existence of the DPB parameter syntax structure and the HRD parameter syntax structure in the VPS include the vps_dpb_hrd_params_present_flag syntax element.
[0240] Clause 8. The method as described in Clause 7, wherein the syntax element specifying the number of layers contained in each OLS includes the each_layer_is_an_ols_flag syntax element.
[0241] Clause 9. A method for decoding video data, the method comprising: decoding in a Video Parameter Set (VPS) of a current bitstream of the video data a syntax element that commonly indicates the presence of a Decoded Picture Buffer (DPB) parameter syntax structure, a Hypothetical Reference Decoder (HRD) parameter syntax structure, and a Profile Hierarchy Level (PTL) parameter syntax structure in the VPS; decoding the DPB, HRD, and PTL parameter syntax structures from the VPS in response to the syntax elements indicating the presence of the DPB, HRD, and PTL parameter syntax structures in the VPS; and reconstructing the video data represented by the current bitstream based on the DPB, HRD, and PTL parameter syntax structures.
[0242] Clause 10. The method as described in Clause 9, wherein decoding a syntax element comprises selectively decoding the syntax element based on the value of the syntax element which specifies the number of layers contained in each Output Layer Set (OLS).
[0243] Item 11. The method as described in Item 10, wherein the syntax elements that commonly indicate the presence of the DPB, HRD, and PTL parameter syntax structures in the VPS include the vps_ptl_dpb_hrd_params_present_flag syntax element.
[0244] Clause 12. The method as described in Clause 11, wherein the syntax element specifying the number of layers contained in each OLS includes the each_layer_is_an_ols_flag syntax element.
[0245] Clause 13. The method described in any of Clauses 1-12, wherein decoding includes decoding.
[0246] Clause 14. The method described in any of Clauses 1-13, wherein decoding includes encoding.
[0247] Clause 15. An apparatus for decoding video data, the apparatus comprising one or more components for performing the methods of any one of Clauses 1-14.
[0248] Clause 16. The device as described in Clause 15, wherein one or more components include one or more processors implemented in a circuit.
[0249] Clause 17. The device as described in any of Clauses 15 and 16 further includes a memory for storing video data.
[0250] Clause 18. The device as described in any one of Clauses 15-17 further includes a display configured to display decoded video data.
[0251] Clause 19. The device as described in any of Clauses 15-18, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.
[0252] Clause 20. The device as described in any one of Clauses 15-19, wherein the device includes a video decoder.
[0253] Clause 21. The device as described in any one of Clauses 15-19, wherein the device includes a video encoder.
[0254] Clause 22. A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform any one of the methods described in Clauses 1-14.
[0255] Clause 23. A method for decoding video data, the method comprising: determining the number of decoded picture buffer (DPB) parameter syntax structures in a video parameter set (VPS) of a current bitstream of the video data; and decoding blocks based on the determined number of DPB parameter syntax structures in the VPS.
[0256] Clause 24. The method as described in Clause 23, wherein determining the number of DPB parameter syntax structures in the VPS includes inferring that the number of DPB parameter syntax structures in the VPS is zero in response to the fact that the syntax element determining the number of DPB parameter syntax structures in the specified VPS does not exist in the bitstream.
[0257] Clause 25. The method as described in Clause 23 further includes determining that the number of DPB parameter syntax structures in the VPS is zero, thereby preventing the signaling notification that the syntax element specifying the number of DPB parameter syntax structures in the VPS does not exist in the bitstream.
[0258] Clause 26. The method as described in any of Clauses 24 and 25, wherein the syntax element specifying the number of DPB parameter syntax structures in the VPS comprises the number of DPB parameter syntax structures in the VPS of the current bitstream of the video data minus 1.
[0259] Clause 27. The method as described in Clause 24, wherein the number of DPB parameter syntax structures in the VPS includes a first number of DPB parameter syntax structures, and wherein the syntax element includes a first instance of the syntax element, the method further comprising determining a second number of DPB parameter syntax structures in the VPS based on a second instance of the syntax element specifying a second number of DPB parameter syntax structures in the VPS.
[0260] Clause 28. The method as described in Clause 25, wherein the number of DPB parameter syntax structures in the VPS includes a first number of DPB parameter syntax structures, and wherein the syntax element includes a first instance of the syntax element, the method further comprising signaling a second number of DPB parameter syntax structures in the VPS based on a second instance of the syntax element specifying a second number of DPB parameter syntax structures in the VPS, wherein the second number of DPB parameter syntax structures is greater than zero.
[0261] Clause 29. A method for decoding video data, the method comprising: decoding a decoded picture buffer (DPB) parameter syntax structure from the SPS when the sequence parameter set (SPS) of the current bitstream of the video data is referenced by a layer as the only layer of the output layer set (OLS); and reconstructing the video data represented by the current bitstream based on the DPB parameter syntax structure.
[0262] Clause 30. The method as described in Clause 29, wherein the DPB parameter syntax structure includes syntax elements that provide information on the DPB size, maximum number of image reorderings, and maximum wait time for one or more OLS.
[0263] Clause 31. The method as described in Clause 29 or 30 further comprises: decoding a syntax element from the video parameter set (VPS) of the current bitstream of video data, minus one number of DPB syntax structures in the specified VPS; and inferring that the number of DPB syntax structures in the VPS is zero in response to determining that the syntax element does not exist in the bitstream.
[0264] Clause 32. The method described in Clause 31 further includes: decoding the DPB parameter syntax structure from the SPS when only one layer is included in any OLS of the VPS.
[0265] Clause 33. The method as described in Clause 31 or 32, wherein the syntax element includes the vps_num_dpb_params_minus1 syntax element.
[0266] Clause 34. The method of any of Clauses 31-33, wherein the syntax element specifying the number of DPB parameter syntax structures in the VPS minus one is the first syntax element, further comprising: decoding from the VPS a second syntax element specifying whether each OLS contains only one layer or is allowed to contain multiple layers, wherein decoding the first syntax element comprises: decoding the first syntax element in response to the second syntax element indicating that each OLS is allowed to contain multiple layers.
[0267] Clause 35. The method as described in Clause 34, wherein the second syntax element includes the each_layer_is_an_ols_flag syntax element.
[0268] Clause 36. The method as described in Clause 34 or 35 further comprises: in response to a second syntax element indicating that each OLS is allowed to contain multiple layers, decoding from the VPS a third syntax element specifying whether the VPS includes a hypothetical reference decoder (HRD) parameter syntax structure.
[0269] Clause 37. The method described in Clause 36, wherein the third syntax element includes hrd_params_present_flag.
[0270] Clause 38. A video decoding apparatus comprising: a memory configured to store at least a portion of a decoded video bitstream; and one or more processors implemented in a circuit and configured to perform the method described in any one of Clauses 29-37.
[0271] Clause 39. A video decoding apparatus comprising components for performing the method described in any one of Clauses 29-37.
[0272] Clause 40. A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the method described in any one of Clauses 29-37.
[0273] Clause 41. Any combination as described in Clauses 1-40.
[0274] It should be recognized that, depending on the example, some actions or events of any of the techniques described herein may be performed in a different sequence, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technique). Furthermore, in some examples, actions or events may be performed concurrently, rather than sequentially, for example, through multithreading, interrupt handling, or multiple processors.
[0275] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, corresponding to a tangible medium such as a data storage medium; or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0276] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that is accessible by a computer. Similarly, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, optical fiber, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, optical fiber, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient, tangible storage media. Disks and optical discs as used herein include compact discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the foregoing should also be included within the scope of computer-readable media.
[0277] Instructions can be executed by one or more processors such as digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the terms "processor" and "processing circuit" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Similarly, the technique can be fully implemented in one or more circuit or logic elements.
[0278] The techniques disclosed herein can be implemented in various devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. More specifically, as described above, the various units can be combined in a codec hardware unit or provided as a collection of interoperable hardware units including one or more processors as described above, combined with appropriate software and / or firmware.
[0279] Various examples have been described. These and other examples are within the scope of the appended claims.
Claims
1. A method for decoding video data, the method comprising: When the sequence parameter set SPS of the current bitstream of video data is referenced by a layer that is the only layer of the output layer set OLS, the parameter syntax structure of the decoded picture buffer DPB is decoded from the SPS. Based on the DPB parameter syntax structure, reconstruct the video data represented by the current bitstream; From the video parameter set VPS of the current bitstream of the video data, decode the syntax elements of the DPB parameter syntax structure in the specified VPS by one less than the number of syntax elements; as well as In response to determining that the syntax element does not exist in the bitstream, it is inferred that the number of DPB syntax structures in the VPS is zero.
2. The method of claim 1, wherein the DPB parameter syntax structure includes syntax elements that provide information on the DPB size, maximum number of image reorderings, and maximum wait time for one or more OLS.
3. The method of claim 1, further comprising: When only one layer is included in any OLS of the VPS, the DPB parameter syntax structure is decoded from the SPS.
4. The method of claim 1, wherein the syntax element includes the vps_num_dpb_params_minus1 syntax element.
5. The method of claim 1, wherein the syntax element specifying the number of DPB parameter syntax structures in the VPS minus one is the first syntax element, further comprising: Decoding the second syntax element from the VPS specifies whether each OLS contains only one layer or is allowed to contain multiple layers, wherein decoding the first syntax element includes: In response to the second syntax element indicating that each OLS is allowed to contain multiple layers, the first syntax element is decoded.
6. The method of claim 5, wherein the second syntax element includes the each_layer_is_an_ols_flag syntax element.
7. The method of claim 5, further comprising: In response to the second syntax element indicating that each OLS is allowed to contain multiple layers, a third syntax element is decoded from the VPS to specify whether the VPS includes a assumed reference decoder HRD parameter syntax structure.
8. The method of claim 7, wherein the third syntax element includes hrd_params_present_flag.
9. A video decoding device, comprising: A memory configured to store at least a portion of a decoded video bitstream; as well as One or more processors implemented in a circuit and configured to perform the following operations: When the sequence parameter set SPS of the decoded video bitstream is referenced by a layer that is the only layer of the output layer set OLS, the parameter syntax structure of the decoded picture buffer DPB is decoded from the SPS. Based on the DPB parameter syntax structure, reconstruct the video data represented by the current bitstream; From the video parameter set VPS of the decoded video bitstream, decode the syntax elements of the DPB parameter syntax structure specified in the VPS minus one; as well as In response to determining that the syntax element does not exist in the bitstream, it is inferred that the number of DPB syntax structures in the VPS is zero.
10. The video decoding device of claim 9, wherein the DPB parameter syntax structure includes syntax elements providing information on the DPB size, maximum number of image reorderings, and maximum latency for one or more OLS.
11. The video decoding device of claim 9, wherein the one or more processors are further configured to: When only one layer is included in any OLS of the VPS, the DPB parameter syntax structure is decoded from the SPS.
12. The video decoding device of claim 9, wherein the syntax element includes the vps_num_dpb_params_minus1 syntax element.
13. The video decoding device of claim 9, wherein the syntax element specifying the number of DPB parameter syntax structures in the VPS minus one is the first syntax element, and wherein the one or more processors are further configured to: The VPS decodes a second syntax element specifying whether each OLS contains only one layer or is allowed to contain multiple layers, wherein, in order to decode the first syntax element, the one or more processors are configured to: In response to a second syntax element indicating that each OLS is allowed to contain multiple layers, the first syntax element is decoded.
14. The video decoding device of claim 13, wherein the second syntax element includes the each_layer_is_an_ols_flag syntax element.
15. The video decoding device of claim 13, wherein the one or more processors are further configured to: In response to a second syntax element indicating that each OLS is allowed to contain multiple layers and from the VPS, a third syntax element is decoded specifying whether the VPS includes a assumed reference decoder HRD parameter syntax structure.
16. The video decoding device of claim 15, wherein the third syntax element includes hrd_params_present_flag.
17. A video decoding device, comprising: The component used to decode the parameter syntax structure of the decoded picture buffer DPB from the sequence parameter set SPS when the current bitstream of video data is referenced by a layer that is the only layer of the output layer set OLS; A component for reconstructing video data represented by the current bitstream based on the DPB parameter syntax structure; A component for decoding, from the current bitstream of video data, the video parameter set VPS, the number of syntax elements in the DPB parameter syntax structure specified in the VPS minus one; as well as A component used to infer that the number of DPB syntax structures in the VPS is zero in response to determining that the syntax element does not exist in the bitstream.
18. The video decoding device of claim 17, further comprising: The video parameter set VPS used to decode the current bitstream of video data specifies the first syntax element of each OLS as either containing only one layer or being allowed to contain multiple layers; A component for responding to a first syntax element indicating that each OLS is allowed to contain multiple layers and from the VPS, decoding a second syntax element specifying the number of DPB parameter syntax structures in the VPS minus one; as well as A component used to infer that the number of DPB syntax structures in the VPS is zero in response to determining that the second syntax element does not exist in the bitstream.
19. A computer-readable storage medium storing instructions, which, when executed, cause one or more processors to: When the sequence parameter set SPS of the current bitstream of video data is referenced by a layer that is the only layer of the output layer set OLS, the parameter syntax structure of the decoded picture buffer DPB is decoded from the SPS. Based on the DPB parameter syntax structure, reconstruct the video data represented by the current bitstream; From the video parameter set VPS of the current bitstream of the video data, decode the syntax elements of the DPB parameter syntax structure in the specified VPS by one less than the number of syntax elements; as well as In response to determining that the syntax element does not exist in the bitstream, it is inferred that the number of DPB syntax structures in the VPS is zero.
20. The computer-readable storage medium of claim 19, further storing instructions that cause the one or more processors to perform the following operations: The first syntax element that specifies whether each OLS contains only one layer or is allowed to contain multiple layers is the video parameter set VPS decoding of the current bitstream of the video data. In response to the first syntax element indicating that each OLS is allowed to contain multiple layers and from the VPS, decode the second syntax element specifying the number of DPB parameter syntax structures in the VPS minus one; as well as In response to determining that the second syntax element does not exist in the bitstream, it is inferred that the number of DPB syntax structures in the VPS is zero.