Encoding and decoding of output layer set data and consistency window data for high-level syntax used in video encoding and decoding.

By inferring OLS PTL index values ​​and simplifying consistency window data processing during video encoding and decoding, the signaling overhead and complexity issues of HLS video data are resolved, improving the performance of video encoders and decoders while reducing bitrate.

CN115152223BActive Publication Date: 2025-10-31QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180016106.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-02-25
Filing Date
2021-02-26
Publication Date
2025-10-31
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies suffer from high signaling overhead and complex correspondence determination when processing High-Level Syntax (HLS) video data, especially in determining the correspondence between the Output Layer Set (OLS) and the Profile-Level-Rank (PTL) data structure, which leads to an increase in processing loops for video encoders and decoders.

Method used

By inferring rather than explicitly decoding the OLS PTL index values, the encoding and decoding are performed directly using the correspondence of the PTL data structure in the VPS. This reduces signaling notifications, simplifies the inference of consistency window data, avoids explicitly notifying the layer of correlation information, and optimizes the bit rate and processing loop of the video bitstream.

Benefits of technology

It reduces signaling overhead and processing loops in the video encoding and decoding process, improves the performance of video encoders and decoders, and reduces the bit rate without compromising video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115152223B_ABST
    Figure CN115152223B_ABST
Patent Text Reader

Abstract

In one example, an apparatus for decoding video data includes one or more processors implemented in a circuit, and the processors are configured to: determine that the value of a syntax element representing the number of grade-level-level (PTL) data structures in a video parameter set (VPS) of a bitstream is equal to the total number of output layer sets (OLS) specified for the VPS; in response to determining that the value of the syntax element representing the number of grade-level-level data structures in the VPS is equal to the total number of OLS specified for the VPS, infer the value of an OLS PTL index value without explicitly decoding the value of the OLS PTL index value; and based on the inferred value of the OLS PTL index value, decode the video data of one or more OLS using the corresponding PTL data structure in the PTL data structures of the VPS.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Application No. 17 / 249,264, filed February 25, 2021; U.S. Provisional Application No. 62 / 983,128, filed February 28, 2020; and U.S. Provisional Application No. 63 / 003,574, filed April 1, 2020, each of which is incorporated herein by reference in its entirety. U.S. Application No. 17 / 249,264 claims the benefits of U.S. Provisional Application Nos. 62 / 983,128 and 63 / 003,574, both filed February 28, 2020. Technical Field

[0003] This disclosure relates to video encoding and decoding, including video encoding and video decoding. Background Technology

[0004] Digital video capabilities can be integrated into a wide range of devices, including digital televisions, digital direct broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite wireless phones, so-called "smartphones," video conferencing equipment, video streaming devices, and more. Digital video devices implement video codec technologies, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Codec (AVC), ITU-T H.265 / High-Efficiency Video Codec (HEVC), and extensions to these standards. By implementing such video codec technologies, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0005] Video coding and decoding techniques include spatial (intra-frame picture) prediction and / or temporal (inter-frame picture) prediction to reduce or eliminate redundancy inherent in video sequences. For block-based video coding and decoding, video slices (e.g., video pictures or portions of video pictures) can be segmented into video blocks, which may also be referred to as codec tree units (CTUs), codec units (CUs), and / or codec nodes. Video blocks in an intra-frame coding (I) slice of a picture are encoded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. Video blocks in an inter-frame coding (P or B) slice of a picture can use either spatial prediction with respect to reference samples in adjacent blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0006] In general, this disclosure describes techniques for processing and encoding / decoding High-Level Syntax (HLS) video data. For example, HLS video data may include Output Layer Sets (OLS) and Prefix-Level-Rank (PTL) data structures. HLS video data may also include data representing the correspondence between OLS and PTL data structures. However, the correspondence between OLS and PTL data structures can be inferred when the number of PTL data structures equals the total number of OLSs specified for a Video Parameter Set (VPS). In particular, the indices used for the OLS and the corresponding PTL data structures can be equal. In this way, the indices do not need to be signaled in this case, which reduces signaling overhead and also simplifies the determination of the correspondence between OLS and PTL data structures.

[0007] As another example, in some cases, where the Picture Parameter Set (PPS) includes data indicating the maximum possible values ​​for picture width and picture height, the consistent cropping window data will not be signaled in the PPS. Instead, the consistent cropping window data for the PPS can be inferred from the corresponding Sequence Parameter Set (SPS), such as the SPS signaling its identifier in the PPS.

[0008] In one example, a method for decoding video data includes: determining that the value of a syntax element representing the number of grade-level-level (PTL) data structures in a video parameter set (VPS) of a bitstream is equal to the total number of output level sets (OLS) specified for the VPS; in response to determining that the value of the syntax element representing the number of grade-level-level data structures in the VPS is equal to the total number of OLS specified for the VPS, inferring the value of an OLS PTL index value without explicitly decoding the value of the OLS PTL index value; and decoding the video data of one or more OLS using the corresponding PTL data structure in the PTL data structures of the VPS based on the inferred value of the OLS PTL index value.

[0009] In another example, an apparatus for decoding video data includes: a memory configured to store video data; and one or more processors implemented in a circuit and configured to: determine that the value of a syntax element representing the number of grade-level-level (PTL) data structures in a video parameter set (VPS) of a bitstream is equal to the total number of output layer sets (OLS) specified for the VPS; in response to determining that the value of the syntax element representing the number of grade-level-level data structures in the VPS is equal to the total number of OLS specified for the VPS, infer the value of an OLS PTL index value without explicitly decoding the value of the OLS PTL index value; and based on the inferred value of the OLS PTL index value, decode the video data of one or more OLS using the corresponding PTL data structure in the PTL data structures of the VPS.

[0010] In another example, instructions are stored on a computer-readable storage medium that, when executed, cause a processor to: determine that the value of a syntax element representing the number of Level-Level-Rank (PTL) data structures in a Video Parameter Set (VPS) representing a bitstream is equal to the total number of Output Level Sets (OLS) specified for the VPS; in response to determining that the value of the syntax element representing the number of Level-Level-Rank (PTL) data structures in the VPS is equal to the total number of OLS specified for the VPS, infer the value of the OLS PTL index without explicitly decoding the value of the OLS PTL index; and, based on the inferred value of the OLS PTL index, decode the video data of one or more OLS using the corresponding PTL data structure in the PTL data structures of the VPS.

[0011] In another example, an apparatus for decoding video data includes: a component for determining that the value of a syntax element representing the number of grade-level-hierarchy (PTL) data structures in a video parameter set (VPS) of a bitstream is equal to the total number of output layer sets (OLS) specified for the VPS; a component for inferring the value of an OLS PTL index value in response to the determination that the value of the syntax element representing the number of grade-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS, without explicitly decoding the value of the OLS PTL index value; and a component for decoding video data of one or more OLS using the corresponding PTL data structure in the PTL data structure of the VPS based on the inferred value of the OLS PTL index value.

[0012] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. Attached Figure Description

[0013] Figure 1This is a block diagram illustrating an example video encoding and decoding system capable of performing the techniques of this disclosure.

[0014] Figure 2A and Figure 2B This is a conceptual diagram illustrating a quadtree-binary tree (QTBT) structure and its corresponding codec tree unit (CTU).

[0015] Figure 3 This is a block diagram illustrating an example video encoder that can perform the techniques of this disclosure.

[0016] Figure 4 This is a block diagram illustrating an example video decoder that can perform the techniques of this disclosure.

[0017] Figure 5 This is a flowchart illustrating an example method for encoding the current block according to the technology of this disclosure.

[0018] Figure 6 This is a flowchart illustrating an example method for decoding the current block according to the technology of this disclosure.

[0019] Figure 7 This is a flowchart illustrating an example method for decoding video data according to the techniques of this disclosure.

[0020] Figure 8 This is a flowchart illustrating another example method for decoding video data according to the techniques of this disclosure. Detailed Implementation

[0021] The Universal Video Codec (VVC) is being developed by the Joint Video Experts Group (JVET) of ITU-T and ISO / IEC to achieve powerful video compression capabilities that surpass ITU-T H.265 High Efficiency Video Codec (HEVC) for a wider range of applications.

[0022] The current draft of VVC specifies the standard bitstream and picture formats, High-Level Syntax (HLS) and semantics, as well as the parsing and decoding process for encoded video data. VVC also specifies in its appendices the Profile / Level / Class (PTL) restrictions, byte stream formats, Hypothetical Reference Decoder (HRD), and Supplemental Enhancement Information (SEI).

[0023] VVC inherits several advanced features from HEVC, such as the concepts of Network Abstraction Layer (NAL) units and parameter sets, slice and wavefront parallel processing, layered encoding and decoding, and the use of SEI messages as supplementary data signaling. VVC introduces even more new advanced features, including the concepts of rectangular slices and subpictures, adaptive picture resolution, mixed NAL unit types, picture headers, Progressive Decoding Refresh (GDR) pictures, virtual boundaries, and a Reference Picture Table (RPL) for reference picture management.

[0024] H.264 / AVC introduced parameter sets to fix the vulnerability of missing picture headers. Parameter sets can be part of the video bitstream or received by the decoder through other means, such as out-of-band transmission, encoder or decoder hardware encoding / decoding, etc. The following is a list of parameter sets currently defined by VVC:

[0025] • Decoding Capability Information (DCI): Includes sublayers and PTL information not needed during the decoding process.

[0026] • Video Parameter Set (VPS): Contains information such as layer dependencies, Output Layer Set (OLS), and PTL information applicable to multiple layers and sublayers.

[0027] • Sequence Parameter Set (SPS): Contains information such as maximum image resolution, consistency window, sub-image layout and ID mapping, RPL, and sequence-level codec parameters applicable to the codec layer video sequence (CLVS).

[0028] • Picture Parameter Set (PPS): Contains information such as picture resolution, consistency window and scaling window, slice and slice segmentation, and picture-level encoding and decoding parameters applicable to multiple pictures.

[0029] • The Adaptive Parameter Set (APS) contains the Adaptive Loop Filter (ALF): parameters applicable to the slice, scaling list parameters, and Luminance Mapping with Chroma Scaling (LMCS) parameters.

[0030] VVC also specifies that the image header carries parameters that can be shared by multiple slices in the image to reduce overhead.

[0031] This disclosure recognizes that certain features of the current VVC HLS design based on VVC Draft 8 can be improved. Such improvements can reduce the bit rate without negatively impacting video fidelity (e.g., increasing distortion). Similarly, such improvements can improve the performance of the video encoder and / or decoder. For example, these improvements can reduce the number of processing loops performed by the video codec without negatively impacting video fidelity.

[0032] As an example, this disclosure describes a technique related to the signaling correspondence between the Output Layer Set (OLS) and the Priority-Level-Rank (PTL) data structure used for the OLS. The OLS comprises a set of layers of video data to be output (which may be equal to or smaller than the set of layers of video data to be decoded). The PTL data structure describes the priority and level of the corresponding OLS, as well as the ranks within those levels. Typically, whenever more than one PTL data structure is signaled in the VPS, an index is encoded and decoded in the VPS to represent the correspondence between the PTL data structure and the OLS.

[0033] However, this disclosure recognizes that when the number of PTL data structures is equal to the number of OLSs on the VPS, the video encoder and video decoder can be configured to infer the correspondence between the PTL data structures and the OLSs. Specifically, the video encoder and video decoder can determine that the i-th PTL data structure corresponds to the i-th OLS (for all i values ​​between 0 and the total number of OLSs). Therefore, the video encoder and video decoder can use the values ​​of the i-th PTL data structure to encode or decode the video data of the i-th OLS. For example, the video encoder and / or video decoder can activate encoding / decoding tools indicated to be used by the i-th PTL data structure and disable encoding / decoding tools not used according to the i-th PTL data structure.

[0034] As another example, the Sequence Parameter Set (SPS) and / or Picture Parameter Set (PPS) can signal the data representing the consistency window. Generally, the consistency window specifies the picture region considered for use in the consistent picture output. The SPS can signal the consistency window for the entire picture sequence, while the PPS can signal the refinement of the consistency window for each picture within the sequence. However, this disclosure recognizes that when the picture size of a picture in a sequence is the maximum possible picture size for that sequence, it is not necessary to signal the consistency window refinement data in the PPS for that picture, because the consistency window can be inferred from the SPS.

[0035] Accordingly, the video encoder and video decoder can infer the value of a syntax element (e.g., pps_conformance_window_flag) indicating whether the conformance window syntax element itself is signaled, and further, infer the value of the conformance window syntax element when the image has a maximum size. Therefore, the video encoder and video decoder can be configured to encode and decode the value of the conformance window syntax element of the PPS only when the image size indicated by the PPS is smaller than the maximum image size. The image size can be smaller than the maximum image size when the signaled image width and / or image height are respectively smaller than the corresponding maximum width and / or height.

[0036] Furthermore, when a PPS syntax element (e.g., pps_conformance_window_flag) indicating whether a conformance window syntax element has been signaled indicates that the conformance window syntax element has not been signaled (e.g., when the image size is equal to the maximum image size), the video encoder and video decoder can be configured to infer that the PPS conformance window syntax element is equal to the corresponding SPS conformance window syntax element.

[0037] Decoding Capability Information (DCI) is currently designated as a non-VCL NAL unit, but this information is not necessary for the decoding process, and neither the parameter set nor the VCL NAL unit relates to DCI. This disclosure describes techniques including removing the DCI parameter set and carrying decoding capability information in the SEI message. Table 1 below shows the DCI syntax structure for VVC.

[0038] Table 1—DCI RBSP Syntax

[0039]

[0040] According to VVC, layer dependency information is signaled in the VPS, subject to the value of `vps_all_independent_layers_flag`. Unless all layers are independent, the dependency of each layer is explicitly signaled. Table 2 below shows a partial VPS syntax structure according to VVC.

[0041] Table 2—VPS RBSP

[0042]

[0043]

[0044] A common scenario for layer encoding / decoding is that the i-th layer may have only one directly related layer, and that directly related layer is the (i-1)-th layer, where i is greater than 0. This is based on the ols_mode_idc semantics of VVC, as detailed below.

[0045] `ols_mode_idc` equal to 0 specifies that the total number of OLS specified by the VPS is equal to `vps_max_layers_minus1+1`, the i-th OLS includes layers with layer indices from 0 to i (inclusive), and for each OLS, only the highest-ranking layer in the OLS is output. `ols_mode_idc` equal to 1 specifies that the total number of OLS specified by the VPS is equal to `vps_max_layers_minus1+1`, the i-th OLS includes layers with layer indices from 0 to i (inclusive), and for each OLS, all layers in the OLS are output.

[0046] This disclosure describes techniques that can be used to avoid explicit signaling notifications of correlation information for each layer, for example, in such a common situation. By avoiding explicit signaling notifications of this correlation information in such a manner, the bitrate of a video bitstream that includes multiple layers in this manner can be reduced without degrading video quality (e.g., without increasing video distortion). Similarly, video encoders and decoders can avoid processing video correlation information with explicit signaling notifications, which can improve the performance of video encoders and decoders.

[0047] Table 2 above includes the `max_tid_il_ref_pics_plus1` syntax element to indicate sublayers that can be used as inter-layer reference pictures (ILRPs) to decode the current layer's picture. It's possible that all layers share the same sublayer attributes, making explicit signaling for each layer potentially unnecessary. Furthermore, sublayer ILRP indications are primarily used to discard lower-layer sublayer pictures that were not used for reference or output. Therefore, notifying each layer of such an indication via signaling may not be effective.

[0048] VVC specifies that the 0th Output Layer Set (OLS) contains only the lowest layer, and in the derivation of LayerIdInOls[i][j], LayerIdInOls[0][0] is inferred to be equal to vps_layer_id[0]. There may be multiple available dependency trees in a VPS. This disclosure describes a technique for assigning each independent non-base layer to a separate OLS.

[0049] According to VVC, the index ols_ptl_idx of the list of profile_tier_level() (PTL) syntax structures and the index ols_dbp_params_index of the list of decoded picture buffer (DPB) dpb_parameters() syntax structures are signaled in the VPS. This disclosure describes techniques for skipping this signaling to reduce signaling overhead when the total number of syntax structures equals the same number of OLS. Therefore, these techniques can reduce the bitrate of the video bitstream without degrading video quality.

[0050] The first 16 bits of the SPS can be equal to "00000000 0000000", which can be used to emulate the start code depending on the following syntax element values. Even though the use of the syntax element `emulation_prevention_three_byte` for encapsulation can prevent the emulation of the start code within the NAL unit, preventing start code emulation in the first place may be more direct. This disclosure describes techniques for preventing start code emulation that can increase the bitrate of the video bitstream without degrading video quality.

[0051] VVC specifies the following constraint for the consistency window syntax element in PPS: When pic_width_in_luma_samples equals pic_width_max_in_luma_samples and pic_height_in_luma_samples equals pic_height_max_in_luma_samples, the bitstream consistency requirement is that pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset are equal to sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset, respectively. An alternative method to simplify the constraint is to directly constrain pps_conformance_window_flag. In VVC, the inference of the values ​​of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset is "when pps_conformance_window_flag equals 0, the values ​​of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset are inferred to be equal to 0." This disclosure describes techniques for improving the inference of these values, for example, to reduce the processing operations performed by the video encoder and video decoder.

[0052] The VVC specifies the SPS Picture Sequence Count (POC) Most Significant Bit (MSB) syntax element for signaling to the POCMSB, for example, to indicate which picture to be included in a reference picture set or list. Specifically, the VVC specifies the value of `sps_poc_msb_flag` to indicate whether the `ph_poc_msb_present_flag` syntax element exists in the picture header (PH) of the reference SPS. When `sps_poc_msb_flag` equals 1, it signals to another syntax element, `poc_msb_len_minus1`, to specify the length of the `poc_msb_val` syntax element of the reference SPS in the PH using bits. Tables 3 and 4 below show some of the relevant SPS and PH syntax structures. This disclosure describes techniques for aggregating two syntax elements into one for simplification, which reduces processing requirements on the video encoder and decoder and also reduces the bitrate of the video bitstream without degrading video quality. This disclosure also describes examples of including signaling to a common constraint flag to allow the POC MSB to be updated at the PH.

[0053] Table 3—SPS RBSP

[0054]

[0055]

[0056] Table 4—Image Headers (RBSP)

[0057]

[0058] VVC specifies data for the slice height deviation of the signaling notification image slice. Specifically, according to VVC, "num_exp_slices_in_tile[i] specifies the number of slice heights explicitly provided in the current slice containing more than one rectangular slice. The value of num_exp_slices_in_tile[i] should be in the range of 0 to RowHeight[tileY]-1 (inclusive), where tileY is the slice row index containing the i-th slice. When it does not exist, the value of num_exp_slices_in_tile[i] is inferred to be equal to 0. When num_exp_slices_in_tile[i] is equal to 0, the value of the variable NumSlicesInTile[i] is inferred to be equal to 1." According to VVC, when a slice contains multiple slices, num_exp_slices_in_tile[i] does not exist, and the value of the variable NumSlicesInTile[i] is inferred to be equal to 1. According to VVC, when a slice contains a tile, the value of num_exp_slices_in_tile[i] is equal to 1, and NumSlicesInTile[i] is also deduced to be equal to 1.

[0059] This disclosure recognizes that these semantics may be problematic in certain scenarios. Specifically, when num_exp_slices_in_tile[i] is greater than 0, for k, the variables NumSlicesInTile[i] and SliceHeightInCtusMinus1[i+k] can be derived as follows, where k is in the range of 0 to NumSlicesInTile[i]-1:

[0060]

[0061]

[0062] This disclosure recognizes that the above derivation may be problematic when there is only one slice in the slice, because SliceHeightInCtusMinus1[i-1] is undefined when numExpSliceInTile equals 1. This disclosure describes techniques that can be used to address these problems.

[0063] Figure 1This is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. The techniques of this disclosure generally refer to encoding and / or decoding video data. Typically, video data includes any data used for processing video. Thus, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.

[0064] like Figure 1 As shown, in this example, system 100 includes source device 102, which provides encoded video data to be decoded and displayed by target device 116. Specifically, source device 102 provides video data to target device 116 via computer-readable medium 110. Source device 102 and target device 116 can include any of a variety of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, mobile phones such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and target device 116 may be equipped for wireless communication and may therefore be referred to as wireless communication devices.

[0065] exist Figure 1 In the example, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Target device 116 includes an input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of target device 116 can be configured to apply techniques for encoding and decoding values ​​of high-level syntax elements. Therefore, source device 102 represents an example of a video encoding device, while target device 116 represents an example of a video decoding device. In other examples, the source and target devices may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, target device 116 may interface with an external display device instead of including an integrated display device.

[0066] like Figure 1The system 100 shown is merely an example. Typically, any digital video encoding and / or decoding device can perform techniques for encoding and decoding values ​​of high-level syntax elements. Source device 102 and target device 116 are merely examples of such encoding / decoding devices, where source device 102 generates encoded video data for transmission to target device 116. This disclosure refers to "encoding / decoding" devices as devices that perform the encoding and / or decoding of data. Thus, video encoder 200 and video decoder 300 represent examples of encoding / decoding devices, specifically a video encoder and a video decoder, respectively. In some examples, source device 102 and target device 116 may operate in a substantially symmetrical manner, such that each of source device 102 and target device 116 includes video encoding and decoding components. Therefore, system 100 can support one-way or two-way video transmission between source device 102 and target device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0067] Typically, video source 104 represents the source of video data (i.e., raw, unencoded video data) and provides a continuous series of pictures (also referred to as "frames") of video data to video encoder 200, which encodes the data of the pictures. Video source 104 of source device 102 may include video capture devices such as cameras, video archives containing previously captured raw video, and / or video feed interfaces that receive video from video content providers. Alternatively, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from the order of receipt (sometimes referred to as "display order") into an encoding / decoding order for encoding and decoding. Video encoder 200 may generate a bitstream comprising encoded video data. The source device 102 can then output encoded video data to a computer-readable medium 110 via an output interface 108 for reception and / or retrieval by an input interface 122 of a target device 116, for example.

[0068] The memory 106 of source device 102 and the memory 120 of target device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data (e.g., raw video from video source 104) and raw decoded video data from video decoder 300. Additionally or alternatively, memories 106 and 120 may store software instructions that can be executed by, for example, video encoder 200 and video decoder 300. Although shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store encoded video data, such as data output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106 and 120 may be allocated as one or more video buffers, for example, to store raw, decoded, and / or encoded video data.

[0069] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded video data from source device 102 to target device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to transmit encoded video data directly to target device 116 in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, output interface 108 can modulate the transmitted signal including the encoded video data, and input interface 122 can demodulate the received transmitted signal. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, switch, base station, or any other equipment that facilitates communication from source device 102 to target device 116.

[0070] In some examples, source device 102 can output encoded data to storage device 112 from output interface 108. Similarly, target device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as hard disks, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0071] In some examples, source device 102 may output encoded video data to file server 114 or another intermediate storage device that may store the encoded video generated by source device 102. Target device 116 may access the stored video data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing encoded video data and sending it to target device 116. File server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Target device 116 may access the encoded video data from file server 114 via any standard data connection, including an internet connection. This may include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or a combination of both suitable for accessing encoded video data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming protocol, a download transfer protocol, or a combination thereof.

[0072] Output interface 108 and input interface 122 may represent a wireless transmitter / receiver, a modem, a wired network component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data such as encoded video data according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, or similar standards. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured according to other wireless standards, such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee). TM ),Bluetooth TM The source device 102 and / or the target device 116 may be respective system-on-a-chip (SoC) devices. For example, the source device 102 may include an SoC device for performing functions belonging to the video encoder 200 and / or output interface 108, and the target device 116 may include an SoC device for performing functions belonging to the video decoder 300 and / or input interface 122.

[0073] The technology disclosed herein can be applied to video encoding and decoding that supports any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as Dynamic Adaptive Streaming (DASH) via HTTP), digital video encoded to a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0074] The input interface 122 of the target device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, which is also used by the video decoder 300, such as syntax elements having values ​​describing the characteristics and / or processing of video blocks or other encoding / decoding units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays a decoded image of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device.

[0075] Although not in Figure 1 As shown, but in some examples, the video encoder 200 and video decoder 300 may each be integrated with the audio encoder and / or audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to handle multiplexed streams that include both audio and video in a common data stream. Where applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as User Datagram Protocol (UDP).

[0076] The video encoder 200 and video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented in part as software, the device may store instructions for software in a suitable non-transitory computer-readable medium and use one or more processors to execute those instructions in hardware to perform the technology of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.

[0077] The video encoder 200 and video decoder 300 may operate according to video codec standards such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC) or its extensions such as Multi-View and / or Scalable Video Coding Extensions. Alternatively, the video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards such as Universal Video Coding (VVC). A draft of the VVC standard is described below: “Versatile Video Coding (Draft 8)” by Bross et al., Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 17th meeting: Brussels, BE, January 7-17, 2020, JVET-Q2001-vA (hereinafter referred to as “VVC Draft 8”). However, the technology disclosed herein is not limited to any specific codec standard.

[0078] Typically, video encoder 200 and video decoder 300 can perform block-based encoding and decoding of images. The term "block" generally refers to a structure that includes data to be processed (e.g., data encoded, decoded, or otherwise used during the encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, video encoder 200 and video decoder 300 can encode and decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, video encoder 200 and video decoder 300 can encode and decode luminance and chrominance components, rather than encoding and decoding the red, green, and blue (RGB) data of the image samples, where the chrominance components may include both red and blue chrominance components. In some examples, video encoder 200 converts the received RGB formatted data to a YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and post-processing units (not shown) may perform these conversions.

[0079] This disclosure can generally refer to the encoding and decoding (e.g., encoding and decoding) of images to include the process of encoding or decoding data for the image. Similarly, this disclosure can refer to the encoding and decoding of blocks of images to include the process of encoding or decoding data for the blocks, such as prediction and / or residual encoding and decoding. Encoded video bitstreams typically include a series of values ​​for syntax elements representing encoding and decoding decisions (e.g., encoding / decoding modes) and the segmentation of images into blocks. Therefore, references to encoding and decoding images or blocks should generally be understood as encoding and decoding the values ​​of the syntax elements that form the images or blocks.

[0080] HEVC defines various blocks, including Codec Units (CUs), Prediction Units (PUs), and Transform Units (TUs). According to HEVC, a video codec (such as a video encoder 200) partitions a Codec Tree Unit (CTU) into CUs based on a quadtree structure. That is, the video codec partitions the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and a CU with such a leaf node can include one or more PUs and / or one or more TUs. The video codec can also partition PUs and TUs. For example, in HEVC, a Residual Quadtree (RQT) represents a partition of a TU. In HEVC, a PU represents inter-frame prediction data, while a TU represents residual data. Intra-predicted CUs include intra-frame prediction information, such as intra-frame mode indications.

[0081] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, the video codec (such as video encoder 200) segments the image into multiple codec tree units (CTUs). Video encoder 200 can segment CTUs according to a tree structure such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concept of multiple segmentation types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure consists of two layers: a first layer segmented according to quadtree segmentation, and a second layer segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to codec units (CUs).

[0082] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning (also known as triplet tree (TT)). A ternary tree or triplet tree partition is a partition that divides a block into three sub-blocks. In some examples, a ternary tree or triplet tree partition divides a block into three sub-blocks without using a center to partition the original block. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0083] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).

[0084] Video encoder 200 and video decoder 300 can be configured to use quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structures based on HEVC. For illustrative purposes, the description of the techniques of this disclosure is presented in relation to QTBT segmentation. However, it should be understood that the techniques of this disclosure can also be applied to video codecs configured to use quadtree segmentation or other types of segmentation.

[0085] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a block can refer to a rectangular area of ​​a row of CTUs within a specific slice in an image. A slice can be a rectangular area of ​​CTUs within a specific slice column and a specific slice row in an image. A slice column is a rectangular area of ​​CTUs with a height equal to the image height and a width specified by a syntax element (e.g., a syntax element in the image parameter set). A slice row is a rectangular area of ​​CTUs with a height specified by a syntax element (e.g., a syntax element in the image parameter set) and a width equal to the image width.

[0086] In some examples, a slice can be divided into multiple tiles, each tile potentially including one or more CTU rows within the slice. A slice that is not divided into multiple tiles can also be referred to as a tile. However, a tile that is a proper subset of a slice cannot be referred to as a slice.

[0087] Tiles in an image can also be arranged into slices. A slice can be an integer number of tiles in the image, which can be exclusively contained within a single Network Abstraction Layer (NAL) unit. In some examples, a slice consists of multiple complete pieces or a continuous sequence of complete tiles consisting of only one piece.

[0088] This disclosure uses "N×N" and "N by N" interchangeably to refer to the sample size of a block (such as a CU or other video block) in the vertical and horizontal dimensions, for example, 16×16 samples or 16 by 16 samples. Typically, a 16×16 CU has 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU may include N×M samples, where M is not necessarily equal to N.

[0089] The video encoder 200 encodes video data used for the control unit (CU) to represent prediction and / or residual information, as well as other information. Prediction information indicates how the CU should be predicted to form a prediction block of the CU. Residual information typically represents the point-by-point difference between the samples of the CU before encoding and the prediction block.

[0090] To predict a Cubicles (CUs), the video encoder 200 typically forms prediction blocks of the CUs through inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU from data of a previously encoded / decoded image, while intra-frame prediction generally refers to predicting the CU from data of a previously encoded / decoded image of the same image. To perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate prediction blocks. The video encoder 200 typically performs a motion search to identify reference blocks that closely match the CU (e.g., in terms of the difference between the CU and a reference block). The video encoder 200 can use sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to compute difference metrics to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.

[0091] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter-frame prediction mode. In affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular motion types).

[0092] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of VVC provide sixty-seven intra-frame prediction modes, including various orientation modes, as well as planar and DC modes. Typically, the video encoder 200 selects an intra-frame prediction mode that describes the neighboring samples of the current block (e.g., a block of a CU) to predict the prediction samples of the current block from them. Assuming the video encoder 200 encodes and decodes CTUs and CUs in raster scan order (from left to right, from top to bottom), such samples are typically located above, above to the left of, or to the left of the current block in the same frame as the current block.

[0093] The video encoder 200 encodes data representing the prediction mode of the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used, as well as motion information for the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may use Advanced Motion Vector Prediction (AMVP) or merging modes to encode motion vectors. The video encoder 200 may use similar modes to encode motion vectors for affine motion compensation modes.

[0094] Following prediction, such as intra-frame or inter-frame prediction of a block, the video encoder 200 can compute residual data for that block. The residual data (e.g., a residual block) represents the sample-by-sample difference between the block and its predicted block, which is formed using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply a Discrete Cosine Transform (DCT), integer transform, wavelet transform, or conceptually similar transforms to the residual video data. Additionally, the video encoder 200 can apply a second transform after the first transform, such as a Mode-dependent Indivisible Quadratic Transform (MDNSST), a Signal-dependent Transform, a Karhunen-Loeve Transform (KLT), etc. The video encoder 200 generates transform coefficients after applying one or more transforms.

[0095] As described above, after any transform that produces the transform coefficients, the video encoder 200 can perform quantization of the transform coefficients. Quantization generally refers to the process in which transform coefficients are quantized to minimize the amount of data used to represent the coefficients, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the coefficients. For example, the video encoder 200 can round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bit-by-bit right shift of the value to be quantized.

[0096] After quantization, the video encoder 200 can scan the transform coefficients to produce a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place higher-energy (and therefore lower-frequency) coefficients before the vector and lower-energy (and therefore higher-frequency) transform coefficients after the vector. In some examples, the video encoder 200 can scan the quantized transform coefficients using a predetermined scan order to produce a serialized vector, and then entropy-encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can, for example, entropy-encode the one-dimensional vector according to context-adaptive binary arithmetic codec (CABAC). The video encoder 200 can also entropy-encode the values ​​of syntax elements used to describe metadata associated with the encoded video data for use by the video decoder 300 in decoding the video data.

[0097] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. Context may involve, for example, whether the neighboring values ​​of a symbol are zero. Probability determination can be based on the context assigned to the symbol.

[0098] The video encoder 200 can also generate syntax data for the video decoder 300, for example, from picture headers, block headers, slice headers, or other syntax data (such as sequence parameter sets (SPS), picture parameter sets (PPS), or video parameter sets (VPS)). These syntax data can be block-based, picture-based, and sequence-based. The video decoder 300 can similarly decode such syntax data to determine how to decode the corresponding video data.

[0099] In this manner, the video encoder 200 can generate a bitstream that includes encoded video data, such as syntax elements describing the segmentation of images into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0100] Typically, the video decoder 300 performs the inverse of the process performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC, in a manner substantially similar to but inverse of the CABAC encoding process of the video encoder 200, to decode the values ​​of the syntax elements of the bitstream. Syntax elements can define segmentation information used to segment the image into CTUs and to segment each CTU according to a corresponding segmentation structure such as a QTBT structure to define the CUs of the CTUs. Syntax elements can also define prediction and residual information for blocks (e.g., CUs) of the video data.

[0101] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inversely transform the quantized transform coefficients of the block to reconstruct the residual block of the block. The video decoder 300 uses the prediction mode (intra-frame prediction or inter-frame prediction) notified by signaling and relevant prediction information (e.g., motion information for inter-frame prediction) to form a prediction block of the block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reconstruct the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the block boundaries.

[0102] This disclosure may generally refer to "signaling notification" of certain information, such as syntax elements. The term "signaling notification" can generally refer to communication of values ​​for syntax elements and / or other data used to decode encoded video data. That is, the video encoder 200 may signal the values ​​of syntax elements in the bitstream. Typically, signaling notification refers to generating values ​​in the bitstream. As described above, the source device 102 may transmit the bitstream to the target device 116 substantially in real time (or non-real time, such as when syntax elements are stored in storage device 112 for later retrieval by the target device 116).

[0103] According to various techniques disclosed herein, video encoder 200 and video decoder 300 can encode and decode (encode and decode, respectively) the values ​​of high-level syntax elements. In one example, video encoder 200 and video decoder 300 can encode and decode Decoding Capability Information (DCI) Supplemental Enhancement Information (SEI) messages. DCI SEI messages can provide video decoder 300 with the decoding capability information required for the associated bitstream, including the maximum number of sublayers and a list of profile / tier / level information. Table 5 below shows an example DCI SEI message syntax, where profile_tier_level() can be based on the current VVC specification.

[0104] Table 5—Example of DCI SEI Information I

[0105]

[0106] The semantics of the syntactic elements in the examples in Table 5 can be defined as follows:

[0107] `dci_max_sublayers` specifies the maximum number of temporal sublayers that may exist in a layer within each CVS of the bitstream. The value of `dci_max_sublayers` should be in the range of 1 to 7 (inclusive).

[0108] dci_num_ptls specifies the number of profile_tier_level() syntax structures in the DCI SEI message.

[0109] Table 6 shows another example DCI SEI message syntax, which includes a list of the maximum general_profile_idc, the maximum general_level_idc, and the general_sub_profile_idc supported by the video decoder 300 in the DCI SEI message.

[0110] Table 6—Example of DCI SEI Information II

[0111]

[0112] The semantics of the syntactic elements in the examples in Table 6 can be defined as follows:

[0113] `dci_max_profile_idc` specifies the highest value of `general_profile_idc` supported by the decoder. The bitstream must not contain any value of `general_profile_idc` other than those specified in Appendix A of VVC.

[0114] dci_max_level_idc specifies the highest value of general_level_idc supported by the decoder. The bitstream must not contain any general_level_idc other than those specified in VVC Appendix A.

[0115] The semantics of dci_max_sublayers, dci_num_sub_profiles, and general_sub_profile_idc can be the same as those specified in VVC.

[0116] In another example, the profile_idc value in the signaling notification of the DCI SEI message can represent the profile that provides the preferred decoding result or preferred bitstream identifier determined by the video encoder 200.

[0117] In this manner, video encoder 200 can encode DCI SEI messages, and video decoder 300 can decode DCI SEI messages including, for example, syntax elements indicating the maximum number of sublayers that may exist in each codec video sequence (CVS) of the bitstream. Video encoder 200 can also encode DCI SEI messages, and video decoder 300 can also decode DCI SEI messages including syntax elements indicating the number of grade / hierarchy / level syntax structures included in the DCI SEI message. Alternatively, another element of source device 102 (e.g., output interface 108 or not in...) Figure 1 The post-processing unit shown in the diagram) and another element of the target device 116 (e.g., input interface 122 or not shown in the diagram) and the target device 116. Figure 1 The preprocessing unit or media data retrieval unit shown can process the DCI SEI message of the video bitstream. The target device 116 can use this data to determine whether the video decoder 300 is capable of decoding the video bitstream. The source device 102 can signal this data to indicate the capability required for the video decoder to decode the corresponding video bitstream.

[0118] When video decoder 300 is unable to decode the video bitstream, target device 116 can use, for example, different corresponding DCI SEI messages that video decoder 300 can decode to select an alternative video bitstream. For example, if multiple versions of a specific video program are available, target device 116 can retrieve the DCI SEI message for each version and select one of the versions that video decoder 300 can decode, as indicated by the information in the corresponding DCI SEI message. In this way, these techniques allow target device 116 to determine whether video decoder 300 can decode the video bitstream without retrieving the DCI parameter set of the video bitstream, which reduces wasted bandwidth consumption and reduces latency associated with retrieving video data that video decoder 300 can decode.

[0119] Video encoder 200 and video decoder 300 may additionally or alternatively be configured to encode and decode layer-related information as discussed below. In the examples below, the layer-related information is specified in a video parameter set (VPS). Table 7 below shows example VPSs including layer-related information according to these technologies, where the annotation [added: "added text"] is used to indicate text added relative to the VVC.

[0120] Table 7—Example VPSs including tier dependency indicators

[0121]

[0122]

[0123] The semantics of the `vps_layer_dependency_idc` syntax element for the VPS in the example in Table 7, and other existing syntax elements for VVC, can be defined as follows. Text added relative to VVC is indicated by the comment `[added: "added text"]`.

[0124] [Added: "vps_layer_dependency_idc equal to 0 indicates that all layers in CVS are independently encoded and decoded, and inter-layer prediction is not used. vps_layer_dependency_idc equal to 1 indicates that all non-base layers in CVS use inter-layer prediction, the layer at index i is the direct reference layer of the layer at index (i+1), and all sub-layers of all layers except the highest layer are used for inter-layer prediction. vps_layer_dependency_idc equal to 2 indicates that one or more layers in CVS can use inter-layer prediction. The value of vps_layer_dependency_idc should be in the range of 0 to 2 (inclusive). The value of ols_mode_idc 3 is reserved by ITU-T|ISO / IEC for future use. When it does not exist, the value of vps_layer_dependency_idc is inferred to be equal to 0."]

[0125] `vps_independent_layer_flag[i]` equal to 1 indicates that the layer at index `i` does not use inter-layer prediction. `vps_independent_layer_flag[i]` equal to 0 indicates that the layer at index `i` can use inter-layer prediction, and the VPS contains the syntax element `vps_direct_ref_layer_flag[i][j]`, where `j` ranges from 0 to `i-1` (inclusive). [added: "When `vps_independent_layer_flag` does not exist, the value of `vps_independent_layer_flag[i]` is inferred to be equal to 1 when `vps_layer_dependency_idc` equals 0; and the value of `vps_independent_layer_flag[i]` is inferred to be equal to 0 when `vps_layer_dependency_idc` equals 1."]

[0126] In the alternative expression, the inference rule can be expressed as: [added: "When vps_independent_layer_flag does not exist, the value of vps_independent_layer_flag[i] is inferred to be equal to 1 - vps_layer_dependency_idc."]

[0127] `vps_direct_ref_layer_flag[i][j]` equal to 0 indicates that the layer at index `j` is not a direct reference layer to the layer at index `i`. `vps_direct_ref_layer_flag[i][j]` equal to 1 indicates that the layer at index `j` is a direct reference layer to the layer at index `i`. [added: "When `vps_direct_ref_layer_flag` does not exist, the following inference is made: When `vps_layer_dependency_idc` equals 0, `vps_direct_ref_layer_flag[i][j]` is inferred to be equal to 0, where `i` and `j` range from 0 to `vps_max_layers_minus1` (inclusive). When `vps_layer_dependency_idc` equals 1, `vps_direct_ref_layer_flag[i][j]` is inferred to be equal to 0.] `_layer_flag[i][j]` is inferred to be 0 when `i` is not equal to (j+1) and inferred to be 1 when `i` is equal to (j+1), where `i` and `j` range from 0 to `vps_max_layers_minus1` (inclusive). When `vps_layer_dependency_idc` is not equal to 0, there should be at least one `j` value in the range 0 to `i-1` (inclusive) such that `vps_direct_ref_layer_flag[i][j]` equals 1.

[0128] `max_tid_il_ref_pics_plus1[i]` equal to 0 indicates that non-IRAP images in layer i do not use inter-layer prediction. `max_tid_il_ref_pics_plus1[i]` greater than 0 indicates that, for decoding images in layer i, no image with a TemporalId greater than `max_tid_il_ref_pics_plus1[i] - 1` is used for ILRP. [added: "When `max_tid_il_ref_pics_plus1` does not exist, the value of `max_tid_il_ref_pics_plus1` is inferred as follows: when `vps_layer_dependency_idc` equals 0, the value of `max_tid_il_ref_pics_plus1[i]` is inferred to be 0; when `vps_layer_dependency_idc` equals 1, the value of `max_tid_il_ref_pics_plus1[i]` is inferred to be 7."]

[0129] When [added: "vps_layer_dependency_idc equals 0"] and each_layer_is_an_ols_flag equals 0, the value of ols_mode_idc is inferred to be equal to 2.

[0130] Accordingly, the video encoder 200 and the video decoder 300 can be configured to encode and decode the values ​​of the VPS layer correlation indicator syntax element, the VPS independent layer syntax element, the VPS direct reference layer syntax element, and the maximum temporal domain identifier interlayer reference picture syntax element according to the example syntax and semantics above.

[0131] Additionally or alternatively, the video encoder 200 and video decoder 300 may encode and decode sublayer interlayer reference picture (ILRP) information as follows. Table 8 below shows an example VPS that includes sublayer ILRP information. Specifically, as in an existing VPS of VVC, max_tid_il_ref_pics_plus1 is used to indicate the sublayer that includes the picture used as the ILRP for decoding the picture of the current layer. However, in the example in Table 8, under certain conditions, the video encoder 200 and video decoder 300 may encode and decode only the value of max_tid_il_ref_pics_plus1.

[0132] Table 8—VPS Examples Including ILRP Indicator Syntax

[0133]

[0134]

[0135] The semantics of the introduced vps_all_layers_same_ilrp_sublayers_flag and vps_default_max_tid_il_ref_pics_plus1 syntax elements can be as follows:

[0136] A value of 1 for `vps_all_layers_same_ilrp_sublayers_flag` indicates that the syntax element `vps_default_max_tid_il_ref_pics_plus1` exists. A value of 0 for `vps_all_layers_same_ilrp_sublayers_flag` indicates that the syntax element `vps_default_max_tid_il_ref_pics_plus1` does not exist.

[0137] A value of 0 for `vps_default_max_tid_il_ref_pics_plus1` indicates that inter-layer predictions are not used by non-IRAP (or as an alternative to non-CLVSS) in layer i. A value greater than 0 for `vps_default_max_tid_il_ref_pics_plus1` indicates that images with a TemporalId greater than `vps_default_max_tid_il_ref_pics_plus1` are used for ILRP. When `vps_default_max_tid_il_ref_pics_plus1` does not exist, its value is inferred to be 7.

[0138] The semantics of the existing VVC VPS syntax elements can be modified as follows, using [added: "added text"] to represent added text relative to the existing VVC.

[0139] A value of 0 for `max_tid_il_ref_pics_plus1[i]` indicates that inter-layer predictions are not used by non-IRAP (or as an alternative non-CLVSS) images of layer i. A value greater than 0 for `max_tid_il_ref_pics_plus1[i]` indicates that, for decoding images of layer i, no images with a TemporalId greater than `max_tid_il_ref_pics_plus1[i]-1` are used for ILRP. [added: "When max_tid_il_ref_pics_plus1 does not exist, and when vps_all_layers_same_ilrp_sublayers_flag equals 1, the value of max_tid_il_ref_pics_plus1[i] is inferred to be equal to vps_default_max_tid_il_ref_pics_plus1; otherwise, the value of max_tid_il_ref_pics_plus1[i] is inferred to be equal to 7."] In an alternative semantic expression, [added: "When max_tid_il_ref_pics_plus1[i] does not exist, the value of max_tid_il_ref_pics_plus1 is inferred to be equal to vps_default_max_tid_il_ref_pics_plus1."]

[0140] In another example, the indication of sublayer ILRP usage can be communicated only to output layer signaling, as shown in Table 9 below:

[0141] Table 9—ILRP Usage Instructions Syntax

[0142]

[0143] The semantics of max_tid_ref_present_flag[i] and max_tid_il_ref_pics_plus1[i] can be as follows:

[0144] `max_tid_ref_present_flag[i]` equal to 1 indicates that the syntax element `max_tid_il_ref_pics_plus1[i]` exists. `max_tid_ref_present_flag[i]` equal to 0 indicates that the syntax element `max_tid_il_ref_pics_plus1[i]` does not exist. When `max_tid_ref_present_flag[i]` does not exist, the value of `max_tid_ref_present_flag[i]` is inferred to be equal to 0. `max_tid_il_ref_pics_plus1[i]` equal to 0 indicates that inter-layer prediction is not used for non-IRAP (or as another alternative non-CLVSS) pictures of layer i. `max_tid_il_ref_pics_plus1[i][j]` greater than 0 indicates that, for decoding pictures of layer i, no picture with a TemporalId greater than `max_tid_il_ref_pics_plus1[i] - 1` is used as ILRP. When it does not exist, the value of max_tid_il_ref_pics_plus1[i][j] is inferred to be equal to 7.

[0145] Therefore, the video encoder 200 and the video decoder 300 can be configured to encode or decode the values ​​of any or all of the syntax elements in Tables 8 and / or 9, and any or both of Tables 8 and / or 9 can be combined with the syntax elements in Table 7, as discussed above.

[0146] The video encoder 200 and video decoder 300 can also be configured to encode and decode data representing an Output Layer Set (OLS). The VVC specifies that the 0th OLS contains only the lowest layer, and for the 0th OLS, the only layer contained is the output. Multiple non-basic independent layers may exist in the bitstream. This disclosure recognizes that it may be beneficial to allocate independent non-basic layers to the OLS, in addition to the 0th OLS.

[0147] The video decoder 300 can derive the number of independent layers (NumIndependentLayers) from the VPS layer correlation signaling, as shown below (e.g., according to the following algorithm pseudocode):

[0148]

[0149] The first set of NumIndependentLayers output layers contains independent layers, and in each OLS, the only layer included is the output layer.

[0150] The variable TotalNumOlss, which specifies the total number of OLS values ​​assigned by the VPS, can be derived as follows:

[0151] In some examples, the 0th OLS can contain all the independent layers.

[0152] The video encoder 200 and video decoder 300 can derive the variable NumLayersInOls[i], which specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[i][j], which specifies the nuh_layer_id value of the j-th layer in the i-th OLS, as shown below:

[0153]

[0154] In other words, the video encoder 200 and the video decoder 300 can determine the total number of OLS as follows: when the maximum number of layers of the VPS minus 1 equals zero, the total number of OLS equals 1; when at least one of the following: 1) each layer of the VPS is an OLS, 2) the OLS mode indicator value equals 0 or 3) the OLS mode indicator value equals 1, the total number of OLS equals the maximum number of layers of the VPS; or when the OLS mode indicator value equals 2, the total number of OLS equals the number of independent layers plus the value of the syntax element of the VPS indicating the number of OLS.

[0155] The video encoder 200 and video decoder 300 can be configured based on the following conditions regarding the profile / tier / level (PTL) and decoded picture buffer (DPB) index signaling. In this example, ols_ptl_idx[i] specifies the index of the list of profile_tier_level() syntax structures in the VPS for the i-th OLS. When the number of profile_tier_level() syntax structures equals TotalNumOlss, ols_ptl_idx[i] can be derived accordingly without explicit signaling notification in the VPS. In this example, ols_dpb_params_idx[i] specifies the index of the list of dpb_parameters() syntax structures in the VPS for the i-th OLS. When the number of dpb_parameters() syntax structures equals TotalNumOlss, ols_dpb_params_idx[i] can be derived accordingly without explicit signaling notification in the VPS.

[0156] Table 10 shows example conditions for signaling notifications of ols_ptl_idx[i] and ols_dpb_params_idx[i]. Typically, when the number of profile_tier_level() syntax structures in the VPS equals the number of OLSs (i.e., one profile_tier_level() syntax structure corresponds to each OLS), it is not necessary to send the index ols_ptl_idx to indicate which profile_tier_level() syntax structure to use. In this case, in some examples, the video decoder 300 can infer that the index ols_ptl_idx is equal to the OLS index.

[0157] Table 10—Example VPS Conditions Including PTL and DPB Index Signaling Notifications

[0158]

[0159]

[0160] Compared to existing VVC proposals, the semantics of the syntax elements affected by this example condition can be modified as follows:

[0161] `ols_ptl_idx[i]` specifies the index of the list of `profile_tier_level()` syntax structures in the VPS for the `profile_tier_level()` syntax structure applied to the `i`th OLS. When present, the value of `ols_ptl_idx[i]` should be in the range of 0 to `vps_num_ptls_minus1` (inclusive). [added: "When `ols_ptl_idx[i]` does not exist, the value of `ols_ptl_idx[i]` is inferred as follows: if the value of `vps_num_ptls_minus1+1` equals `TotalNumOlss`, then the value of `ols_ptl_idx[i]` is inferred to be equal to `i`; otherwise, "] when `vps_num_ptls_minus1` equals 0, then the value of `ols_ptl_idx[i]` is inferred to be equal to 0.]

[0162] When NumLayersInOls[i] equals 1, the profile_tier_level() syntax structure applied to the i-th OLS also exists in the SPS referenced by the layer in the i-th OLS. The requirement for bitstream consistency is that when NumLayersInOls[i] equals 1, the profile_tier_level() syntax structure for the i-th OLS used for signaling notifications in the VPS and SPS should be exactly the same.

[0163] `ols_dpb_params_idx[i]` specifies the index of the list of `dpb_parameters()` syntax structures in the VPS when `NumLayersInOls[i]` is greater than 1. If present, the value of `ols_dpb_params_idx[i]` should be in the range of 0 to `vps_num_dpb_params-1` (inclusive). [added: "When ols_dpb_params_idx[i] does not exist, the value of ols_dpb_params_idx[i] is inferred as follows: if the value of vps_num_dpb_params is equal to TotalNumOlss, then the value of ols_dpb_params_idx[i] is inferred to be equal to i. Otherwise, "]When ols_dpb_params_idx[i] does not exist, then the value of ols_dpb_params_idx[i] is inferred to be equal to 0.]

[0164] When NumLayersInOls[i] equals 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS of the layer reference in the i-th OLS. [added: "The requirement for bitstream consistency is that when NumLayersInOls[i] equals 1, the dpb_parameters() syntax structure for signaling notifications in the VPS and SPS of the i-th OLS should be exactly the same."]

[0165] In this manner, the video encoder 200 and video decoder 300 can encode and decode the values ​​of syntax elements according to the example syntax and semantics in Table 10 above, which can be combined with any or all of the example syntax and semantics in Tables 7-9 above. Furthermore, it should be understood that the DCI SEI messaging technology in Table 6 can be used alone or in combination with the example VPS in Tables 7-10 above.

[0166] In some instances, based on the examples in Table 10 and the discussion above, the video encoder 200 can be configured to determine that the number of PTL data structures in the VPS is equal to the total number of OLS specified for the VPS. In this scenario, the syntax element representing the number of PTL data structures in the VPS will be equal to the total number of OLS specified for the VPS. In response, the video encoder 200 can avoid encoding the values ​​of the OLS PTL index values ​​in the VPS. The video encoder 200 can also ensure that the video data of the OLS conforms to the grade, level, and rank values ​​of the PTL data structures with indexes that match the OLS in sequence. That is, the video encoder 200 can ensure that the video data of the i-th OLS is encoded according to all i values ​​from 0 to the number of OLS by the encoding / decoding tool of the i-th PTL data structure.

[0167] Similarly, the video decoder 300 can determine that the value of the syntax element indicating the number of PTL data structures for the VPS (e.g., vps_num_ptls_minus1) is equal to the total number of OLSs specified for the VPS. In response, the video decoder 300 can infer the value of the index that specifies the correspondence between OLSs and PTL data structures (e.g., ols_ptl_idx[i]). Specifically, as discussed above, the video decoder 300 can infer the value of the index corresponding to the order in which the index will appear, i.e., ols_ptl_idx[i] = i. In this way, the video decoder 300 can determine that when decoding the video data of the i-th OLS, the i-th PTL data structure will be used for all values ​​of i between 0 and the total number of OLSs.

[0168] Additionally or alternatively, the video encoder 200 and video decoder 300 can be configured to prevent SPS emulation in certain situations. Table 11 below shows a partial set of SPS syntax elements. Depending on the value of pic_width_max_in_luma_samples, the values ​​of the syntax elements from sps_seq_parameter_set_id to res_change_in_clvs_allowed_flag in Table 11 may all be equal to 0, which could result in an emulation start code.

[0169] Table 11—SPS RBSP Syntax

[0170]

[0171]

[0172] VVC specifies the start code prefix as the hexadecimal value 0x000001. VVC indicates that the start code prefix is ​​the prefix of the NAL unit, thus signaling that a new NAL unit is about to appear. Therefore, according to the existing structure of VVC as shown in Table 11, if each syntax element from sps_seq_parameter_set_id to res_change_in_clvs_allowed_flag has a value of 0 and the subsequent bits are 1, the start code prefix will be emulated, which may cause errors in the processing of the video decoder 300.

[0173] The video encoder 200 and the video decoder 300 can be configured to prevent such emulation, individually or in any combination, according to any one or all of the following example techniques:

[0174] It is required that all syntax elements that contribute to the simulation of the start code (i.e., those currently having a value of 0) cannot all be equal to 0. In this case, at least one syntax element must have a non-zero value, which prevents start code simulation from occurring. A similar approach can be applied to any parameter set or header.

[0175] Change sps_max_sublayers_minus1 to sps_max_sublayers, so that the value of sps_max_sublayers should be in the range of 1 to vps_max_sublayers_minus1+1.

[0176] Change sps_reserved_zero_4bits to sps_reserved_one_4bits, and the value of sps_reserved_one_4bits should be equal to '1111' (0xF).

[0177] In this way, the semantics of these SPS syntax elements can be defined such that the video encoder 200 and video decoder 300 do not need to encode or decode emulation prevention bytes in SPS. By avoiding the encoding / decoding of emulation prevention bytes, the video encoder 200 and video decoder 300 can avoid additional processing operations, and the video bitstream does not need to include this overhead data.

[0178] In some examples, the video encoder 200 and video decoder 300 can be configured according to the bitstream consistency requirements of the PPS consistency window syntax elements, for example, as described below. Modifications to the VVC specification are indicated by the following [added: "added text"].

[0179] When pic_width_in_luma_samples equals pic_width_max_in_luma_samples and pic_height_in_luma_samples equals pic_height_max_in_luma_samples, [added: "pps_conformance_window_flag equal to 0 is a requirement for bitstream consistency."]

[0180] `pps_conf_win_left_offset`, `pps_conf_win_right_offset`, `pps_conf_win_top_offset`, and `pps_conf_win_bottom_offset` specify the sample points of the image in the CLVS output from the decoding process, according to the rectangular region specified in the coordinates used for the output image. [added: "When `pps_conformance_window_flag` equals 0, the values ​​of `pps_conf_win_left_offset`, `pps_conf_win_right_offset`, `pps_conf_win_top_offset`, and `pps_conf_win_bottom_offset` are inferred to be equal to `sps_conf_win_left_offset`, `sps_conf_win_right_offset`, `sps_conf_win_top_offset`, and `sps_conf_win_bottom_offset`, respectively."]

[0181] Accordingly, the video encoder 200 and video decoder 300 may be configured to, in response to determining that the value of the syntax element representing the picture width (e.g., pic_width_in_luma_samples) in the picture parameter set (PPS) of the bitstream is the maximum picture width value (e.g., pic_width_max_in_luma_samples), and the value of the syntax element representing the picture height (e.g., pic_height_in_luma_samples) in the PPS is the maximum picture height value (e.g., pic_height_max_in_luma_samples), determine that the consistency window value is equal to zero (i.e., the consistency window syntax element of the PPS is not explicitly signaled). Additionally or alternatively, the video encoder 200 and video decoder 300 may be configured to, in response to determining that the consistency window value in the PPS is equal to zero (i.e., the consistency window syntax element of the PPS is not explicitly signaled), infer the value of the consistency window offset of the PPS as a corresponding value equal to the consistency window offset of the SPS.

[0182] Additionally, when pic_width_in_luma_samples and pic_height_in_luma_samples are equal to pic_width_max_in_luma_samples and pic_height_max_in_luma_samples, respectively, they can be signaled to be equal to their default values, such as 0. Then, if pic_width_in_luma_samples or pic_height_in_luma_samples is equal to the default value, the value of pic_width_in_luma_samples or pic_height_in_luma_samples can be replaced with the values ​​of pic_width_max_in_luma_samples and pic_height_max_in_luma_samples, respectively.

[0183] In an alternative solution, video encoder 200 and video decoder 300 can encode and decode the value of the gating flag pic_size_present_flag. Video encoder 200 and video decoder 300 can only encode and decode pic_width_in_luma_samples and pic_height_in_luma_samples if the value of pic_size_present_flag is equal to 1. Inference rules can then be added, whereby when pic_width_in_luma_samples and pic_height_in_luma_samples are not present, video encoder 200 and video decoder 300 can infer that these values ​​are equal to pic_width_max_in_luma_samples and pic_height_max_in_luma_samples, respectively. When pic_width_in_luma_samples and pic_height_in_luma_samples are equal to pic_width_max_in_luma_samples and pic_height_max_in_luma_samples respectively (which may be a typical encoding / decoding case), this method can save the signaling overhead used for them.

[0184] In another example, video encoder 200 and video decoder 300 can encode and decode the values ​​of PPS syntax elements (e.g., pps_res_change_allowed_flag as shown in the example in Table 12 below) to indicate whether the image resolution of the reference PPS image is changeable. Syntax elements (e.g., flags) can be used to adjust the presence of PPS image resolution, PPS consistency window, and PPS scaling window syntax elements. When they are not present, video decoder 300 can infer the values ​​of these syntax elements from the corresponding SPS syntax elements.

[0185] Table 12 - Example PPS Syntax Structure

[0186]

[0187]

[0188] The semantics of the pps_res_change_allowed_flag syntax element in the examples in Table 12 above can be as follows:

[0189] A value of 1 for `pps_res_change_allowed_flag` indicates that the spatial resolution of the image can be changed on the reference PPS image. A value of 0 for `pps_res_change_allowed_flag` indicates that the spatial resolution of the image remains unchanged on the reference PPS image. When the value of `res_change_in_clvs_allowed_flag` is 0, the value of `pps_res_change_allowed_flag` should also be 0.

[0190] The semantics of the other syntactic elements in Table 12 can be modified as follows:

[0191] A value of 1 for pps_conformance_window_flag indicates that the consistency trimming window offset parameter in PPS follows the next one. A value of 0 for pps_conformance_window_flag indicates that the consistency trimming window offset parameter does not exist in PPS. [added: "When it does not exist, the value of pps_conformance_window_flag is inferred to be equal to 0."]

[0192] A scaling_window_explicit_signalling_flag value of 1 indicates that the scaling window offset parameter exists in PPS. A scaling_window_explicit_signalling_flag value of 0 indicates that the scaling window offset parameter does not exist in PPS. [added: "When it does not exist, the value of scaling_window_explicit_signalling_flag is inferred to be 0."]

[0193] According to VVC, the value of `sps_poc_msb_flag` indicates whether the value of the `ph_poc_msb_present_flag` syntax element exists in the picture header (PH) of the reference SPS. When `sps_poc_msb_flag` equals 1, VVC instructs signaling in the SPS to another syntax element, `poc_msb_len_minus1`, to indicate the length of the `poc_msb_val` syntax element in the PH of the reference SPS. Combining two syntax elements into one simplifies SPS syntax design. Table 13 below shows example syntax element changes. In this example, the syntax element `poc_msb_len` replaces `sps_poc_msg_flag` and `poc_msb_len_minus1`. In Table 13, `[removed: "removed text"]` means text removed from VVC, while `[added: "added text"]` means text added to VVC.

[0194] Table 13—Examples of POC MSB Syntax Elements in SPS

[0195]

[0196] The semantics of the `poc_msb_len` syntax element can be defined as follows: When present in the PH of the reference SPS, `poc_msb_len` specifies the length of the `poc_msb_val` syntax element in bits. The value of `poc_msb_len` should be in the range of 0 to 32 - log2_max_pic_order_cnt_lsb_minus4-4 (inclusive). When `poc_msb_len` equals 0, the `ph_poc_msb_present_flag` and `poc_msb_val` syntax elements do not exist in the PH of the reference SPS.

[0197] Table 14 below shows example changes to the image header (PH) corresponding to the example changes in Table 13:

[0198] Table 14—PH POC MSB Syntax

[0199]

[0200]

[0201] The semantics of poc_msb_val in Table 14 can be defined as follows, updated relative to VVC: poc_msb_val specifies the POC MSB value of the current image. [added: "The length of the syntax element poc_msb_val is poc_msb_len bits."]

[0202] Similarly, the signaling notifications of the alf_luma_coeff_abs and alf_luma_coeff_sign syntax elements in the Adaptive Parameter Set (APS) can be converted into a single syntax element alf_luma_coeff signaled in binary code, having a sign, such as a signed exponential Golomb code, e.g., se(v). For example, the syntax element alf_luma_coeff[sfIdx][j] with the descriptor se(v) can replace alf_luma_coeff_abs[sfidx][j] and alf_luma_coeff_sign[sfIdx][j].

[0203] Accordingly, the video encoder 200 and the video decoder 300 can encode and decode the values ​​of the syntax elements in Tables 12 and 13 according to the techniques discussed above. The video encoder 200 and the video decoder 300 can perform these techniques individually or in any combination with the techniques discussed in Tables 6-11 above.

[0204] This disclosure recognizes that problems may occur during the derivation of slice heights within a slice when num_exp_slices_in_tile[i] is equal to 0 or 1, and when the value of the variable NumSlicesInTile[i] is deduced to be equal to 1. The video encoder 200 and video decoder 300 can be configured according to the example changes to VVC discussed below, where [removed: "removed text"] means removed text, and [added: "added text"] means added text.

[0205] `num_exp_slices_in_tile[i]` specifies the slice height explicitly provided in the current slice containing more than one rectangular slice. The value of `num_exp_slices_in_tile[i]` should be in the range [removed: "0" 0][added: "1"] to `RowHeight[tileY][removed: "-1"]` (inclusive), where `tileY` is the slice row index containing the i-th slice. When it does not exist, the value of `num_exp_slices_in_tile[i]` is inferred to be equal to 0. When `num_exp_slices_in_tile[i]` is equal to 0, the value of the variable `NumSlicesInTile[i]` is inferred to be equal to [removed: "1"][added: "0"].

[0206] `exp_slice_height_in_ctus_minus1[j]` incremented by 1 specifies the height of the j-th rectangular slice in the current slice, in units of CTU rows. The value of `exp_slice_height_in_ctus_minus1[j]` should be in the range of 0 to RowHeight[tileY]-1 (inclusive), where tileY is the slice row index of the current slice.

[0207] When num_exp_slices_in_tile[i] is greater than 0, the variables NumSlicesInTile[i][added: "for i in the range of 1 to NumSlicesInTile[i]"] and SliceHeightInCtusMinus1[i+k] (where k is in the range of 0 to NumSlicesInTile[i]-1) are derived as follows:

[0208]

[0209] Figure 2A and Figure 2B This is a conceptual diagram illustrating an example Quadtree Binary Tree (QTBT) structure 130 and its corresponding Code-to-Code-Unit (CTU) 132. Solid lines represent quadtree partitions, and dashed lines indicate binary tree partitions. In each partition node (i.e., a non-leaf node) of the binary tree, a flag is signaled to indicate which partition type (i.e., horizontal or vertical) is used, where in this example, 0 indicates a horizontal partition and 1 indicates a vertical partition. For quadtree partitions, it is not necessary to indicate the partition type because the quadtree node divides the block horizontally and vertically into four equal-sized sub-blocks. Accordingly, the video encoder 200 can encode syntax elements (such as partition information) for the region tree layer (i.e., solid lines) of the QTBT structure 130 and syntax elements (such as partition information) for the prediction tree layer (i.e., dashed lines) of the QTBT structure 130, and the video decoder 300 can decode these syntax elements. The video encoder 200 can encode video data (such as prediction and transform data) of the CU represented by the terminal leaf nodes of the QTBT structure 130, and the video decoder 300 can decode the video data.

[0210] generally, Figure 2B The CTU 132 can be associated with parameters that define the size of the block corresponding to the nodes of the QTBT structure 130 in the first and second layers. These parameters may include the CTU size (the size of the CTU 132 in sample points), the minimum quadtree size (MinQTSize, representing the minimum allowed size of the leaf nodes of the quadtree), the maximum binary tree size (MaxBTSize, representing the maximum allowed size of the root node of the binary tree), the maximum binary tree depth (MaxBTDepth, representing the maximum allowed depth of the binary tree), and the minimum binary tree size (MinBTSize, representing the minimum allowed size of the leaf nodes of the binary tree).

[0211] The root node of the QTBT structure corresponding to the CTU can have four child nodes in the first level of the QTBT structure, each of which can be partitioned according to a quadtree partition. That is, nodes in the first level are either leaf nodes (without child nodes) or have four child nodes. An example of QTBT structure 130 represents such nodes as including a parent node and child nodes with solid lines for branching. Nodes in the first level can also be partitioned by the corresponding binary tree if they are not larger than the maximum allowed binary tree root node size (MaxBTSize). The binary tree partitioning of a node can be iterated until the partitioned nodes reach the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). An example of QTBT structure 130 represents such nodes with dashed lines for branching. Binary tree leaf nodes are called codec units (CUs), which are used for prediction (e.g., intra-frame image prediction or inter-frame image prediction) and transformation without any further partitioning. As discussed above, CUs can also be referred to as “video chunks” or “blocks”.

[0212] In one example of a QTBT partitioning structure, the CTU size is set to 128×128 (luminance samples and two corresponding 64×64 chrominance samples), MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. Quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If a quadtree leaf node is 128×128, it will not be further partitioned into a binary tree because its size exceeds MaxBTSize (i.e., 64×64 in this example). Otherwise, the quadtree leaf node will be further partitioned into a binary tree. Therefore, the quadtree leaf node is also used as the root node of the binary tree, and the binary tree depth is 0. When the depth of a binary tree reaches MaxBTDepth (4 in this example), no further partitioning is permitted. When the width of a binary tree node equals MinBTSize (4 in this example), it means that no further vertical partitioning (width partitioning) is allowed for that node. Similarly, a binary tree node with a height equal to MinBTSize means that no further horizontal partitioning (i.e., height partitioning) is allowed for that node. As mentioned above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further partitioning.

[0213] Figure 3This is a block diagram illustrating an example video encoder 200 capable of performing the techniques of this disclosure. Figure 3 This disclosure is for illustrative purposes and should not be construed as limiting the techniques extensively illustrated and described herein. For illustrative purposes, this disclosure describes a video encoder 200 within the context of video codec standards, such as the HEVC video codec standard and the H.266 video codec standard under development. However, the techniques disclosed herein are not limited to these video codec standards and are generally applicable to other video encoding and decoding standards.

[0214] exist Figure 3 In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy coding unit 220 can be implemented in one or more processors or processing circuits. Additionally, the video encoder 200 may include additional or alternative processors or processing circuits to perform these and other functions.

[0215] The video data storage 230 can store video data to be encoded by components of the video encoder 200. The video encoder 200 can, for example, obtain data from video source 104 (…). Figure 1 The video encoder 200 receives video data stored in video data memory 230. DPB 218 can act as a reference picture memory, storing reference video data for use by the video encoder 200 when predicting subsequent video data. Video data memory 230 and DPB 218 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive random access RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, video data memory 230 can be on-chip with other components of the video encoder 200, as illustrated, or off-chip relative to those components.

[0216] In this disclosure, references to video data memory 230 should not be construed as being limited to memory within video encoder 200 unless specifically described therein, or to memory external to video encoder 200 unless specifically described therein. Rather, references to video data memory 230 should be understood as reference memory storing video data (e.g., video data of the current block to be encoded) received by video encoder 200 for encoding. Figure 1 The memory 106 can also provide temporary storage for the outputs from various units of the video encoder 200.

[0217] The diagram Figure 3 The various units are used to help understand the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit is a circuit that provides a specific function and is pre-programmed to perform certain operations. A programmable circuit is a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations it can perform. For example, a programmable circuit can run software or firmware that causes it to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit can run software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more of the units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more units can be integrated circuits.

[0218] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core, all formed by programmable circuitry. In an example where software running on programmable circuitry is used to perform the operation of the video encoder 200, memory 106 ( Figure 1 The video encoder 200 may store the target code of the software received and run by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.

[0219] The video data storage unit 230 is configured to store the received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be the raw video data to be encoded.

[0220] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units to perform video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-frame block copying unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0221] The mode selection unit 202 typically coordinates multiple coding processes to test combinations of coding parameters and the resulting rate-distortion values ​​for such combinations. Coding parameters may include the CTU-CU split, the prediction mode for the CU, the transformation type of the residual data for the CU, the quantization parameters of the residual data for the CU, etc. The mode selection unit 202 can ultimately select a combination of coding parameters that has a better rate-distortion value than other test combinations.

[0222] In some examples, mode selection unit 202 can be configured to automatically determine which codec tools are enabled and / or disabled for one or more OLSs. For example, mode selection unit 202 can perform rate-distortion optimization (RDO) to calculate the rate-distortion (RD) value for each OLS and various combinations of enabled / disabled codec tools, and then select the codec tool set that produces the optimal RD value for the OLS. Alternatively, an administrator or other user can enable and / or disable the codec tools for a given OLS (i.e., any one or all OLSs). In any case, mode selection unit 202 can determine the appropriate grade, level, and rank value for each OLS in the total number of OLSs in the video bitstream.

[0223] According to the technology disclosed herein, the mode selection unit 202 can also signal the PTL data structure of the OLS. The mode selection unit 202 can also determine the number of PTL data structures to be signaled for the OLS and signal that number in the VPS. When the number of PTL data structures in the VPS is equal to the total number of OLS specified for the VPS, the mode selection unit 202 can cause the video encoder 200 to avoid encoding the values ​​of the OLS PTL index values. Conversely, the mode selection unit 202 can cause the entropy encoding unit 220 to encode the PTL data structures in the VPS in the same order as the OLS, so that the correspondence between the PTL data structures and the OLS can be inferred from the signaling order.

[0224] Additionally or alternatively, the mode selection unit 202 may use RDO technology to determine the image size of the images in the image sequence. Alternatively, the mode selection unit 202 may receive configuration data from a user, such as an administrator, to indicate the image size to be used. In an example where the mode selection unit 202 determines that an image has the maximum image size, the mode selection unit 202 may cause the entropy encoding unit 220 to generate an image parameter set (PPS) for images that do not contain consistency window syntax elements. In some examples, the mode selection unit 202 may cause the entropy encoding unit 220 to encode the PPS to include a consistency window flag, which has a value indicating that other consistency window syntax elements (e.g., offset values) have not been signaled. In this case, the values ​​of the consistency window syntax elements of the PPS can be inferred from the corresponding values ​​in the sequence parameter set (SPS).

[0225] The video encoder 200 can segment an image retrieved from the video data storage 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can segment the CTUs of the image according to a tree structure (such as the QTBT structure or quadtree structure of HEVC described above). As described above, the video encoder 200 can form one or more CUs by segmenting CTUs according to a tree structure. This CU can also generally be referred to as a "video block" or "block".

[0226] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate predicted blocks for the current block (e.g., the overlapping portion of PU and TU in the current CU or HEVC). To perform inter-frame prediction for the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously encoded / decoded pictures stored in DPB 218). Specifically, motion estimation unit 222 may, for example, calculate values ​​representing how similar a potential reference block is to the current block based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. Motion estimation unit 222 may typically perform these calculations using the sample-by-sample difference between the current block and the reference block being considered. Motion estimation unit 222 may identify the reference block with the lowest value obtained from these calculations, indicating the reference block that most closely matches the current block.

[0227] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the position of a reference block in a reference image relative to the position of a current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data for reference blocks. As another example, if the motion vectors have fractional sample precision, motion compensation unit 224 can interpolate the values ​​of the prediction blocks according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by the corresponding motion vectors and combine the retrieved data, for example, by sample-by-sample averaging or weighted averaging.

[0228] As another example, for intra-prediction or intra-prediction codec, intra-prediction unit 226 can generate a prediction block from samples adjacent to the current block. For example, in directional mode, intra-prediction unit 226 can typically mathematically combine the values ​​of adjacent samples and fill these calculated values ​​in a defined direction across the current block to produce a prediction block. As another example, in DC mode, intra-prediction unit 226 can calculate the average of the adjacent samples of the current block and generate a prediction block to include that average value obtained for each sample of the prediction block.

[0229] Mode selection unit 202 provides the prediction block to residual generation unit 204. Residual generation unit 204 receives the raw, unencoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block of the current block. In some examples, residual generation unit 204 can also determine the differences between sample values ​​in the residual block to generate the residual block using Residual Differential Pulse Codec Modulation (RDPCM). In some examples, residual generation unit 204 can be formed using one or more subtractor circuits performing binary subtraction.

[0230] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As mentioned above, the size of a CU can refer to the size of its luma codec block, while the size of a PU can refer to the size of its luma prediction unit. Assuming a specific CU has a size of 2N×2N, video encoder 200 can support PU sizes of 2N×2N or N×N for intra-frame prediction, and symmetrical PU sizes of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric segmentation of PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.

[0231] In the example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luma codec block and a corresponding chroma codec block. Similarly, the size of the CU can refer to the size of the luma codec block of the CU. The video encoder 200 and video decoder 300 can support CU sizes of 2N×2N, 2N×N, or N×2N.

[0232] For other video codec techniques, such as intra-block copy mode codec, affine mode codec, and linear model (LM) mode codec, as some examples, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the codec technique. In some examples, such as palette mode codec, mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block should be reconstructed based on the selected palette. In such a mode, mode selection unit 202 can provide these syntax elements to entropy coding unit 220 for encoding.

[0233] As described above, the residual generation unit 204 receives video data of the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0234] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply discrete cosine transform (DCT), direction transform, Karhunen-Loeve transform (KLT), or conceptually similar transformations to the residual block. In some examples, transform processing unit 206 may perform multiple transformations on the residual block, such as primary and secondary transformations, such as rotation transformations. In some examples, transform processing unit 206 does not apply any transformations to the residual block.

[0235] Quantization unit 208 can quantize the transform coefficients in a transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 can quantize the transform coefficients of the transform coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and therefore, the quantized transform coefficients may have lower precision than the original transform coefficients generated by transform processing unit 206.

[0236] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform to the quantized transform coefficient block, respectively, to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although it may have some degree of distortion). For example, the reconstruction unit 214 can add samples of the reconstructed residual block to corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0237] Filter unit 216 can perform one or more filtering operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce block artifacts along the edges of the CU. The operation of filter unit 216 can be skipped in some examples.

[0238] The video encoder 200 stores reconstructed blocks in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 can store the reconstructed blocks in the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 can store filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve a reference image from the DPB 218, which is formed from reconstructed (and possibly filtered) blocks, to perform inter-frame prediction of blocks in subsequently encoded images. Furthermore, the intra-frame prediction unit 226 can use the reconstructed blocks in the DPB 218 of the current image to perform intra-frame prediction of other blocks in the current image.

[0239] Typically, entropy coding unit 220 can entropy-encode syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy-encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy-encode predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, another example of video data, to generate entropy-encoded data. For example, entropy coding unit 220 can perform context-adaptive variable-length codec (CAVLC), CABAC, variable-to-variable (V2V) length codec, syntax-based context-adaptive binary arithmetic codec (SBAC), probabilistic interval partitioned entropy (PIPE) codec, exponential-Golomb coding, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in bypass mode, in which syntax elements are not entropy-encoded.

[0240] The video encoder 200 can output a bitstream containing entropy-encoded syntax elements required to reconstruct slices or blocks of images. Specifically, the entropy coding unit 220 can output a bitstream.

[0241] The above operations are described relative to blocks. This description should be understood as operations applied to luma and / or chroma codec blocks. As mentioned above, in some examples, the luma and chroma codec blocks are the luma and chroma components of the CU. In some examples, the luma and chroma codec blocks are the luma and chroma components of the PU.

[0242] In some examples, for the chroma codec block, it is not necessary to repeat the operations performed relative to the luma codec block. As an example, the operations for identifying the motion vector (MV) and reference image for the luma codec block do not need to be repeated for identifying the MV and reference image for the chroma block. Instead, the MV for the luma codec block can be scaled to determine the MV for the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma codec blocks.

[0243] Video encoder 200 represents an example of a video codec that can be configured to perform the techniques described above with respect to any of Tables 7 to 13.

[0244] Figure 4 This is a block diagram illustrating an example video decoder 300 capable of performing the techniques of this disclosure. Provided Figure 4 This disclosure is for illustrative purposes and not for limiting the extensive examples and techniques described herein. For illustrative purposes, this disclosure describes a video decoder 300 based on VCC and HEVC techniques. The techniques of this disclosure can be implemented by video codec devices configured for other video codec standards.

[0245] exist Figure 4 In the example, the video decoder 300 includes a codec picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any one or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or processing circuits. Additionally, the video decoder 300 may include additional or alternative processors or processing circuits to perform these and other functions.

[0246] The video decoder 300 can initially process (e.g., decode, parse, and / or interpret) high-level syntax data structures, such as video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS). For example, the video decoder 300 can use data from the VPS to determine the grade, level, and grade (PTL) values ​​of the output layer set (OLS) of the bitstream. According to the techniques of this disclosure, the video decoder 300 can, for example, use the value of vps_num_ptls_minus1 to determine the number of PTL data structures in the VPS. The video decoder 300 can also determine the total number of OLS specified for the VPS, for example, using the pseudocode described above for calculating the value of TotalNumOlss.

[0247] The video decoder 300 can then determine whether the number of PTL data structures is equal to the total number of OLS specified for the VPS. If the number of PTL data structures is equal to the total number of OLS specified for the VPS, the video decoder 300 can determine that the OLSPTL index value will not be explicitly signaled in the VPS. Instead, the video decoder 300 can infer the value of the OLS PTL index value. For example, the video decoder 300 can determine that the i-th PTL data structure describes the PTL data of the i-th OLS for all i values ​​between 0 and the total number of OLS specified for the VPS. On the other hand, if the number of PTL data structures is not equal to the total number of OLS, the video decoder 300 can decode the explicit value of the OLS PTL index value from the VPS.

[0248] Additionally or alternatively, the video decoder 300 may determine whether the values ​​representing the image size of the PPS (e.g., pic_width_in_luma_samples and pic_height_in_luma_samples) indicate that the image has a maximum size (e.g., pic_width_in_luma_samples equals pic_width_max_in_luma_samples and pic_height_in_luma_samples equals pic_height_max_in_luma_samples). When the image has a maximum size, the video decoder 300 may determine that the PPS conformance window syntax element has not been signaled. When the PPS conformance window syntax element is not signaled (e.g., when PPS_conformance_window_flag has an explicit or inferred value of 0), the video decoder 300 can infer the values ​​of other PPS conformance window syntax elements (e.g., pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset) from the corresponding SPS conformance window syntax elements (e.g., sps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset).

[0249] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-prediction unit 318. The prediction processing unit 304 may include additional units to perform predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copying unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0250] CPB memory 320 can store video data to be decoded by components of video decoder 300, such as encoded video bitstreams. For example, it can be stored from computer-readable medium 110 ( Figure 1 The video data stored in the CPB memory 320 is obtained. The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Furthermore, the CPB memory 320 may store video data other than the syntax elements of the encoded / decoded pictures, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores decoded pictures that the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures from the encoded video bitstream. The CPB memory 320 and DPB 314 may be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.

[0251] Additionally or alternatively, in some examples, the video decoder 300 can be drawn from the memory 120 ( Figure 1 Retrieving encoded and decoded video data. That is, memory 120 can store data, as discussed above regarding CPB memory 320. Similarly, when some or all of the functions of video decoder 300 are implemented in software that will be run by the processing circuitry of video decoder 300, memory 120 can store instructions that will be run by video decoder 300.

[0252] Illustration Figure 4 The various units shown aid in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to... Figure 3Fixed-function circuits are circuits that provide a specific function and are pre-programmed to perform certain operations. Programmable circuits are circuits that can be programmed to perform various tasks and provide flexible functionality within the operations they can perform. For example, a programmable circuit can run software or firmware that causes it to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can run software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.

[0253] The video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video decoder 300 is performed by software running on the programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0254] Entropy decoding unit 302 can receive encoded video data from CPB and perform entropy decoding on the video data to reconstruct syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on syntax elements extracted from the bitstream.

[0255] Typically, the video decoder 300 reconstructs the image on a block-by-block basis. The video decoder 300 can perform the reconstruction operation separately for each block (where the block currently being reconstructed, i.e. the decoded block, can be referred to as the "current block").

[0256] Entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients defining the quantized transform coefficient block, as well as transform information (such as quantization parameters (QP) and / or (one or more) transform mode indications). Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization to be applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bit-by-bit left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 can thus form a transform coefficient block including the transform coefficients.

[0257] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply the inverse DCT, inverse integer transform, inverse Karhunen-Loeve transform (KLT), inverse rotation transform, inverse direction transform, or another inverse transform to the transform coefficient block.

[0258] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-frame predicted, the motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements may indicate a reference image in the DPB 314 from which the reference block is retrieved, and a motion vector identifying the position of the reference block in the reference image relative to the position of the current block in the current image. The motion compensation unit 316 can typically be configured with respect to the motion compensation unit 224 ( Figure 3 The method described is essentially the same as the method used to perform the inter-frame prediction process.

[0259] As another example, if the prediction information syntax element indicates that the current block is intra-predictable, then intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, intra-prediction unit 318 can typically be configured with respect to intra-prediction unit 226 ( Figure 3 The intra-prediction process is performed in a manner substantially similar to that described above. The intra-prediction unit 318 can retrieve data of neighboring samples of the current block from the DPB 314.

[0260] Reconstruction unit 310 can use prediction blocks and residual blocks to reconstruct the current block. For example, reconstruction unit 310 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block.

[0261] Filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce block artifacts along the edges of the reconstructed block. The operation of filter unit 312 is not necessarily performed in all examples.

[0262] The video decoder 300 can store reconstructed blocks in the DPB 314. As described above, the DPB 314 can provide reference information to the prediction processing unit 304, such as samples of the current image for intra-frame prediction and previously decoded images for subsequent motion compensation. Furthermore, the video decoder 300 can output decoded images from the DPB 314 for subsequent rendering in applications such as... Figure 1 The display device 118 is on the display device.

[0263] Video decoder 300 represents an example of a video codec that can be configured to perform the techniques described above with respect to any of Tables 7 to 13.

[0264] Figure 5 This is a flowchart illustrating an example method for encoding a current block according to the techniques of this disclosure. The current block may include the current CU. Although regarding video encoder 200 ( Figure 1 and 3 This description is provided, but it should be understood that other devices can be configured to perform similar actions. Figure 5 The method.

[0265] In this example, the video encoder 200 initially predicts the current block (350). For example, the video encoder 200 may form a prediction block for the current block. The video encoder 200 may then compute the residual block for the current block (352). To compute the residual block, the video encoder 200 may compute the difference between the original, unencoded block and the prediction block for the current block. The video encoder 200 may then transform and quantize the coefficients of the residual block (354). Next, the video encoder 200 may scan the transform coefficients of the quantized residual block (356). During or after the scan, the video encoder 200 may entropy encode the coefficients (358). For example, the video encoder 200 may encode the coefficients using CAVLC or CABAC. The video encoder 200 may then output the entropy-encoded data of the block (360).

[0266] Figure 6 This is a flowchart illustrating an example method for decoding the current block according to the techniques of this disclosure. The current block may include the current CU. Although regarding the video decoder 300 ( Figure 1 and 4 This description is provided, but it should be understood that other devices can be configured to perform similar actions. Figure 6 The method.

[0267] The video decoder 300 can receive entropy-coded data of the current block, such as entropy-coded prediction information and entropy-coded data of the coefficients of the residual block corresponding to the current block (370). The video decoder 300 can entropy decode the entropy-coded data to determine the prediction information of the current block and reproduce the coefficients of the residual block (372). The video decoder 300 can predict the current block (374), for example, using an intra-frame or inter-frame prediction mode indicated by the prediction information of the current block to compute a prediction block for the current block. The video decoder 300 can then inversely scan the reproduced coefficients (376) to create a block of quantized transform coefficients. The video decoder 300 can then inversely quantize and inverse transform the coefficients to produce a residual block (378). The video decoder 300 can finally decode the current block by combining the prediction block and the residual block (380).

[0268] Figure 7 This is a flowchart illustrating an example method for decoding video data according to the techniques of this disclosure. The video decoder 300 can perform the above... Figure 6 Execute the method before Figure 7 The video encoder 200 can perform a substantially similar method, although in some examples described below, a reversible technique is used.

[0269] Initially, the video decoder 300 can receive a video parameter set (VPS). The video decoder 300 can decode the data of the VPS to determine the number (400) of PTL data structures in the VPS. For example, the video decoder 300 can decode the value of the vps_num_ptls_minus1 syntax element of the VPS and determine the number of PTL data structures in the VPS from the value of the vps_num_ptls_minus1 syntax element.

[0270] The video decoder 300 can also determine the total number of OLS specified for the VPS (402). For example, as described above, the video decoder 300 can calculate the value of the variable TotalNumOlss as follows:

[0271]

[0272] In other words, when the maximum number of layers in the VPS minus 1 equals zero, the video decoder 300 can determine that the total number of OLS is equal to 1; when at least one of the following is true: 1) each layer in the VPS is an OLS, 2) the OLS mode indicator value is equal to 0, or 3) the OLS mode indicator value is equal to 1, the video decoder 300 can determine that the total number of OLS is equal to the maximum number of layers in the VPS; or when the OLS mode indicator value is equal to 2, the video decoder 300 can determine that the total number of OLS is equal to the number of independent layers plus the value of the syntax element of the VPS indicating the number of OLS.

[0273] The video decoder 300 can then determine whether the number of PTL data structures equals the total number of OLS specified for the VPS. Figure 7 In the example, suppose video decoder 300 determines that the number of PTL data structures is equal to the total number of OLS specified for the VPS (404). As a result, video decoder 300 can infer the value of the OLS PTL index (406). Specifically, video decoder 300 can infer the value of the OLS PTL index (e.g., ols_ptl_idx[i]) without explicitly decoding the value representing the index value from the VPS. For example, for all values ​​of i between 0 and the total number of OLS, video decoder 300 can determine that the i-th OLS PTL index value in the OLS PTL index values ​​is equal to i.

[0274] The video decoder 300 can also use the corresponding PTL data structure to decode OLS video data (408). For example, the video decoder 300 can allocate an appropriate amount of memory in the storage device, initialize the codec tools, and avoid initializing unused codec tools according to the PTL data structure. The video decoder 300 can then use the initialized codec tools (e.g., according to the above description) Figure 6 The method is used to decode blocks of video data.

[0275] In this way, Figure 7 The method represents an example of a method for decoding video data, including: determining that the value of a syntax element representing the number of grade-level-level (PTL) data structures in a video parameter set (VPS) in a bitstream is equal to the total number of output layer sets (OLS) specified for the VPS; in response to determining that the value of the syntax element representing the number of grade-level-level data structures in the VPS is equal to the total number of OLS specified for the VPS, inferring the value of the OLS PTL index value without explicitly decoding the value of the OLS PTL index value; and decoding the video data of one or more OLS using the corresponding PTL data structure in the PTL data structure in the VPS based on the inferred value of the OLS PTL index value.

[0276] Figure 8This is a flowchart illustrating another example method for decoding video data according to the techniques of this disclosure. The video decoder 300 can decode the image width value (420) of a Picture Parameter Set (PPS). The image width value can be, for example, pic_width_in_luma_samples. The video decoder 300 can also decode the maximum image width value (422) of a Sequence Parameter Set (SPS) referenced by the PPS. For example, the maximum image width value can be pic_width_max_in_luma_samples. Similarly, the video decoder 300 can decode the image height value (424) of the PPS and the maximum image height value (426) of the SPS.

[0277] The video decoder 300 can then determine whether the image corresponding to the PPS has a maximum image height. For example, the video decoder 300 can determine whether the image width value is equal to the maximum image width value, and whether the image height value is equal to the maximum image height value. In this example, assume that the video decoder 300 determines that the image has a maximum image size (428).

[0278] As a result, the video decoder 300 can determine that the PPS conformance window value has not been explicitly encoded or decoded (430). For example, the video decoder 300 can determine that the pps_conformance_window_flag of the PPS has a value of 0 (e.g., by inference of the value or explicit decoding). Therefore, according to the techniques of this disclosure, the video decoder 300 can infer the conformance window offset value from the corresponding SPS value of the SPS (432). For example, the video decoder 300 can infer the values ​​of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset of the PPS from the sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and pps_conf_win_bottom_offset of the SPS, respectively. Specifically, the video decoder 300 can infer the consistency window offset value without explicitly decoding data from the PPS values, such as pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset.

[0279] In this way, Figure 8The method represents an example of a method that includes: in response to determining that the value of the syntax element representing the image width in the picture parameter set (PPS) of the bitstream is a maximum image width value, and the value of the syntax element representing the image height in the PPS is a maximum image height value, determining that the consistency window value is equal to zero. Furthermore, Figure 8 The method represents an example of a method that includes: in response to determining that the consistency window value in the picture parameter set (PPS) is equal to zero, inferring that the value of the consistency window offset of the PPS is equal to the corresponding value of the consistency window offset of the sequence parameter set (SPS).

[0280] Certain techniques of this disclosure are described in the following terms:

[0281] Clause 1: A method for processing video data, the method comprising: processing a Decoder Capability Information (DCI) Supplemental Enhancement Information (SEI) message of a video bitstream, the DCI SEI message including data indicating information representing the capabilities that a video decoder must have to decode the video bitstream; and providing the video bitstream to the video decoder when the video decoder has the capability.

[0282] Clause 2: The method of Clause 1, wherein the information representing capability includes data indicating the maximum number of temporal sublayers that may exist in each encoded video sequence (CVS) of the video bitstream.

[0283] Clause 3: The method of any of Clauses 1 or 2, wherein the information representing capability includes data indicating the number of grade, level, and rank syntax structures included in the DCI SEI message.

[0284] Clause 4: The method of any of Clauses 1-3, wherein the information representing capability includes data indicating the highest value of a generic level indicator to be supported by the video decoder.

[0285] Clause 5: The method of any of Clauses 1-4, wherein the information representing capability includes data indicating the highest value of a generic level indicator to be supported by the video decoder.

[0286] Clause 6: A method for encoding and decoding video data, the method comprising: encoding and decoding a value of a syntax element indicating whether all layers in an encoded and decoded video sequence (CVS) are independently encoded and decoded without using inter-layer prediction, whether all non-base layers in the CVS use inter-layer prediction, whether each layer i is a direct reference layer of layer i+1, whether sub-layers of all layers except the highest layer are used for inter-layer prediction, or whether one or more layers in the CVS can use inter-layer prediction; and encoding and decoding a picture of the layers of the CVS according to the value of the syntax element.

[0287] Clause 7: The methods of Clause 6 shall also include the methods of any of Clauses 1-5.

[0288] Clause 8: A method as described in any of Clauses 6 or 7, wherein the syntax element includes a first syntax element, the method further comprising inferring the value of a second syntax element, the second syntax element indicating whether the layer uses inter-layer prediction when the second syntax element is not encoded or decoded based on the value of the first syntax element.

[0289] Clause 9: The method of Clause 8, wherein inferring the value of the second syntax element comprises: inferring the value of the second syntax element to be 0 when the value of the first syntax element is 1; or inferring the value of the second syntax element to be 1 when the value of the first syntax element is 0.

[0290] Clause 10: The method of Clause 8, wherein inferring the value of the second syntax element includes inferring the value of the second syntax element as 1 minus the value of the first syntax element.

[0291] Clause 11: The method of any of Clauses 8-10, wherein the second syntax element includes vps_independent_layer_flag[i].

[0292] Clause 12: The method of any of Clauses 6-11, wherein encoding or decoding the value of a syntax element includes encoding or decoding the value in the Video Parameter Set (VPS).

[0293] Clause 13: A method for encoding and decoding video data, the method comprising: encoding and decoding the value of a syntax element indicating whether a default maximum temporal layer identifier for an interlayer reference picture of the video data exists in the video bitstream; and encoding and decoding a picture of the video bitstream based on the value of the syntax element.

[0294] Clause 14: The method of Clause 13 shall also include the method of any of Clauses 1-12.

[0295] Clause 15: The method of any of Clauses 13 or 14 further includes encoding and decoding the value of the syntax element indicating the actual maximum temporal layer identifier of the interlayer reference picture when the value of the syntax element indicating the default maximum temporal layer identifier of the interlayer reference picture is not present in the video bitstream.

[0296] Clause 16: The method of any of Clauses 13-15 further includes encoding and decoding the value of the default maximum temporal layer identifier when the value of the syntax element indicates that the value of the default maximum temporal layer identifier is present in the video bitstream.

[0297] Clause 17: A method for encoding and decoding video data, the method comprising: encoding and decoding data of an output layer set configured as a non-basic independent codec layer for the video data; and encoding and decoding the video data according to the indicated output layer set.

[0298] Clause 18: The methods of Clause 17 shall also include the methods of any of Clauses 1-16.

[0299] Clause 19: A method for encoding and decoding video data, the method comprising: determining that the number of output layer sets (OLS) is equal to the number of grade, level, rank (PTL) data structures of the video bitstream; in response to determining that the number of OLS is equal to the number of PTL data structures, inferring an index between the OLS and the PTL data structures without encoding or decoding the value of an index; and encoding and decoding the video data based on the inferred index.

[0300] Clause 20: The methods of Clause 19 shall also include the methods of any of the Clauses 1-18.

[0301] Clause 21: A method for encoding and decoding video data, the method comprising: encoding and decoding the values ​​of syntax elements of a sequence parameter set (SPS) of a video bitstream to prevent emulation of byte encoding and decoding; and encoding and decoding the video bitstream according to the SPS.

[0302] Clause 22: The methods of Clause 21 shall also include the methods of any of the Clauses 1-20.

[0303] Clause 23: The method of any of Clauses 21 or 22, wherein encoding / decoding the value of a syntax element of the SPS includes encoding / decoding the non-zero value of at least one syntax element.

[0304] Clause 24: The method of any of Clauses 21-23, wherein encoding and decoding the value of a syntax element of the SPS comprises encoding and decoding the value of a syntax element of the SPS that represents the SPS-specific maximum sublayer number in the range of values ​​of a syntax element of the Video Parameter Set (VPS) indicating the VPS-specific maximum sublayer number.

[0305] Clause 25: The method of any of Clauses 21-23, wherein encoding or decoding the value of a syntax element of the SPS includes encoding or decoding the binary value “1111” following the first 11 bits of the SPS.

[0306] Clause 26: A method for encoding and decoding video data, the method comprising: inferring values ​​of consistency window syntax elements of a picture parameter set (PPS) from corresponding consistency window syntax elements of a sequence parameter set (SPS) without encoding or decoding the values ​​of consistency window syntax elements of the PPS of a video bitstream; and encoding or decoding video data of the video bitstream based on the inferred values.

[0307] Clause 27: The methods of Clause 26 shall also include the methods of any of Clauses 1-25.

[0308] Clause 28: The method of any of Clauses 26 or 27, wherein the inferred value includes inferring the value when the value of the syntax element indicating the picture size of the video bitstream indicates that the picture size is equal to the maximum possible picture size of the video bitstream.

[0309] Clause 29: The method of Clause 27 also includes encoding and decoding the value of the syntax element that defines the maximum possible image size.

[0310] Clause 30: A method for encoding and decoding video data, the method comprising: encoding and decoding the value of a syntax element of a Picture Parameter Set (PPS), the syntax element indicating whether the picture resolution of a picture referenced to the PPS can be changed; when the value of the syntax element indicates that the picture resolution cannot be changed, inferring the value of a syntax element of the PPS defining the picture resolution of the picture referenced to the PPS from the corresponding value of a syntax element of a corresponding Sequence Parameter Set (SPS), without encoding or decoding the value of the syntax element of the PPS; and encoding and decoding the picture referenced to the PPS based on the value of the syntax element of the PPS.

[0311] Clause 31: The method of Clause 30 shall also include the method of any of the clauses 1-29.

[0312] Clause 32: The method as in any of Clauses 30 and 31, wherein the syntax element of PPS includes pps_res_change_allowed_flag.

[0313] Clause 33: The method of any of Clauses 30-32, wherein inferring the value of a syntax element includes inferring the values ​​of the pps_conformance_window_flag and scaling_window_explicit_signaling flags.

[0314] Clause 34: A method for encoding and decoding video data, the method comprising: encoding and decoding the value of a single syntax element indicating the bit length of the most significant bit (MSB) of the picture sequence count (POC) present in a picture header of a video bitstream; encoding and decoding the POC MSB of the picture header based on the value of the single syntax element; and encoding and decoding the picture corresponding to the picture header using the POC MSB.

[0315] Clause 35: The methods of Clause 34 shall also include the methods of any of Clauses 1-33.

[0316] Clause 36: A method for encoding and decoding video data, the method comprising: encoding and decoding a value of a syntax element indicating the number of slice heights explicitly provided in a current slice containing video data of more than one rectangular slice, the value being in the range of 1 to the height of the current slice in a row; and encoding and decoding the video data of the current slice based on the value of the syntax element.

[0317] Clause 37: The methods of Clause 36 shall also include the methods of any of Clauses 1-35.

[0318] Clause 38: The method of any of the clauses 6-37, wherein encoding / decoding includes decoding.

[0319] Clause 39: The method of any of the clauses 6-38, wherein encoding / decoding includes encoding.

[0320] Clause 40: An apparatus for processing or encoding / decoding video data, the apparatus comprising one or more components for performing the methods of any one of Clauses 1-39.

[0321] Clause 41: A device as described in Clause 40, wherein one or more components include one or more processors implemented in a circuit.

[0322] Clause 42: The device as described in Clause 40 also includes a display configured to display decoded video data.

[0323] Clause 43: Devices as described in Clause 40, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device, or set-top box.

[0324] Clause 44: A device as described in Clause 40 also includes a memory configured to store video data.

[0325] Clause 45: A computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor of an apparatus for processing or encoding / decoding video data to perform any of the clauses 1-39.

[0326] Clause 46: A method for decoding video data, the method comprising: determining that the value of a syntax element representing the number of grade-level-hierarchy (PTL) data structures in a video parameter set (VPS) of a bitstream is equal to the total number of output layer sets (OLS) specified for the VPS; in response to determining that the value of the syntax element representing the number of grade-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS, inferring the value of an OLS PTL index value without explicitly decoding the value of the OLS PTL index value; and decoding the video data of one or more OLS using the corresponding PTL data structure in the PTL data structures of the VPS based on the inferred value of the OLS PTL index value.

[0327] Clause 47: The method of Clause 46, wherein inferring the value of the OLS PTL index includes: for all i values ​​between 0 and the total number of OLS, determining that the i-th OLS PTL index value is equal to i.

[0328] Clause 48: The method described in any of Clauses 46 and 47, wherein decoding video data of one or more OLSs comprises: for any value i between 0 and the total number of OLSs, using the i-th PTL data structure to decode the i-th OLS.

[0329] Clause 49: As in any of Clauses 46-48, where the syntax element representing the number of PTL data structures in the VPS includes vps_num_ptls_minus1.

[0330] Clause 50: The method of any of Clauses 46-49 further includes determining the total number of OLS, including: determining the total number of OLS equal to 1 when the maximum number of tiers of the VPS minus 1 equals zero; determining the total number of OLS equal to the maximum number of tiers of the VPS when at least one of the following: 1) each tier of the VPS is an OLS, 2) the OLS mode indicator value is equal to 0, or 3) the OLS mode indicator value is equal to 1; or determining the total number of OLS equal to the number of independent tiers plus the value of the syntax element of the VPS indicating the number of OLS when the OLS mode indicator value is equal to 2.

[0331] Clause 51: The method of any of Clauses 46-50 further includes: in response to determining that the value of a syntax element representing the picture width in the Picture Parameter Set (PPS) of the bitstream is a maximum picture width value, and determining that the value of a syntax element representing the picture height in the PPS is a maximum picture height value, determining that the consistency window value is equal to zero.

[0332] Clause 52: The method of Clause 51 further includes, in response to determining that the consistency window value in the picture parameter set (PPS) is equal to zero, inferring that the consistency window offset of the PPS is equal to the corresponding value of the consistency window offset of the sequence parameter set (SPS).

[0333] Clause 53: An apparatus for decoding video data, the apparatus comprising: a memory configured to store video data; and one or more processors implemented in a circuit and configured to: determine that the value of a syntax element representing the number of grade-level-hierarchy (PTL) data structures in a video parameter set (VPS) of a bitstream is equal to the total number of output layer sets (OLS) specified for the VPS; in response to determining that the value of the syntax element representing the number of grade-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS, inferring the value of an OLS PTL index value without explicitly decoding the value of the OLS PTL index value; and decoding the video data of one or more OLS using the corresponding PTL data structure in the PTL data structures of the VPS based on the inferred value of the OLS PTL index value.

[0334] Clause 54: A device as described in Clause 53, wherein, in order to infer the value of an OLS PTL index value, one or more processors are configured to: determine, for all i values ​​between 0 and the total number of OLS, that the i-th OLS PTL index value is equal to i.

[0335] Clause 55: A device as described in any of Clauses 53 and 54, wherein, in order to decode video data of one or more OLS, one or more processors are configured to: for any value i between 0 and the total number of OLS, use the i-th data structure in the PTL to decode the i-th data in the OLS.

[0336] Clause 56: For any of the clauses 53-55, the syntax element representing the number of PTL data structures in the VPS includes vps_num_ptls_minus1.

[0337] Clause 57: For any of Clauses 53-56, one or more processors are further configured to determine the total number of OLS, including: determining that the total number of OLS is equal to 1 when the maximum number of tiers of the VPS minus 1 equals zero; determining that the total number of OLS is equal to the maximum number of tiers of the VPS when at least one of the following: 1) each tier of the VPS is an OLS, 2) the OLS mode indicator value is equal to 0, or 3) the OLS mode indicator value is equal to 1; or determining that the total number of OLS is equal to the number of independent tiers plus the value of the syntax element of the VPS indicating the number of OLS when the OLS mode indicator value is equal to 2.

[0338] Clause 58: A device as described in any of Clauses 53-57, wherein one or more processors are further configured to determine that the consistency window value is zero in response to determining that the value of the syntax element representing the picture width in the picture parameter set (PPS) of the bitstream is a maximum picture width value and the value of the syntax element representing the picture height in the PPS is a maximum picture height value.

[0339] Clause 59: In a device as described in Clause 58, one or more processors are further configured to, in response to determining that the consistency window value in the picture parameter set (PPS) is equal to zero, infer a value of the consistency window offset of the PPS as equal to a corresponding value of the consistency window offset of the sequence parameter set (SPS).

[0340] Clause 60: Devices as described in Clause 53 also include displays configured to display decoded video data.

[0341] Clause 61: Devices as described in Clause 53, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device, or set-top box.

[0342] Clause 62: A computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor to: determine that the value of a syntax element representing the number of grade-level-hierarchy (PTL) data structures in a video parameter set (VPS) of a bitstream is equal to the total number of output layer sets (OLS) specified for the VPS; in response to determining that the value of the syntax element representing the number of grade-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS, infer the value of an OLS PTL index value without explicitly decoding the value of the OLS PTL index value; and, based on the inferred value of the OLS PTL index value, decode the video data of one or more OLS using the corresponding PTL data structure in the PTL data structures of the VPS.

[0343] Clause 63: A computer-readable storage medium as described in Clause 62, wherein the instructions that cause a processor to infer the value of an OLS PTL index include instructions that cause the processor to: determine, for all i values ​​between 0 and the total number of OLS, determine that the i-th OLS PTL index value is equal to i.

[0344] Clause 64: A computer-readable storage medium as described in any of Clauses 62 and 63, wherein instructions causing a processor to decode one or more OLSs include: for any value i between 0 and the total number of OLSs, causing the processor to use the i-th PTL data structure to decode the instructions of the i-th OLS.

[0345] Clause 65: A computer-readable storage medium as described in any of Clauses 62-64, wherein the syntax element representing the number of PTL data structures in the VPS includes vps_num_ptls_minus1.

[0346] Clause 66: A computer-readable storage medium as described in any of Clauses 62-65 further includes instructions that cause a processor to determine the total number of OLS, comprising instructions that cause the processor to: determine that the total number of OLS is equal to 1 when the maximum number of tiers of the VPS minus 1 equals zero; determine that the total number of OLS is equal to the maximum number of tiers of the VPS when at least one of the following: 1) each tier of the VPS is an OLS, 2) the OLS mode indicator value is equal to 0, or 3) the OLS mode indicator value is equal to 1; or determine that the total number of OLS is equal to the number of independent tiers plus the value of the syntax element of the VPS indicating the number of OLS when the OLS mode indicator value is equal to 2.

[0347] Clause 67: A computer-readable storage medium as described in any of Clauses 62-66 further includes instructions that cause a processor to perform the following operation: in response to determining that the value of a syntax element representing the picture width in the picture parameter set (PPS) of the bitstream is a maximum picture width value and the value of a syntax element representing the picture height in the PPS is a maximum picture height value, determining that the consistency window value is equal to zero.

[0348] Clause 68: The computer-readable storage medium of Clause 67 further includes instructions that cause the processor to perform the following operation: in response to determining that the consistency window value in the picture parameter set (PPS) is equal to zero, infer that the consistency window offset of the PPS is equal to the corresponding value of the consistency window offset of the sequence parameter set (SPS).

[0349] Clause 69: An apparatus for decoding video data, the apparatus comprising: means for determining that the value of a syntax element representing the number of grade-level-hierarchy (PTL) data structures in a video parameter set (VPS) of a bitstream is equal to the total number of output layer sets (OLS) specified for the VPS; means for inferring the value of an OLS PTL index value in response to determining that the value of the syntax element representing the number of grade-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS, without explicitly decoding the value of the OLS PTL index value; and means for decoding video data of one or more OLS using the corresponding PTL data structure in the PTL data structures of the VPS based on the inferred value of the OLS PTL index value.

[0350] Clause 70: The apparatus of Clause 69, wherein the component for inferring the value of the OLS PTL index value includes a component for: determining, for all i values ​​between 0 and the total number of OLS, that the i-th OLS PTL index value is equal to i.

[0351] Clause 71: A device as described in any of Clauses 69 and 70, wherein the component for decoding one or more OLSs includes a component for: using the i-th PTL data structure to decode the i-th OLS for any i value between 0 and the total number of OLSs.

[0352] Clause 72: For any of the clauses 69-71, the syntax element representing the number of PTL data structures in the VPS includes vps_num_ptls_minus1.

[0353] Clause 73: The apparatus of any of Clauses 69-72 further includes components for determining the total number of OLS, including: components for determining that the total number of OLS is equal to 1 when the maximum number of tiers of the VPS minus 1 equals zero; components for determining that the total number of OLS is equal to the maximum number of tiers of the VPS when at least one of the following: 1) each tier of the VPS is an OLS, 2) the OLS mode indicator value is equal to 0, or 3) the OLS mode indicator value is equal to 1; or components for determining that the total number of OLS is equal to the number of independent tiers plus the value of the syntax element of the VPS indicating the number of OLS when the OLS mode indicator value is equal to 2.

[0354] Clause 74: The apparatus of any of Clauses 69-73 further includes a component for determining that, in response to determining that the value of the syntax element representing the picture width in the picture parameter set (PPS) of the bitstream is a maximum picture width value and the value of the syntax element representing the picture height in the PPS is a maximum picture height value, the consistency window value is equal to zero.

[0355] Clause 75: The apparatus of Clause 74 further includes a component for inferring, in response to determining that a consistency window value in the Picture Parameter Set (PPS) is equal to zero, a value equal to the corresponding value of the consistency window offset in the Sequence Parameter Set (SPS). It should be recognized that, depending on the example, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technique). Moreover, in some examples, actions or events may be performed concurrently, for example through multithreaded processing, interrupt handling, or multiprocessor, rather than sequentially.

[0356] In one or more examples, the described functionality can be implemented using hardware, software, firmware, or any combination thereof. If implemented in software, these functions can be stored or transmitted as one or more instructions or code to a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium can include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium that includes, for example, any medium facilitating the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium can be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products can include computer-readable media.

[0357] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks generally reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0358] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuitry" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be implemented entirely within one or more circuit or logic elements.

[0359] The techniques disclosed herein can be implemented in a variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to highlight functional aspects of a device configured to perform the disclosed techniques, but implementation by different hardware units is not necessarily required. Rather, as described above, various units can be combined in a codec hardware unit or provided by a set of interoperable hardware units including one or more processors as described above, combined with suitable software and / or firmware.

[0360] Various examples have been described. These and other examples are all within the scope of the appended claims.

Claims

1. A method for decoding video data, the method comprising: The value of the syntax element that determines the number of tier-level-grade PTL data structures in the video parameter set (VPS) representing the bitstream is equal to the total number of output layer sets (OLS) specified for the VPS. In response to determining that the value of the syntax element representing the number of tier-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS, the value of the OLS PTL index value is inferred without explicitly decoding the value of the OLS PTL index value; and Based on the inferred value of the OLS PTL index, the corresponding PTL data structure in the VPS is used to decode one or more video data from the OLS.

2. The method of claim 1, wherein inferring the value of the OLS PTL index includes: For all values ​​of i between 0 and the total number of OLS, determine that the i-th OLS PTL index value is equal to i.

3. The method of claim 1, wherein decoding the video data of the one or more OLS comprises: For any value of i between 0 and the total number in OLS, the i-th element in the PTL data structure is used to decode the i-th element in the OLS.

4. The method of claim 1, wherein the syntax element representing the number of PTL data structures in the VPS includes vps_num_ptls_minus1.

5. The method of claim 1, further comprising determining the total number of OLS, including: When the maximum number of VPS layers minus 1 equals zero, the total number of OLS is determined to be equal to 1; The total number of OLS is determined to be equal to the maximum number of tiers of the VPS when at least one of the following conditions is met: 1) each tier of the VPS is an OLS, 2) the OLS mode indicator value is equal to 0, or 3) the OLS mode indicator value is equal to 1. or When the OLS mode indicator value is equal to 2, the total number of OLS is determined to be equal to the number of independent layers plus the value of the syntax element of the VPS that indicates the number of OLS.

6. The method according to claim 1, further comprising: In response to determining that the value of the syntax element representing the image width in the image parameter set PPS of the bitstream is the maximum image width value, and determining that the value of the syntax element representing the image height in the PPS is the maximum image height value, the consistency window value is determined to be zero.

7. The method of claim 1, further comprising, in response to determining that the consistency window value in the image parameter set PPS is equal to zero, inferring that the consistency window offset of the PPS is equal to the corresponding value of the consistency window offset of the sequence parameter set SPS.

8. An apparatus for decoding video data, the apparatus comprising: A memory configured to store video data; as well as One or more processors implemented in the circuit, and configured to: The value of the syntax element that determines the number of tier-level-grade PTL data structures in the video parameter set (VPS) representing the bitstream is equal to the total number of output layer sets (OLS) specified for the VPS. In response to determining that the value of the syntax element representing the number of tier-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS, the value of the OLS PTL index value is inferred without explicitly decoding the value of the OLS PTL index value; and Based on the inferred value of the OLS PTL index, the corresponding PTL data structure in the PTL data structure in the VPS is used to decode the video data of one or more OLS.

9. The device according to claim 8, wherein, In order to infer the value of the OLS PTL index value, the one or more processors are configured to: for all i values ​​between 0 and the total number of OLS, determine that the i-th OLS PTL index value is equal to i.

10. The apparatus of claim 8, wherein, in order to decode the video data of the one or more OLS, the one or more processors are configured to: for any i value between 0 and the total number of OLS, use the i-th data structure in the PTL to decode the i-th data in the OLS.

11. The device of claim 8, wherein the syntax element representing the number of PTL data structures in the VPS includes vps_num_ptls_minus1.

12. The device according to claim 8, wherein, The one or more processors are further configured to determine the total number of OLS, including: When the maximum number of VPS layers minus 1 equals zero, the total number of OLS is determined to be equal to 1; The total number of OLS is determined to be equal to the maximum number of tiers of the VPS when at least one of the following conditions is met: 1) each tier of the VPS is an OLS; 2) the OLS mode indicator value is equal to 0; or 3) the OLS mode indicator value is equal to 1. When the OLS mode indicator value is equal to 2, the total number of OLS is determined to be equal to the number of independent layers plus the value of the syntax element of the VPS that indicates the number of OLS.

13. The apparatus of claim 8, wherein the one or more processors are further configured to: determine that the value of the syntax element representing the image width in the image parameter set PPS of the bitstream is a maximum image width value, and the value of the syntax element representing the image height in the PPS is a maximum image height value, thereby determining that the consistency window value is equal to zero.

14. The apparatus of claim 8, wherein the one or more processors are further configured to: in response to determining that the consistency window value in the picture parameter set PPS is equal to zero, infer a value of the consistency window offset of the PPS as equal to a corresponding value of the consistency window offset of the sequence parameter set SPS.

15. The device of claim 8, further comprising a display configured to display decoded video data.

16. The device of claim 8, wherein the device comprises one or more of a camera, computer, mobile device, broadcast receiver device, or set-top box.

17. A computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor to: The value of the syntax element that determines the number of tier-level-grade PTL data structures in the video parameter set (VPS) representing the bitstream is equal to the total number of output layer sets (OLS) specified for the VPS. In response to determining that the value of the syntax element representing the number of tier-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS, the value of the OLS PTL index value is inferred without explicitly decoding the value of the OLS PTL index value; and Based on the inferred value of the OLS PTL index, the corresponding PTL data structure in the PTL data structure in the VPS is used to decode video data from one or more OLS.

18. The computer-readable storage medium of claim 17, wherein the instructions causing the processor to infer the value of the OLSPTL index value include instructions causing the processor to: determine, for all i values ​​between 0 and the total number of OLS, that the i-th OLS PTL index value of the OLS PTL index value is equal to i.

19. The computer-readable storage medium of claim 17, wherein causing the processor to decode the instructions of the one or more OLS comprises causing the processor to: for any i value between 0 and the total number of OLS, use the i-th instruction in the PTL data structure to decode the i-th instruction in the OLS.

20. The computer-readable storage medium of claim 17, wherein the syntax element representing the number of PTL data structures in the VPS includes vps_num_ptls_minus1.

21. The computer-readable storage medium of claim 17, further comprising instructions for causing the processor to determine the total number of OLS, including instructions for causing the processor to perform the following operations: When the maximum number of VPS layers minus 1 equals zero, the total number of OLS is determined to be equal to 1; The total number of OLS is determined to be equal to the maximum number of tiers of the VPS when at least one of the following conditions is met: 1) each tier of the VPS is an OLS, 2) the OLS mode indicator value is equal to 0, or 3) the OLS mode indicator value is equal to 1. or When the OLS mode indicator value is equal to 2, the total number of OLS is determined to be equal to the number of independent layers plus the value of the syntax element of the VPS that indicates the number of OLS.

22. The computer-readable storage medium of claim 17, further comprising an instruction that causes the processor to: determine a consistency window value equal to zero in response to determining that the value of a syntax element representing the image width in the image parameter set PPS of the bitstream is a maximum image width value and the value of a syntax element representing the image height in the PPS is a maximum image height value.

23. The computer-readable storage medium of claim 17, further comprising an instruction that causes the processor to: in response to determining that the consistency window value in the picture parameter set PPS is equal to zero, infer a value of the consistency window offset of the PPS as equal to the corresponding value of the consistency window offset of the sequence parameter set SPS.

24. An apparatus for decoding video data, the apparatus comprising: The value of the syntax element used to determine the number of grade-level-hierarchy PTL data structures in the video parameter set (VPS) representing the bitstream is equal to the total number of output layer sets (OLS) specified for the VPS. A component for inferring the value of an OLS PTL index value without explicitly decoding it, in response to determining that the value of the syntax element representing the number of tier-level-hierarchy data structures in the VPS is equal to the total number of OLS specified for the VPS. as well as A component used to decode video data of one or more OLS using the corresponding PTL data structure in the PTL data structure in the VPS, based on the inferred value of the OLS PTL index value.

25. The device of claim 24, wherein the component for inferring the value of the OLS PTL index value includes a component for: determining, for all i values ​​between 0 and the total number of OLS, that the i-th OLS PTL index value is equal to i.

26. The apparatus of claim 24, wherein the component for decoding the one or more OLSs includes a component for: decoding the i-th OLS using the i-th data structure in the PTL for any i-value between 0 and the total number of OLSs.

27. The device of claim 24, wherein the syntax element representing the number of PTL data structures in the VPS includes vps_num_ptls_minus1.

28. The apparatus of claim 24, further comprising components for determining the total number of OLS, including: A component used to determine that the total number of OLS is equal to 1 when the maximum number of layers of the VPS minus 1 equals zero; A component for determining that the total number of OLS is equal to the maximum number of tiers of the VPS when at least one of the following conditions is met: 1) each tier of the VPS is an OLS, 2) the OLS mode indicator value is equal to 0, or 3) the OLS mode indicator value is equal to 1. or A component for determining, when the OLS mode indicator value is equal to 2, that the total number of OLS is equal to the number of independent layers plus the value of the syntax element indicating the number of OLS in the VPS.

29. The apparatus of claim 24, further comprising a component for determining a consistency window value equal to zero in response to determining that the value of a syntax element representing the image width in the image parameter set PPS of the bitstream is a maximum image width value and the value of a syntax element representing the image height in the PPS is a maximum image height value.

30. The apparatus of claim 24, further comprising a component for inferring, in response to determining that the consistency window value in the picture parameter set PPS is equal to zero, that the value of the consistency window offset of the PPS is equal to the corresponding value of the consistency window offset of the sequence parameter set SPS.