Method for encoding or decoding a video parameter set or a sequence parameter set - Patents.com

By optimizing the signaling of video parameter sets and sequence parameter sets based on the equality of syntax structures, the method enhances compression efficiency and reduces unnecessary data transmission in video coding, addressing inefficiencies in existing standards.

JP7739579B2Active Publication Date: 2025-09-16ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024228716
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-26
Filing Date
2024-12-25
Publication Date
2025-09-16
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing video coding standards face inefficiencies in signaling video parameter sets and sequence parameter sets, particularly in scenarios where the number of profile/tier/level syntax structures and output layer sets are equal, leading to unnecessary signaling of indices and flags, which can impact compression efficiency.

Method used

A method for encoding and decoding video that determines the equality of profile/tier/level syntax structures and decoded picture buffer parameter structures with output layer sets, and adjusts signaling of syntax elements accordingly, including conditional encoding and decoding of flags and indices based on the presence and values of these structures within the bitstream.

Benefits of technology

Enhances compression efficiency by optimizing the signaling of video parameter sets and sequence parameter sets, reducing unnecessary data transmission and improving bandwidth utilization in video coding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007739579000001
    Figure 0007739579000001
  • Figure 0007739579000002
    Figure 0007739579000002
  • Figure 0007739579000003
    Figure 0007739579000003
Patent Text Reader

Abstract

To provide a computer-implemented method for encoding video.SOLUTION: A method includes: determining whether a coded video sequence (CVS) contains equal numbers of profile, tier and level (PTL) syntax structures and output layer sets (OLSs); and, in response to the CVS containing the equal numbers of PTL syntax structures and OLSs, coding a bitstream without signaling a first PTL syntax element specifying an index, to a list of PTL syntax structures in a VPS, of a PTL syntax structure that applies to a corresponding OLS in the VPS.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 994,995, filed March 26, 2020, the entire contents of which are incorporated herein by reference.

[0002] Technical Field FIELD OF THE DISCLOSURE

[0002] The present disclosure relates generally to video processing, and more particularly to a video processing method for signaling video parameter sets (VPS) and sequence parameter sets (SPS). [Background technology]

[0003] background

[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is usually referred to as encoding, and the decompression process is usually referred to as decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards that specify specific video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard. As more and more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]

[0004] Disclosure Overview

[0004] Embodiments of the present disclosure provide a computer-implemented method for encoding video. In some embodiments, the method includes determining whether a coded video sequence (CVS) includes an equal number of profile / tier / level (PTL) syntax structures and an output layer set (OLS), and, in response to the CVS including an equal number of PTL syntax structures and an OLS, encoding a bitstream without signaling a first PTL syntax element specifying an index into a list of PTL syntax structures in the VPS of a PTL syntax structure that applies to a corresponding OLS in the VPS.

[0005]

[0005] In some embodiments, the method includes determining whether a coded video sequence (CVS) has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLSs), and in response to the CVS having an equal number of DPB parameter syntax structures and OLSs, encoding the bitstream without signaling a first DPB syntax element specifying an index to a list of DPB parameter syntax structures in the VPS of the DPB parameter syntax structure that applies to the corresponding OLS.

[0006]

[0006] In some embodiments, the method includes determining whether at least one of a profile / tier / level (PTL) syntax structure, a decoded picture buffer (DPB) parameter syntax structure, or a hypothetical reference decoder (HRD) parameter syntax structure is within a sequence parameter set (SPS) of the bitstream, determining whether a first value is greater than 1, the first value specifying a maximum number of temporal sub-layers within a coded layer video sequence (CLVS) that references the SPS, and signaling a flag configured to control the presence of syntax elements within the DPB parameter syntax structure within the SPS if at least one of the PTL syntax structure, the DPB parameter syntax structure, or the HRD parameter syntax structure is within the SPS and the first value is greater than 1.

[0007]

[0007] In some embodiments, the method includes determining a value of a first sequence parameter set (SPS) syntax element, wherein the first SPS syntax element specifies an identifier of a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero; in response to the value of the first SPS syntax element being greater than zero, assigning a range of a second SPS syntax element specifying a maximum number of temporal sub-layers in each coded layer video sequence (CLVS) that references the SPS based on the corresponding VPS syntax element; and in response to the value of the first SPS syntax element being equal to zero, assigning a range of the second SPS syntax element specifying the maximum number of temporal sub-layers in each CLVS that references the SPS to be in the range greater than or equal to zero and less than or equal to a fixed value.

[0008]

[0008] In some embodiments, the method includes encoding one or more profile / tier / level (PTL) syntax elements that specify PTL-related information, and signaling one or more PTL syntax elements having variable lengths within a video parameter set (VPS) or sequence parameter set (SPS) of the bitstream.

[0009]

[0009] In some embodiments, the method includes encoding a video parameter set (VPS) syntax element having a variable length and signaling the VPS syntax element within the VPS, wherein the VPS syntax element is associated with the number of output layer sets (OLS) included in a coded video sequence (CVS) that references the VPS.

[0010]

[0010] Embodiments of the present disclosure provide a computer-implemented method for decoding video. In some embodiments, the method includes receiving a bitstream including a Coded Video Sequence (CVS), determining whether the CVS has an equal number of Profile / Tier / Level (PTL) syntax structures and Output Layer Sets (OLS), and, in response to the number of PTL syntax structures being equal to the number of OLSs, when decoding a Video Picture Sequence (VPS), skipping decoding of a first PTL syntax element that specifies an index into a list of PTL syntax structures in the Video Picture Sequence (VPS) of a PTL syntax structure that applies to a corresponding OLS.

[0011]

[0011] In some embodiments, the method includes receiving a bitstream including a coded video sequence (CVS), determining whether the CVS has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLSs), and in response to the CVS having an equal number of DPB parameter syntax structures and OLSs, skipping decoding of a first DPB syntax element that specifies an index to a list of DPB parameter syntax structures in the VPS of a DPB parameter syntax structure that applies to the corresponding OLS.

[0012]

[0012] In some embodiments, the method includes receiving a bitstream including a video parameter set (VPS) and a sequence parameter set (SPS); in response to at least one of a profile / tier / level (PTL) syntax structure, a decoded picture buffer (DPB) parameter syntax structure, or a hypothetical reference decoder (HRD) parameter syntax structure being in the SPS, determining whether a first value specifying a maximum number of temporal sub-layers in each coded layer video sequence (CLVS) that references the SPS is greater than 1; and in response to the first value being greater than 1, decoding a flag configured to control the presence of syntax elements in the DPB parameter syntax structure in the SPS.

[0013]

[0013] In some embodiments, the method includes determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero, and decoding a second SPS syntax element specifying a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS, the range of the second SPS syntax element being based on a corresponding VPS syntax element of the VPS referenced by the SPS when the value of the first SPS syntax element is greater than zero, or based on a fixed value when the value of the first SPS syntax element is equal to zero.

[0014]

[0014] In some embodiments, the method includes receiving a bitstream including a video parameter set (VPS) or a sequence parameter set (SPS), and decoding one or more profile / tier / level (PTL) syntax elements within the VPS or SPS, wherein the one or more PTL syntax elements specify PTL-related information.

[0015]

[0015] In some embodiments, the method includes receiving a bitstream including a video parameter set (VPS) and decoding a VPS syntax element having a variable length within the VPS, wherein the VPS syntax element is associated with the number of output layer sets (OLS) included within a coded video sequence (CVS) that references the VPS.

[0016]

[0016] Embodiments of the present disclosure provide an apparatus. In some embodiments, the apparatus includes a memory configured to store instructions and a processor coupled to the memory, the processor configured to execute the instructions to perform a computer-implemented method for encoding video. In some embodiments, the apparatus includes a memory configured to store instructions and a processor coupled to the memory, the processor configured to execute the instructions to perform a computer-implemented method for decoding video.

[0017]

[0017] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of an apparatus to cause the apparatus to perform a method for encoding video.

[0018]

[0018] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium that stores a set of instructions executable by one or more processors of a device to cause the device to perform a method for decoding video.

[0019] BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]

[0020] [Figure 1]

[0020] FIG. 1 is a schematic diagram illustrating the structure of an example video sequence, according to some embodiments of the present disclosure. [Figure 2A]

[0021] 1 is a schematic diagram illustrating an exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 2B]

[0022] FIG. 2 is a schematic diagram illustrating another exemplary encoding process for a hybrid video encoding system, according to an embodiment of the present disclosure. [Figure 3A]

[0023] FIG. 2 is a schematic diagram illustrating an exemplary decoding process of a hybrid video coding system, according to an embodiment of the present disclosure. [Figure 3B]

[0024] FIG. 2 is a schematic diagram illustrating another exemplary decoding process of a hybrid video coding system, according to an embodiment of the present disclosure. [Figure 4]

[0025] 1 is a block diagram of an example device for encoding or decoding video in accordance with some embodiments of the present disclosure. [Figure 5]

[0026] 1 is a schematic diagram of an exemplary bitstream in accordance with some embodiments of the present disclosure. [Figure 6A]

[0027] 1 illustrates an example encoding syntax table of a PTL syntax structure according to some embodiments of the present disclosure. [Figure 6B]

[0028] 1 illustrates an example encoding syntax table of a DPB parameter syntax structure according to some embodiments of the present disclosure. [Figure 7A]

[0029] 1 illustrates an example encoding syntax table of an HRD parameter syntax structure according to some embodiments of the present disclosure. [Figure 7B]

[0030] 10 illustrates another example encoding syntax table of an HRD parameter syntax structure according to some embodiments of the present disclosure. [Figure 8A]

[0031] 1 illustrates an example encoding syntax table of a portion of a VPS Raw Byte Sequence Payload (RBSP) syntax structure, according to some embodiments of the present disclosure. [Figure 8B]

[0031] An example encoding syntax table of a portion of a VPS Raw Byte Sequence Payload (RBSP) syntax structure is shown in accordance with some embodiments of the present disclosure. [Figure 9]

[0032] 10 illustrates an example encoding syntax table of a portion of an SPS RBSP syntax structure according to some embodiments of the present disclosure. [Figure 10A]

[0033] 1 shows a flowchart of an exemplary video encoding method in accordance with some embodiments of the present disclosure. [Figure 10B]

[0034] 10B shows a flowchart of an example video decoding method corresponding to the video encoding method of FIG. 10A, in accordance with some embodiments of the present disclosure. [Figure 10C]

[0035] 1 illustrates a portion of an example VPS syntax structure, according to some embodiments of the present disclosure. [Figure 10D]

[0036] 1 illustrates a portion of an example VPS syntax structure, according to some embodiments of the present disclosure. [Figure 11A]

[0037] 1 shows a flowchart of an exemplary video encoding method in accordance with some embodiments of the present disclosure. [Figure 11B]

[0038] 11B shows a flowchart of an example video decoding method corresponding to the video encoding method of FIG. 11A, in accordance with some embodiments of the present disclosure. [Figure 11C]

[0039] 1 illustrates a portion of an example VPS syntax structure, according to some embodiments of the present disclosure. [Figure 11D]

[0040] 1 illustrates a portion of an example VPS syntax structure, according to some embodiments of the present disclosure. [Figure 12A]

[0041] 1 shows a flowchart of an exemplary video encoding method in accordance with some embodiments of the present disclosure. [Figure 12B]

[0042] 12B shows a flowchart of an example video decoding method corresponding to the video encoding method of FIG. 12A, in accordance with some embodiments of the present disclosure. [Figure 12C]

[0043] 1 illustrates a portion of an example SPS syntax structure, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0021] Detailed Description

[0044] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which like reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations set forth in the following description of exemplary embodiments do not represent all implementations in accordance with the present invention. Rather, they are merely examples of apparatus and methods in accordance with aspects related to the present invention as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.

[0022]

[0045] The ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.

[0023]

[0046] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing technology beyond HEVC using the Joint Search Model (JEM) reference software. Because the coding technology has been incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.

[0024]

[0047] The VVC standard is a recent development and continues to incorporate more coding techniques that result in better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0025]

[0048] A video is a set of static pictures (or "frames") arranged in time sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display such pictures in time sequence. In some applications, a video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as supervision, conferencing, or live broadcasting.

[0026]

[0049] To reduce the storage space and transmission bandwidth required by such applications, video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be performed by software executed by a processor (e.g., a processor in a general-purpose computer) or by specialized hardware. A module for compression is commonly referred to as an “encoder,” and a module for decompression is commonly referred to as a “decoder.” Collectively, the encoder and decoder can be referred to as a “codec.” The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression may be performed by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, or the like. In some applications, a codec may decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard. In this case, the codec may be referred to as a "transcoder."

[0027]

[0050] A video coding process can identify and retain useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such a coding process may be called "lossy." Otherwise, it may be called "lossless." Most coding processes are lossy; this is a tradeoff to reduce the required storage space and transmission bandwidth.

[0028]

[0051] Useful information about the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include changes in pixel position, brightness, or color, of which position changes are the most important. Changes in the position of a group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.

[0029]

[0052] A picture that is coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." A picture that is coded using a previous picture as a reference picture is called a "P-picture." A picture that is coded using both a previous picture and a future picture as a reference picture (i.e., the references are "bidirectional") is called a "B-picture."

[0030]

[0053] 1 illustrates the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video sequence 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.

[0031]

[0054] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with additional pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) need not immediately precede or follow that picture. For example, the reference picture of picture 104 may be a picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the examples shown in FIG.

[0032]

[0055] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, video codecs may divide a picture into elementary segments and encode or decode the picture segment by segment. Such elementary segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units can have variable sizes in pictures, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for a picture based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit.

[0033]

[0056] A basic processing unit may be a logical unit that can include groups of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) that represents colorless luminance information, one or more chroma components (e.g., Cb and Cr) that represent color information, and related syntax elements, where the luma and chroma components may have the same size of the basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTBs") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be performed repeatedly on each of its luma and chroma components.

[0034]

[0057] Video coding has multiple computational stages, examples of which are shown in Figures 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still become too large for processing and therefore may be further divided into segments referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunits may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits may have the same or smaller size than the basic processing units. Similar to the basic processing units, the basic processing subunits are also logical units that can contain groups of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit may be repeatedly performed on each of its luma and chroma components. Note that such division may be performed to further levels as needed for processing. Also, note that different stages may use different schemes to divide the basic processing units.

[0035]

[0058] For example, in a mode decision stage (an example of which is shown in FIG. 2B ), an encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs, as in the case of H.265 / HEVC or H.266 / VVC) and decide the type of prediction for each individual basic processing sub-unit.

[0036]

[0059] As another example, in the prediction stage (an example of which is shown in FIGS. 2A-2B), the encoder may perform prediction operations at the level of basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder may further divide the basic processing subunits into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level the prediction operations may be performed.

[0037]

[0060] As another example, in the transform stage (an example of which is shown in FIGS. 2A-2B), the encoder may perform transform operations for residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder may further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which levels the transform operations may be performed. Note that the division scheme of the same basic processing subunit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0038]

[0061] 1, the basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown as dotted lines. Different basic processing units of the same picture may be divided into basic processing sub-units in different ways.

[0039]

[0062] In some implementations, to provide parallel processing and error resilience capabilities to video encoding and decoding, a picture can be divided into regions for processing, so that the encoding or decoding process does not rely on information about one region of the picture from any other region of the picture. In other words, each region of a picture can be processed independently. Doing so allows a codec to process different regions of a picture in parallel, thereby increasing coding efficiency. Also, when data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thereby providing error resilience. Some video coding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different pictures in video sequence 100 can have different partitioning schemes for dividing the picture into regions.

[0040]

[0063] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing subunits, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.

[0041]

[0064] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 in FIG. 1, video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in temporal order. Similar to structure 110 in FIG. 1, each original picture in video sequence 202 may be divided into basic processing units, basic processing sub-units, or regions for processing by the encoder. In some embodiments, the encoder may perform process 200A at the level of basic processing units for each original picture in video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for a region (eg, regions 114-118) of each original picture of video sequence 202.

[0042]

[0065] 2A , an encoder may provide a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add reconstructed residual BPU 222 to prediction BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0043]

[0066] The encoder may perform process 200A iteratively to encode each original BPU of the original picture (in the forward path) and generate (in the reconstruction path) a prediction reference 224 for encoding the next original BPU of the original picture. After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in video sequence 202.

[0044]

[0067] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, acquiring, obtaining, getting, reading, accessing, or any act by any method for inputting data.

[0045]

[0068] In the prediction step 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224, perform a prediction operation, and generate predicted data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction step 204 is to reduce information redundancy by extracting predicted data 206, which can be used to reconstruct the original BPU as a predicted BPU 208 from the predicted data 206 and the prediction reference 224.

[0046]

[0069] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder can subtract it from the original BPU to generate residual BPU 210. For example, the encoder can subtract pixel values ​​(e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values ​​of the original BPU. Each pixel of residual BPU 210 can have a residual value that is the result of such a subtraction between the corresponding pixel of the original BPU and predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 can have fewer bits, which can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.

[0047]

[0070] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a change frequency (e.g., a frequency of luminance change) component of the residual BPU 210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the changes in the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0048]

[0071] Different transform algorithms can use different basis patterns. For example, various transform algorithms can be used in transform stage 212, such as a discrete cosine transform, a discrete sine transform, or the like. The transform in transform stage 212 is invertible. That is, an encoder can recover residual BPU 210 by inverting the transform (referred to as an "inverse transform"). For example, to recover pixels of residual BPU 210, the inverse transform can multiply the values ​​of corresponding pixels in the basis pattern by their associated coefficients and add the products to generate a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Compared to residual BPU 210, the transform coefficients can have fewer bits, which can be used to reconstruct residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0049]

[0072] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different change frequencies (e.g., luminance change frequencies). Because the human eye is generally better at perceiving low-frequency changes, the encoder can ignore high-frequency change information without significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process can also be inverted, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients in an inverse operation of quantization (referred to as "dequantization").

[0050]

[0073] Because the encoder ignores the remainder of such a division in a rounding operation, the quantization stage 214 may be lossy. Typically, the quantization stage 214 may contribute the greatest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0051]

[0074] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transform in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. The encoder may generate a video bitstream 228 using the output data of the binary encoding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0052]

[0075] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0053]

[0076] It should be noted that other variations of process 200A may also be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages in FIG. 2A.

[0054]

[0077] 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0055]

[0078] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels from one or more already-encoded neighboring BPUs within the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions from one or more already-encoded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include an encoded picture. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0056]

[0079] Referring to process 200B, within the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs within the same picture that are coded (in the forward path) and reconstructed (in the reconstruction path). The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating, for each pixel of the predicted BPU 208, the value of the corresponding pixel. The neighboring BPUs used for extrapolation can be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below-left, below-right, above-left, or above-right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the locations (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientations of the neighboring BPUs used relative to the original BPU, or the like.

[0057]

[0080] As another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference picture. The location of the search window in the reference picture may be determined based on the location of the original BPU of the current picture. For example, the search window may be centered at a location in the reference picture that has the same coordinates as the original BPU in the current picture and may extend outward over a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., by using a pixel-recursive algorithm, a block-matching algorithm, or the like), the encoder can determine such a region as a matching region. The matching region can have different dimensions than the original BPU (e.g., smaller than, equal to, larger than, or a different shape than the original BPU). Because the reference picture and the current picture are temporally separated in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" toward the location of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." When multiple reference pictures are used (e.g., as in picture 106 in FIG. 1), the encoder can search for the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching region in each matching reference picture.

[0058]

[0081] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, zooming, or the like. For inter prediction, prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, or the like.

[0059]

[0082] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift the matching region of the reference picture according to the motion vector, in which case the encoder may predict the original BPU of the current picture. When multiple reference pictures are used (e.g., as in picture 106 in FIG. 1), the encoder may shift the matching region of the reference picture according to each motion vector and average the pixel values ​​of the matching region. In some embodiments, if the encoder weights the pixel values ​​of the matching region of each matching reference picture, the encoder may add a weighted sum of the pixel values ​​to the shifted matching region.

[0060]

[0083] In some embodiments, inter-prediction can be unidirectional or bidirectional. Unidirectional inter-prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which a reference picture (e.g., picture 102) precedes picture 104. Bidirectional inter-prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture in which reference pictures (e.g., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0061]

[0084] Still referring to the forward path of process 200B, after spatial prediction step 2042 and temporal prediction step 2044, in mode decision step 230, the encoder may select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique. In this technique, the encoder may select a prediction mode to minimize the value of a cost function that depends on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.

[0062]

[0085] Within the reconstruction path of process 200B, if an intra-prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current BPU coded and reconstructed in the current picture), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If an inter-prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current picture coded and reconstructed in all BPUs), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by the inter prediction. The encoder can apply various loop filter techniques within the loop filter stage 232, such as deblocking, sample adaptive offset (SAO), adaptive loop filter (ALF), or the like. The loop-filtered reference picture may be stored in a buffer 234 (or "decoded picture buffer (DPB)") for later use (e.g., to be used as an inter-prediction reference picture for a future picture in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in the binary encoding stage 226.

[0063]

[0086] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 in FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, in which case the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0064]

[0087] In FIG. 3A , a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as a “coding BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.

[0065]

[0088] The decoder may perform process 300A iteratively to decode each coded BPU of a coded picture and generate a prediction reference 224 for encoding the next coded BPU of the coded picture. After decoding all coded BPUs of a coded picture, the decoder may output the picture to video stream 304 for display and proceed to decode the next coded picture in video bitstream 228.

[0066]

[0089] In binary decoding stage 302, the decoder may perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if video bitstream 228 is transmitted in packets over a network, the decoder may depacketize video bitstream 228 before providing it to binary decoding stage 302.

[0067]

[0090] 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.

[0068]

[0091] In process 300B, prediction data 206 decoded by the decoder from binary decoding stage 302 for a coding basic processing unit (referred to as the “current BPU”) of a coding picture being decoded (referred to as the “current picture”) may include various types of data, depending on what prediction mode was used by the encoder to encode the current BPU. For example, if intra-prediction was used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-prediction, parameters of the intra-prediction operation, or the like. Parameters of the intra-prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, the orientation of the neighboring BPUs relative to the original BPU, or the like. As another example, if inter-prediction was used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-prediction, parameters of the inter-prediction operation, or the like. Parameters for the inter prediction operation may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, or the like.

[0069]

[0092] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in spatial prediction step 2042 or temporal prediction (e.g., inter prediction) in temporal prediction step 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. The decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in FIG. 3A.

[0070]

[0093] In process 300B, the decoder may provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may provide the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may provide the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer (DPB) in computer memory) for later use (e.g., to be used as an inter-prediction reference picture for a future coded picture of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include loop filter parameters (e.g., loop filter strength). The reconstructed picture from the buffer 234 may also be transmitted to a display, such as a TV, PC, smartphone, or tablet, for viewing by an end user.

[0071]

[0094] FIG. 4 is a block diagram of an exemplary device 400 for encoding or decoding video in accordance with an embodiment of the present disclosure. As shown in FIG. 4, device 400 may include a processor 402. When processor 402 executes instructions described herein, device 400 can become a specialized machine for video encoding or decoding. Processor 402 can be any type of circuitry capable of manipulating or processing information. For example, processor 402 can include any number and combination of a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a system-on-chip (SoC), an application-specific integrated circuit (ASIC), or the like. In some embodiments, processor 402 may be a set of processors grouped as a single logical entity. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0072]

[0095] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for performing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access (e.g., via bus 410) the program instructions and data for processing, execute the program instructions, and perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any number or combination of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, or the like. Memory 404 may also be a group of memories (not shown in FIG. 4) grouped as a single logical entity.

[0073]

[0096] Bus 410 may be a communication device that transfers data between components internal to device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.

[0074]

[0097] For ease of explanation and without ambiguity, the processor 402 and other data processing circuitry will be collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely in hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, stand-alone module or may be fully or partially combined with any other component of the device 400.

[0075]

[0098] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, network interface 406 may include any number and combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near field communication ("NFC") adapter, a cellular network chip, or the like.

[0076]

[0099] In some embodiments, optionally, apparatus 400 may further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., an input interface coupled to a camera or a video archive), or the like.

[0077]

[0100] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that may be loaded into memory 404. As another example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as specialized data processing circuitry (e.g., FPGA, ASIC, NPU, or the like).

[0078]

[0101] Figure 5 is a schematic diagram of an example bitstream 500 encoded by an encoder in accordance with some embodiments of the present disclosure. In some embodiments, the structure of bitstream 500 can be applied to the video bitstream 228 shown in Figures 2A-2B and 3A-3B. In Figure 5, bitstream 500 includes a video parameter set (VPS) 510, a sequence parameter set (SPS) 520, a picture parameter set (PPS) 530, a picture header 540, and slices 550-570, which are separated by synchronization markers M1-M7. Each slice 550-570 includes a corresponding header block (e.g., header 552) and data block (e.g., data 554), and each data block includes one or more CTUs (e.g., CTU1-CTUn in data 554).

[0079]

[0102] According to some embodiments, a bitstream 500, which is a sequence of bits in the form of a network abstraction layer (NAL) unit or byte stream, forms one or more coded video sequences (CVS). A CVS includes one or more coding layer video sequences (CLVS). In some embodiments, a CLVS is a sequence of picture units (PUs), with each PU including one coded picture. Among other things, a PU includes a picture header syntax structure as payload, one coded picture including one or more video coding layer (VCL) NAL units, and zero or one picture header NAL unit (e.g., picture header 540), which optionally includes one or more other non-VCL NAL units. A VCL NAL unit is a generic term for coded slice NAL units (e.g., slices 550-570) and, in some embodiments, a subset of NAL units with a reserved value for the NAL unit type that are classified as VCL NAL units. A coded slice NAL unit includes a slice header and a slice data block (e.g., header 552 and data 554).

[0080]

[0103] In other words, in some embodiments of the present disclosure, a layer may be a set of video coding layer (VCL) NAL units and associated non-VCL NAL units that have a particular value of NAL layer ID. Within these layers, inter-layer prediction can be applied between various layers to achieve high compression performance.

[0081]

[0104] In some embodiments, an output layer set (OLS) may be specified to support decoding of some, but not all, layers. The OLS is a set of layers that includes a designated set of layers, where one or more layers in the set of layers are designated to be output layers. Thus, the OLS may include one or more output layers and other layers necessary to decode the output layer for inter-layer prediction.

[0082]

[0105] In some embodiments, "profiles," "tiers," and "levels" (collectively known as "PTLs") are used to specify constraints on a bitstream and therefore limits on the capabilities required to decode the bitstream. Profiles, tiers, and levels can also be used to indicate interoperability points between individual decoder implementations. A "profile" is a subset of bitstream syntax that specifies a subset of algorithmic features and limits that may be supported by decoders conforming to that profile. Within the boundaries imposed by the syntax of a given profile, it is possible to determine the variation in encoder and decoder performance depending on the values ​​taken by syntax elements in the bitstream, such as the specified size of a decoded picture. In some applications, it may not be practical or economical to implement a decoder that can handle all hypothetical uses of the syntax in a particular profile.

[0083]

[0106] Additionally, "tiers" and "levels" are specified within each profile. A tier's level is a set of specified constraints placed on the values ​​of syntax elements in the bitstream. These constraints may be simple restrictions on values. Alternatively, they may take the form of constraints on arithmetic combinations of values ​​(e.g., picture width x picture height x number of decoded pictures per second). In some embodiments, the same set of tier and level definitions is used for all profiles. Some implementations may support different tiers and different levels within a tier for each supported profile. For a given profile, the tier's level generally corresponds to the processing load and memory capabilities of a particular decoder. The levels specified for lower tiers may be more constrained than the levels specified for higher tiers.

[0084]

[0107] 6A shows an example coding syntax table of a PTL syntax structure 600A signaled in a VPS or SPS, highlighted in bold, according to some embodiments of the present disclosure. In some embodiments, PTL-related information for each operation point identified by parameters TargetOlsIdx and Htid may be indicated by parameters 610A and 620A (e.g., general_profile_idc and general_tier_flag) in the PTL syntax structure 600A signaled in a video parameter set (VPS) or sequence parameter set (SPS) and a parameter 630A (e.g., sublayer_level_idc[Htid]) found in or derived from the PTL syntax structure. TargetOlsIdx is a variable used to identify the OLS index of the target OLS to be decoded, and Htid is a variable used to identify the highest temporal sublayer to be decoded.

[0085]

[0108] As described in the above paragraph, the decoded picture buffer (DPB) includes picture storage buffers for storing decoded pictures. Each picture storage buffer may contain decoded pictures that are marked as "used for reference purposes" or that are kept for future output. In some embodiments, the process is specified to be applied sequentially, starting from the lowest layer in the OLS and increasing order of the nuh_layer_id values ​​of the layers in the OLS, and is applied separately for each layer, where nuh_layer_id is a parameter that specifies the identifier of the layer to which a VCL NAL unit belongs or to which a non-VCL NAL unit applies. In some embodiments, the value of nuh_layer_id should be the same for VCL NAL units of a coded picture or PU. The process includes outputting and removing pictures from the DPB before decoding the current picture, marking and storing the currently decoded picture, and additional bumping.

[0086]

[0109] 6B shows an example encoding syntax table of a DPB parameter syntax structure 600B signaled within a VPS or SPS, highlighted in bold, in accordance with some embodiments of the present disclosure. As shown in FIG. 6B, DPB parameters 610B, 620B, and 630B, which are the parameters necessary to apply the above process and check bitstream conformance, can be found within or derived from the DPB parameter syntax structure 600B signaled within a VPS or SPS. These parameters can include max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid].

[0087]

[0110] In some embodiments, parameter 610B (e.g., max_dec_pic_buffering_minus1[i]) plus 1 specifies the maximum required size of the DPB in units of picture storage buffers when Htid is equal to index i. The value of parameter 610B may be in the range of 0 to (MaxDpbSize-1), inclusive. MaxDpbSize is a parameter that specifies the maximum decoded picture buffer size in units of picture storage buffers. If index i is greater than 0, max_dec_pic_buffering_minus1[i] shall be greater than or equal to max_dec_pic_buffering_minus1[i-1]. If max_dec_pic_buffering_minus1[i] is absent for an index i in the range 0 to (maxSubLayersMinus1-1), then parameter 610B is inferred to be equal to max_dec_pic_buffering_minus1[maxSubLayersMinus1] due to subLayerInfoFlag equal to 0.

[0088]

[0111] Parameter 620B (e.g., max_num_reorder_pics[i]) specifies the maximum allowed number of pictures in the OLS and may precede any picture in the OLS in decoding order and follow that picture in output order if Htid is equal to index i. The value of parameter 620B may be in the range of 0 to max_dec_pic_buffering_minus1[i], inclusive. For index i greater than 0, parameter 620B may be greater than or equal to max_num_reorder_pics[i-1]. If parameter 620B is absent for index i in the range of 0 to maxSubLayersMinus1 minus 1, inclusive, then parameter 620B is inferred to be equal to max_num_reorder_pics[maxSubLayersMinus1] due to subLayerInfoFlag equal to 0.

[0089]

[0112] Parameter 630B can be used to calculate the value of MaxLatencyPictures[i], which specifies the maximum number of pictures in the OLS, that may precede any picture in the OLS in output order and follow that picture in decoding order if Htid is equal to index i. If parameter 630B is equal to 0, the corresponding limit is not expressed. If parameter 630B is not equal to 0, the value of MaxLatencyPictures[i] can be determined based on the following formula: MaxLatencyPictures[i]=max_num_reorder_pics[i]+max_latency_increase_plus1[i]-1 (Formula 1)

[0090]

[0113] The value of parameter 630B must be greater than or equal to 0 (2 32 If parameter 630B is absent for index i in the range 0 to maxSubLayersMinus1 minus 1, inclusive, due to subLayerInfoFlag equal to 0, parameter 630B is inferred to be equal to max_latency_increase_plus1[maxSubLayersMinus1].

[0091]

[0114] A Hypothetical Reference Decoder (HRD) is a hypothetical decoder model that specifies constraints on the diversity of compliant NAL unit streams or compliant byte streams that the encoding process can result in and can be used to check bitstream and decoder conformance. Two types of bitstreams or bitstream subsets are subject to HRD conformance checking for VVC. The first type, named Type I bitstream, is a NAL unit stream that contains VCL NAL units and NAL units with nal_unit_type equal to filler data NAL units (FD_NUT) for all access units (AUs) in the bitstream. The second type, named a Type II bitstream, includes VCL NAL units and filler data NAL units for all AUs in the bitstream, plus at least one of: (1) additional non-VCL NAL units other than filler data NAL units; or (2) all leading_zero_8bits, zero_byte, start_code_prefix_one_3bytes, and trailing_zero_8bits syntax elements that form a byte stream from the NAL unit stream. In some embodiments, leading_zero_8bits is a byte equal to 0x00, zero_byte is a single byte equal to 0x00, start_code_prefix_one_3bytes, called the start code prefix, is a three-byte fixed-value sequence equal to 0x000001, and trailing_zero_8bits is a byte equal to 0x00. In some embodiments, the leading_zero_8bits syntax element is in the first byte stream NAL unit of the bitstream.According to the NAL unit syntax structure, any byte equal to 0x00 (which should be interpreted as a zero_byte syntax element followed by a start_code_prefix_one_3bytes syntax element) preceding the 4-byte sequence 0x00000001 is considered to be a trailing_zero_8bits syntax element that is part of the preceding byte stream NAL unit.

[0092]

[0115] 7A and 7B show example encoding syntax tables for HRD parameter syntax structures 700A and 700B, respectively, highlighted in bold, according to some embodiments of the present disclosure. 700A and 700B may be signaled within a VPS or an SPS. As shown in FIGS. 7A and 7B, two sets of HRD parameters (NAL HRD parameters and VCL HRD parameters) may be used. The HRD parameters may be signaled by the generic HRD parameter syntax structure 700A of FIG. 7A and the OLS HRD parameter syntax structure 700B of FIG. 7B, which may be part of a VPS or an SPS.

[0093]

[0116] As described above, the OLS, PTL syntax structure (e.g., PTL syntax structure 600A in FIG. 6A), DPB parameter syntax structure (e.g., DPB parameter syntax structure 600B in FIG. 6B), and HRD parameter syntax structure (e.g., syntax structures 700A and 700B in FIGS. 7A and 7B) may be signaled within a VPS or SPS. For each of these syntax structures, the VPS or SPS signals the set of syntax structures that apply to each OLS and the index of the syntax structure.

[0094]

[0117] FIG. 8 illustrates an example encoding syntax table of a portion of a VPS Raw Byte Sequence Payload (RBSP) syntax structure 800 signaled within a VPS, highlighted in bold, in accordance with some embodiments of the present disclosure. As shown in FIG. 8, VPS parameter 810 (vps_max_layers_minus1) plus 1 specifies the maximum allowable number of layers within each CVS that references the VPS. VPS parameter 812 (vps_max_sublayers_minus1) plus 1 specifies the maximum number of temporal sublayers allowed within a layer within each CVS that references the VPS. In some embodiments, the value of VPS parameter 812 is in the range of 0 to a predefined static value (e.g., 6). VPS parameter 814 (each_layer_is_an_ols_flag) equal to 1 indicates that each OLS contains one layer and that each layer in a CVS that references the VPS is itself an OLS with a single contained layer that is the only output layer. VPS parameter 814 equal to 0 indicates that the OLS may contain multiple layers. If the VPS parameter 810 is equal to 0, the value of the VPS parameter 814 is inferred to be equal to 1. In some embodiments, if the corresponding flag specifies that one or more of the layers specified by the VPS can use inter-layer prediction (e.g., if vps_all_independent_layers_flag is equal to 0), the value of the VPS parameter 814 is inferred to be equal to 0.

[0095]

[0118] In some embodiments, the value of the VPS parameter 816 (ols_mode_idc) can be in the range of 0 to 2, inclusive, with a value of 3 for the VPS parameter 816 being reserved for future use. In some embodiments, a VPS parameter 816 equal to 0 specifies that the total number of OLSs (TotalNumOlss) specified by the VPS is equal to the value of the VPS parameter 810 plus 1, the i-th OLS includes layers with layer indices from 0 to i, and for each OLS, only the top layer in the OLS is output. A VPS parameter 816 equal to 1 specifies that the total number of OLSs specified by the VPS is equal to the VPS parameter 810 plus 1, the i-th OLS includes layers with layer indices from 0 to i, and for each OLS, all layers in the OLS are output. A VPS parameter 816 equal to 2 specifies that the total number of OLSs specified by the VPS is explicitly signaled using the VPS parameter 818, the layer that is output for each OLS is explicitly signaled, and the other layers are direct or indirect reference layers of the output layer of the OLS.

[0096]

[0119] In some embodiments, if a corresponding flag specifies that all layers specified by the VPS are coded independently without using inter-layer prediction (e.g., vps_all_independent_layers_flag is equal to 1) and the VPS parameter 814 is equal to 0, the value of the VPS parameter 816 is inferred to be equal to 2.

[0097]

[0120] When the VPS parameter 816 is equal to 2, the VPS parameter 818 (num_output_layer_sets_minus1) is signaled to indicate the total number of OLSs minus 1. In other words, the VPS parameter 818 plus 1 specifies the total number of OLSs (e.g., the variable TotalNumOlss) specified by the VPS when the VPS parameter 816 is equal to 2. The variable TotalNumOlss can be derived and calculated from the code as follows: if(vps_max_layers_minus1==0) TotalNumOlss=1 else if(each_layer_is_an_ols_flag||ols_mode_idc==0||ols_mode_idc==1) TotalNumOlss=vps_max_layers_minus1+1 else if(ols_mode_idc==2) TotalNumOlss=num_output_layer_sets_minus1+1

[0098]

[0121] Additionally, if VPS parameter 816 is equal to 2, then a VPS parameter 820 (ols_output_layer_flag[i][j]) equal to 1 specifies that the jth layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[j]) is the output layer of the i-th OLS, and a VPS parameter 820 equal to 0 specifies that the jth layer is not an output layer of the i-th OLS. In other words, each OLS can be defined for each layer in a CVS to indicate whether the layer is an output layer or not within that OLS by signaling a corresponding flag (e.g., VPS parameter 820).

[0099]

[0122] The VPS parameter 822 (vps_num_ptls_minus1) plus 1 specifies the number of PTL syntax structures in the VPS. The value of the VPS parameter 822 can be less than the total number of OLSs (i.e., TotalNumOlss).

[0100]

[0123] A VPS parameter 824 (pt_present_flag[i]) equal to 1 specifies that profile, tier, and general constraint information is present in the i-th PTL syntax structure in the VPS. A VPS parameter 824 equal to 0 specifies that profile, tier, and general constraint information is not present in the i-th PTL syntax structure in the VPS. In some embodiments, a value of pt_present_flag[0] is inferred to be equal to 1. When VPS parameter 824 is equal to 0, the profile, tier, and general constraint information for the i-th PTL syntax structure in the VPS is inferred to be the same as that of the (i-1)-th PTL syntax structure in the VPS.

[0101]

[0124] The VPS parameter 826 (ptl_max_temporal_id[i]) specifies the temporal identifier (TemporalId) of the highest sub-layer representation for which level information is found in the i-th PTL syntax structure within the VPS. The value of the VPS parameter 826 can be in the range of 0 or greater and 0 or less than the VPS parameter 812. If the VPS parameter 812 is equal to 0, the value of the VPS parameter 826 is inferred to be equal to 0. If the VPS parameter 812 is greater than 0 and the corresponding flag (e.g., vps_all_layers_same_num_sublayers_flag) is equal to 1, the value of the VPS parameter 826 is inferred to be equal to the VPS parameter 812.

[0102]

[0125] In some embodiments, VPS parameter 828 (vps_ptl_alignment_zero_bit) may be equal to 0. In some embodiments, VPS parameter 830 (ols_ptl_idx[i]) specifies the index into the list of PTL syntax structures in the VPS of the PTL syntax structure that applies to the i-th OLS. If present, the value of VPS parameter 830 may be in the range greater than or equal to 0 and less than or equal to VPS parameter 822. If VPS parameter 828 is equal to 0, the value of VPS parameter 830 is inferred to be equal to 0.

[0103]

[0126] If NumLayersInOls[i], a variable specifying the number of layers in the i-th OLS, is equal to 1, then the PTL syntax structures that apply to the i-th OLS are also in the SPS referenced by the layers in the i-th OLS. In some embodiments, it is a bitstream compatibility requirement that when the number of layers in the i-th OLS is equal to 1, the PTL syntax structures signaled in the VPS and in the SPS for the i-th OLS can be identical.

[0104]

[0127] The VPS parameter 832 (vps_num_dpb_params) specifies the number of DPB parameter syntax structures in the VPS. In some embodiments, the value of the VPS parameter 832 can be in the range of 0 to 16, inclusive. If absent, the value of the VPS parameter 832 is inferred to be equal to 0.

[0105]

[0128] The VPS parameter 834 (vps_sublayer_dpb_params_present_flag) is used to control the presence of syntax elements (e.g., max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[]) in the DPB parameter syntax structure in the VPS. If not present, the VPS parameter 834 is inferred to be equal to 0.

[0106]

[0129] The VPS parameter 836 (dpb_max_temporal_id[i]) specifies the TemporalId of the highest sublayer representation at which a DPB parameter may be in the i-th DPB parameter syntax structure within the VPS. The value of the VPS parameter 836 may be in the range of 0 or greater and 0 or less than the VPS parameter 812. If the VPS parameter 812 is equal to 0, the value of the VPS parameter 836 is inferred to be equal to 0. If the VPS parameter 812 is greater than 0 and the corresponding flag (e.g., vps_all_layers_same_num_sublayers_flag) is equal to 1, the value of the VPS parameter 836 is inferred to be equal to the VPS parameter 812.

[0107]

[0130] VPS parameter 838 (ols_dpb_pic_width[i]) and VPS parameter 840 (ols_dpb_pic_height[i]) specify the width and height, respectively, in luma samples of each picture storage buffer for the i-th OLS.

[0108]

[0131] The VPS parameter 842 (ols_dpb_params_idx[i]) specifies an index into the list of DPB parameter syntax structures in the VPS for the DPB parameter syntax structures that apply to the i-th OLS when NumLayersInOls[i] is greater than 1. If present, the value of the VPS parameter 842 can be in the range of 0 to VPS parameter 832 minus 1, inclusive. If the VPS parameter 842 is absent, the value of the VPS parameter 842 is inferred to be equal to 0. In some embodiments, if NumLayersInOls[i] is equal to 1, the DPB parameter syntax structures that apply to the i-th OLS are in the SPS referenced by the layers in the i-th OLS.

[0109]

[0132] A VPS parameter 844 (vps_general_hrd_params_present_flag) equal to 1 specifies that the general HRD parameter syntax structure (e.g., syntax structure 700A of FIG. 7A) and other HRD parameters are present in the VPS RBSP syntax structure. A VPS parameter 844 equal to 0 specifies that the general HRD parameter syntax structure and other HRD parameters are not present in the VPS RBSP syntax structure. If not, the value of VPS parameter 844 is inferred to be equal to 0. If NumLayersInOls[i] is equal to 1, the general HRD parameter syntax structure that applies to the ith OLS is in the SPS referenced by the layer in the ith OLS.

[0110]

[0133] A VPS parameter 846 (vps_sublayer_cpb_params_present_flag) equal to 1 specifies that the i-th OLS HRD parameters syntax structure in the VPS (e.g., syntax structure 700B of FIG. 7B) includes HRD parameters for a sublayer representation with a TemporalId in the range greater than or equal to 0 and less than or equal to the value of VPS parameter 850. A VPS parameter 846 equal to 0 specifies that the i-th OLS HRD parameters syntax structure in the VPS includes HRD parameters for a sublayer representation with a TemporalId equal to the value of VPS parameter 850 only. If VPS parameter 812 is equal to 0, the value of VPS parameter 846 is inferred to be equal to 0.

[0111]

[0134] If the VPS parameter 846 is equal to 0, then the HRD parameters for sublayer representations with TemporalIds in the range of 0 to the value of the VPS parameter 850 minus 1 are inferred to be the same as those for sublayer representations with TemporalIds equal to the VPS parameter 850. In some embodiments, under the conditional statement "if(general_vcl_hrd_params_present_flag)" in the OLS HRD parameter syntax structure 700B of FIG. 7B , the HRD parameters are inferred to be the same as those for sublayer representations with TemporalIds starting from the fixed_pic_rate_general_flag[i] syntax element up to and including the sublayer_hrd_parameters(i) syntax structure.

[0112]

[0135] The VPS parameter 848 (num_ols_hrd_params_minus1) plus 1 specifies the number of OLS HRD parameter syntax structures within the generic HRD parameter syntax structure when the VPS parameter 844 is equal to 1. The value of the VPS parameter 848 can be in the range of 0 or greater and the value of TotalNumOlss minus 1 or less.

[0113]

[0136] The VPS parameter 850 (hrd_max_tid[i]) specifies the TemporalId of the highest sublayer representation whose HRD parameters are included in the i-th OLS HRD parameter syntax structure. The value of the VPS parameter 850 can be in the range of 0 or greater and 0 or less than the VPS parameter 812. If the VPS parameter 812 is equal to 0, the value of the VPS parameter 850 is inferred to be equal to 0. If the VPS parameter 812 is greater than 0 and the corresponding flag (e.g., vps_all_layers_same_num_sublayers_flag) is equal to 1, the value of the VPS parameter 850 is inferred to be equal to the VPS parameter 812.

[0114]

[0137] The VPS parameter 852 (ols_hrd_idx[i]) specifies an index into the list of OLS HRD parameter syntax structures in the VPS for the OLS HRD parameter syntax structures that apply to the i-th OLS when NumLayersInOls[i] is greater than 1. The value of the VPS parameter 852 can be in the range of 0 to VPS parameter 848, inclusive. If NumLayersInOls[i] is equal to 1, the OLS HRD parameter syntax structures that apply to the i-th OLS are in the SPS referenced by the layer in the i-th OLS. In some embodiments, if the value of VPS parameter 848 plus 1 equals TotalNumOlss, the value of VPS parameter 852 is inferred to be equal to index i. In some embodiments, if NumLayersInOls[i] is greater than 1 and VPS parameter 848 is equal to 0, the value of VPS parameter 852 is inferred to be equal to 0.

[0115]

[0138] When a sequence parameter set (SPS) layer is an independent layer, the PTL parameter syntax structure, the DPB parameter syntax structure, and the HRD parameter syntax structure can be signaled within the SPS.

[0116]

[0139] 9 shows an example coding syntax table of a portion of an SPS RBSP syntax structure 900 signaled within an SPS, highlighted in bold, in accordance with some embodiments of the present disclosure. As shown in FIG. 9, an SPS parameter 910 (sps_seq_parameter_set_id) provides an identifier for the SPS for reference by other syntax elements. Regardless of the value of nuh_layer_id, SPS NAL units share the same value space of the SPS parameter 910. Let spsLayerId be the value of nuh_layer_id for a particular SPS NAL unit, and let vclLayerId be the value of nuh_layer_id for a particular VCL NAL unit. A particular VCL NAL unit cannot reference a particular SPS NAL unit unless spsLayerId is less than or equal to vclLayerId and a layer with nuh_layer_id equal to spsLayerId is included in at least one OLS that includes a layer with nuh_layer_id equal to vclLayerId.

[0117]

[0140] The SPS parameter 912 (sps_video_parameter_set_id) specifies the value of vps_video_parameter_set_id of the VPS referenced by the SPS when the SPS parameter 912 is greater than 0. In some embodiments, if the SPS parameter 912 is equal to 0, the corresponding SPS does not reference a VPS, and the VPS is not referenced when decoding each CLV that references the SPS. Additionally, the value of the corresponding VPS parameter 810 is inferred to be equal to 0, the CLV contains only one layer (i.e., the VCL NAL units in the CLV have the same value of nuh_layer_id), the value of GeneralLayerIdx[nuh_layer_id] is inferred to be equal to 0, and the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1. If vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, then an SPS referenced by CLVS with a particular nuh_layer_id value, nuhLayerId, can have nuh_layer_id equal to nuhLayerId. In some embodiments, the value of SPS parameter 912 can be the same within an SPS referenced by CLVS within a CVS.

[0118]

[0141] The SPS parameter 914 (sps_max_sublayers_minus1) plus 1 specifies the maximum number of temporal sublayers that may be in each CLVS that references the SPS. The value of the SPS parameter 914 may range from 0 to the corresponding VPS parameter 812, inclusive.

[0119]

[0142] In some embodiments, the SPS parameter 916 (sps_reserved_zero_4bits) may be equal to 0 in the bitstream. Other values ​​of the SPS parameter 916 may be reserved for future use.

[0120]

[0143] An SPS parameter 918 (sps_ptl_dpb_hrd_params_present_flag) equal to 1 specifies that the PTL syntax structure and the DPB parameter syntax structure are in the SPS, and that the general HRD parameter syntax structure and the OLS HRD parameter syntax structure may also be in the SPS. An SPS parameter 918 equal to 0 specifies that none of these four syntax structures are in the SPS. The value of the SPS parameter 918 may be equal to vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]].

[0121]

[0144] The SPS parameter 920 (sps_sublayer_dpb_params_flag) is used to control the presence of syntax elements (e.g., max_dec_pic_buffering_minus1[i], max_num_reorder_pics[i], and max_latency_increase_plus1[i]) in the DPB parameter syntax structure in the SPS. If not present, the value of the SPS parameter 920 can be inferred to be equal to 0.

[0122]

[0145] An SPS parameter 922 (sps_general_hrd_params_present_flag) equal to 1 specifies that the SPS contains the general HRD parameter syntax structure and the OLS HRD parameter syntax structure. An SPS parameter 922 equal to 0 specifies that the SPS does not contain the general HRD parameter syntax structure or the OLS HRD parameter syntax structure.

[0123]

[0146] An SPS parameter 924 (sps_sublayer_cpb_params_present_flag) equal to 1 specifies that the OLS HRD parameter syntax structures within the SPS include HRD parameters for sublayer representations with TemporalIds in the range greater than or equal to 0 and less than or equal to SPS parameter 914. An SPS parameter 924 equal to 0 specifies that the OLS HRD parameter syntax structures within the SPS include HRD parameters for sublayer representations with TemporalIds equal to SPS parameter 914 only. If SPS parameter 914 is equal to 0, the value of SPS parameter 924 is inferred to be equal to 0. If SPS parameter 924 is equal to 0, the HRD parameters for sublayer representations with TemporalIds in the range greater than or equal to 0 and less than or equal to SPS parameter 914 minus 1 are inferred to be the same as those for sublayer representations with TemporalIds equal to SPS parameter 914. In some embodiments, under the conditional statement "if(general_vcl_hrd_params_present_flag)" in the OLS HRD parameters syntax structure 700B of FIG. 7B, these HRD parameters begin with the fixed_pic_rate_general_flag[i] syntax element and include up to and including the sublayer_hrd_parameters(i) syntax structure.

[0124]

[0147] There are some issues regarding the signaling of the PTL, DPB and HRD parameters within the VPS or SPS and the signaling of the OLS.

[0125]

[0148] For example, as shown in FIG. 8, to signal DPB and HRD parameters, the number of syntax structures (VPS parameter 832 for DPB parameters and VPS parameter 848 for HRD parameters) and the index of the syntax structure applied to the i-th OLS (VPS parameter 842 for DPB parameters and VPS parameter 852 for HRD parameters) are signaled in the VPS by the "ue(v)" coding method, which is a variable-length coding method in which small values ​​are coded with fewer bits and large values ​​are coded with more bits. On the other hand, to signal PTL parameters, the VPS parameter 822, which is the number of syntax structures, and the VPS parameter 830, which is the index of the syntax structure applied to the i-th OLS, are coded by "u(8)," which is a fixed-length coding method that uses 8 bits for values ​​in the range of 0 to 255.

[0126]

[0149] Therefore, the number of syntax structures and the coding method of the syntax structure indexes between the DPB, HRD, and PTL parameters may differ. In addition, in many cases, the number of PTL syntax structures is relatively small in practical applications. Using 8 bits for these syntax elements (e.g., VPS parameters 822 and 830) unnecessarily increases signaling overhead.

[0127]

[0150] Furthermore, to specify which PTL, DPB, and HRD syntax structures apply to each OLS, an index (e.g., VPS parameter 830, 842, or 852) is signaled within the VPS for each OLS. However, if the number of OLSs is equal to the number of syntax structures, an efficient encoder can perform a one-to-one mapping from syntax structures to OLSs and avoid wasting bits resulting from unused syntax structures. Therefore, the encoder can signal the syntax structures in the order of the OLSs to which they apply without signaling the index. At the decoder side, it can be inferred that the i-th syntax structure applies to the i-th OLS. Skipping the signaling of the index can reduce the number of bits and thus improve coding efficiency. In some embodiments, this mechanism is used for the HRD parameters. In some embodiments, this mechanism is also used for the PTL and DPB parameters, preventing increased signaling overhead or inconsistent signaling designs.

[0128]

[0151] Furthermore, in an SPS, SPS parameter 920, which is signaled when SPS parameter 914 is greater than 0, controls the presence of syntax elements including parameters 610B, 620B, and 630B in DPB parameter syntax structure 600B within the SPS. SPS parameter 920. When SPS parameter 918 is equal to 0, the DPB parameter syntax structure is not signaled. In other words, when SPS parameter 914 is greater than 0 and SPS parameter 918 is equal to 0, signaling of SPS parameter 920 is redundant.

[0129]

[0152] 8, the VPS parameter 818 used to specify the total number of OLSs is fixed-length coded with a code length of 8. Therefore, the maximum value of the VPS parameter 818 is 255. An OLS is defined for each layer in the CVS by the VPS parameter 820, and the total number of OLSs is less than or equal to the number of combinations of the VPS parameter 820. If the maximum number of layers in the CVS (i.e., the value of the VPS parameter 810 plus 1) is less than 8, using 8-bit fixed-length coding for the VPS parameter 818 may be unnecessary.

[0130]

[0153] 10A shows a flowchart of an exemplary video encoding method 1000A according to some embodiments of the present disclosure. In some embodiments, the video encoding method 1000A may be performed by an encoder (e.g., an encoder performing process 200A of FIG. 2A or process 200B of FIG. 2B) to generate the bitstream 500 shown in FIG. 5. For example, the encoder may be implemented as one or more software or hardware components of an apparatus (e.g., apparatus 400 in FIG. 4) for encoding or transcoding a video sequence (e.g., video sequence 202 of FIG. 2A or 2B) to generate a bitstream for the video sequence (e.g., video bitstream 228 of FIG. 2A or 2B). For example, a processor (e.g., processor 402 of FIG. 4) may perform the video encoding method 1000A.

[0131]

[0154] Referring to video encoding method 1000A, in step 1010a, an encoder encodes one or more PTL syntax elements (e.g., VPS parameters 822 or VPS parameters 830 of FIG. 8) that specify PTL-related information having a variable length. In step 1020a, the encoder signals the one or more PTL syntax elements having a variable length within a VPS (e.g., VPS 510 of FIG. 5) or SPS (e.g., SPS 520 of FIG. 5) of a bitstream (e.g., bitstream 500 of FIG. 5).

[0132]

[0155] FIG. 10B shows a flowchart of an example video decoding method 1000B corresponding to the video encoding method 1000A of FIG. 10A, according to some embodiments of the present disclosure. In some embodiments, the video decoding method 1000B may be performed by a decoder (e.g., a decoder performing the decoding process 300A of FIG. 3A or the decoding process 300B of FIG. 3B) to decode the bitstream 500 of FIG. 5. For example, the decoder may be implemented as one or more software or hardware components of a device (e.g., device 400 of FIG. 4) for decoding a bitstream (e.g., video bitstream 228 of FIG. 3A or 3B) to reconstruct a video stream of the bitstream (e.g., video stream 304 of FIG. 3A or 3B). For example, a processor (e.g., processor 402 of FIG. 4) may perform the video decoding method 1000B. 10B, within video decoding method 1000B, in step 1010b, a decoder receives a bitstream (e.g., bitstream 500 of FIG. 5) containing a VPS or SPS to be decoded. In step 1020b, the decoder decodes one or more PTL syntax elements that specify PTL-related information having a variable length within the VPS or SPS.

[0133]

[0156] 10C and 10D show example VPS syntax structures 1000C and 1000D, respectively, according to some embodiments of the present disclosure. VPS syntax structures 1000C and 1000D may be used in methods 1000A and 1000B, respectively. VPS syntax structures 1000C and 1000D are modified based on syntax structure 800 of FIG. 8.

[0134]

[0157] As shown in FIG. 10C , in some embodiments, PTL syntax elements encoded or decoded using variable length encoding may include a first PTL syntax element (e.g., VPS parameter 830) that specifies an index of a PTL syntax structure, or a second PTL syntax element (e.g., VPS parameter 822) that specifies the number of PTL syntax structures in a VPS or SPS. The first PTL syntax element and the second PTL syntax element may be encoded or decoded using exponential-Golomb coding. For example, the number of PTL syntax structures (e.g., VPS parameter 822) and the index of a PTL syntax structure (e.g., VPS parameter 830) may be encoded using the “ue(v)” encoding method, which is a zeroth-order exponential-Golomb code. From the perspective of design consistency, signaling PTL parameters using the “ue(v)” encoding method allows PTL parameters to be signaled in the same manner as DPB and HDR parameters. As shown in Figure 10C, descriptors 822d and 830d indicate that the encoding method of VPS parameters 822 and 830 is changed from u(8) to ue(v), respectively. The semantics shown in syntax structure 1000C of Figure 10C (left column of the table) are unchanged and therefore the same as in syntax structure 800 of Figure 8.

[0135]

[0158] As shown in syntax structure 1000D of Figure 10D, in some other embodiments, when encoding a first PTL syntax element (e.g., VPS parameters 830), the length of the first PTL syntax element can be set to be the smallest integer greater than or equal to the logarithm to the base 2 of the number of PTL syntax structures in the VPS or SPS. In some other embodiments, when encoding a second PTL syntax element (e.g., VPS parameters 822), the length of the second PTL syntax element can be a fixed length, such as u(8).

[0136]

[0159] For example, an index of a PTL syntax structure (e.g., VPS parameters 830) can be encoded using the "u(v)" encoding method, which does not have a fixed number of bits. The number of bits depends on the values ​​of other syntax elements, such as the value associated with the number of PTL syntax structures (e.g., VPS parameters 822), which are still encoded in u(v). Thus, after parsing the value of VPS parameters 822, the length of VPS parameters 830 is calculated as Ceil(log2(vps_num_ptls_minus1+1)), where log2(x) is the base 2 logarithm of x and Ceil(x) is the smallest integer greater than or equal to x. Then, VPS parameters 830 are parsed with Ceil(log2(vps_num_ptls_minus1+1)) bits. Specifically, for VPS parameters 830, the encoding method can be changed from u(v) to u(v). 8, if present, the value of VPS parameter 830 may range from greater than or equal to 0 to less than or equal to VPS parameter 822. If VPS parameter 828 is equal to 0, then the value of VPS parameter 830 is inferred to be equal to 0.

[0137]

[0160] In some embodiments, if the number of PTL or DPB syntax structures is equal to the number of OLSs, the index (e.g., VPS parameters 830 and 842) can be inferred without signaling. By eliminating the signaling of the index, signaling costs can be reduced.

[0138]

[0161]

[0033] Figure 11A shows a flowchart of an exemplary video encoding method 1100A according to some embodiments of the present disclosure. Figure 11B shows a flowchart of an exemplary video decoding method 1100B corresponding to the video encoding method 1100A of Figure 11A according to some embodiments of the present disclosure. Similar to methods 1000A and 1000B of Figures 10A and 10B, the video encoding method 1100A and the video decoding method 1100B may be performed by an encoder and a decoder implemented as one or more software or hardware components of a device (e.g., a processor executing a set of instructions).

[0139]

[0162] 11A, in step 1110a, the encoder determines whether the number of PTL syntax structures in the VPS (vps_num_ptls_minus1+1) is equal to the number of OLSs (TotalNumOlss). In other words, the encoder determines whether the coded video sequence (CVS) contains an equal number of PTL syntax structures and OLSs.

[0140]

[0163] In response to the CVS containing an equal number of PTL syntax structures and OLSs (step 1110a - yes), the encoder bypasses steps 1120a and 1130a and encodes the bitstream without signaling the first PTL syntax element (e.g., ols_ptl_idx[i]) that specifies the index into the list of PTL syntax structures in the VPS of the PTL syntax structure that applies to the i-th OLS.

[0141]

[0164] In response to the number of PTL syntax structures being different from the number of OLSs (step 1110a—no), in step 1120a, the encoder determines whether the number of PTL syntax structures is equal to 1 (e.g., by determining whether parameter vps_num_ptls_minus1 is greater than zero). In response to the number of PTL syntax structures being less than or equal to 1 (step 1120a—no), the encoder bypasses step 1130a and encodes the bitstream without signaling the first PTL syntax element in the VPS.

[0142]

[0165] In response to the number of PTL syntax structures being greater than one and different from the number of OLSs (step 1110a—no, step 1120a—yes), the encoder performs step 1130a, signaling a first PTL syntax element within the VPS (or SPS). In some embodiments, the first PTL syntax element is signaled with a fixed length.

[0143]

[0166] In some other embodiments, step 1120a is performed before step 1110a. In response to the number of PTL syntax structures being greater than one and different from the number of OLSs (step 1120a—yes, step 1110a—no), step 1130a is performed; otherwise, step 1130a is skipped. For example, in response to the number of PTL syntax structures being equal to one (step 1120a—no), the encoder may bypass both steps 1110a and 1130a.

[0144]

[0167] Similar to the coding of PTL syntax elements, the coding of indices (eg, VPS parameters 842) for DPB syntax elements can be inferred without signaling if certain conditions are met.

[0145]

[0168] In step 1140a, the encoder determines whether the number of DPB parameter syntax structures (vps_num_dpb_params) in the VPS is the same as the number of OLSs (TotalNumOlss). In other words, the encoder determines whether the coded video sequence (CVS) includes an equal number of DPB parameter syntax structures and OLSs.

[0146]

[0169] In response to the number of DPB parameter syntax structures being equal to the number of OLSs (step 1140a - yes), the encoder bypasses steps 1150a and 1160a and encodes the bitstream without signaling the first DPB syntax element (e.g., ols_dpb_params_idx[i]) that specifies the index into the list of DPB parameter syntax structures in the VPS of the DPB parameter syntax structure that applies to the i-th OLS.

[0147]

[0170] In response to the number of DPB parameter syntax structures being different from the number of OLSs (step 1140a—no), in step 1150a, the encoder determines whether the number of DPB parameter syntax structures is equal to 1 (e.g., by determining whether the parameter vps_num_dpb_params is greater than 1). In response to the number of DPB parameter syntax structures being less than or equal to 1 (step 1150a—no), the encoder bypasses step 1160a and encodes the bitstream without signaling the first DPB syntax element.

[0148]

[0171] In response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs (step 1140a—no, step 1150a—yes), the encoder performs step 1160a and signals a first DPB syntax element in the VPS. In some embodiments, the first DPB syntax element is signaled with a variable length using ue(v).

[0149]

[0172] In some embodiments, step 1150a is performed before step 1140a. In response to the number of DPB syntax structures being greater than one and different from the number of OLSs (step 1150a—yes, step 1140a—no), step 1160a is performed; otherwise, step 1160a is skipped. For example, in response to the number of DPB syntax structures being equal to one (step 1150a—no), the encoder may bypass both steps 1140a and 1160a.

[0150]

[0173] Referring to video decoding method 1100B shown in Figure 11B, on the decoder side, in step 1110b, the decoder receives a bitstream (e.g., video bitstream 500 of Figure 5) including a coded video sequence (CVS). In step 1120b, the decoder determines whether the number of PTL syntax structures in the VPS (vps_num_ptls_minus1+1) is the same as the number of OLSs (TotalNumOlss). In some embodiments, the decoder decodes one or more VPS syntax elements associated with the OLSs to obtain information about the number of OLSs (TotalNumOlss).

[0151]

[0174] In response to the number of PTL syntax structures being the same as the number of OLSs (step 1120b—yes), the decoder infers in step 1125b that the first PTL syntax element (e.g., ols_ptl_idx[i]) that specifies the index into the list of PTL syntax structures in the VPS of the PTL syntax structure that applies to the i-th OLS is equal to the sequence number (e.g., index i) of the i-th OLS. In response to the number of PTL syntax structures being different from the number of OLSs (step 1120b—no), in step 1130b the decoder determines whether the number of PTL syntax structures is equal to 1 (e.g., by determining whether parameter vps_num_ptls_minus1 is greater than 0). In response to the number of PTL syntax structures being one or less (step 1130b—no), the decoder infers in step 1135b that the first PTL syntax element is zero.

[0152]

[0175] In response to the number of PTL syntax structures being greater than one and different from the number of OLSs (step 1120b—no, step 1130b—yes), the decoder performs step 1140b, decoding the first PTL syntax element encoded in the VPS or SPS. In some embodiments, the first PTL syntax element is signaled with a fixed length in the VPS or SPS.

[0153]

[0176] In some embodiments, step 1130b is performed before step 1120b. In response to the number of PTL syntax structures being greater than one and different from the number of OLSs (step 1130b—yes, step 1120b—no), step 1140a is performed; otherwise, step 1140a is skipped.

[0154]

[0177] Similarly, in step 1150b, the decoder determines whether the number of DPB parameter syntax structures in the VPS (vps_num_dpb_params) and the number of OLSs (TotalNumOlss) are the same.

[0155]

[0178] In response to the number of DPB parameter syntax structures being the same as the number of OLSs (step 1150b—Yes), in step 1155b, the decoder infers that the first DPB syntax element (e.g., ols_dpb_params_idx[i]) that specifies the index into the list of DPB parameter syntax structures in the VPS of the DPB parameter syntax structure that applies to the i-th OLS is equal to the sequence number (e.g., index i) of the i-th OLS.

[0156]

[0179] In response to the number of DPB parameter syntax structures being different from the number of OLSs (step 1150b—no), in step 1160b, the decoder determines whether the number of DPB parameter syntax structures is equal to 1 (e.g., by determining whether the parameter vps_num_dpb_params is greater than 1). In response to the number of DPB parameter syntax structures being less than or equal to 1 (step 1160b—no), in step 1165b, the decoder infers the first DPB syntax element to be zero.

[0157]

[0180] In response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs (step 1150b—no; step 1160b—yes), the decoder performs step 1170b and decodes the first DPB syntax element in the VPS. In some embodiments, the first DPB syntax element is signaled with a variable length using ue(v).

[0158]

[0181] In some embodiments, step 1160b is performed before step 1150b. In response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs (step 1160b—yes, step 1150b—no), step 1170b is performed; otherwise, step 1170b is skipped.

[0159]

[0182] 11C illustrates a portion of an example VPS syntax structure 1100C related to a possible implementation of the proposed methods 1100A and 1100B according to some embodiments of the present disclosure. The VPS syntax structure 1100C in FIG. 11C is modified based on the syntax structure 800 in FIG. 8.

[0160]

[0183] 8, as shown in conditional statement 1110c of VPS syntax structure 1100C, VPS parameter 830 is signaled if the value of VPS parameter 822 (vps_num_ptls_minus1) plus 1 is not equal to the total number of OLSs specified by the VPS (TotalNumOlss) and is not equal to 0. If the value of VPS parameter 822 plus 1 is equal to TotalNumOlss, the value of VPS parameter 830 is inferred to be equal to index i. If VPS parameter 822 is equal to 0, the value of VPS parameter 830 is inferred to be equal to 0.

[0161]

[0184] Similarly, as shown in conditional statement 1120c of VPS syntax structure 1100C, VPS parameter 842 (ols_dpb_params_idx[i]) is signaled if VPS parameter 832 is not equal to TotalNumOlss and is not equal to 0. When VPS parameter 842 is not signaled, if the value of VPS parameter 832 is equal to TotalNumOlss, the value of VPS parameter 842 is inferred to be index i. If VPS parameter 832 is equal to 0, the value of VPS parameter 842 is inferred to be equal to 0.

[0162]

[0185] As described above, during the encoding or decoding process, an encoder or decoder may need to determine the number of OLSs (TotalNumOlss) in a VPS. In some embodiments, decoders and encoders can derive the number of OLSs according to the OLS mode indicated by a VPS syntax element (e.g., VPS parameters 816) and the maximum number of layers in a CVS indicated by another VPS syntax element (e.g., VPS parameters 810). In some embodiments, an encoder can encode a VPS syntax element (e.g., VPS parameters 818) having a fixed length related to the number of OLSs specified by the VPS, and a decoder can decode the VPS syntax element having the fixed length to determine the number of OLSs specified by the VPS. In some embodiments, an encoder can encode a VPS syntax element (e.g., VPS parameters 818) having a variable length related to the number of OLSs included in a CVS that references the VPS, and a decoder can decode the VPS syntax element having the variable length to determine the number of OLSs included in a CVS that references the VPS. If the maximum allowable number of layer combinations in the CVS that reference the VPS is less than the default length value, the length of the VPS syntax element may be equal to the maximum allowable number.

[0163]

[0186] 11D illustrates a portion of an example VPS syntax structure 1100D that is modified based on the syntax structure 800 of FIG. 8 and relates to a possible implementation for encoding or decoding VPS syntax elements with variable lengths, according to some embodiments of the present disclosure. In some cases, the total number of OLSs is limited by the number of combinations of VPS parameters 820 (ols_output_layer_flag) of all layers in the CVS, and TotalNumOlss satisfies the following inequality: TotalNumOlss≦2 vps_max_layers_minus1+1

[0164]

[0187] When VPS parameter 816 is equal to 2, VPS parameter 818 plus 1 specifies the total number of OLSs specified by the VPS. Thus, VPS parameter 818 can be represented by (vps_max_layers_minus1+1) bits (e.g., the value of VPS parameter 810 plus 1).

[0165]

[0188] 11D, the italicized descriptor 818d indicates that the encoding method for the VPS parameters 818 may be variable length encoding, where the length of the VPS parameters 818 is the smaller of a default value (e.g., 8) or the VPS parameters 810 plus 1 (i.e., the maximum allowable number of layers in the CVS that reference the VPS). For example, in some embodiments, the length of the VPS parameters 818 may be determined based on the following min function: Min(8,vps_max_layers_minus1+1)

[0166]

[0189] Referring again to Figure 9, as explained above, the value of an SPS parameter 914 can range from 0 to the corresponding VPS parameter 812. In some embodiments, when the SPS parameter 912 is equal to 0, the SPS does not reference a VPS, and therefore there is no corresponding VPS parameter 812, and therefore the range of the SPS parameter 914 is undefined. The syntax can be modified to solve this problem by assigning an independent range of values ​​for the SPS parameter 914 to the VPS parameter 812 when the SPS parameter 912 (sps_video_parameter_set_id) is equal to 0.

[0167]

[0190]

[0013] Figure 12A shows a flowchart of an exemplary video encoding method 1200A according to some embodiments of the present disclosure. Figure 12B shows a flowchart of an exemplary video decoding method 1200B corresponding to the video encoding method 1200A of Figure 12A according to some embodiments of the present disclosure. Similar to methods 1000A and 1000B of Figures 10A and 10B, the video encoding method 1200A and the video decoding method 1200B may be performed by an encoder and a decoder implemented as one or more software or hardware components of a device (e.g., a processor executing a set of instructions).

[0168]

[0191] Referring to video encoding method 1200A shown in FIG. 12A, in step 1210a, the encoder encodes a first SPS syntax element (e.g., SPS parameter 914 sps_max_sublayers_minus1) associated with the maximum number of temporal sublayers that may be present in each CLVS that references the SPS. In step 1220a, the encoder determines whether a PTL syntax structure, a DPB parameter syntax structure, and / or an HRD parameter syntax structure is present in the SPS. For example, the encoder may make the determination and then set the value of SPS parameter 918 (sps_ptl_dpb_hrd_params_present_flag) based on the determination. If the determination is true (step 1220a—yes), the encoder performs step 1230a to determine whether the maximum number of temporal sublayers in each coded layered video sequence (CLVS) that references the SPS is greater than one by determining whether SPS parameter 914 is greater than zero.

[0169]

[0192] If both conditions are met (step 1220a—YES, step 1230a—YES), in step 1240a, the encoder signals a flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag) configured to control the presence of a syntax element in the DPB parameter syntax structure in the SPS. The encoder then performs step 1250a to signal one or more syntax elements in the DPB parameter syntax structure.

[0170]

[0193] If the maximum number of temporal sub-layers is less than or equal to 1 (step 1220a—yes, step 1230a—no), the encoder bypasses step 1240a and performs step 1250a for signaling one or more syntax elements in the DPB parameter syntax structure by not signaling a flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag) but setting it equal to 0.

[0171]

[0194] If there is no PTL syntax structure, DPB parameter syntax structure, or HRD parameter syntax structure in the SPS (step 1220a - No), the encoder bypasses steps 1230a to 1250a and encodes the SPS without signaling flags (e.g., SPS parameters 920, sps_sublayer_dpb_params_flag) and DPB syntax elements.

[0172]

[0195] In some embodiments, with regard to encoding the SPS parameter 914 in the video encoding method 1200A, when an SPS references a VPS, the range of the SPS parameter 914 (sps_max_sublayers_minus1) can be set to the same as that of the VPS parameter 812 (vps_max_sublayers_minus1). If an SPS does not reference a VPS, and therefore there is no corresponding VPS parameter 812 (vps_max_sublayers_minus1), the range of the SPS parameter 914 can still be determined by setting a default value. For example, the value of the VPS parameter 812 can be in the range of 0 to a default static value (e.g., 6). Thus, the value of the SPS parameter 914 can be in the range of 0 to MaxSubLayer minus 1, and the value of the parameter MaxSubLayer can be derived and calculated from the code using the ternary operator as follows: MaxSubLayer=(sps_video_parameter_set_id==0?6:vps_max_sublayers_minus1)+1

[0173]

[0196] In other words, the encoder may first determine the value of the SPS parameter 912, which specifies the identifier of the VPS referenced by the SPS when the value of the SPS parameter 912 is greater than zero, and indicates that the SPS does not reference a VPS when the value of the SPS parameter 912 is equal to zero. Then, in response to the value of the SPS parameter 912 being greater than zero, the encoder assigns a range for the SPS parameter 914 that specifies the maximum number of temporal sublayers in each CLVS that references the SPS based on the corresponding VPS parameter 812. That is, the value of the SPS parameter 914 (e.g., sps_max_sublayers_minus1) is within the range of 0 to the VPS parameter 812 (e.g., vps_max_sublayers_minus1), inclusive. On the other hand, in response to the value of the SPS parameter 912 being equal to zero, the encoder assigns the range of the SPS parameter 914 to be greater than or equal to zero and less than or equal to a fixed value (e.g., 6). That is, the value of an SPS parameter 914 (eg, sps_max_sublayers_minus1) is in the range of 0 to 6 (inclusive).

[0174]

[0197] Thus, when the SPS parameter 912 is equal to 0, an inferred value can be given to the SPS parameter 914. Thus, the maximum value of the SPS parameter 914 is available when the SPS does not reference any VPS. In some embodiments, the semantics can be modified such that when the SPS parameter 912 is equal to 0, the value of the VPS parameter 812 (vps_max_sublayers_minus1) is inferred to be equal to a default static value (e.g., 6).

[0175]

[0198] Referring to video decoding method 1200B shown in FIG. 12B, on the decoder side, in step 1210b, the decoder receives a bitstream (e.g., video bitstream 500 of FIG. 5) including an SPS to be decoded. In step 1220b, the decoder decodes a first SPS syntax element (e.g., SPS parameter 914 sps_max_sublayers_minus1). In some embodiments, when decoding an SPS on the decoder side, the decoder may first determine the value of SPS parameter 912 and then decode SPS parameter 914. The range of SPS parameter 914 is based on the corresponding VPS parameter 812 of the VPS referenced by the SPS if the value of SPS parameter 912 is greater than zero, or is based on a fixed value if the value of SPS parameter 912 is equal to zero.

[0176]

[0199] In step 1230b, the decoder determines whether there are PTL syntax structures, DPB parameter syntax structures, and / or HRD parameter syntax structures in the SPS. For example, the decoder may make the determination based on whether the SPS parameter 918 (sps_ptl_dpb_hrd_params_present_flag) is equal to 1. If the determination is true (step 1230b—yes), the decoder performs step 1240b to determine whether the maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS is greater than 1 by determining whether the SPS parameter 914 is greater than zero.

[0177]

[0200] If both conditions are met (step 1230b—YES, step 1240b—YES), in step 1250b, the decoder decodes a flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag) signaled in the SPS and configured to control the presence of syntax elements in the DPB parameter syntax structure. The decoder then performs step 1260b to decode one or more syntax elements in the DPB parameter syntax structure based on the value of the flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag).

[0178]

[0201] If the maximum number of temporal sub-layers is less than or equal to one (step 1230b - yes, step 1240b - no), the decoder bypasses step 1250b, infers that the flag (e.g., SPS parameters 920, sps_sublayer_dpb_params_flag) is equal to zero, and then performs step 1260b to decode one or more syntax elements in the DPB parameter syntax structure based on the value of the flag (e.g., SPS parameters 920, sps_sublayer_dpb_params_flag).

[0179]

[0202] If there is no PTL syntax structure, DPB parameter syntax structure, or HRD parameter syntax structure in the SPS (step 1230b - No), the decoder bypasses steps 1240b to 1260b and decodes the SPS without decoding flags and DPB syntax elements.

[0180]

[0203] 12C shows a portion of an example SPS syntax structure 1200C associated with a possible implementation of the proposed methods 1200A and 1200B according to some embodiments of the present disclosure. The SPS syntax structure 1200C of FIG. 12C may be modified based on the syntax structure 900 of FIG. 9. As explained above, in some embodiments, the signaling of the SPS parameter 920 may be based on both the SPS parameter 918 and the SPS parameter 914. Thus, as shown in FIG. 12C, the syntax element SPS parameter 920 may be signaled (or decoded) under the condition that the SPS parameter 918 is equal to 1 and the SPS parameter 914 is greater than 0. Otherwise, the SPS parameter 920 is not signaled (or decoded).

[0181]

[0204] In view of the above, as proposed in various embodiments of the present disclosure, by applying variable length encoding or decoding of PTL syntax elements, the number of syntax structures and the encoding method of syntax structure indices between DPB, HRD, and PTL parameters can be made consistent and efficient, thereby reducing the signaling overhead caused by using fixed lengths for these syntax elements. In addition, by appropriately inferring the value of a syntax element when it is not signaled, it is possible to skip signaling of the index in some cases, which reduces the number of output bits and thus improves coding efficiency. This method can be used not only for PTL and DPB parameters but also for HRD parameters to reduce signaling overhead and ensure consistency in signaling design.

[0182]

[0205] The embodiments may be further described using the following clauses: 1. A computer-implemented method for encoding video, comprising: determining whether the coded video sequence (CVS) includes an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS); and and encoding the bitstream, in response to the CVS including an equal number of PTL syntax structures and OLSs, without signaling a first PTL syntax element specifying an index into a list of PTL syntax structures in the VPS of a PTL syntax structure that applies to a corresponding OLS in the VPS. 20. A computer-implemented method comprising: 2. Determining whether the number of PTL syntax structures is equal to 1; and and encoding the bitstream without signaling a first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to one. 2. The method of clause 1, further comprising: 3. In response to the number of PTL syntax structures being greater than one and different from the number of OLSs, signaling a first PTL syntax element having a fixed length within the VPS. 2. The method of clause 1, further comprising: 4. A computer-implemented method for encoding video, comprising: Determining whether the coded video sequence (CVS) has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS); and In response to the CVS having an equal number of DPB parameter syntax structures and OLSs, encoding the bitstream without signaling a first DPB syntax element specifying an index, relative to a list of DPB parameter syntax structures in the VPS, of a DPB parameter syntax structure that applies to a corresponding OLS. 20. A computer-implemented method comprising: 5. Determining whether the number of DPB parameter syntax structures is less than or equal to one; and and encoding the bitstream without signaling a first DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being equal to one. 5. The method of clause 4, further comprising: 6. In response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs, signaling a first DPB syntax element in the VPS having a variable length. 5. The method of clause 4, further comprising: 7. A computer-implemented method for encoding video, comprising: determining whether at least one of a profile / tier / level (PTL) syntax structure, a decoded picture buffer (DPB) parameters syntax structure, or a hypothetical reference decoder (HRD) parameters syntax structure is in a sequence parameter set (SPS) of the bitstream; determining whether a first value is greater than one, the first value specifying a maximum number of temporal sub-layers in a coded layered video sequence (CLVS) that references the SPS; and signaling a flag configured to control the presence of a syntax element in the DPB parameter syntax structure in the SPS when at least one of the PTL syntax structure, the DPB parameter syntax structure, or the HRD parameter syntax structure is in the SPS and the first value is greater than 1; 20. A computer-implemented method comprising: 8. If the first value is less than or equal to 1, signaling one or more syntax elements in the DPB parameter syntax structure without signaling a flag in the SPS. 8. The method of clause 7, further comprising: 9. A computer-implemented method for encoding video, comprising: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero; In response to a value of the first SPS syntax element being greater than zero, assigning a range of a second SPS syntax element based on the corresponding VPS syntax element, the second SPS syntax element specifying a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS; and In response to the value of the first SPS syntax element being equal to zero, assigning a range of a second SPS syntax element specifying a maximum number of temporal sublayers in each CLVS that references the SPS to be greater than or equal to zero and less than or equal to a fixed value. 20. A computer-implemented method comprising: 10. A computer-implemented method for encoding video, comprising: encoding one or more Profile / Tier / Level (PTL) syntax elements that specify PTL-related information; and Signaling one or more PTL syntax elements with variable length within a video parameter set (VPS) or sequence parameter set (SPS) of a bitstream 20. A computer-implemented method comprising: 11. Encoding one or more PTL syntax elements: Encoding a first PTL syntax element having a variable length, the first PTL syntax element specifying an index of a PTL syntax structure. 11. The method of clause 10, comprising: 12. The VPS or SPS includes N PTL syntax structures, where N is an integer, and encoding one or more PTL syntax elements includes: Setting the length of the first PTL syntax element to be the smallest integer greater than or equal to the base 2 logarithm of N 12. The method of clause 11, further comprising: 13. Encoding one or more PTL syntax elements: Encoding the first PTL syntax element by using an exponential-Golomb code 12. The method according to clause 11, comprising: 14. The VPS or SPS includes N PTL syntax structures, where N is an integer, and encoding one or more PTL syntax elements includes: encoding a second PTL syntax element having a variable length, the second PTL syntax element specifying N; 11. The method of clause 10, comprising: 15. A computer-implemented method for encoding video, comprising: Encoding a video parameter set (VPS) syntax element having a variable length; and Signaling a VPS syntax element within the VPS, the VPS syntax element relating to the number of Output Layer Sets (OLS) contained within a Coded Video Sequence (CVS) that references the VPS. 20. A computer-implemented method comprising: 16. If the maximum allowable number of layers in a Coded Video Sequence (CVS) that references a VPS is less than the default length value, set the length of the VPS syntax element equal to the maximum allowable number. 16. The method of clause 15, further comprising: 17. A computer-implemented method for decoding video, comprising: receiving a bitstream including a coded video sequence (CVS); Determining whether the CVS has an equal number of Profile / Tier / Level (PTL) syntax structures and Output Layer Sets (OLS); and and in response to the number of PTL syntax structures being equal to the number of OLSs, when decoding the VPS, skipping decoding of a first PTL syntax element that specifies an index into a list of PTL syntax structures in the VPS of a PTL syntax structure that applies to the corresponding OLS. 20. A computer-implemented method comprising: 18. In response to the number of PTL syntax structures being equal to the number of OLSs, determining a first PTL syntax element such that the first PTL syntax element is equal to an index of the OLS to which the PTL syntax structure specified by the first PTL syntax element applies. 18. The method of clause 17, further comprising: 19. Determining whether the number of PTL syntax structures is equal to 1; and and in response to the number of PTL syntax structures being equal to one, skipping decoding of the first PTL syntax element when decoding the VPS. 18. The method of clause 17, further comprising: 20. In response to the number of PTL syntax structures being equal to one, determining the first PTL syntax element to be zero. 20. The method of clause 19, further comprising: 21. In response to the number of PTL syntax structures being greater than one and different from the number of OLSs, decoding a first PTL syntax element having a fixed length in the VPS. 18. The method of clause 17, further comprising: 22. A computer-implemented method for decoding video, comprising: receiving a bitstream including a coded video sequence (CVS); Determining whether the CVS has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS); and and in response to the CVS having an equal number of DPB parameter syntax structures and OLSs, skipping decoding of a first DPB syntax element that specifies an index into a list of DPB parameter syntax structures in the VPS of a DPB parameter syntax structure that applies to a corresponding OLS. 20. A computer-implemented method comprising: 23. Determining whether the number of DPB parameter syntax structures is less than or equal to one; and and skipping decoding of the first DPB syntax element when decoding the VPS in response to the number of DPB parameter syntax structures being one or less. 23. The method of clause 22, further comprising: 24. In response to the number of DPB parameter syntax structures being less than or equal to one, determining the first DPB syntax element to be zero; and In response to the number of DPB syntax structures being equal to the number of OLSs, determining a first DPB syntax element such that the DPB parameter syntax structure specified by the first DPB syntax element is equal to an index of the OLS to which it applies. 24. The method of clause 23, further comprising: 25. In response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs, decoding a first DPB syntax element having a variable length in the VPS. 24. The method of clause 23, further comprising: 26. A computer-implemented method for decoding video, comprising: receiving a bitstream including a video parameter set (VPS) and a sequence parameter set (SPS); In response to at least one of a Profile / Tier / Level (PTL) syntax structure, a Decoded Picture Buffer (DPB) parameters syntax structure, or a Hypothetical Reference Decoder (HRD) parameters syntax structure being in the SPS, determining whether a first value specifying a maximum number of temporal sub-layers in each Coded Layered Video Sequence (CLVS) that references the SPS is greater than one; and decoding a flag configured to control a presence of a syntax element within a DPB parameter syntax structure within the SPS in response to the first value being greater than one; 20. A computer-implemented method comprising: 27. In response to the first value being less than or equal to one, inferring that the flag is equal to zero and decoding one or more syntax elements in the DPB parameter syntax structure without decoding the flag in the SPS. 27. The method of clause 26, further comprising: 28. A computer-implemented method for decoding video, comprising: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero; and decoding a second SPS syntax element that specifies a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS, wherein the range of the second SPS syntax element is based on a corresponding VPS syntax element of a VPS referenced by the SPS if the value of the first SPS syntax element is greater than zero, or is based on a fixed value if the value of the first SPS syntax element is equal to zero; 20. A computer-implemented method comprising: 29. A computer-implemented method for decoding video, comprising: receiving a bitstream including a video parameter set (VPS) or a sequence parameter set (SPS); and Decoding one or more Profile / Tier / Level (PTL) syntax elements in a VPS or SPS, where the one or more PTL syntax elements specify PTL-related information. 20. A computer-implemented method comprising: 30. Decoding one or more PTL syntax elements Decoding a first PTL syntax element having a variable length, the first PTL syntax element specifying an index of a PTL syntax structure. 29. The method of claim 29, comprising: 31. The method of clause 30, wherein the VPS or SPS includes N PTL syntax structures, N is an integer, and the length of the first PTL syntax element is the smallest integer greater than or equal to the logarithm to the base 2 of N. 32. The method of clause 30, wherein the first PTL syntax element is encoded by using an exponential-Golomb code. 33. The VPS or SPS includes N PTL syntax structures, where N is an integer, and decoding one or more PTL syntax elements includes: Decoding a second PTL syntax element having a variable length, the second PTL syntax element specifying N. 29. The method of claim 29, comprising: 34. A computer-implemented method for decoding video, comprising: receiving a bitstream including a video parameter set (VPS); and Decoding a VPS syntax element having a variable length in the VPS, the VPS syntax element being related to the number of Output Layer Sets (OLS) contained in the Coded Video Sequence (CVS) that references the VPS. 20. A computer-implemented method comprising: 35. The method of clause 34, wherein if the maximum allowed number of layers in a coded video sequence (CVS) that references the VPS is less than the predetermined length value, the length of the VPS syntax element is equal to the maximum allowed number. 36. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: determining whether the coded video sequence (CVS) includes an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS); and and encoding the bitstream, in response to the CVS including an equal number of PTL syntax structures and OLSs, without signaling a first PTL syntax element specifying an index into a list of PTL syntax structures in the VPS of a PTL syntax structure that applies to a corresponding OLS in the VPS. 10. An apparatus configured to execute instructions to cause the apparatus to: 37. The processor shall: determining whether the number of PTL syntax structures is equal to 1; and and encoding the bitstream without signaling a first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to one. 36. An apparatus as described in clause 36 configured to execute instructions to: 38. The processor shall: signaling a first PTL syntax element having a fixed length in the VPS in response to the number of PTL syntax structures being greater than one and different from the number of OLSs; 36. An apparatus as described in clause 36 configured to execute instructions to: 39. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: Determining whether the coded video sequence (CVS) has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS); and In response to the CVS having an equal number of DPB parameter syntax structures and OLSs, encoding the bitstream without signaling a first DPB syntax element specifying an index, relative to a list of DPB parameter syntax structures in the VPS, of a DPB parameter syntax structure that applies to a corresponding OLS. 10. An apparatus configured to execute instructions to cause the apparatus to: 40. The processor shall: determining whether the number of DPB parameter syntax structures is less than or equal to one; and and encoding the bitstream without signaling a first DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being equal to one. 39. An apparatus as described in clause 39 configured to execute instructions to: 41. The processor shall: signaling a first DPB syntax element having a variable length in the VPS in response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs; 39. An apparatus as described in clause 39 configured to execute instructions to: 42. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: determining whether at least one of a profile / tier / level (PTL) syntax structure, a decoded picture buffer (DPB) parameters syntax structure, or a hypothetical reference decoder (HRD) parameters syntax structure is in a sequence parameter set (SPS) of the bitstream; determining whether a first value is greater than one, the first value specifying a maximum number of temporal sub-layers in a coded layered video sequence (CLVS) that references the SPS; and signaling a flag configured to control the presence of a syntax element in the DPB parameter syntax structure in the SPS when at least one of the PTL syntax structure, the DPB parameter syntax structure, or the HRD parameter syntax structure is in the SPS and the first value is greater than 1; 10. An apparatus configured to execute instructions to cause the apparatus to: 43. The processor shall: If the first value is less than or equal to 1, signaling one or more syntax elements in the DPB parameter syntax structure without signaling a flag in the SPS. 42. An apparatus as described in clause 42 configured to execute instructions to: 44. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero; In response to a value of the first SPS syntax element being greater than zero, assigning a range of a second SPS syntax element based on the corresponding VPS syntax element, the second SPS syntax element specifying a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS; and In response to the value of the first SPS syntax element being equal to zero, assigning a range of a second SPS syntax element specifying a maximum number of temporal sublayers in each CLVS that references the SPS to be greater than or equal to zero and less than or equal to a fixed value. 10. An apparatus configured to execute instructions to cause the apparatus to: 45. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: encoding one or more Profile / Tier / Level (PTL) syntax elements that specify PTL-related information; and Signaling one or more PTL syntax elements with variable length within a video parameter set (VPS) or sequence parameter set (SPS) of a bitstream 10. An apparatus configured to execute instructions to cause the apparatus to: 46. ​​The processor shall: Encoding a first PTL syntax element having a variable length, the first PTL syntax element specifying an index of a PTL syntax structure. 46. ​​The apparatus of clause 45, configured to execute instructions to encode one or more PTL syntax elements by: 47. The VPS or SPS includes N PTL syntax structures, where N is an integer, and the processor: Setting the length of the first PTL syntax element to be the smallest integer greater than or equal to the base 2 logarithm of N 47. The apparatus of clause 46, configured to execute instructions to encode one or more PTL syntax elements by: 48. The processor shall: Encoding the first PTL syntax element by using an exponential-Golomb code 47. The apparatus of clause 46, configured to execute instructions to encode one or more PTL syntax elements by: 49. The VPS or SPS includes N PTL syntax structures, where N is an integer, and the processor: encoding a second PTL syntax element having a variable length, the second PTL syntax element specifying N; 46. ​​The apparatus of clause 45, configured to execute instructions to encode one or more PTL syntax elements by: 50. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: Encoding a video parameter set (VPS) syntax element having a variable length; and Signaling a VPS syntax element within the VPS, the VPS syntax element relating to the number of Output Layer Sets (OLS) contained within a Coded Video Sequence (CVS) that references the VPS. 10. An apparatus configured to execute instructions to cause the apparatus to: 51. The processor shall: If the maximum allowable number of layers in a Coded Video Sequence (CVS) that references a VPS is less than the default length value, set the length of the VPS syntax element equal to the maximum allowable number. 50. An apparatus as described in clause 50 configured to execute instructions to: 52. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: receiving a bitstream including a coded video sequence (CVS); Determining whether the CVS has an equal number of Profile / Tier / Level (PTL) syntax structures and Output Layer Sets (OLS); and and in response to the number of PTL syntax structures being equal to the number of OLSs, when decoding the VPS, skipping decoding of a first PTL syntax element that specifies an index into a list of PTL syntax structures in the VPS of a PTL syntax structure that applies to the corresponding OLS. 10. An apparatus configured to execute instructions to cause the apparatus to: 53. The processor shall: and determining a first PTL syntax element such that the number of PTL syntax structures is equal to the number of OLSs to which the PTL syntax structure specified by the first PTL syntax element applies, in response to the number of PTL syntax structures being equal to the number of OLSs. 52. An apparatus as described in clause 52 configured to execute instructions to: 54. The processor shall: determining whether the number of PTL syntax structures is equal to 1; and and in response to the number of PTL syntax structures being equal to one, skipping decoding of the first PTL syntax element when decoding the VPS. 52. An apparatus as described in clause 52 configured to execute instructions to: 55. The processor shall: determining a first PTL syntax element to be zero in response to the number of PTL syntax structures being equal to one; 54. An apparatus as described in clause 54 configured to execute instructions to: 56. The processor shall: decoding a first PTL syntax element having a fixed length in the VPS in response to the number of PTL syntax structures being greater than one and different from the number of OLSs; 52. An apparatus as described in clause 52 configured to execute instructions to: 57. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: receiving a bitstream including a coded video sequence (CVS); Determining whether the CVS has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS); and and in response to the CVS having an equal number of DPB parameter syntax structures and OLSs, skipping decoding of a first DPB syntax element that specifies an index into a list of DPB parameter syntax structures in the VPS of a DPB parameter syntax structure that applies to a corresponding OLS. 10. An apparatus configured to execute instructions to cause the apparatus to: 58. The processor shall: determining whether the number of DPB parameter syntax structures is less than or equal to one; and and skipping decoding of the first DPB syntax element when decoding the VPS in response to the number of DPB parameter syntax structures being one or less. 57. An apparatus as described in clause 57 configured to execute instructions to: 59. The processor shall: determining the first DPB syntax element to be zero in response to the number of DPB parameter syntax structures being less than or equal to one; and In response to the number of DPB syntax structures being equal to the number of OLSs, determining a first DPB syntax element such that the DPB parameter syntax structure specified by the first DPB syntax element is equal to an index of the OLS to which it applies. 58. An apparatus as described in clause 58 configured to execute instructions to: 60. The processor shall: decoding a first DPB syntax element having a variable length in the VPS in response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs; 58. An apparatus as described in clause 58 configured to execute instructions to: 61. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: receiving a bitstream including a video parameter set (VPS) and a sequence parameter set (SPS); In response to at least one of a Profile / Tier / Level (PTL) syntax structure, a Decoded Picture Buffer (DPB) parameters syntax structure, or a Hypothetical Reference Decoder (HRD) parameters syntax structure being in the SPS, determining whether a first value specifying a maximum number of temporal sub-layers in each Coded Layered Video Sequence (CLVS) that references the SPS is greater than one; and decoding a flag configured to control a presence of a syntax element within a DPB parameter syntax structure within the SPS in response to the first value being greater than one; 10. An apparatus configured to execute instructions to cause the apparatus to: 62. The processor shall: in response to the first value being less than or equal to one, inferring that the flag is equal to zero and decoding one or more syntax elements in the DPB parameter syntax structure without decoding the flag in the SPS. 6. An apparatus as described in clause 61 configured to execute instructions to: 63. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero; and decoding a second SPS syntax element that specifies a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS, wherein the range of the second SPS syntax element is based on a corresponding VPS syntax element of a VPS referenced by the SPS if the value of the first SPS syntax element is greater than zero, or is based on a fixed value if the value of the first SPS syntax element is equal to zero; 10. An apparatus configured to execute instructions to cause the apparatus to: 64. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: receiving a bitstream including a video parameter set (VPS) or a sequence parameter set (SPS); and Decoding one or more Profile / Tier / Level (PTL) syntax elements in a VPS or SPS, where the one or more PTL syntax elements specify PTL-related information. 10. An apparatus configured to execute instructions to cause the apparatus to: 65. The processor shall: Decoding a first PTL syntax element having a variable length, the first PTL syntax element specifying an index of a PTL syntax structure. 65. The apparatus of claim 64, configured to execute instructions to decode one or more PTL syntax elements by: 66. The device described in clause 65, wherein the VPS or SPS includes N PTL syntax structures, where N is an integer, and the length of the first PTL syntax element is the smallest integer greater than or equal to the base 2 logarithm of N. 67. The device of clause 65, wherein the first PTL syntax element is encoded by using an exponential-Golomb code. 68. The VPS or SPS includes N PTL syntax structures, where N is an integer, and the processor: Decoding a second PTL syntax element having a variable length, the second PTL syntax element specifying N. 65. The apparatus of claim 64, configured to execute instructions to decode one or more PTL syntax elements by: 69. A memory configured to store instructions; a processor coupled to the memory, the processor comprising: receiving a bitstream including a video parameter set (VPS); and Decoding a VPS syntax element having a variable length in the VPS, the VPS syntax element being related to the number of Output Layer Sets (OLS) contained in the Coded Video Sequence (CVS) that references the VPS. 10. An apparatus configured to execute instructions to cause the apparatus to: 70. The apparatus of clause 69, wherein if the maximum allowable number of layers in a coded video sequence (CVS) that references the VPS is less than the predetermined length value, the length of the VPS syntax element is equal to the maximum allowable number. 71. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for encoding video, the method comprising: determining whether the coded video sequence (CVS) includes an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS); and and encoding the bitstream, in response to the CVS including an equal number of PTL syntax structures and OLSs, without signaling a first PTL syntax element specifying an index into a list of PTL syntax structures in the VPS of a PTL syntax structure that applies to a corresponding OLS in the VPS. 1. A non-transitory computer-readable storage medium comprising: 72. A set of instructions executable by one or more processors of a device comprises: determining whether the number of PTL syntax structures is equal to 1; and and encoding the bitstream without signaling a first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to one. 72. The non-transitory computer-readable storage medium of claim 71, further causing the device to: 73. A set of instructions executable by one or more processors of a device comprises: signaling a first PTL syntax element having a fixed length in the VPS in response to the number of PTL syntax structures being greater than one and different from the number of OLSs; 72. The non-transitory computer-readable storage medium of claim 71, further causing the device to: 74. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for encoding video, the method comprising: Determining whether the coded video sequence (CVS) has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS); and In response to the CVS having an equal number of DPB parameter syntax structures and OLSs, encoding the bitstream without signaling a first DPB syntax element specifying an index, relative to a list of DPB parameter syntax structures in the VPS, of a DPB parameter syntax structure that applies to a corresponding OLS. 1. A non-transitory computer-readable storage medium comprising: 75. A set of instructions executable by one or more processors of a device comprises: determining whether the number of DPB parameter syntax structures is less than or equal to one; and and encoding the bitstream without signaling a first DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being equal to one. 75. The non-transitory computer-readable storage medium of clause 74, further causing the device to: 76. A set of instructions executable by one or more processors of a device comprises: signaling a first DPB syntax element having a variable length in the VPS in response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs; 75. The non-transitory computer-readable storage medium of clause 74, further causing the device to: 77. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for encoding video, the method comprising: determining whether at least one of a profile / tier / level (PTL) syntax structure, a decoded picture buffer (DPB) parameters syntax structure, or a hypothetical reference decoder (HRD) parameters syntax structure is in a sequence parameter set (SPS) of the bitstream; determining whether a first value is greater than one, the first value specifying a maximum number of temporal sub-layers in a coded layered video sequence (CLVS) that references the SPS; and signaling a flag configured to control the presence of a syntax element in the DPB parameter syntax structure in the SPS when at least one of the PTL syntax structure, the DPB parameter syntax structure, or the HRD parameter syntax structure is in the SPS and the first value is greater than 1; 1. A non-transitory computer-readable storage medium comprising: 78. A set of instructions executable by one or more processors of a device comprises: If the first value is less than or equal to 1, signaling one or more syntax elements in the DPB parameter syntax structure without signaling a flag in the SPS. 78. The non-transitory computer-readable storage medium of claim 77, further causing the device to: 79. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for encoding video, the method comprising: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero; In response to a value of the first SPS syntax element being greater than zero, assigning a range of a second SPS syntax element based on the corresponding VPS syntax element, the second SPS syntax element specifying a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS; and In response to the value of the first SPS syntax element being equal to zero, assigning a range of a second SPS syntax element specifying a maximum number of temporal sublayers in each CLVS that references the SPS to be greater than or equal to zero and less than or equal to a fixed value. 1. A non-transitory computer-readable storage medium comprising: 80. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for encoding video, the method comprising: encoding one or more Profile / Tier / Level (PTL) syntax elements that specify PTL-related information; and Signaling one or more PTL syntax elements with variable length within a video parameter set (VPS) or sequence parameter set (SPS) of a bitstream 1. A non-transitory computer-readable storage medium comprising: 81. A set of instructions executable by one or more processors of a device comprises: Encoding a first PTL syntax element having a variable length, the first PTL syntax element specifying an index of a PTL syntax structure. 81. The non-transitory computer-readable storage medium of clause 80, causing a device to encode one or more PTL syntax elements by: 82. The VPS or SPS includes N PTL syntax structures, where N is an integer, and the set of instructions executable by one or more processors of the device is: Setting the length of the first PTL syntax element to be the smallest integer greater than or equal to the base 2 logarithm of N 82. The non-transitory computer-readable storage medium of claim 81, causing a device to encode one or more PTL syntax elements by: 83. A set of instructions executable by one or more processors of a device comprises: Encoding the first PTL syntax element by using an exponential-Golomb code 82. The non-transitory computer-readable storage medium of claim 81, causing a device to encode one or more PTL syntax elements by: 84. The VPS or SPS includes N PTL syntax structures, where N is an integer, and the set of instructions executable by one or more processors of the device is: encoding a second PTL syntax element having a variable length, the second PTL syntax element specifying N; 81. The non-transitory computer-readable storage medium of clause 80, causing a device to encode one or more PTL syntax elements by: 85. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for encoding video, the method comprising: Encoding a video parameter set (VPS) syntax element having a variable length; and Signaling a VPS syntax element within the VPS, the VPS syntax element relating to the number of Output Layer Sets (OLS) contained within a Coded Video Sequence (CVS) that references the VPS. 1. A non-transitory computer-readable storage medium comprising: 86. A set of instructions executable by one or more processors of a device comprises: If the maximum allowable number of layers in a Coded Video Sequence (CVS) that references a VPS is less than the default length value, set the length of the VPS syntax element equal to the maximum allowable number. 86. The non-transitory computer-readable storage medium of claim 85, further causing the device to: 87. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for decoding video, the method comprising: receiving a bitstream including a coded video sequence (CVS); Determining whether the CVS has an equal number of Profile / Tier / Level (PTL) syntax structures and Output Layer Sets (OLS); and and in response to the number of PTL syntax structures being equal to the number of OLSs, when decoding the VPS, skipping decoding of a first PTL syntax element that specifies an index into a list of PTL syntax structures in the VPS of a PTL syntax structure that applies to the corresponding OLS. 1. A non-transitory computer-readable storage medium comprising: 88. A set of instructions executable by one or more processors of a device comprises: and determining a first PTL syntax element such that the number of PTL syntax structures is equal to the number of OLSs to which the PTL syntax structure specified by the first PTL syntax element applies, in response to the number of PTL syntax structures being equal to the number of OLSs. 88. The non-transitory computer-readable storage medium of claim 87, further causing the device to: 89. A set of instructions executable by one or more processors of a device comprises: determining whether the number of PTL syntax structures is equal to 1; and and in response to the number of PTL syntax structures being equal to one, skipping decoding of the first PTL syntax element when decoding the VPS. 88. The non-transitory computer-readable storage medium of claim 87, further causing the device to: 90. A set of instructions executable by one or more processors of a device comprises: determining a first PTL syntax element to be zero in response to the number of PTL syntax structures being equal to one; 89. The non-transitory computer-readable storage medium of claim 89, further causing the device to: 91. A set of instructions executable by one or more processors of a device comprises: decoding a first PTL syntax element having a fixed length in the VPS in response to the number of PTL syntax structures being greater than one and different from the number of OLSs; 88. The non-transitory computer-readable storage medium of claim 87, further causing the device to: 92. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for decoding video, the method comprising: receiving a bitstream including a coded video sequence (CVS); Determining whether the CVS has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS); and and in response to the CVS having an equal number of DPB parameter syntax structures and OLSs, skipping decoding of a first DPB syntax element that specifies an index into a list of DPB parameter syntax structures in the VPS of a DPB parameter syntax structure that applies to a corresponding OLS. 1. A non-transitory computer-readable storage medium comprising: 93. A set of instructions executable by one or more processors of a device comprises: determining whether the number of DPB parameter syntax structures is less than or equal to one; and and skipping decoding of the first DPB syntax element when decoding the VPS in response to the number of DPB parameter syntax structures being one or less. 93. The non-transitory computer-readable storage medium of claim 92, further causing the device to: 94. A set of instructions executable by one or more processors of a device comprises: determining the first DPB syntax element to be zero in response to the number of DPB parameter syntax structures being less than or equal to one; and In response to the number of DPB syntax structures being equal to the number of OLSs, determining a first DPB syntax element such that the DPB parameter syntax structure specified by the first DPB syntax element is equal to an index of the OLS to which it applies. 94. The non-transitory computer-readable storage medium of claim 93, further causing the device to: 95. A set of instructions executable by one or more processors of a device comprises: decoding a first DPB syntax element having a variable length in the VPS in response to the number of DPB parameter syntax structures being greater than one and different from the number of OLSs; 94. The non-transitory computer-readable storage medium of claim 93, further causing the device to: 96. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for decoding video, the method comprising: receiving a bitstream including a video parameter set (VPS) and a sequence parameter set (SPS); In response to at least one of a Profile / Tier / Level (PTL) syntax structure, a Decoded Picture Buffer (DPB) parameters syntax structure, or a Hypothetical Reference Decoder (HRD) parameters syntax structure being in the SPS, determining whether a first value specifying a maximum number of temporal sub-layers in each Coded Layered Video Sequence (CLVS) that references the SPS is greater than one; and decoding a flag configured to control a presence of a syntax element within a DPB parameter syntax structure within the SPS in response to the first value being greater than one; 1. A non-transitory computer-readable storage medium comprising: 97. A set of instructions executable by one or more processors of a device comprises: in response to the first value being less than or equal to one, inferring that the flag is equal to zero and decoding one or more syntax elements in the DPB parameter syntax structure without decoding the flag in the SPS. 97. The non-transitory computer-readable storage medium of claim 96, further causing the device to: 98. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for decoding video, the method comprising: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero; and decoding a second SPS syntax element that specifies a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS, wherein the range of the second SPS syntax element is based on a corresponding VPS syntax element of a VPS referenced by the SPS if the value of the first SPS syntax element is greater than zero, or is based on a fixed value if the value of the first SPS syntax element is equal to zero; 1. A non-transitory computer-readable storage medium comprising: 99. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for decoding video, the method comprising: receiving a bitstream including a video parameter set (VPS) or a sequence parameter set (SPS); and Decoding one or more Profile / Tier / Level (PTL) syntax elements in a VPS or SPS, where the one or more PTL syntax elements specify PTL-related information. 1. A non-transitory computer-readable storage medium comprising: 100. A set of instructions executable by one or more processors of a device comprising: Decoding a first PTL syntax element having a variable length, the first PTL syntax element specifying an index of a PTL syntax structure. 99. A non-transitory computer-readable storage medium as described in clause 99, which causes a device to decode one or more PTL syntax elements by: 101. The non-transitory computer-readable storage medium of clause 100, wherein the VPS or SPS includes N PTL syntax structures, where N is an integer, and the length of the first PTL syntax element is the smallest integer greater than or equal to the base 2 logarithm of N. 102. The non-transitory computer-readable storage medium of clause 100, wherein the first PTL syntax element is encoded by using an Exponential-Golomb code. 103. The VPS or SPS includes N PTL syntax structures, where N is an integer, and the set of instructions executable by one or more processors of the device is: Decoding a second PTL syntax element having a variable length, the second PTL syntax element specifying N. 99. A non-transitory computer-readable storage medium as described in clause 99, which causes a device to decode one or more PTL syntax elements by: 104. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for decoding video, the method comprising: receiving a bitstream including a video parameter set (VPS); and Decoding a VPS syntax element having a variable length in the VPS, the VPS syntax element being related to the number of Output Layer Sets (OLS) contained in the Coded Video Sequence (CVS) that references the VPS. 1. A non-transitory computer-readable storage medium comprising: 105. The non-transitory computer-readable storage medium of clause 104, wherein if the maximum allowable number of layers in a coded video sequence (CVS) that references the VPS is less than a predetermined length value, the length of the VPS syntax element is equal to the maximum allowable number.

[0183]

[0206] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which can be executed by a device (such as the encoders and decoders of the present disclosure) to perform the above-described methods. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. A device can include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0184]

[0207] It should be noted that relational terms herein, such as "first" and "second," are used merely to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between those entities or operations. Furthermore, the words "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and open-ended in that the element or elements following any of these words are not meant to be an exclusive listing of such elements or elements, or to be limited to only the listed element or elements.

[0185]

[0208] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless impracticable. For example, if it is stated that a database can include A or B, then the database can include A, B, A and B unless specifically stated otherwise or impracticable. As a second example, if it is stated that a database can include A, B or C, then the database can include A, B, C, A and B, A and C, B and C, A and B and C, unless specifically stated otherwise or impracticable.

[0186]

[0209] It is understood that the above-described embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the methods of the present disclosure. The computational units and other functional units described in the present disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0187]

[0210] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the above-described embodiments may be made. Other embodiments may be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the invention being indicated by the appended claims. It is also intended that the sequences of steps depicted in the figures are for illustrative purposes only and are not intended to be limited to any particular sequence of steps. As such, one skilled in the art will recognize that these steps may be performed in different orders while implementing the same method.

[0188]

[0211] Illustrative embodiments have been disclosed in the drawings and herein. However, many variations and modifications to these embodiments may be made. Thus, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. A device, a memory configured to store instructions; one or more processors, the processors comprising: determining whether the coded video sequence (CVS) includes an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS); and and encoding the bitstream without signaling a first PTL syntax element specifying an index into a list of PTL syntax structures in a video parameter set (VPS) of a PTL syntax structure that applies to a corresponding OLS in the VPS, in response to the CVS including an equal number of PTL syntax structures and OLSs. an apparatus configured to execute the instructions to cause the apparatus to perform

2. The one or more processors: determining whether the number of PTL syntax structures is equal to 1; and and encoding the bitstream without signaling the first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to one.

2. The device of claim 1, configured to execute the instructions to cause the device to:

3. The one or more processors: signaling the first PTL syntax element having a fixed length within the VPS in response to the number of PTL syntax structures being greater than one and different from the number of OLSs.

2. The device of claim 1, configured to execute the instructions to cause the device to:

4. The one or more processors: determining whether at least one of the PTL syntax structure, the decoded picture buffer (DPB) parameters syntax structure, or the hypothetical reference decoder (HRD) parameters syntax structure is in a sequence parameter set (SPS) of the bitstream; determining whether a first value is greater than one, the first value specifying a maximum number of temporal sub-layers in a coded layer video sequence (CLVS) that references the SPS; and signaling a flag configured to control the presence of syntax elements in the DPB parameter syntax structure in the SPS when at least one of the PTL syntax structure, the DPB parameter syntax structure, or the HRD parameter syntax structure is in the SPS and the first value is greater than 1; 2. The device of claim 1, configured to execute the instructions to cause the device to:

5. The one or more processors: signaling one or more syntax elements in the DPB parameter syntax structure without signaling the flag in the SPS if the first value is less than or equal to one.

5. The device of claim 4, configured to execute the instructions to cause the device to:

6. The one or more processors: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of the VPS referenced by the SPS when the value of the first SPS syntax element is greater than zero; In response to the value of the first SPS syntax element being greater than zero, assigning a range of a second SPS syntax element based on a corresponding VPS syntax element, the range specifying a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS; and assigning the range of the second SPS syntax element specifying the maximum number of temporal sub-layers in each CLVS referencing the SPS to be greater than or equal to zero and less than or equal to a fixed value in response to the value of the first SPS syntax element being equal to zero.

2. The device of claim 1, configured to execute the instructions to cause the device to:

7. 1. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for encoding video, the method comprising: determining whether the coded video sequence (CVS) includes an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS); and and encoding the bitstream without signaling a first PTL syntax element specifying an index into a list of PTL syntax structures in a video parameter set (VPS) of a PTL syntax structure that applies to a corresponding OLS in the VPS, in response to the CVS including an equal number of PTL syntax structures and OLSs.

1. A non-transitory computer-readable storage medium comprising:

8. The set of instructions executable by the one or more processors of the device comprises: determining whether the number of PTL syntax structures is equal to 1; and and encoding the bitstream without signaling the first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to one. The non-transitory computer-readable storage medium of claim 7 , further causing the device to:

9. The set of instructions executable by the one or more processors of the device comprises: signaling the first PTL syntax element having a fixed length within the VPS in response to the number of PTL syntax structures being greater than one and different from the number of OLSs. The non-transitory computer-readable storage medium of claim 7 , further causing the device to:

10. The set of instructions executable by the one or more processors of the device comprises: determining whether at least one of the PTL syntax structure, the decoded picture buffer (DPB) parameters syntax structure, or the hypothetical reference decoder (HRD) parameters syntax structure is in a sequence parameter set (SPS) of the bitstream; determining whether a first value is greater than one, the first value specifying a maximum number of temporal sub-layers in a coded layer video sequence (CLVS) that references the SPS; and signaling a flag configured to control the presence of syntax elements in the DPB parameter syntax structure in the SPS when at least one of the PTL syntax structure, the DPB parameter syntax structure, or the HRD parameter syntax structure is in the SPS and the first value is greater than 1; The non-transitory computer-readable storage medium of claim 7 , further causing the device to:

11. The set of instructions executable by the one or more processors of the device comprises: signaling one or more syntax elements in the DPB parameter syntax structure without signaling the flag in the SPS if the first value is less than or equal to one. The non-transitory computer-readable storage medium of claim 10 , further causing the device to:

12. The set of instructions executable by the one or more processors of the device comprises: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of the VPS referenced by the SPS when the value of the first SPS syntax element is greater than zero; In response to the value of the first SPS syntax element being greater than zero, assigning a range of a second SPS syntax element based on a corresponding VPS syntax element, the range specifying a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS; and assigning the range of the second SPS syntax element specifying the maximum number of temporal sub-layers in each CLVS referencing the SPS to be greater than or equal to zero and less than or equal to a fixed value in response to the value of the first SPS syntax element being equal to zero. The non-transitory computer-readable storage medium of claim 7 , further causing the device to:

13. 1. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions executable by one or more processors of a device to cause the device to perform a method for decoding video, the method comprising: receiving a bitstream including a coded video sequence (CVS) and a video parameter set (VPS); determining whether the number of Profile / Tier / Level (PTL) syntax structures signaled in the VPS is equal to the number of Output Layer Sets (OLS) specified in the VPS; and and in response to the number of the PTL syntax structures being equal to the number of the OLSs, when decoding the VPS, skipping decoding of a first PTL syntax element that specifies an index into a list of PTL syntax structures in the VPS of a PTL syntax structure that applies to a corresponding OLS.

1. A non-transitory computer-readable storage medium comprising:

14. The set of instructions executable by the one or more processors of the device comprises: and in response to the number of the PTL syntax structures being equal to the number of the OLSs, determining the first PTL syntax element to be equal to an index of the OLS to which the PTL syntax structure specified by the first PTL syntax element applies. The non-transitory computer-readable storage medium of claim 13 , further causing the device to:

15. The set of instructions executable by the one or more processors of the device comprises: determining whether the number of PTL syntax structures is equal to one; and and in response to the number of PTL syntax structures being equal to one, skipping decoding of the first PTL syntax element when decoding the VPS. The non-transitory computer-readable storage medium of claim 13 , further causing the device to:

16. The set of instructions executable by the one or more processors of the device comprises: determining the first PTL syntax element to be zero in response to the number of PTL syntax structures being equal to one; The non-transitory computer-readable storage medium of claim 15 , further causing the device to:

17. The set of instructions executable by the one or more processors of the device comprises:

14. The non-transitory computer-readable storage medium of claim 13, further causing the device to decode the first PTL syntax element having a fixed length in the VPS in response to a number of PTL syntax structures being greater than one and different from the number of OLSs.

18. The set of instructions executable by the one or more processors of the device comprises: The bitstream includes the VPS and a sequence parameter set (SPS), and the method comprises: In response to at least one of the PTL syntax structure, a decoded picture buffer (DPB) parameters syntax structure, or a hypothetical reference decoder (HRD) parameters syntax structure being in the SPS, determining whether a first value specifying a maximum number of temporal sub-layers in each coded layer video sequence (CLVS) that references the SPS is greater than one; and decoding a flag configured to control a presence of a syntax element in the DPB parameter syntax structure in the SPS in response to the first value being greater than one; The non-transitory computer-readable storage medium of claim 13 , further causing the device to:

19. The set of instructions executable by the one or more processors of the device comprises: in response to the first value being less than or equal to one, inferring that the flag is equal to zero and decoding one or more syntax elements in the DPB parameter syntax structure without decoding the flag in the SPS.

20. The non-transitory computer-readable storage medium of claim 18, further causing the device to:

20. The set of instructions executable by the one or more processors of the device comprises: determining a value of a first sequence parameter set (SPS) syntax element, the first SPS syntax element specifying an identifier of the VPS referenced by an SPS when the value of the first SPS syntax element is greater than zero; and decoding a second SPS syntax element that specifies a maximum number of temporal sub-layers in each coded layered video sequence (CLVS) that references the SPS, wherein the range of the second SPS syntax element is based on a corresponding VPS syntax element of the VPS referenced by the SPS if the value of the first SPS syntax element is greater than zero, or is based on a fixed value if the value of the first SPS syntax element is equal to zero; The non-transitory computer-readable storage medium of claim 13 , further causing the device to:

Citation Information

Patent Citations

  • Video decoding method and device using the same

    JP2017508417A

  • Profile, Tier, Level for the 0th Output Layer Set in Video Coding

    JP2017523683A

  • Profile, tier, level for the 0-th output layer set in video coding

    US20150373361A1