Method for encoding or decoding a video parameter set or sequence parameter set

By optimizing video parameter set signaling in VVC/H.266, the method addresses inefficiencies in existing standards, achieving reduced bandwidth and memory usage through selective encoding and decoding of syntax elements, aligning with the standard's compression efficiency goals.

JP7911614B2Active Publication Date: 2026-08-26ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025145902
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-26
Filing Date
2025-09-03
Publication Date
2026-08-26
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing video encoding standards face inefficiencies in signaling video parameter sets, leading to increased bandwidth and memory requirements, particularly with evolving standards like VVC/H.266, which aim for higher compression efficiency.

Method used

The method involves determining equal numbers of profile/tier/level (PTL) and decoded picture buffer (DPB) parameter syntax structures, and selectively encoding or decoding syntax elements within video parameter sets (VPS) and sequence parameter sets (SPS) to optimize bitstream signaling, thereby reducing unnecessary information transmission.

Benefits of technology

This approach enhances encoding and decoding efficiency by minimizing redundant signaling, aligning with the higher compression goals of VVC/H.266, thus reducing bandwidth and memory requirements while maintaining video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007911614000001
    Figure 0007911614000001
  • Figure 0007911614000002
    Figure 0007911614000002
  • Figure 0007911614000003
    Figure 0007911614000003
Patent Text Reader

Abstract

To provide a computer-implemented method for encoding video.SOLUTION: A method includes: determining whether a coded video sequence (CVS) contains equal numbers of profile, tier and level (PTL) syntax structures and output layer sets (OLSs); and, in response to the CVS containing the equal numbers of PTL syntax structures and OLSs, coding a bitstream without signaling a first PTL syntax element specifying an index, to a list of PTL syntax structures in a VPS, of a PTL syntax structure that applies to a corresponding OLS in the VPS.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications

[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 994,995, filed Mar. 26, 2020, which is hereby incorporated by reference in its entirety.

[0002] Technical Field

[0002] This disclosure generally relates to video processing, and more specifically, to a video processing method for signaling a video parameter set (VPS) and a sequence parameter set (SPS).

Background Art

[0003] Background

[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce memory and transmission bandwidth, video can be compressed before storage or transmission and restored before display. The compression process is usually referred to as encoding, and the restoration process is usually referred to as decoding. Most commonly, there are various video encoding formats that use standardized video encoding techniques based on prediction, transformation, quantization, entropy encoding, and in - loop filtering. Video encoding standards such as the High - Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify a particular video encoding format, have been developed by standardization organizations. As evolving video encoding techniques are successively adopted in video standards, the encoding efficiency of new video encoding standards becomes even higher.

Summary of the Invention

Means for Solving the Problems

[0004] Summary of the Disclosure

[0004] Embodiments of the present disclosure provide a computer implementation for encoding video. In some embodiments, the method includes determining whether an encoded video sequence (CVS) contains an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS), and encoding a bitstream without signaling a first PTL syntax element that specifies an index of the PTL syntax structures in the VPS to a list of PTL syntax structures in the VPS that apply to the corresponding OLS in the VPS.

[0005]

[0005] In some embodiments, the method includes determining whether an encoded video sequence (CVS) has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS), and encoding the bitstream without signaling a first DPB syntax element that specifies an index of the DPB parameter syntax structures in the VPS to a list of DPB parameter syntax structures in the VPS, in response to the CVS having an equal number of DPB parameter syntax structures and OLS.

[0006]

[0006] In some embodiments, the method includes determining whether at least one of a profile / tier / level (PTL) syntax structure, a decoded picture buffer (DPB) parameter syntax structure, or a virtual reference decoder (HRD) parameter syntax structure is in a bitstream sequence parameter set (SPS), determining whether a first value is greater than 1, the first value specifying the maximum number of time-axis sublayers in an encoded layer video sequence (CLVS) referencing the SPS, and signaling a flag configured to control the presence of syntax elements in the DPB parameter syntax structure within the SPS if at least one of the PTL syntax structure, DPB parameter syntax structure, or HRD parameter syntax structure is in the SPS and the first value is greater than 1.

[0007]

[0007] In some embodiments, this method includes determining the value of a first sequence parameter set (SPS) syntax element, wherein the first SPS syntax element specifies an identifier for a video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero; assigning a range of a second SPS syntax element, based on the corresponding VPS syntax element, that specifies the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) referencing the SPS, in response to the value of the first SPS syntax element being equal to zero; and assigning the range of the second SPS syntax element, which specifies the maximum number of time-axis sublayers in each CLVS referencing the SPS, to be between zero and a fixed value in response to the value of the first SPS syntax element being equal to zero.

[0008]

[0008] In some embodiments, the method includes encoding one or more profile / tier / level (PTL) syntax elements that specify PTL-related information, and signaling one or more PTL syntax elements having a variable length within a video parameter set (VPS) or sequence parameter set (SPS) of a bitstream.

[0009]

[0009] In some embodiments, the method comprises encoding a variable-length video parameter set (VPS) syntax element and signaling the VPS syntax element within the VPS, wherein the VPS syntax element relates to the number of output layer sets (OLS) contained within an encoded video sequence (CVS) that references the VPS.

[0010]

[0010] Embodiments of the present disclosure provide a computer implementation for decoding video. In some embodiments, the method includes receiving a bitstream containing an encoded video sequence (CVS), determining whether the CVS has an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS), and, in response that the number of PTL syntax structures is equal to the number of OLS, skipping the decoding of a first PTL syntax element that specifies an index of the PTL syntax structures in the VPS that apply to the corresponding OLS, against a list of PTL syntax structures in the VPS.

[0011]

[0011] In some embodiments, the method includes receiving a bitstream containing an encoded video sequence (CVS), determining whether the CVS has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS), and, in response that the CVS has an equal number of DPB parameter syntax structures and OLS, skipping the decoding of a first DPB syntax element that specifies an index of the DPB parameter syntax structure that fits in the corresponding OLS against a list of DPB parameter syntax structures in the VPS.

[0012]

[0012] In some embodiments, the method includes receiving a bitstream containing a video parameter set (VPS) and a sequence parameter set (SPS), determining whether a first value specifying the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) referencing the SPS is greater than 1 in response that at least one of a profile / tier / level (PTL) syntax structure, a decoded picture buffer (DPB) parameter syntax structure, or a virtual reference decoder (HRD) parameter syntax structure is in the SPS, and decoding a flag configured to control the presence of syntax elements in the DPB parameter syntax structure in the SPS in response that the first value is greater than 1.

[0013]

[0013] In some embodiments, the method involves determining the value of a first sequence parameter set (SPS) syntax element, wherein the first SPS syntax element, when the value of the first SPS syntax element is greater than zero, specifies an identifier for the video parameter set (VPS) referenced by the SPS, and decodes a second SPS syntax element, which specifies the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) that references the SPS, wherein the range of the second SPS syntax element is based on the corresponding VPS syntax element of the VPS referenced by the SPS when the value of the first SPS syntax element is greater than zero, or based on a fixed value when the value of the first SPS syntax element is equal to zero.

[0014]

[0014] In some embodiments, the method includes receiving a bitstream containing a video parameter set (VPS) or sequence parameter set (SPS), and decoding one or more profile / tier / level (PTL) syntax elements within the VPS or SPS, wherein one or more PTL syntax elements specify PTL-related information.

[0015]

[0015] In some embodiments, the method includes receiving a bitstream containing a video parameter set (VPS) and decoding a variable-length VPS syntax element within the VPS, wherein the VPS syntax element relates to the number of output layer sets (OLS) contained within an encoded video sequence (CVS) that references the VPS.

[0016]

[0016] Embodiments of the present disclosure provide a device. In some embodiments, the device includes a memory configured to store instructions and a processor coupled to the memory, the processor being configured to execute instructions to perform a computer implementation for encoding video. In some embodiments, the device includes a memory configured to store instructions and a processor coupled to the memory, the processor being configured to execute instructions to perform a computer implementation for decoding video.

[0017]

[0017] Embodiments of the present disclosure provide a non-temporary computer-readable storage medium for storing a set of instructions that can be executed by one or more processors of a device to cause the device to perform a method for encoding video.

[0018]

[0018] Embodiments of the present disclosure provide a non-temporary computer-readable storage medium for storing a set of instructions that can be executed by one or more processors of a device to cause the device to perform a method for decoding video.

[0019] Brief explanation of the drawing

[0019] Embodiments and various aspects of the present disclosure are shown in the following detailed description and accompanying drawings. Various features shown in the figures are not drawn to scale. [Brief explanation of the drawing]

[0020] [Figure 1]

[0020] This is a schematic diagram showing the structure of an example video sequence according to some embodiments of the present disclosure. [Figure 2A]

[0021] This is a schematic diagram illustrating an exemplary coding process of a hybrid video coding system according to embodiments of the present disclosure. [Figure 2B]

[0022] This is a schematic diagram illustrating another exemplary coding process for a hybrid video coding system according to embodiments of the present disclosure. [Figure 3A]

[0023] Schematic diagram showing an exemplary decoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 3B]

[0024] Schematic diagram showing another exemplary decoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 4]

[0025] Block diagram of an exemplary device for encoding or decoding video according to some embodiments of the present disclosure. [Figure 5]

[0026] Schematic diagram of an exemplary bitstream according to some embodiments of the present disclosure. [Figure 6A]

[0027] Exemplary encoding syntax table of PTL syntax structure according to some embodiments of the present disclosure is shown. [Figure 6B]

[0028] Exemplary encoding syntax table of DPB parameter syntax structure according to some embodiments of the present disclosure is shown. [Figure 7A] [

[0029] Exemplary encoding syntax table of HRD parameter syntax structure according to some embodiments of the present disclosure is shown. [Figure 7B]

[0030] Another exemplary encoding syntax table of HRD parameter syntax structure according to some embodiments of the present disclosure is shown. [Figure 8A]

[0031] Exemplary encoding syntax table of a part of VPS raw byte sequence payload (RBSP) syntax structure according to some embodiments of the present disclosure is shown. [Figure 8B]

[0031] Exemplary encoding syntax table of a part of VPS raw byte sequence payload (RBSP) syntax structure according to some embodiments of the present disclosure is shown. [Figure 9]

[0032] Exemplary encoding syntax table of a part of SPS RBSP syntax structure according to some embodiments of the present disclosure is shown. [Figure 10A]

[0033] A flowchart of an exemplary video encoding method according to some embodiments of this disclosure is shown. [Figure 10B]

[0034] A flowchart of an exemplary video decoding method corresponding to the video encoding method of Figure 10A, according to some embodiments of the present disclosure, is shown. [Figure 10C]

[0035] The following are some exemplary VPS syntax structures according to some embodiments of this disclosure. [Figure 10D]

[0036] The following are some exemplary VPS syntax structures according to some embodiments of this disclosure. [Figure 11A]

[0037] A flowchart of an exemplary video encoding method according to some embodiments of this disclosure is shown. [Figure 11B]

[0038] A flowchart of an exemplary video decoding method corresponding to the video encoding method in Figure 11A, according to some embodiments of the present disclosure, is shown. [Figure 11C]

[0039] The following are some exemplary VPS syntax structures according to some embodiments of this disclosure. [Figure 11D]

[0040] The following are some exemplary VPS syntax structures according to some embodiments of this disclosure. [Figure 12A]

[0041] A flowchart of an exemplary video encoding method according to some embodiments of this disclosure is shown. [Figure 12B]

[0042] A flowchart of an exemplary video decoding method corresponding to the video encoding method in Figure 12A, according to some embodiments of the present disclosure, is shown. [Figure 12C]

[0043] The following are some exemplary SPS syntax structures according to some embodiments of this disclosure. [Modes for carrying out the invention]

[0021] Detailed explanation

[0044] Herein, we refer in detail to exemplary embodiments illustrated in the accompanying drawings. The following description refers to the accompanying drawings, where the same reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations shown in the following description of exemplary embodiments do not represent all implementations according to the present invention. Rather, they are merely examples of devices and methods according to aspects relating to the present invention as enumerated in the accompanying claims. Specific aspects of this disclosure are described in more detail below. In the event of any conflict between terms and / or definitions incorporated by reference and those provided herein, the terms and definitions provided herein shall prevail.

[0022]

[0045] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Multipurpose Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0023]

[0046] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing techniques that surpass HEVC using the Joint Search Model (JEM) reference software. Because the coding techniques are incorporated into JEM, JEM achieves substantially higher coding performance than HEVC.

[0024]

[0047] The VVC standard is a relatively recent development and continues to incorporate more encoding techniques to deliver better compression performance. VVC is based on the same hybrid video encoding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0025]

[0048] Video is a set of static pictures (or "frames") arranged chronologically to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures chronologically, and a video playback device (e.g., a television, computer, smartphone, tablet computer, video player, or any end-user terminal with display capabilities) can be used to display such pictures chronologically. In some applications, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as supervision, conference hosting, or live broadcasting.

[0026]

[0049] To reduce the memory space and transmission bandwidth required for such applications, video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be performed by software or specialized hardware executed by a processor (e.g., a general-purpose computer processor). The module for compression is generally called an "encoder," and the module for decompression is generally called a "decoder." Encoders and decoders can be collectively called a "codec." Encoders and decoders can be implemented as any of various suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuit mechanisms such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed within a computer-readable medium. Video compression and decompression can be performed using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, or similar. In some applications, a codec can decompress video from a first encoding standard and then recompress the decompressed video using a second encoding standard. In this case, the codec may be referred to as a "transcoder."

[0027]

[0050] A video encoding process can identify and maintain useful information that can be used to reconstruct a picture, while ignoring information that is not important for reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process may be called "lossy." Otherwise, it may be called "lossy." Most encoding processes are lossy, which is a trade-off to reduce the required memory space and transmission bandwidth.

[0028]

[0051] Useful information about the picture being encoded (referred to as the "current picture") includes changes relative to the reference picture (e.g., a previously encoded and reconstructed picture). Such changes can include changes in pixel position, brightness, or color, with position being the most important. Changes in the position of groups of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.

[0029]

[0052] A picture encoded without referencing another picture (i.e., it is its own reference picture) is called an "I picture". A picture encoded using a previous picture as a reference picture is called a "P picture". A picture encoded using both a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B picture".

[0030]

[0053] Figure 1 shows the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video sequence 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface for receiving video from a video content provider (e.g., a video broadcast transceiver).

[0031]

[0054] As shown in Figure 1, the video sequence 100 may include a series of pictures arranged chronologically along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with further pictures between pictures 106 and 108. In Figure 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102, as indicated by the arrows. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately before or after that picture. For example, the reference picture of picture 104 may be the picture before picture 102. Please note that the reference pictures 102-106 are merely examples, and this disclosure does not limit the embodiments of the reference pictures to the examples shown in Figure 1.

[0032]

[0055] Typically, due to the computational complexity of such tasks, video codecs do not encode or decode an entire picture at once. Rather, video codecs can divide a picture into basic segments and encode or decode the picture segment by segment. Such basic segments are referred to in this disclosure as basic processing units ("BPUs"). For example, structure 110 in Figure 1 shows an exemplary structure of a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, their boundaries shown as dashed lines. In some embodiments, basic processing units may be referred to as "macroblocks" in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as "encoding tree units" ("CTUs") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have a variable size in the picture, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for the picture based on a balance between encoding efficiency and the level of detail that should be maintained in the basic processing unit.

[0033]

[0056] A basic processing unit can be a logical unit that can contain groups of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luminance component (Y) representing colorless luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chroma components may have the same size as the basic processing unit. The luminance and chroma components may be referred to as an "encoded tree block" ("CTB") in some video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be repeated on each of its luminance and chroma components.

[0034]

[0057] Video encoding involves multiple computational stages, examples of which are shown in Figures 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still be too large for processing and can therefore be further divided into segments referred to in this disclosure as “basic processing subunits.” In some embodiments, a basic processing subunit may be referred to as a “block” in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as an “encoding unit” (“CU”) in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing subunit may be the same size as or smaller than a basic processing unit. Like a basic processing unit, a basic processing subunit is a logical unit that can contain groups of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing subunit can be repeated on each of its luma and chroma components. Note that such divisions can be carried out to further levels as needed for processing. Also note that different stages can divide the basic processing unit using different methods.

[0035]

[0058] For example, during the mode determination phase (an example of which is shown in Figure 2B), the encoder can determine which prediction mode (e.g., intrapicture prediction or interpicture prediction) to use for the basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs, as in the case of H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each basic processing subunit.

[0036]

[0059] As another example, in the prediction phase (an example of which is shown in Figures 2A and 2B), the encoder can perform prediction calculations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and prediction calculations can be performed at that level.

[0037]

[0060] As another example, in the conversion stage (an example of which is shown in Figures 2A and 2B), the encoder can perform conversion operations for residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as "conversion blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and conversion operations can be performed at that level. Note that the division method for the same basic processing subunit may differ in the prediction and conversion stages. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and conversion blocks of the same CU may have different sizes and numbers.

[0038]

[0061] In the structure 110 of Figure 1, the basic processing unit 112 is further divided into 3x3 basic processing subunits, their boundaries shown as dotted lines. Different basic processing units of the same picture may be divided into basic processing subunits in different ways.

[0039]

[0062] In some implementations, to bring parallel processing and error tolerance to video encoding and decoding, a picture can be divided into processing regions, so that the encoding or decoding process does not have to rely on information from any other regions of the picture for any given region of the picture. In other words, each region of the picture can be processed independently. This allows the codec to process different regions of the picture in parallel, thereby increasing encoding efficiency. Furthermore, if the data in a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error tolerance. Some video encoding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC offer two types of regions: "slices" and "tiles". It should also be noted that different pictures in video sequence 100 may have different partitioning schemes for dividing the picture into regions.

[0040]

[0063] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, their boundaries shown as solid lines within structure 110. Region 114 contains four basic processing units. Regions 116 and 118 each contain six basic processing units. Note that the basic processing units, basic processing subunits, and regions of structure 110 in Figure 1 are merely examples, and this disclosure does not limit its embodiments.

[0041]

[0064] Figure 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As shown in Figure 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 in Figure 1, video sequence 202 may include a set of pictures (referred to as “original pictures”) arranged in chronological order. Similar to structure 110 in Figure 1, each original picture in video sequence 202 may be divided by the encoder into a basic processing unit, basic processing subunit, or region for processing. In some embodiments, the encoder may perform process 200A at the level of a basic processing unit for each original picture in video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in a single iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for each region of the original picture in the video sequence 202 (e.g., regions 114-118).

[0042]

[0065] In Figure 2A, the encoder can supply the basic processing unit (referred to as the "original BPU") of the original picture of the video sequence 202 to the prediction stage 204, generating prediction data 206 and prediction BPU 208. The encoder can subtract the prediction BPU 208 from the original BPU to generate residual BPU 210. The encoder can supply the residual BPU 210 to the conversion stage 212 and the quantization stage 214, generating quantization conversion coefficients 216. The encoder can supply the prediction data 206 and quantization conversion coefficients 216 to the binary coding stage 226, generating video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after the quantization stage 214, the encoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224, which will be used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the “reconstruction path”. The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.

[0043]

[0066] The encoder can iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and to generate a prediction criterion 224 (in the reconstruction path) for encoding the next original BPU of the original picture. After encoding all the original BPUs of the original picture, the encoder can proceed to encode the next picture in the video sequence 202.

[0044]

[0067] Referring to process 200A, the encoder can receive a video sequence 202 generated by a video acquisition device (e.g., a camera). As used herein, the term “receive” can mean receiving, inputting, acquiring, obtaining, getting, reading, accessing, or any act by any means for inputting data.

[0045]

[0068] In prediction stage 204, in the current iteration, the encoder can receive the original BPU and prediction criterion 224, perform prediction calculations, and generate prediction data 206 and prediction BPU 208. The prediction criterion 224 may be generated from the reconstruction path of a previous iteration of process 200A. The objective of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, which can be used to reconstruct the original BPU as prediction BPU 208 from the prediction data 206 and prediction criterion 224.

[0046]

[0069] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is generally slightly different from the original BPU. To record such differences, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate the residual BPU 210. For example, the encoder can subtract the pixel values ​​(e.g., grayscale values ​​or RGB values) of the predicted BPU 208 from the corresponding pixel values ​​of the original BPU. Each pixel of the residual BPU 210 may have a residual value resulting from such a subtraction between the original BPU and the corresponding pixels of the predicted BPU 208. Compared to the original BPU, the predicted data 206 and residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.

[0047]

[0070] To further compress the residual BPU 210, in the transformation step 212, the encoder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns," each basis pattern associated with a "transformation coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent the frequency component of the residual BPU 210's changes (e.g., the frequency of the brightness changes). None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition allows the changes in the residual BPU 210 to be decomposed into the frequency domain. Such a decomposition is analogous to the discrete Fourier transform of a function, in which case the basis patterns are analogous to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are analogous to the coefficients associated with the basis functions.

[0048]

[0071] Different transformation algorithms can use different basis patterns. For example, various transformation algorithms such as discrete cosine transform, discrete sine transform, or similar can be used in transformation stage 212. The transformation in transformation stage 212 is inversely operable. That is, the encoder can recover the residual BPU 210 by performing the inverse operation of the transformation (referred to as "inverse transform"). For example, to recover the pixels of the residual BPU 210, the inverse transform can be performed by multiplying the values ​​of the corresponding pixels in the basis pattern by their respective associated coefficients, adding the products, and generating a weighted sum. For video coding standards, both the encoder and decoder can use the same transformation algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basis pattern from the encoder. Compared to the residual BPU 210, the transformation coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU210 is further compressed.

[0049]

[0072] The encoder can further compress the conversion coefficients in the quantization stage 214. In the conversion process, different basis patterns can represent different change frequencies (e.g., brightness change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore information about high-frequency changes without causing significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate quantization conversion coefficients 216 by dividing each conversion coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some conversion coefficients of high-frequency basis patterns may be converted to 0, and conversion coefficients of low-frequency basis patterns may be converted to smaller integers. The encoder can ignore quantization conversion coefficients 216 that are 0, thereby further compressing the conversion coefficients. The quantization process is also inversely operable, in which case the quantization conversion coefficients 216 can be reconstructed into conversion coefficients in the inverse operation of quantization (referred to as "inverse quantization").

[0050]

[0073] Because the encoder rounds off the remainder of such divisions, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantization conversion coefficient 216 may require. To obtain different levels of information loss, the encoder can use different values ​​for the quantization parameters or any other parameters of the quantization process.

[0051]

[0074] In the binary coding stage 226, the encoder can encode the prediction data 206 and the quantization conversion coefficients 216 using binary coding techniques such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantization conversion coefficients 216, the encoder can encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of conversion in the conversion stage 212, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), or the like. The encoder can generate a video bitstream 228 using the output data from the binary coding stage 226. In some embodiments, the video bitstream 228 can be further packetized for network transmission.

[0052]

[0075] Referring to the reconstruction path of process 200A, in the inverse quantization step 218, the encoder can perform inverse quantization on the quantization transformation coefficients 216 to generate reconstruction transformation coefficients. In the inverse transformation step 220, the encoder can generate reconstruction residual BPU 222 based on the reconstruction transformation coefficients. The encoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction criterion 224, which will be used in the next iteration of process 200A.

[0053]

[0076] It should be noted that other variations of process 200A may also be used to encode the video sequence 202. In some embodiments, the steps of process 200A may be performed in a different order by the encoder. In some embodiments, one or more steps of process 200A may be combined into a single step. In some embodiments, a single step of process 200A may be divided into multiple steps. For example, the transformation step 212 and the quantization step 214 may be combined into a single step. In some embodiments, process 200A may include additional steps. In some embodiments, process 200A may omit one or more steps in Figure 2A.

[0054]

[0077] Figure 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used with an encoder compliant with a hybrid video encoding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode determination stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0055]

[0078] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels from one or more already encoded neighboring BPUs within the same picture to predict the current BPU. That is, the prediction criterion 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions from one or more already encoded pictures to predict the current BPU. That is, the prediction criterion 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0056]

[0079] Referring to process 200B, in the forward path, the encoder performs prediction calculations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-prediction. For the original BPU of the picture being encoded, the prediction criterion 224 may include one or more neighboring BPUs encoded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate a prediction BPU 208 by extrapolating neighboring BPUs. The extrapolation technique may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or similar. In some embodiments, the encoder may perform extrapolation at the pixel level, for example, by extrapolating the values ​​of the corresponding pixels for each pixel of the prediction BPU 208. The adjacent BPUs used for extrapolation can be positioned relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below left, below right, above left, or above right of the original BPU), or any direction defined in the video encoding standard used. For intra-prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the extrapolation parameters, the orientation of the adjacent BPUs used relative to the original BPU, or similar.

[0057]

[0080] As another example, in the temporal prediction stage 2044, the encoder can perform interpretation. For the original BPU of the current picture, the prediction criterion 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs for the same picture have been generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for matching regions within the range of the reference picture (referred to as the "search window"). The location of the search window in the reference picture may be determined based on the location of the original BPU of the current picture. For example, the search window may have its center in the reference picture at a location with the same coordinates as the original BPU in the current picture and may extend outward over a predetermined distance. When the encoder identifies a region in the search window similar to the original BPU (for example, by using a pixel recursion algorithm, a block matching algorithm, or similar), the encoder can determine such a region to be a matching region. The matching region may have different dimensions from the original BPU (for example, smaller than, equal to, larger than, or of a different shape than the original BPU). Because the reference picture and the current picture are temporally separated in the timeline (for example, as shown in Figure 1), the matching region can be considered to "move" to the location of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector". When multiple reference pictures are used (for example, as picture 106 in Figure 1), the encoder can search for a matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching region for each matching reference picture.

[0058]

[0081] Motion estimation can be used to identify various types of motion, such as translation, rotation, zooming, or similar. For interpretation, the prediction data 206 may include, for example, the location of the matching region (e.g., coordinates), the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, or similar.

[0059]

[0082] To generate a predicted BPU 208, the encoder can perform a “motion compensation” operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on prediction data 206 (e.g., motion vectors) and prediction criteria 224. For example, the encoder can move the matching region of a reference picture according to the motion vector, in which case the encoder can predict the original BPU of the current picture. When multiple reference pictures are used (e.g., picture 106 in Figure 1), the encoder can move the matching region of each reference picture according to its respective motion vector and average the pixel values ​​of the matching region. In some embodiments, if the encoder weights the pixel values ​​of the matching region of each matching reference picture, the encoder can add the weighted sum of the pixel values ​​to the moved matching region.

[0060]

[0083] In some embodiments, interpretation can be unidirectional or bidirectional. Unidirectional interpretation can use one or more reference pictures in the same time direction relative to the current picture. For example, picture 104 in Figure 1 is a unidirectional interpretation picture in which the reference picture (e.g., picture 102) precedes picture 104. Bidirectional interpretation can use one or more reference pictures in both time directions relative to the current picture. For example, picture 106 in Figure 1 is a bidirectional interpretation picture in which the reference pictures (e.g., pictures 104 and 108) are in both time directions relative to picture 104.

[0061]

[0084] Still referring to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, in the mode determination stage 230, the encoder may select a prediction mode (e.g., either intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique. In this technique, the encoder may select a prediction mode to minimize a cost function value that depends on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate the corresponding prediction BPU 208 and prediction data 206.

[0062]

[0085] If the intra-prediction mode is selected in the forward path within the reconstruction path of process 200B, after generating the prediction criterion 224 (e.g., the current BPU encoded and reconstructed in the current picture), the encoder can directly feed the prediction criterion 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the inter-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current picture with all BPUs encoded and reconstructed), the encoder can feed the prediction criterion 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction criterion 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by inter-prediction. The encoder can apply various loop filtering techniques in the loop filter stage 232, such as deblocking, sample-adaptive offset (SAO), adaptive loop filtering (ALF), or similar. Loop-filtered reference pictures may be stored in buffer 234 (or “Decoded Picture Buffer (DPB)”) for later use (e.g., to be used as interpredictive reference pictures for future pictures in video sequence 202). The encoder may store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with quantization transformation coefficients 216, prediction data 206, and other information in the binary coding stage 226.

[0063]

[0086] Figure 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to the compression process 200A in Figure 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in Figures 2A-2B), the video stream 304 is generally not identical to the video sequence 202. Similar to processes 200A and 200B in Figures 2A-2B, the decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded within the video bitstream 228. For example, the decoder can perform process 300A in an iterative manner, in which case the decoder can decode the basic processing unit in a single iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for each region of picture encoded in the video bitstream 228 (e.g., regions 114-118).

[0064]

[0087] In Figure 3A, the decoder can supply a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode this portion into prediction data 206 and quantization conversion coefficients 216. The decoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The decoder can supply the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224. In some embodiments, the prediction criterion 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can supply the prediction criterion 224 to the prediction stage 204 to perform the prediction calculation in the next iteration of process 300A.

[0065]

[0088] The decoder can iteratively perform process 300A to decode each encoding BPU of the encoded picture and generate a prediction criterion 224 for encoding the next encoding BPU of the encoded picture. After decoding all encoding BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.

[0066]

[0089] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and quantization conversion coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, type of conversion, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if the video bitstream 228 is transmitted in the form of packets over the network, the decoder can depacket the video bitstream 228 before supplying it to the binary decoding stage 302.

[0067]

[0090] Figure 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used with a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0068]

[0091] In process 300B, the prediction data 206 decoded by the decoder from binary decoding stage 302 for the encoding base processing unit ("current BPU") of the encoded picture being decoded ("current picture") may include various types of data, depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra-prediction is used by the encoder to encode the current BPU, the prediction data 206 may include intra-prediction, parameters of the intra-prediction operation, or prediction mode indicators (e.g., flag values) indicating the same. Parameters of the intra-prediction operation may include, for example, the locations (e.g., coordinates) of one or more adjacent BPUs used as reference, the size of the adjacent BPUs, extrapolation parameters, the orientation of the adjacent BPUs relative to the original BPU, or the same. As another example, if inter-prediction is used by the encoder to encode the current BPU, the prediction data 206 may include inter-prediction, parameters of the inter-prediction operation, or prediction mode indicators (e.g., flag values) indicating the same. The parameters for the interpretation calculation may include, for example, the number of reference pictures associated with the current BPU, the weights associated with each reference picture, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each matching region, or similar.

[0069]

[0092] Based on the prediction mode indicator, the decoder can determine whether to perform a spatial prediction (e.g., intra-prediction) in the spatial prediction stage 2042 or a temporal prediction (e.g., inter-prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal predictions are shown in Figure 2B and will not be repeated below. After performing such spatial or temporal predictions, the decoder can generate a prediction BPU 208. The decoder can then add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction criterion 224, as described in Figure 3A.

[0070]

[0093] In process 300B, the decoder may supply the prediction criterion 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform the prediction calculation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-prediction in the spatial prediction stage 2042, after generating the prediction criterion 224 (e.g., the decoded current BPU), the decoder may supply the prediction criterion 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter-prediction in the temporal prediction stage 2044, after generating the prediction criterion 224 (e.g., the reference picture with all BPUs decoded), the encoder may supply the prediction criterion 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply the loop filter to the prediction criterion 224 in the manner described in Figure 2B. Loop-filtered reference pictures may be stored in buffer 234 (e.g., a decoded picture buffer (DPB) in computer memory) for later use (e.g., to be used as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data 206 may further include loop filter parameters (e.g., loop filter strength) when a prediction mode indicator indicates that inter-prediction was used to encode the current BPU. The reconstructed picture from buffer 234 may also be transmitted to a display such as a TV, PC, smartphone, or tablet for viewing by an end user.

[0071]

[0094] Figure 4 is a block diagram of an exemplary device 400 for encoding or decoding video according to an embodiment of the present disclosure. As shown in Figure 4, the device 400 may include a processor 402. When the processor 402 executes instructions as described herein, the device 400 can become a specialized machine for video encoding or decoding. The processor 402 may be any kind of circuit mechanism having the ability to manipulate or process information. For example, the processor 402 may include a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a composite programmable logic unit (CPLD), a field-programmable gate array (FPGA), a system-on-a-chip (SoC), an application-specific integrated circuit (ASIC), or any number or any combination of the same. In some embodiments, the processor 402 may also be a set of processors grouped as a single logical component. For example, as shown in Figure 4, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0072]

[0095] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for carrying out steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and data for processing (e.g., via bus 410), execute the program instructions, and perform arithmetic or operations on the data for processing. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include random-access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, CompactFlash® (CF) cards, or any number or any combination of the like. Memory 404 can also be a group of memories grouped together as a single logical component (not shown in Figure 4).

[0073]

[0096] Bus 410 may be a communication device that transfers data between internal components of the device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or similar.

[0074]

[0097] To facilitate explanation without creating ambiguity, the processor 402 and other data processing circuits are collectively referred to as “data processing circuits” in this disclosure. Data processing circuits may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. In addition, data processing circuits may be a single, standalone module, or may be fully or partially integrated with any other component of the device 400.

[0075]

[0098] The device 400 may further include a network interface 406 for providing wired or wireless communication to a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the same). In some embodiments, the network interface 406 may include a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a Near Field Communication ("NFC") adapter, a cellular network chip, or any combination of any number of the same.

[0076]

[0099] In some embodiments, the device 400 may optionally further include a peripheral interface 408 for providing connectivity to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to video archives), or similar.

[0077]

[0100] It should be noted that the video codec (for example, the codec that performs processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules within the device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of the device 400, such as special data processing circuits (e.g., FPGA, ASIC, NPU, or similar).

[0078]

[0101] Figure 5 is a schematic diagram of an example of a bitstream 500 encoded by an encoder according to some embodiments of the present disclosure. In some embodiments, the structure of bitstream 500 can be applied to the video bitstream 228 shown in Figures 2A-2B and 3A-3B. In Figure 5, bitstream 500 includes a video parameter set (VPS) 510, a sequence parameter set (SPS) 520, a picture parameter set (PPS) 530, a picture header 540, and slices 550-570, separated by synchronization markers M1-M7. Slices 550-570 each include a corresponding header block (e.g., header 552) and a data block (e.g., data 554), each data block containing one or more CTUs (e.g., CTU1-CTUn in data 554).

[0079]

[0102] According to some embodiments, a bitstream 500, which is a network abstraction layer (NAL) unit or a sequence of bits in byte stream format, forms one or more encoded video sequences (CVS). A CVS includes one or more encoded layer video sequences (CLVS). In some embodiments, a CLVS is a sequence of picture units (PUs), each PU containing one encoded picture. In particular, a PU includes zero or one picture header NAL unit (e.g., picture header 540), which includes a picture header syntax structure as a payload, one encoded picture containing one or more video encoding layer (VCL) NAL units, and optionally one or more other non-VCL NAL units. A VCL NAL unit is a collective term for an encoded slice NAL unit (e.g., slices 550-570), and a subset of NAL units having reserved values ​​of the type of NAL unit classified as a VCL NAL unit in some embodiments. An encoded slice NAL unit includes a slice header and a slice data block (e.g., header 552 and data 554).

[0080]

[0103] In other words, in some embodiments of this disclosure, a layer may be a set of video coding layer (VCL) NAL units having a specific value for the NAL layer ID, and associated non-VCL NAL units. Among these layers, inter-layer predictions can be applied between various layers to achieve high compression performance.

[0081]

[0104] In some embodiments, an output layer set (OLS) may be specified to support the decoding of some, but not all, layers. The OLS is a set of layers that includes a set of specified layers, in which one or more layers in the set of layers are specified as output layers. Thus, the OLS may include one or more output layers and other layers necessary to decode the output layers for inter-layer prediction.

[0082]

[0105] In some embodiments, “profiles,” “tiers,” and “levels” (collectively known as “PTLs”) are used to specify constraints on the bitstream, and therefore limits on the capabilities required to decode the bitstream. Profiles, tiers, and levels can also be used to indicate points of interoperability between individual decoder implementations. A “profile” is a subset of bitstream syntax that specifies a subset of algorithmic features and limitations that can be supported by decoders conforming to that profile. Within the boundaries imposed by the syntax of a given profile, it is possible to determine variations in encoder and decoder performance depending on the values ​​taken by the syntax elements in the bitstream, such as the size of the decoded picture specification. In some applications, it may not be practical or economical to implement a decoder that can handle all hypothetical uses of the syntax within a particular profile.

[0083]

[0106] In addition, "tiers" and "levels" are specified within each profile. Tier levels are a set of constraints imposed on the values ​​of syntax elements in the bitstream. These constraints may be simple restrictions on values. Alternatively, they may take the form of constraints on arithmetic combinations of values ​​(e.g., picture width × picture height × number of decoded pictures per second). In some embodiments, the same set of tier and level definitions is used for all profiles. Some implementations may support different tiers and different levels within tiers for each supported profile. With respect to a given profile, tier levels generally correspond to the processing load and memory capabilities of a particular decoder. Levels specified for lower tiers may be more constrained than levels specified for higher tiers.

[0084]

[0107] Figure 6A shows an exemplary coded syntax table, highlighted in bold, of a PTL syntax structure 600A signaled within a VPS or SPS, according to some embodiments of the present disclosure. In some embodiments, PTL-related information for each operation point identified by the parameters TargetOlsIdx and Htid may be indicated by parameters 610A and 620A (e.g., general_profile_idc and general_tier_flag) within the PTL syntax structure 600A signaled within the video parameter set (VPS) or sequence parameter set (SPS), and parameter 630A (e.g., sublayer_level_idc[Htid]) found within or derived from the PTL syntax structure. TargetOlsIdx is a variable used to identify the OLS index of the target OLS to be decoded, and Htid is a variable used to identify the top-level time-axis sublayer to be decoded.

[0085]

[0108] As described in the paragraph above, the Decoded Picture Buffer (DPB) includes a picture storage buffer for storing the decoded picture. Each picture storage buffer may contain a decoded picture that is marked "for reference use" or held for future output. In some embodiments, the process is specified to be applied sequentially, starting from the lowest layer in the OLS and in ascending order of the nuh_layer_id values ​​of the layers in the OLS, applied separately for each layer, where nuh_layer_id is a parameter that specifies the identifier of the layer to which the VCL NAL unit belongs or the identifier of the layer to which the non-VCL NAL unit belongs. In some embodiments, the value of nuh_layer_id should be the same with respect to the VCL NAL unit of the encoded picture or the PU. The process includes the process of outputting and removing the picture from the DPB before decoding the current picture, the process of marking and storing the current decoded picture, and the process of additional bumping.

[0086]

[0109] Figure 6B shows an exemplary coded syntax table of a DPB parameter syntax structure 600B signaled within a VPS or SPS, highlighted in bold, according to some embodiments of the present disclosure. As shown in Figure 6B, the DPB parameters 610B, 620B, and 630B, which are necessary parameters for applying the above process and checking the conformance of the bitstream, may be found within or derived from the DPB parameter syntax structure 600B signaled within a VPS or SPS. These parameters may include max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid].

[0087]

[0110] In some embodiments, parameter 610B (e.g., max_dec_pic_buffering_minus1[i]) plus 1 specifies the maximum required size of the DPB in units of picture storage buffers, when Htid is equal to index i. The value of parameter 610B can be in the range of 0 or greater and (MaxDpbSize-1) or less. MaxDpbSize is a parameter that specifies the maximum decoded picture buffer size in units of picture storage buffers. If index i is greater than 0, max_dec_pic_buffering_minus1[i] shall be greater than or equal to max_dec_pic_buffering_minus1[i-1]. If there is no max_dec_pic_buffering_minus1[i] for index i in the range of 0 or greater and (maxSubLayersMinus1-1) or less, then parameter 610B is inferred to be equal to max_dec_pic_buffering_minus1[maxSubLayersMinus1] because subLayerInfoFlag is equal to 0.

[0088]

[0111] Parameter 620B (e.g., max_num_reorder_pics[i]) specifies the maximum number of pictures allowed in the OLS, which, if Htid is equal to index i, may precede any picture in the OLS in decoding order and follow that picture in output order. The value of parameter 620B can be between 0 and max_dec_pic_buffering_minus1[i]. If index i is greater than 0, parameter 620B can be greater than or equal to max_num_reorder_pics[i-1]. If parameter 620B is not found for an index i between 0 and maxSubLayersMinus1 minus 1, parameter 620B is inferred to be equal to max_num_reorder_pics[maxSubLayersMinus1] because subLayerInfoFlag is equal to 0.

[0089]

[0112] Parameter 630B can be used to calculate the value of MaxLatencyPictures[i], which specifies the maximum number of pictures in the OLS. If Htid is equal to index i, it can precede any picture in the OLS in output order and follow that picture in decoding order. If parameter 630B is equal to 0, the corresponding limit is not expressed. If parameter 630B is not equal to 0, the value of MaxLatencyPictures[i] can be calculated based on the following formula. MaxLatencyPictures[i]=max_num_reorder_pics[i]+max_latency_increase_plus1[i]-1 (Formula 1)

[0090]

[0113] The value of parameter 630B is 0 or greater, (2 32 -2) It can be within the following range. If there is no parameter 630B for index i in the range of 0 or greater and maxSubLayersMinus1 minus 1 or less, then parameter 630B is inferred to be equal to max_latency_increase_plus1[maxSubLayersMinus1] because subLayerInfoFlag is equal to 0.

[0091]

[0114] The Virtual Reference Decoder (HRD) is a virtual decoder model that specifies constraints on the diversity of compliant NAL unit streams or compliant byte streams that an encoding process may produce, and can be used to check the conformance of bitstreams and decoders. Two types of bitstreams or bitstream subsets are subject to HRD conformance checks for VVCs. The first type, called Type I bitstream, is a NAL unit stream containing VCL NAL units and NAL units, where for all access units (AUs) in the bitstream, nal_unit_type is equal to the filled data NAL unit (FD_NUT). A second type, named Type II bitstream, includes, in addition to the VCL NAL units and fill data NAL units for all AUs in the bitstream, (1) additional non-VCL NAL units other than the fill data NAL units, or (2) at least one of the leading_zero_8bits syntax elements, zero_byte syntax element, start_code_prefix_one_3bytes syntax element and trailing_zero_8bits syntax element that form a byte stream from the NAL unit stream. In some embodiments, leading_zero_8bits is a byte equal to 0x00, zero_byte is a single byte equal to 0x00, start_code_prefix_one_3bytes, called the start code prefix, is a fixed value sequence of 3 bytes equal to 0x000001, and trailing_zero_8bits is a byte equal to 0x00. In some embodiments, the leading_zero_8bits syntax element is in the first byte stream NAL unit of the bitstream.According to the NAL unit syntax structure, any byte equal to 0x00 preceding the 4-byte sequence 0x00000001 (which should be interpreted as a zero_byte syntax element followed by a start_code_prefix_one_3bytes syntax element) is considered a trailing_zero_8bits syntax element, which is part of the preceding byte stream NAL unit.

[0092]

[0115] Figures 7A and 7B show exemplary coded syntax tables, highlighted in bold, of HRD parameter syntax structures 700A and 700B, respectively, according to some embodiments of the present disclosure. 700A and 700B may be signaled within a VPS or SPS. Two sets of HRD parameters (NAL HRD parameters and VCL HRD parameters) may be used, as shown in Figures 7A and 7B. HRD parameters may be signaled by the generic HRD parameter syntax structure 700A in Figure 7A and the OLS HRD parameter syntax structure 700B in Figure 7B, which may be part of a VPS or part of an SPS.

[0093]

[0116] As described above, the OLS, PTL syntax structure (e.g., PTL syntax structure 600A in Figure 6A), DPB parameter syntax structure (e.g., DPB parameter syntax structure 600B in Figure 6B), and HRD parameter syntax structure (e.g., syntax structures 700A and 700B in Figures 7A and 7B) can be signaled within the VPS or SPS. For each of these syntax structures, the VPS or SPS signals the set of syntax structures applicable to each OLS, and the index of the syntax structure.

[0094]

[0117] Figure 8 shows an exemplary encoded syntax table of a portion of the VPS Raw Byte Sequence Payload (RBSP) syntax structure 800 signaled within the VPS, highlighted in bold, according to some embodiments of the present disclosure. As shown in Figure 8, the VPS parameter 810 (vps_max_layers_minus1) plus 1 specifies the maximum allowed number of layers in each CVS referencing the VPS. The VPS parameter 812 (vps_max_sublayers_minus1) plus 1 specifies the maximum number of time-axis sublayers that may be in a layer in each CVS referencing the VPS. In some embodiments, the value of the VPS parameter 812 is in the range of 0 or greater and less than or equal to a default static value (e.g., 6). The VPS parameter 814 (each_layer_is_an_ols_flag) equal to 1 indicates that each OLS contains one layer and that each layer in the CVS referencing the VPS itself is an OLS having a single contained layer which is the only output layer. The VPS parameter 814 equal to 0 indicates that an OLS may contain multiple layers. If VPS parameter 810 is equal to 0, the value of VPS parameter 814 is inferred to be equal to 1. In some embodiments, if the corresponding flag indicates that one or more of the layers specified by the VPS can use inter-layer prediction (for example, if vps_all_independent_layers_flag is equal to 0), the value of VPS parameter 814 is inferred to be equal to 0.

[0095]

[0118] In some embodiments, the value of VPS parameter 816 (ols_mode_idc) can be in the range of 0 to 2, and a value of 3 for VPS parameter 816 can be reserved for future use. In some embodiments, VPS parameter 816 equal to 0 specifies that the total number of OLS specified by the VPS (TotalNumOlss) is equal to the value of VPS parameter 810 plus 1, the i-th OLS contains layers with layer indices from 0 to index i, and for each OLS, only the top layer within the OLS is output. VPS parameter 816 equal to 1 specifies that the total number of OLS specified by the VPS is equal to VPS parameter 810 plus 1, the i-th OLS contains layers with layer indices from 0 to index i, and for each OLS, all layers within the OLS are output. A VPS parameter 816 equal to 2 specifies that the total number of OLS specified by the VPS is clearly signaled using VPS parameter 818, that the output layer for each OLS is clearly signaled, and that other layers are direct or indirect reference layers of the output layer of the OLS.

[0096]

[0119] In some embodiments, the corresponding flag specifies that all layers specified by the VPS are encoded independently without using inter-layer prediction (for example, vps_all_independent_layers_flag is equal to 1), and if VPS parameter 814 is equal to 0, then the value of VPS parameter 816 is inferred to be equal to 2.

[0097]

[0120] If VPS parameter 816 is equal to 2, VPS parameter 818 (num_output_layer_sets_minus1) is signaled to indicate the total number of OLS minus 1. In other words, VPS parameter 818 plus 1 specifies the total number of OLS specified by VPS when VPS parameter 816 is equal to 2 (e.g., the variable TotalNumOlss). The variable TotalNumOlss can be derived and calculated from the code as follows. if(vps_max_layers_minus1==0) TotalNumOlss=1 else if(each_layer_is_an_ols_flag||ols_mode_idc==0||ols_mode_idc==1) TotalNumOlss=vps_max_layers_minus1+1 else if(ols_mode_idc==2) TotalNumOlss=num_output_layer_sets_minus1+1

[0098]

[0121] In addition, if VPS parameter 816 is equal to 2, VPS parameter 820 (ols_output_layer_flag[i][j]) equal to 1 specifies that the j-th layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[j]) is the output layer of the i-th OLS, and VPS parameter 820 equal to 0 specifies that the j-th layer is not the output layer of the i-th OLS. In other words, by signaling the corresponding flag (e.g., VPS parameter 820), each OLS can be defined for each layer in CVS to indicate whether a layer is an output layer or not within a given OLS.

[0099]

[0122] The VPS parameter 822 (vps_num_ptls_minus1) plus 1 specifies the number of PTL syntax structures within the VPS. The value of VPS parameter 822 can be less than the total number of OLS (i.e., TotalNumOlss).

[0100]

[0123] A VPS parameter 824 (pt_present_flag[i]) equal to 1 indicates that the profile, tier, and general constraint information is present in the i-th PTL syntax structure within the VPS. A VPS parameter 824 equal to 0 indicates that the profile, tier, and general constraint information is not present in the i-th PTL syntax structure within the VPS. In some embodiments, the value of pt_present_flag[0] is inferred to be equal to 1. If VPS parameter 824 is equal to 0, the profile, tier, and general constraint information for the i-th PTL syntax structure within the VPS is inferred to be the same as that for the (i-1)-th PTL syntax structure within the VPS.

[0101]

[0124] The VPS parameter 826 (ptl_max_temporal_id[i]) specifies the temporal identifier (TemporalId) of the top-level sublayer representation in the i-th PTL syntax structure within the VPS, where level information resides. The value of VPS parameter 826 can be between 0 and VPS parameter 812. If VPS parameter 812 is equal to 0, the value of VPS parameter 826 is inferred to be equal to 0. If VPS parameter 812 is greater than 0 and the corresponding flag (e.g., vps_all_layers_same_num_sublayers_flag) is equal to 1, the value of VPS parameter 826 is inferred to be equal to VPS parameter 812.

[0102]

[0125] In some embodiments, the VPS parameter 828 (vps_ptl_alignment_zero_bit) may be equal to 0. In some embodiments, the VPS parameter 830 (ols_ptl_idx[i]) specifies the index of the PTL syntax structure that fits the i-th OLS to a list of PTL syntax structures in the VPS. If present, the value of VPS parameter 830 can be between 0 and VPS parameter 822. If VPS parameter 828 is equal to 0, the value of VPS parameter 830 is inferred to be equal to 0.

[0103]

[0126] If NumLayersInOls[i], a variable specifying the number of layers in the i-th OLS, is equal to 1, then the PTL syntax structure applicable to the i-th OLS is also in the SPS referenced by the layers in the i-th OLS. In some embodiments, if the number of layers in the i-th OLS is equal to 1, then it is a bitstream compatibility requirement that the PTL syntax structure signaled in the VPS and SPS with respect to the i-th OLS may be identical.

[0104]

[0127] The VPS parameter 832 (vps_num_dpb_params) specifies the number of DPB parameter syntax structures within the VPS. In some embodiments, the value of VPS parameter 832 can be in the range of 0 to 16. If none of these apply, the value of VPS parameter 832 is inferred to be equal to 0.

[0105]

[0128] The VPS parameter 834 (vps_sublayer_dpb_params_present_flag) is used to control the presence of syntax elements (e.g., max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[]) within the DPB parameter syntax structure in the VPS. If not present, VPS parameter 834 is inferred to be equal to 0.

[0106]

[0129] The VPS parameter 836 (dpb_max_temporal_id[i]) specifies the TemporalId of the top-level sublayer representation in the i-th DPB parameter syntax structure within the VPS where a DPB parameter may exist. The value of VPS parameter 836 can be between 0 and VPS parameter 812. If VPS parameter 812 is equal to 0, the value of VPS parameter 836 is inferred to be equal to 0. If VPS parameter 812 is greater than 0 and the corresponding flag (e.g., vps_all_layers_same_num_sublayers_flag) is equal to 1, the value of VPS parameter 836 is inferred to be equal to VPS parameter 812.

[0107]

[0130] VPS parameters 838 (ols_dpb_pic_width[i]) and 840 (ols_dpb_pic_height[i]) specify the width and height in luma samples of each picture storage buffer for the i-th OLS, respectively.

[0108]

[0131] The VPS parameter 842 (ols_dpb_params_idx[i]) specifies the index of the DPB parameter syntax structure applicable to the i-th OLS in the list of DPB parameter syntax structures within the VPS, when NumLayersInOls[i] is greater than 1. If any, the value of VPS parameter 842 can be between 0 and VPS parameter 832 minus 1. If VPS parameter 842 is not present, its value is inferred to be equal to 0. In some embodiments, when NumLayersInOls[i] is equal to 1, the DPB parameter syntax structure applicable to the i-th OLS is located in the SPS referenced by the layer in the i-th OLS.

[0109]

[0132] A VPS parameter 844 (vps_general_hrd_params_present_flag) equal to 1 indicates that the generic HRD parameter syntax structure (e.g., syntax structure 700A in Figure 7A) and other HRD parameters are located within the VPS RBSP syntax structure. A VPS parameter 844 equal to 0 indicates that the generic HRD parameter syntax structure and other HRD parameters are not located within the VPS RBSP syntax structure. If not, the value of VPS parameter 844 is inferred to be equal to 0. If NumLayersInOls[i] is equal to 1, the generic HRD parameter syntax structure applicable to the i-th OLS is located within the SPS referenced by the layer in the i-th OLS.

[0110]

[0133] A VPS parameter 846 (vps_sublayer_cpb_params_present_flag) equal to 1 specifies that the i-th OLS HRD parameter syntax structure in the VPS (e.g., syntax structure 700B in Figure 7B) contains HRD parameters for a sublayer representation whose TemporalId is between 0 and the value of VPS parameter 850. A VPS parameter 846 equal to 0 specifies that the i-th OLS HRD parameter syntax structure in the VPS contains HRD parameters for a sublayer representation whose TemporalId is equal to only the value of VPS parameter 850. If VPS parameter 812 is equal to 0, the value of VPS parameter 846 is inferred to be equal to 0.

[0111]

[0134] If VPS parameter 846 is equal to 0, then the HRD parameters for a sublayer representation with a TemporalId in the range of 0 or greater and less than or equal to VPS parameter 850 minus 1 are inferred to be the same as those for a sublayer representation with a TemporalId equal to VPS parameter 850. In some embodiments, under the conditional statement "if(general_vcl_hrd_params_present_flag)" in the OLS HRD parameter syntax structure 700B in Figure 7B, the HRD parameters include those starting from the fixed_pic_rate_general_flag[i] syntax element and extending to the sublayer_hrd_parameters(i) syntax structure.

[0112]

[0135] The VPS parameter 848 (num_ols_hrd_params_minus1) plus 1 specifies the number of OLS HRD parameter syntax structures within the generic HRD parameter syntax structure, when the VPS parameter 844 is equal to 1. The value of VPS parameter 848 can be between 0 and TotalNumOlss minus 1.

[0113]

[0136] The VPS parameter 850 (hrd_max_tid[i]) specifies the TemporalId of the top-level sublayer representation containing the HRD parameter within the i-th OLS HRD parameter syntax structure. The value of VPS parameter 850 can be between 0 and VPS parameter 812. If VPS parameter 812 is equal to 0, the value of VPS parameter 850 is inferred to be equal to 0. If VPS parameter 812 is greater than 0 and the corresponding flag (e.g., vps_all_layers_same_num_sublayers_flag) is equal to 1, the value of VPS parameter 850 is inferred to be equal to VPS parameter 812.

[0114]

[0137] The VPS parameter 852 (ols_hrd_idx[i]) specifies the index of the OLS HRD parameter syntax structure applicable to the i-th OLS in the list of OLS HRD parameter syntax structures within the VPS, when NumLayersInOls[i] is greater than 1. The value of VPS parameter 852 can be between 0 and VPS parameter 848. If NumLayersInOls[i] is equal to 1, the OLS HRD parameter syntax structure applicable to the i-th OLS is located in the SPS referenced by the layer in the i-th OLS. In some embodiments, if the value of VPS parameter 848 plus 1 is equal to TotalNumOls, the value of VPS parameter 852 is inferred to be equal to index i. In some embodiments, if NumLayersInOls[i] is greater than 1 and VPS parameter 848 is equal to 0, the value of VPS parameter 852 is inferred to be equal to 0.

[0115]

[0138] If the layers of the Sequence Parameter Set (SPS) are independent layers, the PTL parameter syntax structure, DPB parameter syntax structure, and HRD parameter syntax structure can be signaled within the SPS.

[0116]

[0139] Figure 9 shows an exemplary coded syntax table of some of the SPS RBSP syntax structures 900 signaled within the SPS, highlighted in bold, according to some embodiments of the present disclosure. As shown in Figure 9, the SPS parameter 910 (sps_seq_parameter_set_id) provides an identifier for the SPS for reference by other syntax elements. Regardless of the value of nuh_layer_id, SPS NAL units share the same value space for the SPS parameter 910. Let spsLayerId be the value of nuh_layer_id for a particular SPS NAL unit, and vclLayerId be the value of nuh_layer_id for a particular VCL NAL unit. A particular VCL NAL unit cannot reference a particular SPS NAL unit unless spsLayerId is less than or equal to vclLayerId and a layer having a nuh_layer_id equal to spsLayerId is contained within at least one OLS that contains a layer having a nuh_layer_id equal to vclLayerId.

[0117]

[0140] The SPS parameter 912 (sps_video_parameter_set_id) specifies the value of vps_video_parameter_set_id of the VPS referenced by the SPS when SPS parameter 912 is greater than 0. In some embodiments, if SPS parameter 912 is equal to 0, the corresponding SPS does not reference a VPS, and when decoding each CLVS that references an SPS, the VPS is not referenced. In addition, the value of the corresponding VPS parameter 810 is inferred to be equal to 0, the CVS contains only one layer (i.e., the VCL NAL units in the CVS have the same value of nuh_layer_id), the value of GeneralLayerIdx[nuh_layer_id] is inferred to be equal to 0, and the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1. If vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, then a specific nuh_layer_id value, an SPS referenced by CLVS along with nuhLayerId, can have a nuh_layer_id equal to nuhLayerId. In some embodiments, the value of the SPS parameter 912 may be the same within the SPS referenced by CLVS in the CVS.

[0118]

[0141] The SPS parameter 914 (sps_max_sublayers_minus1) plus 1 specifies the maximum number of time-axis sublayers that may exist within each CLVS referencing the SPS. The value of the SPS parameter 914 can be between 0 and the corresponding VPS parameter 812.

[0119]

[0142] In some embodiments, the SPS parameter 916 (sps_reserved_zero_4bits) may be equal to 0 in the bitstream. Other values ​​for the SPS parameter 916 can be reserved for future use.

[0120]

[0143] An SPS parameter 918 (sps_ptl_dpb_hrd_params_present_flag) equal to 1 specifies that the PTL syntax structure and DPB parameter syntax structure are present in the SPS, and that the general HRD parameter syntax structure and OLS HRD parameter syntax structure may also be present in the SPS. An SPS parameter 918 equal to 0 specifies that none of these four syntax structures are present in the SPS. The value of SPS parameter 918 can be equal to vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]].

[0121]

[0144] The SPS parameter 920 (sps_sublayer_dpb_params_flag) is used to control the presence of syntax elements within the DPB parameter syntax structure in SPS (e.g., max_dec_pic_buffering_minus1[i], max_num_reorder_pics[i], and max_latency_increase_plus1[i]). If they are not present, the value of SPS parameter 920 can be inferred to be equal to 0.

[0122]

[0145] An SPS parameter 922 (sps_general_hrd_params_present_flag) equal to 1 specifies that the SPS includes a general HRD parameter syntax structure and an OLS HRD parameter syntax structure. An SPS parameter 922 equal to 0 specifies that the SPS does not include a general HRD parameter syntax structure or an OLS HRD parameter syntax structure.

[0123]

[0146] An SPS parameter 924 (sps_sublayer_cpb_params_present_flag) equal to 1 specifies that the OLS HRD parameter syntax structure in the SPS includes HRD parameters for sublayer representations with TemporalIds in the range of 0 or greater and less than or equal to SPS parameter 914. An SPS parameter 924 equal to 0 specifies that the OLS HRD parameter syntax structure in the SPS includes HRD parameters for sublayer representations with TemporalIds equal only to SPS parameter 914. If SPS parameter 914 is equal to 0, the value of SPS parameter 924 is inferred to be equal to 0. If SPS parameter 924 is equal to 0, the HRD parameters for sublayer representations with TemporalIds in the range of 0 or greater and less than or equal to SPS parameter 914 minus 1 are inferred to be the same as those for sublayer representations with TemporalIds equal to SPS parameter 914. In some embodiments, under the conditional statement "if(general_vcl_hrd_params_present_flag)" in the OLS HRD parameter syntax structure 700B in Figure 7B, these HRD parameters include those starting from the fixed_pic_rate_general_flag[i] syntax element and extending to the sublayer_hrd_parameters(i) syntax structure.

[0124]

[0147] There are several issues regarding the signaling of PTL, DPB, and HRD parameters within the above-mentioned VPS or SPS, as well as the signaling of OLS.

[0125]

[0148] For example, as shown in Figure 8, in order to signal the DPB and HRD parameters, the number of syntax structures (VPS parameter 832 for the DPB parameter and VPS parameter 848 for the HRD parameter) and the index of the syntax structure applied to the i-th OLS (VPS parameter 842 for the DPB parameter and VPS parameter 852 for the HRD parameter) are signaled in the VPS using the "ue(v)" encoding method, which is a variable-length encoding method in which small values ​​are encoded with fewer bits and large values ​​are encoded with more bits. On the other hand, in order to signal the PTL parameters, the VPS parameter 822, which is the number of syntax structures, and the VPS parameter 830, which is the index of the syntax structure applied to the i-th OLS, are encoded by "u(8)", which is a fixed-length encoding method that uses 8 bits for values ​​in the range of 0 to 255.

[0126]

[0149] Therefore, the number of syntax structures and the encoding method for the syntax structure indices between DPB, HRD, and PTL parameters can differ. In addition, in many practical applications, the number of PTL syntax structures is relatively small. Using 8 bits for these syntax elements (e.g., VPS parameters 822 and 830) unnecessarily increases signaling overhead.

[0127]

[0150] Furthermore, an index (e.g., VPS parameters 830, 842, or 852) is signaled within the VPS for each OLS to specify which PTL, DPB, and HRD syntax structures apply to each OLS. However, if the number of OLSs is equal to the number of syntax structures, an efficient encoder can perform a one-to-one mapping from syntax structures to OLSs, avoiding the wasted bits resulting from unused syntax structures. Thus, the encoder can signal syntax structures in the order to which they apply to OLSs without signaling an index. On the decoder side, it can be inferred that the i-th syntax structure applies to the i-th OLS. Skipping the signaling of the index can reduce the number of bits and thus improve encoding efficiency. In some embodiments, this mechanism is used for HRD parameters. In some embodiments, this mechanism is also used for PTL and DPB parameters to prevent increased signaling overhead or inconsistencies in the signaling design.

[0128]

[0151] Furthermore, in SPS, SPS parameter 920, which is signaled when SPS parameter 914 is greater than 0, controls the presence of syntax elements including parameters 610B, 620B, and 630B within the DPB parameter syntax structure 600B in SPS. If SPS parameter 918 is equal to 0, the DPB parameter syntax structure is not signaled. In other words, if SPS parameter 914 is greater than 0 and SPS parameter 918 is equal to 0, the signaling of SPS parameter 920 is redundant.

[0129]

[0152] In addition, in the syntax structure 800 of Figure 8, the VPS parameter 818, used to specify the total number of OLS, is fixed-length encoded with a code length of 8. Therefore, the maximum value of the VPS parameter 818 is 255. For each layer in the CVS, an OLS is determined by the VPS parameter 820, and the total number of OLS is less than or equal to the number of combinations of the VPS parameter 820. If the maximum number of layers in the CVS (i.e., the value of the VPS parameter 810 plus 1) is less than 8, then using 8-bit fixed-length encoding for the VPS parameter 818 may be unnecessary.

[0130]

[0153] Figure 10A shows a flowchart of an exemplary video encoding method 1000A according to some embodiments of the present disclosure. In some embodiments, the video encoding method 1000A may be performed by an encoder (e.g., an encoder that performs process 200A in Figure 2A or process 200B in Figure 2B) to generate the bitstream 500 shown in Figure 5. For example, the encoder may be implemented as one or more software or hardware components of a device (e.g., device 400 in Figure 4) for encoding or transcoding a video sequence (e.g., video sequence 202 in Figure 2A or Figure 2B) to generate a bitstream for the video sequence (e.g., video bitstream 228 in Figure 2A or Figure 2B). For example, a processor (e.g., processor 402 in Figure 4) may perform the video encoding method 1000A.

[0131]

[0154] Referring to video encoding method 1000A, in step 1010a, the encoder encodes one or more PTL syntax elements (e.g., VPS parameter 822 or VPS parameter 830 in Figure 8) that specify PTL-related information having a variable length. In step 1020a, the encoder signals one or more PTL syntax elements having a variable length within the VPS (e.g., VPS510 in Figure 5) or SPS (e.g., SPS520 in Figure 5) of the bitstream (e.g., bitstream 500 in Figure 5).

[0132]

[0155] Figure 10B shows a flowchart of an exemplary video decoding method 1000B corresponding to the video encoding method 1000A of Figure 10A, according to some embodiments of the present disclosure. In some embodiments, the video decoding method 1000B may be performed by a decoder (e.g., a decoder that performs the decoding process 300A of Figure 3A or the decoding process 300B of Figure 3B) to decode the bitstream 500 of Figure 5. For example, the decoder may be implemented as one or more software or hardware components of a device (e.g., the device 400 of Figure 4) for decoding a bitstream (e.g., the video bitstream 228 of Figure 3A or Figure 3B) in order to reconstruct the video stream of the bitstream (e.g., the video stream 304 of Figure 3A or Figure 3B). For example, a processor (e.g., the processor 402 of Figure 4) may perform the video decoding method 1000B. Referring to Figure 10B, in the video decoding method 1000B, step 1010b is when the decoder receives a bitstream containing the VPS or SPS to be decoded (for example, bitstream 500 in Figure 5). In step 1020b, the decoder decodes one or more PTL syntax elements that specify variable-length PTL-related information within the VPS or SPS.

[0133]

[0156] Figures 10C and 10D show exemplary VPS syntax structures 1000C and 1000D, respectively, according to some embodiments of the present disclosure. Each of the VPS syntax structures 1000C and 1000D may be used within methods 1000A and 1000B. The VPS syntax structures 1000C and 1000D are modified based on the syntax structure 800 in Figure 8.

[0134]

[0157] As shown in Figure 10C, in some embodiments, PTL syntax elements encoded or decoded using a variable length may include a first PTL syntax element (e.g., VPS parameter 830) specifying an index of a PTL syntax structure, or a second PTL syntax element (e.g., VPS parameter 822) specifying the number of PTL syntax structures within a VPS or SPS. The first and second PTL syntax elements can be encoded or decoded using exponential-Golomb codes. For example, the number of PTL syntax structures (e.g., VPS parameter 822) and the index of a PTL syntax structure (e.g., VPS parameter 830) can be encoded using the "ue(v)" encoding method, which is a zero-order exponential-Golomb code. From a design consistency standpoint, signaling PTL parameters using the "ue(v)" encoding method allows PTL parameters to be signaled in the same way as DPB and HDR parameters. As shown in Figure 10C, descriptors 822d and 830d indicate that the encoding method for VPS parameters 822 and 830 is changed from u(8) to ue(v), respectively. The semantics shown in syntax structure 1000C in Figure 10C (left column of the table) remain unchanged and are therefore the same as those in syntax structure 800 in Figure 8.

[0135]

[0158] As shown in the syntax structure 1000D of Figure 10D, in some other embodiments, when encoding a first PTL syntax element (e.g., VPS parameter 830), the length of the first PTL syntax element can be set to the smallest integer of the number of PTL syntax structures in the VPS or SPS that is greater than or equal to the base 2 logarithm. In some other embodiments, when encoding a second PTL syntax element (e.g., VPS parameter 822), the length of the second PTL syntax element may be a fixed length such as u(8).

[0136]

[0159] For example, the index of a PTL syntax structure (e.g., VPS parameter 830) can be encoded using the "u(v)" encoding method, which does not have a fixed number of bits. The number of bits still depends on the values ​​of other syntax elements that are encoded with u(8), such as the value related to the number of PTL syntax structures (e.g., VPS parameter 822). Thus, after the value of VPS parameter 822 is syntax-parsed, the length of VPS parameter 830 is calculated as Ceil(log2(vps_num_ptls_minus1+1)), where log2(x) is the base-2 logarithm of x, and Ceil(x) is the smallest integer greater than or equal to x, and then VPS parameter 830 is syntax-parsed with Ceil(log2(vps_num_ptls_minus1+1)) bits. Specifically, for VPS parameter 830, the encoding method can be changed from u(8) to u(v). Similar to the embodiment shown in Figure 8, if present, the value of VPS parameter 830 can be in the range of 0 or greater and less than or equal to VPS parameter 822. If VPS parameter 828 is equal to 0, the value of VPS parameter 830 is inferred to be equal to 0.

[0137]

[0160] In some embodiments, if the number of PTL or DPB syntax structures is equal to the number of OLS, the index (e.g., VPS parameters 830 and 842) can be inferred without signaling. By omitting the signaling of the index, the signaling cost can be reduced.

[0138]

[0161] Figure 11A shows a flowchart of an exemplary video encoding method 1100A according to some embodiments of the present disclosure. Figure 11B shows a flowchart of an exemplary video decoding method 1100B corresponding to the video encoding method 1100A of Figure 11A, according to some embodiments of the present disclosure. Similar to method 1000A in Figure 10A and method 1000B in Figure 10B, the video encoding method 1100A and the video decoding method 1100B may be performed by an encoder and a decoder implemented as one or more software or hardware components of a device (e.g., a processor that executes a set of instructions).

[0139]

[0162] Referring to the video encoding method 1100A shown in Figure 11A, in step 1110a, the encoder determines whether the number of PTL syntax structures in the VPS (vps_num_ptls_minus1+1) and the number of OLS (TotalNumOlss) are the same. In other words, the encoder determines whether the encoded video sequence (CVS) contains an equal number of PTL syntax structures and OLS.

[0140]

[0163] In response that the CVS contains an equal number of PTL syntax structures and OLS (step 1110a - yes), the encoder bypasses steps 1120a and 1130a and encodes the bitstream without signaling a first PTL syntax element (e.g., ols_ptl_idx[i]) that specifies the index of the PTL syntax structure that fits the i-th OLS against the list of PTL syntax structures in the VPS.

[0141]

[0164] In response to the number of PTL syntax structures being different from the number of OLS (step 1110a - no), in step 1120a, the encoder determines whether the number of PTL syntax structures is equal to 1 (for example, by determining whether the parameter vps_num_ptls_minus1 is greater than zero). In response to the number of PTL syntax structures being 1 or less (step 1120a - no), the encoder bypasses step 1130a and encodes the bitstream without signaling the first PTL syntax element in the VPS.

[0142]

[0165] In response to the number of PTL syntax structures being greater than 1 and different from the number of OLS (step 1110a - no, step 1120a - yes), the encoder performs step 1130a and signals a first PTL syntax element within the VPS (or SPS). In some embodiments, the first PTL syntax element is signaled with a fixed length.

[0143]

[0166] In some other embodiments, step 1120a is performed before step 1110a. In response to the number of PTL syntax structures being greater than 1 and different from the number of OLS (step 1120a - yes, step 1110a - no), step 1130a is performed, otherwise step 1130a is skipped. For example, in response to the number of PTL syntax structures being equal to 1 (step 1120a - no), the encoder can bypass both steps 1110a and 1130a.

[0144]

[0167] Similar to the coding of PTL syntax elements, the coding of indices related to DPB syntax elements (e.g., VPS parameter 842) can be inferred without signaling if certain conditions are met.

[0145]

[0168] In step 1140a, the encoder determines whether the number of DPB parameter syntax structures in the VPS (vps_num_dpb_params) and the number of OLS (TotalNumOlss) are the same. In other words, the encoder determines whether the encoded video sequence (CVS) contains an equal number of DPB parameter syntax structures and OLS.

[0146]

[0169] In response to the fact that the number of DPB parameter syntax structures is equal to the number of OLS (step 1140a - yes), the encoder bypasses steps 1150a and 1160a and encodes the bitstream without signaling a first DPB syntax element (e.g., ols_dpb_params_idx[i]) that specifies the index of the DPB parameter syntax structure applicable to the i-th OLS in the list of DPB parameter syntax structures in the VPS.

[0147]

[0170] In response to the number of DPB parameter syntax structures being different from the number of OLS (step 1140a - no), in step 1150a, the encoder determines whether the number of DPB parameter syntax structures is equal to 1 (for example, by determining whether the parameter vps_num_dpb_params is greater than 1). In response to the number of DPB parameter syntax structures being 1 or less (step 1150a - no), the encoder bypasses step 1160a and encodes the bitstream without signaling the first DPB syntax element.

[0148]

[0171] In response to the fact that the number of DPB parameter syntax structures is greater than 1 and different from the number of OLS (step 1140a - no, step 1150a - yes), the encoder performs step 1160a and signals the first DPB syntax element within the VPS. In some embodiments, the first DPB syntax element is signaled with a variable length using ue(v).

[0149]

[0172] In some embodiments, step 1150a is performed before step 1140a. In response to the number of DPB syntax structures being greater than 1 and different from the number of OLS (step 1150a - yes, step 1140a - no), step 1160a is performed, otherwise step 1160a is skipped. For example, in response to the number of DPB syntax structures being equal to 1 (step 1150a - no), the encoder may bypass both steps 1140a and 1160a.

[0150]

[0173] Referring to the video decoding method 1100B shown in Figure 11B, in step 1110b, the decoder receives a bitstream containing the encoded video sequence (CVS) (for example, the video bitstream 500 in Figure 5). In step 1120b, the decoder determines whether the number of PTL syntax structures in the VPS (vps_num_ptls_minus1+1) is the same as the number of OLS (TotalNumOlss). In some embodiments, the decoder decodes one or more VPS syntax elements related to the OLS to obtain information on the number of OLS (TotalNumOlss).

[0151]

[0174] In response to the number of PTL syntax structures being equal to the number of OLS (step 1120b - yes), in step 1125b, the decoder infers that the first PTL syntax element (e.g., ols_ptl_idx[i]) specifying the index of the PTL syntax structure applicable to the i-th OLS against the list of PTL syntax structures in the VPS is equal to the sequence number of the i-th OLS (e.g., index i). In response to the number of PTL syntax structures being different from the number of OLS (step 1120b - no), in step 1130b, the decoder determines whether the number of PTL syntax structures is equal to 1 (e.g., by determining whether the parameter vps_num_ptls_minus1 is greater than 0). In response to the number of PTL syntax structures being 1 or less (step 1130b - no), in step 1135b, the decoder infers that the first PTL syntax element is zero.

[0152]

[0175] In response to the fact that the number of PTL syntax structures is greater than 1 and different from the number of OLS (step 1120b - no, step 1130b - yes), the decoder performs step 1140b and decodes the first PTL syntax elements encoded in the VPS or SPS. In some embodiments, the first PTL syntax elements are signaled with a fixed length in the VPS or SPS.

[0153]

[0176] In some embodiments, step 1130b is performed before step 1120b. In response to the number of PTL syntax structures being greater than 1 and different from the number of OLS (step 1130b - yes, step 1120b - no), step 1140a is performed, otherwise step 1140a is skipped.

[0154]

[0177] Similarly, in step 1150b, the decoder determines whether the number of DPB parameter syntax structures in the VPS (vps_num_dpb_params) is the same as the number of OLS (TotalNumOlss).

[0155]

[0178] In response to the fact that the number of DPB parameter syntax structures is the same as the number of OLS (step 1150b - yes), in step 1155b the decoder infers that the first DPB syntax element (e.g., ols_dpb_params_idx[i]) specifying the index of the DPB parameter syntax structure applicable to the i-th OLS against the list of DPB parameter syntax structures in the VPS is equal to the sequence number of the i-th OLS (e.g., index i).

[0156]

[0179] In response to the number of DPB parameter syntax structures being different from the number of OLS (step 1150b - no), in step 1160b, the decoder determines whether the number of DPB parameter syntax structures is equal to 1 (for example, by determining whether the parameter vps_num_dpb_params is greater than 1). In response to the number of DPB parameter syntax structures being 1 or less (step 1160b - no), in step 1165b, the decoder infers that the first DPB syntax element is zero.

[0157]

[0180] In response to the fact that the number of DPB parameter syntax structures is greater than 1 and different from the number of OLS (step 1150b - no, step 1160b - yes), the decoder performs step 1170b and decodes the first DPB syntax elements in the VPS. In some embodiments, the first DPB syntax elements are signaled with a variable length using ue(v).

[0158]

[0181] In some embodiments, step 1160b is performed before step 1150b. In response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS (step 1160b - yes, step 1150b - no), step 1170b is performed, otherwise step 1170b is skipped.

[0159]

[0182] Figure 11C shows a portion of exemplary VPS syntax structure 1100C related to possible implementations of the proposed methods 1100A and 1100B according to some embodiments of the present disclosure. The VPS syntax structure 1100C in Figure 11C is modified based on the syntax structure 800 in Figure 8.

[0160]

[0183] Compared to the embodiment shown in Figure 8, as shown in the conditional statement 1110c of the VPS syntax structure 1100C, VPS parameter 830 is signaled if the value of VPS parameter 822 (vps_num_ptls_minus1) plus 1 is not equal to the total number of OLS specified by the VPS (TotalNumOlss) and is not equal to 0. If the value of VPS parameter 822 plus 1 is equal to TotalNumOlss, the value of VPS parameter 830 is inferred to be equal to index i. If VPS parameter 822 is equal to 0, the value of VPS parameter 830 is inferred to be equal to 0.

[0161]

[0184] Similarly, as shown in the conditional statement 1120c of the VPS syntax structure 1100C, the VPS parameter 842 (ols_dpb_params_idx[i]) is signaled if the VPS parameter 832 is not equal to TotalNumOlss and is not equal to 0. When the VPS parameter 842 is not signaled, if the value of the VPS parameter 832 is equal to TotalNumOlss, the value of the VPS parameter 842 is inferred to be index i. If the VPS parameter 832 is equal to 0, the value of the VPS parameter 842 is inferred to be equal to 0.

[0162]

[0185] As described above, during the encoding or decoding process, the encoder or decoder may need to determine the number of OLS in the VPS (TotalNumOlss). In some embodiments, the decoder and encoder can derive the number of OLS according to the OLS mode indicated by a VPS syntax element (e.g., VPS parameter 816) and the maximum number of layers in the CVS indicated by another VPS syntax element (e.g., VPS parameter 810). In some embodiments, the encoder can encode a VPS syntax element (e.g., VPS parameter 818) having a fixed length related to the number of OLS specified by the VPS, and the decoder can decode the fixed-length VPS syntax element to determine the number of OLS specified by the VPS. In some embodiments, the encoder can encode a VPS syntax element (e.g., VPS parameter 818) having a variable length related to the number of OLS contained in the CVS referencing the VPS, and the decoder can decode the variable-length VPS syntax element to determine the number of OLS contained in the CVS referencing the VPS. If the maximum allowed number of layer combinations in CVS that reference a VPS is less than the default length value, the length of the VPS syntax element can be equal to the maximum allowed number.

[0163]

[0186] Figure 11D shows a portion of an exemplary VPS syntax structure 1100D, which is modified based on the syntax structure 800 of Figure 8 according to some embodiments of the present disclosure and relates to possible implementations for encoding or decoding VPS syntax elements having a variable length. In some cases, the total number of OLS is limited by the number of combinations of VPS parameters 820 (ols_output_layer_flag) for all layers in the CVS, and TotalNumOls satisfies the following inequality: TotalNumOlss≦2 vps_max_layers_minus1+1

[0164]

[0187] When VPS parameter 816 is equal to 2, VPS parameter 818 plus 1 specifies the total number of OLS specified by the VPS. Therefore, VPS parameter 818 can be represented by (vps_max_layers_minus1+1) bits (for example, the value of VPS parameter 810 plus 1).

[0165]

[0188] In some embodiments, as shown in Figure 11D, descriptor 818d, highlighted in italics, indicates that the encoding method for the VPS parameter 818 may be variable-length encoding, and the length of the VPS parameter 818 is the smaller of either a default value (e.g., 8) or the VPS parameter 810 plus 1 (i.e., the maximum allowed number of layers in CVS that reference the VPS). For example, in some embodiments, the length of the VPS parameter 818 can be determined based on the following min function. Min(8,vps_max_layers_minus1+1)

[0166]

[0189] Refer to Figure 9 again. As explained above, the value of the SPS parameter 914 can be in the range of 0 or greater and less than or equal to the corresponding VPS parameter 812. In some embodiments, if the SPS parameter 912 is equal to 0, the SPS does not refer to the VPS, and therefore there is no corresponding VPS parameter 812, and thus the range of the SPS parameter 914 is undefined. The syntax can be modified to solve this problem by assigning a range of values ​​for the independent SPS parameter 914 to the VPS parameter 812 when the SPS parameter 912 (sps_video_parameter_set_id) is equal to 0.

[0167]

[0190] Figure 12A shows a flowchart of an exemplary video encoding method 1200A according to some embodiments of the present disclosure. Figure 12B shows a flowchart of an exemplary video decoding method 1200B corresponding to the video encoding method 1200A of Figure 12A, according to some embodiments of the present disclosure. Similar to method 1000A in Figure 10A and method 1000B in Figure 10B, the video encoding method 1200A and the video decoding method 1200B may be performed by an encoder and a decoder implemented as one or more software or hardware components of a device (e.g., a processor that executes a set of instructions).

[0168]

[0191] Referring to the video encoding method 1200A shown in Figure 12A, in step 1210a, the encoder encodes a first SPS syntax element (e.g., SPS parameter 914 sps_max_sublayers_minus1) related to the maximum number of time-axis sublayers that may be in each CLVS referencing the SPS. In step 1220a, the encoder determines whether the SPS contains a PTL syntax structure, a DPB parameter syntax structure, and / or an HRD parameter syntax structure. For example, the encoder can make a determination and then set the value of SPS parameter 918 (sps_ptl_dpb_hrd_params_present_flag) based on that determination. If the determination is true (step 1220a - yes), the encoder performs step 1230a and determines whether the maximum number of time-axis sublayers in each encoded layer video sequence (CLVS) referencing the SPS is greater than 1 by determining whether SPS parameter 914 is greater than zero.

[0169]

[0192] If both conditions are met (step 1220a - yes, step 1230a - yes), in step 1240a, the encoder signals a flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag) configured to control the presence of syntax elements in the DPB parameter syntax structure within the SPS. The encoder then performs step 1250a to signal one or more syntax elements in the DPB parameter syntax structure.

[0170]

[0193] If the maximum number of sublayers in the time axis direction is 1 or less (step 1220a - yes, step 1230a - no), the encoder bypasses step 1240a and performs step 1250a to signal one or more syntax elements in the DPB parameter syntax structure by setting the flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag) to equal to 0, but not signaling it.

[0171]

[0194] If the SPS does not contain any PTL syntax structure, DPB parameter syntax structure, or HRD parameter syntax structure (step 1220a - no), the encoder bypasses steps 1230a to 1250a and encodes the SPS without signaling any flags (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag) and DPB syntax elements.

[0172]

[0195] In some embodiments, with respect to the encoding of the SPS parameter 914 in the video encoding method 1200A, when SPS refers to VPS, the range of SPS parameter 914 (sps_max_sublayers_minus1) can be set to be the same as that of VPS parameter 812 (vps_max_sublayers_minus1). If SPS does not refer to VPS, and therefore there is no corresponding VPS parameter 812 (vps_max_sublayers_minus1), the range of SPS parameter 914 can still be defined by setting a default value. For example, the value of VPS parameter 812 can be in the range of 0 to a default static value (e.g., 6). Thus, the value of SPS parameter 914 can be in the range of 0 or greater and MaxSubLayer minus 1 or less, and the value of parameter MaxSubLayer can be derived and calculated from the code using a ternary operator as follows. MaxSubLayer=(sps_video_parameter_set_id==0?6:vps_max_sublayers_minus1)+1

[0173]

[0196] In other words, the encoder can first determine the value of the SPS parameter 912, which, when its value is greater than zero, specifies the identifier of the VPS referenced by the SPS, and when its value is equal to zero, indicates that the SPS does not reference a VPS. Then, in response to the value of the SPS parameter 912 being greater than zero, the encoder assigns a range of the SPS parameter 914, based on the corresponding VPS parameter 812, that specifies the maximum number of time-axis sublayers within each CLVS that references the SPS. That is, the value of the SPS parameter 914 (e.g., sps_max_sublayers_minus1) is within the range of 0 to the VPS parameter 812 (e.g., vps_max_sublayers_minus1) (including the endpoints). On the other hand, in response to the value of the SPS parameter 912 being equal to zero, the encoder assigns a range of the SPS parameter 914 such that it is greater than or equal to zero and less than or equal to a fixed value (e.g., 6). In other words, the value of the SPS parameter 914 (e.g., sps_max_sublayers_minus1) is within the range of 0 to 6 (including the endpoints).

[0174]

[0197] Therefore, if SPS parameter 912 is equal to 0, an inferred value can be given to SPS parameter 914. Thus, the maximum value of SPS parameter 914 is available when SPS does not refer to any VPS. In some embodiments, the semantics can be modified so that if SPS parameter 912 is equal to 0, the value of VPS parameter 812 (vps_max_sublayers_minus1) is inferred to be equal to a default static value (e.g., 6).

[0175]

[0198] Referring to the video decoding method 1200B shown in Figure 12B, on the decoder side, in step 1210b, the decoder receives a bitstream containing the SPS to be decoded (e.g., video bitstream 500 in Figure 5). In step 1220b, the decoder decodes a first SPS syntax element (e.g., SPS parameter 914 sps_max_sublayers_minus1). In some embodiments, when decoding the SPS on the decoder side, the decoder may first determine the value of SPS parameter 912 and then decode SPS parameter 914. The range of SPS parameter 914 is based on the corresponding VPS parameter 812 of the VPS referenced by the SPS if the value of SPS parameter 912 is greater than zero, or on a fixed value if the value of SPS parameter 912 is equal to zero.

[0176]

[0199] In step 1230b, the decoder determines whether the SPS contains a PTL syntax structure, a DPB parameter syntax structure, and / or an HRD parameter syntax structure. For example, the decoder can make this determination based on whether the SPS parameter 918 (sps_ptl_dpb_hrd_params_present_flag) is equal to 1. If the determination is true (step 1230b - yes), the decoder performs step 1240b and determines whether the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) referencing the SPS is greater than 1 by determining whether the SPS parameter 914 is greater than zero.

[0177]

[0200] If both conditions are met (step 1230b - yes, step 1240b - yes), in step 1250b, the decoder decodes a flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag) that is signaled within the SPS and configured to control the presence of syntax elements in the DPB parameter syntax structure. The decoder then performs step 1260b to decode one or more syntax elements in the DPB parameter syntax structure based on the value of the flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag).

[0178]

[0201] If the maximum number of sublayers in the time axis direction is 1 or less (step 1230b - yes, step 1240b - no), the decoder bypasses step 1250b and infers that the flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag) is equal to zero, and then performs step 1260b to decode one or more syntax elements in the DPB parameter syntax structure based on the value of the flag (e.g., SPS parameter 920, sps_sublayer_dpb_params_flag).

[0179]

[0202] If the SPS does not contain any PTL syntax structure, DPB parameter syntax structure, or HRD parameter syntax structure (step 1230b - no), the decoder bypasses steps 1240b-1260b and decodes the SPS without decoding the flags and DPB syntax elements.

[0180]

[0203] Figure 12C shows a portion of an exemplary SPS syntax structure 1200C relating to possible implementations of the proposed methods 1200A and 1200B according to some embodiments of the present disclosure. The SPS syntax structure 1200C in Figure 12C may be modified based on the syntax structure 900 in Figure 9. As described above, in some embodiments, the signal of SPS parameter 920 can be based on both SPS parameter 918 and SPS parameter 914. Thus, as shown in Figure 12C, the syntax element SPS parameter 920 can be signaled (or decoded) under the condition that SPS parameter 918 is equal to 1 and SPS parameter 914 is greater than 0. Otherwise, SPS parameter 920 is not signaled (or decoded).

[0181]

[0204] In view of the above, by applying encoding or decoding of PTL syntax elements using variable length, as proposed in various embodiments of this disclosure, the encoding method for the number of syntax structures and the indices of syntax structures between DPB, HRD, and PTL parameters can be made consistent and efficient, reducing the signaling overhead that would result from using fixed length for those syntax elements. In addition, by appropriately inferring the value of a syntax element when it is not signaled, it is possible to skip the signaling of the index in some cases, which reduces the number of output bits and thus improves encoding efficiency. This method can be used not only for PTL and DPB parameters but also for HRD parameters to reduce signaling overhead and ensure consistency in signaling design.

[0182]

[0205] Embodiments can be further described using the following clauses: 1. A computer method for encoding video, To determine whether an encoded video sequence (CVS) contains an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS), and In response to CVS containing an equal number of PTL syntax structures and OLS, the bitstream is encoded without signaling a first PTL syntax element that specifies an index of the PTL syntax structure in the VPS that applies to the corresponding OLS in the VPS, within a list of PTL syntax structures in the VPS. A computer implementation method including 2. Determine whether the number of PTL syntax structures is equal to 1, and Encoding a bitstream without signaling the first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to 1. The method described in Clause 1, further including the method described in Clause 1. 3. In response to the number of PTL syntax structures being greater than 1 and different from the number of OLS, signal a first PTL syntax element with a fixed length within the VPS. The method described in Clause 1, further including the method described in Clause 1. 4. A computer method for encoding video, To determine whether the encoded video sequence (CVS) has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS), and In response to the CVS having an equal number of DPB parameter syntax structures and OLS, the bitstream is encoded without signaling a first DPB syntax element that specifies an index of the corresponding DPB parameter syntax structure in the OLS against a list of DPB parameter syntax structures in the VPS. A computer implementation method including 5. Determine whether the number of DPB parameter syntax structures is 1 or less, and Encoding the bitstream without signaling the first DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being equal to 1. The method described in Clause 4, further including the method described in Clause 4. 6. In response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS, signal a first DPB syntax element with variable length within the VPS. The method described in Clause 4, further including the method described in Clause 4. 7. A computer method for encoding video, Determining whether at least one of the Profile / Tier / Level (PTL) syntax structure, Decoded Picture Buffer (DPB) parameter syntax structure, or Virtual Reference Decoder (HRD) parameter syntax structure is present in the bitstream's Sequence Parameter Set (SPS). The first value is determined to be greater than 1, the first value specifying the maximum number of time-axis sublayers in the coded layer video sequence (CLVS) that references the SPS, and If at least one of the PTL syntax structure, DPB parameter syntax structure, or HRD parameter syntax structure is present in the SPS and the first value is greater than 1, then signal a flag configured to control the presence of syntax elements in the DPB parameter syntax structure within the SPS. A computer implementation method including 8. If the first value is 1 or less, signal one or more syntax elements in the DPB parameter syntax structure without signaling the flag within the SPS. The method described in Clause 7, further including the method described in Clause 7. 9. A computer method for encoding video, The value of a first sequence parameter set (SPS) syntax element is determined, wherein the first SPS syntax element specifies an identifier for the video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero. In response to the value of the first SPS syntax element being greater than zero, a range of the second SPS syntax element is assigned, based on the corresponding VPS syntax element, which specifies the maximum number of time-axis sublayers within each coded layer video sequence (CLVS) that references the SPS, and In response to the value of the first SPS syntax element being equal to zero, assign a range to the second SPS syntax element, which specifies the maximum number of time-axis sublayers within each CLVS that references an SPS, such that it is between zero and a fixed value. A computer implementation method including 10. A computer method for encoding video, Encoding one or more profile / tier / level (PTL) syntax elements that specify PTL-related information, and Signaling one or more variable-length PTL syntax elements within a bitstream's video parameter set (VPS) or sequence parameter set (SPS). A computer implementation method including 11. Encoding one or more PTL syntax elements is: Encoding a first PTL syntax element having a variable length, wherein the first PTL syntax element specifies an index of the PTL syntax structure. The method described in Clause 10, including the method described in Clause 10. 12. A VPS or SPS contains N PTL syntax structures, where N is an integer and one or more PTL syntax elements are encoded. Set the length of the first PTL syntax element such that it is the smallest integer greater than or equal to the base 2 logarithm of N. The method described in Clause 11, further including the method described in Clause 11. 13. Encoding one or more PTL syntax elements is: Encoding the first PTL syntax element by using exponential-Golomb coding. The method described in Article 11, including the method described in Article 11. 14. A VPS or SPS contains N PTL syntax structures, where N is an integer and one or more PTL syntax elements are encoded. Encoding a second PTL syntax element having a variable length, wherein the second PTL syntax element specifies N. The method described in Clause 10, including the method described in Clause 10. 15. A computer method for encoding video, Encoding variable-length video parameter set (VPS) syntax elements, and The VPS syntax element is signaled within the VPS, and the VPS syntax element is related to the number of output layer sets (OLS) contained within the encoded video sequence (CVS) that references the VPS. A computer implementation method including 16. If the maximum number of layers in an encoded video sequence (CVS) that references a VPS is less than the default length value, set the length of the VPS syntax element to be equal to the maximum number of layers. The method described in Clause 15, further including the method described in Clause 15. 17. A computer method for decoding video, Receiving a bitstream containing an encoded video sequence (CVS), To determine whether the CVS has an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS), and In response to the number of PTL syntax structures being equal to the number of OLS, when decrypting a VPS, the decoding of the first PTL syntax element that specifies the index of the PTL syntax structure that applies to the corresponding OLS against the list of PTL syntax structures in the VPS is skipped. A computer implementation method including 18. In response to the number of PTL syntax structures being equal to the number of OLS, determine the first PTL syntax element such that the PTL syntax structure specified by the first PTL syntax element is equal to the index of the OLS to which it applies. The method described in Article 17, further including the method described in Article 17. 19. Determine whether the number of PTL syntax structures is equal to 1, and When decrypting the VPS, skip decrypting the first PTL syntax element in response to the number of PTL syntax structures being equal to 1. The method described in Article 17, further including the method described in Article 17. 20. Determine the first PTL syntax element to be zero in response to the number of PTL syntax structures being equal to 1. The method described in Article 19, further including the method described in Article 19. 21. Decode a first fixed-length PTL syntax element in the VPS in response to the number of PTL syntax structures being greater than 1 and different from the number of OLS. The method described in Article 17, further including the method described in Article 17. 22. A computer method for decoding video, Receiving a bitstream containing an encoded video sequence (CVS), Determining whether the CVS has an equal number of Decoded Picture Buffer (DPB) parameter syntax structures and Output Layer Sets (OLS), and In response to the CVS having an equal number of DPB parameter syntax structures and OLS, the decoding of the first DPB syntax element that specifies the index of the DPB parameter syntax structure that fits the corresponding OLS against the list of DPB parameter syntax structures in the VPS is skipped. A computer implementation method including 23. Determine whether the number of DPB parameter syntax structures is 1 or less, and When decoding the VPS, skip decoding the first DPB syntax element in response to the number of DPB parameter syntax structures being 1 or less. The method described in Clause 22, further including the method described in Clause 22. 24. In response to the number of DPB parameter syntax structures being 1 or less, determine that the first DPB syntax element is zero, and In response to the fact that the number of DPB syntax structures is equal to the number of OLS, the first DPB syntax element is determined such that the DPB parameter syntax structure specified by the first DPB syntax element is equal to the index of the OLS to which it applies. The method described in Clause 23, further including the method described in Clause 23. 25. Decoding a first variable-length DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS. The method described in Clause 23, further including the method described in Clause 23. 26. A computer method for decoding video, Receiving a bitstream containing a video parameter set (VPS) and a sequence parameter set (SPS), In response to the presence of at least one of the Profile / Tier / Level (PTL) syntax structure, Decoded Picture Buffer (DPB) parameter syntax structure, or Virtual Reference Decoder (HRD) parameter syntax structure within the SPS, it is determined whether a first value specifying the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) referencing the SPS is greater than 1, and Decode a flag configured to control the presence of syntax elements in the DPB parameter syntax structure within the SPS in response to the first value being greater than 1. A computer implementation method including 27. Inferring that the flag is equal to zero in response to the first value being 1 or less, and decoding one or more syntax elements in the DPB parameter syntax structure without decoding the flag in the SPS. The method described in Article 26, further including the method described in Article 26. 28. A computer method for decoding video, Determining the value of a first sequence parameter set (SPS) syntax element, wherein the first SPS syntax element specifies an identifier for the video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero, and Decoding a second SPS syntax element that specifies the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) that references an SPS, wherein the range of the second SPS syntax element is based on the corresponding VPS syntax element of the VPS referenced by the SPS if the value of the first SPS syntax element is greater than zero, or based on a fixed value if the value of the first SPS syntax element is equal to zero. A computer implementation method including 29. A computer method for decoding video, Receiving a bitstream containing a video parameter set (VPS) or sequence parameter set (SPS), and Decrypting one or more profile / tier / level (PTL) syntax elements within a VPS or SPS, wherein one or more PTL syntax elements specify PTL-related information. A computer implementation method including 30. Decoding one or more PTL syntax elements is Decoding a first PTL syntax element having a variable length, wherein the first PTL syntax element specifies an index of the PTL syntax structure. The method described in Article 29, including the method described in Article 29. 31. The method according to clause 30, wherein the VPS or SPS comprises N PTL syntax structures, where N is an integer, and the length of the first PTL syntax element is the smallest integer greater than or equal to the base 2 logarithm of N. 32. The first PTL syntax element is encoded by using the exponential-Golomb code, as described in Clause 30. 33. A VPS or SPS contains N PTL syntax structures, where N is an integer, and decoding one or more PTL syntax elements is possible. Decoding a second PTL syntax element having a variable length, wherein the second PTL syntax element specifies N. The method described in Article 29, including the method described in Article 29. 34. A computer method for decoding video, Receiving a bitstream containing a video parameter set (VPS), and Decoding a variable-length VPS syntax element within a VPS, wherein the VPS syntax element relates to the number of output layer sets (OLS) contained within the encoded video sequence (CVS) referencing the VPS. A computer implementation method including 35. If the maximum allowed number of layers in an encoded video sequence (CVS) that references a VPS is less than the default length value, the length of the VPS syntax element shall be equal to the maximum allowed number, as described in Clause 34. 36. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is To determine whether an encoded video sequence (CVS) contains an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS), and In response to CVS containing an equal number of PTL syntax structures and OLS, the bitstream is encoded without signaling a first PTL syntax element that specifies an index of the PTL syntax structure in the VPS that applies to the corresponding OLS in the VPS, within a list of PTL syntax structures in the VPS. A device configured to execute commands in order to cause another device to perform a certain action. 37. The processor is, Determine whether the number of PTL syntax structures is equal to 1, and Encoding a bitstream without signaling the first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to 1. The equipment described in Clause 36, configured to carry out instructions in order to do so. 38. The processor is, In response to the number of PTL syntax structures being greater than 1 and different from the number of OLS, a first PTL syntax element with a fixed length is signaled within the VPS. The equipment described in Clause 36, configured to carry out instructions in order to do so. 39. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is To determine whether the encoded video sequence (CVS) has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLS), and In response to the CVS having an equal number of DPB parameter syntax structures and OLS, the bitstream is encoded without signaling a first DPB syntax element that specifies an index of the corresponding DPB parameter syntax structure in the OLS against a list of DPB parameter syntax structures in the VPS. A device configured to execute commands in order to cause another device to perform a certain action. 40. The processor is, Determine whether the number of DPB parameter syntax structures is 1 or less, and Encoding the bitstream without signaling the first DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being equal to 1. The equipment described in Clause 39, configured to carry out instructions in order to do so. 41. The processor is, In response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS, a first variable-length DPB syntax element is signaled within the VPS. The equipment described in Clause 39, configured to carry out instructions in order to do so. 42. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is Determining whether at least one of the Profile / Tier / Level (PTL) syntax structure, Decoded Picture Buffer (DPB) parameter syntax structure, or Virtual Reference Decoder (HRD) parameter syntax structure is present in the bitstream's Sequence Parameter Set (SPS). The first value is determined to be greater than 1, the first value specifying the maximum number of time-axis sublayers in the coded layer video sequence (CLVS) that references the SPS, and If at least one of the PTL syntax structure, DPB parameter syntax structure, or HRD parameter syntax structure is present in the SPS and the first value is greater than 1, then signal a flag configured to control the presence of syntax elements in the DPB parameter syntax structure within the SPS. A device configured to execute commands in order to cause another device to perform a certain action. 43. The processor is, If the first value is 1 or less, signal one or more syntax elements in the DPB parameter syntax structure without signaling the flag within the SPS. The equipment described in Clause 42, configured to carry out instructions in order to do so. 44. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is The value of a first sequence parameter set (SPS) syntax element is determined, wherein the first SPS syntax element specifies an identifier for the video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero. In response to the value of the first SPS syntax element being greater than zero, a range of the second SPS syntax element is assigned, based on the corresponding VPS syntax element, which specifies the maximum number of time-axis sublayers within each coded layer video sequence (CLVS) that references the SPS, and In response to the value of the first SPS syntax element being equal to zero, assign a range to the second SPS syntax element, which specifies the maximum number of time-axis sublayers within each CLVS that references an SPS, such that it is between zero and a fixed value. A device configured to execute commands in order to cause another device to perform a certain action. 45. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is Encoding one or more profile / tier / level (PTL) syntax elements that specify PTL-related information, and Signaling one or more variable-length PTL syntax elements within a bitstream's video parameter set (VPS) or sequence parameter set (SPS). A device configured to execute commands in order to cause another device to perform a certain action. 46. ​​The processor is, Encoding a first PTL syntax element having a variable length, wherein the first PTL syntax element specifies an index of the PTL syntax structure. The device described in Clause 45, configured to execute instructions to encode one or more PTL syntax elements. 47. A VPS or SPS contains N PTL syntax structures, where N is an integer and the processor is... Set the length of the first PTL syntax element such that it is the smallest integer greater than or equal to the base 2 logarithm of N. The device described in Clause 46, configured to execute instructions to encode one or more PTL syntax elements. 48. The processor is, Encoding the first PTL syntax element by using exponential-Golomb coding. The device described in Clause 46, configured to execute instructions to encode one or more PTL syntax elements. 49. A VPS or SPS contains N PTL syntax structures, where N is an integer and the processor is... Encoding a second PTL syntax element having a variable length, wherein the second PTL syntax element specifies N. The device described in Clause 45, configured to execute instructions to encode one or more PTL syntax elements. 50. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is Encoding variable-length video parameter set (VPS) syntax elements, and The VPS syntax element is signaled within the VPS, and the VPS syntax element is related to the number of output layer sets (OLS) contained within the encoded video sequence (CVS) that references the VPS. A device configured to execute commands in order to cause another device to perform a certain action. 51. The processor is, If the maximum number of layers in an Encoded Video Sequence (CVS) that references a VPS is less than the default length value, set the length of the VPS syntax element to be equal to the maximum number of layers. The equipment described in Clause 50, configured to carry out instructions in order to do so. 52. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is Receiving a bitstream containing an encoded video sequence (CVS), To determine whether the CVS has an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS), and In response to the number of PTL syntax structures being equal to the number of OLS, when decrypting a VPS, the decoding of the first PTL syntax element that specifies the index of the PTL syntax structure that applies to the corresponding OLS against the list of PTL syntax structures in the VPS is skipped. A device configured to execute commands in order to cause another device to perform a certain action. 53. The processor is, In response to the fact that the number of PTL syntax structures is equal to the number of OLS, the first PTL syntax element is determined such that the PTL syntax structure specified by the first PTL syntax element is equal to the index of the OLS to which it applies. The equipment described in Clause 52, configured to carry out instructions in order to do so. 54. The processor is, Determine whether the number of PTL syntax structures is equal to 1, and When decrypting the VPS, skip decrypting the first PTL syntax element in response to the number of PTL syntax structures being equal to 1. The equipment described in Clause 52, configured to carry out instructions in order to do so. 55. The processor is, Determine the first PTL syntax element to be zero in response to the number of PTL syntax structures being equal to 1. The equipment described in Clause 54, configured to carry out instructions in order to do so. 56. The processor is, Decoding a first fixed-length PTL syntax element within a VPS in response to the number of PTL syntax structures being greater than 1 and different from the number of OLS. The equipment described in Clause 52, configured to carry out instructions in order to do so. 57. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is Receiving a bitstream containing an encoded video sequence (CVS), Determining whether the CVS has an equal number of Decoded Picture Buffer (DPB) parameter syntax structures and Output Layer Sets (OLS), and In response to the CVS having an equal number of DPB parameter syntax structures and OLS, the decoding of the first DPB syntax element that specifies the index of the DPB parameter syntax structure that fits the corresponding OLS against the list of DPB parameter syntax structures in the VPS is skipped. A device configured to execute commands in order to cause another device to perform a certain action. 58. The processor is, Determine whether the number of DPB parameter syntax structures is 1 or less, and When decoding the VPS, skip decoding the first DPB syntax element in response to the number of DPB parameter syntax structures being 1 or less. The equipment described in Clause 57, configured to carry out instructions in order to do so. 59. The processor is, In response to the number of DPB parameter syntax structures being 1 or less, the first DPB syntax element is determined to be zero, and In response to the fact that the number of DPB syntax structures is equal to the number of OLS, the first DPB syntax element is determined such that the DPB parameter syntax structure specified by the first DPB syntax element is equal to the index of the OLS to which it applies. The equipment described in Clause 58, configured to carry out instructions in order to do so. 60. The processor is, Decoding a first variable-length DPB syntax element within a VPS in response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS. The equipment described in Clause 58, configured to carry out instructions in order to do so. 61. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is Receiving a bitstream containing a video parameter set (VPS) and a sequence parameter set (SPS), In response to the presence of at least one of the Profile / Tier / Level (PTL) syntax structure, Decoded Picture Buffer (DPB) parameter syntax structure, or Virtual Reference Decoder (HRD) parameter syntax structure within the SPS, it is determined whether a first value specifying the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) referencing the SPS is greater than 1, and Decode a flag configured to control the presence of syntax elements in the DPB parameter syntax structure within the SPS in response to the first value being greater than 1. A device configured to execute commands in order to cause another device to perform a certain action. 62. The processor is, In response to the first value being 1 or less, infer that the flag is equal to zero, and decode one or more syntax elements in the DPB parameter syntax structure without decoding the flag in the SPS. The equipment described in Clause 61, configured to carry out instructions in order to do so. 63. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is, Determining the value of a first sequence parameter set (SPS) syntax element, wherein the first SPS syntax element specifies an identifier for the video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero, and Decoding a second SPS syntax element that specifies the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) that references an SPS, wherein the range of the second SPS syntax element is based on the corresponding VPS syntax element of the VPS referenced by the SPS if the value of the first SPS syntax element is greater than zero, or based on a fixed value if the value of the first SPS syntax element is equal to zero. A device configured to execute commands in order to cause another device to perform a certain action. 64. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is Receiving a bitstream containing a video parameter set (VPS) or sequence parameter set (SPS), and Decrypting one or more profile / tier / level (PTL) syntax elements within a VPS or SPS, wherein one or more PTL syntax elements specify PTL-related information. A device configured to execute commands in order to cause another device to perform a certain action. 65. The processor is, Decoding a first PTL syntax element having a variable length, wherein the first PTL syntax element specifies an index of the PTL syntax structure. The device described in Clause 64, configured to execute instructions to decode one or more PTL syntax elements. 66. The equipment described in Clause 65, wherein the VPS or SPS comprises N PTL syntax structures, where N is an integer, and the length of the first PTL syntax element is the smallest integer greater than or equal to the base 2 logarithm of N. 67. The first PTL syntax element is encoded by using the exponential-Golomb code, as described in Clause 65. 68. A VPS or SPS contains N PTL syntax structures, where N is an integer and the processor is... Decoding a second PTL syntax element having a variable length, wherein the second PTL syntax element specifies N. The device described in Clause 64, configured to execute instructions to decode one or more PTL syntax elements. 69. Memory configured to store instructions, It includes a processor coupled to memory, and the processor is Receiving a bitstream containing a video parameter set (VPS), and Decoding a variable-length VPS syntax element within a VPS, wherein the VPS syntax element relates to the number of output layer sets (OLS) contained within the encoded video sequence (CVS) referencing the VPS. A device configured to execute commands in order to cause another device to perform a certain action. 70. If the maximum allowed number of layers in an Encoded Video Sequence (CVS) that references a VPS is less than the default length value, the length of the VPS syntax element shall be equal to the maximum allowed number, as described in Clause 69. 71. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for encoding an image, and the method is To determine whether an encoded video sequence (CVS) contains an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS), and In response to CVS containing an equal number of PTL syntax structures and OLS, the bitstream is encoded without signaling a first PTL syntax element that specifies an index of the PTL syntax structure in the VPS that applies to the corresponding OLS in the VPS, within a list of PTL syntax structures in the VPS. Non-temporary computer-readable storage media, including [specific type of storage medium]. 72. The set of instructions that can be executed by one or more processors of a device is: Determine whether the number of PTL syntax structures is equal to 1, and Encoding a bitstream without signaling the first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to 1. A non-temporary computer-readable storage medium as described in Clause 71, which causes the device to perform the following actions. 73. The set of instructions that can be executed by one or more processors of a device is: In response to the number of PTL syntax structures being greater than 1 and different from the number of OLSs, signaling a first PTL syntax element having a fixed length within the VPS The non - transitory computer - readable storage medium according to clause 71, which further causes the device to perform the above 74. A non - transitory computer - readable storage medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to perform a method for encoding video, the method comprising determining whether an encoded video sequence (CVS) has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLSs), and encoding a bitstream without signaling a first DPB syntax element that specifies an index of a DPB parameter syntax structure within the VPS that corresponds to an OLS, in response to the CVS having an equal number of DPB parameter syntax structures and OLSs The non - transitory computer - readable storage medium comprising the above 75. A set of instructions executable by one or more processors of a device, the set of instructions comprising determining whether the number of DPB parameter syntax structures is less than or equal to 1, and encoding a bitstream without signaling a first DPB syntax element within the VPS, in response to the number of DPB parameter syntax structures being equal to 1 The non - transitory computer - readable storage medium according to clause 74, which further causes the device to perform the above 76. A set of instructions executable by one or more processors of a device, the set of instructions comprising in response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLSs, signaling a first DPB syntax element having a variable length within the VPS The non - transitory computer - readable storage medium according to clause 74, which further causes the device to perform the above 77. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for encoding an image, Determining whether at least one of the Profile / Tier / Level (PTL) syntax structure, Decoded Picture Buffer (DPB) parameter syntax structure, or Virtual Reference Decoder (HRD) parameter syntax structure is present in the bitstream's Sequence Parameter Set (SPS). The first value is determined to be greater than 1, the first value specifying the maximum number of time-axis sublayers in the coded layer video sequence (CLVS) that references the SPS, and If at least one of the PTL syntax structure, DPB parameter syntax structure, or HRD parameter syntax structure is present in the SPS and the first value is greater than 1, then signal a flag configured to control the presence of syntax elements in the DPB parameter syntax structure within the SPS. Non-temporary computer-readable storage media, including [specific type of storage medium]. 78. The set of instructions that can be executed by one or more processors of a device is: If the first value is 1 or less, signal one or more syntax elements in the DPB parameter syntax structure without signaling the flag within the SPS. A non-temporary computer-readable storage medium as described in Clause 77, which allows the device to perform the following actions. 79. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for encoding an image, The value of a first sequence parameter set (SPS) syntax element is determined, wherein the first SPS syntax element specifies an identifier for the video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero. In response to the value of the first SPS syntax element being greater than zero, a range of the second SPS syntax element is assigned, based on the corresponding VPS syntax element, which specifies the maximum number of time-axis sublayers within each coded layer video sequence (CLVS) that references the SPS, and In response to the value of the first SPS syntax element being equal to zero, assign a range to the second SPS syntax element, which specifies the maximum number of time-axis sublayers within each CLVS that references an SPS, such that it is between zero and a fixed value. Non-temporary computer-readable storage media, including [specific type of storage medium]. 80. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for encoding an image, Encoding one or more profile / tier / level (PTL) syntax elements that specify PTL-related information, and Signaling one or more variable-length PTL syntax elements within a bitstream's video parameter set (VPS) or sequence parameter set (SPS). Non-temporary computer-readable storage media, including [specific type of storage medium]. 81. The set of instructions that can be executed by one or more processors of a device is: Encoding a first PTL syntax element having a variable length, wherein the first PTL syntax element specifies an index of the PTL syntax structure. A non-temporary computer-readable storage medium as described in Clause 80, which causes the device to encode one or more PTL syntax elements. 82. A VPS or SPS contains N PTL syntax structures, where N is an integer and the set of instructions executable by one or more processors in the device is: Set the length of the first PTL syntax element such that it is the smallest integer greater than or equal to the base 2 logarithm of N. A non-temporary computer-readable storage medium as described in Clause 81, which causes the device to encode one or more PTL syntax elements. 83. The set of instructions that can be executed by one or more processors of a device is: Encoding the first PTL syntax element by using exponential-Golomb coding. A non-temporary computer-readable storage medium as described in Clause 81, which causes the device to encode one or more PTL syntax elements. 84. A VPS or SPS contains N PTL syntax structures, where N is an integer and the set of instructions executable by one or more processors in the device is: Encoding a second PTL syntax element having a variable length, wherein the second PTL syntax element specifies N. A non-temporary computer-readable storage medium as described in Clause 80, which causes the device to encode one or more PTL syntax elements. 85. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for encoding an image, Encoding variable-length video parameter set (VPS) syntax elements, and The VPS syntax element is signaled within the VPS, and the VPS syntax element is related to the number of output layer sets (OLS) contained within the encoded video sequence (CVS) that references the VPS. Non-temporary computer-readable storage media, including [specific type of storage medium]. 86. The set of instructions that can be executed by one or more processors of a device is: If the maximum number of layers in an Encoded Video Sequence (CVS) that references a VPS is less than the default length value, set the length of the VPS syntax element to be equal to the maximum number of layers. A non-temporary computer-readable storage medium as described in Clause 85, which causes the device to perform the following actions. 87. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for decoding an image, Receiving a bitstream containing an encoded video sequence (CVS), To determine whether the CVS has an equal number of profile / tier / level (PTL) syntax structures and output layer sets (OLS), and In response to the number of PTL syntax structures being equal to the number of OLS, when decrypting a VPS, the decoding of the first PTL syntax element that specifies the index of the PTL syntax structure that applies to the corresponding OLS against the list of PTL syntax structures in the VPS is skipped. Non-temporary computer-readable storage media, including [specific type of storage medium]. 88. The set of instructions that can be executed by one or more processors of a device is: In response to the fact that the number of PTL syntax structures is equal to the number of OLS, the first PTL syntax element is determined such that the PTL syntax structure specified by the first PTL syntax element is equal to the index of the OLS to which it applies. A non-temporary computer-readable storage medium as described in Clause 87, which allows the device to perform the following actions. 89. The set of instructions that can be executed by one or more processors of a device is: Determine whether the number of PTL syntax structures is equal to 1, and When decrypting the VPS, skip decrypting the first PTL syntax element in response to the number of PTL syntax structures being equal to 1. A non-temporary computer-readable storage medium as described in Clause 87, which allows the device to perform the following actions. 90. The set of instructions that can be executed by one or more processors of a device is: Determine the first PTL syntax element to be zero in response to the number of PTL syntax structures being equal to 1. The non - transitory computer - readable storage medium according to clause 89, which further causes the apparatus to perform 91. The set of instructions executable by one or more processors of the apparatus In response to the number of PTL syntax structures being greater than 1 and different from the number of OLSs, decoding a first PTL syntax element having a fixed length within the VPS The non - transitory computer - readable storage medium according to clause 87, which further causes the apparatus to perform 92. A non - transitory computer - readable storage medium storing a set of instructions, the set of instructions being executable by one or more processors of the apparatus to cause the apparatus to perform a method for decoding video, the method Receiving a bitstream including an encoded video sequence (CVS), Determining whether the CVS has an equal number of decoded picture buffer (DPB) parameter syntax structures and output layer sets (OLSs), and In response to the CVS having an equal number of DPB parameter syntax structures and OLSs, skipping the decoding of a first DPB syntax element that specifies an index of the DPB parameter syntax structure corresponding to the OLS with respect to the list of DPB parameter syntax structures within the VPS A non - transitory computer - readable storage medium including 93. The set of instructions executable by one or more processors of the apparatus Determining whether the number of DPB parameter syntax structures is less than or equal to 1, and In response to the number of DPB parameter syntax structures being less than or equal to 1, skipping the decoding of the first DPB syntax element when decoding the VPS The non - transitory computer - readable storage medium according to clause 92, which further causes the apparatus to perform 94. The set of instructions executable by one or more processors of the apparatus [[ID=A]] The set of instructions executable by one or more processors of the apparatus In response to the number of DPB parameter syntax structures being 1 or less, the first DPB syntax element is determined to be zero, and In response to the fact that the number of DPB syntax structures is equal to the number of OLS, the first DPB syntax element is determined such that the DPB parameter syntax structure specified by the first DPB syntax element is equal to the index of the OLS to which it applies. A non-temporary computer-readable storage medium as described in Clause 93, which causes the device to perform the following actions. 95. The set of instructions that can be executed by one or more processors of a device is: Decoding a first variable-length DPB syntax element within a VPS in response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS. A non-temporary computer-readable storage medium as described in Clause 93, which causes the device to perform the following actions. 96. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for decoding an image, Receiving a bitstream containing a video parameter set (VPS) and a sequence parameter set (SPS), In response to the presence of at least one of the Profile / Tier / Level (PTL) syntax structure, Decoded Picture Buffer (DPB) parameter syntax structure, or Virtual Reference Decoder (HRD) parameter syntax structure within the SPS, it is determined whether a first value specifying the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) referencing the SPS is greater than 1, and Decode a flag configured to control the presence of syntax elements in the DPB parameter syntax structure within the SPS in response to the first value being greater than 1. Non-temporary computer-readable storage media, including [specific type of storage medium]. 97. The set of instructions that can be executed by one or more processors of a device is: In response to the first value being 1 or less, infer that the flag is equal to zero, and decode one or more syntax elements in the DPB parameter syntax structure without decoding the flag in the SPS. A non-temporary computer-readable storage medium as described in Clause 96, which causes the device to perform the following actions. 98. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for decoding an image, Determining the value of a first sequence parameter set (SPS) syntax element, wherein the first SPS syntax element specifies an identifier for the video parameter set (VPS) referenced by the SPS when the value of the first SPS syntax element is greater than zero, and Decoding a second SPS syntax element that specifies the maximum number of time-axis sublayers in each coded layer video sequence (CLVS) that references an SPS, wherein the range of the second SPS syntax element is based on the corresponding VPS syntax element of the VPS referenced by the SPS if the value of the first SPS syntax element is greater than zero, or based on a fixed value if the value of the first SPS syntax element is equal to zero. Non-temporary computer-readable storage media, including [specific type of storage medium]. 99. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for decoding an image, Receiving a bitstream containing a video parameter set (VPS) or sequence parameter set (SPS), and Decrypting one or more profile / tier / level (PTL) syntax elements within a VPS or SPS, wherein one or more PTL syntax elements specify PTL-related information. Non-temporary computer-readable storage media, including [specific type of storage medium]. 100. The set of instructions that can be executed by one or more processors of a device is: Decoding a first PTL syntax element having a variable length, wherein the first PTL syntax element specifies an index of the PTL syntax structure. A non-temporary computer-readable storage medium as described in Clause 99, which causes a device to decode one or more PTL syntax elements. 101. A non-temporary computer-readable storage medium as described in Clause 100, wherein the VPS or SPS comprises N PTL syntax structures, where N is an integer, and the length of the first PTL syntax element is the smallest integer greater than or equal to the base 2 logarithm of N. 102. The first PTL syntax element is encoded by using exponential-Golomb code in the non-temporary computer-readable storage medium as described in Clause 100. 103. A VPS or SPS contains N PTL syntax structures, where N is an integer and the set of instructions executable by one or more processors in the device is: Decoding a second PTL syntax element having a variable length, wherein the second PTL syntax element specifies N. A non-temporary computer-readable storage medium as described in Clause 99, which causes a device to decode one or more PTL syntax elements. 104. A non-temporary computer-readable storage medium for storing a set of instructions, wherein the set of instructions is executable by one or more processors of the device to cause the device to perform a method for decoding an image, Receiving a bitstream containing a video parameter set (VPS), and Decoding a variable-length VPS syntax element within a VPS, wherein the VPS syntax element relates to the number of output layer sets (OLS) contained within the encoded video sequence (CVS) referencing the VPS. Non-temporary computer-readable storage media, including [specific type of storage medium]. 105. If the maximum allowed number of layers in an encoded video sequence (CVS) that references a VPS is less than the default length value, the length of the VPS syntax elements shall be equal to the maximum allowed number, as described in Clause 104 for non-temporary computer-readable storage media.

[0183]

[0206] In some embodiments, non-temporary computer-readable storage media containing instructions are also provided, which can be executed by a device (such as an encoder and decoder of this disclosure) to carry out the methods described above. Common forms of non-temporary media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROMs and EPROMs, FLASH®-EPROMs or any other flash memory, NVRAMs, caches, registers, any other memory chips or cartridges, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces and / or memory.

[0184]

[0207] It should be noted that relational terms in this specification, such as "First" and "Second," are used solely to distinguish one entity or action from another, and do not imply or require any actual relationship or order between these entities or actions. Furthermore, the words "contain," "have," "contain," and "include," as well as other similar forms, are intended to be synonymous and open-ended in that any element or group of elements following any of these words is not meant to be an exhaustive enumeration of such elements or groups of elements, nor is it meant to be limited to only the enumerated elements or groups of elements.

[0185]

[0208] As used herein, unless otherwise specified, the term "or" encompasses all possible combinations, except in cases where it is not feasible. For example, if it is stated that a database may include A or B, then unless otherwise specified or it is not feasible, the database may include A, B, A and B. As a second example, if it is stated that a database may include A, B, or C, then unless otherwise specified or it is not feasible, the database may include A, B, C, A and B, A and C, B and C, A and B and C.

[0186]

[0209] It is understood that the embodiments described above may be implemented by hardware or software (program code) or a combination of hardware and software. When implemented by software, it may be stored in the computer-readable medium described above. When the software is executed by a processor, it can perform the methods of this disclosure. The computing units and other functional units described in this disclosure may be implemented by hardware or software or a combination of hardware and software. Those skilled in the art will also understand that several of the above modules / units may be combined into a single module / unit, and that each of the above modules / units may be further divided into several submodules / subunits.

[0187]

[0210] In this specification described above, embodiments have been described with reference to many specific details that may differ depending on the implementation. Specific adaptations and modifications of the embodiments described above may be made. Other embodiments may become apparent to those skilled in the art from the considerations herein and the implementation of the invention disclosed herein. This specification and the examples are intended to be considered as examples only, and the true scope and spirit of the invention are indicated by the appended claims. Furthermore, the arrangement of steps shown in the figures is for illustrative purposes only and is not intended to limit to any particular arrangement of steps. Thus, those skilled in the art will understand that these steps may be performed in different orders while carrying out the same method.

[0188]

[0211] Exemplary embodiments are disclosed in the drawings and this specification. However, many variations and modifications can be made to these embodiments. Accordingly, where certain terms are used, they are used merely for general descriptive purposes and not for limiting purposes.

Claims

1. A method for encoding video, In an encoded video sequence (CVS), it is determined whether the number of decoded picture buffer (DPB) parameter syntax structures is equal to the number of output layer sets (OLS), and In the aforementioned CVS, the bitstream is encoded without setting a first DPB syntax element that specifies the index of the DPB parameter syntax structure corresponding to the OLS in the list of DPB parameter syntax structures in the video parameter set (VPS), in response to the number of DPB parameter syntax structures being equal to the number of OLS. Methods that include...

2. To determine whether the number of DPB parameter syntax structures is 1 or less, and Encode the bitstream without setting the first DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being equal to 1. The method according to claim 1, further comprising:

3. In response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS, a first DPB syntax element having a variable length is set within the VPS. The method according to claim 1, further comprising:

4. In the CVS, in response to the number of profile / tier / level (PTL) syntax structures being equal to the number of OLS, the bitstream is encoded without setting a first PTL syntax element that specifies an index of the PTL syntax structure corresponding to the OLS in the VPS against a list of PTL syntax structures in the VPS. The method according to claim 1, further comprising:

5. To determine whether the number of PTL syntax structures is equal to 1, and Encode the bitstream without setting the first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to 1. The method according to claim 4, further comprising:

6. In response to the number of PTL syntax structures being greater than 1 and different from the number of OLS, a first PTL syntax element having a fixed length is set within the VPS. The method according to claim 4, further comprising:

7. A method for decoding video, Receiving a bitstream containing an encoded video sequence (CVS), In the CVS, it is determined whether the number of decoded picture buffer (DPB) parameter syntax structures and the number of output layer sets (OLS) are equal, and In the CVS, in response to the number of DPB parameter syntax structures being equal to the number of OLS, the decoding of a first DPB syntax element that specifies the index of the DPB parameter syntax structure corresponding to the OLS in the list of DPB parameter syntax structures in the video parameter set (VPS) is selectively bypassed. Methods that include...

8. To determine whether the number of DPB parameter syntax structures is 1 or less, and In response to the number of DPB parameter syntax structures being 1 or less, when decoding the VPS, selectively bypass decoding the first DPB syntax element. The method according to claim 7, further comprising:

9. In response to the number of DPB parameter syntax structures being 1 or less, the first DPB syntax element is determined to be zero, and In response to the number of DPB syntax structures being equal to the number of OLS, the first DPB syntax element is determined such that the DPB parameter syntax structure specified by the first DPB syntax element is equal to the index of the OLS to which it applies. The method according to claim 8, further comprising:

10. Decode the first variable-length DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS. The method according to claim 7, further comprising:

11. In the CVS, in response to the number of profile / tier / level (PTL) syntax structures being equal to the number of OLS, the decoding of a first PTL syntax element specifying an index of the PTL syntax structure corresponding to the OLS in the VPS against a list of PTL syntax structures in the VPS is selectively bypassed. The method according to claim 7, further comprising:

12. To determine whether the number of PTL syntax structures is equal to 1, and Selectively bypass decoding the first PTL syntax element in the VPS in response to the number of PTL syntax structures being equal to 1. The method according to claim 11, further comprising:

13. Decode the first fixed-length PTL syntax element in the VPS in response to the number of PTL syntax structures being greater than 1 and different from the number of OLS. The method according to claim 11, further comprising:

14. A method for singling a bitstream, Receiving a video sequence, The aforementioned video sequence, In an encoded video sequence (CVS), it is determined whether the number of decoded picture buffer (DPB) parameter syntax structures is equal to the number of output layer sets (OLS), In the aforementioned CVS, in response to the number of DPB parameter syntax structures being equal to the number of OLS, the bitstream is encoded without setting a first DPB syntax element that specifies the index of the DPB parameter syntax structure corresponding to the OLS in the list of DPB parameter syntax structures in the video parameter set (VPS). Encoding by, and Signaling the bitstream generated based on the aforementioned encoding. Methods that include...

15. The aforementioned video sequence is Determine whether the number of DPB parameter syntax structures is 1 or less, Encoding the bitstream without setting the first DPB syntax element in the VPS in response to the number of DPB parameter syntax structures being equal to 1. The method according to claim 14, which is encoded by

16. The aforementioned video sequence is In response to the number of DPB parameter syntax structures being greater than 1 and different from the number of OLS, a first DPB syntax element having a variable length is set within the VPS. The method according to claim 14, which is encoded by

17. The aforementioned video sequence is In the CVS, in response to the number of profile / tier / level (PTL) syntax structures being equal to the number of OLS, the bitstream is encoded without setting a first PTL syntax element that specifies an index of the PTL syntax structure corresponding to the OLS in the VPS against a list of PTL syntax structures in the VPS. The method according to claim 14, which is encoded by

18. The aforementioned video sequence is Determine whether the number of PTL syntax structures is equal to 1, Encoding the bitstream without setting the first PTL syntax element within the VPS in response to the number of PTL syntax structures being equal to 1. The method according to claim 17, which is encoded by

19. The aforementioned video sequence is In response to the number of PTL syntax structures being greater than 1 and different from the number of OLS, a first PTL syntax element having a fixed length is set within the VPS. The method according to claim 17, which is encoded by

20. The bitstream is stored in a non-temporary computer-readable storage medium. The method according to claim 14, further comprising:

Citation Information

Patent Citations

  • Video decoding method and device using the same

    JP2017508417A

  • Profile, Tier, Level for the 0th Output Layer Set in Video Coding

    JP2017523683A

  • Profile, tier, level for the 0-th output layer set in video coding

    US20150373361A1