A method, an apparatus and a computer program product for video coding and decoding using one or two-byte header

The method and apparatus efficiently derive scalability and temporal sublayer identifiers from video bitstreams using a single header byte, addressing inefficiencies in existing standards and optimizing video encoding and decoding processes.

WO2026114532A1PCT designated stage Publication Date: 2026-06-04NOKIA TECHNOLOGIES OY

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2025-07-11
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing video coding standards face inefficiencies in deriving scalability layer and temporal sublayer identifiers from video bitstreams due to variable header structures, leading to increased complexity and resource utilization.

Method used

A method and apparatus that derive scalability layer and temporal sublayer identifiers using a first byte of the header, determining the presence of a second byte, and utilizing pre-defined information and syntax elements to efficiently determine these identifiers without requiring the second byte when it is absent.

Benefits of technology

This approach simplifies the header processing, reduces resource usage, and enhances the efficiency of video decoding and encoding by accurately determining scalability and temporal sublayer information, thereby optimizing video bitstream handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025069955_04062026_PF_FP_ABST
    Figure EP2025069955_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments concern a method comprising receiving a bitstream comprising one or more units each having a header; receiving a first byte of the header; determining, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream; deriving at least one parameter value depending on the determination if the header does not comprise the second byte, the apparatus is caused to derive the at least one parameter value based on at least one syntax element in the first byte and pre-defined information; wherein the at least one parameter value comprises one or both of a scalability layer identifier and a temporal sublayer identifier The embodiments also concern technical equipment for implementing the method.
Need to check novelty before this filing date? Find Prior Art

Description

A METHOD, AN APPARATUS AND A COMPUTER PROGRAM PRODUCT FOR VIDEO CODING AND DECODINGTechnical Field

[0001] The present solution generally relates to a method and an apparatus and a computer program product for video coding and decoding.Background

[0002] This section is intended to provide a background or context to the invention that is recited in the claims. The description herein may include concepts that could be pursued but are not necessarily ones that have been previously conceived or pursued. Therefore, unless otherwise indicated herein, what is described in this section is not prior art to the description and claims in this application and is not admitted to be prior art by inclusion in this section.

[0003] A video codec comprises an encoder and a decoder. The encoder transforms the input video into a compressed representation suited for storage / transmission. The decoder can un-compress the compressed video representation back into a viewable form. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e., at lower bitrate).Summary

[0004] An aim of some embodiments of the present solution is to provide an improved header for units of a video bitstream.

[0005] The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention.

[0006] Various aspects include a method, an apparatus and a computer readable medium comprising a computer program stored therein, which are characterized by what is stated in the independent claims. Various embodiments are disclosed in the dependent claims.

[0007] According to a first aspect, there is provided an apparatus comprising at least one processor and at least one memory, said at least one memory stored with code thereon, which when executed by said at least one processor, causes the apparatus to:- receive a bitstream comprising one or more units each having a header;- receive a first byte of the header;- determine, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream;- derive at least one parameter value depending on the determination:• if the header does not comprise the second byte, the apparatus is caused to derive the at least one parameter value based on at least one syntax element in the first byte and pre-defined information;- wherein the at least one parameter value comprises one or both of a scalability layer identifier and a temporal sublayer identifier.

[0008] According to a second aspect, there is provided a method comprising- receiving a bitstream comprising one or more units each having a header;- receiving a first byte of the header;- determining, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream;- deriving at least one parameter value depending on the determination:• if the header does not comprise the second byte, the apparatus is caused to derive the at least one parameter value based on at least one syntax element in the first byte and pre-defined information;- wherein the at least one parameter value comprises one or both of a scalability layer identifier and a temporal sublayer identifier.

[0009] According to a third aspect, there is provided an apparatus comprising at least- means for receiving a bitstream comprising one or more units each having a header;- means for receiving a first byte of the header;- means for determining, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream;- means for deriving at least one parameter value depending on the determination:• if the header does not comprise the second byte, the apparatus is caused to derive the at least one parameter value based on at least one syntax element in the first byte and pre-defined information;- wherein the at least one parameter value comprises one or both of a scalability layer identifier and a temporal sublayer identifier.

[0010] According to a fourth aspect, there is provided computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or a system to- receive a bitstream comprising one or more units each having a header;- receive a first byte of the header;- determine, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream;- derive at least one parameter value depending on the determination:• if the header does not comprise the second byte, the apparatus is caused to derive the at least one parameter value based on at least one syntax element in the first byte and pre-defined information;- wherein the at least one parameter value comprises one or both of a scalability layer identifier and a temporal sublayer identifier.

[0011] According to an embodiment, if the header comprises the second byte, the at least one parameter value is derived based on the at least one syntax element in the first byte and information in the second byte.

[0012] According to an embodiment, the header is a network abstraction layer (NAL) unit header.

[0013] According to an embodiment, it is determined from an extension flag in the first byte of the header whether the header comprises the second byte.

[0014] According to an embodiment, the at least one parameter value is derived based on a non-reference flag in the first byte of the header.

[0015] According to an embodiment, the at least one parameter value is derived based on a first flag in the first byte of the header and the extension flag.

[0016] According to an embodiment, when the header does not comprise the second byte, it is determined if a first flag in the first byte is enabled, wherein based on the determination that the first flag is enabled, the temporal sublayer identifier is determined to be equal to a first value, and based on the determination that the first flag is not enabled, the temporal sublayer identifier is determined to be equal to a second value; and wherein when the header comprises the second byte, the temporal sublayer identifier is determined based on the temporal indicator in the second byte and the first flag in the first byte.

[0017] According to an embodiment, it is determined from a no extension flag in a first bit of the header whether the header comprises a second byte.

[0018] According to an embodiment, when the header does not comprise the second byte, the temporal sublayer identifier is derived based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0019] According to an embodiment, when the header comprises the second byte, the temporal sublayer identifier is determined based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on a parameter specifying the one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0020] According to an embodiment, when the header comprises the second byte, the scalability layer identifier is determined from the second byte.

[0021] According to an embodiment, when the header does not comprise the second byte, the scalability layer identifier is determined based on a parameter specifying one or more least significant bits defined in the first byte.

[0022] According to an embodiment, when the header comprises the second byte, the scalability layer identifier is determined based on a parameter specifying one or more least significant bits for the scalability layer identifier defined in the first byte and a parameterspecifying one or more most significant bits for the scalability layer identifier defined in the second byte.

[0023] According to an embodiment, when the header does not comprise the second byte, the temporal sublayer identifier is derived based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte and the scalability layer identifier is derived based on a parameter specifying one or more least significant bits of the scalability layer identifier defined in the first byte.

[0024] According to an embodiment, when the header comprises the second byte, the scalability layer identifier is determined based on a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte and a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte and the temporal sublayer identifier is determined based on a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte and a parameter specifying one or more most significant bits for the temporal sublayer identifier defined in the second byte.

[0025] According to an embodiment, when the header does not comprise the second byte, the temporal sublayer identifier is derived based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0026] According to an embodiment, when the header comprises the second byte, the temporal sublayer identifier is determined based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on the parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0027] According to an embodiment, when the header does not comprise the second byte, the temporal sublayer identifier is determined based on a temporal identifier indicator in the first byte and a unit type is determined based on the temporal identifier indicator and a unit type indicator in the first byte.

[0028] According to an embodiment, when the header comprises the second byte, the temporal sublayer identifier is determined based on a parameter specifying one or more most significant bits of the temporal identifier indicator in the second byte and a temporal identifier indicator in the first byte and a unit type is determined based on the temporal identifier indicator and a unit type indicator in the first byte.

[0029] According to an embodiment, the computer program product is embodied on a non- transitory computer readable medium.Description of the Drawings

[0030] In the following, various embodiments will be described in more detail with reference to the appended drawings, in which

[0031] Fig. 1 shows a simplified example of an encoding process;

[0032] Fig. 2 shows a simplified example of a decoding process;

[0033] Fig. 3 is a flowchart of a method according to an embodiment;

[0034] Fig. 4 is a flowchart of a method according to another embodiment;

[0035] Fig. 5 shows an example of an apparatus;

[0036] Fig. 6 shows a schematic diagram of an example multimedia communication system within which various embodiments may be implemented;

[0037] Fig 7 is a flowchart of a method according to yet another embodiment; and

[0038] Fig. 8 is a flowchart of a method according to yet another embodiment.Embodiments

[0039] The following description and drawings are illustrative to discuss embodiments of the present solution with examples. The specific details are provided for understanding purposes. However, in certain instances, well-known or conventional details are not described in order to avoid obscuring the description. Reference in this specification to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure.

[0040] Before describing the embodiments further, a brief reference to evolution of video coding standardization is given. The present embodiments are suited within the context of next generation video coding standardization, e.g., H.267 video coding standard and the ECM (Enhanced Compression Model) exploration, but are not limited to any particular standard.

[0041] A video codec comprises an encoder and a decoder. The encoder transforms the input video into a compressed representation suited for storage / transmission. The decoder can un-compress the compressed video representation back into a viewable form. The encodermay discard some information in the original video sequence in order to represent the video in a more compact form (i.e., at lower bitrate).

[0042] Figure 1 shows a simple example of a principle of an encoding process for 2D pictures, and Figure 2 shows a simple example of a principle of a decoding process for 2D pictures. In Figure 1, the following have been illustrated:- an image to be encoded (In);- a predicted representation of an image block (P'n);- a prediction error signal (Dn);- a reconstructed prediction error signal (D'n);- a preliminary reconstructed image (I'n);- a final reconstructed image (R'n);- a transform (T) and inverse transform (T-l);- a quantization (Q) and inverse quantization (Q-l);- entropy encoding (E);- a reference frame memory (RFM);- inter prediction (Pinter);- intra prediction (Pintra);- mode selection (MS), and- filtering (F).

[0043] In Figure 2 the following have been illustrated:- a predicted representation of an image block (P'n);- a reconstructed prediction error signal (D'n);- a preliminary reconstructed image (I'n);- a final reconstructed image (R'n); an inverse transform (T-l);- an inverse quantization (Q-l);- an entropy decoding (E-l);- a reference frame memory (RFM);- a prediction (either inter or intra) (P);- and filtering (F).

[0044] An elementary unit for the input to an encoder and the output of a decoder, respectively, in many cases is a picture (also referred to as “an image”). A picture given asan input to an encoder may also be referred to as a source picture, and a picture decoded by a decoded may be referred to as a decoded picture or a reconstructed picture.

[0045] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays:- Luma (Y) only (monochrome).- Luma and two chroma (YCbCr or YCgCo).- Green, Blue and Red (GBR, also known as RGB).- Arrays representing other unspecified monochrome or tri-stimulus color samplings (for example, YZX, also known as XYZ).

[0046] The sample arrays of a picture may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr; regardless of the actual color representation method in use. The actual color representation method in use can be indicated, e.g., in a coded bitstream e.g., using the Video Usability Information (VUI) syntax of HEVC or alike. A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) or the array or a single sample of the array that compose a picture in monochrome format.

[0047] Samples of a sample array have a certain bit depth, such as 8 bits per sample or 10 bits per sample. A bit depth implicitly specifies a value range, which may be referred to as the full range. For example, the full range is from 0 to 255, inclusive, for 8 bits per sample, or from 0 to 1023, inclusive, for 10 bits per sample. The source video may use allocate a narrower sample value range than the full range. A specific value range, sometimes referred to as the studio range, has been specified in the ITU-T H.273 standard specifying codingindependent code points for video. A source value range may interchangeably be referred to as a source sample value range, and may be defined as the sample value range of the video that is given as input to a video encoder to be encoded.

[0048] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays.

[0049] The Advanced Video Coding standard (which may be abbreviated AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). There have been multiple versions of the H.264 / AVC standard, each integrating several extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).

[0050] The High Efficiency Video Coding standard (which may be abbreviated HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. Extensions to H.265 / HEVC include scalable, multiview, three- dimensional, and fidelity range extensions, which may be referred to as SHVC, MV- HEVC, 3D-HEVC, and REXT, respectively.

[0051] Versatile Video Coding (which may be abbreviated VVC, H.266, or H.266 / VVC) is a video compression standard developed as the successor to HEVC. VVC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG-I Part 3.

[0052] ECM was developed by JVET (Joint Video Experts Team) of ITU-T VCEG and ISO / IEC MPEG, to provide a future video coding technology the compression capability of which would exceed that of the VVC.

[0053] A specification of the AVI bitstream format and decoding process were developed by the Alliance for Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.

[0054] Some key definitions, bitstream and coding structures, and concepts of H.264 / AVC, HEVC, VVC, and / or AVI and some of their extensions are described in this section as an example of a video encoder, decoder, encoding method, decoding method, and a bitstream structure, wherein the embodiments may be implemented. The aspects of various embodiments are not limited to H.264 / AVC, HEVC, VVC, and / or AVI or their extensions, but rather the description is given for one possible basis on top of which the present embodiments may be partly or fully realized.

[0055] Hybrid video codecs, for example ITU-T H.263, H.264 / AVC, HEVC, and VVC, may encode the video information in two phases. At first, pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means or by spatial means.

[0056] In motion compensation based prediction (which may be referred to as inter prediction, temporal prediction or motion-compensated temporal prediction or motion-compensated prediction or MCP) an area in one of the previously coded frames that corresponds closely to the block being coded is found and used for prediction. Inter prediction may reduce temporal redundancy.

[0057] In spatial prediction pixel values around the block to be coded are used. In the first phase, predictive coding may be applied, for example, as so-called sample prediction and / or so-called syntax prediction. In the sample prediction, pixel or sample values in a certain picture area or "block" are predicted. These pixel or sample values can be predicted, for example, using one or more of motion compensation or intra prediction mechanisms.

[0058] Intra prediction, where pixel or sample values can be predicted by spatial mechanisms, involve finding and indicating a spatial region relationship. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.

[0059] In motion vector prediction, motion vectors e.g., for inter and / or inter-view prediction may be coded differentially with respect to a block-specific predicted motion vector. In many video codecs, the predicted motion vectors are created in a predefined way, for example by calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions, sometimes referred to as advanced motion vector prediction (AMVP), is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signalling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or co-located blocks in temporal reference picture. Differential coding of motion vectors is typically disabled across slice boundaries.

[0060] The block partitioning, e.g., from a coding tree unit (CTU) to coding units (CUs) and down to prediction units (PUs), may be predicted.

[0061] In filter parameter prediction, the filtering parameters e.g., for sample adaptive offset may be predicted. Prediction approaches using image information from a previously coded image can also be called as inter prediction methods which may also be referred to as temporal prediction and motion compensation. Prediction approaches using image information within the same image can also be called as intra prediction methods.

[0062] In the second phase of encoding, the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This may be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size of transmission bitrate).

[0063] In some video codecs, such as H.265 / HEVC, video pictures are divided into coding units (CU) covering the area of the picture. A CU consists of one or more prediction units (PU) defining the prediction process for the samples within the CU and one or more transform units (TU) defining the prediction error coding process for the samples in the said CU. A CU may consist of a square block of samples with a size selectable from a predefined set of possible CU sizes. A CU with the maximum allowed size is typically named as LCU (largest coding unit) or CTU (coding tree unit) and the video picture is divided into non-overlapping CTUs. A CTU can be further split into a combination of smaller CUs, e.g., by recursively splitting the CTU and resultant CUs. Each resulting CU typically has at least one PU and at least one TU associated with it. Each PU and TU can be further split into smaller PUs and TUs in order to increase granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it defining what kind of a prediction is to be applied for the pixels within that PU (e.g., motion vector information for inter predicted PUs and intra prediction directionality information for intra predicted PUs). Similarly, each TU is associated with information describing the prediction error decoding process for the samples within the said TU (including e.g., DCT coefficient information). It is typically signaled at CU level whether prediction error coding is applied or not for each CU. In the case there is noprediction error residual associated with the CU, it can be considered there are no TUs for the said CU. The division of the image into CUs, and division of CUs into PUs and TUs is typically signaled in the bitstream allowing the decoder to reproduce the intended structure of these units.

[0064] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.

[0065] In many video codecs, including H.264 / AVC, HEVC, and VVC, motion information is indicated by motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder) or decoded (at the decoder) and the prediction source block in one of the previously coded or decoded images (or pictures). In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index may be predicted from adjacent blocks and / or or co-located blocks in temporal reference picture. Moreover, high efficiency video codecs may employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without anymodification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.

[0066] Video codecs may support motion compensated prediction from one source image (uni -prediction) and two sources (bi-prediction). In the case of uni -prediction a single motion vector is applied whereas in the case of bi-prediction two motion vectors are signaled and the motion compensated predictions from two sources are averaged to create the final sample prediction. In the case of weighted prediction, the relative weights of the two predictions can be adjusted, or a signaled offset can be added to the prediction signal

[0067] In addition to applying motion compensation for inter picture prediction, similar approach can be applied to intra picture prediction. In this case the displacement vector indicates where from the same picture a block of samples can be copied to form a prediction of the block to be coded or decoded. This kind of intra block copying methods can improve the coding efficiency substantially in presence of repeating structures within the frame - such as text or other graphics.

[0068] In video codecs the prediction residual after motion compensation or intra prediction may be first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.

[0069] Many video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g., the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + ZR (Eq. 1)Where C is the Lagrangian cost to be minimized, D is the image distortion (e.g., Mean Squared Error) with the mode and motion vectors considered, and R the number of bitsneeded to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).

[0070] A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.

[0071] A syntax element may be defined as an element of data represented in a bitstream. A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.

[0072] Syntax structures may be specified, for example, using arithmetic, logical, relational, bit-wise, and assignment operators similar to those available in many programming languages. For example, & may indicate a bit-wise ‘AND’ operation. Furthermore, syntax structures may be specified with reference to mathematical functions.

[0073] Syntax structures and semantics may use the values of variables derived from the values of syntax elements. Naming conventions may be defined for variables. For example, variables may be named by a mixture of lower case and upper-case letter and without any underscore characters. Variables starting with an upper-case letter may be derived for the decoding of the current syntax structure and all depending syntax structures. Variables starting with an upper-case letter may, in some cases, be used in the decoding process for later syntax structures without mentioning the originating syntax structure of the variable. Variables starting with a lower-case letter may only be used in relation to the syntax structure or function they have been defined for.

[0074] Binary notation may be indicated by suffixing a bitstring with the letter b. For example, 10111100b represents an 8-bit unsigned integer value equal to 188.

[0075] Hexadecimal notation may be indicated by prefixing the hexadecimal number by "Ox". Hexadecimal numbers may be used notation when the number of bits to represent the number is an integer multiple of 4. For example, 0x41 represents an eight-bit string having only its second and its last bits (counted from the most to the least significant bit) equal to 1.

[0076] The phrase along the bitstream (e.g., indicating along the bitstream) may be defined to refer to out-of-band transmission, signaling, or storage in a manner that the out-of-band data is associated with the bitstream. The phrase decoding along the bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream. For example, anindication along the bitstream may refer to metadata in a container file that encapsulates the bitstream.

[0077] Video coding specifications may define an elementary unit that for the output an of an encoder and / or for the input to a decoder. For example, such an elementary unit may be an open bitstream unit (OBU), as specified e.g. in AVI, or a Network Abstraction Layer (NAL) unit, as specified e.g. in HEVC or VVC.

[0078] In some video codecs, an elementary unit for the output of an encoder and the input of a decoder, respectively, is a Network Abstraction Layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format has been specified in some video coding specifications for transmission or storage environments that do not provide framing structures. The bytestream format separates NAL units from each other by attaching a start code in front of each NAL unit. The start code may, for example, comprise a 24-bit byte- aligned value 0x0000001, which may be preceded by a byte that is equal to 0x00. To avoid false detection of NAL unit boundaries, encoders run a byte-oriented start code emulation prevention algorithm, which adds an emulation prevention byte to the NAL unit payload if a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet- and stream-oriented systems, start code emulation prevention may always be performed regardless of whether the bytestream format is in use or not. A NAL unit may be defined as a syntax structure containing an indication of the type of data to follow and bytes containing that data in the form of an RBSP interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure containing an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits containing syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.

[0079] In some video coding formats, a bitstream may comprise a sequence of units, such as NAL units or OBUs.

[0080] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure.

[0081] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.

[0082] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, the header may comprise a size of the payload in bytes.

[0083] In some coding standards, NAL units include a header (also known as a NAL unit header, which may be abbreviated as NUH) and payload. A NAL unit starts with the header, which is followed by the payload. In some coding standards, the NAL unit header indicates the type of the NAL unit. In some coding standards, the NAL unit header indicates a scalability layer identifier (e.g., called nuh layer id), which may be used, e.g., for indicating spatial or quality layers, views of a multiview video, or auxiliary layers (such as depth maps or alpha planes). In some coding standards, the NAL unit header includes a temporal sublayer identifier, which may be used for indicating temporal subsets of the bitstream, such as a 30-frames-per-second subset of a 60-frames-per-second bitstream.

[0084] Bitstreams or coded video sequences may be encoded to be temporally scalable as follows. Each picture may be assigned to a particular temporal sub-layer. A temporal sublayer may be equivalently called a sub-layer, temporal sublayer, sublayer, or temporal level. Temporal sub-layers may be enumerated, e.g., from 0 (zero) upwards. The lowest temporal sub-layer, sub-layer 0, may be decoded independently. Pictures at temporal sublayer 1 may be predicted from reconstructed pictures at temporal sub-layers 0 and 1. Pictures at temporal sub-layer 2 may be predicted from reconstructed pictures at temporal sub-layers 0, 1, and 2, and so on. In other words, a picture at temporal sub-layer N does not use any picture at temporal sub-layer greater than N as a reference for inter prediction. The bitstream created by excluding all pictures greater than or equal to a selected sub-layer value and including pictures remains conforming.

[0085] Each picture of a temporally scalable bitstream may be assigned with a temporal identifier (also known as TID, temporal layer identifier, sub-layer identifier, sublayer identifier, temporal sub-layer identifier, temporal sublayer identifier, or temporal layer ID), which may be, for example, assigned to a variable Temporalld. The temporal identifier may, for example, be indicated in a NAL unit header or in an OBU extension header.Temporalld equal to 0 corresponds to the lowest temporal level. The bitstream created by excluding all coded pictures having a Temporalld greater than or equal to a selected value and including all other coded pictures remains conforming. Consequently, a picture having Temporalld equal to tid value does not use any picture having a Temporalld greater than tid value as a prediction reference.

[0086] In some video coding standards, a sub-layer or a temporal sub-layer may be defined to be a temporal scalable layer (or a temporal layer, TL) of a temporal scalable bitstream, consisting of VCL NAL units with a particular value of the Temporalld variable and the associated non-VCL NAL units.

[0087] The NAL unit header enables use of NAL-unit-structured video bitstreams, such as AVC, HEVC, and VVC bitstreams, in a variety of transmission systems, including MPEG- 2 transport stream and Real-time Transport Protocol (RTP).

[0088] NAL units may be categorized into Video Coding Layer (VCL) NAL units and non- VCL NAL units. In some coding formats, VCL NAL units may be coded slice NAL units.

[0089] The NAL unit header provides information that enables a media aware network element (MANE) or a decoder to determine how to treat individual NAL units under different situations. For example, a MANE may to discard some NAL units to reduce bitrate. As another example, a decoder may select when to begin decoding when it joins an in-progress broadcast. As yet another example, a decoder may select to skip decoding some pictures to maintain real-time operation. As yet another example, a single layer can be extracted from a multi-layer bitstream.

[0090] Some coding formats, such as HEVC and WC, support temporal scalability functionality. When temporal scalability is used, higher temporal sublayers may be discarded without impacting decoding of lower temporal sublayers. The HEVC and VVC reference software encoders support a hierarchical prediction structure which can be used with temporal scalability.

[0091] Scalable video coding may refer to coding structure where one bitstream may include multiple representations of the content, for example, at different bitrates, resolutions or frame rates. In these cases, the receiver can extract the desired representation depending on its characteristics (e.g., resolution that matches best the display device). Alternatively, a server or a network element may extract the portions of the bitstream to be transmitted to the receiver depending on, e.g., the network characteristics or processing capabilities of thereceiver. A meaningful decoded representation may be produced by decoding only certain parts of a scalable bitstream. A scalable bitstream typically include of a ‘base layer’ providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer typically depends on the lower layers. For example, the motion and mode information of the enhancement layer can be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer.

[0092] A scalable bitstream may include a ‘base layer’ providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer may depend on the lower layers. E.g., the motion and mode information of the enhancement layer may be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer. A scalable video codec for quality scalability (also known as signal-to-noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non-scalable video encoder and decoder is used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer for an enhancement layer. In codecs using reference picture list(s) for inter prediction, the base layer decoded pictures may be inserted into a reference picture list(s) for coding / decoding of an enhancement layer picture similarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a base-layer reference picture as inter prediction reference and indicate its use, e.g., with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture is used as inter prediction reference for the enhancement layer. When a decoded base-layer picture is used as prediction reference for an enhancement layer, it is referred to as an inter-layer reference picture.

[0093] It needs to be understood that the description of scalable video coding may be generalized to any scalability hierarchy with more than two layers. In this case, a second enhancement layer may depend on a first enhancement layer in encoding and / or decoding processes, and the first enhancement layer may therefore be regarded as the base layer for the encoding and / or decoding of the second enhancement layer. Furthermore, it needs tobe understood that there may be inter-layer reference pictures from more than one layer in a reference picture buffer or reference picture lists of an enhancement layer, and each of these inter-layer reference pictures may be considered to reside in a base layer or a reference layer for the enhancement layer being encoded and / or decoded. Furthermore, it needs to be understood that other types of inter-layer processing than reference-layer picture upsampling may take place instead or additionally. For example, the bit-depth of the samples of the reference-layer picture may be converted to the bit depth of the enhancement layer and / or the sample values may undergo a mapping from the color space of the reference layer to the color space of the enhancement layer.

[0094] A scalable video coding and / or decoding scheme may use multi-loop coding and / or decoding, which may be characterized as follows. In the encoding / decoding, a base layer picture may be reconstructed / decoded to be used as a motion-compensation reference picture for subsequent pictures, in coding / decoding order, within the same layer or as a reference for inter-layer (or inter-view or inter-component) prediction. The reconstructed / decoded base layer picture may be stored in the decoded picture buffer (DPB). An enhancement layer picture may likewise be reconstructed / decoded to be used as a motion-compensation reference picture for subsequent pictures, in coding / decoding order, within the same layer or as reference for inter-layer (or interview or intercomponent) prediction for higher enhancement layers, when any. In addition to reconstructed / decoded sample values, syntax element values of the base / reference layer or variables derived from the syntax element values of the base / reference layer may be used in the inter-layer / inter-component / inter-view prediction.

[0095] A multi-layer bitstream is a bitstream comprising multiple layers, which may be, but are not limited to, base and enhancement layers as discussed above for scalable video coding. A multilayer bitstream may additionally or alternatively comprise independent layers that do not have inter-layer prediction relationship between each other and may even represent different types of content. Any multi-layer bitstream may be regarded as a scalable video bitstream.

[0096] Some coding formats, such as HEVC and VVC, support multi-layer functionality, e.g. for scalability, multiview and auxiliary pictures. Some coding formats, such as HEVC and VVC, include a video parameter set (VPS) that indicates for multi-layer bitstreams the maximum number of layers in the bitstream and any inter-layer coding dependenciesamong the layers. For example, with 2-layer spatial scalable coding an enhancement layer uses inter-layer coding with a dependency upon a base layer. Enhancement layer NAL units may be discarded without impacting decoding of base layer NAL units. Conversely, discarding of base layer NAL units negatively impacts decoding of an enhancement layer that uses inter-layer prediction from the base layer. Multi-layer bitstreams may also include independent layers, that do not use inter-layer prediction and may be decoded without any other layer.

[0097] In some coding formats, picture unit (PU) may be defined as a set of data units, such as NAL units, that are associated with each other, are consecutive in decoding order, and contain exactly one coded picture. For example, certain non-video-coding data units, such as non-VCL NAL units, may be next to coded video data units in decoding order and the respective picture unit may comprise both these non-video-coding data units and the video coding data units of a coded picture.

[0098] Video bitstreams, such as AVC, HEVC and VVC bitstreams, may be carried in an MPEG-2 program stream, as defined in ITU-T Rec. H.222.0 | ISO / IEC 13818-1, or use other transport mechanisms such as RTP.

[0099] MPEG-2 Program Elementary Stream (PES) start codes comprise a 24-bit prefix (0x000001) and an 8-bit stream id. Values of stream id in the range of 188 to 255, inclusive, are defined in ITU-T Rec. H.222.0 | ISO / IEC 13818-1 for purposes other than indicating units of a video bitstream.

[0100] Start code emulation avoidance in AVC, HEVC, and VVC:

[0101] The first byte in the NUH in some video coding formats, such as AVC, HEVC, and VVC, is designed to avoid MPEG-2 program stream stream id values in the range of 188 to 255.

[0102] The NUH designs in some video coding formats, such as AVC, HEVC, and VVC, avoid MPEG-2 PES start code emulation, using the forbidden zero bit. The first syntax element, forbidden zero bit is always equal to 0, making it impossible for the first byte to have values greater or equal than 0x80, hence avoiding the values 188 to 255 which may emulate stream id values in MPEG-2 PES start codes.

[0103] forbidden zero bit causes the value of the first byte of the NAL unit header to be in the range of 0 to 127, inclusive. Hence, the combination of a start code prefix 0x000001, which is used in the bytestream format of some video coding specifications, such as VVC,and the first byte of the NAL unit header does not cause an emulation of MPEG-2 PES start code. However, it is asserted that MPEG-2 PES start code emulation can be avoided by disallowing values 188 to 255, inclusive, in the first byte of the NAL unit header, while values in the range of 0 to 187, inclusive, could be allowed.

[0104] The syntax of the second byte in the NUH of some video coding formats, such as VVC, is designed to avoid the value of the byte being equal to 0. The second byte of the VVC NUH includes a temporal_id_plusl syntax element, which is used to signal the temporal ID for the NAL unit plus 1. The value of temporal_id_plusl shall not be equal to 0. Values of temporal ID between 0 and 6 are supported.

[0105] Signaling temporal_id_plusl in the second byte of the VVC NUH rather than directly signalling the temporal ID value causes the syntax element to have a value of at least 1, avoiding the syntax element and the second byte from having a value of 0. It is useful to avoid the value 0 in the second byte because emulation prevention byte stuffing is required if the first three bytes of the NAL unit, including the NUH, is equal to 0x000000, 0x000001 , or 0x000002. Consequently, encoders need not perform start code emulation prevention for NUHs, start code emulation bytes do not appear within NUHs, and decoders or other entities may use a NUH prior to performing start code emulation prevention byte removal for the respective NAL unit and may perform start code emulation prevention byte removals for NAL unit payloads (excluding NUHs).

[0106] Bytestream format and MPEG-2 Program stream:

[0107] MPEG-2 Program Elementary Stream (PES) start codes typically consist of a 24-bit prefix (0x000001) and an 8-bit stream id. The values for stream id typically are in the range of 10111100b to 11111111b, i.e. 188 to 255.

[0108] Some video coding specifications, such as AVC, HEVC, and VVC, contain an optional byte stream format. The byte stream NAL unit syntax and semantics describe how to identify emulation prevention bytes, utilizing all bytes in the NAL unit, including the NUH.

[0109] The byte stream format has been formulated as follows:

[0110] “Some systems (e.g., H.320 and MPEG-2 / H.222.0 systems) require delivery of the entire or partial NAL unit stream as an ordered stream of bytes or bits within which the locations of NAL unit boundaries need to be identifiable from patterns within the coded data itself. For use in such systems, the H.264 / AVC specification defines a byte stream format. In the byte stream format, each NAL unit is prefixed by a specific pattern of threebytes called a start code prefix. The boundaries of the NAL unit can then be identified by searching the coded data for the unique start code prefix pattern. The use of emulation prevention bytes guarantees that start code prefixes are unique identifiers of the start of a new NAL unit. A small amount of additional data (one byte per video picture) is also added to allow decoders that operate in systems that provide streams of bits without alignment to byte boundaries to recover the necessary alignment from the data in the stream. Additional data can also be inserted in the byte stream format that allows expansion of the amount of data to be sent and can aid in achieving more rapid byte alignment recovery, if desired.”

[0111] The start code prefix of the VVC byte stream format is 0x000001.

[0112] To avoid accidental emulation of a start code, an emulation prevention byte is inserted into the NAL unit when a start code prefix would otherwise occur. An emulation prevention byte has value 0x03 in some video coding specifications, such as VVC, but it could generally have any pre-defined value that avoids a start code emulation. Insertion of emulation prevention bytes to avoid start code emulation causes an increase in bitrate. When video is coded at high bitrates, the overhead is small, but becomes more impactful at lower bitrates.

[0113] The syntax of the byte stream NAL unit in some video coding specifications, such as VVC, is as follows:

[0114] The leading_zero_8bits syntax element could only be present in the first byte streamNAL unit of the bitstream, because any bytes equal to 0x00 that follow a NAL unit syntaxstructure and precede the four-byte sequence 0x00000001 (which is to be interpreted as a zero byte followed by a start_code_prefix_one_3bytes) would be considered to be trailing_zero_8bits syntax elements that are part of the preceding byte stream NAL unit.

[0115] zero_byte is a single byte equal to 0x00.

[0116] In VVC, when one or more of the following conditions are true, the zero byte syntax element shall be present:- The nal unit type within the nal_unit( ) syntax structure is equal to DCI NUT (13), OPI NUT (12), VPS NUT (14), SPS NUT (15), PPS_NUT(16), PREFIX_APS_NUT(17), or SUFFIX_APS_NUT(18).- The byte stream NAL unit syntax structure contains the first NAL unit of an AU in decoding order.

[0117] NAL unit:

[0118] NAL units in some video coding formats, such as AVC, HEVC, and VVC, contain a NAL unit header, raw byte sequence payload (RBSP) bytes, and emulation prevention bytes.

[0119] VVC provides the following definitions of emulation prevention byte and raw byte sequence payload (RBSP): emulation prevention byte: A byte equal to 0x03 that is present within a NAL unit when the syntax elements of the bitstream form certain patterns of byte values in a manner that ensures that no sequence of consecutive byte-aligned bytes in the NAL unit can contain a start code prefix. raw byte sequence payload (RBSP): A syntax structure containing an integer number of bytes that is encapsulated in a NAL unit and is either empty or has the form of a string of data bits containing syntax elements followed by an RBSP stop bit and zero or more subsequent bits equal to 0.

[0120] The NAL unit syntax and semantics according to VVC are provided below.

[0121] emulation_prevention_three_byte is a byte equal to 0x03. When an emulation_prevention_three_byte is present in the NAL unit, it shall be discarded by the decoding process.

[0122] The last byte of the NAL unit shall not be equal to 0x00.

[0123] Within the NAL unit, the following three-byte sequences shall not occur at any byte- aligned position:- 0x000000;- 0x000001;- 0x000002.

[0124] Within the NAL unit, any four-byte sequence that starts with 0x000003 other than the following sequences shall not occur at any byte-aligned position:- 0x00000300;- 0x00000301;- 0x00000302;- 0x00000303.

[0125] NAL unit types in AVC, HEVC, and VVC:

[0126] VVC, HEVC, and AVC each define a list of NAL unit type (NUT) values, as shown in the following sections.

[0127] The NUH design of AVC includes a nal ref idc syntax element that indicates that that the slice in the NAL unit is in a non-reference picture, meaning that it cannot be used as a reference picture for any other pictures in the coded video sequence (CVS).

[0128] In HEVC, a distinction is made between NUTs that indicate that the slice in the NAL unit is in a picture that may be used as a reference picture for subsequent pictures in the same temporal sublayer or in a picture that may not be used as a reference picture for subsequent pictures in the same temporal sublayer.

[0129] JVET contributions during development of VVC NAL unit header:

[0130] JVET-N0067 “AHG17: On the first byte of the NAL unit header,” proposed to remove forbidden zero bit from the NAL unit header and asserted that MPEG-2 Program Elementary Stream (PES) start code emulation is avoided with the following proposed syntax element arrangement in the first byte of the NAL unit header.

[0131] NalUnitType is set equal to ( zero tid required flag « 4 ) + nal unit type lsb.

[0132] The assignment of NAL unit type values to types is rearranged, zero tid required flag shall be equal to 1 for DPS, SPS, EOS, EOB, and IRAP NAL units, which therefore have NalUnitType in the range of 16 to 31, inclusive. The proposed number and assignment of reserved NAL unit types differs from that in VVC draft 4.

[0133] When zero tid required flag is equal to 1, nuh_temporal_id_plusl shall be equal to 1 (i.e., Temporalld shall be equal to 0).

[0134] JVET-O0237, “AHG17: Compact NAL unit header,” proposes a NAL unit header design with a length of either one or two bytes, based on an extension flag. The proposeddesign includes a temporal ID and a 2 -bit NAL unit type LSB in the first byte and a layer ID and 1 -bit NAL unit type extension in the second byte.

[0135] vcl flag equal to 1 indicates that the NAL unit is a VCL NAL unit or that the NAL unit type is unspecified, vcl flag equal to 0 indicates that the NAL unit is a non-VCL NAL unit.

[0136] nal_unit_type_lsb specifies the least significant bits for the NAL unit type.

[0137] nuh extension flag equal to 1 specifies that the NAL unit header has an extension, nuh extension flag equal to 0 specifies that there is no NAL unit header extension.

[0138] nal_unit_type_ext_bit specifies the extension bit for the NAL unit type. If not present, nal unit type ext bit is inferred to be equal to 0.

[0139] The variable NalUnitType, which specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit as specified in Table 7 1, is derived as follows:

[0140] NalUnitType = ( zero_tid_required_flag « 4 ) + ( vcl_flag « 3 ) + ( nal unit type ext bit « 2 ) + nal unit type lsb

[0141] VVC NAL unit header:

[0142] The NUH syntax table and semantics for VVC is shown below. It uses 2 bytes.

[0143] forbidden_zero_bit shall be equal to 0.

[0144] nuh_reserved_zero_bit shall be equal to 0. The value 1 of nuh reserved zero bit could be specified in the future by ITU-T | ISO / IEC. Although the value of nuh reserved zero bit is required to be equal to 0 in this version of this Specification, decoders conforming to this version of this Specification shall also allow the value of nuh reserved zero bit equal to 1 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh reserved zero bit equal to 1.

[0145] nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of a layer to which a non-VCL NAL unit applies. The value of nuh layer id shall be in the range of 0 to 55, inclusive. Other values for nuh layer id are reserved for future use by ITU-T | ISO / IEC. Although the value of nuh layer id is required to be the range of 0 to 55, inclusive, in this version of this Specification, decoders conforming to this version of this Specification shall also allow the value of nuh layer id to be greater than 55 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_layer_id greater than 55.

[0146] The value of nuh layer id shall be the same for all VCL NAL units of a coded picture. The value of nuh layer id of a coded picture or a PU is the value of the nuh layer id of the VCL NAL units of the coded picture or the PU.

[0147] When nal unit type is equal to PH NUT, or FD NUT, nuh layer id shall be equal to the nuh layer id of associated VCL NAL unit.

[0148] When nal unit type is equal to EOS NUT, nuh layer id shall be equal to one of the nuh layer id values of the layers present in the CVS.

[0149] It is to be noted that the value of nuh layer id for DCI, OPI, VPS, AUD, and EOB NAL units is not constrained.

[0150] nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit as specified in the table below.

[0151] NAL units that have nal_unit_type in the range of UNSPEC_28..UNSPEC_31, inclusive, for which semantics are not specified, shall not affect the decoding process of the present embodiments.

[0152] It is to be noted that NAL unit types in the range of UNSPEC_28..UNSPEC_31 could be used as determined by the present embodiments. No decoding process for these values of nal unit type is specified in this Specification. Since different applications might use these NAL unit types for different purposes, particular care is expected to be exercised in the design of encoders that generate NAL units with these nal unit type values, and in the design of decoders that interpret the content of NAL units with these nal unit type values. This Specification does not define any management for these values. These nal unit type values might only be suitable for use in contexts in which "collisions" of usage (i.e., different definitions of the meaning of the NAL unit content for the same nal unit type value) are unimportant, or not possible, or are managed - e.g., defined or managed in the controlling application or transport specification, or by controlling the environment in which bitstreams are distributed.

[0153] For purposes other than determining the amount of data in the DUs of the bitstream, decoders shall ignore (remove from the bitstream and discard) the contents of all NAL units that use reserved values of nal unit type.

[0154] It is to be noted that this requirement allows future definition of compatible extensions to the present embodiments.

[0155] It is to be noted that a clean random access (CRA) picture could have associated RASL or RADL pictures present in the bitstream.

[0156] It is to be noted that an instantaneous decoding refresh (IDR) picture having nal unit type equal to IDR N LP does not have associated leading pictures present in the bitstream. An IDR picture having nal unit type equal to IDR W RADL does not have associated RASL pictures present in the bitstream, but could have associated RADL pictures in the bitstream.

[0157] The value of nal unit type shall be the same for all VCL NAL units of a subpicture. A subpicture is referred to as having the same NAL unit type as the VCL NAL units of the subpicture.

[0158] For VCL NAL units of any particular picture, the following applies:- If pps_mixed_nalu_types_in_pic_flag is equal to 0, the value of nal unit type shall be the same for all VCL NAL units of a picture, and a picture or a PU is referred to as having the same NAL unit type as the coded slice NAL units of the picture or PU.- Otherwise (pps_mixed_nalu_types_in_pic_flag is equal to 1), all of the following constraints apply:- The picture shall have at least two subpictures.- VCL NAL units of the picture shall have two or more different nal unit type values.- There shall be no VCL NAL unit of the picture that has nal unit type equal to GDR NUT.- When a VCL NAL unit of the picture has nal unit type equal to nalUnitTypeA that is equal to IDR W RADL, IDR N LP, or CRA NUT, other VCL NAL units of the picture shall all have nal unit type equal to nalUnitTypeA or TRAIL NUT.

[0159] The value of nal unit type shall be the same for all pictures in an IRAP or GDR AU.

[0160] When sps_video_parameter_set_id is greater than 0, vps_max_tid_il_ref_pics_plusl[ i ][ j ] is equal to 0 for j equal toGeneralLayerldxf nuh layer id ] and any value of i in the range of j + 1 to vps max layers minusl, inclusive, and pps_mixed_nalu_types_in_pic_flag is equal to 1, the value of nal unit type shall not be equal to IDR W RADL, IDR N LP, or CRA NUT.

[0161] It is a requirement of bitstream conformance that the following constraints apply:- When a picture is a leading picture of an IRAP picture, it shall be a RADL or RASL picture.- When a subpicture is a leading subpicture of an IRAP subpicture, it shall be a RADL or RASL subpicture.- When a picture is not a leading picture of an IRAP picture, it shall not be a RADL or RASL picture.- When a subpicture is not a leading subpicture of an IRAP subpicture, it shall not be a RADL or RASL subpicture.- No RASL pictures shall be present in the bitstream that are associated with an IDR picture.- No RASL subpictures shall be present in the bitstream that are associated with an IDR subpicture.- No RADL pictures shall be present in the bitstream that are associated with an IDR picture having nal unit type equal to IDR N LP.It is to be noted that it is possible to perform random access at the position of an IRAP AU by discarding all PUs before the IRAP AU (and to correctly decode the non-RASL pictures in the IRAP AU and all the subsequent AUs in decoding order), provided each parameter set is available (either in the bitstream or by external means not specified in the present embodiments) when it is referenced.- No RADL subpictures shall be present in the bitstream that are associated with an IDR subpicture having nal unit type equal to IDR N LP.- Any picture, with nuh layer id equal to a particular value layerld, that precedes an IRAP picture with nuh layer id equal to layerld in decoding order shall precede the IRAP picture in output order and shall precede any RADL picture associated with the IRAP picture in output order.Any subpicture, with nuh layer id equal to a particular value layerld and subpicture index equal to a particular value subpicldx, that precedes, in decoding order, an IRAP subpicture with nuh layer id equal to layerld and subpicture index equal to subpicldx shall precede, in output order, the IRAP subpicture and all its associated RADL subpictures.Any picture, with nuh layer id equal to a particular value layerld, that precedes a recovery point picture with nuh layer id equal to layerld in decoding order shall precede the recovery point picture in output order.Any subpicture, with nuh layer id equal to a particular value layerld and subpicture index equal to a particular value subpicldx, that precedes, in decoding order, a subpicture with nuh layer id equal to layerld and subpicture index equal to subpicldx in a recovery point picture shall precede that subpicture in the recovery point picture in output order.Any RASL picture associated with a CRA picture shall precede any RADL picture associated with the CRA picture in output order.Any RASL subpicture associated with a CRA subpicture shall precede any RADL subpicture associated with the CRA subpicture in output order.Any RASL picture, with nuh layer id equal to a particular value layerld, associated with a CRA picture shall follow, in output order, any IRAP or GDR picture with nuh layer id equal to layerld that precedes the CRA picture in decoding order.Any RASL subpicture, with nuh layer id equal to a particular value layerld and subpicture index equal to a particular value subpicldx, associated with a CRA subpicture shall follow, in output order, any IRAP or GDR subpicture, with nuh layer id equal to layerld and subpicture index equal to subpicldx, that precedes the CRA subpicture in decoding order.If sps field seq flag is equal to 0, the following applies: when the current picture, with nuh layer id equal to a particular value layerld, is a leading picture associated with an IRAP picture, it shall precede, in decoding order, all non-leading pictures that are associated with the same IRAP picture. Otherwise (sps field seq flag is equal to 1), let picA and picB be the first and the last leading pictures, in decoding order, associated with an IRAP picture, respectively, there shall be at most onenon-leading picture with nuh layer id equal to layerld preceding picA in decoding order, and there shall be no non-leading picture with nuh layer id equal to layerld between picA and picB in decoding order.- If sps field seq flag is equal to 0, the following applies: when the current subpicture, with nuh layer id equal to a particular value layerld and subpicture index equal to a particular value subpicldx, is a leading subpicture associated with an IRAP subpicture, it shall precede, in decoding order, all non-leading subpictures that are associated with the same IRAP subpicture. Otherwise (sps field seq flag is equal to 1), let subpicA and subpicB be the first and the last leading subpictures, in decoding order, associated with an IRAP subpicture, respectively, there shall be at most one non-leading subpicture with nuh layer id equal to layerld and subpicture index equal to subpicldx preceding subpicA in decoding order, and there shall be no non-leading picture with nuh layer id equal to layerld and subpicture index equal to subpicldx between picA and picB in decoding order.

[0162] nuh_temporal_id_plusl minus 1 specifies a temporal identifier for the NAL unit.

[0163] The value of nuh_temporal_id_plusl shall not be equal to 0.

[0164] The variable Temporalld is derived as follows:Temporalld = nuh_temporal_id_plusl - 1

[0165] When nal unit type is in the range of IDR W RADL to RSV IRAP l l, inclusive, Temporalld shall be equal to 0.

[0166] When nal unit type is equal to STSA NUT and vps_independent_layer_flag[ GeneralLayerIdx[ nuh layer id ] ] is equal to 1, Temporalld shall be greater than 0.

[0167] The value of Temporalld shall be the same for all VCL NAL units of an AU. The value of Temporalld of a coded picture, a PU, or an AU is the value of the Temporalld of the VCL NAL units of the coded picture, PU, or AU. The value of Temporalld of a sublayer representation is the greatest value of Temporalld of all VCL NAL units in the sublayer representation.

[0168] The value of Temporalld for non-VCL NAL units is constrained as follows:If nal unit type is equal to DCI_NUT, OPI_NUT, VPS_NUT, or SPS_NUT, Temporalld shall be equal to 0 and the Temporalld of the AU containing the NAL unit shall be equal to 0.- Otherwise, if nal unit type is equal to PH NUT, Temporalld shall be equal to the Temporalld of the PU containing the NAL unit.- Otherwise, if nal unit type is equal to EOS NUT or EOB NUT, Temporalld shall be equal to 0.- Otherwise, if nal_unit_type is equal to AUD NUT, FD_NUT, PREFIX SEI NUT, or SUFFIX SEI NUT, Temporalld shall be equal to the Temporalld of the AU containing the NAL unit.- Otherwise, when nal unit type is equal to PPS NUT, PREFIX APS NUT, or SUFFIX APS NUT, Temporalld shall be greater than or equal to the Temporalld of the PU containing the NAL unit.It is to be noted that when the NAL unit is a non-VCL NAL unit, the value of Temporalld is equal to the minimum value of the Temporalld values of all AUs to which the non-VCL NAL unit applies. When nal unit type is equal to PPS NUT, PREFIX APS NUT, or SUFFIX APS NUT, Temporalld could be greater than or equal to the Temporalld of the containing AU, as all PPSs and APSs could be included in the beginning of the bitstream (e.g., when they are transported out-of-band, and the receiver places them at the beginning of the bitstream), wherein the first coded picture has Temporalld equal to 0.

[0169] HEVC NAL unit header

[0170] The NUH syntax table and semantics for HEVC is shown below. It uses 2 bytes.

[0171] forbidden_zero_bit shall be equal to 0.

[0172] nal_unit_type specifies the type of RBSP data structure contained in the NAL unit as specified in the table below.

[0173] NAL units that have nal_unit_type in the range of UNSPEC48..UNSPEC63, inclusive, for which semantics are not specified, shall not affect the decoding process specified by the present embodiments.

[0174] It is to be noted that NAL unit types in the range of UNSPEC48..UNSPEC63 can be used as determined by the present embodiments. No decoding process for these values ofnal unit type is specified by the present embodiments. Since different applications might use these NAL unit types for different purposes, particular care is expected to be exercised in the design of encoders that generate NAL units with these nal unit type values, and in the design of decoders that interpret the content of NAL units with these nal unit type values. The present embodiments does not define any management for these values. These nal unit type values might only be suitable for use in contexts in which "collisions" of usage (i.e., different definitions of the meaning of the NAL unit content for the same nal unit type value) are unimportant, or not possible, or are managed - e.g., defined or managed in the controlling application or transport specification, or by controlling the environment in which bitstreams are distributed.

[0175] For purposes other than determining the amount of data in the decoding units of the bitstream, decoders shall ignore (remove from the bitstream and discard) the contents of all NAL units that use reserved values of nal_unit_type.

[0176] It is to be noted that this requirement allows future definition of compatible extensions to the present embodiments.

[0177] It is to be noted that a clean random access (CRA) picture can have associated random access skipped leading (RASL) or random access decodable leading (RADL) pictures present in the bitstream.

[0178] It is to be noted that a broken link access (BLA) picture having nal unit type equal to BLA W LP can have associated RASL or RADL pictures present in the bitstream. A BLA picture having nal unit type equal to BLA W RADL does not have associated RASL pictures present in the bitstream, but can have associated RADL pictures in the bitstream. A BLA picture having nal unit type equal to BLA N LP does not have associated leading pictures present in the bitstream.

[0179] It is to be noted that an instantaneous decoding refresh (IDR) picture having nal unit type equal to IDR N LP does not have associated leading pictures present in the bitstream. An IDR picture having nal unit type equal to IDR W RADL does not have associated RASL pictures present in the bitstream, but can have associated RADL pictures in the bitstream.

[0180] It is to be noted that a sub-layer non-reference (SLNR) picture is not included in any of RefPicSetStCurrBefore, RefPicSetStCurrAfter and RefPicSetLtCurr of any picture with the same value of Temporalld, and can be discarded without affecting the decodability of other pictures with the same value of Temporalld.

[0181] All coded slice segment NAL units of an access unit shall have the same value of nal unit type. A picture or an access unit is also referred to as having a nal unit type equal to the nal unit type of the coded slice segment NAL units of the picture or access unit.

[0182] If a picture has nal_unit_type equal to TRAIL_N, TSA_N, STSA_N, RADL_N, RASL_N, RSV VCL N10, RSV_VCL_N12 or RSV_VCL_N14, the picture is an SLNR picture. Otherwise, the picture is a sub-layer reference picture.

[0183] Each picture, other than the first picture in the bitstream in decoding order, is considered to be associated with the previous intra random access point (IRAP) picture in decoding order.

[0184] When a picture is a leading picture, it shall be a RADL or RASL picture.

[0185] When a picture is a trailing picture, it shall not be a RADL or RASL picture.

[0186] When a picture is a leading picture, it shall precede, in decoding order, all trailing pictures that are associated with the same IRAP picture.

[0187] No RASL pictures shall be present in the bitstream that are associated with a BL A picture having nal unit type equal to BLA W RADL or BLA N LP.

[0188] No RASL pictures shall be present in the bitstream that are associated with an IDR picture.

[0189] No RADL pictures shall be present in the bitstream that are associated with a BL A picture having nal unit type equal to BLA N LP or that are associated with an IDR picture having nal unit type equal to IDR N LP.

[0190] It is to be noted that it is possible to perform random access at the position of an IRAP access unit by discarding all access units before the IRAP access unit (and to correctly decode the IRAP picture and all the subsequent non-RASL pictures in decoding order), provided each parameter set is available (either in the bitstream or by external means not specified in this Specification) when it needs to be activated.

[0191] Any picture that has PicOutputFlag equal to 1 that precedes an IRAP picture in decoding order shall precede the IRAP picture in output order and shall precede any RADL picture associated with the IRAP picture in output order.

[0192] Any RASL picture associated with a CRA or BLA picture shall precede any RADL picture associated with the CRA or BLA picture in output order.

[0193] Any RASL picture associated with a CRA picture shall follow, in output order, any IRAP picture that precedes the CRA picture in decoding order.

[0194] When sps temporal id nesting flag is equal to 1 and Temporalld is greater than 0, the nal unit type shall be equal to TSA R, TSA N, RADL R, RADL N, RASL R or RASL N.

[0195] nuh layer id specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of a layer to which a non-VCL NAL unit applies. The value of nuh layer id shall be in the range of 0 to 62, inclusive. The value of 63 may be specified in the futureby ITU-T | ISO / IEC. For purposes other than determining the amount of data in the decoding units of the bitstream, decoders shall ignore all data that follow the value 63 for nuh_layer_id in a NAL unit.

[0196] It is to be noted that the value of 63 for nuh layer id could be used to indicate an extended layer identifier in a future extension of this Specification.

[0197] The value of nuh layer id shall be the same for all VCL NAL units of a coded picture. The value of nuh layer id of a coded picture is the value of the nuh layer id of the VCL NAL units of the coded picture.

[0198] When nal unit type is equal to EOB NUT, the value of nuh layer id shall be equal to 0.

[0199] nuh_temporal_id_plusl minus 1 specifies a temporal identifier for the NAL unit. The value of nuh_temporal_id_plusl shall not be equal to 0.

[0200] The variable Temporalld is specified as follows:Temporalld = nuh_temporal_id_plusl - 1

[0201] When nal unit type is in the range of BLA W LP to RSV IRAP VCL23, inclusive, i.e., the coded slice segment belongs to an HEAP picture, Temporalld shall be equal to 0.

[0202] When nal unit type is equal to TSA R or TSA N, Temporalld shall not be equal to 0.

[0203] When nuh layer id is equal to 0 and nal unit type is equal to STSA R or STSA N, Temporalld shall not be equal to 0.

[0204] The value of Temporalld shall be the same for all VCL NAL units of an access unit. The value of Temporalld of a coded picture or an access unit is the value of the Temporalld of the VCL NAL units of the coded picture or the access unit. The value of Temporalld of a sub-layer representation is the greatest value of Temporalld of all VCL NAL units in the sub-layer representation.

[0205] The value of Temporalld for non-VCL NAL units is constrained as follows:- If nal unit type is equal to VPS NUT or SPS NUT, Temporalld shall be equal to 0 and the Temporalld of the access unit containing the NAL unit shall be equal to 0.- Otherwise if nal unit type is equal to EOS NUT or EOB NUT, Temporalld shall be equal to 0.- Otherwise, if nal unit type is equal to AUD NUT or FD NUT, Temporalld shall be equal to the Temporalld of the access unit containing the NAL unit.- Otherwise, Temporalld shall be greater than or equal to the Temporalld of the access unit containing the NAL unit.

[0206] It is to be noted that when the NAL unit is a non-VCL NAL unit, the value of Temporalld is equal to the minimum value of the Temporalld values of all access units to which the non-VCL NAL unit applies. When nal unit type is equal to PPS NUT, Temporalld can be greater than or equal to the Temporalld of the containing access unit, as all picture parameter sets (PPSs) can be included in the beginning of a bitstream, wherein the first coded picture has Temporalld equal to 0. When nal unit type is equal to PREFIX SEI NUT or SUFFIX SEI NUT, Temporalld can be greater than or equal to the Temporalld of the containing access unit, as an SEI NAL unit can contain information, e.g., in a buffering period SEI message or a picture timing SEI message, that applies to a bitstream subset that includes access units for which the Temporalld values are greater than the Temporalld of the access unit containing the SEI NAL unit.

[0207] AVC NAL unit header:

[0208] AVC uses a single byte NUH design for single-layer bitstreams. For its multi-view and scalability extensions, particular values of nal unit type are used to indicate additional bytes in NUH.

[0209] The single byte design contain a nal ref idc syntax element which is used to indicate that slices contained in the NAL unit are in a non-reference picture, which is not used for inter-prediction of any other picture.

[0210] nal_ref_idc not equal to 0 specifies that the content of the NAL unit contains a sequence parameter set, a sequence parameter set extension, a subset sequence parameter set, a picture parameter set, a slice of a reference picture, a slice data partition of a reference picture, or a prefix NAL unit preceding a slice of a reference picture.

[0211] For coded video sequences conforming to one or more of the profiles that are decoded using the decoding process, nal ref idc equal to 0 for a NAL unit containing a slice or slice data partition indicates that the slice or slice data partition is part of a non-reference picture.

[0212] nal ref idc shall not be equal to 0 for sequence parameter set or sequence parameter set extension or subset sequence parameter set or picture parameter set NAL units. When nal ref idc is equal to 0 for one NAL unit with nal unit type in the range of 1 to 4, inclusive, of a particular picture, it shall be equal to 0 for all NAL units with nal unit type in the range of 1 to 4, inclusive, of the picture.

[0213] nal ref idc shall not be equal to 0 for NAL units with nal unit type equal to 5.

[0214] nal ref idc shall be equal to 0 for all NAL units having nal unit type equal to 6, 9, 10, 11, or 12.

[0215] AVC NUH extension for SVC

[0216] AVC NUH extension for 3D AVC:

[0217] AVI OBU header:

[0218] obu forbidden bit must be set to 0.

[0219] obu type specifies the type of data structure contained in the OBU payload: obu type Name of obu type0 Reserved1 OBU SEQUENCE HEADER2 OBU TEMPORAL DELIMITER3 OBU FRAME HEADER4 OBU TILE GROUP5 OBU METADATA6 OBU FRAME7 OBU REDUNDANT FRAME HEADER8 OBU TILE LIST9-14 Reserved15 OBU PADDING

[0220] obu extension flag indicates if the optional obu extension header is present.

[0221] obu has size field equal to 1 indicates that the obu size syntax element will be present, obu has size field equal to 0 indicates that the obu size syntax element will not be present.

[0222] obu reserved lbit must be set to 0. The value is ignored by a decoder.

[0223] temporal id specifies the temporal level of the data contained in the OBU. When temporal id is not present, temporal id is inferred to be equal to 0.

[0224] spatial id specifies the spatial level of the data contained in the OBU. When spatial id is not present, spatial id is inferred to be equal to 0.

[0225] The term “spatial” refers to the fact that the enhancement here occurs in the spatial dimension: either as an increase in spatial resolution, or an increase in spatial fidelity (increased SNR).

[0226] Tile group OBU data associated with spatial id and temporal id equal to 0 are referred to as the base layer, whereas tile group OBU data that are associated with spatial id greater than 0 or temporal id greater than 0 are referred to as enhancement layer(s).

[0227] Coded video data of a temporal level with temporal id T and spatial level with spatial id S are only allowed to reference previously coded video data of temporal id T’ and spatial id S’, where T’ <= T and S’ <= S.

[0228] extension_header_reserved_3bits must be set to 0. The value is ignored by a decoder.

[0229] As video coding standards become more advanced and efficient, the bitrate used for high-level syntax such as NAL unit headers (NUHs) becomes a more significant percentage of the overall bitrate.

[0230] The initial AVC NUH had a length of 1 byte, which was extended an additional 2 or 3 bytes for the scalable video coding (SVC) and 3D extensions in order to carry additional information. The HEVC and VVC NUHs have a fixed length of 2 bytes.

[0231] The AVC, HEVC, and VVC NUH designs avoid MPEG-2 start code emulation and can be used with the byte stream format defined in those standards.

[0232] To avoid accidental emulation of a start code in some video coding standards, such as VVC, an emulation prevention byte is inserted into the NAL unit when a start code prefix would otherwise occur. Insertion of emulation prevention bytes to avoid start code emulation causes an increase in bitrate. When video is coded at high bitrates, the overhead is small, but becomes more impactful at lower bitrates.

[0233] The NUH design for a future video coding standard should reduce bitrate while avoiding MPEG-2 start code emulation and providing flexibility.

[0234] Embodiments that are presented in the following are described with reference to NUH. It is appreciated that the embodiments may apply similarly to any unit header, such as OBU header.

[0235] Some embodiments that are presented in the following are described to avoid a set of disallowed values of the first byte of a NUH, which may be selected to avoid MPEG-2 PESstart code emulation and may comprise values from 188 to 255, inclusive. However, it is to be understood that embodiments may be applied to any byte stream format with a different set of disallowed values to be avoided in the first byte of a NUH.

[0236] Some embodiments that are presented in the following are described to cause the first byte of a NUH to have a value among a set of allowed values, which may be selected to exclude values that may cause MPEG-2 PES start code emulation. For example, a set of allowed values may comprise values from 0 to 187, inclusive. However, it is to be understood that embodiments may be applied to any byte stream format with a different set of allowed values for the first byte of a NUH.

[0237] Some embodiments propose reducing a bitrate overhead for the NUH for a future video standard while enabling support for multi-layer and temporal scalability use cases and avoiding MPEG-2 start code emulation.

[0238] The NUH according to present embodiments has a length of either one byte or two bytes. A single byte is used whenever possible, with the second byte only used when necessary to convey information that cannot be derived. Reducing the length to one byte for some NAL units and / or some use cases reduces bitrate. It is appreciated that the embodiments may apply similarly to NAL unit headers that include more than two bytes. For example, embodiments apply to a NUH design that may include only one byte as described in embodiments, or only two bytes as described in embodiments, but which may be further extended to have a third byte, for example. The presence of a third byte in a NUH may, for example, be indicated by a second extension flag, which may, for example, be present in the second byte of the NUH. In an example embodiment, the third byte of a NUH comprises syntax element(s) for parameter(s) that are not indicated by the first or second byte. For example, the third byte may comprise a subpicture identifier. In this case, the third byte may be present for VCL NAL units when multiple subpictures are in use, and needs not be present for NAL units that are independent of subpictures. In another example embodiment, the third byte of a NUH comprises an extension of a layer identifier. In this case, the third byte may be present for NAL units that have a layer identifier value greater than what can be indicated by the first and second bytes of the NUH.

[0239] Several options are presented in which the second byte is only used when necessary to convey information that cannot be derived.

[0240] Some of several options being presented provide limited support for multi-layer and / or temporal scalability while using one byte. However, additional layers and / or temporal sublayers can be supported with the second byte, whose presence is indicated through use of an extension flag or a “no extension” flag.

[0241] In an embodiment, an extension flag is present in the first byte of the NUH.

[0242] In an embodiment, an extension flag replaces the forbidden zero bit in the first byte of the NUH.

[0243] When an extension flag is equal to 0, the second byte of the NUH is absent. When an extension flag is equal to 1, the second byte of the NUH is present.

[0244] In an embodiment, a no extension flag is present in the first byte of the NUH.

[0245] In an embodiment, a no extension flag replaces the forbidden zero bit in the first byte of the NUH.

[0246] When a no extension flag is equal to 0, the second byte of the NUH is present. When a no extension flag is equal to 1, the second byte of the NUH is absent.

[0247] It is to be understood that many embodiments described with reference to an extension flag may be similarly realized by replacing the extension flag with a no extension flag. Likewise, it is to be understood that many embodiments described with reference to a no extension flag may be similarly realized by replacing the no extension flag with tan extension flag.

[0248] Some of the examples given below do not include the additional layers or temporal sublayers that would require the second byte, in which case the overhead for the NUH is cut in half.

[0249] Bitrate savings can still be achieved when the second byte is only required for some NAL units in the bitstream

[0250] In some options, a “no extension” flag is used instead of a forbidden zero flag to avoid MPEG-2 start code emulation.

[0251] Figure 3 illustrates a flowchart of a method for decoding according to an embodiment. The method comprises receiving 310 a bitstream comprising one or more units, each having a header; receiving a first byte of the header 320; determining 330, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream; deriving 340 at least one parameter depending on the determination, wherein the at least one parameter value comprises one or both of a scalability layer identifier and a temporalsublayer identifier; if the header does not comprise the second byte, deriving 350 the at least one parameter value based on at least one syntax element in the first byte and predefined information. If the header comprises the second byte, the method alternatively comprises deriving 360 the least one parameter value based on at least one syntax element in the first byte and information in the second byte.

[0252] Figure 4 illustrates a flowchart of a method for encoding according to an embodiment. The method comprises determining 410 at least one parameter value, wherein the at least one parameter value comprises one or more of the following: a unit type, a scalability layer identifier, a temporal sublayer identifier; deriving 420 a one-byte header for a unit comprised in a bitstream, wherein the at least one parameter value is represented by at least one syntax element and pre-defined information; if the one-byte header has an allowed value among one or more allowed values, determining 430 a header for the unit to be the one-byte header, wherein the header comprises the first byte; alternatively, if the one-byte header has a disallowed value among one or more disallowed values, deriving 440 a header comprising a first byte and a second byte, wherein the first byte is indicative that the second byte is present, and wherein a value of the at least one syntax element in the first byte and information on the second byte are set based on at least one parameter value; and including 450 the header (either one-byte or two-byte) in the bitstream.

[0253] In the following, several options are proposed for aNUH design for a new video codec, such as H.267, which use either one or two bytes while maintaining support for start code emulation avoidance and the byte stream format provided in HEVC and VVC. Using one byte instead of two bytes in the NUH saves bitrate. Several options are provided in which the second byte is only used when necessary to convey information that cannot be derived.

[0254] In the following options, a scalability layer identifier is referred to as Layer ID, and a temporal sublayer identifier is referred to as Temporal ID.

[0255] Option 1:

[0256] In an embodiment, an encoder indicates the temporal sublayer identifier using a nonreference flag in the first byte of the NUH.

[0257] In an embodiment, a decoder derives the temporal sublayer identifier based on a nonreference flag in the first byte of the NUH.

[0258] In an embodiment, the first byte of the NUH contains a forbidden zero bit, the NAL unit type, a first flag, such as a non-reference flag, and a second flag, such as an extension flag. When the extension flag is set, the second byte is present.

[0259] The non-reference flag is used to indicate that a picture is not used as a reference for prediction from any other coded picture in the CVS. If the non-reference flag is set, a MANE or decoder can discard the NAL unit without impacting the decoding of future pictures in decoding order. This allows temporal scalability with two temporal sublayers to be supported with a one-byte NUH.

[0260] If the bitstream contains more than one layer, or when the temporal scalability with more than two temporal sublayers is used, Temporal ID and Layer ID are included in the second byte. When the second byte is not present in the NUH, the values of Temporal ID and Layer ID are derived.

[0261] When the second byte is present, for signaling of Temporal ID, the temporal_id_plusl syntax element is signaled. Plus 1 may be used to avoid the value of 0, to avoid emulation prevention byte stuffing. In the present option, 2 bits are allocated to signalling of temporal_id_plusl in the second byte of the NUH, meaning that values of nuh_temporal_id_plusl in the range of 1 to 3 can be signaled. The nuh non reference flag can be used to provide another value of Temporal ID. The allowable range of Temporal ID for this option is 0 to 3. It is noted that this design supports fewer Temporal ID values than supported in VVC, which has an allowable range of 0 to 6, because 3 bits are allocated to temporal_id_plusl in VVC.

[0262] The design assumption is that large numbers of Temporal IDs are not frequently used in practice, and indicate of non-reference should be sufficient in most use cases, so that the single byte design can frequently be used.

[0263] It is to be noted that different orders of syntax elements can be used, including swapping the order of nal unit type and nuh non reference flag in the first byte and / or swapping the order of nuh_temporal_id_plusl and nuh layer id in the second byte.

[0264] The semantics differences from VVC for this embodiment are shown below:

[0265] forbidden_zero_bit is the same as defined in WC.

[0266] nuh extension flag equal to 1 specifies that the nuh_temporal_id_plusl and nuh layer id syntax elements are present, nuh extension flag equal to 0 specifies that the nuh_temporal_id_plusl and nuh layer id syntax elements are not present.

[0267] nal_unit_type is the same as defined in VVC.

[0268] nuh non reference flag equal to 1 specifies that the PU containing the NAL unit is not a reference picture, nuh non reference flag equal to 0 specifies that the PU containing the NAL unit may be a reference picture.

[0269] nuh_temporal_id_plusl minus 1 specifies a temporal identifier for the NAL unit.

[0270] The value of nuh_temporal_id_plusl shall not be equal to 0.

[0271] The variable of the temporal sublayer identifier, Temporalld, may be derived as follows: if (nuh non reference flag)Temporalld = 3 elseTemporalld = nuh_temporal_id_plusl - 1

[0272] The remaining semantics may be the same as defined in VVC.

[0273] The scalability layer identifier, nuh_layer_id, specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of a layer to which a non-VCL NAL unit applies.

[0274] When not present, the value of nuh layer id is inferred to be equal to 0.

[0275] The remaining semantics may be the same as defined in VVC.

[0276] Option 2:

[0277] According to an option 2, during the encoding phase, when the header does not comprise the second byte, the at least one parameter value may be indicated using the first flag, such as non reference flag, in the first byte of the header and the extension flag. When the header comprises the second byte, the temporal sublayer identifier is indicated using a temporal indicator in the second byte and a first flag in the first byte.

[0278] Similarly, during the decoding phase, when the header does not comprise the second byte, it may be determined if the first flag, such as non reference flag, is enabled, wherein based on the determination that the first flag is enabled, it may be determined that the temporal sublayer identifier is equal to a first value, e.g., a maximum value, and based on the determination that the first flag is not enabled, it may be determined that the temporal sublayer identifier is equal to a second value, e.g., a minimum value; and wherein when the header comprises the second byte, the temporal sublayer identifier may be determined based on the temporal indicator in the second byte and the first flag in the first byte.

[0279] In an embodiment, an encoder indicates the temporal sublayer identifier using a first flag in the first byte of the NUH and the extension flag. The first flag may be a nonreference flag. The sublayer identifier for a non-reference picture may be indicated using the non-reference flag equal to 1 and the extension flag indicating the second byte of the NUH to be absent. The sublayer identifier for a reference picture at the lowest temporal sublayer may be indicated using the non-reference flag equal to 0 and the extension flag indicating the second byte of the NUH to be absent.

[0280] In an embodiment, a decoder derives the temporal sublayer identifier based on a first flag in the first byte of the NUH and the extension flag. The first flag may be a non- reference flag. When the extension flag indicates the second byte of the NUH to be absent, the sublayer identifier may be derived to be equal to a first pre-defined value, when the first flag has a particular value, and equal to a second pre-defined value, when the first flag has another particular value.

[0281] In an embodiment, when the NUH comprises the second byte, an encoder indicates the temporal sublayer identifier using a temporal indicator in the second byte and the first flag in the first byte.

[0282] In an embodiment, when the NUH comprises the second byte, a decoder derives the temporal sublayer identifier based on a temporal indicator in the second byte and the first flag in the first byte.

[0283] Mapping the temporal sublayer identifier to a combination the first flag in the first byte and the temporal indicator in the second byte may include, but may not be limited to, any of the following:The temporal sublayer identifier is based on the sum of the temporal indicator and the first flag.- When the first flag is equal to 0, the temporal indicator values are mapped to a first range of temporal sublayer identifier values. When the first flag is equal to 1, the temporal indicator values are mapped to a second range of temporal sublayer identifier values, which is non-overlapping with the first range. For example, when the first flag is a non-reference flag and the temporal indicator indicates a value range from 0 to 2, inclusive, the temporal sublayer identifier value is equal to value indicated by the temporal indicator value (when the non-reference flag is equal to 0) or 3 + the value indicated by the temporal indicator value (when the non-reference flag is equal to 1). The first flag forms the least significant bit of the temporal sublayer identifier, while the other bits of the temporal sublayer identifier are represented by the temporal indicator.

[0284] The second option is similar to Optionl, but enables to have one more allowable value in the range of Temporal ID for some bitstreams, e.g., for single-layer bitstreams with Layer ID of 0. In those bitstreams, the second byte of the NUH is not required to signal the Layer ID value, which can be derived to be zero, so the presence of the second byte can be used in combination with the nuh non reference flag in the derivation the value of Temporal ID.

[0285] In an example, the allowable range of Temporal ID of this option is 0 to 5.

[0286] In this option, nuh_temporal_id_plusl may be renamed as nuh temporal idc because its interpretation differs.

[0287] In an example embodiment, the following syntax and semantics may be used.

[0288] forbidden_zero_bit is the same as defined in WC.

[0289] nuh extension flag equal to 1 specifies that the nuh temporal idc and nuh layer id syntax elements are present, nuh extension flag equal to 0 specifies that the nuh temporal idc and nuh layer id syntax elements are not present.

[0290] nal_unit_type is the same as defined in VVC.

[0291] nuh non reference flag equal to 1 specifies that the PU containing the NAL unit is not a reference picture, nuh non reference flag equal to 0 specifies that the PU containing the NAL unit may be a reference picture.

[0292] nuh_temporal_idc specifies a temporal indicator for the NAL unit. The value of nuh temporal idc shall not be equal to 0.

[0293] In one embodiment of the option 2, the variable of the temporal sublayer identifier, Temporalld is derived as follows: if (Inuh extension flag)Temporalld = nuh non reference flag ? 4 : 0 elseTemporalld = nuh temporal idc - 1 + nuh non reference flag

[0294] In another embodiment of the option 2, Temporalld values 0 to 2, inclusive, indicate a reference picture and Temporalld values 3 to 5, inclusive, indicate a non-reference picture, and Temporalld is derived as follows: if (!nuh_extension_flag)Temporalld = nuh non reference flag ? 5 : 0 elseTemporalld = nuh non reference flag ?( nuh temporal idc - 1 ) : 2 + nuh temporal idc

[0295] In yet another embodiment of the option 2, the semantics of nuh non reference flag are described as above only when nuh extension flag is equal to 0. When nuh extension flag is equal to 1, nuh non reference flag are used in deriving the Temporalld value, for example as follows, and does not necessarily indicate whether the PU is a reference picture or a non-reference picture. if (!nuh_extension_flag)Temporalld = nuh non reference flag ? 5 : 0 elseTemporalld =( ( nuh temporal idc - 1 ) « 1 )+nuh_non_reference_flag

[0296] The remaining semantics may be the same as VVC semantics for temporal_id_plusl

[0297] The scalability layer identifier, nuh layer id specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of a layer to which a non-VCL NAL unit applies. When not present, the value of nuh layer id is inferred to be equal to 0. The remaining semantics may be the same as VVC.

[0298] Option 3:

[0299] According to option 3, during the encoding phase, when the header does not comprise the second byte, a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte is set based on the at least one parameter value. Alternatively, if the header comprises the second byte, a value of a parameter specifying one or more most significant bits of the temporal sublayer identifier in the second byte and a value of a parameter specifying the one or more least significant bits of the temporal sublayer identifier in the first byte are set based on the at least one parameter value.

[0300] During the decoding phase, when the header does not comprise the second byte, the temporal sublayer identifier is derived based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte. Alternatively, when the header comprises the second byte, the temporal sublayer identifier is determined based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on the parameter specifying the one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0301] In an embodiment, an encoder sets a no extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 1 and sets a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte based on the determined temporal sublayer identifier.

[0302] In an embodiment, a decoder decodes a no extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the no extension flag is equal to 1, the decoder derives the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0303] In an embodiment, an encoder sets a no extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 0 and sets a value of a parameter specifying one or more most significant bits of the temporal sublayer identifier in the second byte and a valueof a parameter specifying the one or more least significant bits of the temporal sublayer identifier in the first byte based on the determined temporal sublayer identifier.

[0304] In an embodiment, a decoder decodes a no extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the no extension flag is equal to 0, the decoder derives the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on a parameter specifying the one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0305] In the third option, the forbidden zero bit, i.e., the first bit, is replaced with a third flag, such as a no extension flag, to indicate that the NUH extension is not present. The flag’s value can be set to 1 when it is needed to avoid start code emulation, even when the second byte is not required for signaling the syntax elements in the second byte.

[0306] Thus, nuh no extension flag can be used to avoid start code emulation, in place of the forbidden zero bit. nuh no extension flag can be set to 0 if the contents of the first byte would other emulate a start code, even when second byte wouldn’t otherwise be needed for scalability information.

[0307] In an embodiment of this option, signaling of the Temporal ID may be split between two least significant bits (LSB) in the first byte and two most significant bits (MSB) in the second byte. Putting the LSBs in the first byte allows for temporal scalability with 4 temporal sublevels without requiring a second byte. A plus 1 is used for signaling the MSBs in the second byte, to avoid the need for emulation prevention byte stuffing. The range of Temporal ID values supported is 0 to 11.

[0308] The nuh no extension flag bit being set to 0 indicates that the second byte is present, when needed to avoid start code emulation values in first byte for specific values of NUT and / or when needed for scalability.

[0309] In an example embodiment, the following syntax and semantics may be used.

[0310] Start code emulation may occur when the value of the byte is in the range of 188 to 255. Therefore, in the example syntax, nuh no extension flag bit is required to be set to 0 when the nal unit type value is 14 or higher, which in combination with the bits allocated for nuh temporal id lsb, could allow a first byte value of 188 or higher.

[0311] nuh no extension flag equal to 0 specifies that the nuh_temporal_id_msb_plusl and nuh layer id syntax elements are present, nuh no extension flag equal to 1 specifies thatthe nuh_temporal_id_msb_plusl and nuh layer id syntax elements are not present.

[0312] nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit. The NAL unit type values may be ordered so that the most common values are in the range of 0 to 13, so that it is less likely that the second byte will be present. When nuh no extension flag equal to 1, the value of nal unit type shall be in the range of 0 to 13. It is to be noted that to prevent the possible emulation of start codes when nal unit type is equal or greater than 14, nuh no extension flag may be set to 0.

[0313] nuh_temporal_id_lsb specifies the least significant bits of the temporal ID.

[0314] nuh temporal id msb plusl minus 1 specifies the most significant bits of the temporal ID.

[0315] The value of nuh_temporal_id_msb_plusl shall not be equal to 0.

[0316] The variable of the temporal sublayer identifier, Temporalld is derived as follows: if (nuh no extension flag)Temporalld = nuh temporal id lsb elseTemporalld = (nuh_temporal_id_msb_plusl - 1) « 2 + nuh temporal id lsb

[0317] The remaining semantics may be the same as VVC semantics for nuh_temporal_id_plusl .

[0318] The scalability layer identifier nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of a layer to which a non-VCL NAL unit applies. When not present, the value of nuh layer id is inferred to be equal to 0. The remaining semantics may be the same as VVC.

[0319] Possible alternatives are to reorder the syntax elements in the first byte and / or reorder the syntax elements in the second byte.

[0320] The + operator in the Temporal Id calculation can be replaced with a bitwise or operator |Temporalld = (nuh_temporal_id_msb_plusl - 1) « 2 | nuh temporal id lsb

[0321] Option 4

[0322] The option 4 differs from option 3 in such a manner that in addition to the option 3, in the encoding phase according to option 4 when the header comprises the second byte, the scalability layer identifier is indicated in the second byte. When the header does not comprise the second byte, a value of a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte is set based on the at least one parameter value. On the other hand, when the header comprises the second byte, a value of a parameter specifying one or more most significant bits of the temporal identifier in the second byte and a value of a parameter specifying the one or more least significant bits of the temporal sublayer identifier in the first byte are set based on the at least one parameter value.

[0323] In decoding phase, similarly when the header comprises the second byte, the scalability layer identifier is determined from the second byte and the temporal sublayer identifier is determined based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on the parameter specifying the one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0324] Option 4 is similar to option 3 but has a different parameter in the second byte for representing one or more most significant bits of the temporal sublayer identifier.

[0325] Another option is to have nuh temporal id msb and nuh_layer_id_plusl (and a reserved bit) in the second byte.

[0326] In the fourth option, as in option 3, the third flag, such as nuh no extension flag, can be used to avoid start code emulation. This option is similar to option 3 with the following changes: reduces the temporal ID MSB in the second byte to a single bit and remove the plusl, include a plusl for the layer ID, and add a reserved bit.

[0327] In an example embodiment, the following syntax and semantics may be used.

[0328] The example syntax of the fourth option supports Temporal ID in the range of 0 to 7. The example syntax of the fourth option also supports Layer ID in the range of 0 to 62, which is less than the 0 to 63 supported in VVC. Note that in VVC, values of nuh_layer_id above 55 are reserved for future use.

[0329] nuh no extension flag equal to 0 specifies that the nuh temporal id msb, nuh_layer_id_plusl, and nuh reserved zero bit syntax elements are present, nuh no extension flag equal to 1 specifies that the nuh temporal id msb, nuh_layer_id_plusl, and nuh reserved zero bit syntax elements are not present.

[0330] nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit. When nuh no extension flag is equal to 1, the value of nal unit type shall be in the range of 0 to 13. It is to be noted that to prevent the possible emulation of start codes when nal unit type is equal or greater than 14, nuh no extension flag may be set to 0.

[0331] nuh_temporal_id_lsb specifies the least significant bits of the temporal ID.

[0332] nuh_temporal_id_msb specifies the most significant bits of the temporal ID.

[0333] The variable of the temporal sublayer identifier, Temporalld may be derived as follows: if (nuh no extension flag)Temporalld = nuh temporal id lsb elseTemporalld = nuh temporal id msb « 2 + nuh temporal id lsb

[0334] The remaining semantics may be the same as VVC semantics for nuh_temporal_id_plusl .

[0335] nuh_layer_id_plusl minus 1 specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of a layer to which a non-VCL NAL unit applies.

[0336] The value of nuh_layer_id_plusl shall not be equal to 0.

[0337] When not present, the value of nuh layer id is inferred to be equal to 1.

[0338] The variable for the scalability layer identifier, Layerld, is derived as follows:

[0339] Layerld = nuh_layer_id_msb_plusl - 1

[0340] The remaining semantics may be the same as VVC for nuh layer id, but replacing nuh layer id in further semantics by Layerld.

[0341] nuh_reserved_zero_bit is the same as in VVC.

[0342] Option 5:

[0343] According to option 5, during the encoding phase, when the header does not comprise the second byte, a value of a parameter specifying one or more least significant bits for the scalability layer identifier in the first byte is set based on at least one parameter value, when the header comprises the second byte, a value of a parameter specifying one or more least significant bits for the scalability layer identifier in the first byte and a value of a parameter specifying one or more most significant bits for the scalability layer identifier in the second byte are set based on at least one parameter value.

[0344] In decoding when the header does not comprise the second byte, the scalability layer identifier is determined based on a parameter specifying one or more least significant bits defined in the first byte. Alternatively, when the header comprises the second byte, the scalability layer identifier is determined based on a parameter specifying one or more least significant bits for the layer identifier defined in the first byte and a parameter specifying one or more most significant bits for the layer identifier defined in the second byte.

[0345] In an embodiment, an encoder sets a no extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 1 and sets a value of a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte based on the determined scalability layer identifier.

[0346] In an embodiment, a decoder decodes a no extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the no extension flag is equal to 1, the decoder derives the scalability layer identifier based on a parameter specifying one or more least significant bits of the scalability layer identifier defined in the first byte.

[0347] In an embodiment, an encoder sets a no extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 0 and sets a value of a parameter specifying one or more least significant bits for the scalability layer identifier in the first byte and a value of a parameter specifying one or more most significant bits for the scalability layer identifier in the second byte based on the determined scalability layer identifier.

[0348] In an embodiment, a decoder decodes a no extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the no extension flag is equal to 0, the decoder derives the scalability layer identifier based on a parameter specifying one or more least significant bits for the scalability layer identifier defined in the first byte and a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte.

[0349] In the fifth option, as in options 3 and 4, a third flag, such as nuh no extension flag, can be used to avoid start code emulation

[0350] In an example embodiment, the following syntax and semantics may be used.

[0351] The example syntax of the fifth option includes two least significant bits for Layer ID in the first byte. This allows up to 4 layers to be included in the bitstream without requiring the presence of the second byte.

[0352] In the example syntax, four most significant bits of Layer ID may be included in the second byte. The allowable range of Layer ID values is 0 to 63, the same as supported in VVC, although VVC reserves all values greater than 55.

[0353] In the example syntax, Temporal ID is also in the second byte with a plus 1. The allowable range of Temporal ID of this option is 0 to 6, which is the same as supported in VVC.

[0354] The second byte also contains a reserved bit.

[0355] nuh no extension flag equal to 0 specifies that the nuh layer id msb, nuh_temporal_id_plusl, and nuh reserved zero bit are present, nuh no extension flag equal to 1 specifies that the nuh layer id msb, nuh_temporal_id_plusl, and nuh reserved zero bit syntax elements are not present.

[0356] nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit. When nuh no extension flag equal to 1, the value of nal unit type shall be in the range of 0 to 13. It is to be noted that to prevent the possible emulation of start codes when nal unit type is equal or greater than 14, nuh no extension flag may be set to 0.

[0357] nuh_layer_id_lsb specifies the least significant bits of the layer ID.

[0358] nuh_layer_id_msb specifies the most significant bits of the layer ID. When not present, the value of nuh layer id msb is inferred to be equal to 0.

[0359] The variable for the scalability layer identifier, Layerlld, is derived as follows:

[0360] Layerld = nuh layer id msb « 2 + nuh layer id lsb

[0361] The remaining semantics may be the same as VVC semantics for nuh_layer_id_plusl

[0362] nuh_temporal_id_plusl minus 1 specifies a temporal identifier for the NAL unit.The value of nuh_temporal_id_plusl shall not be equal to 0.

[0363] The variable for the temporal sublayer identifier, Temporalld, is derived as follows: Temporalld = nuh no reference flag ? 3 : nuh_temporal_id_plusl - 1

[0364] nuh_reserved_zero_bit is the same as defined in VVC.

[0365] Option 6:

[0366] According to option 6, during encoding phase, when the header does not comprise the second byte, a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte is set based on the at least one parameter value, and a value of a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte is set based on the at least one parameter value. When the header comprises the second byte, a value of a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte is set based on the at least one parameter value, and a value of a parameter specifying one or more most significant bits for the temporal sublayer identifier in the second byte is set based on the at least one parameter value.

[0367] During decoding phase, when the header does not comprise the second byte, the temporal sublayer identifier is derived based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte and the scalability layer identifier is derived based on a parameter specifying one or more least significant bits of the layer identifier defined in the first byte. When the header comprises the second byte, the scalability layer identifier is determined based on a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte and a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte and the temporal sublayer identifier is determined based on a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte and a parameter specifying one or more most significant bits for the temporal sublayer identifier defined in the second byte.

[0368] In an embodiment, an encoder sets a no extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 1 and sets a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte based on the determined temporal sublayer identifier and sets a value of a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte based on the determined scalability layer identifier.

[0369] In an embodiment, a decoder decodes a no extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the no extension flag is equal to 1, the decoder derives the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte and derives the scalability layer identifier based on a parameter specifying one or more least significant bits of the scalability layer identifier defined in the first byte.

[0370] In an embodiment, an encoder sets a no extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 0; sets a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte based on the determined temporal sublayer identifier;sets a value of a parameter specifying one or more most significant bits for the temporal sublayer identifier in the second byte based on the determined temporal sublayer identifier; sets a value of a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte based on the determined scalability layer identifier; and sets a value of a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte based on the determined scalability layer identifier.

[0371] In an embodiment, a decoder decodes a no extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the no extension flag is equal to 1, the decoder derives the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte and a parameter specifying one or more most significant bits for the temporal sublayer identifier defined in the second byte; and derives the scalability layer identifier based on a parameter specifying one or more least significant bits of the scalability layer identifier defined in the first byte and a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte.

[0372] In an example embodiment, the following syntax and semantics may be used.

[0373] The example syntax of the sixth option is similar to option 5, but includes a single least significant bit for the scalability layer identifier, Layer ID, and a single least significant bit for the temporal sublayer identifier, Temporal ID, in the first byte.

[0374] A third flag, such as nuh no extension flag, can be used to avoid start code emulation. This option includes in the first byte a first flag, such as a non-reference flag, and a single bit for the layer ID LSB. This allows a single byte NUH to support simple temporal scalability with two temporal sublayers and two layers for multi-layer bitstreams.

[0375] Five most significant bits of Layer ID are included in the second byte. The allowable range of Layer ID values is 0 to 63, the same as supported in VVC, although WC reserves all values greater than 55.

[0376] Two most significant bits of Temporal ID are also in the second byte, including a plus 1. The allowable range of Temporal ID of this option is 0 to 6, which is the same as supported in VVC.

[0377] The second byte also contains a reserved bit.

[0378] nuh no extension flag equal to 0 specifies that the nuh_temporal_id_msb_plusl, nuh layer id msb, and nuh reserved zero bit syntax elements are present, nuh no extension flag equal to 1 specifies that the nuh_temporal_id_msb_plusl, nuh layer id msb, and nuh reserved zero bit syntax elements are not present.

[0379] nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit. When nuh no extension flag equal to 1, the value of nal unit type shall be in the range of 0 to 13. It is to be noted that to prevent the possible emulation of start codes when nal unit type is equal or greater than 14, nuh no extension flag may be set to 0.

[0380] nuh_temporal_id_lsb specifies the least significant bits of the temporal ID.

[0381] nuh_temporal_id_msb_plusl minus 1 specifies the most significant bits of the temporal ID. The value of nuh_temporal_id_msb_plusl shall not be equal to 0.

[0382] The variable for the temporal sublayer identifier, Temporalld, is derived as follows: if (nuh no extension flag)Temporalld = nuh temporal id lsb elseTemporalld = (nuh_temporal_id_msb_plusl - 1) « 1 + nuh temporal id lsb

[0383] The remaining semantics may be the same as VVC semantics for nuh_temporal_id_plusl .

[0384] nuh layer id msb specifies the most signficant bits of the layer ID.

[0385] When not present, the value of nuh layer id msb is inferred to be equal to 0.

[0386] The variable for the scalability layer identifier, Layerlld, is derived as follows: Layerld = nuh layer id msb « 1 + nuh layer id lsb

[0387] The remaining semantics may be the same as VVC for nuh layer id, but replacing nuh layer id in further semantics by Layerld.

[0388] nuh_reserved_zero_bit is same as in VVC.

[0389] Option 7:

[0390] According to option 7, during encoding phase, when the header does not comprise the second byte, a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte is set based on the at least one parameter value. When the header comprises the second byte, a value of a parameter specifying one or more most significant bits of the temporal sublayer identifier in the second byte and a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte are set based on the at least one parameter value.

[0391] During decoding phase, when the header does not comprise the second byte, the temporal sublayer identifier is derived based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte. When the header comprises the second byte, the temporal sublayer identifier is derived based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on the parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0392] In an embodiment, an encoder sets an extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 0 and sets a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte based on the determined temporal sublayer identifier.

[0393] In an embodiment, a decoder decodes an extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the extension flag is equal to 0, the decoder derives the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0394] In an embodiment, an encoder sets an extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 1 and sets a value of a parameter specifying one or more most significant bits of the temporal sublayer identifier in the second byte and a value of a parameter specifying the one or more least significant bits of the temporal sublayer identifier in the first byte based on the determined temporal sublayer identifier.

[0395] In an embodiment, a decoder decodes an extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the extension flag is equal to 1, the decoder derives the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on a parameter specifying the one or more least significant bits of the temporal sublayer identifier defined in the first byte.

[0396] The seventh includes in the first byte of the NUH the following: a second flag, such as an extension flag (!=No extension flag), certain number, such as two, least significant bits of the temporal ID, and the NAL unit type.

[0397] In an example embodiment, the following syntax and semantics may be used.

[0398] In the example syntax, if the NUH contains the second byte, the second byte contains two most significant bits for the temporal sublayer identifier with a plusl for byte stuffing avoidance. The second byte also contains scalability layer identifier.

[0399] In the example syntax of this option, signaling of the temporal sublayer identifier, Temporal ID, may be split between two least significant bits in the first byte and two significant bits in the second byte. Putting the LSBs in the first byte allows for temporal scalability with four temporal sublevels without requiring a second byte. A plus 1 is used for signaling the MSBs in the second byte, to avoid the need for emulation prevention byte stuffing. The range of Temporal ID values supported is 0 to 11.

[0400] nuh extension flag equal to 1 specifies that the nuh_temporal_id_msb_plusl and nuh layer id syntax elements are present, nuh extension flag equal to 0 specifies that the nuh_temporal_id_msb_plusl and nuh layer id syntax elements are not present.

[0401] nuh_temporal_id_lsb specifies the least significant bits of the temporal ID.

[0402] When nuh extension flag is equal to 1, nuh temporal id lsb shall be in the range of 0 to 1, inclusive.

[0403] nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit.

[0404] When nuh extension flag is equal to 1 and nuh temporal id lsb is equal to 1, nal unit type shall be in the range of 0 to 27, inclusive.

[0405] It is to be noted that in order to avoid MPEG-2 PES start code emulation, the NAL unit type values are reordered such that values in the range of 28 to 31 require that temporal ID be equal to 0, such as those associated with IRAP pictures, video parameter sets and sequence parameter sets. Therefore, when nuh extension flag is equal to 1 and nal_unit_type is in the range of 28 to 31, (i.e., l l lxx), the value of the first byte is 1001 l lxx, i.e., less than 188 and hence does not cause a MPEG-2 PES start code emulation problem.

[0406] nuh_temporal_id_msb_plusl minus 1 specifies the most significant bits of the temporal sublayer ID.

[0407] The value of nuh_temporal_id_msb_plusl shall not be equal to 0.

[0408] The variable for the temporal sublayer identifier, Temporalld, is derived as follows: if (!nuh_extension_flag)Temporalld = nuh temporal id lsb elseTemporalld = (nuh_temporal_id_msb_plusl - 1) « 2 + nuh temporal i d_l sb .

[0409] The remaining semantics may be the same as VVC semantics for nuh_temporal_id_plusl .

[0410] nuh_layer_id specifies the scalability layer identifier. When not present, the value of nuh layer id is inferred to be equal to 0. The remaining semantics may be the same as VVC.

[0411] Option 8:

[0412] According to option 8, during the encoding phase, when the header does not comprise the second byte, a value of a temporal sublayer identifier in the first byte is set based on at least one parameter value and a value of a unit type indicator in the first byte is set based on the at least one parameter value. When the header comprises the second byte, a value of a parameter specifying one or more most significant bits of the temporal identifier indicator in the second byte and a value of the temporal identifier indicator in the first byte are set based on the at least one parameter value.

[0413] During the decoding phase, when the header does not comprise the second byte, the temporal sublayer identifier is determined based on the temporal identifier indicator in the first byte to determine a unit type based on the temporal identifier indicator and a unit type indicator in the first byte. When the header comprises the second byte, the temporal sublayer identifier is determined by a parameter specifying one or more most significant bits of the temporal identifier indicator in the second byte and the temporal identifier indicator in the first byte.

[0414] In an embodiment, an encoder sets a no extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 1 and sets a value of a temporal identifier indicator in the first byte based on the determined temporal sublayer identifier and sets a value of a NAL unit type indicator in the first byte based on the determined NAL unit type.

[0415] In an embodiment, a decoder decodes a no extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the no extension flag is equal to 1, the decoder derives the temporal sublayer identifier based on a temporal identifier indicator in the first byte and derives the NAL unit type based on the temporal identifier indicator and a NAL unit type indicator in the first byte.

[0416] In an embodiment, an encoder sets a no extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 0; sets a value of a temporal identifier indicator in the first byte based on the determined temporal sublayer identifier; sets a value of a parameter specifying one or more most significant bits of the temporal identifier indicator in the second byte based on the determined temporal sublayer identifier; andsets a value of a NAL unit type indicator in the first byte based on the determined NAL unit type.

[0417] In an embodiment, a decoder decodes a no extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the no extension flag is equal to 0, the decoder derives the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal identifier indicator in the second byte and a temporal identifier indicator in the first byte and derives a NAL unit type based on the temporal identifier indicator and a NAL unit type indicator in the first byte.

[0418] In an example embodiment, the following syntax and semantics may be used.

[0419] The example syntax of the eighth option includes in the first byte of the NUH the following: a third flag, such as nuh no extension flag, 3-bit temporal ID indicator (nuh tid idc), and 4-bit NAL unit type indicator (nal unit type idc). Value 0 of nuh tid idc is used to indicate a category of NAL unit types that imply temporal ID equal to 0.

[0420] nuh no extension flag equal to 0 specifies that the nuh_tid_plusl_msb, nuh layer id and nuh reserved zero bit syntax elements are present, nuh extension flag equal to 1 specifies that the nuh tid msb, nuh layer id and nuh reserved zero bit syntax elements are not present.

[0421] nuh_tid_idc affects the derivation of Temporalld and NalUnitType. When nuh no extension flag is equal to 1, nuh tid idc shall be in the range of 0 to 3, inclusive.

[0422] nal_unit_type_idc affects the derivation of NalUnitType. When nuh no extension flag is equal to 1 and nuh tid idc is equal to 3, the value of nal unit type idc shall be in the range of 0 to 9, inclusive.

[0423] NalUnitType is set equal to ( nuh_tid_idc > 0 ) ? 16 : 0 ) + nal_unit_type_idc.

[0424] An example embodiment for the mapping of NalUnitType values to their names and NAL unit type classes is specified in the table below.

[0425] It is a requirement of bitstream conformance that when the value of NalUnitType is in the range of 0 to 3, inclusive, nuh no extension flag shall be equal to 1.

[0426] nuh_tid_plusl_msb affects the derivation of temporal sublayer identifier, Temporalld.

[0427] In an embodiment, the variable for the temporal sublayer identifier, Temporalld is derived as follows: if ( nuh tid idc = = 0 ) Temporalld = 0 else if( nuh no extension flag )Temporalld = nuh tid idc - 1 elseTemporalld = ( ( nuh_tid_plusl_msb « 2 ) + nuh tid idc ) - 1

[0428] In another embodiment, nuh_tid_plus_msb may be renamed as nuh tid lsb and the variable for the temporal sublayer identifier, Temporalld is derived as follows: if ( nuh tid idc = = 0 )Temporalld = 0 else if( nuh no extension flag )Temporalld = nuh tid idc - 1 elseTemporalld = ( nuh tid idc « 1 ) + nuh tid lsb

[0429] nuh_reserved_zero_bit has the same semantics as nuh reserved zero bit in VVC.

[0430] nuh_layer_id specifies the scalability layer identifier.

[0431] When not present, the value of nuh layer id is inferred to be equal to 0.

[0432] nuh_reserved_one_bit shall be equal to 1. The value 0 of nuh reserved one bit could be specified in the future by ITU-T | ISO / IEC. Although the value of nuh reserved one bit is required to be equal to 1 in this version of this Specification,decoders conforming to this version of this Specification shall also allow the value of nuh reserved one bit equal to 0 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh reserved one bit equal to 0.

[0433] Option 9

[0434] According to an embodiment of option 9, during the encoding phase, when the header does not comprise the second byte, a value of a parameter specifying one more least significant bits of at least one parameter value in the first byte is set based on the determined at least one parameter value. When the header comprises the second byte, a value of the at least one parameter in the second byte is set based on the determined at least one parameter.

[0435] According to an embodiment, during the decoding phase, when the header does not comprise the second byte, at least one parameter is determined based on a parameter specifying one or more least significant bits of the at least one parameter in the first byte. When the header comprises the second byte, the at least one parameter is determined by a parameter specifying the at least one parameter in the second byte.

[0436] According to an embodiment of option 9, during the encoding phase, when the header does not comprise the second byte, a value of a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte is set based on the determined temporal sublayer identifier. When the header comprises the second byte, a value of a parameter specifying temporal sublayer identifier in the second byte is set based on the determined temporal sublayer identifier.

[0437] According to an embodiment, during the decoding phase, when the header does not comprise the second byte, the temporal sublayer identifier is determined based on a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte. When the header comprises the second byte, the temporal sublayer identifier is determined by a parameter specifying a temporal sublayer identifier in the second byte.

[0438] In an example embodiment, syntax and semantics of the OBU header may be specified as follows.

[0439] obu extension flag equal to 1 specifies that the OBU header contains obu_extension_header(). obu_ extension flag equal to 0 specifies that the OBU header does not contain obu_extension_header().

[0440] obu reserved zero idc, when present, shall be equal to 0.

[0441] Consequently, when obu_extension_flag is equal to 1, the first byte will be lOOxxxxxb, with each x equal to 0 or 1, and hence the first byte will have a value from 128 to 159, which is within the set of allowed values that avoids MPEG-2 PES start code emulation. When obu extension flag is equal to 0, the first byte will have a value from 0 to 127, which is within the set of allowed values that avoids MPEG-2 PES start code emulation.

[0442] temporal id lsb specifies that temporal id is inferred to be equal to temporal id lsb.

[0443] The remaining semantics may be the same as in AVI except that the temporal id, when not present, is inferred to be equal to temporal id lsb.

[0444] Option 10:

[0445] According to option 10, the order and / or the bit position of a unit type indicator and / or a temporal identifier indicator are determined based on whether the header comprises the second byte. The temporal sublayer identifier is determined based on the temporal identifier indicator and a unit type is determined based on the temporal identifier indicator and a unit type indicator.

[0446] Figure 7 illustrates an embodiment of a decoding method. According to such embodiment, the decoding comprises:- receiving 710 a bitstream comprising one or more units each having a header;- receiving 720 one or more syntax elements from a first byte of the header;- determining 730, based on the one or more syntax elements, if the header comprises a second byte that follows the first byte in the bitstream;- in response to the header not comprising the second byte, determining 740 a first order and first bit positions of remaining syntax elements in the header;- in response to the header comprising the second byte, determining 750 a second order and second bit positions of remaining syntax elements in the header, wherein the first order differs from the second order and / or the first bit positions differ at least partly from the second bit positions;- receiving 760 the remaining syntax elements depending on the determination of the first order and the first bit positions, or the second order and the second bit positions; and- deriving 770 at least one parameter value from the remaining syntax elements.

[0447] According to an embodiment, receiving one or more syntax elements from a first byte of the header comprises receiving an extension flag from a first byte of the header. The extension flag may be the first bit (i.e., the most significant bit) of the first byte of the header.

[0448] According to an embodiment, when the extension flag is equal to 0, the header does not comprise a second byte, and when the extension flag is equal to 1, the header comprises a second byte.

[0449] According to an embodiment, when the extension flag is equal to 1, decoding comprises reading a reserved zero bit from the header. The reserved zero bit may be the second bit (i.e., the second most significant bit) of the first byte. In an embodiment, if the reserved zero bit is equal to 1, a unit is determined to be erroneous. Decoding may comprise means to deal with erroneous units, such as skipping an erroneous unit and / or requesting retransmission of an erroneous unit.

[0450] According to an embodiment, the remaining syntax elements comprise either or both of a unit type indicator and / or a temporal identifier indicator.

[0451] According to an embodiment, the bit position(s) of either or both of the unit type indicator and / or the temporal identifier indicator depend on whether the header comprises a second byte.

[0452] According to an embodiment, the order of unit type indicator and a temporal identifier indicator depends on whether the header comprises a second byte.

[0453] According to an embodiment, the at least one parameter value comprises a unit type (e.g., a NAL unit type or NalUnitType variable) and / or a temporal sublayer identifier (e.g., Temporalld variable).

[0454] Figure 8 illustrates an embodiment of an encoding method. According to an embodiment, the encoding comprises:- encoding 810 a bitstream comprising one or more units each having a header;- determining 820 if a second byte of a header is needed for a unit;- setting 830 values of one or more syntax elements to indicate if the header comprises the second byte that follows the first byte in the bitstream;- encoding 840 the one or more syntax elements to a first byte of the header;- in response to the header not comprising the second byte, determining 850 a first order and first bit positions of remaining syntax elements in the header;- in response to the header comprising the second byte, determining 860 a second order and second bit positions of remaining syntax elements in the header, wherein the first order differs from the second order or the first bit positions differ at least partly from the second bit positions;- setting 870 values of remaining syntax elements based on at least parameter value associated with the unit; and- encoding 880 the remaining syntax elements to the header depending on the determination of the first order and the first bit positions, or the second order and the second bit positions.

[0455] According to an embodiment, "setting values of one or more syntax elements to indicate if the header comprises the second byte that follows the first byte in the bitstream" comprises setting an extension flag of a first byte of the header to a first value (e.g., 0) to indicate that the header does not comprise the second byte or to a second value (e.g., 1) to indicate that the header comprises the second byte. The extension flag may be the first bit (i.e., the most significant bit) of the first byte of the header.

[0456] According to an embodiment, the remaining syntax elements comprise either or both of a unit type indicator and / or a temporal identifier indicator.

[0457] According to an embodiment, the bit position(s) of either or both of the unit type indicator and / or the temporal identifier indicator depend on whether the header comprises a second byte.

[0458] According to an embodiment, the order of unit type indicator and a temporal identifier indicator depends on whether the header comprises a second byte.

[0459] According to an embodiment, the at least one parameter value comprises a unit type (e.g., a NAL unit type or NalUnitType variable) and / or a temporal sublayer identifier (e.g., Temporalld variable).

[0460] In an embodiment, an encoder signals an extension flag in the first byte of the NUH, and, when the extension flag is equal to 1, additionally signals in the first byte a zero bit, and when the extension flag is equal to 0, not signaling a zero bit in the first byte.

[0461] In an embodiment, a decoder decodes an extension flag in the first byte of the NUH, and, when the extension flag is equal to 1, additionally receives in the first byte a zero bit, and when the extension flag is equal to 0, does not receive a zero bit in the first byte.

[0462] In an embodiment, an encoder sets a value of a temporal identifier indicator based on the determined temporal sublayer identifier and a value of a NAL unit type indicator in based on the determined NAL unit type.

[0463] In an embodiment, an encoder sets an extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 0 and sets a value of a temporal identifier indicator in the first byte based on the determined temporal sublayer identifier and sets a value of a NAL unit type indicator in the first byte based on the determined NAL unit type.

[0464] In an embodiment, a decoder decodes an extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the extension flag is equal to 0, the decoder derives the temporal sublayer identifier based on a temporal identifier indicator in the first byte and derives the NAL unit type based on the temporal identifier indicator and a NAL unit type indicator in the first byte.

[0465] In an embodiment, an encoder sets an extension flag that replaces the forbidden zero bit in the first byte of the NUH equal to 1 ; sets a second most significant bit in the first byte of the NUH equal to 0; sets a value of a NAL unit type indicator based on the determined NAL unit type; andsets a value of a temporal identifier indicator based on the determined temporal sublayer identifier.

[0466] In an embodiment, a decoder decodes an extension flag that replaces the forbidden zero bit in the first byte of the NUH, and when the extension flag is equal to 1, the decoder derives the temporal sublayer identifier based on a temporal identifier indicator and derives the NAL unit type based on the temporal identifier indicator and a NAL unit type indicator.

[0467] In an embodiment for encoding and / or decoding, one or more values or value ranges of the temporal identifier indicator are used to indicate offsets for the determined NAL unit type value. For example, a temporal identifier indicator equal to 0 may indicate an offset equal to 0 and a temporal identifier indicator greater than 0 may indicate an offset equal to 16 (assuming a unit type indicator in the range of 0 to 15, inclusive). A unit type may be determined as a sum of the offset and the unit type indicator. It is to be understood that embodiments may be realized with different assignments of the association between temporal identifier indicator value(s) and the offset(s). For example, embodiments may likewise be realized when a temporal identifier indicator equal to 0 indicates an offset equal to 16 (assuming a unit type indicator in the range of 0 to 15, inclusive) and a temporal identifier indicator greater than 0 indicates an offset equal to 0. In another example, a 3 -bit temporal identifier indicator equal to 7 may indicate an offset equal to 0 and Temporalld equal to 0 and a temporal identifier indicator less than 7 may indicate an offset equal to 16 (assuming a unit type indicator in the range of 0 to 15, inclusive) and Temporalld equal to the temporal identifier indicator.

[0468] In an example embodiment, the example syntax of the tenth option includes in the first byte of the NUH the following: an extension flag, such as nuh extension flag, 3-bit temporal ID indicator (nuh tid a idc and nuh tid b idc), and 4-bit NAL unit type indicator (nuh nalu type a idc and nuh nalu type b idc). Value 0 of the NAL unit type indicator is used to indicate a category of NAL unit types that imply temporal ID equal to 0.

[0469] In an example embodiment, the following syntax and semantics may be used.

[0470] In another example embodiment, the following syntax and semantics may be used.

[0471] nuh extension flag equal to 1 specifies that the NAL unit header comprises a second byte, nuh extension flag equal to 0 specifies that the NAL unit header comprises only one byte.

[0472] nuh_reserved_zero_bit shall be equal to 0.

[0473] nuh_tid_a_idc or nuh_tid_b_idc, whichever is present, affects the derivation of Temporalld and NalUnitType.

[0474] nuh nalu type a idc or nuh nalu type b idc, whichever is present, affects the derivation of NalUnitType.

[0475] NalUnitType is derived as follows: if ( nuh_extension_flag = = 1 ) NalUnitType = ( nuh_tid_a_idc >0)? 16 : 0) + nuh_nalu_type_a_idc elseNalUnitType = ( nuh_tid_b_idc >0)? 16 : 0) + nuh_nalu_type_b_idc

[0476] An example embodiment for the mapping of NalUnitType values to their names and NAL unit type classes is specified in the table below.

[0477] In an embodiment, the variable for the temporal sublayer identifier, Temporalld is derived as follows: if ( nuh extension flag = = 1 ) { if( nuh tid a idc > 0 )Temporalld = nuh tid a idc - 1 elseTemporalld = 0} else if( nuh tid b idc > 0 )Temporalld = nuh tid b idc - 1 elseTemporalld = 0

[0478] In another example embodiment, the following syntax and semantics may be used.

[0479] nuh extension flag equal to 1 specifies that the NAL unit header comprises a second byte, nuh extension flag equal to 0 specifies that the NAL unit header comprises only one byte.

[0480] nuh_reserved_zero_bit shall be equal to 0.

[0481] nuh_nalu_type_idc specifies the four least significant bits of the NAL unit type value, which is assigned to the variable NalUnitType as specified below.

[0482] nuh_tid_idc equal to 0 specifies that NalUnitType is in the range of 0 to 15, inclusive, and Temporalld is equal to 0. nuh tid idc greater than 0 specifies that NalUnitType is in the range of 16 to 31, inclusive, and Temporalld is equal to nuh tid idc - 1.

[0483] NalUnitType is set equal to ( ( nuh tid idc > 0 ) « 4 ) + nuh nalu type idc.

[0484] In an embodiment, if the extension flag is equal to 0 and the unit type indicator (e.g., nuh nalu type idc) is equal to 15, an encoder sets the temporal identifier indicator to 0 or 1 to avoid the first byte to have the values 188 to 191 which may emulate stream id values in MPEG-2 PES start codes.

[0485] In an embodiment, if the extension flag is equal to 0 and the unit type indicator (e.g., nuh nalu type idc) is equal to 15, a decoder concludes an erroneous unit if the temporal identifier indicator is greater than 1.

[0486] In another example embodiment, the following syntax may be used with the syntax as described above.

[0487] In some of the embodiments above, the following semantics apply.

[0488] nuh_layer_id specifies the scalability layer identifier.

[0489] nuh_reserved_one_bit shall be equal to 1. The value 0 of nuh reserved one bit could be specified in the future by ITU-T | ISO / IEC. Although the value of nuh reserved one bit is required to be equal to 1 in this version of this Specification, decoders conforming to this version of this Specification shall also allow the value of nuh reserved one bit equal to 0 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh reserved one bit equal to 0.

[0490] When not present, the value of nuh layer id is inferred to be equal to 0.

[0491] In some of the embodiments above, the following association of NalUnitType to the NAL unit type names may be used. It is to be understood that embodiments may likewise be realized with other assignments.

[0492] According to an embodiment, which may be used together with or independently of embodiments describing conditional decoding of the second byte of a header, decoding comprises: - receiving a bitstream comprising one or more units each having a header;- receiving a unit type syntax element and a temporal identifier indicator syntax element of the header;- in response to a value of the temporal identifier indicator being equal to a first value, o determining a temporal identifier to be equal to a predefined value (such as 0) and o determining a unit type to be equal to the sum of a first offset value and the value of the unit type syntax element;- in response to the value of the temporal identifier indicator being within a first range not including the first value, o determining the temporal identifier to be equal to the value of the temporal identifier indicator subtracted by a second offset value and o determining the unit type to be equal to the sum of a third offset value and the value of the unit type syntax element.

[0493] According to an embodiment, which may be used together with or independently of embodiments describing conditional encoding of the second byte of a header, encoding comprises:- encoding a bitstream comprising one or more units each having a header;- determining a unit type and a temporal identifier for a unit;- in response to a unit type being among a first set of unit types that always have temporal identifier equal to a pre-defined value (such as 0): o setting a value of a temporal identifier indicator syntax element in a header of the unit to be equal to a first value, and o setting a value of a unit type indicator syntax element in the header to be equal to the unit type subtracted by a first offset value;- in response to a unit type not being among the first set of unit types: o setting the value of a temporal identifier indicator syntax element in a header of the unit to be equal to the sum of the temporal identifier and a second offset value, ando setting the value of a unit type indicator syntax element in the header to be equal to the unit type subtracted by a third offset value.

[0494] In some examples, the above-described embodiments use the following:- The first value is 0. For example, nuh tid idc equal to 0 indicates a category of NAL unit types for which Temporalld is always equal to 0.- The first offset value is 0. As an example consequence, NAL unit types for which Temporalld is always equal to 0 occupy NAL unit types in the range of 0 to 15, inclusive.- The first range is from 1 to 7, inclusive.- The second offset value is 1.- The third offset value is 16. As an example consequence, NAL unit types occupy NAL unit types in the range of 16 to 31, inclusive.

[0495] It is to be understood that embodiments could likewise be realized with mathematically equivalent expressions. For example, offsets in some embodiments may be realized by bitshifting or vice versa. For example, setting NalUnitType equal to ( ( nuh tid idc > 0 ) « 4 ) + nuh nalu type idc could likewise be realized with any of the following: ( nuh tid idc > 0 ) ? 16 : 0 ) + nuh nalu type idc; or((nuh tid idc > 0) * 16) + nuh nalu type idc; or((nuh tid idc > 0) « 4) | nuh nalu type idc; or((nuh tid idc > 0) * 16) | nuh nalu type idc; or(((nuh_tid_idc | (nuh_tid_idc » 1) | (nuh_tid_idc » 2)) & 1) « 4) | nuh_nalu_type_idc

[0496] It is to be understood that embodiments may likewise be realized with any choices for syntax element names. For example, nuh tid idc in some embodiments could be alternatively named as nuh_temporal_id_plusl.

[0497] An apparatus for decoding according to an embodiment comprises means for receiving a bitstream comprising one or more units each having a header; means for receiving a first byte of the header; means for determining, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream; means for deriving at least one parameter value depending on the determination: if the header does not comprise the second byte, the apparatus comprises means for deriving the at least one parameter value based on at least one syntax element in the first byte and pre-defined information; whereinthe at least one parameter value comprises one or both of a scalability layer identifier and a temporal sublayer identifier.

[0498] As a further improvement, if the header comprises the second byte, the apparatus may comprise means for deriving the at least one parameter value based on the at least one syntax element in the first byte and information in the second byte.

[0499] The means discussed in previous, comprises at least one processor, and a memory including a computer program code, wherein the processor may further comprise processor circuitry. The memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform the method of Figure 3 according to various embodiments.

[0500] An apparatus for encoding according to an embodiment comprises means for determining at least one parameter value, wherein the at least one parameter value comprises one or more of following: a unit type, a scalability layer identifier, and a temporal sublayer identifier; means for deriving a one-byte header for a unit comprised in a bitstream, wherein the at least one parameter value is represented by at least one syntax element and pre-defined information; wherein if the one-byte header has an allowed value among one or more allowed values, the apparatus comprises means for determining a header for the unit to be the one-byte header, wherein the header comprises a first byte, and means for including the header in the bitstream.

[0501] As a further improvement, if the one-byte header has a disallowed value among one or more disallowed values the apparatus may comprise means for deriving a header comprising a first byte and a second byte, wherein the first byte is indicative that the second byte is present and wherein and a value of the at least one syntax element in the first byte and information in the second byte are set based on the at least one parameter value; and means for including the header in a video bitstream, wherein the second byte follows the first byte in the video bitstream.

[0502] According to an embodiment, the apparatus for decoding comprises means for receiving a bitstream comprising one or more units each having a header; means for receiving one or more syntax elements from a first byte of the header; means for determining, based on the one or more syntax elements, if the header comprises a second byte that follows the first byte in the bitstream; in response to the header not comprising the second byte, means for determining a first order and first bit positions of remainingsyntax elements in the header; in response to the header comprising the second byte, means for determining a second order and second bit positions of remaining syntax elements in the header, wherein the first order differs from the second order and / or the first bit positions differ at least partly from the second bit positions; means for receiving the remaining syntax elements depending on the determination of the first order and the first bit positions, or the second order and the second bit positions; and means for deriving at least one parameter value from the remaining syntax elements.

[0503] According to an embodiment, the apparatus for encoding comprises means for encoding a bitstream comprising one or more units each having a header; means for determining if a second byte of a header is needed for a unit; means for setting values of one or more syntax elements to indicate if the header comprises the second byte that follows the first byte in the bitstream; means for encoding the one or more syntax elements to a first byte of the header; in response to the header not comprising the second byte, means for determining a first order and first bit positions of remaining syntax elements in the header; in response to the header comprising the second byte, means for determining a second order and second bit positions of remaining syntax elements in the header, wherein the first order differs from the second order or the first bit positions differ at least partly from the second bit positions; means for setting values of remaining syntax elements based on at least parameter value associated with the unit; and means for encoding the remaining syntax elements to the header depending on the determination of the first order and the first bit positions, or the second order and the second bit positions.

[0504] According to a yet further embodiment, which may be used together with or independently of embodiments describing conditional decoding of the second byte of a header, the apparatus for decoding comprises:- means for receiving a bitstream comprising one or more units each having a header;- means for receiving a unit type syntax element and a temporal identifier indicator syntax element of the header;- in response to a value of the temporal identifier indicator being equal to a first value, o means for determining a temporal identifier to be equal to a predefined value (such as 0) ando means for determining a unit type to be equal to the sum of a first offset value and the value of the unit type syntax element;- in response to the value of the temporal identifier indicator being within a first range not including the first value, o means for determining the temporal identifier to be equal to the value of the temporal identifier indicator subtracted by a second offset value and o means for determining the unit type to be equal to the sum of a third offset value and the value of the unit type syntax element.

[0505] According to yet another embodiment, which may be used together with or independently of embodiments describing conditional encoding of the second byte of a header, the apparatus for encoding comprises:- means for encoding a bitstream comprising one or more units each having a header;- means for determining a unit type and a temporal identifier for a unit;- in response to a unit type being among a first set of unit types that always have temporal identifier equal to a predefined value (such as 0): o means for setting a value of a temporal identifier indicator syntax element in a header of the unit to be equal to a first value, and o means for setting a value of a unit type indicator syntax element in the header to be equal to the unit type subtracted by a first offset value;- in response to a unit type not being among the first set of unit types: o means for setting the value of a temporal identifier indicator syntax element in a header of the unit to be equal to the sum of the temporal identifier and a second offset value, and o means for setting the value of a unit type indicator syntax element in the header to be equal to the unit type subtracted by a third offset value.

[0506] The means discussed in any of the previous examples and embodiments, comprises at least one processor, and a memory including a computer program code, wherein the processor may further comprise processor circuitry. The memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform the method of Figure 4 according to various embodiments.

[0507] Figure 5 illustrates an example of an electronic apparatus 500, being an example of video coding system where the present embodiments can be implemented. In some embodiments, the apparatus may be a mobile terminal or a user equipment of a wireless communication system or a camera device. The apparatus 500 may also be comprised at a local or a remote server or a graphic processing unit of a computer. The apparatus may also be comprised as part of a head-mounted display device.

[0508] The apparatus may be configured to perform various functions, such as for example, gathering information by one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information gathered or received by the apparatus. An apparatus configured to encode a video scene may optionally comprise one or more microphones for capturing the scene and / or one or more cameras for capturing information about the physical environment in which the scene is captured. Alternatively, the apparatus configured for encoding may be configured to receive information about an environment in which a scene is captured and / or a simulated environment. An apparatus configured to decode and / or render the video scene may be configured to receive a bitstream comprising encoded video. An apparatus configured to decode and / or render the video scene may comprise one or more speakers / audio transducers and / or displays, and / or may be configured to transmit a decoded scene or signals to a device comprising one or more speakers / audio transducers and / or displays. An apparatus configured to decode and / or render the video scene may comprise a user equipment, a head-mounted display, or another device capable of rendering to a user an AR; VR and / or MR experience.

[0509] The apparatus 500 comprises one or more processors 510 and one or more memories 520 and one or more transceivers interconnected through one or more buses. The one or more memories 520 store computer instructions, for example in respective modules (Module 1, Module2, ModuleN). The one or more memories may store data in the form of image, video and / or audio data, and / or may also store instructions to be executed by the processors or the processor circuitry. The one or more processors may comprise a centralprocessing unit (CPU) and / or a graphical processing unit (GPU). The one or more buses may be address, data or control buses, and may include interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment. The apparatus also comprises a codec 630 that is configured to implement various embodiments relating to present solution. According to some embodiments, the apparatus may comprise an encoder or a decoder. The apparatus 600 also comprises a communication interface 540 which is suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network, and thus enabling data transfer over data transfer network 550.

[0510] The apparatus 500 may comprise a display in the form of a liquid crystal display. In other embodiments of the invention the display may be any suitable display technology suitable to display an image or video. The apparatus 500 may further comprise a keypad. In other embodiments of the invention any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display. The apparatus 500 may comprise a microphone or any suitable audio input which may be a digital or analogue signal input. The apparatus 500 may further comprise an audio output device which in embodiments of the invention may be any one of: an earpiece, speaker, or an analogue audio or digital audio output connection. The apparatus 500 may also comprise a battery (or in other embodiments of the invention the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera capable of recording or capturing images and / or video. The camera may be a multi-lens camera system having at least two camera sensors. The camera is capable of recording or detecting individual frames which are then passed to the codec 530 or to processor 510. The apparatus may receive the video and / or image data for processing from another device prior to transmission and / or storage.

[0511] The apparatus 500 may further comprise e.g., the other functional units disclosed in any of the Figures 1 - 2 for implementing any of the present embodiments.

[0512] The apparatus may operate in a system, comprising multiple communication devices, which can communicate through one or more networks. The system may comprise any combination of wired or wireless networks including, but not limited to a wireless cellulartelephone network (such as a GSM, UMTS, CDMA network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.

[0513] For example, the system can be a mobile telephone network enabling a connection to the internet. The connection can form, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0514] The example communication devices operating in the system may include, but are not limited to, an electronic device or apparatus, a combination of a personal digital assistant (PDA) and a mobile telephone, a PDA, an integrated messaging device (IMD), a desktop computer, a notebook computer, each of which can be a representative of the apparatus according to present embodiments. The apparatus according to present embodiments may be stationary or mobile when carried by an individual who is moving. The apparatus may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle, or any similar suitable mode of transport.

[0515] The apparatus may also be a set-top box; i.e. a digital TV receiver, which may / may not have a display or wireless capabilities, a tablet or (laptop) a personal computer (PC), which have hardware or software or combination of the encoder / decoder implementations, in various operating systems, or a chipset, processor, DSP and / or embedded system offering hardware / software based coding.

[0516] The apparatus according to present embodiments may send and receive calls and messages and communicate with service providers through a wireless connection to a base station. The base station may be connected to a network server that allows communication between the mobile telephone network and the internet.

[0517] The apparatus may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP -IP), short messaging service (SMS),multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11 and any similar wireless communication technology. A communications device involved in implementing various embodiments of the present invention may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0518] Figure 6 is a graphical representation of an example multimedia communication system within which various embodiments may be implemented. A data source 1510 provides a source signal in an analog, uncompressed digital, or compressed digital format, or any combination of these formats. An encoder 1520 may include or be connected with a pre-processing, such as data format conversion and / or filtering of the source signal. The encoder 1520 encodes the source signal into a coded media bitstream. It should be noted that a bitstream to be decoded may be received directly or indirectly from a remote device located within virtually any type of network. Additionally, the bitstream may be received from local hardware or software. The encoder 1520 may be capable of encoding more than one media type, such as audio and video, or more than one encoder 1520 may be required to code different media types of the source signal. The encoder 1520 may also get synthetically produced input, such as graphics and text, or it may be capable of producing coded bitstreams of synthetic media. In the following, only processing of one coded media bitstream of one media type is considered to simplify the description. It should be noted, however, that typically real-time broadcast services comprise several streams (typically at least one audio, video and text sub-titling stream). It should also be noted that the system may include many encoders, but in the figure only one encoder 1520 is represented to simplify the description without a lack of generality. It should be further understood that, although text and examples contained herein may specifically describe an encoding process, one skilled in the art would understand that the same concepts and principles also apply to the corresponding decoding process and vice versa.

[0519] The coded media bitstream may be transferred to a storage 1530. The storage 1530 may comprise any type of mass memory to store the coded media bitstream. The format of the coded media bitstream in the storage 1530 may be an elementary self-contained bitstream format, or one or more coded media bitstreams may be encapsulated into a container file, or the coded media bitstream may be encapsulated into a Segment format suitable for DASH (or a similar streaming system) and stored as a sequence of Segments.If one or more media bitstreams are encapsulated in a container file, a file generator (not shown in the figure) may be used to store the one more media bitstreams in the file and create file format metadata, which may also be stored in the file. The encoder 1520 or the storage 1530 may comprise the file generator, or the file generator is operationally attached to either the encoder 1520 or the storage 1530. Some systems operate “live”, i.e. omit storage and transfer coded media bitstream from the encoder 1520 directly to the sender 1540. The coded media bitstream may then be transferred to the sender 1540, also referred to as the server, on a need basis. The format used in the transmission may be an elementary self-contained bitstream format, a packet stream format, a Segment format suitable for DASH (or a similar streaming system), or one or more coded media bitstreams may be encapsulated into a container file. The encoder 1520, the storage 1530, and the server 1540 may reside in the same physical device or they may be included in separate devices. The encoder 1520 and server 1540 may operate with live real-time content, in which case the coded media bitstream is typically not stored permanently, but rather buffered for small periods of time in the content encoder 1520 and / or in the server 1540 to smooth out variations in processing delay, transfer delay, and coded media bitrate.

[0520] The server 1540 sends the coded media bitstream using a communication protocol stack. The stack may include but is not limited to one or more of Real-Time Transport Protocol (RTP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Transmission Control Protocol (TCP), and Internet Protocol (IP). When the communication protocol stack is packet-oriented, the server 1540 encapsulates the coded media bitstream into packets. For example, when RTP is used, the server 1540 encapsulates the coded media bitstream into RTP packets according to an RTP payload format. Typically, each media type has a dedicated RTP payload format. It should be again noted that a system may contain more than one server 1540, but for the sake of simplicity, the following description only considers one server 1540.

[0521] If the media content is encapsulated in a container file for the storage 1530 or for inputting the data to the sender 1540, the sender 1540 may comprise or be operationally attached to a “sending file parser” (not shown in the figure). In particular, if the container file is not transmitted as such but at least one of the contained coded media bitstream is encapsulated for transport over a communication protocol, a sending file parser locates appropriate parts of the coded media bitstream to be conveyed over the communicationprotocol. The sending file parser may also help in creating the correct format for the communication protocol, such as packet headers and payloads. The multimedia container file may contain encapsulation instructions, such as hint tracks in the ISOBMFF, for encapsulation of the at least one of the contained media bitstream on the communication protocol.

[0522] The server 1540 may or may not be connected to a gateway 1550 through a communication network, which may e.g. be a combination of a CDN, the Internet and / or one or more access networks. The gateway may also or alternatively be referred to as a middle-box. For DASH, the gateway may be an edge server (of a CDN) or a web proxy. It is noted that the system may generally comprise any number gateways or alike, but for the sake of simplicity, the following description only considers one gateway 1550. The gateway 1550 may perform different types of functions, such as translation of a packet stream according to one communication protocol stack to another communication protocol stack, merging and forking of data streams, and manipulation of data stream according to the downlink and / or receiver capabilities, such as controlling the bit rate of the forwarded stream according to prevailing downlink network conditions. The gateway 1550 may be a server entity in various embodiments.

[0523] The system includes one or more receivers 1560, typically capable of receiving, demodulating, and de-capsulating the transmitted signal into a coded media bitstream. The coded media bitstream may be transferred to a recording storage 1570. The recording storage 1570 may comprise any type of mass memory to store the coded media bitstream. The recording storage 1570 may alternatively or additively comprise computation memory, such as random-access memory. The format of the coded media bitstream in the recording storage 2170 may be an elementary self-contained bitstream format, or one or more coded media bitstreams may be encapsulated into a container file. If there are multiple coded media bitstreams, such as an audio stream and a video stream, associated with each other, a container file is typically used and the receiver 1560 comprises or is attached to a container file generator producing a container file from input streams. Some systems operate “live,” i.e. omit the recording storage 1570 and transfer coded media bitstream from the receiver 1560 directly to the decoder 1580. In some systems, only the most recent part of the recorded stream, e.g., the most recent 10-minute excerption of the recordedstream, is maintained in the recording storage 1570, while any earlier recorded data is discarded from the recording storage 1570.

[0524] The coded media bitstream may be transferred from the recording storage 1570 to the decoder 1580. If there are many coded media bitstreams, such as an audio stream and a video stream, associated with each other and encapsulated into a container file or a single media bitstream is encapsulated in a container file e.g. for easier access, a file parser (not shown in the figure) is used to decapsulate each coded media bitstream from the container file. The recording storage 1570 or a decoder 1580 may comprise the file parser, or the file parser is attached to either recording storage 1570 or the decoder 1580. It should also be noted that the system may include many decoders, but here only one decoder 1580 is discussed to simplify the description without a lack of generality.

[0525] The coded media bitstream may be processed further by a decoder 1580, whose output is one or more uncompressed media streams. Finally, a Tenderer 1590 may reproduce the uncompressed media streams with a loudspeaker or a display, for example. The receiver 1560, recording storage 1570, decoder 1580, and Tenderer 1590 may reside in the same physical device or they may be included in separate devices.

[0526] A sender 1540 and / or a gateway 1550 may be configured to perform switching between different representations e.g. for switching between different viewports of 360- degree video content, view switching, bitrate adaptation and / or fast start-up, and / or a sender 1540 and / or a gateway 1550 may be configured to select the transmitted representation(s). Switching between different representations may take place for multiple reasons, such as to respond to requests of the receiver 1560 or prevailing conditions, such as throughput, of the network over which the bitstream is conveyed. In other words, the receiver 1560 may initiate switching between representations. A request from the receiver can be, e.g., a request for a Segment or a Subsegment from a different representation than earlier, a request for a change of transmitted scalability layers and / or sub-layers, or a change of a rendering device having different capabilities compared to the previous one. A request for a Segment may be an HTTP GET request. A request for a Subsegment may be an HTTP GET request with a byte range. Additionally, or alternatively, bitrate adjustment or bitrate adaptation may be used for example for providing so-called fast start-up in streaming services, where the bitrate of the transmitted stream is lower than the channel bitrate after starting or random-accessing the streaming in order to start playbackimmediately and to achieve a buffer occupancy level that tolerates occasional packet delays and / or retransmissions. Bitrate adaptation may include multiple representation or layer up- switching and representation or layer down-switching operations taking place in various orders.

[0527] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.

[0528] A decoder 1580 may be configured to perform switching between different representations e.g. for switching between different viewports of 360-degree video content, view switching, bitrate adaptation and / or fast start-up, and / or a decoder 1580 may be configured to select the transmitted representation(s). Switching between different representations may take place for multiple reasons, such as to achieve faster decoding operation or to adapt the transmitted bitstream, e.g. in terms of bitrate, to prevailing conditions, such as throughput, of the network over which the bitstream is conveyed. Faster decoding operation might be needed for example if the device including the decoder 1580 is multi-tasking and uses computing resources for other purposes than decoding the video bitstream. In another example, faster decoding operation might be needed when content is played back at a faster pace than the normal playback speed, e.g. twice or three times faster than conventional real-time playback rate.

[0529] In the above, some embodiments have been described with reference to and / or using terminology of HEVC and / or VVC. It needs to be understood that embodiments may be similarly realized with any video encoder and / or video decoder.

[0530] For example, some embodiments have been described with reference to NAL unit header or NUH. It is to be understood that embodiments can be similarly realized with any unit header, such as an OBU header.

[0531] Some embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.

[0532] Some example embodiments have been described with the help of certain syntax elements the bitstream. It needs to be understood, however, that embodiments may be similarly realized with different sets of syntax elements that partly or fully cover the semantics of one or more syntax elements described in example embodiments.

[0533] In the above, where the example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder may have corresponding elements in them. Likewise, where the example embodiments have been described with reference to a decoder, it needs to be understood that the encoder may have structure and / or computer program for generating the bitstream to be decoded by the decoder.

[0534] The embodiments of the invention described above describe the codec in terms of separate encoder and decoder apparatus in order to assist the understanding of the processes involved. However, it would be appreciated that the apparatus, structures and operations may be implemented as a single encoder-decoder apparatus / structure / operation. Furthermore, it is possible that the coder and decoder may share some or all common elements.

[0535] Although the above examples describe embodiments of the invention operating within a codec within an electronic device, it would be appreciated that the invention as defined in the claims may be implemented as part of any video codec. Thus, for example, embodiments of the invention may be implemented in a video codec which may implement video coding over fixed or wired communication paths.

[0536] Thus, user equipment may comprise a video codec such as those described in embodiments of the invention above. It shall be appreciated that the term user equipment is intended to cover any suitable type of wireless user equipment, such as mobile telephones, portable data processing devices or portable web browsers.

[0537] Furthermore, elements of a public land mobile network (PLMN) may also comprise video codecs as described above.

[0538] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects ofthe invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0539] The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.

[0540] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general- purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multi-core processor architecture, as non-limiting examples.

[0541] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0542] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus,GDSII, or the like) may be transmitted to a semiconductor fabrication facility or “fab” for fabrication.

[0543] The various embodiments can be implemented with the help of computer program code that resides in a memory and causes the relevant apparatuses to carry out the method. For example, a device may comprise circuitry and electronics for handling, receiving, and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the device to carry out the features of an embodiment. Yet further, a network device like a server may comprise circuitry and electronics for handling, receiving, and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the network device to carry out the features of various embodiments.

[0544] If desired, the different functions discussed herein may be performed in a different order and / or concurrently with other. Furthermore, if desired, one or more of the abovedescribed functions and embodiments may be optional or may be combined.

[0545] Although various aspects of the embodiments are set out in the independent claims, other aspects comprise other combinations of features from the described embodiments and / or the dependent claims with the features of the independent claims, and not solely the combinations explicitly set out in the claims.

[0546] It is also noted herein that while the above describes example embodiments, these descriptions should not be viewed in a limiting sense. Rather, there are several variations and modifications, which may be made without departing from the scope of the present disclosure as, defined in the appended claims.

Claims

CLAIMS:

1. An apparatus comprising at least one processor and at least one memory, said at least one memory stored with code thereon, which when executed by said at least one processor, causes the apparatus to:- receive a bitstream comprising one or more units each having a header;- receive a first byte of the header;- determine, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream;- derive at least one parameter value depending on the determination:• if the header does not comprise the second byte, the apparatus is caused to derive the at least one parameter value based on at least one syntax element in the first byte and pre-defined information;- wherein the at least one parameter value comprises one or both of a scalability layer identifier and a temporal sublayer identifier.

2. The apparatus according to claim 1, wherein if the header comprises the second byte, the apparatus is caused to derive the at least one parameter value based on the at least one syntax element in the first byte and information in the second byte.

3. The apparatus according to claim 1, wherein the header is a network abstraction layer (NAL) unit header.

4. The apparatus according to claim 1, wherein the apparatus is caused to determine from an extension flag in the first byte of the header whether the header comprises the second byte.

975. The apparatus according to claim 1, wherein the apparatus is caused to derive the at least one parameter value based on a non-reference flag in the first byte of the header.

6. The apparatus according to claim 4, wherein the apparatus is caused to derive the at least one parameter value based on a first flag in the first byte of the header and the extension flag.

7. The apparatus according to claim 1, wherein when the header does not comprise the second byte, the apparatus is caused to determine if a first flag in the first byte is enabled, wherein based on the determination that the first flag is enabled, the apparatus is caused to determine the temporal sublayer identifier to be equal to a first value, and based on the determination that the first flag is not enabled, the apparatus is caused to determine the temporal sublayer identifier to be equal to a second value; and wherein when the header comprises the second byte, the apparatus is caused to determine the temporal sublayer identifier based on the temporal indicator in the second byte and the first flag in the first byte.

8. The apparatus according to claim 1, wherein the apparatus is caused to determine from a no extension flag in a first bit of the header whether the header comprises a second byte.

9. The apparatus according to claim 8, wherein when the header does not comprise the second byte, the apparatus is caused to derive the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

10. The apparatus according to claim 8, wherein when the header comprises the second byte, the apparatus is caused to determine the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on a parameter specifying the oneor more least significant bits of the temporal sublayer identifier defined in the first byte.

11. The apparatus according to claim 8, wherein when the header comprises the second byte, the apparatus is caused to determine the scalability layer identifier from the second byte.

12. The apparatus according to claim 8, wherein when the header does not comprise the second byte, the apparatus is caused to determine the scalability layer identifier based on a parameter specifying one or more least significant bits defined in the first byte.

13. The apparatus according to claim 8, wherein when the header comprises the second byte, the apparatus is caused to determine the scalability layer identifier based on a parameter specifying one or more least significant bits for the scalability layer identifier defined in the first byte and a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte.

14. The apparatus according to claim 8, wherein when the header does not comprise the second byte, the apparatus is caused to derive the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte and to derive the scalability layer identifier based on a parameter specifying one or more least significant bits of the scalability layer identifier defined in the first byte.

15. The apparatus according to claim 8, wherein when the header comprises the second byte, the apparatus is caused to determine the scalability layer identifier based on a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte and a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte and to determine the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte and a parameter99specifying one or more most significant bits for the temporal sublayer identifier defined in the second byte.

16. The apparatus according to claim 4, wherein when the header does not comprise the second byte, the apparatus is caused to derive the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

17. The apparatus according to claim 4, wherein when the header comprises the second byte, the apparatus is caused to determine the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on the parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

18. The apparatus according to claim 8, wherein when the header does not comprise the second byte, the apparatus is caused to determine the temporal sublayer identifier based on a temporal identifier indicator in the first byte and to determine a unit type based on the temporal identifier indicator and a unit type indicator in the first byte.

19. The apparatus according to claim 8, wherein when the header comprises the second byte, the apparatus is caused to determine the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal identifier indicator in the second byte and a temporal identifier indicator in the first byte and to determine a unit type based on the temporal identifier indicator and a unit type indicator in the first byte.

20. A method comprising- receiving a bitstream comprising one or more units each having a header;- receiving a first byte of the header;- determining, based on the first byte, if the header comprises a second byte that follows the first byte in the bitstream;- deriving at least one parameter value depending on the determination:100• if the header does not comprise the second byte, the method comprises deriving the at least one parameter value based on at least one syntax element in the first byte and predefined information;- wherein the at least one parameter value comprises one or both of a scalability layer identifier and a temporal sublayer identifier.

21. The method according to claim 20, wherein if the header comprises the second byte, the method comprises deriving the at least one parameter value based on the at least one syntax element in the first byte and information in the second byte.

22. The method according to claim 20, further comprising determining from an extension flag in the first byte of the header whether the header comprises the second byte.

23. The method according to claim 20, further comprising deriving the at least one parameter value based on a non-reference flag in the first byte of the header.

24. The method according to claim 23, further comprising deriving the at least one parameter value based on a first flag in the first byte of the header and the extension flag.

25. The method according to claim 20, wherein when the header does not comprise the second byte, the method comprises determining if a first flag in the first byte is enabled, wherein based on the determination that the first flag is enabled, the method comprises determining the temporal sublayer identifier to be equal to a first value, and based on the determination that the first flag is not enabled, the method comprises determining the temporal sublayer identifier to be equal to a second value; and wherein when the header comprises the second byte, the method comprises determining the temporal sublayer identifier based on the temporal indicator in the second byte and the first flag in the first byte.

26. The method according to claim 20, further comprising determining from a no extension flag in a first bit of the header whether the header comprises a second byte.

27. The method according to claim 26, wherein when the header does not comprise the second byte, the method comprises deriving the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

28. The method according to claim 26, wherein when the header comprises the second byte, the method comprises determining the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on a parameter specifying the one or more least significant bits of the temporal sublayer identifier defined in the first byte.

29. The method according to claim 26, wherein when the header comprises the second byte, the method comprises determining the scalability layer identifier from the second byte.

30. The method according to claim 26, wherein when the header does not comprise the second byte, the method comprises determining the scalability layer identifier based on a parameter specifying one or more least significant bits defined in the first byte.

31. The method according to claim 26, wherein when the header comprises the second byte, the method comprises determining the scalability layer identifier based on a parameter specifying one or more least significant bits for the scalability layer identifier defined in the first byte and a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte.

32. The method according to claim 26, wherein when the header does not comprise the second byte, the method comprises deriving the temporal sublayer identifier102based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte and deriving the scalability layer identifier based on a parameter specifying one or more least significant bits of the scalability layer identifier defined in the first byte.

33. The method according to claim 26, wherein when the header comprises the second byte, the method comprises determining the scalability layer identifier based on a parameter specifying one or more least significant bits of the scalability layer identifier in the first byte and a parameter specifying one or more most significant bits for the scalability layer identifier defined in the second byte and determining the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier in the first byte and a parameter specifying one or more most significant bits for the temporal sublayer identifier defined in the second byte.

34. The method according to claim 22, wherein when the header does not comprise the second byte, the method comprises deriving the temporal sublayer identifier based on a parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

35. The method according to claim 22, wherein when the header comprises the second byte, the apparatus is caused to determine the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal sublayer identifier defined in the second byte and on the parameter specifying one or more least significant bits of the temporal sublayer identifier defined in the first byte.

36. The method according to claim 26, wherein when the header does not comprise the second byte, the method comprises determining the temporal sublayer identifier based on a temporal identifier indicator in the first byte and to determine a unit type based on the temporal identifier indicator and a unit type indicator in the first byte.

37. The method according to claim 26, wherein when the header comprises the second byte, the method comprises determining the temporal sublayer identifier based on a parameter specifying one or more most significant bits of the temporal identifier indicator in the second byte and a temporal identifier indicator in the first byte and determining a unit type based on the temporal identifier indicator and a unit type indicator in the first byte.

38. An apparatus comprising at least one processor and at least one memory, said at least one memory stored with code thereon, which when executed by said at least one processor, causes the apparatus to:- receive a bitstream comprising one or more units each having a header;- receive one or more syntax elements from a first byte of the header;- determine, based on the one or more syntax elements, if the header comprises a second byte that follows the first byte in the bitstream;- in response to the header not comprising the second byte, determine a first order and first bit positions of remaining syntax elements in the header;- in response to the header comprising the second byte, determine a second order and second bit positions of remaining syntax elements in the header, wherein the first order differs from the second order and / or the first bit positions differ at least partly from the second bit positions;- receive the remaining syntax elements depending on the determination of the first order and the first bit positions, or the second order and the second bit positions; and- derive at least one parameter value from the remaining syntax elements.

39. An apparatus according to claim 38, wherein the apparatus is caused to receive an extension flag from the first byte of the header for said one or more syntax elements from the first byte.

40. An apparatus according to claim 39, when the extension flag is equal to zero, the apparatus is caused to determine that the header does not comprise a second byte,and when the extension flag is equal to one, the apparatus is caused to determine that the header comprises a second byte.

41. The apparatus according to claim 39, when the extension flag is equal to one, the apparatus is caused to read a reserved zero bit from the header during decoding.

42. The apparatus according to claim 38, wherein the remaining syntax elements comprise either or both of a unit type indicator and / or a temporal identifier indicator.

43. The apparatus according to claim 42, wherein the bit position(s) of either or both of the unit type indicator and / or the temporal identifier indicator depend on whether the header comprises the second byte.

44. The apparatus according to claim 42, wherein the order of unit type indicator and a temporal identifier indicator depends on whether the header comprises the second byte.

45. The apparatus according to claim 38, wherein the at least one parameter value comprises a unit type and / or a temporal sublayer identifier.

46. The apparatus according to claim 38, further being caused to- receive a unit type syntax element and a temporal identifier indicator syntax element of the header;- in response to a value of the temporal identifier indicator being equal to a first value, o determine a temporal sublayer identifier to be equal to 0 and o determine a unit type to be equal to the sum of a first offset value and the value of the unit type syntax element;- in response to the value of the temporal identifier indicator being within a first range not including the first value,o determine the temporal sublayer identifier to be equal to the value of the temporal identifier indicator subtracted by a second offset value and o determine the unit type to be equal to the sum of a third offset value and the value of the unit type syntax element.

47. A method comprising- receiving a bitstream comprising one or more units each having a header;- receiving one or more syntax elements from a first byte of the header;- determining, based on the one or more syntax elements, if the header comprises a second byte that follows the first byte in the bitstream;- in response to the header not comprising the second byte, determining a first order and first bit positions of remaining syntax elements in the header;- in response to the header comprising the second byte, determining a second order and second bit positions of remaining syntax elements in the header, wherein the first order differs from the second order and / or the first bit positions differ at least partly from the second bit positions;- receiving the remaining syntax elements depending on the determination of the first order and the first bit positions, or the second order and the second bit positions; and- deriving at least one parameter value from the remaining syntax elements.

48. The method according to claim 47, further being caused to- receiving a unit type syntax element and a temporal identifier indicator syntax element of the header;- in response to a value of the temporal identifier indicator being equal to a first value, o determining a temporal sublayer identifier to be equal to 0 and o determining a unit type to be equal to the sum of a first offset value and the value of the unit type syntax element;- in response to the value of the temporal identifier indicator being within a first range not including the first value, o determining the temporal sublayer identifier to be equal to the value of the temporal identifier indicator subtracted by a second offset value and o determining the unit type to be equal to the sum of a third offset value and the value of the unit type syntax element.

49. An apparatus comprising at least one processor and at least one memory, said at least one memory stored with code thereon, which when executed by said at least one processor, causes the apparatus to:- receive a bitstream comprising one or more units each having a header;- receive a unit type syntax element and a temporal identifier indicator syntax element of the header;- in response to a value of the temporal identifier indicator being equal to a first value, o determine a temporal sublayer identifier to be equal to 0 and o determine a unit type to be equal to the sum of a first offset value and the value of the unit type syntax element;- in response to the value of the temporal identifier indicator being within a first range not including the first value, o determine the temporal sublayer identifier to be equal to the value of the temporal identifier indicator subtracted by a second offset value and o determine the unit type to be equal to the sum of a third offset value and the value of the unit type syntax element.

50. A method comprising:- receiving a bitstream comprising one or more units each having a header;- receiving a unit type syntax element and a temporal identifier indicator syntax element of the header;- in response to a value of the temporal identifier indicator being equal to a first value, o determining a temporal sublayer identifier to be equal to 0 and o determining a unit type to be equal to the sum of a first offset value and the value of the unit type syntax element;- in response to the value of the temporal identifier indicator being within a first range not including the first value, o determining the temporal sublayer identifier to be equal to the value of the temporal identifier indicator subtracted by a second offset value and o determining the unit type to be equal to the sum of a third offset value and the value of the unit type syntax element.108