Configurable NAL and Segment Code Point Mechanism for Stream Fusion

MX431505BActive Publication Date: 2026-02-25FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MX2022002644
Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-03
Filing Date
2022-03-03
Publication Date
2026-02-25
Estimated Expiration
2040-09-03

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in merging multiple encoded video streams due to restrictions on Network Abstraction Layer (NAL) unit types, particularly when combining Random Access Point (RAP) and non-RAP units, which can cause prediction errors and require excessive bit rates for IDR images.

Method used

A mechanism for mapping NAL unit types within a syntax structure, allowing flexible conversion between RAP and non-RAP units, and utilizing parameter sets to indicate additional information in segment headers, enabling efficient merging of video streams without transcoding.

Benefits of technology

Enhances the efficiency of video stream merging by reducing bit rate requirements for IDR images and allowing seamless integration of different NAL unit types, improving decoding performance and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure MX431505B0
    Figure MX431505B0
Patent Text Reader

Abstract

Video decoder configured to decode a video comprising a plurality of images from a video data stream by decoding each image from one or more video encoding units within an access unit of the video data stream that is associated with the respective image; reading a surrogate encoding unit type from a parameter set unit of the video data stream; for each default video encoding unit, reading an encoding unit type identifier (100) from the respective video encoding unit;verify whether the encoding unit identifier identifies an encoding unit type from a first subset of one or more encoding unit types (102) or from a second subset of encoding unit types (104); if the encoding unit identifier identifies an encoding unit type from the first subset of one or more encoding unit types, attribute the respective video encoding unit to the substitute encoding unit type; if the encoding unit identifier identifies an encoding unit type from the second subset of encoding unit types, attribute the respective video encoding unit to the encoding unit type from the second subset of encoding unit types identified by the encoding unit identifier.
Need to check novelty before this filing date? Find Prior Art

Description

NAL CONFIGURE AND SEGMENT CODE POINT MECHANISM FOR STREAM MERGING Μταζηη / ζζηζ / Β / γίΛΐ Field of Invention This application relates to a data structure to indicate a type of encoding unit and characteristics of a video encoding unit of a video data stream. Background of the Invention Image types are known to be specified in the NAL drive headers of the NAL drives that carry the image segments. Therefore, essential NAL drive payload properties are readily available at a high level for application use. The types of images include the following: Random Access Point (RAP) images are where a decoder can begin decoding an encoded video sequence. These are referred to as Intra-Random Access (IRAP) images. There are three types of IRAP images: Instant Decoder Update (IDR), Clean Random Access (CRA), and Broken Link Access (BLA). The decoding process for an encoded video sequence always begins at an IRAP. - Front images, which precede a random access point image in output order, but are Ref. 332239 encodes after it in the encoded video sequence. Forward frames that are independent of the frames preceding the random access point in encoding order are called Random Access Decodable Forward Frames (RADLs). Forward frames that use frames preceding the random access point in encoding order for prediction could be corrupted if decoding begins at the corresponding IRAP. These are called Random Access Skipped Forward Frames (RASLs). - Rear Images (TRAIL), which follow the IRAP and the front images, both in order of output and display. - Images in which the decoder can switch the temporal resolution of the encoded video sequence: Temporal Sublayer Access (TSA) and Stepped Temporal Sublayer Access (STSA). Therefore, the data structure of the NAL unit is an important factor for flow fusion. Brief Description of the Invention The subject matter of this application is to provide a decoder that derives the necessary information from a video encoding unit from a video data stream by reading an identifier that indicates a type of substituted encoding unit and a decoder that derives features from a video data stream. The additional object of the subject matter of this application is to provide an encoder that indicates a type of encoding unit substituted for a video encoding unit by using an identifier and an encoder that indicates characteristics of a video damage stream. This objective is achieved through the objective matter of the claims of the requested presence. According to the modalities of the present application, a video decoder configured to decode a video comprising a plurality of images from a video data stream by decoding each image from one or more video encoding units within an access unit of the video data stream that is associated with the respective image; reading a substitute encoding unit type from a parameter set unit of the video data stream; for each predetermined video encoding unit, reading an encoding unit type identifier (100), for example, a syntax element included in a final unit header, from the respective encoding unit; verifying whether the encoding unit identifier identifies an encoding unit type from a first subset of one or more types of Μταζηη / ζζηζ / Β / γίΛΐ encoding unit (102), for example, indicate whether the final unit is of the VCL (video encoding layer) unit type unappealable or not, or of a second subset of encoding unit types (104), for example, indicate the final unit type, if the encoding unit identifier identifies an encoding unit type from the first subset of one or more encoding unit types, attribute the respective default video encoding unit to the substitute encoding unit type; if the encoding unit identifier identifies an encoding unit type from the second subset of encoding unit types, attribute the respective default video encoding unit to the encoding unit type from the second subset of encoding unit types identified by the encoding unit identifier.In other words, the respective final unit type is indicated by the identifier, the first subset of encoding unit types, and the second subset of encoding unit types; that is, the final unit type is rewritten according to the indication of the first and second subsets of encoding unit types. Therefore, it is possible to improve merging efficiency. According to the terms of this application, the video decoder is configured to decode, from each video encoding unit, the region associated with the respective video encoding unit in a manner that depends on the encoding unit type assigned to that unit. The video decoder may be configured so that the substitute encoding unit type is outside the second subset of video encoding types. The video decoder may also be configured so that the substitute encoding unit type is outside a third subset of video encoding types, for example, a non-VCL unit type, comprising at least one video encoding type not included in the second subset of video encoding types. According to this application, it is possible to improve encoding efficiency. According to the terms of this application, the default video coding units carry image block partitioning data, block-related prediction parameters, and residual prediction data. When an image contains one or more video coding units, for example, segments, with a coding unit type from the first subset and one or more video coding units, for example, segments, with a coding unit type from the second subset, these latter video coding units are of a unit type of The encoding is equal to the type of substitute encoding unit. The type of substitute encoding unit is a random access point (RAP) encoding type. The type of substitute encoding unit is a different encoding type than a random access point (RAP) encoding type. That is, the type of substitute encoding unit is identified, and video encoding units that have the same type of substitute encoding unit are merged, thus appropriately improving merging efficiency. According to the terms of this application, each of the default video encoding units is associated with a different region of the image, as is the access unit within which the respective default video encoding unit is located. The parameter set unit of the video data stream has a scope that covers a sequence of images, an image, or a set of image segments. The parameter set unit is indicative of the type of substitute encoding unit for a specific video data stream profile. That is, it is possible to efficiently merge segments and thus improve encoding efficiency. In accordance with the modalities of the present In the request, the parameter set unit of the video data stream is; the parameter set unit that has a scope covering a sequence of images, or an access unit delimiter that has a scope covering one or more of the images associated with the access unit. That is, the sequence of images is properly indicated and, therefore, it is possible to efficiently decode the images that are required to be rendered. According to the terms of this application, the parameter set unit indicates the type of substitute encoding unit in a video data stream. This is true whether the default video encoding unit is used as the updated starting point of the video sequence for decoding a video, for example, a RAP type (i.e., it includes an instantaneous decoding update, IDR), or as the continuous starting point of the video sequence for decoding a video, for example, a non-RAP type (i.e., it does not include IDR). In other words, it is possible to indicate whether the encoding unit is the first frame of the video sequence or not when using the parameter set unit. According to the modalities of the present application, a video decoder configured to decode a video comprising a plurality of images from a video data stream by decoding each image from one or more Μταζηη / ζζηζ / Β / γίΛΐ video coding units within an access unit of the video data stream that is associated with the respective image, wherein each video coding unit carries image block partitioning data, block-related prediction parameters, and residual prediction data and is associated with a different region of the image than the access unit within which the respective default video coding unit is located; reading, from each of the default video coding units, an n-ary set of one or more syntax elements, e.g., two flags, each being 2-ary so that the pair is 4-ary, map(200), e.g., the mapping can be set by default;Alternatively, the data stream is signaled, or both by splitting the range of values, the nary set of one or more syntax elements into an m-ary set of one or more features (202), e.g., three binary features, each of which is therefore 2-ary so that the triplet is 8-ary, each feature redundantly describing with corresponding data in the default video encoding unit, i.e., the features can be deduced from a deeper inspection of encoding data, as to how the video is encoded in the video data stream with respect to the picture with which the access unit is associated; Μταζηη / ζζηζ / Β / γίΛΐ within which is the default video encoding unit, where m>n, or read, from each of the default video encoding units, N syntax elements (210), for example, N=2 flags, each of which is 2-ary, with N>0, read association information from the video data stream, associate, i.e. treat them as a variable of the associated feature, depending on the association information, each of the N syntax elements with information about one of the M features, for example, M=3 binary features, each of which is, therefore, 2-ary, the association information would have 3 possibilities of associating the two flags with 2 out of 3, i.e., UP features,Each feature that redundantly describes the corresponding data in the default video encoding unit regarding how the video is encoded in the video data stream with respect to the image associated with the access unit within which the default video encoding unit is located, where M>N. That is, for example, the video data stream condition, i.e., how the video is encoded in the video data stream with respect to the image in the access unit, is indicated by the map and flags, making it possible to efficiently provide additional information. Μταζηη / ζζηζ / Β / γίΛΐ According to the terms of this application, the map is included in the parameter set unit and indicates the location of the mapped features. The map is flagged in the data flow and indicates the location of the mapped features. The N syntax elements are indicative of the presence of the features. That is, by combining the flag and the mapping, there is flexibility in indicating the flags in the parameter set. According to the form of the present application, a video encoder configured to encode a video comprising a plurality of images in a video data stream by encoding each image in one or more video encoding units within an access unit of the video data stream that is associated with the respective image; indicating a substitute encoding unit type in a parameter set unit of the video data stream; for each default video encoding unit, encoding in the video data stream an encoding unit type identifier (100) for the respective encoding unit, wherein the encoding unit identifier identifies an encoding unit type from a first subset of one or more encoding unit types (102) or from a second subset of encoding unit types (104),where if the encoding unit identifier identifies an encoding unit type from the first subset of one or more encoding unit types, the respective default video encoding unit is to be assigned to the substitute encoding unit type; if the encoding unit identifier identifies an encoding unit type from the second subset of encoding unit types, the respective default video encoding unit is to be assigned to the encoding unit type from the second subset of encoding unit types identified by the encoding unit identifier, where the substitute encoding unit type is a RAP type and the video encoder is configured to identify RAP image video encoding units as the default video encoding units and, for example,Directly encode the encoding unit type identifier for purely intracoded video encoding units of non-RAP images that identify a RAP type. That is, the encoding unit type is indicated in the parameter set unit of the video data stream and, therefore, it is possible to improve encoding efficiency; i.e., it is not necessary to encode each sector with an IDR image. Μταζηη / ζζηζ / Ε / γίΛΐ According to the modalities of the present application, a video compositor configured to compose a video data stream having a video comprising a plurality of images encoded therein, each image being in one or more video encoding units within an access unit of the video data stream, one or more video encoding units being associated with the respective image for each of the tiles into which the images are subdivided; changing a substitute encoding unit type in a parameter set unit of the video data stream from indicating a RAP type to indicating a non-RAP type; identifying in the vds images exclusively encoded video encoding units whose identifier (100) encoded in the video data stream a encoding unit type identifies a RAP image;wherein for each of the predetermined video coding units of the video data stream, an identifier (100) for the respective video coding unit p encoded in the video data stream, a coding unit type identifies a coding unit type from a first subset of one or more coding unit types (102) or from a second subset of coding unit types (104), wherein if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, the respective predetermined video coding unit is to be attributed to the substitute coding unit type;If the encoding unit identifier identifies an encoding unit type from the second subset of encoding unit types, the respective default video encoding unit will be assigned to the encoding unit type from the second subset of encoding unit types identified by the encoding unit identifier. The video encoding unit type is identified by using the identifier, a first and a second subset of encoding unit types, and thus the video image, for example, constructed from a plurality of tiles, is efficiently composited. According to the modalities of the present application, a video encoder configured to encode a video comprising a plurality of images in a video data stream by encoding each image in one or more video encoding units within an access unit of the video data stream that is associated with the respective image, wherein each video encoding unit carries image block partitioning data, block-related prediction parameters, and residual prediction data and is associated with a different region of the image than the access unit within which the default video encoding unit is located. respective Mταζηη / ζζηζ / Ε / γίΛΐ; indicate, in each of the default video encoding unit, an n-ary set of one or more syntax elements, e.g., two flags, each being 2-ary so that the pair is 4-ary, map (200), the mapping can be set by default; alternatively, signal in the data stream, or both by splitting the range of values, the n-ary set of one or more syntax elements into an m-ary set of one or more features (202), e.g., three binary features, each being therefore 2-ary so that the triplet is 8-ary, each feature being redundantly described with corresponding data in the default video encoding unit, i.e., the features can be deduced from an inspection of deeper encoding data,Regarding how the video is encoded in the video data stream with respect to the image with which the access unit is associated, within which the default video encoding unit is located, where m>n, or indicate, in each of the default video encoding units, N syntax elements (210), for example, N=2 flags, each of which is 2-ary, with N>0, indicate association information in the video data stream, associate, i.e., treat them as a variable of the associated feature, depending on the association information, each of the, Μταζηη / ζζηζ / Ε / γίΛΐ N syntax elements with information about one of the M features, for example, M=3 binary features, each of which is, therefore, 2-ary association information would have 3 possibilities of associating the two flags with 2 out of 3, i.e., features, each feature being described in a redundant manner with the corresponding data in the default video encoding unit as to how the video is encoded in the video data stream with respect to the image with which the access unit within which the default video encoding unit is located is associated, where M>N. That is, for example, the features of each video encoding unit of an encoded video sequence are indicated by using a flag and, therefore, it is possible to provide additional information efficiently. According to the modalities of the present application, a method comprising decoding a video comprising a plurality of images from a video data stream by decoding each image from one or more video coding units within an access unit of the video data stream that is associated with the respective image; reading a type of coding unit that substitutes for a parameter set unit of the video data stream; for each unit of Μταζηη / ζζηζ / Β / γίΛΐ default video encoding, read an encoding unit type identifier from the respective encoding unit; check if the encoding unit identifier identifies an encoding unit type from a first subset of one or more encoding unit types or from a second subset of encoding unit types, if the encoding unit identifier identifies an encoding unit type from the first subset of one or more encoding unit types, attribute the respective video encoding unit to the substitute encoding unit type;If the encoding unit identifier identifies an encoding unit type from the second subset of encoding unit types, attribute the respective video encoding unit to the encoding unit type from the second subset of encoding unit types identified by the encoding unit identifier. According to embodiments of the present application, a method comprising decoding a video comprising a plurality of images from a video data stream by decoding each image from one or more video encoding units within an access unit of the video data stream that is associated with the respective image, wherein each video encoding unit carries image block partitioning data, block-related prediction parameters, and residual prediction data and is associated with a different region of the image than the access unit within which the respective default video encoding unit is located; reading, from each of the default video encoding unit, an n-ary set of one or more syntax elements, for example, two flags, each being 2-ary so that the pair is 4-ary, the mapping being fixed by default;Alternatively, the data stream is signaled, or both by splitting the range of values, the n-ary set of one or more syntax elements into an m-ary set of one or more features, e.g., three binary features, each of which is therefore 2-ary so that the triplet is 8-ary, each feature being redundantly described by the corresponding data in the default video encoding unit [i.e., the features can be deduced from an inspection of deeper encoding data] as to how the video is encoded in the video data stream with respect to the picture with which the access unit within which the default video encoding unit is located is associated, wherein m>n, or read, from each of the default video encoding units, N syntax elements, e.g., N=2 flags, each of which is 2-ary, with N>0, read one; Μταζηη / ζζηζ / Β / γίΛΐ video data stream association information, associate, i.e. treat them as a variable of the associated feature, depending on the association information, each of the N syntax elements with information about one of the M features, e.g., M=3 binary features, each of which is, therefore, 2-ary association information would have 3 possibilities of associating the two flags with 2 out of 3, i.e., features, each feature being described in a redundant manner with the corresponding data in the default video encoding unit as to how the video is encoded in the video data stream with respect to the picture with which the access unit within which the default video encoding unit is located is associated, where M>N. According to the modalities of the present application, a method comprising encoding a video comprising a plurality of images in a video data stream by encoding each image in one or more video encoding units within an access unit of the video data stream that is associated with the respective image; indicating a substitute encoding unit type in a parameter set unit of the video data stream;For each default video encoding unit, define an encoding unit type identifier (100) for the respective encoding unit, wherein the encoding unit identifier identifies an encoding unit type from a first subset of one or more encoding unit types (102) or from a second subset of encoding unit types (104), if the encoding unit identifier identifies an encoding unit type from the first subset of one or more encoding unit types, attribute the respective video encoding unit to the substitute encoding unit type;If the encoding unit identifier identifies an encoding unit type from the second subset of encoding unit types, attribute the respective video encoding unit to the encoding unit type from the second subset of encoding unit types identified by the encoding unit identifier. According to the modalities of the present application, a method comprising composing a video data stream having a video comprising a plurality of images encoded therein, each image being in one or more video encoding units within an access unit of the video data stream, one or more video encoding units being associated with the respective image for each of the tiles into which the images are subdivided; changing a substitute encoding unit type in a parameter set unit of the video data stream from indicating a RAP type to indicating a non-RAP type; identifying in the images vds exclusively encoded video encoding units whose identifier (100) encoded in the video data stream a encoding unit type identifies a RAP image;wherein for each of the default video coding units of the video data stream, an identifier (100) for the respective video coding unit p encoded in the video data stream a coding unit type identifies a coding unit type from a first subset of one or more coding unit types (102) or from a second subset of coding unit types (104), wherein if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, the respective default video coding unit is to be attributed to the substitute coding unit type;If the encoding unit identifier identifies an encoding unit type from the second subset of encoding unit types, the respective default video encoding unit will be attributed to the encoding unit type from the second subset of unit types; Μταζηη / ζζηζ / Β / γίΛΐ coding identified by the coding unit identifier. According to embodiments of the present application, a method comprising encoding, a video comprising a plurality of images in a video data stream by encoding each image in one or more video encoding units within an access unit of the video data stream that is associated with the respective image, wherein each video encoding unit carries image block partitioning data, block-related prediction parameters, and residual prediction data and is associated with a different region of the image than the access unit within which the respective default video encoding unit is located; indicating, in each of the default video encoding unit, an n-ary set of one or more syntax elements, for example, two flags, each being 2-ary so that the pair is 4-ary, map(200), the mapping being fixed by default;Alternatively, it is signaled in the data stream, or both by splitting the range of values, the n-ary set of one or more syntax elements into an m-ary set of one or more features (202), e.g., three binary features, each of which is therefore 2-ary so that the triplet is 8-ary, each; The characteristic is redundantly described with the corresponding data in the default video encoding unit, i.e., the characteristics can be deduced from a deeper inspection of the encoding data, as to how the video is encoded in the video data stream with respect to the image with which the access unit within which the default video encoding unit is located is associated, where m>n, or indicating, in each of the default video encoding units, N syntax elements (210), e.g., N=2 flags, each of which is 2-ary, with N>0, indicating association information in the video data stream, associating, i.e., treating them as a variable of the associated characteristic, depending on the association information, each of the N syntax elements with information about one of the M characteristics, e.g., M=3 binary characteristics,each one that is, therefore, 2-ary association information would have 3 possibilities of associating the two flags with 2 out of 3, i.e., Lv / features, each feature being described in a redundant manner with the corresponding data in the default video encoding unit as to how the video is encoded in the video data stream with respect to the picture with which the access unit is associated within which the default video encoding unit is located, where M>N., Brief Description of the Figures The following describes preferred forms of this application with respect to the figures, including: Figure 1 shows a schematic diagram illustrating a client-server system for virtual reality applications as an example of where the modalities shown in the following figures can be used advantageously; Figures 2a and 2b show a schematic illustration that indicates an example of a 360-degree video in a cube map projection at two resolutions and in 6x4 tiled tiles, which can be adjusted to the system in Figure 1; Figures 3a-3c show a schematic illustration indicating an example of user viewpoint and tile selections for real-time 360-degree video streaming as shown in Figure 2; Figure 4 shows a schematic illustration indicating an example of a resulting tiling arrangement (packing) of tiles indicated in Figures 3a-3c into a single bitstream after the operation of Μταζηη / ζζηζ / Β / γίΛΐ fusion; Figures 5a-5c show a schematic illustration that indicates an example of a tiling with a low-resolution backing as an individual tile for real-time transmission of 360-degree video; Figures 6a and 6b show a schematic illustration indicating an example of tile selection in tile-based real-time streaming; Figures 7a and 7b show a schematic illustration that indicates another example of tile selection in real-time transmission based on unequal RAP (random access point) tile-based ... Figure 8 shows a schematic diagram indicating an example of a NAL (Network Adaptive Layer) unit header according to the modalities of this application; Figure 9 shows an example of a table indicating the type of RBSP (Row Byte Sequence Payloads) data structure contained in the NAL unit according to the modalities of this application; Figure 10 shows a schematic diagram indicating an example of a sequence parameter set indicating that the NAL unit type is mapped according to the modalities of the present application; Figure 11 shows a schematic diagram indicating an example of the characteristics of a segment indicated in a segment header; Figure 12 shows a schematic diagram indicating an example of a map that indicates characteristics of a segment in the sequence parameter set according to the modalities of the present application; Figure 13 shows a schematic diagram indicating an example of features indicated by the map in the sequence parameter set of Figure 12 according to the modalities of this application; Figure 14 shows a schematic diagram indicating an example of association information indicated when using the syntax structure according to the modalities of this application; Figure 15 shows a schematic diagram indicating another example of the map that indicates characteristics of a segment in the sequence parameter set according to the modalities of the present application; Figure 16 shows a schematic diagram indicating another example of features indicated by the map in the parameter set of Figure 15 according to the modalities of this application; Figure 17 shows a schematic diagram indicating an additional example of the map indicating segment characteristics in a segment sector header according to the modalities of this application; Μταζηη / ζζηζ / Β / γίΛΐ and Figure 18 shows a schematic diagram indicating an additional example of features indicated by the map in the segment sector header of Figure 17 according to the modalities of this application. Equal or equivalent elements or elements with equal or equivalent functionality are denoted in the following description by equal or equivalent reference numbers. Detailed Description of the Invention The following description sets out several details to provide a more complete explanation of the methods described herein. However, it will be evident to a person skilled in the art that the methods described herein can be practiced without these specific details. In other cases, well-known structures and devices are shown in block diagram form rather than in detail to avoid complicating the methods described herein. Furthermore, the features of the different methods described below can be combined with each other, unless specifically stated otherwise. Introductory Observations In what follows, it should be noted that the individual aspects described The terms Μταζηη / ζζηζ / Β / γΐΛ in this document can be used individually or in combination. Therefore, details can be added to each of these individual aspects without adding details to any other of these aspects. It should also be noted that this description explicitly or implicitly describes features usable in a video decoder (a device for providing a decoded representation of a video signal based on an encoded representation). Therefore, any of the features described herein may be used in the context of a video decoder. Furthermore, the features and functionalities described herein in relation to a method can also be used in a device (configured to perform this functionality). Likewise, any feature and functionality described herein with respect to a device can also be used in a corresponding method. In other words, the methods described herein can be supplemented with any of the features and functionalities described with respect to devices. To facilitate understanding of the modalities described in this application with respect to its various aspects, Figure 1 shows an example of an environment where the modalities described below can be advantageously applied and used. In particular, the Figure 1 shows a system composed of client 10 and server 20 interacting through adaptive real-time streaming. For example, Dynamic Adaptive Real-Time Streaming over HTTP (DASH) can be used for communication 22 between client 10 and server 20. However, the modalities detailed later should not be interpreted as restricted to the use of DASH, and similarly, terms such as Media Presentation Description (MPD) should be understood as being broad enough to also cover manifest files defined differently than in DASH. Figure 1 illustrates a system configured to implement a virtual reality application. Specifically, the system is configured to present a user using a front display 24—that is, via an internal display 26 of the front display 24—with a view section 28 of a time-varying spatial scene 30. This section 28 corresponds to an orientation of the front display 24 measured, for example, by an internal orientation sensor 32, such as an inertial sensor of the front display 24. In other words, the section 28 presented to the user forms a section of the spatial scene 30 whose spatial position corresponds to the orientation of the front display 24. In the case of Figure 1, the time-varying spatial scene 30 is represented as an omnidirectional video. The spherical Mταζηη / ζζηζ / Β / γίΛΐ description of Figure 1 and the modalities explained later can also be easily transferred to other examples, such as presenting a section of a video with a spatial position of section 28 determined by the intersection of a facial or visual access point with a virtual or real projector wall or similar. Furthermore, the sensor 32 and the display 26 can, for example, be comprised of different devices such as a remote control and corresponding television, respectively, or they can be part of a portable device such as a mobile device like a tablet or a mobile phone.Finally, it should be noted that some of the modalities described below can also be applied to scenarios where the area 28 presented to the user constantly covers the entire time-varying spatial scene 30 with the inequality in the presentation of the time-varying spatial scene related, for example, to an uneven distribution of quality across the spatial scene. Figure 1 illustrates further details regarding server 20, client 10, and how spatial content 30 is delivered on server 20, as described below. These details, however, should not be considered limitations of the modalities explained later, but rather serve as an example of how to implement any of them. In particular, as shown in Figure 1, the server 20 may comprise a storage device 34 and a controller 36, such as a suitably programmed computer, an application-specific integrated circuit, or the like. The storage device 34 has media sectors stored therein that represent the time-varying spatial scene 30. A specific example will be discussed in more detail later with reference to the illustration in Figure 1. The controller 36 responds to requests sent by the client 10 by forwarding the requested media sectors to the client 10, along with a media presentation description, and may also send the client 10 additional information on its own. Further details regarding this are also discussed later. The controller 36 can retrieve requested media sectors from the storage device 34.Within this storage, other information can also be stored, such as the description of media presentation or parts thereof, in the other signals sent from server 20 to client 10. As shown in Figure 1, the server 20 may optionally further comprise a flow modifier 38 that modifies the media segments sent from the server 20 to the client 10 in response to requests from The latter, to result in client 10 receiving a media data stream that forms an individual media stream decodable by an associated decoder, although, for example, the media segments retrieved by client 10 in this way are actually aggregated from several media streams. However, the existence of this stream modifier 38 is optional. The client 10 in Figure 1 is illustrated as an example comprising a client device or controller 40 or more decoders 42 and a reprojector 44. The client device 40 may be a suitably programmed computer, a microprocessor, a programmed hardware device such as an FPGA or an application-specific integrated circuit, or the like. The client device 40 is responsible for selecting sectors to be retrieved from the server 20 from the plurality 46 of media sectors offered on the server 20. To this end, the client device 40 first retrieves a media presentation description or manifest from the server 20. From this, the client device 40 obtains a computational rule for calculating the addresses of the media sectors from the plurality 46 that correspond to certain required spatial portions of the spatial scene 30.The media sectors selected in this way are retrieved by the client device 40 from the server 20 when sending. Μταζηη / ζζηζ / Β / γίΛΐ requests respective to server 20. These requests contain calculated addresses. The media sectors retrieved in this way by the client device 40 are forwarded by the latter to one or more decoders 42 for decoding. In the example in Figure 1, the media sectors retrieved and decoded in this way represent, for each unit of time, merely a spatial section 48 outside the time-varying spatial scene 30, but as previously stated, this may differ according to other factors, where, for example, the view section 28 to be presented constantly covers the entire scene. The reprojector 44 can optionally reproject and crop the view section 28 to be displayed to the user from the scene content retrieved and decoded from the selected, retrieved, and decoded media sectors.To this end, as shown in Figure 1, the client device 40 can, for example, continuously track and update a spatial position of view section 28 sensitive to user orientation data from sensor 32 and inform the reprojector 44, for example, about this current spatial position of scene section 28, as well as the reprojection mapping to be applied to retrieved and decoded media content to be mapped onto the area forming view section 28. The reprojector 44. Μταζηη / ζζηζ / Β / γίΛΐ can, therefore, apply mapping and interpolation to a regular grid of pixels, for example, to be displayed on screen 26. Figure 1 illustrates the case where a cubic mapping has been used to map spatial scene 30 onto 50 tiles. The tiles are thus represented as rectangular subregions of a cube onto which scene 30, which is in the shape of a sphere, has been projected. The reprojector 44 inverts this projection. However, other examples can also be applied. For instance, instead of a cubic projection, a projection onto a truncated pyramid or a pyramid without truncation can be used. Furthermore, although the tiles in Figure 1 are represented as non-overlapping in terms of spatial scene 30 coverage, the subdivision into tiles may involve a mutual overlap of tiles. And, as will be described in more detail later, the subdivision of scene 30 into 50 tiles spatially, with each tile forming a representation as explained later, is also not mandatory. Therefore, as illustrated in Figure 1, the entire spatial scene 30 is spatially subdivided into 50 tiles. In the example in Figure 1, each of the six faces of the cube is subdivided into 4 tiles. For illustrative purposes, the tiles are numbered. For each tile 50, the server Μταζηη / ζζηζ / Β / γίΛΐ offers a video 52 as illustrated in Figure 1. To be more precise, server 20 even offers more than one video 52 per tile 50, these videos differing in quality Q#. Furthermore, the videos 52 are temporarily subdivided into temporary sectors 54. The temporary sectors 54 of all the videos 52 of all the tiles T# form, or are encoded in, respectively, one of the media sectors of the plurality 46 of media sectors stored in storage 34 of server 20. It is emphasized again that even the example of a tile-based real-time streaming transmission illustrated in Figure 1 is merely one example from which many deviations are possible. For example, although Figure 1 seems to suggest that media sectors belonging to a higher-quality representation of Scene 30 refer to tiles that match tiles containing media sectors with Scene 30 encoded at quality Q1, this match is not necessary, and tiles of different qualities may even correspond to tiles from a different projection of Scene 30. Furthermore, although not yet analyzed, it is possible that the media sectors corresponding to different quality levels represented in Figure 1 differ in spatial resolution and / or signal-to-noise ratio and / or temporal resolution. Similar Μταζηη / ζζηζ / Β / γίΛΐ . Finally, unlike a mosaic-based real-time streaming concept, according to which the media sectors that can be individually retrieved by the server's device 40 refer to tiles 50 into which the scene 30 is spatially subdivided, the media sectors offered on the server 20 can alternatively, for example, each have the scene 30 encoded in it in a spatially complete manner with a spatially variable sampling resolution, yet having a maximum sampling resolution at different spatial positions in the scene 30. For example, this could be achieved by offering on the server 20 sequences of sectors 54 related to a projection of the scene 30 onto truncated pyramids whose truncated tips are oriented in mutually different directions, thus leading to differently oriented resolution peaks. Furthermore, regarding the optionally present flow modifier 38, it is noted that it can alternatively be part of client 10, or it can even be placed between them, within a network device through which client 10 and server 20 exchange the signals described herein. There is a certain video-based application in which Multiple encoded video bitstreams are decoded together, i.e., merged into a single bitstream and fed into a single decoder, such as: • Multiparty conferences: where encoded video streams from multiple participants are processed at a single endpoint • or tile-based real-time streaming: for example, for 360-degree tiled video playback in VR applications In the latter, a 360-degree video is spatially segmented, and each spatial sector is offered to streaming clients in real time in multiple representations of different spatial resolutions, as illustrated in Figures 2 and 2b. Figure 2a shows high-resolution mosaics, and Figure 2b shows low-resolution mosaics. Figures 2a and 2b depict a cube map projected onto a 360-degree video divided into 6x4 spatial sectors at two resolutions. For simplicity, these independently decodable spatial sectors are referred to as mosaics in this description. A user typically views only a subset of the tiles that make up the complete 360-degree video when using prior art head-mounted displays, as illustrated in Figure 3a through a solid viewing window boundary 80 representing a Μταζηη / ζζηζ / Β / γίΛΐ 90x90 degree field of view. The corresponding mosaics are indicated by a reference number 82 in figure 3b, and are downloaded at the highest resolution. However, the client application will also need to download and decode a representation of the other tiles outside the current viewing window, indicated by reference number 84 in Figure 3c, in order to handle sudden changes in the user's orientation. Therefore, a client in this application would download tiles covering its current viewing window at the highest resolution and tiles outside its current viewing window at a comparatively lower resolution, as shown in Figure 3c, while the selection of tile resolutions constantly adapts to the user's orientation. After downloading on the client side, merging the downloaded tiles into a single bitstream to be processed by a single decoder is a means of addressing the limitations of typical mobile devices with limited computational and power resources.Figure 4 illustrates a possible tiling arrangement in a joint bitstream for the previous examples. The merging operations to generate a joint bitstream must be performed through compressed domain processing, that is, avoiding pixel-domain processing through transcoding. While the example in Figure 4 illustrates the case where all tiles (high and low resolution) cover the entire 360-degree space and no tiles repeatedly cover the same regions, another set of tiles can also be used, as illustrated in Figures 5a–5c. Define the entire low-resolution portion of the video as a low-resolution backing layer, as shown in Figure 5b, which can be merged with the high-resolution tiles in Figure 5a that cover a subset of the 360-degree video. The entire low-resolution backing video can then be encoded as a single tile, as shown in Figure 5c, while the high-resolution tiles are rendered as an overlay on the low-resolution portion of the video in the final stage of the rendering process. A client begins a real-time streaming session according to their tile selection by downloading all the desired tile tracks as illustrated in Figures 6a and 6b, where a client starts the session with tile 0 indicated by reference number 90 and tile 1 indicated by reference number 92 in Figure 6a. Whenever a viewing window change occurs (i.e., the user turns their head to look elsewhere), the tile selection is changed in the next time sector that occurs, i.e., tile 0 Μταζηη / ζζηζ / Β / γίΛΐ and mosaic 2 indicated by a reference number 94 in figure 6a and in the next available sector, the client changes the position of mosaic 2 and replaces mosaic 0 with mosaic 1 as indicated in figure 6b. It is important to note that all sectors must start with an IDR (Instant Decoder Update) image, i.e., prediction chain reset images, as for any new mosaic selection and mosaic position change, otherwise it will cause prediction errors, artifacts, and deviation. Encoding each sector with an IDR image is costly in terms of bit rate. Sectors can potentially be very short-lived, for example, to react quickly to orientation changes, so it is desirable to encode multiple variants with variable IDR periods (or RAPs: Random Access Points), as illustrated in Figures 7a-7b. For example, as shown in Figure 7b, at time instance ti, there is no reason to break the prediction chain for tile 0 since tile 0 has already been downloaded for that time instance and placed in the same position. Therefore, a client can choose a sector that does not start with a RAP that is available on the server. However, a remaining problem is that the segments (tiles) within an encoded image must Mtazen / zziz / B / GLA must obey certain restrictions. One of these is that an image cannot contain Network Abstraction Layer (NAL) units of both RAP and non-RAP unit types simultaneously. Therefore, for requests, there are only two less desirable options to address the above problem. First, clients can rewrite the NAL unit type of RAP images when they are merged with non-RAP NAL units in an image. Second, servers can hide the RAP characteristic of these images by using non-RAP from the outset. However, this makes it difficult to detect RAP characteristics on systems that must handle these encoded videos, for example, for file format packaging. The invention is a NAL unit type mapping, which allows mapping one NAL unit type to another NAL unit type through an easily rewritable syntax structure. In one embodiment of the invention, a type of NAL unit is specified as mappable and the mapped type is specified in a set of parameters, for example, as follows based on draft 6 V14 of the WC (Versatile Video Coding) specification with highlighted edits. Figure 8 shows an NAL unit header syntax. The syntax nal_unit_type, that is, the identifier 100, specifies the NAL unit type. Μταζηη / ζζηζ / Β / γίΛΐ, that is, the type of data structure RBSP (Row Byte Sequence Payloads) contained in the NAL unit as specified in the table shown in Figure 9. The NalUnitType variable is defined as follows: When nal_unit_type != MAP_NUT NalUnitType is equal to nal_unit_type Otherwise (nal unit type == MAP NUT) NalUnitType equals mapped_nut All references to the syntax element nal_unit_type in the specification are replaced with references to the variable NalUnitType, for example, as in the following constraint: The value of NalUnitType must be the same for all encoded segment NAL units in an image. An image or layer access unit is referred to as having the same NAL unit type as the encoded segment NAL units of the image or layer access unit. That is, as illustrated in Figure 9, a first subset of encoding unit types 102 indicates nal_unit_type 12, namely MAP_NUT and VCL as the NAL unit type class. Therefore, a second subset of encoding unit types 104 indicates VCL as the NAL unit type class, meaning that all encoded segment NAL units in an image, as indicated by identifier 100 number 0 to 15, have the same class. Μταζηη / ζζηζ / Β / γίΛΐ NAL unit type of the coding unit type of the first subset of coding unit types 102, i.e., VCL. Figure 10 shows a set of RBSP syntax sequence parameters that includes mapped_nut, as indicated by reference sign 106, which indicates that the NalUnitType of NAL units with nal unit type is equal to MAP_NUT. In another mode, that mapped_nut syntax element is carried in the access unit delimiter, AUD. In another modality, it is a bitstream conformance requirement that the value of mapped_nut must be a VCL NAL unit type. In another approach, the mapping of the NalUnitType of NAL units with nal_unit_type equal to MAP_NUT is performed using profiling information. This mechanism could allow for more than one assignable NAL unit type instead of a single MAP_NUT, and specify the required interpretation of the NALUnitTypes of assignable NAL units within a simple profiling mechanism or an individual syntax element mapped_nut_space_idc. In another mode, the mapping mechanism is used to extend the range of values ​​of NALUnitTypes currently limited to 32 (since it is a u(5) , for example, as Μταζηη / ζζηζ / Β / γίΛΐ indicates in figure 10. The mapping mechanism could indicate any unlimited value as long as the number of NALUnitTypes required does not exceed the number of values ​​reserved for mappable NAL units. In one mode, when a simultaneous image contains segments of the surrogate encoding unit type and segments of the regular encoding unit types (e.g., existing NAL units of the VCL category), the mapping is performed in a way that results in all segments of the image effectively having the same encoding unit type properties; that is, the surrogate encoding unit type is equal to the encoding unit type of the non-substitute segments of the regular encoding types. Furthermore, the above mode is valid only for images with random access properties or for images without random access properties. In addition to the problems described regarding NAL drive types in fusion scenarios and NAL drive type extensibility and corresponding solutions, there are several video applications where information related to the video and how the video has been encoded is required for system integration and transmission or manipulation, such as on-the-fly adaptation. There is certain common information that has been established in recent years that is widely used in the industry and is clearly specified, using specific bit values ​​for this purpose. Examples of these are: • Temporary ID in the NAL unit header • Types of NAL units, including IDR, CRA, TRAIL, ... or SPS (Sequence Parameter Set), PPS (Image Parameter Set), etc. However, there are several scenarios where additional information could be useful. Other types of NAL units, not widely used but found useful in some cases, include BLA, partially RAP NAL units for sub-images, and non-reference sublayer NAL units. Some of these NAL unit types could be implemented using the extensibility mechanism described earlier. Alternatively, however, certain fields within the segment headers can be used. In the past, additional information has been reserved in segment headers used to indicate a particular characteristic of a segment: • discardable flag: specifies that the encoded image is not used as a reference image for interprediction and is not used as a source image for interlayer prediction. • cross layer bla flag: affects the derivation of the Μταζηη / ζζηζ / Β / γίΛΐ output images for layer encoding, where the image preceding the RAP in the upper layers may not be produced. A similar mechanism could be envisioned for future video codec standards. However, a limitation of such mechanisms is that the defined flags occupy a specific position within the segment header. Figure 11 below illustrates the use of these flags in HEVC. As seen previously, the problem with this solution is that the position of the extra segment header bits is allocated progressively, and for applications that use information less frequently, the flag would likely end up in a later position, increasing the number of bits that need to be sent in the extra bits (e.g., discardable_flag and crosslayer bla flag in the case of HEVC). Alternatively, following a mechanism similar to that described for NAL unit types, the mapping of flags to the additional segment header bits in the segment header could be defined in parameter sets. An example is shown in Figure 12. Figure 12 shows an example of a sequence parameter set that includes a map to indicate Μταζηη / ζζηζ / Β / γίΛΐ association information using extra_slice_header_bits_mapping_space_idc, i.e., a map, as indicated by the reference sign 200, indicating the mapping space for the extra bits in the segment header. Figure 13 shows the bit mapping with the flags present, as indicated by `extra_slice_header_bits_mapping_space_idc 200` in Figure 12. As illustrated in Figure 13, the binary features 202 redundantly describe the corresponding data in the default video encoding unit. In Figure 13, three binary features 202 are represented: 0, 1, and 2. The number of binary features could vary depending on the number of flags. In another mode, this mapping is carried out in a syntax structure (for example, as illustrated in Figure 14, which indicates the presence of a syntax element in the extra bits in the segment header (for example, as illustrated in Figure 11). That is, for example, as illustrated in Figure 11, a condition on a presence flag controls the presence of a syntax element in the extra bits in the segment header, i.e., num_extra_slice_header_bits>iy ;i <num_extra_slice_header_bits; 1 + +. En la figura 11, cada The syntax element in the extra bits is placed in a particular position in the segment header as explained above; however, in this mode, it is not necessary for the syntax element, for example, discardable flag, cross layer bla flag, or slice_reserved_flag[i], to occupy that particular position. Instead, when it is indicated that a first syntax element (for example, discardable flag in Figure 11) is not present when the condition is checked on the value of the particular presence flag (for example, discardable_flag_present_flag in Figure 14), a subsequent second syntax element takes the position of the first syntax element in the segment header when it is present. Furthermore, the syntax element in the extra bits could be present in the image header when indicating the flag, for example, sps_extra_ph_bit_present_flag[i].Furthermore, the syntax structure, for example, the number of syntax elements (i.e., the number of flags presented), indicates the existence / presence of particular features or a number of syntax elements presented in the extra bits. That is, the number of particular features presented or syntax elements in the extra bits is indicated by counting how many syntax elements (flags) are presented. In Figure 14, the presence of each syntax element is indicated by flags 210. In other words, each flag in 210 indicates the presence of a segment header indication for particular features of the default video encoding unit. Additionally, an extra syntax [...] / / additional flags as shown in Figure 14 and corresponding to the syntax elements slice_reserved_flag[i]' ' in the segment header of Figure 11 is used as a placeholder indicating the existence / presence of a syntax element in the additional bit or is used as an indication of the existence / presence of additional flags. In another mode, flag-type mapping is signaled by each additional segment header bit in a parameter set, for example, as shown in Figure 15. As indicated in Figure 15, an extra_slice_header_bit_mapping_idc syntax, i.e., map 200, is signaled in the sequence parameter set and indicates the location of the mapped features. Figure 16 shows the bit mapping with the flags present indicated by extra_slice_header_bits_mapping_space_idc in Figure 15. That is, the binary features 202 corresponding to the map 200 indicated in Figure 15 are illustrated in Figure 16. In another mode, the segment header extension bits are replaced by an ido signaling that Μταζηη / ζζηζ / Β / γίΛΐ represents a specific flag value combination, for example, as shown in Figure 17. As illustrated in Figure 17, map 200, i.e., extra slice header bit idc, is indicated in the segment sector header, i.e., map 200 indicates the presence of features as shown in Figure 18. Figure 18 shows that the flag values, i.e., the 202 binary features, represented by a certain value of extra_slice_header_bit_idc, are signaled in a set of parameters or predefined in the specification (known a priori). In one mode, the value space of extra_slice_header_bit_idc, i.e., the value space for map 200, is divided into two intervals. One interval represents the combinations of flag values ​​known a priori, and the other interval represents the combinations of flag values ​​signaled in the parameter sets. Although some aspects have been described in the context of an apparatus, it is evident that these aspects also represent a description of the corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, the aspects described in the context of a method step also represent a description of a Μταζηη / ζζηζ / Β / γίΛΐ block or element or corresponding feature of a corresponding device. Some or all of the method steps can be executed by (or using) a hardware device such as a microprocessor, a programmable computer, or an electronic circuit. In some modalities, one or more of the most important method steps can be executed by this device. The data stream of the invention can be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet. Depending on certain implementation requirements, the application's modalities can be implemented in hardware or software. Implementation can be carried out using a digital storage medium, such as a floppy disk, DVD, Blu-ray disc, CD, ROM, PROM, EPROM, EEPROM, or flash memory, that has electronically readable control signals stored within it, which cooperate (or are capable of cooperating) with a programmable computer system to perform the respective method. Therefore, the digital storage medium must be computer-readable. Some embodiments according to the invention comprise a data carrier having control signals Electronically readable Μταζηη / ζζηζ / Β / γίΛΐ, which are capable of cooperating with a programmable computer system, so that one of the methods described herein is carried out. In general, the embodiments of this application can be implemented as a computer program product with program code, the program code being operational for performing one of the methods when the computer program product is executed on a computer. The program code can, for example, be stored on a machine-readable medium. Other modalities include a computer program for performing one of the methods described herein, stored on a machine-readable carrier In other words, a modality of the inventive method is, therefore, a computer program that has program code to perform one of the methods described herein, when the computer program is executed on a computer. A further form of inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for carrying out one of the methods described herein. The data carrier, the digital storage medium, or the recorded medium is typically tangible or non-tangible. Μταζηη / ζζηζ / Ε / γίΛΐ transients . An additional form of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for carrying out one of the methods described herein. The data stream or the sequence of signals can, for example, be configured to be transferred over a data communication connection, such as the Internet. An additional modality comprises a processing means, for example, a computer or a programmable logic device, configured to, or adapted to, perform one of the methods described herein. One modality also includes a computer that has installed on it the computer program to perform one of the methods described herein. A further embodiment according to the invention comprises an apparatus or system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device, or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver. In some forms, a logical device A field-programmable gate array (e.g., a field-programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably implemented using any hardware device. The apparatus described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer. The apparatus described herein, or any component of the apparatus described herein, may be implemented at least partially in hardware and / or software. It is hereby stated that, as of this date, the best method known to the applicant for putting the present invention into practice is the one that is clear from the present description of the invention.

Claims

1. A decoding apparatus, characterized in that it comprises a microprocessor and memory, the memory comprising a computer program, which, when executed by the microprocessor, causes the decoding apparatus to: receive a sequence parameter set (SPS) from a video data stream, the SPS including an additional segment header bitmap SPS comprising a plurality of bits, wherein each bit of the plurality of bits comprised by the additional segment header bitmap SPS indicates a value of a corresponding syntax element presence flag, wherein each value indicates either the presence or non-presence of the corresponding syntax element in a segment header;analyze each bit in the SPS additional segment header bitmap of the SPS, wherein a first bit in the SPS additional segment header bitmap corresponds to a first syntax element and the first bit indicates the absence of the first syntax element in the segment header, and wherein a second bit in the SPS additional segment header bitmap corresponds to a second syntax element, the second bit indicates the presence of the second syntax element in the segment header, and the second bit follows the first bit in the SPS additional segment header bitmap; receive the segment header from the video data stream;analyze a first additional segment header bit of the segment header, the first additional segment header bit that is included in a set of one or more additional segment header bits in a position that is found first in the set of one or more additional segment header bits; and apply the first additional segment header bit to the second syntax element based on the first bit and the second bit.

2. The decoding apparatus according to claim 1, characterized in that the application of the first additional segment header bit to the second syntax element is based on the first bit indicating the absence of the first syntax element in the segment header and the second bit indicating the presence of the second syntax element in the segment header.

3. The decoding apparatus according to claim 1, characterized in that the computer program further causes the decoding apparatus to determine a number of bits included in the additional segment header bit map SPS.

4. The decoding apparatus according to claim 1, characterized in that the computer program further causes the decoding apparatus to determine a number of additional segment header bits included in the segment header based on the additional segment header bit map SPS.

5. The decoding apparatus according to claim 1, characterized in that the computer program further causes the decoding apparatus to determine that the segment header is not for a dependent segment, wherein the analysis of the first additional segment header bit of the segment header is performed in response to the determination that the segment header is not for a dependent segment.

6. A decoding method, characterized in that it comprises: receiving a sequence parameter set (SPS) from a video data stream, the SPS including an additional segment header bitmap SPS comprising a plurality of bits, wherein each bit of the plurality of bits comprised by the additional segment header bitmap SPS indicates a value of a corresponding syntax element presence flag, wherein each value indicates either the presence or non-presence of the corresponding syntax element in a segment header;analyze each bit in the SPS additional segment header bitmap of the SPS, wherein a first bit in the SPS additional segment header bitmap corresponds to a first syntax element and the first bit indicates the absence of the first syntax element in the segment header, and wherein a second bit in the SPS additional segment header bitmap corresponds to a second syntax element, the second bit indicates the presence of the second syntax element in the segment header, and the second bit follows the first bit in the SPS additional segment header bitmap; receive the segment header from the video data stream;analyze a first additional segment header bit of the segment header, the first additional segment header bit that is included in a set of one or more additional segment header bits in a position that is found first in the set of one or more additional segment header bits; and apply the first additional segment header bit to the second syntax element based on the first bit and the second bit.; 7. The decoding method according to claim 6, characterized in that the application of the first additional segment header bit to the second syntax element is based on the first bit indicating the absence of the first syntax element in the segment header and the second bit indicating the presence of the second syntax element in the segment header.

8. The decoding method according to claim 6, characterized in that it further comprises determining a number of bits included in the additional segment SPS header bitmap.

9. The decoding method according to claim 6, characterized in that it further comprises determining a number of additional segment header bits included in the segment header based on the additional segment header bit map SPS.

10. The decoding method according to claim 6, characterized in that it further comprises determining that the segment header is not for a dependent segment, wherein the analysis of the first additional segment header bit of the segment header is performed in response to the determination that the segment header is not for a dependent segment.

11. A non-transient video data stream, characterized in that it comprises: a sequence parameter set (SPS) including an additional segment header bitmap SPS comprising a plurality of bits, wherein each bit of the plurality of bits comprised by the additional segment header bitmap SPS indicates a value of a corresponding syntax element presence flag, wherein each value indicates either the presence or absence of the corresponding syntax element in a segment header, wherein a first bit in the additional segment header bitmap SPS corresponds to a first syntax element and the first bit indicates the absence of the first syntax element in the segment header, and wherein a second bit in the additional segment header bitmap SPS corresponds to a second syntax element,the second bit indicates the presence of the second syntax element in the segment header, and the second bit follows the first bit in the additional segment header bitmap SPS; and the segment header, which includes a set of one or more additional segment header bits and includes a first additional segment header bit in a position that is first in the set of one or more additional segment header bits, wherein the first bit and the second bit indicate that the first additional segment header bit applies to the second syntax element.

12. The video data stream according to claim 11, characterized in that the first bit indicating the absence of the first syntax element in the segment header and the second bit indicating the presence of the second syntax element in the segment header, together indicate that the additional first segment header bit applies to the second syntax element.

13. The video data stream according to claim 11, characterized in that the segment header is not for a dependent segment, and the additional first segment header bit is included in the segment header on the basis that the segment header is not for a dependent segment.

14. An encoding apparatus characterized in that it comprises a microprocessor and memory, the memory comprising a computer program, which, when executed by the microprocessor, causes the encoding apparatus to: encode a sequence parameter set (SPS) into a video data stream, the SPS including an additional segment header bitmap SPS comprising a plurality of bits, wherein each bit of the plurality of bits comprising the additional segment header bitmap SPS indicates a value of a corresponding syntax element presence flag, wherein each value indicates either the presence or non-presence of the corresponding syntax element in a segment header,wherein a first bit in the additional segment header bitmap SPS corresponds to a first syntax element and the first bit indicates the absence of the first syntax element in the segment header, and wherein a second bit in the additional segment header bitmap SPS corresponds to a second syntax element, the second bit indicates the presence of the second syntax element in the segment header, and the second bit follows the first bit in the additional segment header bitmap SPS; and encode the segment header, which includes a set of one or more additional segment header bits and includes a first additional segment header bit in a position that is first in the set of one or more additional segment header bits, wherein the first bit and the second bit indicate that the first additional segment header bit applies to the second syntax element.

15. The encoding apparatus according to claim 14, characterized in that the first bit indicating the absence of the first syntax element in the segment header and the second bit indicating the presence of the second syntax element in the segment header, together indicate that the additional first segment header bit applies to the second syntax element.

16. The encoding apparatus according to claim 14, characterized in that the segment header is not for a dependent segment, and the additional first segment header bit is included in the segment header on the basis that the segment header is not for a dependent segment.

17. A coding method, characterized in that it comprises: encoding a sequence parameter set (SPS) in a video data stream, the SPS including an additional segment header bitmap SPS comprising a plurality of bits, wherein each bit of the plurality of bits comprised by the additional segment header bitmap SPS indicates a value of a corresponding syntax element presence flag, wherein each value Mταζηη / ζζηζ / Β / γίΛΐ indicates either the presence or non-presence of the corresponding syntax element in a segment header, wherein a first bit in the additional segment header bitmap SPS corresponds to a first syntax element and the first bit indicates the non-presence of the first syntax element in the segment header, and wherein a second bit in the additional segment header bitmap SPS corresponds to a second syntax element,the second bit indicates the presence of the second syntax element in the segment header, and the second bit follows the first bit in the additional segment header bitmap SPS; and encode the segment header, which includes a set of one or more additional segment header bits and includes a first additional segment header bit in a position that is first in the set of one or more additional segment header bits, wherein the first bit and the second bit indicate that the first additional segment header bit applies to the second syntax element.

18. The encoding method according to claim 17, characterized in that the first bit indicating the absence of the first syntax element in the segment header and the second bit indicating the presence of the second syntax element in the segment header, together indicate that the additional first segment header bit applies to the second syntax element.

19. The encoding method according to claim 17, characterized in that the segment header is not for a dependent segment, and the additional first segment header bit is included in the segment header on the basis that the segment header is not for a dependent segment.