Decoding device, decoding method, encoding device, and encoding method

By indicating the alternative encoding unit type in the parameter set unit of the video data stream and determining the type of the video encoding unit using the encoding unit type identifier, the problem of inefficient video data stream merging and decoding in the prior art is solved, and efficient video data stream processing is achieved.

CN114600458BActive Publication Date: 2025-05-27FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080076836.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-03
Filing Date
2020-09-03
Publication Date
2025-05-27
Estimated Expiration
2040-09-03

AI Technical Summary

Technical Problem

In the prior art, when processing multiple encoded video bitstreams, it is difficult to efficiently merge and decode. Especially in virtual reality applications, it is necessary to frequently switch video clips of different resolutions, resulting in insolable encoding efficiency.

Method used

By indicating the alternative encoding unit type in the parameter set unit of the video data stream, the type of the video encoding unit is determined using the encoding unit type identifier, and merging and decoding are performed according to the type, rewriting and mapping of the NAL unit type is realized to improve the merging efficiency.

Benefits of technology

It improves the merging efficiency and encoding efficiency of video data streams, can efficiently handle the switching and merging of multiple video clips, and is suitable for high-demand applications such as virtual reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114600458B_ABST
    Figure CN114600458B_ABST
Patent Text Reader

Abstract

A video decoder, configured to decode a video including a plurality of pictures from a video data stream by decoding each picture of one or more video coding units within an access unit from the video data stream associated with the respective picture; read an alternative coding unit type from a parameter set unit of the video data stream; for each predetermined video coding unit, read a coding unit type identifier (100) from the respective video coding unit; check whether the coding unit identifier identifies a coding unit type among a first subset (102) of one or more coding unit types or a coding unit type among a second subset (104) of coding unit types, and if the coding unit identifier identifies a coding unit type among the first subset of one or more coding unit types, assign the respective video coding unit to the alternative coding unit type; if the coding unit identifier identifies a coding unit type among the second subset of coding unit types, assign the respective video coding unit to the coding unit type among the second subset of coding unit types identified by the coding unit identifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a data structure for indicating the coding unit type and characteristics of video coding units of a video data stream. Background Art

[0002] It is known that the picture type is indicated in the NAL unit header of the NAL unit carrying the picture slice. Thus, the basic attributes of the NAL unit payload are available for use by applications at a very high level.

[0003] The picture types include the following:

[0004] - Random Access Point (RAP) pictures, where the decoder can start decoding the encoded video sequence. These are called Intra Random Access Pictures (IRAP). There are three IRAP picture types: Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA). The decoding process of an encoded video sequence always starts at an IRAP.

[0005] - Leading pictures, which are before the random access point pictures in output order, but are encoded after it in the encoded video sequence. Leading pictures that are independent of the pictures before the random access point in the encoding order are called Random Access Decodable Leading pictures (RADL). Leading pictures that are predicted using pictures before the random access point in the encoding order may be corrupted if decoding starts at the corresponding IRAP. These are called Random Access Skipped Leading pictures (RASL).

[0006] - Trailing (TRAIL) pictures, which follow both the IRAP and the leading pictures in both output and display order.

[0007] - Pictures at which the decoder can switch the temporal resolution of the encoded video sequence: Temporal Sub-layer Access (TSA) and Stepwise Temporal Sub-layer Access (STSA).

[0008] Therefore, the data structure of the nal unit is an important factor in stream merging. Summary of the Invention

[0009] The object of the subject matter of this application is to provide a decoder for deriving the necessary information of video coding units of a video data stream by reading an identifier indicating an alternative coding unit type, and a decoder for deriving the characteristics of a video data stream.

[0010] Another object of the subject matter of this application is to provide an encoder for indicating an alternative coding unit type of a video coding unit by using an identifier and an encoder for indicating the characteristics of a video data stream.

[0011] This object is achieved by the subject matter of the claims of this application.

[0012] According to an embodiment of the present application, a video decoder is configured to decode a video including a plurality of pictures from a video data stream by decoding each picture of one or more video coding units within an access unit from the video data stream associated with the corresponding picture; read an alternative coding unit type from a parameter set unit of the video data stream; for each predetermined video coding unit, read a coding unit type identifier (100) from the corresponding video coding unit, such as a syntax element included in the NAL unit header; check whether the coding unit identifier identifies a coding unit type among a first subset (102) of one or more coding unit types - for example, indicating whether the NAL unit is a mappable VCL (Video Coding Layer) unit type, or identifies a coding unit type among a second subset (104) of coding unit types - for example, indicating the NAL unit type. If the coding unit identifier identifies a coding unit type among a first subset of one or more coding unit types, attribute the corresponding predetermined video coding unit to the alternative coding unit type; if the coding unit identifier identifies a coding unit type among a second subset of coding unit types, attribute the corresponding predetermined video coding unit to a coding unit type among the second subset of coding unit types identified by the coding unit identifier. That is, the corresponding NAL unit type is indicated by the identifier, the first subset of coding unit types, and the second subset of coding unit types, that is, the NAL unit type is rewritten after the indication of the first and second subsets of coding unit types. Therefore, it is possible to improve the merging efficiency.

[0013] According to an embodiment of the present application, the video decoder is configured to: decode a region associated with the corresponding video coding unit from each video coding unit in a manner depending on the coding unit type attributed to the corresponding video coding unit. The video decoder may be configured such that the alternative coding unit type is among the second subset of video coding types. The video decoder may be configured such that the alternative coding unit type is among a third subset of video coding types, such as non-VCL unit types, which include at least one video coding type not included in the second subset of video coding types. According to the present application, it is possible to improve the coding efficiency.

[0014] According to an embodiment of the present application, a predetermined video coding unit carries tile partition data, block-related prediction parameters, and prediction residual data. When a picture includes both one or more video coding units (e.g., slices) having a first subset of coding unit types and one or more video coding units (e.g., slices) having a second subset of coding unit types, the latter video coding units have a coding unit type equal to an alternative coding unit type. The alternative coding unit type is a random access point (RAP) coding type. The alternative coding unit type is a coding type different from the random access point RAP coding type. That is, the alternative coding unit type is identified, and video coding units having the same alternative coding unit type are merged, and thus, the merging efficiency is appropriately improved.

[0015] According to an embodiment of the present application, each predetermined video coding unit is associated with a different region of a picture, and the picture is associated with an access unit in which the corresponding predetermined video coding unit is located. A parameter set unit of a video data stream has a scope covering a picture sequence, a picture, or a set of slices within a picture. The parameter set unit indicates the alternative coding unit type in a video data stream profile-specific manner. That is, it is possible to efficiently merge slices, and thus improve the coding efficiency.

[0016] According to an embodiment of the present application, the parameter set unit of the video data stream is: a parameter set unit having a scope covering a picture sequence; or an access unit delimiter having a scope covering one or more pictures associated with an access unit. That is, the order of pictures is appropriately indicated, and thus, it is possible to efficiently decode the pictures that need to be rendered.

[0017] According to an embodiment of the present application, the parameter set unit indicates the alternative coding unit type in the video data stream, and whether the predetermined video coding unit is a refresh start point for a video sequence for decoding a video, such as the RAP type, i.e., including instantaneous decoding refresh IDR, or a continuous start point for a video sequence for decoding a video, such as a non-RAP type, i.e., not including IDR. That is, it is possible to indicate whether a coding unit is the first picture of a video sequence by using the parameter set unit.

[0018] According to an embodiment of the present application, a video decoder is configured to decode a video including a plurality of pictures from a video data stream by decoding each picture of one or more video coding units within an access unit from the video data stream associated with a corresponding picture, wherein each video coding unit carries tile partition data, block-related prediction parameters, and prediction residual data, and is associated with a different region of the picture, the picture being associated with the access unit in which the corresponding predetermined video coding unit is located; read an n - tuple set of a plurality of syntax elements, such as two flags, each being 2 - tuple, from each of the predetermined video coding units such that the pair is a 4 - tuple mapping (200), for example, the mapping may default to be fixed; alternatively, it is signaled in the data stream, or both, by: splitting the value range, the n - tuple set of a plurality of syntax elements into an m - tuple set of one or more characteristics (202), such as three binary characteristics, each binary characteristic thus being 2 - tuple, such that the triple is 8 - tuple, each characteristic being described in a redundant manner in the case of the corresponding data in the predetermined video coding unit, that is, the characteristic can be deduced from an examination of deeper - encoded data regarding how the video is encoded relative to the picture into the video data stream, the picture being associated with the access unit in which the predetermined video coding unit is located, where m>n; or read N syntax elements (210) from each of the predetermined video coding units, such as N = 2 flags, each being 2 - tuple, where N>0, read association information from the video data stream, associate depending on the association information, that is, treat them as variables of associated characteristics, each of the N syntax elements having information about one of the M characteristics, such as M = 3 binary characteristics, thus each being 2 - tuple → the association information will have 3 possibilities to associate the two flags with 2 of the 3 characteristics, that is associated with the characteristics, each characteristic being described in a redundant manner in the case of the corresponding data in the predetermined video coding unit regarding how the video is encoded relative to the picture into the video data stream, the picture being associated with the access unit in which the predetermined video coding unit is located, where M>N. That is, for example, the video data stream condition - that is, how the video is encoded relative to the picture in the access unit into the video data stream - is indicated by the mapping and the flags, and it is possible to efficiently provide additional information.

[0019] According to an embodiment of the present application, the mapping is included in a parameter set unit and indicates the position of the mapped characteristics. The mapping is signaled in the data stream and indicates the position of the mapped characteristics. The N syntax elements indicate the existence of the characteristics. That is, by combining the flags and the mapping, there is flexibility in indicating the flags at the parameter set.

[0020] According to an embodiment of the present application, a video encoder is configured to encode a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with the corresponding picture; indicate an alternative coding unit type in a parameter set unit of the video data stream; for each predetermined video coding unit, encode a coding unit type identifier (100) of the corresponding video coding unit into the video data stream, where the coding unit identifier identifies a coding unit type among a first subset (102) of one or more coding unit types or among a second subset (104) of coding unit types, wherein if the coding unit identifier identifies a coding unit type among the first subset of one or more coding unit types, the corresponding predetermined video coding unit will be assigned to the alternative coding unit type; if the coding unit identifier identifies a coding unit type among the second subset of coding unit types, the corresponding predetermined video coding unit will be assigned to a coding unit type among the second subset of coding unit types identified by the coding unit identifier, wherein the alternative coding unit type is a RAP type, and the video encoder is configured to identify video coding units of RAP pictures as predetermined video coding units, and for example, directly encode a coding unit type identifier for a pure intra-coded video coding unit that identifies a non-RAP picture of the RAP type. That is, the coding unit type is indicated in the parameter set unit of the video data stream, and thus, it is possible to improve the coding efficiency, that is, it is not necessary to encode each slice in the case of an IDR picture.

[0021] According to an embodiment of the present application, a video synthesizer is configured to synthesize a video data stream having a video that includes a plurality of pictures encoded therein, each picture being divided into one or more video coding units within an access unit of the video data stream, the one or more video coding units being associated with a respective picture of each tile into which the picture is subdivided; change an alternative coding unit type in a parameter set unit of the video data stream from indicating a RAP type to indicate a non-RAP type; identify, in a vds picture, a specifically coded video coding unit whose identifier (100) is encoded into the video data stream, the coding unit type identifying a RAP picture; wherein for each predetermined video coding unit of the video data stream, the identifier (100) of the corresponding p video coding unit is encoded into the video data stream, the coding unit type identifying a coding unit type among a first subset (102) of one or more coding unit types or among a second subset (104) of coding unit types, wherein if the coding unit identifier identifies a coding unit type among the first subset of one or more coding unit types, the corresponding predetermined video coding unit will belong to the alternative coding unit type; if the coding unit identifier identifies a coding unit type among the second subset of coding unit types, the corresponding predetermined video coding unit will belong to the coding unit type among the second subset of coding unit types identified by the coding unit identifier. By using the identifier and the first and second subsets of coding unit types to identify the types of video coding units, and thus, for example, a video picture constructed from a plurality of tiles is efficiently synthesized.

[0022] According to an embodiment of the present application, a video encoder is configured to encode a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with the corresponding picture, wherein each video coding unit carries tile partition data, block-related prediction parameters, and prediction residual data, and is associated with a different region of the picture associated with the access unit in which the corresponding predetermined video coding unit is located; indicating an n-ary set of one or more syntax elements into each of the predetermined video coding units, such as two flags, each being 2-ary, such that the pair is a 4-ary mapping (200), the mapping may be fixed by default; alternatively, it is signaled in the data stream, or both, by: splitting a value range, an n-ary set of one or more syntax elements into an m-ary set of one or more characteristics (202), such as three binary characteristics, each binary characteristic thus being 2-ary, such that the triplet is 8-ary, each characteristic being described in a redundant manner in the case of the corresponding data in the predetermined video coding unit, i.e., the characteristic can be derived from an examination of deeper encoded data regarding how the video is encoded relative to the picture into the video data stream, the picture being associated with the access unit in which the predetermined video coding unit is located, where m > n; or indicating N syntax elements (210) into each of the predetermined video coding units, such as N = 2 flags, each being 2-ary, where N > 0, indicating association information into the video data stream, associating depending on the association information, i.e., treating them as variables of associated characteristics, each of the N syntax elements having information regarding one of the M characteristics, such as M = 3 binary characteristics, thus each being 2-ary → the association information will have 3 possibilities to associate two flags with 2 of the 3 characteristics, i.e., associated with two characteristics, each characteristic being described in a redundant manner in the case of the corresponding data in the predetermined video coding unit regarding how the video is encoded relative to the picture into the video data stream, the picture being associated with the access unit in which the predetermined video coding unit is located, where M > N. That is, for example, the characteristics of each video coding unit of an encoded video sequence are indicated by using flags, and thus, it is possible to efficiently provide additional information.

[0023] According to an embodiment of the present application, a method includes: decoding a video including a plurality of pictures from a video data stream by decoding each picture in one or more video coding units within an access unit of the video data stream associated with a corresponding picture; reading an alternative coding unit type from a parameter set unit of the video data stream; for each predetermined video coding unit, reading a coding unit type identifier from the corresponding video coding unit; checking whether the coding unit identifier identifies a coding unit type among a first subset of one or more coding unit types or a coding unit type among a second subset of coding unit types, and if the coding unit identifier identifies a coding unit type among a first subset of one or more coding unit types, attributing the corresponding video coding unit to the alternative coding unit type; if the coding unit identifier identifies a coding unit type among a second subset of coding unit types, attributing the corresponding video coding unit to a coding unit type among the second subset of coding unit types identified by the coding unit identifier.

[0024] According to an embodiment of the present application, a method includes: decoding a video including a plurality of pictures from a video data stream by decoding each picture in one or more video coding units within an access unit of the video data stream associated with a corresponding picture, wherein each video coding unit carries tile partition data, block-related prediction parameters, and prediction residual data and is associated with a different region of a picture associated with the access unit in which the corresponding predetermined video coding unit is located; reading an n-ary set of a plurality of syntax elements, such as two flags, each being 2-ary, from each of the predetermined video coding units such that the pair is a 4-ary mapping, the mapping being default fixed; alternatively, it is signaled in the data stream, or both, by: splitting a value range, an n-ary set of a plurality of syntax elements into an m-ary set of one or more characteristics, such as three binary characteristics, each binary characteristic thus being 2-ary such that the triple is 8-ary, each characteristic describing in a redundant manner in the case of the corresponding data in the predetermined video coding unit - that is, the characteristic can be derived from an examination of deeper encoded data - regarding how the video is encoded relative to the picture into the video data stream, the picture being associated with the access unit in which the predetermined video coding unit is located, where m > n; or reading N syntax elements, such as N = 2 flags, each being 2-ary, from each of the predetermined video coding units, where N > 0, reading association information from the video data stream, associating depending on the association information, that is, treating them as variables of association characteristics, each of the N syntax elements having information about one of the M characteristics, such as M = 3 binary characteristics, each thus being 2-ary → the association information will have 3 possibilities to associate the two flags with 2 out of 3 characteristics, that is associated with each characteristic, each characteristic describing how video is encoded into a video data stream relative to a picture redundantly in the case of corresponding data in a predetermined video coding unit, the picture being associated with an access unit in which the predetermined video coding unit is located, where M > N.

[0025] According to an embodiment of the present application, a method includes: encoding a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with the corresponding picture; indicating an alternative coding unit type in a parameter set unit of the video data stream; for each predetermined video coding unit, defining a coding unit type identifier (100) of the corresponding video coding unit, where the coding unit identifier identifies a coding unit type among a first subset (102) of one or more coding unit types or among a second subset (104) of coding unit types, and if the coding unit identifier identifies a coding unit type among the first subset of one or more coding unit types, then the corresponding video coding unit belongs to the alternative coding unit type; if the coding unit identifier identifies a coding unit type among the second subset of coding unit types, then the corresponding video coding unit belongs to a coding unit type among the second subset of coding unit types identified by the coding unit identifier.

[0026] According to an embodiment of the present application, a method includes: synthesizing a video data stream having a video, the video including a plurality of pictures encoded therein, each picture being divided into one or more video coding units within an access unit of the video data stream, the one or more video coding units being associated with the corresponding picture of each tile into which the picture is divided; changing an alternative coding unit type in a parameter set unit of the video data stream from indicating a RAP type to indicating a non - RAP type; identifying, in a vds picture, a specially - coded video coding unit whose identifier (100) is encoded into the video data stream, the coding unit type identifying a RAP picture; where for each predetermined video coding unit of the video data stream, the identifier (100) of the corresponding p video coding unit is encoded into the video data stream, the coding unit type identifying a coding unit type among a first subset (102) of one or more coding unit types or among a second subset (104) of coding unit types, where if the coding unit identifier identifies a coding unit type among the first subset of one or more coding unit types, then the corresponding predetermined video coding unit will belong to the alternative coding unit type; if the coding unit identifier identifies a coding unit type among the second subset of coding unit types, then the corresponding predetermined video coding unit will belong to a coding unit type among the second subset of coding unit types identified by the coding unit identifier.

[0027] According to an embodiment of the present application, a method includes: encoding a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with the corresponding picture, wherein each video coding unit carries tile partition data, block-related prediction parameters, and prediction residual data, and is associated with a different region of the picture associated with the access unit in which the corresponding predetermined video coding unit is located; indicating an n-ary set of one or more syntax elements into each of the predetermined video coding units, such as two flags, each being 2-ary, such that the pair is a 4-ary mapping (200), which mapping may be fixed by default; alternatively, it is signaled in the data stream, or both, by: splitting a value range, an n-ary set of one or more syntax elements into an m-ary set of one or more characteristics (202), such as three binary characteristics, each binary characteristic thus being 2-ary, such that the triplet is 8-ary, each characteristic describing redundantly in the case of the corresponding data in the predetermined video coding unit - i.e., the characteristic can be deduced from an examination of deeper encoded data - regarding how the video is encoded relative to the picture into the video data stream, the picture being associated with the access unit in which the predetermined video coding unit is located, where m > n; or indicating N syntax elements (210) into each of the predetermined video coding units, such as N = 2 flags, each being 2-ary, where N > 0, indicating association information into the video data stream, associating depending on the association information, i.e., treating them as variables of associated characteristics, each of the N syntax elements having information about one of the M characteristics, such as M = 3 binary characteristics, and thus each being 2-ary → the association information will have 3 possibilities of associating the two flags with 2 of the 3 characteristics, i.e., characteristics, each characteristic describing redundantly in the case of the corresponding data in the predetermined video coding unit regarding how the video is encoded relative to the picture into the video data stream, the picture being associated with the access unit in which the predetermined video coding unit is located, where M > N. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Preferred embodiments of the present application are described below with reference to the various figures, wherein:

[0029] Figure 1 FIG. shows a schematic diagram illustrating a client and server system for virtual reality applications, as an example of where the embodiments set forth in the following figures can be advantageously used;

[0030] FIGS. 2(a) and 2(b) show schematic illustrations of examples of 360-degree videos indicating in two resolutions in a cube map projection and tiled into 6×4 tiles of a system that can be assembled into Figure 1 the system.

[0031] FIG. 3(a), FIG. 3(b), and FIG. 3(c) show schematic illustrations indicating a user's viewpoint and tile selection of a 360-degree video stream as indicated in FIGS. 2(a) and 2(b);

[0032] Figure 4 show a schematic illustration indicating an example of a resulting tile arrangement (encapsulation) of the tiles indicated in FIGS. 3(a), 3(b), and 3(c) in a combined bitstream after a merge operation;

[0033] FIG. 5(a), FIG. 5(b), and FIG. 5(c) show schematic illustrations indicating examples of tiling with low-resolution fallback as individual tiles of 360-degree video streaming;

[0034] Figure 6 show a schematic illustration indicating an example of tile selection in tile-based streaming;

[0035] Figure 7 show a schematic illustration indicating another example of tile selection in tile-based streaming, having unequal RAP (Random Access Point) periods for each tile;

[0036] Figure 8 show a schematic diagram indicating an example of an NAL (Network Adaptation Layer) unit header according to an embodiment of the present application;

[0037] Figure 9 show an example of a table indicating the type of RBSP (Raw Byte Sequence Payload) data structure included in an NAL unit according to an embodiment of the present application;

[0038] Figure 10 show a schematic diagram indicating an example of a sequence parameter set to which an NAL unit type is mapped according to an embodiment of the present application;

[0039] Figure 11 show a schematic illustration indicating an example of characteristics of a slice indicated in a slice header;

[0040] Figure 12 show a schematic diagram according to an embodiment of the present application, which indicates an example of a mapping indicating characteristics of a slice in a sequence parameter set;

[0041] Figure 13 show according to an embodiment of the present application an indication Figure 12 of an example of characteristics indicated by a mapping in a sequence parameter set;

[0042] Figure 14 show a schematic diagram according to an embodiment of the present application, which indicates an example of associated information indicated by using a syntax structure;

[0043] Figure 15 A schematic diagram showing another example of a mapping indicating slice characteristics in a sequence parameter set according to an embodiment of the present application;

[0044] Figure 16 A schematic diagram showing an indication according to an embodiment of the present application Figure 15 Another example of a characteristic indicated by a mapping in a parameter set;

[0045] Figure 17 A schematic diagram showing another example of a mapping indicating slice characteristics in a slice segment header according to an embodiment of the present application; and

[0046] Figure 18 A schematic diagram showing an indication according to an embodiment of the present application Figure 17 Another example of a characteristic indicated by a mapping in a slice segment header. Detailed Description

[0047] In the following description, the same or equivalent elements or elements having the same or equivalent functions are denoted by the same or equivalent reference numerals.

[0048] In the following description, a number of details are set forth to provide a more thorough explanation of embodiments of the present application. However, it will be clear to those skilled in the art that the embodiments of the present application may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, to avoid obscuring the embodiments of the present application. Additionally, features of different embodiments described below may be combined with each other unless otherwise specifically noted.

[0049] Introduction

[0050] In the following, it should be noted that the individual aspects described herein may be used alone or in combination. Thus, details may be added to each of the individual aspects without adding details to another of the aspects.

[0051] It should also be noted that the present disclosure explicitly or implicitly describes features that may be used in a video decoder (a device for providing a decoded representation of a video signal based on an encoded representation). Thus, any feature described herein may be used in the context of a video decoder.

[0052] Furthermore, the features and functions related to methods disclosed herein may also be used in a device (configured to perform such functions). Additionally, any feature and function regarding a device disclosed herein may also be used in a corresponding method. In other words, the methods disclosed herein may be supplemented by any characteristics and functions described regarding the device.

[0053] To facilitate understanding of the description of embodiments of the present application with respect to various aspects of the present application, Figure 1 an example of an environment in which the subsequently described embodiments of the present application can be applied and advantageously used is shown. In particular, Figure 1 a system consisting of a client 10 and a server 20 is shown, where the client 10 and the server 20 interact via adaptive streaming. For example, Dynamic Adaptive Streaming over HTTP (DASH) can be used for the communication 22 between the client 10 and the server 20. However, the embodiments outlined subsequently should not be construed as limited to the use of DASH, and similarly, terms such as Media Presentation Description (MPD) should be understood broadly so as to also cover manifest files defined differently in DASH.

[0054] Figure 1 illustrates a system configured to implement a virtual reality application. That is, the system is configured to present a view segment 28 within a time-varying spatial scene 30 to a user wearing a head-up display 24, i.e., via an internal display 26 of the head-up display 24, where the segment 28 corresponds to the orientation of the head-up display 24 exemplarily measured by an internal orientation sensor 32 such as an inertial sensor of the head-up display 24. That is, the segment 28 presented to the user forms a segment of the spatial scene 30 whose spatial position corresponds to the orientation of the head-up display 24. In Figure 1 the case where, the time-varying spatial scene 30 is depicted as an omnidirectional video or a spherical video, but Figure 1 the description and the embodiments explained subsequently can also be easily transferred to other examples, such as presenting a segment within a video, where the spatial position of the segment 28 is determined by the intersection of a face entry or an eye entry with a virtual or real projector wall, etc. Additionally, the sensor 32 and the display 26 can be included, for example, by different devices such as a remote control and a corresponding television, or they can be part of a handheld device such as a mobile device (such as a tablet computer or a mobile phone). Finally, it should be noted that some of the embodiments described later can also be applied to the scenario where the segment 28 presented to the user continuously covers the entire time-varying spatial scene 30, where the non-uniformity in presenting the time-varying spatial scene is related to, for example, a non-uniform distribution of quality over the spatial scene.

[0055] Additional details regarding the server 20, the client 10, and the manner of providing the spatial scene 30 at the server 20 are illustrated in Figure 1 and described below. However, these details should not be regarded as limiting the embodiments explained subsequently, but rather as examples of how to implement any of the embodiments explained subsequently.

[0056] In particular, as Figure 1As shown in FIG. 1 , the server 20 may include a storage device 34 and a controller 36, such as a suitably programmed computer, an application specific integrated circuit, etc. The storage device 34 has stored thereon media clips representing the time-varying spatial scene 30. Figure 1 The controller 36 responds to the request sent by the client 10 by resending the requested media segments, the media presentation description to the client 10, and may send additional information to the client 10 itself. Details of this are also set out below. The controller 36 may retrieve the requested media segments from the storage device 34. In this storage device, other information, such as the media presentation description or parts thereof, may also be stored in other signals sent from the server 20 to the client 10.

[0057] like Figure 1 As shown in , the server 20 may optionally further include a stream modifier 38 which, in response to a request from the client 10, modifies the media segments sent from the server 20 to the client 10 so as to produce a media data stream at the client 10 which forms a single media stream decodable by one associated decoder, even though, for example, the media segments retrieved in this way by the client 10 are actually aggregated from several media streams. However, the presence of such a stream modifier 38 is optional.

[0058] Figure 1 The client 10 of is exemplarily depicted as comprising a controller or client device 40 or a plurality of decoders 42 and a reprojector 44. The client device 40 may be a suitably programmed computer, a microprocessor, a programmed hardware device such as an FPGA or an application specific integrated circuit. The client device 40 bears the responsibility of selecting the segments to be retrieved from the server 20 from among a plurality of media segments 46 provided at the server 20. To this end, the client device 40 first retrieves a manifest or a media presentation description from the server 20. From the server 20, the client device 40 obtains a calculation rule for calculating the addresses of media segments from among the plurality of media segments 46, which media segments correspond to certain desired spatial portions of the spatial scene 30. The client device 40 retrieves the media segments thus selected from the server 20 by sending corresponding requests to the server 20. These requests contain the calculated addresses.

[0059] The media segments thus retrieved by the client device 40 are forwarded by the latter to one or more decoders 42 for decoding. Figure 1In the example, for each time unit, the media segments retrieved and decoded in this way only represent a spatial section 48 within the time-varying spatial scene 30, but as already pointed out above, this can vary according to other aspects, where for example the view section 28 to be presented continuously covers the entire scene. The reprojection unit 44 can optionally reproject and clip out the view section 28 to be displayed to the user among the retrieved and decoded scene content of the selected, retrieved, and decoded media segments. For this purpose, as Figure 1 shown in, the client device 40 can continuously track and update the spatial position of the view section 28, for example in response to user orientation data from the sensor 32, and for example notify the reprojection unit 44 of this current spatial position of the scene section 28 and the reprojection mapping to be applied to the retrieved and decoded media content in order to be mapped onto the area forming the view section 28. Thus, the reprojection unit 44 can apply the mapping and interpolation to, for example, a regular pixel grid to be displayed on the display 26.

[0060] Figure 1 Illustrates the case where the spatial scene 30 has been mapped onto the tiles 50 using cube mapping. Thus, the tiles are depicted as rectangular sub-regions of a cube onto which the scene 30 in the form of a sphere has been projected. The reprojection unit 44 reverses this projection. However, other examples can also be applied. For example, a projection onto a frustum or an untruncated pyramid can be used instead of the cube projection. Furthermore, although Figure 1 the tiles are depicted as not overlapping in terms of the coverage of the spatial scene 30, the subdivision into tiles may involve mutual tile overlap. And as will be outlined in more detail below, as further explained below, it is not mandatory to spatially subdivide the scene 30 into tiles 50, where each tile forms a representation.

[0061] Thus, as Figure 1 depicted in, the entire spatial scene 30 is spatially subdivided into tiles 50. In Figure 1 the example, each of the six faces of the cube is subdivided into 4 tiles. For illustrative purposes, the tiles are enumerated. For each tile 50, the server 20 provides a video 52 as Figure 1 depicted in. More precisely, the server 20 even provides more than one video 52 for each tile 50, which differ in terms of the quality Q#. Furthermore, the videos 52 are temporally subdivided into time segments 54. The time segments 54 of all the videos 52 of all the tiles T# respectively form or are encoded into one of the media segments 46 of the plurality of media segments stored in the storage device 34 of the server 20.

[0062] It is again emphasized that even Figure 1The example of tile - based streaming illustrated in the figure is only one example, and many deviations from this example are possible. For example, although Figure 1 seems to imply that the media segments related to the representation of scene 30 at a higher quality involve tiles that are consistent with the tiles to which the media segment of scene 30 encoded with quality Q1 belongs, such a consistency is not necessary, and tiles of different qualities can even correspond to tiles of different projections of scene 30. In addition, although not discussed so far, the media segments corresponding to Figure 1 the different quality levels depicted in may differ in terms of spatial resolution and / or signal - to - noise ratio and / or temporal resolution, etc.

[0063] Finally, different from the tile - based streaming concept - according to which the media segments that can be retrieved individually by the client device 40 from the server 20 are related to the tiles 50 into which scene 30 is spatially subdivided - the media segments provided at the server 20 can alternatively, for example, each have scene 30 encoded into them in a spatially complete manner with a spatially varying sampling resolution, however, having the maximum sampling resolution at different spatial positions in scene 30. For example, this can be achieved by providing at the server 20 a sequence of segments 54 related to the projections of scene 30 on a frustum, the apexes of which frustum will be oriented in mutually different directions, resulting in resolution peaks of different orientations.

[0064] In addition, regarding the optionally present stream modifier 38, it should be noted that the stream modifier 38 can alternatively be part of the client 10, or the stream modifier 38 can even be located within the network device through which the client 10 and the server 20 exchange the signals described herein, between the client 10 and the server 20.

[0065] There are certain video - based applications where multiple encoded video bitstreams will be jointly decoded, i.e., merged into a joint bitstream and fed into a single decoder, such as:

[0066] · Multi - party conferencing: where the encoded video streams from multiple participants are processed at a single endpoint

[0067] · Or tile - based streaming: For example, for 360 - degree tiled video playback in VR applications. In the latter case, the 360 - degree video is spatially segmented, and each spatial segment is provided to the streaming client with multiple representations at different spatial resolutions, as illustrated in FIGS. 2(a) and 2(b). FIG. 2(a) shows high - resolution tiles, and FIG. 2(b) shows low - resolution tiles. FIGS. 2(a) and (b) show a cube - map projected 360 - degree video divided into 6×4 spatial segments at two resolutions. For simplicity, these independently decodable spatial segments are referred to as tiles in this specification.

[0068] When using a prior - art head - mounted display, the user typically views only a subset of the tiles that make up the entire 360 - degree video through a stereoscopic viewport boundary 80 representing a 90×90 - degree field of view, as illustrated in FIG. 3(a). The corresponding tiles are indicated by reference numeral 82 in FIG. 3(b) and are downloaded at the highest resolution.

[0069] However, the client application will also have to download and decode representations of other tiles outside the current viewport, as indicated by reference numeral 84 in FIG. 3(c), in order to handle sudden orientation changes of the user. As indicated in FIG. 3(c), the client in such an application will thus download the tiles covering its current viewport at the highest resolution and download the tiles outside its current viewport at a relatively lower resolution, while the tile - resolution selection continuously adapts to the user's orientation. After downloading on the client side, merging the downloaded tiles into a single bit - stream for processing with a single decoder is a means to address the constraints of typical mobile devices with limited computational and power resources. Figure 4 Illustrated is a possible tile arrangement in the combined bit - stream of the above example. The merging operation to generate the combined bit - stream must be carried out through compressed - domain processing, i.e., avoiding processing the pixel domain through transcoding.

[0070] Although the examples from Figure 4 illustrate the case where all tiles (high - resolution and low - resolution) cover the entire 360 - degree space and no tile overlaps the same area repeatedly, another tile grouping can also be used as depicted in FIGS. 5(a), 5(b), and 5(c). It defines the entire low - resolution portion of the video as a "low - resolution fallback" layer, as indicated in FIG. 5(b), which can be merged with the high - resolution tiles of FIG. 5(a) covering a subset of the 360 - degree video. As indicated in FIG. 5(c), the entire low - resolution fallback video can be encoded as a single tile, while the high - resolution tiles are rendered as an overlay on the low - resolution portion of the video in the final stage of the rendering process.

[0071] As Figure 6As illustrated, the client starts a streaming session according to his tile selection by downloading all desired tile tracks, where the client starts the session from tile 0 indicated by reference numeral 90 in (a) of Figure 6 and tile 1 indicated by reference numeral 92. Whenever a viewport change occurs (i.e., the user turns his head to look in another direction), the tile selection is changed at the next occurring time segment, i.e., tile 0 and tile 2 indicated by reference numeral 94 in (a) of Figure 6 , and at the next available segment, the client changes the position of tile 2 and replaces tile 0 with tile 1, as indicated in (b) of Figure 6 . It is important to note that for any new tile selection and change in the position of the tile, all segments need to start with an IDR (Instantaneous Decoder Refresh) picture - i.e., a picture that resets the prediction chain, otherwise it will cause prediction mismatches, artifacts, and drift.

[0072] In terms of bitrate, encoding each segment in the case of an IDR picture is costly. The segments can potentially be very short in duration, e.g., to quickly respond to orientation changes, which is why it is desirable to encode multiple variants with varying IDR (or RAP: Random Access Point) periods, as illustrated in Figure 7 . For example, as indicated in (b) of Figure 7 , at time instance t 1 , there is no reason to interrupt the prediction chain of tile 0 because tile 0 has already been downloaded and placed in the same position at time instance t 0 , which is why the client can choose a segment that does not start with a RAP available at the server.

[0073] However, one problem that still remains is that the slices (tiles) within the encoded picture are subject to certain constraints. One of them is that the picture cannot contain both RAP and non - RAP NAL (Network Abstraction Layer) unit types of NAL units at the same time. Thus, for the application, only two less desirable options exist to solve the above problem. First, the client can rewrite the NAL unit type of the RAP picture when the RAP picture is merged with non - RAP NAL units into a picture. Second, the server can obscure the RAP characteristics of these pictures from the start by using non - RAP. However, this hinders the detection of the RAP characteristics in systems that process these encoded videos, e.g., for file format encapsulation.

[0074] The present invention is a NAL unit type mapping that allows mapping one NAL unit type to another through a syntax structure that can be easily rewritten.

[0075] In one embodiment of the present invention, the NAL unit type is designated as mappable, and the mapped type is designated in the parameter set - for example, as follows based on Draft 6 V14 of the VVC (Versatile Video Coding) specification - with highlighted edits.

[0076] Figure 8 The NAL unit header syntax is shown. The syntax nal_unit_type (i.e., identifier 100) specifies the NAL unit type, i.e., the type of the RBSP (Raw Byte Sequence Payload) data structure contained in the NAL unit as specified in the table indicated as follows Figure 9 in the table.

[0077] The variable NalUnitType is defined as follows:

[0078] When nal_unit_type != MAP_NUT

[0079] NalUnitType is equal to nal_unit_type

[0080] Otherwise (nal_unit_type == MAP_NUT)

[0081] NalUnitType is equal to mapped_nut

[0082] All references to the syntax element nal_unit_type in this specification are replaced with references to the variable NalUnitType, such as, for example, in the following constraints:

[0083] For all coded slice NAL units of a picture, the value of NalUnitType shall be the same. A picture or layer access unit is said to have the same NAL unit type as the coded slice NAL units of the picture or layer access unit. That is, as Figure 9 depicted, the first subset 102 of the coded unit types indicates "nal_unit_type" 12, i.e., "MAP_NUT" and "VCL" as NAL unit type classes. Thus, the second subset 104 of the coded unit types indicates "VCL" as the NAL unit type class, i.e., all coded slice NAL units of a picture indicated by the numbers 0 to 15 of identifier 100 have the same NAL unit type class as the coded unit types of the first subset 102 of the coded unit types, i.e., VCL.

[0084] Figure 10 The sequence parameter set RBSP syntax including mapped_nut is shown, as indicated by reference symbol 106, which indicates the NalUnitType of a NAL unit having a nal_unit_type equal to MAP_NUT.

[0085] In another embodiment, the mapped_nut syntax element is carried in the Access Unit Delimiter (AUD).

[0086] In another embodiment, the requirement for bitstream conformance is that the value of mapped_nut must be a VCL NAL unit type.

[0087] In another embodiment, the mapping of the NalUnitType of a NAL unit with nal_unit_type equal to MAP_NUT is effected by profile information. Such a mechanism can allow for more than one mappable NAL unit type instead of having a single MAP_NUT and indicate the required interpretation of the NALUnitTypes of the mappable NAL units within a simple profile mechanism or a single syntax element, mapped_nut_space_idc.

[0088] In another embodiment, a mapping mechanism is used to extend the value range of NALUnitTypes, which is currently limited to 32 (since it is u(5), as indicated, for example, Figure 10 ). The mapping mechanism can indicate any unrestricted value as long as the number of required NALUnitTypes does not exceed the number of values reserved for mappable NAL units.

[0089] In one embodiment, when a picture contains both slices of an alternative coding unit type and slices of a regular coding unit type (e.g., existing NAL units of the VCL category), the mapping is effected in such a way that all slices of the picture effectively have the same coding unit type attribute, i.e., the alternative coding unit type is equal to the coding unit type of the non - alternative slices of the regular coding type. Additionally, the above - mentioned embodiment applies only to pictures with random access attributes or pictures without random access attributes.

[0090] In addition to the issues described regarding NAL unit types and NAL unit type scalability and corresponding solutions in the merge scenario, there are several video applications where system integration and transmission or manipulation require information related to the video and how the video is encoded, such as real - time adaptation.

[0091] Over the past few years, some general information has been established, which is widely used in the industry, clearly specified, and specific bit values are used for such purposes. Examples thereof are:

[0092] · Temporary ID at the NAL unit header

[0093] · NAL unit types, including IDR, CRA, TRAIL, … or SPS (Sequence Parameter Set), PPS (Picture Parameter Set), etc.

[0094] However, there are several scenarios where additional information can be helpful. Other types of NAL units are not widely used, but some usefulness has been found in some cases, such as BLA, partial RAP NAL units for sub-pictures, sub-layer non-reference NAL units, etc. If the above scalability mechanism is used, some of those NAL unit types can be implemented. However, another alternative is to use some fields within the slice header.

[0095] In the past, additional information was reserved at the slice header for indicating specific characteristics of the slice:

[0096] · discardable_flag: Specifies that the coded picture is not used as a reference picture for inter prediction and is not used as a source picture for inter-layer prediction.

[0097] · cross_layer_bla_flag: Affects the derivation of the output picture of hierarchical coding, where pictures before the RAP at the higher layer may not be output.

[0098] Similar mechanisms can be envisioned for upcoming video codec standards. However, one limitation of those mechanisms is that the defined flags occupy specific positions within the slice header. In the following, the use of those flags in HEVC is shown in Figure 11 as follows.

[0099] As seen above, the problem with such a solution is that the positions of the extra slice header bits are allocated progressively, and for applications using less information, the flags will likely appear at later positions, thus increasing the number of bits to be sent in the extra bits (e.g., "discardable_flag" and "cross_layer_bla_flag" in the case of HEVC).

[0100] Alternatively, following a similar mechanism described for NAL unit types, the mapping of the flags in the extra slice header bits in the slice header can be defined at the parameter set. An example is shown Figure 12 as follows.

[0101] Figure 12 Shows an example of a sequence parameter set including a mapping using "extra_slice_header_bits_mapping_space_idc" to indicate the mapping of associated information, i.e., the mapping indicated by reference symbol 200, which indicates the mapping space of the extra bits in the slice header.

[0102] Figure 13 Shown by Figure 12The mapping of bits indicated by “extra_slice_header_bits_mapping_space_idc” 200 to the presence of flags. As Figure 13 depicted in, the binary feature 202 is described in a redundant manner in the case of corresponding data in a predetermined video coding unit. In Figure 13 three binary features 202, namely “0”, “1” and “2”, are depicted. The number of binary features may vary depending on the number of flags.

[0103] In another embodiment, the mapping is implemented in a syntax structure (e.g., as depicted in Figure 14 ), which indicates the presence of a syntax element in the extra bits in the slice header (e.g., as depicted in Figure 11 ). That is, for example, as depicted in Figure 11 , there is conditional control over the presence of syntax elements in the extra bits in the slice header on the flag, namely “num_extra_slice_header_bits > i” and “i < num_extra_slice_header_bits; i++”. In Figure 11 , as explained above, each syntax element in the extra bits is placed in a specific position in the slice header. However, in this embodiment, syntax elements such as “discardable_flag”, “cross_layer_bla_flag” or “slice_reserved_flag[i]” do not have to occupy that specific position. Instead, when the first syntax element (e.g., “discardable_flag” in Figure 11 ) is indicated as not present when checking the condition on the value of a specific presence flag (e.g., “discardable_flag_present_flag” in Figure 14 ), the subsequent second syntax element, when present, occupies the position of the first syntax element in the slice header. Additionally, by indicating a flag such as “sps_extra_ph_bit_present_flag[i]”, the syntax element in the extra bits can be present in the picture header. Furthermore, the syntax structure (e.g., the number of syntax elements, i.e., the number of flags presented) indicates the presence / presentation of a specific feature or the number of syntax elements presented in the extra bits. That is, the number of a specific feature or syntax element presented in the extra bits is indicated by counting how many syntax elements (flags) are presented. In Figure 14 , the presence of each syntax element 210 is indicated by a flag. That is, each flag in 210 indicates the presence indicated by the slice header of a specific feature of a predetermined video coding unit. Additionally, as Figure 14as indicated in and corresponding to Figure 11 the additional syntax "[…] / / further flags" for the "slice_reserved_flag[i]" syntax element in the slice header is used as a placeholder indicating the presence / presentation of a syntax element in the additional bits, or as an indication of the presence / presentation of additional flags.

[0104] In another embodiment, as for example Figure 15 shown in, the flag type mapping is signaled for each additional slice header bit in the parameter set. As Figure 15 indicated in, the syntax "extra_slice_header_bit_mapping_idc" (i.e., mapping 200) is signaled in the sequence parameter set and indicates the location of the mapping characteristics.

[0105] Figure 16 shows the mapping of bits indicated by Figure 15 the "extra_slice_header_bits_mapping_space_idc" to the presence of flags. That is, in Figure 16 is depicted the binary characteristic 202 corresponding to the mapping 200 indicated in Figure 15 as indicated in.

[0106] In another embodiment, the slice header extension bits are replaced by idc signaling representing a certain combination of flag values, e.g., as Figure 17 shown in. As Figure 17 depicted in, the mapping 200 (i.e., "extra_slice_header_bit_idc") is indicated in the slice segment header, i.e., the mapping 200 indicates the presence of the characteristics shown in Figure 18 as shown in.

[0107] Figure 18 shows that the flag values (i.e., the binary characteristic 202) represented by a certain value of "extra_slice_header_bit_idc" are either signaled in the parameter set or predefined (known a priori) in the specification.

[0108] In one embodiment, the value space of "extra_slice_header_bit_idc" (i.e., the value space of mapping 200) is divided into two ranges. One range represents a priori known combinations of flag values, and one range represents combinations of flag values signaled in the parameter set.

[0109] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0110] The data stream of the present invention can be stored on a digital storage medium or can be transmitted on a transmission medium, such as a wireless transmission medium or a wired transmission medium, such as the Internet.

[0111] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or software. The implementation can be performed using a digital storage medium (e.g., a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory) on which an electronically readable control signal is stored, which cooperates with (or is capable of cooperating with) a programmable computer system such that the corresponding method is performed. Thus, the digital storage medium can be computer-readable.

[0112] Some embodiments according to the present invention include a data carrier having an electronically readable control signal, which is capable of cooperating with a programmable computer system such that one of the methods described herein is performed.

[0113] Generally, embodiments of the present application can be implemented as a computer program product having program code that is operable to perform one of the methods when the computer program product runs on a computer. For example, the program code can be stored on a machine-readable carrier.

[0114] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the methods described herein.

[0115] In other words, thus, embodiments of the method of the present invention are a computer program having program code that, when the computer program runs on a computer, is used to perform one of the methods described herein.

[0116] Therefore, further embodiments of the method of the present invention are a data carrier (or a digital storage medium, or a computer-readable medium) that includes a computer program recorded thereon for performing one of the methods described herein. The data carrier, the digital storage medium, or the recording medium is generally tangible and / or non-transitory.

[0117] Accordingly, a further embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured, for example, to be transmitted via a data communication connection, such as via the Internet.

[0118] A further embodiment includes a processing device, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.

[0119] A further embodiment includes a computer having installed thereon a computer program for performing one of the methods described herein.

[0120] A further embodiment according to the present invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver can be a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transmitting the computer program to the receiver.

[0121] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0122] The apparatus described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0123] The apparatus described herein or any component of the apparatus described herein can be implemented at least in part in hardware and / or software.

Claims

1. A decoding device, comprising a microprocessor and a memory, the memory including a computer program which, when executed by the microprocessor, causes the decoding device to: Receive a sequence parameter set (SPS) from a video data stream, the SPS including an SPS extra slice header bit map comprising a plurality of bits, wherein each of the plurality of bits included in the SPS extra slice header bit map indicates a value of a corresponding syntax element presence flag, and wherein each value indicates the presence or absence of a corresponding syntax element in a slice header; Parse each bit in the SPS extra slice header bit map from the SPS, wherein a first bit in the SPS extra slice header bit map corresponds to a first syntax element, and the first bit indicates the absence of the first syntax element in the slice header, and wherein a second bit in the SPS extra slice header bit map corresponds to a second syntax element, the second bit indicates the presence of the second syntax element in the slice header, and the second bit follows the first bit in the SPS extra slice header bit map; Receive a slice header from the video data stream; Parse a first extra slice header bit from the slice header, the first extra slice header bit being included at a first position in a set of one or more extra slice header bits; and Apply the first extra slice header bit to the second syntax element based on the first bit and the second bit.

2. The decoding device according to claim 1, wherein applying the first extra slice header bit to the second syntax element is based on the first bit indicating the absence of the first syntax element in the slice header and the second bit indicating the presence of the second syntax element in the slice header.

3. The decoding device according to claim 1, wherein, the computer program further causes the decoding device to determine the number of bits included in the SPS extra slice header bit map.

4. The decoding device according to claim 1, wherein, the computer program further causes the decoding device to determine the number of extra slice header bits included in the slice header based on the SPS extra slice header bit map.

5. The decoding device according to claim 1, wherein, the computer program further causes the decoding device to determine that the slice header is not for a dependent slice, and wherein, in response to the determination that the slice header is not for a dependent slice, perform the parsing of the first extra slice header bit from the slice header.

6. A decoding method: Receive a sequence parameter set (SPS) from a video data stream, the SPS including an SPS extra slice header bit map comprising a plurality of bits, wherein each of the plurality of bits included in the SPS extra slice header bit map indicates a value of a corresponding syntax element presence flag, and wherein each value indicates the presence or absence of a corresponding syntax element in a slice header; Parse each bit in the SPS extra slice header bit map from the SPS, wherein a first bit in the SPS extra slice header bit mapping corresponds to a first syntax element, and the first bit indicates the absence of the first syntax element in the slice header, and wherein a second bit in the SPS extra slice header bit mapping corresponds to a second syntax element, the second bit indicates the presence of the second syntax element in the slice header, and the second bit follows the first bit in the SPS extra slice header bit mapping; receive a slice header from a video data stream; parse a first extra slice header bit from the slice header, the first extra slice header bit being included at a first position in a set of one or more extra slice header bits; and apply the first extra slice header bit to the second syntax element based on the first bit and the second bit.

7. The decoding method according to claim 6, wherein applying the first extra slice header bit to the second syntax element is based on the first bit indicating the absence of the first syntax element in the slice header and the second bit indicating the presence of the second syntax element in the slice header.

8. The decoding method according to claim 6, further comprising determining the number of bits included in the SPS extra slice header bit mapping.

9. The decoding method according to claim 6, further comprising determining the number of extra slice header bits included in the slice header based on the SPS extra slice header bit mapping.

10. The decoding method according to claim 6, further comprising determining that the slice header is not for a dependent slice, wherein, in response to the determination that the slice header is not for a dependent slice, perform parsing of the first extra slice header bit from the slice header.

11. An encoding apparatus, comprising a microprocessor and a memory, the memory including a computer program which, when executed by the microprocessor, causes the encoding apparatus to: encode a sequence parameter set SPS into a video data stream, the SPS including an SPS extra slice header bit mapping comprising a plurality of bits, wherein each of the plurality of bits included in the SPS extra slice header bit mapping indicates a value of a corresponding syntax element presence flag, wherein each value indicates the presence or absence of a corresponding syntax element in the slice header, wherein a first bit in the SPS extra slice header bit mapping corresponds to a first syntax element, and the first bit indicates the absence of the first syntax element in the slice header, and wherein a second bit in the SPS extra slice header bit mapping corresponds to a second syntax element, the second bit indicates the presence of the second syntax element in the slice header, and the second bit follows the first bit in the SPS extra slice header bit mapping; and encode the slice header, the slice header including a set of one or more extra slice header bits and including a first extra slice header bit at a first position in the set of one or more extra slice header bits, wherein the first bit and the second bit indicate that the first extra slice header bit is applied to the second syntax element.

12. The encoding device according to claim 11, wherein, a first bit indicating the non-existence of a first syntax element in the slice header and a second bit indicating the existence of a second syntax element in the slice header together indicate that a first extra slice header bit is applied to the second syntax element.

13. The encoding device according to claim 11, wherein, the slice header is not for a dependent slice, and based on the slice header not being for a dependent slice, a first extra slice header bit is included in the slice header.

14. An encoding method, comprising: encoding a sequence parameter set (SPS) into a video data stream, the SPS including an SPS extra slice header bit mapping including a plurality of bits, wherein each of the plurality of bits included in the SPS extra slice header bit mapping indicates a value of a corresponding syntax element presence flag, and each value indicates the existence or non-existence of a corresponding syntax element in the slice header, wherein a first bit in the SPS extra slice header bit mapping corresponds to a first syntax element, and the first bit indicates the non-existence of the first syntax element in the slice header, and wherein a second bit in the SPS extra slice header bit mapping corresponds to a second syntax element, the second bit indicates the existence of the second syntax element in the slice header, and the second bit follows the first bit in the SPS extra slice header bit mapping; and encoding the slice header, the slice header including a set of one or more extra slice header bits and including a first extra slice header bit at a first position in the set of one or more extra slice header bits, wherein the first bit and the second bit indicate that the first extra slice header bit is applied to the second syntax element.

15. The encoding method according to claim 14, wherein, a first bit indicating the non-existence of a first syntax element in the slice header and a second bit indicating the existence of a second syntax element in the slice header together indicate that a first extra slice header bit is applied to the second syntax element.

16. The encoding method according to claim 14, wherein, the slice header is not for a dependent slice, and based on the slice header not being for a dependent slice, a first extra slice header bit is included in the slice header.

Citation Information

Patent Citations

  • Video coding

    CN109691103A

  • Method for processing reference image, application processor, and mobile terminal

    JP2017123645A