Configurable nal and slice code point mechanism for stream merging

By using encoding unit type identifiers in the parameter set units of the video data stream and identifying and assigning different encoding unit types, the problem of low coding unit merging efficiency in the prior art is solved, and more efficient encoding and decoding fluency is achieved.

JP2025072660AActive Publication Date: 2025-05-09FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025024939
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-03
Filing Date
2025-02-19
Publication Date
2025-05-09
Estimated Expiration
2040-09-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently extract and merge different types of video encoding units from video data streams, resulting in limited encoding efficiency and fluency.

Method used

Efficient merging and decoding of encoding units is achieved by using encoding unit type identifiers in the parameter set units of the video data stream, and different encoding unit types, including random access point (RAP) encoding unit types and non-RAP encoding unit types.

Benefits of technology

Improve the encoding efficiency and decoding fluency of video data streams, and through effective encoding unit type management and merging, the dependence on Instantaneous Decoder Refresh (IDR) images is reduced, and the encoding cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072660000001_ABST
    Figure 2025072660000001_ABST
Patent Text Reader

Abstract

To provide a method for deriving necessary information and characteristics of a video coding unit of a video data stream by reading an identifier indicative of a substitute coding unit type, and a non-transitory computer readable medium.SOLUTION: A method includes: decoding each picture from one or more video coding units within an access unit of a video data stream to read a coding unit type identifier 100; checking whether the coding unit identifier identifies a coding unit type out of a first subset of one or more coding unit types 102 or out of a second subset of coding unit types 104; if the coding unit identifier identifies a coding unit type out of the first subset, attributing each of predetermined video coding units to substitute coding unit type, or if the coding unit identifier identifies a coding unit type out of the second subset, attributing each of the predetermined respective video coding units to the coding unit type out of the second subset of coding unit types identified by the coding unit identifier.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application relates to a data structure that indicates coding unit types and characteristics of video coding units of a video data stream. [Background technology]

[0002] It is known that the picture type is indicated in the NAL unit header of a NAL unit that carries a slice of a picture, whereby the essential characteristics of the NAL unit payload are available at a very high level for use by applications.

[0003] Picture types include the following: - Random Access Point (RAP) pictures: at which a decoder may start decoding a coded video sequence. These are referred to as Intra Random Access Pictures (IRAP). There are three IRAP picture types: Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA) and Broken Link Access (BLA). The decoding process of a coded video sequence always starts with an IRAP. - Leading picture: It precedes the random access point picture in output order but is coded after the random access point picture in the coded video sequence. Leading pictures that are independent of the pictures preceding the random access point in coding order are called Random Access Decodable Leading (RADL) pictures. Leading pictures that use pictures preceding the random access point in coding order for prediction may be corrupted if decoding starts at the corresponding IRAP. These are called Random Access Skipped Leading (RASL) pictures. - TRAIL picture: It follows the IRAP picture and the leading picture in both output order and display order. - Pictures for which the temporal resolution of a coded video sequence can be switched by the decoder: Temporal Sublayer Access and Step-by-Step Temporal Sublayer Access (STSA)

[0004] Therefore, the data structure of the NAL units is an important factor for stream merging. Summary of the Invention

[0005] The subject matter of the present application is to provide a decoder that derives necessary information of a video coding unit of a video data stream by reading an identifier indicating an alternative coding unit type, and a decoder that derives characteristics of the video data stream.

[0006] It is a further object of the present subject matter to provide an encoder that utilizes an identifier to indicate alternative coding unit types of video coding units and an encoder that indicates characteristics of a video data stream.

[0007] This object is achieved by the subject matter of the present claims.

[0008] According to an embodiment of the present application, a method is provided for decoding a video comprising a plurality of pictures from a video data stream by decoding each picture from one or more video coding units within an access unit of the video data stream associated with each of the pictures, reading alternative coding unit types from a parameter set unit of the video data stream, and for each given video coding unit, reading a coding unit type identifier (100) from each of the video coding units, such as a syntax element included in a nal unit header, and determining whether the coding unit identifier is a coding unit type identifier, such as a syntax element included in a Video Coding Class (VCL) to which the nal unit can be mapped. a first subset (102) of one or more coding unit types, e.g., indicating whether the coding unit identifier is a (Layer) unit type, or a coding unit type from a second subset (104) of coding unit types, e.g., indicating a nal unit type, and making each of the given video coding units the alternative coding unit type if the coding unit identifier identifies a coding unit type from the first subset of the one or more coding unit types, and making each of the given video coding units the coding unit type from the second subset of coding unit types identified by the coding unit identifier if the coding unit identifier identifies a coding unit type from the second subset of coding unit types. That is, each of the nal unit types is indicated by an identifier, the first subset of coding unit types and the second subset of coding unit types, i.e., the nal unit type is rewritten according to the indication of the first and second coding unit types. Thus, merging efficiency can be improved.

[0009] According to an embodiment of the present disclosure, a video decoder configured to encode, from each video coding unit, a region associated with each of the video coding units in a manner dependent on a coding unit type associated with each of the video coding units. The video decoder may be configured such that the alternative coding unit type is from a second subset of video coding types. The video decoder may be configured such that the alternative coding unit type is from a third subset of video coding types, such as a non-VCL unit type, that includes at least one video coding type not included by the second subset of video coding types.

[0010] According to an embodiment of the present disclosure, a given video coding unit carries picture block partitioning data, block-related prediction parameters, and prediction residual data. When a picture includes both one or more video coding units, such as a slice having a coding unit type of a first subset and a slice having a coding unit type of a second subset, the latter video coding unit has a coding unit type equal to an alternative coding unit type. The alternative coding unit type is a random access point (RAP) coding type. The alternative coding unit type is a coding type other than the random access point (RAP) coding type. That is, an alternative coding unit type is identified, and video coding units having the same alternative coding unit type are merged, thus improving merging efficiency appropriately.

[0011] According to an embodiment of the present application, each of the given video coding units is associated with a different region of the picture to which the access unit in which each of the given video coding units resides is associated. A parameter set unit of a video data stream has a scope covering a sequence of pictures, one picture or a set of slices from one picture. The parameter set unit indicates alternative coding unit types in a manner specific to the profile of the video data stream, i.e. slices can be efficiently merged, thus improving coding efficiency.

[0012] According to an embodiment of the present application, a parameter set unit of a video data stream is either a parameter set unit with a range covering an order of pictures, or an access unit delimiter with a range covering one or more pictures associated with the access unit, i.e., the order of pictures is properly indicated, and therefore pictures that need to be rendered can be efficiently decoded.

[0013] According to an embodiment of the present application, the parameter set unit indicates alternative coding unit types in a video data stream, and indicates whether a given video coding unit is used as a starting point of a refreshed video sequence for video decoding, e.g., RAP type, i.e., includes an immediate decoding refresh (IDR), or as a continuous starting point of a video sequence for video decoding, e.g., non-RAP type, i.e., does not include an IDR. That is, the parameter set unit can be used to indicate whether a coding unit is the first picture of a video sequence or not.

[0014] According to an embodiment of the present application, a video decoder is configured to decode a video including a plurality of pictures from a video data stream by decoding each picture from one or more video coding units within an access unit of the video data stream associated with each picture, each video coding unit carrying picture block partitioning data, block related prediction parameters and prediction residual data, each of the given video coding units being associated with a different region of a picture within which the access unit in which it resides is associated, and reading from each of the given video coding units a map (200) from an n-ary set of one or more syntax elements, such as two flags each 2-ary and a pair 4-ary, where for example the mapping may be fixed by default or it may be modified by decoding the n-ary set of one or more syntax elements into a range of values, such as three binary characteristics, for example an m-ary set of one or more characteristics. The characteristics may be signaled in the data stream by dividing the data into sets, each of which is dyadic such that the triplet becomes 8-ary, and each characteristic describes in an overlapping manner with corresponding data in the given video coding unit, i.e., the characteristics may be inferred from looking at deeper encoded data with respect to how video is encoded in the video data stream with respect to the picture with which the access unit within which the given video coding unit resides is associated (m>n); or by reading N syntax elements (210) from each of the given video coding units, each of which is dyadic, such as N=2 flags (N>0), and reading the associated information from the video data stream and treating them as variables of the associated characteristic depending on the associated information, such as each of the M=3 binary characteristics being dyadic → and the associated information being 3, i.e.,

number

[0015] According to an embodiment of the present application, a map is included in each parameter set unit to indicate the location of the mapped characteristic. A map is signaled in the data stream to indicate the location of the mapped characteristic. N syntax elements indicate the presence or absence of a characteristic. That is, there is flexibility to combine flags and mappings to indicate a flag in the parameter set.

[0016] According to an embodiment of the present application, a video encoder is configured to encode a video having a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with each picture, indicate alternative coding unit types in parameter set units of the video data stream, and for each given video coding unit, encode a coding unit type identifier (100) for each video coding unit into the video data stream, the coding unit identifier identifying a coding unit type from a first subset of one or more coding unit types (102) or from a second subset of coding unit types (104), the coding unit identifier identifying one or more coding unit types from a first subset of one or more coding unit types (105), If the coding unit identifier identifies a coding unit type from a first subset of unit types, then each given video coding unit should be attributed to an alternative coding unit type, and if the coding unit identifier identifies a coding unit type from a second subset of coding unit types, then each given video coding unit should be attributed to a coding unit type from the second subset of coding unit types identified by the coding unit identifier, the alternative coding unit type being a RAP type, and the video encoder is configured to identify the video coding unit of the RAP picture as the given video coding unit, and directly encode the coding unit type identifier for the purely intra-coded video coding unit of the non-RAP picture that identifies the RAP type, i.e., the coding unit type is indicated in a parameter set unit of the video data stream, thus improving coding efficiency, i.e., without the need to encode each segment by an IDR picture.

[0017] According to an embodiment of the present application, a video composer composes a video encoded video data stream including a plurality of pictures, each picture being associated with one or more video coding units for each tile into which the picture is divided, and for each tile into which the picture is divided, a video composer modifies an alternative coding unit type in a parameter set unit of the video data stream to indicate a non-RAP type, and identifies an exclusively coded video coding unit in a picture of the video data stream, the coding unit type identifying an RAP picture, and for each predetermined picture of the video data stream, the video composer modifies an alternative coding unit type in a parameter set unit of the video data stream to indicate a non-RAP type, and for each predetermined picture of the video data stream, the video composer modifies an alternative coding unit type in a parameter set unit of the video data stream to indicate a non-RAP type. For a video data stream of a coding unit type, an identifier (100) for each video coding unit encoded into the video data stream identifies the coding unit type from a first subset (102) of one or more coding unit types or from a second subset (104) of coding unit types, where if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, then each given video coding unit should be assigned to the alternative coding unit type, and if the coding unit identifier identifies a coding unit type from the second subset of coding unit types, then each given video coding unit should be assigned to a coding unit type from the second subset (104) of coding unit types identified by the coding unit identifier. The types of video coding units are identified by utilizing the identifiers to efficiently organize the first and second subsets of coding unit types, and a picture of the video organized, for example, by a plurality of tiles.

[0018] According to an embodiment of the present application, a video encoder encodes a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with each picture, each video coding unit carrying picture block partitioning data, block related prediction parameters, and prediction residual data, the access unit being associated with a different region of a picture within which each given video coding unit resides, and for each given video coding unit, an n-ary set of one or more syntax elements is indicated, e.g., two flags, each being 2-ary, such that the pair is a 4-ary map (200), the mapping may be fixed by default or it may be signaled in the data stream, or both may be signaled by dividing a range of values, and the n-ary set of one or more syntax elements is divided into an m-ary set of one or more characteristics, e.g., three binary characteristics, Each is dyadic, so that the triplet is 8-ary, and each characteristic describes in a redundant way the corresponding data in the given video coding unit, i.e., the access unit is associated with, and the characteristics may be inferred from an inspection of deeper encoded data, with respect to how the video is encoded into the video data stream for the picture the given video coding unit resides within, where m>n, or where m>n indicates N syntax elements (210) in each of the given video coding units, e.g., each of the N=2 flags is dyadic (N>0) and indicates association information to the video data stream, i.e., associates or treats them as variables of the associated characteristic depending on the association information, and each of the N syntax elements carries information about one of the M characteristics, e.g., one of the M=3 binary characteristics, and thus each is dyadic, with three possibilities for associating the two flags from three to two, i.e.,

number

[0019] According to an embodiment of the present application, a method includes: video decoding comprising a plurality of pictures from a video data stream by decoding each picture from one or more video coding units within an access unit of the video data stream associated with each picture; reading alternative coding unit types from a parameter set unit of the video data stream; for each given video coding unit, reading a coding unit type identifier from each coding unit; ascertaining whether the coding unit identifier identifies a coding unit type from a first subset of one or more coding unit types or from a second subset of coding unit types; if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, attributing each video coding unit to the alternative coding unit type; and attributing the coding unit identifier to a coding unit type from the second subset of coding unit types from the second subset of coding unit types identified by the coding unit identifier.

[0020] According to an embodiment of the present application, a method includes decoding a video comprising a plurality of pictures from a video data stream by decoding each picture from one or more video coding units within an access unit of the video data stream associated with each picture, each video coding unit carrying picture block partitioning data, block related prediction parameters, and prediction residual data, associated with a different region of a picture to which the access unit is associated and within which each given video coding unit resides; reading from each given video coding unit an n-ary set of one or more syntax elements, e.g., each of two flags is binary and thus the pair is quaternary, the mapping may be fixed by default or it may be signaled in the data stream or both may be signaled by dividing a range of values; dividing the n-ary set of one or more syntax elements into an m-ary set of one or more characteristics, e.g., each of three binary characteristics is , 2-ary, so that the triplet is 8-ary, and each characteristic describes in a redundant way the corresponding data in the given video coding unit, i.e., with which the access unit is associated, and with respect to how the video is encoded into the video data stream with respect to the picture within which the given video coding unit resides, the characteristics may be inferred from an inspection of deeper encoded data, where m>n, or by reading N syntax elements (210) from each of the given video coding units, e.g., each of N=2 flags is 2-ary (N>0) and indicates association information to the video data stream, i.e., by associating, i.e., treating them as variables of the associated characteristic depending on the association information, and each of the N syntax elements having information about one of M characteristics, e.g., one of M=3 binary characteristics, and thus each is 2-ary, with three possibilities for associating the two flags from three to two, i.e.

number

[0021] According to an embodiment of the present application, a method includes encoding a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with each picture; indicating alternative coding unit types within a parameter set unit of the video data stream; and defining a respective coding unit type identifier (100) for each given video coding unit, the coding unit identifier identifying a coding unit type from a first subset (102) of one or more coding unit types or from a second subset (104) of coding unit types, attributing each video coding unit to a coding unit type if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, and attributing each video coding unit to a coding unit type from the second subset of coding unit types identified by the coding unit identifier if the coding unit identifier identifies a coding unit type from the second subset of coding unit types.

[0022] According to an embodiment of the present application, the method comprises: constructing a video encoded video data stream including a plurality of pictures, each picture being in one or more video coding units in an access unit of the video data stream, one or more coding units for each tile into which the picture is divided being associated with each picture; changing an alternative coding unit type in a parameter set unit of the video data stream to indicate a RAP type to indicate a non-RAP type; identifying in a picture of the video data stream an exclusively coded video coding unit whose coding unit type identifies a RAP picture, an identifier (100) encoded in the video data stream; and, for each predetermined video coding unit of the video data stream, For each coding unit, an identifier (100) for each video coding unit encoded in the video data stream identifies a coding unit type from a first subset (102) of one or more coding unit types or from a second subset (104) of coding unit types, where if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, then each given video coding unit is to be assigned an alternative coding unit type, and if the coding unit identifier identifies a coding unit type from the second subset of coding unit types, then each given video coding unit is to be assigned a coding unit type from the second subset of coding unit types identified by the coding unit identifier.

[0023] According to an embodiment of the present application, a method includes encoding a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with each picture, each video coding unit carrying picture block partitioning data, block related prediction parameters, and prediction residual data, the access unit being associated with a different region of a picture within which each given video coding unit resides, indicating to each given video coding unit an n-ary set of one or more syntax elements, e.g., each of two flags is binary and thus the pair is quaternary, the mapping (200) may be fixed by default or it may be signaled in the data stream or both may be signaled by dividing a range of values, the n-ary set of one or more syntax elements being divided into an m-ary set of one or more characteristics (20), e.g., three binary characteristics: each is dyadic, so that the triplet is 8-ary, and each characteristic describes in a redundant way the corresponding data in the given video coding unit, i.e., the access unit is associated with and the characteristics may be inferred from an inspection of deeper encoded data regarding how video is encoded into the video data stream for the picture the given video coding unit resides within, where m>n, or by reading N syntax elements (210) from each of the given video coding units, e.g., each of N=2 flags is dyadic (N>0) and indicates association information to the video data stream, i.e., by associating, i.e., treating them as variables of the associated characteristic depending on the association information, and each of the N syntax elements having information about one of the M characteristics, e.g., one of M=3 binary characteristics, and thus each is dyadic, with three possibilities for associating the two flags from three to two, i.e.,

number

[0024] Preferred embodiments of the present application are described below with reference to the drawings. [Brief description of the drawings]

[0025] [Figure 1] A schematic diagram showing a client and server system for a virtual reality application is shown as an example in which the embodiments given in the following figures may be advantageously utilized. [Diagram 2] FIG. 2 shows a schematic diagram illustrating an example of 360-degree video in a dual resolution cube map projection tiled into 6×4 tiles that can be fitted into the system of FIG. [Diagram 3] FIG. 3 shows a schematic diagram illustrating an example of user viewpoint and tile selection for 360-degree video streaming as shown in FIG. 2. [Figure 4] FIG. 4 shows a schematic diagram illustrating an example of the arrangement (packing) of tiles obtained as shown in FIG. 3 in a joint bitstream after a merging process. [Diagram 5] FIG. 13 shows a schematic diagram illustrating an example of tiling with low-resolution fallback as a single tile for 360-degree video streaming. [Figure 6] 1 shows a schematic diagram illustrating an example of tile selection in tile-based streaming. [Figure 7] FIG. 13 shows a schematic diagram illustrating another example of tile selection in tile-based streaming with non-uniform Random Access Point (RAP) duration per tile. [Figure 8] 1 shows a schematic diagram illustrating an example of a Network Adaptive Layer (NAL) unit header according to an embodiment of the present application. [Figure 9]1 illustrates a specific example of a table showing types of Row Byte Sequence Payload (RBSP) data structures included in a NAL unit according to an embodiment of the present application. [Figure 10] 1 shows a schematic diagram illustrating an example of a sequence parameter set indicating which NAL unit types are mapped according to an embodiment of the present application. [Figure 11] A schematic diagram showing an example of slice characteristics indicated in a slice header is shown. [Figure 12] FIG. 13 shows a schematic diagram illustrating an example of a map indicating characteristics of slices in a sequence parameter set according to an embodiment of the present application. [Figure 13] FIG. 13 shows a schematic diagram illustrating an example of characteristics shown by a map in the sequence parameter set of FIG. 12 according to an embodiment of the present application. [Figure 14] FIG. 2 shows a schematic diagram illustrating an example of association information indicated by utilizing a syntax structure according to an embodiment of the present application. [Figure 15] FIG. 13 shows a schematic diagram illustrating another example of a map indicating characteristics of slices in a sequence parameter set according to an embodiment of the present application. [Figure 16] FIG. 16 shows a schematic diagram illustrating another example of characteristics shown by a map in the parameter set of FIG. 15 according to an embodiment of the present application. [Figure 17] FIG. 13 shows a schematic diagram illustrating a further example of a map indicating slice characteristics in a slice segment header according to an embodiment of the present application. [Figure 18] 20 shows a schematic diagram illustrating further examples of properties indicated by a map in the slice segment header of FIG. 17 according to an embodiment of the present application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0026] In the following description, equal or equivalent elements having equal or equivalent functions are designated by equal or equivalent reference numbers.

[0027] In the following description, numerous details are set forth to provide a more thorough description of the embodiments of the present application. However, it will be apparent to those skilled in the art that the embodiments of the present application may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the embodiments of the present application. In addition, features of the embodiments described below may be combined with each other unless otherwise specified.

[0028] Introduction In what follows, it should be noted that the individual aspects described herein may be used individually or in combination, and thus details may be added to each of the individual aspects without adding details to any other one of the aspects.

[0029] It should also be noted that this disclosure explicitly or implicitly describes features that can be used in a video decoder (an apparatus for providing a decoded representation of a video signal based on an encoded representation), and thus any of the features described herein can be used in the context of a video decoder.

[0030] Moreover, features and functions disclosed herein in relation to a method may also be used in an apparatus (configured to perform such functions). Moreover, any feature and function disclosed herein in relation to an apparatus may also be used in the corresponding method. In other words, a method disclosed herein may be supplemented by any of the features and functions described in relation to an apparatus.

[0031] To facilitate understanding of the description of the embodiments of the present application with respect to various aspects thereof, Fig. 1 illustrates an example of an environment in which the embodiments described below of the present application may be applied and effectively utilized. In particular, Fig. 1 illustrates a system consisting of a client 10 and a server 20 interacting via adaptive streaming. For example, Dynamic Adaptive Streaming over HTTP (DASH) may be utilized for communication 22 between the client 10 and the server 20. However, the embodiments outlined below should not be construed as being limited to the use of DASH, and similarly, terms such as Media Presentation Description (MPD) should be understood to be broad to include manifest files, which are defined differently from DASH.

[0032] Fig. 1 shows a system configured to realize a virtual reality application, i.e. the system is configured to present to a user wearing a head-up display 24, i.e. via an internal display 26 of the head-up display 24, a view portion 28 from a time-varying spatial scene 30, the portion 28 corresponding to the orientation of the head-up display 24, exemplarily measured by an internal orientation sensor 32, such as an inertial sensor of the head-up display 24. That is, the portion 28 presented to the user forms a portion of the spatial scene 30 at a spatial position corresponding to the orientation of the head-up display 24. In the case of Fig. 1, the time-varying spatial scene 30 is depicted as an omnidirectional or spherical video, but the description of Fig. 1 and the embodiments described below can be easily transferred to other examples, such as presenting a portion from a video with the spatial position of the portion 28 determined by the intersection of a face access or eye access with a virtual or real projector wall, etc. Furthermore, the sensor 32 and the display 26 may be constituted by different devices, e.g. a remote control and a corresponding television, respectively, or they may be part of a portable device, such as a mobile device, such as a tablet or a mobile phone. Finally, it should be noted that some of the embodiments described below may also be applied to scenarios in which the region 28 presented to the user always covers the entire time-varying spatial scene 30 when presenting a time-varying spatial scene, e.g. with non-uniformity associated with an uneven distribution of quality across the spatial scene.

[0033] Further details regarding the server 20, the client 10, and the manner in which the spatial content 30 is provided at the server 20 are shown in Figure 1 and described below, but these details should also not be treated as limiting the embodiments described below, but rather should serve as examples of how to implement any of the embodiments described below.

[0034] In particular, as shown in Fig. 1, the server 20 may comprise a storage 34 and a controller 36, such as a suitably programmed computer, an application specific integrated circuit or the like. The storage 34 is stored with media portions representing a time-varying spatial scene 30. An example is outlined in more detail below with respect to the diagram of Fig. 1. The controller 36 responds to requests sent by the client 10 by resending a media presentation description for the media portions requested by the client 10, and may also send further information of its own to the client 10. More details in this regard are provided below. The controller 36 may fetch the requested media portions from the storage 34. Within this storage, other information, such as the media presentation description or parts thereof, may be stored in other signals sent from the server 20 to the client 10.

[0035] 1, the server 20 may optionally further comprise a stream modifier 38 for modifying the media portions sent from the server 20 to the client 10 in response to a request from the latter, so that, for example, the media portions thus retrieved by the client 10 result in a media data stream at the client 10 which is in fact aggregated from several media streams but which forms one single media stream decodable by one associated decoder. However, the presence of such a stream modifier 38 is optional.

[0036] The client 10 of Fig. 1 is exemplarily shown as including a client device or controller 40 or more decoders 42 and a reprojector 44. The client device 40 may be a suitably programmed computer, a microprocessor, a programmed hardware device such as an FPGA, or an application specific integrated circuit, etc. The client device 40 assumes the responsibility for selecting the media portions to be extracted from the server 20 out of the plurality 46 of media portions provided in the server 20. For this, the client device 40 first extracts a manifest or media presentation description from the server 20. From the same, the client device 40 obtains calculation rules for calculating the address of the media portion out of the plurality of media portions 46 that corresponds to a particular required spatial portion of the spatial scene 30. The media portions thus selected are extracted from the server 20 by the client device 40 by sending respective requests to the server 20. These requests include the calculated addresses.

[0037] The media portions thus extracted by the client device 40 are forwarded by the latter to one or more decoders 42 for decoding. In the example of Fig. 1, the thus extracted and decoded media portions represent only a spatial portion 48 from the time-varying spatial scene 30 for each time unit in time, but as already mentioned above, this may differ according to other aspects, for example where the presented view portion 28 always covers the whole scene. The reprojector 44 may optionally reproject and cut out the view portion 28 displayed to the user from the extracted and decoded scene content of the selected, extracted and decoded media portion. To this end, as shown in Fig. 1, the client device 40 may continuously track and update the spatial position of the view portion 28, for example in response to user orientation data from the sensor 32, and may inform the reprojector 44 of, for example, said current spatial position of the scene portion 28 and the reprojection mapping to be applied to the extracted and decoded media content to be mapped to the area forming the view portion 28. In response, reprojector 44 may apply a mapping and interpolation onto a regular pixel grid, for example, as displayed on display 26.

[0038] FIG. 1 shows the case where cubic mapping is used to map the spatial scene 30 onto the tiles 50. Thus, the tiles are depicted as rectangular sub-regions of a cube onto which the scene 30, which has the form of a sphere, is projected. The reprojector 44 reverses this projection. However, other examples may be applied as well. For example, instead of a cubic projection, a projection onto a truncated pyramid or a pyramid without truncation may be used. Furthermore, although the tiles in FIG. 1 are shown as non-overlapping in terms of coverage of the spatial scene 30, the division into tiles may include mutual tile overlap. Also, as outlined in more detail below, the spatial division of the scene 30 into tiles 50, with each tile forming a representation as described below, is also not mandatory.

[0039] Thus, as depicted in Fig. 1, the entire spatial scene 30 is spatially subdivided into tiles 50. In the example of Fig. 1, each of the six faces of a cube is subdivided into four tiles. For illustrative purposes, the tiles are enumerated. For each tile 50, the server 20 provides a video 52 as shown in Fig. 1. More precisely, the server 20 provides multiple videos 52 per tile 50, which videos have different quality Q#. Furthermore, the videos 52 are temporally subdivided into time segments 54. The time segments 54 of all the videos 52 of all the tiles T# form one of multiple media portions 46 stored in the storage 34 of the server 20 or are encoded respectively.

[0040] It is emphasized again that even the example of tile-based streaming depicted in Fig. 1 merely forms an example for which many variations are possible. For example, although Fig. 1 seems to suggest that media portions associated with a higher quality representation of scene 30 are associated with tiles that match the tiles to which the media portions in which scene 30 is encoded with quality Q1 belong, this correspondence is not necessary and tiles of different qualities may correspond to tiles of different projections of scene 30. Furthermore, although not discussed so far, media portions corresponding to different quality levels depicted in Fig. 1 may differ in spatial resolution, signal-to-noise ratio and / or temporal resolution, etc.

[0041] Finally, unlike the tile-based streaming concept, in which the media portions that may be individually extracted by the device 40 from the server 20 relate to tiles 50 into which the scene 30 is spatially subdivided, the media portions provided at the server 20 may alternatively have a full sampling resolution at different spatial locations in the scene 30, e.g. each having the scene 30 encoded therein in a spatially complete manner with spatially varying sampling resolutions. For example, this can be achieved by providing at the server 20 a sequence of segments 54 that relate to a projection of the scene 30 onto truncated pyramids, the tips of which are oriented in different directions relative to one another, thereby leading to resolution peaks oriented in different directions.

[0042] Further, it should be noted that, optionally, with respect to the current stream modification device 38, the same may be part of the client 10 or may be located between the client 10 and the server 20 in the network devices through which the client 10 and the server 20 exchange signals as described herein.

[0043] There are certain video-based applications in which multiple encoded video bitstreams are jointly decoded, i.e., merged into a joint bitstream and fed to a single decoder as follows. · Multi-party conferencing: Encoded video streams from multiple participants are processed on a single endpoint. Or tile-based streaming: e.g. 360 degree tiled video playback in VR applications

[0044] In the latter, a 360-degree video is spatially segmented and each spatial portion is provided to the streaming client in multiple representations at various spatial resolutions, as shown in Figure 2. Figure 2(a) shows a high-resolution tile and Figure 2(b) shows a low-resolution tile. Figure 2, i.e., Figures 2(a) and (b), show a cube map projecting a 360-degree video divided into 6x4 spatial portions at two resolutions. For simplicity, these independently decodable spatial portions are referred to herein as tiles.

[0045] A user typically, when using a state-of-the-art head mounted display, sees only a subset of the tiles that make up the entire 360 ​​degree video through a solid viewport boundary 80 that represents a 90x90 degree field of view, as shown in Figure 3(a). The corresponding tiles, indicated by reference numeral 82 in Figure 3(b), are downloaded in full resolution.

[0046] However, the client application must also download and decode representations of other tiles outside the current viewport, as indicated by reference numeral 84 in FIG. 3(c), to handle sudden changes in user orientation. Thus, a client of such an application downloads tiles covering the current viewport in the highest resolution and tiles outside the current viewport in a relatively lower resolution, as shown in FIG. 3(c), while the choice of tile resolution is always adapted to the user's orientation. After client-side downloading, merging the downloaded tiles into a single bitstream processed by a single decoder is a way to address the constraints of a typical mobile device with limited computational and power resources. FIG. 4 shows possible tile arrangements in a joint bitstream for the above example. The merging process to generate the joint bitstream needs to be performed by avoiding compressed domain processing, i.e., processing on the pixel domain by transcoding.

[0047] While the example from Fig. 4 shows the case where all tiles (high resolution and low resolution) cover the entire 360 ​​degree space and no tile covers the same area repeatedly, another tile grouping is also available, as shown in Fig. 5. It defines the entire low resolution part of the video as a "low resolution fallback" layer, as shown in Fig. 5(b), which can be merged with the high resolution tiles of Fig. 5(a) that cover a subset of the 360 ​​degree video. The entire low resolution fallback video can be coded as a single tile, as shown in Fig. 5(c), while the high resolution tiles are rendered as an overlay on the low resolution part of the video in the final stage of the rendering process.

[0048] The client starts a streaming session according to the user's tile selection by downloading all desired tile tracks as shown in Fig. 6, where the client starts the session with tile 0, indicated by reference number 90 in Fig. 6(a), and tile 1, indicated by reference number 92. Whenever a viewport change is made (i.e., the user turns his head to look in another direction), the tile selection is changed in the next time portion, i.e., tile 0 and tile 2, indicated by reference number 94 in Fig. 6(a), and in the next available segment, the client changes the position of tile 2 and replaces tile 0 with tile 1, as shown in Fig. 6(b). It is important to note that every segment needs to start with an Instantaneous Decoder Refresh (IDR) picture, i.e., a prediction chain reset picture, for new tile selection and tile position change, which would otherwise cause prediction mismatch, artifacts and drift.

[0049] Encoding each part with an IDR picture is costly in terms of bitrate. Parts can potentially be very short in duration, e.g. to react quickly to orientation changes, which is why it is desirable to encode multiple variants with varying IDR (or RAP: Random Access Point) durations, as shown in Fig. 7. For example, as shown in Fig. 7(b), at time t1, there is no reason to break the prediction chain for tile 0, since tile 0 has already been downloaded for time t0 and is located in the same position that the client can select a part that does not start with a RAP available at the server.

[0050] However, one remaining problem is that slices (tiles) in a coded picture are subject to certain constraints. One of them is that a picture may not simultaneously contain Network Abstract Layer (NAL) units of RAP and non-RAP NAL unit types. Thus, applications have only two less desirable options to address the above problem. First, the client can rewrite the NAL unit type of RAP pictures when they are merged into pictures with non-RAP NAL units. Second, the server can obscure the RAP nature of these pictures by using non-RAP from the start. However, this prevents detection of the RAP nature in systems that should process these coded videos, for example for file format packaging.

[0051] The present invention is a NAL unit type mapping that allows mapping of one NAL unit type to another NAL unit type via an easily rewritable syntax structure.

[0052] In one embodiment of the present invention, a NAL unit type is specified as mappable, and the mapped type is specified in a parameter set, for example based on Draft 6 V14 of the Versatile Video Coding (VVC) specification with highlighted editing, as follows:

[0053] Figure 8 shows the NAL unit header syntax. The syntax nal_unit_type, identifier 100, specifies the NAL unit type, i.e., the type of Row Byte Sequence Payload (RBSP) data structure contained in the NAL unit as specified in the table shown in Figure 9.

[0054] The variable NalUnitType is defined as follows:

number

[0055] All references to the syntax element nal_unit_type in the specification are replaced with references to the variable NalUnitType as well as the following constraints:

[0056] The value of NalUnitType is the same for all coded slice NAL units of a picture. A picture or layer access unit is referred to as having the same NAL unit type as the coded slice NAL units of the picture or layer access unit. That is, as shown in Fig. 9, a first subset of coding unit types 102 indicates "nal_unit_type" 12, i.e., "MAP_NUT" and "VCL" as NAL unit type classes. Accordingly, a second subset of coding unit types 104 indicates "VCL" as NAL unit type class, i.e., all coded slice NAL units of a picture, indicated by numbers 0 to 15 of identifier 100, have the same NAL unit type class of the first subset of coding unit types 102, i.e., coding unit type of VCL.

[0057] FIG. 10 shows a sequence parameter set RBSP syntax including mapped_nut as indicated by reference numeral 106, which indicates a NalUnitType of NAL units with nal_unit_type equal to MAP_NUT.

[0058] In another embodiment, the mapped_nut syntax element is carried in the access unit delimiter AUD.

[0059] In other embodiments, it is a bitstream conformance requirement that the value of mapped_nut must be a VCL NAL unit type.

[0060] In another embodiment, the mapping of NalUnitTypes of NAL units with nal_unit_type equal to MAP_NUT is performed by profiling information. Such a mechanism makes it possible to have more than one mappable NAL unit type instead of having a single MAP_NUT and to indicate the required interpretation of the NALUnitTypes of the mappable NAL units in a simple profiling mechanism or in a single syntax element mapped_nut_space_idc.

[0061] In other embodiments, the mapping mechanism is used to extend the range of values ​​of NAL unit types, which is currently limited to 32 (e.g., because it is u(5) as shown in FIG. 10). The mapping mechanism can indicate any open-ended value, as long as the number of required NAL unit types does not exceed the number of values ​​reserved for mappable NAL units.

[0062] In one embodiment, when a picture simultaneously contains slices of alternative coding unit type and slices of normal coding unit type (e.g., existing NAL units of the VCL category), a mapping is performed to effectively result in all slices of the picture having the same coding unit type characteristics, i.e., the alternative coding unit type is equal to the coding unit type of the non-alternative slice of normal coding type. Furthermore, the above embodiment only holds for pictures with or without the random access nature.

[0063] In addition to the described problems regarding NAL unit types and NAL unit type extensibility in merge scenarios and corresponding solutions, there are some video applications where information related to the video and how the video was coded, such as on-the-fly-adaptation, is required for system integration and transmission or operation.

[0064] There are some common pieces of information established within the last few years that are widely used in the industry and are clearly designated, and specific bit values ​​are used for such purposes. Examples include: -Temporary ID in NAL unit header NAL unit type, including IDR, CRA, TRAIL, ... or SPS (Sequence Parameter Set), PPS (Picture Parameter Set), etc. It is.

[0065] However, there are some scenarios where additional information can be useful. Although not widely used, some usefulness has been found in some cases, e.g., BLA, partial RAP NAL units for subpictures, sub-layer non-reference NAL units, etc. If the extensibility mechanisms mentioned above are used, some of these NAL unit types are feasible. However, as an alternative, some fields in the slice header can also be used.

[0066] Conventionally, additional information is reserved in the slice header that is used to signal specific characteristics of the slice. Discardable flag: indicates that the coded picture can be used as a reference picture for inter prediction and cannot be used as a source picture for inter-layer prediction. · cross_layer_bla_flag: Affects the derivation of output pictures for layered coding, such that pictures preceding the RAP in higher layers may not be output.

[0067] Similar mechanisms may be envisaged for upcoming video codec standards. However, one constraint of these mechanisms is that the defined flags occupy specific positions in the slice header. In the following, the use of these flags in HEVC is shown in Figure 11.

[0068] As mentioned above, the problem with such a solution is that the positions of the extra slice header bits are assigned in a phased manner, and for applications that utilize rarer information, flags will likely come in later positions, increasing the number of bits that need to be transmitted in the extra bits (e.g., in the case of HEVC, “discardable_flag” and “cross_layer_bla_flag”).

[0069] Alternatively, following a similar mechanism described for the NAL unit type, the mapping of flags in extra slice header bits in the slice header can be defined in the parameter set. An example is shown as FIG.

[0070] FIG. 12 shows an example of a sequence parameter set including a map indicating association information utilizing "extra_slice_header_bits_mapping_space_idc", ie a map indicated by reference numeral 200 indicating the mapping space for extra bits of the slice header.

[0071] Figure 13 shows the mapping of bits to the flags indicated by "extra_slice_header_bits_mapping_space_idc" 200 in Figure 12. As shown in Figure 13, the binary characteristics 202 redundantly describe the corresponding data in a given video coding unit. In Figure 13, three binary characteristics 202 are shown, namely "0", "1" and "2". The number of binary characteristics is variable depending on the number of flags.

[0072] In other embodiments, the mapping is performed in a syntax structure (e.g., as shown in FIG. 14) that indicates the presence or absence of syntax elements in the extra bits in the slice header (e.g., as shown in FIG. 11). That is, for example, as shown in FIG. 11, the condition on the presence flag controls the presence of syntax elements in the additional bits in the slice header, i.e., "num_extra_slice_header_bits>i" and "i<num_extra_slice_header_bits i++". In FIG. 11, as described above, each syntax element in the additional bits is placed at a specific position in the slice header. However, in this embodiment, it is not necessary for a syntax element, e.g., "discardable_flag", "cross_layer_bla_flag" or "slice_reserved_flag[i]", to occupy a specific position. Instead, when checking the condition regarding the value of a specific presence flag (e.g., "discardable_flag_present_flag" in FIG. 14), when it is shown that the first syntax element (e.g., "discardable_flag" in FIG. 11) does not exist, the following second syntax element, if it exists, takes the position of the first syntax element in the slice header. Also, the syntax elements in the additional bits can be present in the picture header by indicating a flag, e.g., by indicating "sps_extra_ph_bit_present_flag[i]". Also, the syntax structure, e.g., the number of syntax elements, i.e., the number of flags presented, indicates the presence / absence of specific characteristics or the number of syntax elements presented in the additional bits. That is, the number of specific characteristics or syntax elements presented in the additional bits is indicated by counting the number of syntax elements (flags) presented. In FIG. 14, the presence of each syntax element is indicated by flag 210. That is, each flag in 210 indicates the presence of a slice header indication for a specific characteristic of a given video coding unit.Furthermore, the further syntax "[...] / / additional flag" shown in Figure 14 and corresponding to the "slice_reserved_flag[i]" syntax element in the slice header of Figure 11 is used as a placeholder to indicate the presence / absence of a syntax element in the additional bits or as an indication of the presence / absence of an additional flag.

[0073] In another embodiment, the flag type mapping is signaled for each extra slice header bit in the parameter set, for example as shown in Figure 15. As shown in Figure 15, the syntax "extra_slice_header_bit_mapping_idc", i.e., map 200, is signaled in the sequence parameter set and indicates the location of the mapped property.

[0074] Figure 16 shows the mapping of bits to the flag indicated by "extra_slice_header_bits_mapping_space_idc" in Figure 15. That is, a binary characteristic 202 corresponding to the map 200 shown in Figure 15 is shown in Figure 16.

[0075] In another embodiment, the slice header extension bit is replaced by idc signaling, which indicates a particular flag value combination, for example as shown in Figure 17. As shown in Figure 17, a map 200, namely "extra_slice_header_bit_idc", is indicated in the slice segment header, i.e. the map 200 indicates the presence of a property as shown in Figure 18.

[0076] FIG. 18 shows that the flag value represented by a particular value of "extra_slice_header_bit_idc", i.e., binary characteristic 202, is either signaled in the parameter set or predefined in the specification (known in advance).

[0077] In one embodiment, the value space of "extra_slice_header_bit_idc", i.e. the value space of map 200, is divided into two ranges: one range represents the combinations of flag values ​​that are known a priori, and one range represents the combinations of flag values ​​that are signaled in the parameter set.

[0078] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with blocks or devices corresponding to method steps or features of method steps. Similarly, aspects described in the context of a method step also represent a description of a corresponding block, item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0079] The data stream of the present invention may be stored on a digital storage medium or may be transmitted over a transmission medium, such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0080] Depending on specific implementation requirements, the embodiments of the present application can be realized by hardware or software. The realization can be performed using a digital storage medium, such as, for example, a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM or flash memory, on which electronically readable control signals are stored that cooperate (or can cooperate) with a programmable computer system to execute the respective methods. Thus, the digital storage medium may be computer readable.

[0081] Some embodiments according to the invention comprise a data carrier having electronically readable control signals capable of cooperating with a programmable computer system to cause one of the methods described herein to be performed.

[0082] Generally, the embodiments of the present application can be realized as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. The program code may for example be stored on a machine readable carrier.

[0083] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0084] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0085] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer readable medium) comprising recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.

[0086] A further embodiment of the inventive method is therefore a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or the sequence of signals being adapted to be transmitted via a data communication connection, for example the Internet.

[0087] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0088] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0089] Further embodiments according to the invention include an apparatus or system configured to transmit (e.g. electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0090] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be utilized to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0091] The apparatus described herein may be implemented using a hardware apparatus, using a computer, or using a combination of a hardware apparatus and a computer.

[0092] The apparatus described herein, or any components of the apparatus described herein, may be implemented at least in part in hardware and / or software.

Claims

1. 1. A method of video decoding, comprising: receiving a sequence parameter set (SPS) from a video data stream; Parsing an extra slice header bitmap from the SPS; mapping one or more bits in the extra slice header bitmap to one or more flags indicating the presence or absence of corresponding syntax elements in a slice header; determining a number of extra slice header bits based on a number of the one or more flags indicating presence; parsing syntax elements in the slice header based on the determined number of extra slice header bits from the video data stream; The method comprising:

2. The method of claim 1 , further comprising: determining a position of the syntax element in the slice header based on the determined number of extra slice header bits.

3. The method of claim 1 , wherein the number of the one or more flags indicating presence is determined based on counting the number of the one or more flags indicating presence.

4. The method of claim 1 , wherein each of the one or more flags has a one-bit value.

5. 1. A method of video encoding, comprising: providing a sequence parameter set (SPS) via a video data stream; encoding an extra slice header bitmap into the SPS, the extra slice header bitmap mapping one or more bits in the extra slice header bitmap to one or more flags indicating the presence or absence of corresponding syntax elements in the slice header; determining a number of extra slice header bits based on a number of the one or more flags indicating presence; providing the syntax element in the slice header based on the determined number of extra slice header bits over the video data stream; The method comprising:

6. The method of claim 5 , further comprising determining a location of the syntax element in the slice header based on the determined number of extra slice header bits.

7. The method of claim 5 , wherein the number of the one or more flags indicating presence is determined based on counting the number of the one or more flags indicating presence.

8. The method of claim 5 , wherein each of the one or more flags has a one-bit value.

9. A video coding device comprising a processing circuit configured to perform the method according to any one of claims 1 to 8.

10. A non-transitory computer readable medium having instructions which, when executed by a processing circuit, perform the method of any one of claims 1 to 8.