Configurable NAL and slice codepoint mechanism for stream merging

The decoder and encoder system addresses inefficiencies in stream merging by using identifiers to map NAL unit types, enhancing efficiency and flexibility in decoding and encoding processes for video streams with varying Random Access Points.

JP7760777B2Active Publication Date: 2025-10-27FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025024939
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-03
Filing Date
2025-02-19
Publication Date
2025-10-27
Estimated Expiration
2040-09-03

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in stream merging due to limitations in indicating alternative coding unit types and characteristics, leading to constraints on decoding and encoding processes, particularly in applications like 360-degree video streaming where sudden changes in user orientation require rapid tile selection and merging.

Method used

A decoder and encoder system that utilizes identifiers to indicate alternative coding unit types and characteristics, allowing for efficient merging of video coding units by mapping NAL unit types, enabling flexible decoding and encoding processes.

Benefits of technology

Improves merging efficiency and coding efficiency by allowing flexible mapping of NAL unit types, reducing bitrate costs and enabling seamless decoding of video streams with varying Random Access Point durations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007760777000006
    Figure 0007760777000006
  • Figure 0007760777000007
    Figure 0007760777000007
  • Figure 0007760777000008
    Figure 0007760777000008
Patent Text Reader

Abstract

To provide a method for deriving necessary information and characteristics of a video coding unit of a video data stream by reading an identifier indicative of a substitute coding unit type, and a non-transitory computer readable medium.SOLUTION: A method includes: decoding each picture from one or more video coding units within an access unit of a video data stream to read a coding unit type identifier 100; checking whether the coding unit identifier identifies a coding unit type out of a first subset of one or more coding unit types 102 or out of a second subset of coding unit types 104; if the coding unit identifier identifies a coding unit type out of the first subset, attributing each of predetermined video coding units to substitute coding unit type, or if the coding unit identifier identifies a coding unit type out of the second subset, attributing each of the predetermined respective video coding units to the coding unit type out of the second subset of coding unit types identified by the coding unit identifier.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to a data structure that indicates coding unit types and characteristics of video coding units of a video data stream. [Background technology]

[0002] It is known that the picture type is indicated in the NAL unit header of the NAL unit that carries the slice of the picture, whereby the essential characteristics of the NAL unit payload are available at a very high level for use by applications.

[0003] Picture types include: Random Access Point (RAP) pictures: where a decoder can start decoding a coded video sequence. These are referred to as Intra Random Access Pictures (IRAP). There are three IRAP picture types: Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA). The decoding process of a coded video sequence always starts with an IRAP. Leading picture: It precedes the random access point picture in output order but is coded after the random access point picture in the coded video sequence. Leading pictures that are independent of the picture preceding the random access point in coding order are called Random Access Decodable Leading (RADL) pictures. Leading pictures that use a picture preceding the random access point in coding order for prediction may be corrupted if decoding starts at the corresponding IRAP. These are called Random Access Skipped Leading (RASL) pictures. - TRAIL picture: It follows the IRAP picture and the leading picture in both output order and display order. - Pictures whose temporal resolution of the coded video sequence can be switched by the decoder: Temporal Sublayer Access and Step-by-Step Temporal Sublayer Access (STSA)

[0004] Therefore, the data structure of the NAL units is an important factor in stream merging. Summary of the Invention

[0005] The subject matter of the present application is to provide a decoder that derives necessary information of video coding units of a video data stream by reading identifiers indicating alternative coding unit types, and a decoder that derives characteristics of the video data stream.

[0006] A further object of the present subject matter is to provide an encoder that utilizes identifiers to indicate alternative coding unit types of video coding units and to indicate characteristics of a video data stream.

[0007] This object is achieved by the subject matter of the present claims.

[0008] According to an embodiment of the present application, a video comprising a plurality of pictures from a video data stream is decoded by decoding each picture from one or more video coding units within an access unit of the video data stream associated with each of the pictures, reading alternative coding unit types from parameter set units of the video data stream, and for each given video coding unit, reading a coding unit type identifier (100), such as a syntax element included in a nal unit header, from each of the video coding units, and determining whether the coding unit identifier is a coding unit type identifier, such as a syntax element included in a VCL (Video Coding Classification Language) to which the nal unit can be mapped. the video decoder is configured to: determine whether a coding unit identifier is a first subset (102) of one or more coding unit types, e.g., indicating whether the coding unit identifier is a (Layer) unit type, or identify a coding unit type from a second subset (104) of coding unit types, e.g., indicating an nal unit type; and, if the coding unit identifier identifies a coding unit type from the first subset of the one or more coding unit types, render each of the given video coding units the alternative coding unit type; and, if the coding unit identifier identifies a coding unit type from the second subset of coding unit types, render each of the given video coding units the coding unit type from the second subset of coding unit types identified by the coding unit identifier. That is, each nal unit type is indicated by an identifier, the first subset of coding unit types, and the second subset of coding unit types, i.e., the nal unit type is rewritten according to the indication of the first and second coding unit types. Thus, merging efficiency can be improved.

[0009] According to an embodiment of the present disclosure, a video decoder is configured to encode, from each video coding unit, a region associated with each of the video coding units in a manner dependent on a coding unit type associated with each of the video coding units. The video decoder may be configured such that the alternative coding unit type is from a second subset of video coding types. The video decoder may be configured such that the alternative coding unit type is from a third subset of video coding types, such as a non-VCL unit type, that includes at least one video coding type not included by the second subset of video coding types.

[0010] According to an embodiment of the present disclosure, a given video coding unit carries picture block partitioning data, block-related prediction parameters, and prediction residual data. When a picture includes both one or more video coding units, such as a slice having a coding unit type of a first subset and a slice having a coding unit type of a second subset, the latter video coding unit has a coding unit type equal to an alternative coding unit type. The alternative coding unit type is a random access point (RAP) coding type. The alternative coding unit type is a coding type other than the random access point (RAP) coding type. That is, an alternative coding unit type is identified, and video coding units having the same alternative coding unit type are merged, thereby appropriately improving merging efficiency.

[0011] According to an embodiment of the present application, each of the given video coding units is associated with a different region of the picture to which the access unit within which each of the given video coding units is associated. A parameter set unit of a video data stream has a scope covering a sequence of pictures, a single picture, or a set of slices from a single picture. The parameter set unit indicates alternative coding unit types in a manner specific to the profile of the video data stream, i.e., slices can be efficiently merged, thus improving coding efficiency.

[0012] According to an embodiment of the present application, a parameter set unit of a video data stream is either a parameter set unit having a range covering the order of pictures, or an access unit delimiter having a range covering one or more pictures associated with the access unit, i.e., the order of pictures is properly indicated, and therefore pictures that need to be rendered can be efficiently decoded.

[0013] According to an embodiment of the present application, the parameter set unit indicates alternative coding unit types in a video data stream, and indicates whether a given video coding unit is used as a starting point of a refreshed video sequence for video decoding, e.g., a RAP type, i.e., includes an instantaneous decoding refresh (IDR), or as a continuous starting point of a video sequence for video decoding, e.g., a non-RAP type, i.e., does not include an IDR. That is, the parameter set unit can be used to indicate whether a coding unit is the first picture of a video sequence.

[0014] According to an embodiment of the present application, a video decoder is configured to decode a video including a plurality of pictures from a video data stream by decoding each picture from one or more video coding units within an access unit of the video data stream associated with each picture, each video coding unit carrying picture block partitioning data, block-related prediction parameters, and prediction residual data, each of which is associated with a different region of a picture within which the access unit in which it resides is associated, and reads from each of the given video coding units a map (200) from an n-ary set of one or more syntax elements, such as two flags, each of which is 2-ary and a pair of which is 4-ary, e.g., the mapping may be fixed by default, or it may be configured to map the n-ary set of one or more syntax elements to m-ary sets of one or more properties, e.g., three binary properties, or other ranges of values. The characteristics are signaled in the data stream by dividing the data stream into sets, each of which is binary so that the triplet becomes 8 terms, and each characteristic describes the corresponding data in a given video coding unit in an overlapping manner, i.e., the characteristics may be inferred from examining deeper coded data regarding how video is coded in the video data stream for the picture to which the access unit within which the given video coding unit is associated (m>n), or by reading N syntax elements (210) from each of the given video coding units, such as N=2 flags (N>0), each of which is binary, and reading the associated information from the video data stream and treating them as variables of the associated characteristic depending on the associated information, such as M=3 binary characteristics, each of which is binary → and the associated information is 3, i.e.,

number

[0015] According to an embodiment of the present application, a map is included in each parameter set to indicate the location of the mapped feature. The map is signaled in the data stream to indicate the location of the mapped feature. N syntax elements indicate the presence or absence of a feature. That is, there is flexibility to combine flags and mappings to indicate flags in the parameter set.

[0016] According to an embodiment of the present application, a video encoder is configured to encode a video having a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with each picture, indicate alternative coding unit types in parameter set units of the video data stream, and, for each given video coding unit, encode a coding unit type identifier (100) for each video coding unit into the video data stream, the coding unit identifier identifying a coding unit type from a first subset (102) of one or more coding unit types or from a second subset (104) of coding unit types, the coding unit identifier identifying one or more coding unit types. When the coding unit identifier identifies a coding unit type from a first subset of unit types, each given video coding unit should be assigned an alternative coding unit type, and when the coding unit identifier identifies a coding unit type from a second subset of coding unit types, each given video coding unit should be assigned a coding unit type from the second subset of coding unit types identified by the coding unit identifier, the alternative coding unit type being a RAP type, and the video encoder is configured to identify the video coding unit of the RAP picture as the given video coding unit and directly encode the coding unit type identifier for, for example, a purely intra-coded video coding unit of a non-RAP picture that identifies the RAP type. That is, the coding unit type is indicated in a parameter set unit of the video data stream, thereby improving coding efficiency, i.e., there is no need to encode each segment by an IDR picture.

[0017] According to an embodiment of the present application, a video composer composes a video-encoded video data stream including a plurality of pictures, each picture being within one or more video coding units within an access unit of the video data stream, one or more video coding units for each tile into which the picture is divided, associated with each picture, and modifies an alternative coding unit type in a parameter set unit of the video data stream to indicate a non-RAP type, and identifies exclusively coded video coding units in a picture of the video data stream, the coding unit type being an identifier (100) encoded into the video data stream, identifying a predetermined video coding unit in each picture of the video data stream, the coding unit type identifying a RAP picture. For each video coding unit type, an identifier (100) for each video coding unit encoded into the video data stream identifies the coding unit type from either a first subset (102) of one or more coding unit types or from a second subset (104) of coding unit types, where if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, then each given video coding unit should be assigned to the alternative coding unit type, and if the coding unit identifier identifies a coding unit type from the second subset of coding unit types, then each given video coding unit should be assigned to a coding unit type from the second subset (104) of coding unit types identified by the coding unit identifier. The video coding unit types are identified using the identifiers, and the first and second subsets of coding unit types, and pictures of the video, for example, comprised of multiple tiles, are efficiently constructed.

[0018] According to an embodiment of the present application, a video encoder encodes a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with each picture, each video coding unit carrying picture block partitioning data, block-related prediction parameters, and prediction residual data, the access unit being associated with a different region of the picture within which each given video coding unit resides, and for each given video coding unit, an n-ary set of one or more syntax elements, e.g., two flags, each binary, indicating a pair such that the pair is a 4-ary map (200), the mapping may be fixed by default, or it may be signaled in the data stream, or both may be signaled by dividing a range of values, and the n-ary set of one or more syntax elements may be divided into an m-ary set of one or more characteristics, e.g., three binary characteristics Each is dyadic, so that the triplet is 8-ary, and each characteristic describes in a redundant way the corresponding data in a given video coding unit, i.e., the access unit is associated with, and the characteristics may be inferred from an examination of deeper coded data regarding how video is coded into a video data stream for the picture within which the given video coding unit resides, where m>n or indicates N syntax elements (210) in each given video coding unit, e.g., each of N=2 flags is dyadic (N>0) and indicates association information to the video data stream, i.e., associates them, i.e., treats them as variables of the associated characteristic depending on the association information, and each of the N syntax elements carries information about one of M characteristics, e.g., one of M=3 binary characteristics, and therefore each is dyadic, with three possibilities for associating two flags from three to two, i.e.,

number

[0019] According to an embodiment of the present application, a method includes: decoding video comprising a plurality of pictures from a video data stream by decoding each picture from one or more video coding units within an access unit of the video data stream associated with each picture; reading alternative coding unit types from a parameter set unit of the video data stream; for each given video coding unit, reading a coding unit type identifier from each coding unit; determining whether the coding unit identifier identifies a coding unit type from a first subset of one or more coding unit types or from a second subset of coding unit types; if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, attributing each video coding unit to the alternative coding unit type; and attributing the coding unit identifier to a coding unit type from the second subset of coding unit types identified by the coding unit identifier.

[0020] According to an embodiment of the present application, a method includes decoding a video including a plurality of pictures from a video data stream by decoding each picture from one or more video coding units within an access unit of the video data stream associated with each picture, each video coding unit carrying picture block partitioning data, block-related prediction parameters, and prediction residual data, the access unit associated with a different region of a picture within which each given video coding unit resides; reading from each given video coding unit an n-ary set of one or more syntax elements, e.g., two flags each being binary and thus the pair being quaternary; the mapping may be fixed by default or signaled in the data stream or both may be signaled by dividing a range of values; and dividing the n-ary set of one or more syntax elements into an m-ary set of one or more characteristics, e.g., three binary characteristics each being , 2 terms, so that the triplet is 8 terms, and each characteristic describes in a redundant way the corresponding data in the given video coding unit, i.e., the access unit is associated, and the characteristics may be inferred from an examination of deeper coded data regarding how the video is coded into the video data stream for the picture the given video coding unit resides within, where m>n, or by reading N syntax elements (210) from each of the given video coding units, e.g., each of N=2 flags is 2 terms (N>0) and indicates association information to the video data stream, i.e., by associating them, i.e., treating them as variables of the associated characteristics depending on the association information, and each of the N syntax elements having information about one of M characteristics, e.g., one of M=3 binary characteristics, and therefore each is 2 terms, with three possibilities for associating two flags from three to two, i.e.,

number

[0021] According to an embodiment of the present application, a method includes encoding a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with each picture; indicating alternative coding unit types within a parameter set unit of the video data stream; and defining, for each given video coding unit, a respective coding unit type identifier (100), the coding unit identifier identifying a coding unit type from a first subset (102) of one or more coding unit types or from a second subset (104) of coding unit types, and attributing each video coding unit to a coding unit type if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, and attributing each video coding unit to a coding unit type from the second subset of coding unit types identified by the coding unit identifier if the coding unit identifier identifies a coding unit type from the second subset of coding unit types.

[0022] According to an embodiment of the present application, a method comprises constructing a video-encoded video data stream including a plurality of pictures, each picture being in one or more video coding units within an access unit of the video data stream, one or more coding units for each tile into which the picture is divided, associated with each picture; changing an alternative coding unit type in a parameter set unit of the video data stream to indicate a non-RAP type; identifying exclusively coded video coding units in pictures of the video data stream, the coding unit type being an identifier (100) encoded into the video data stream, identifying a RAP picture; and, for each predetermined video coding unit of the video data stream, identifying an exclusively coded video coding unit in the video data stream, the coding unit type being an RAP picture. For each coding unit, the identifier (100) of each video coding unit encoded in the video data stream identifies a coding unit type from a first subset (102) of one or more coding unit types or from a second subset (104) of coding unit types, where if the coding unit identifier identifies a coding unit type from the first subset of one or more coding unit types, then each given video coding unit should be assigned to the alternative coding unit type, and if the coding unit identifier identifies a coding unit type from the second subset of coding unit types, then each given video coding unit should be assigned to a coding unit type from the second subset of coding unit types identified by the coding unit identifier.

[0023] According to an embodiment of the present application, a method includes encoding a video including a plurality of pictures into a video data stream by encoding each picture into one or more video coding units within an access unit of the video data stream associated with each picture, each video coding unit carrying picture block partitioning data, block-related prediction parameters, and prediction residual data, the access unit being associated with a different region of the picture within which each given video coding unit resides, indicating an n-ary set of one or more syntax elements for each given video coding unit, e.g., two flags each being binary and thus the pair being quaternary, the mapping (200) may be fixed by default or signaled in the data stream or both signaled by dividing a range of values, the n-ary set of one or more syntax elements being divided into an m-ary set of one or more characteristics (20), e.g., three binary characteristics Each is binary, resulting in an 8-ary triplet, and each characteristic describes in a redundant way the corresponding data in a given video coding unit, i.e., the access unit is associated with, and the characteristics may be inferred from an examination of deeper encoded data regarding how video is encoded into the video data stream for the picture within which the given video coding unit resides, where m>n, or by reading N syntax elements (210) from each of the given video coding units, e.g., each of N=2 flags is binary (N>0) and indicates association information to the video data stream, i.e., by associating, i.e., treating them as variables of the associated characteristics depending on the association information, and each of the N syntax elements having information about one of M characteristics, e.g., one of M=3 binary characteristics, and therefore each is binary, with three possibilities for associating two flags from three to two, i.e.,

number

[0024] Preferred embodiments of the present application are described below with reference to the drawings. [Brief explanation of the drawings]

[0025] [Figure 1] A schematic diagram showing a client and server system for a virtual reality application is shown as an example in which the embodiments given in the following figures may be advantageously utilized. [Figure 2] 2 shows a schematic diagram illustrating an example of 360-degree video in a dual-resolution cube-map projection tiled into 6x4 tiles that can be adapted to the system of FIG. 1. [Figure 3] 3 shows a schematic diagram illustrating an example of user viewpoint and tile selection for 360-degree video streaming as shown in FIG. 2. [Figure 4] 4 shows a schematic diagram illustrating an example of the arrangement (packing) of tiles obtained as shown in FIG. 3 in the joint bitstream after the merging process. [Figure 5] FIG. 10 shows a schematic diagram illustrating an example of tiling with low-resolution fallback as a single tile for 360-degree video streaming. [Figure 6] 1 shows a schematic diagram illustrating an example of tile selection in tile-based streaming. [Figure 7] FIG. 10 shows a schematic diagram illustrating another example of tile selection in tile-based streaming with non-uniform Random Access Point (RAP) duration per tile. [Figure 8] 1 shows a schematic diagram illustrating an example of a Network Adaptive Layer (NAL) unit header according to an embodiment of the present application; [Figure 9]1 illustrates an example of a table showing types of Row Byte Sequence Payload (RBSP) data structures included in NAL units according to an embodiment of the present application. [Figure 10] 1 shows a schematic diagram illustrating an example of a sequence parameter set indicating which NAL unit types are mapped according to an embodiment of the present application; [Figure 11] 10 shows a schematic diagram illustrating an example of slice characteristics indicated in a slice header. [Figure 12] 1 shows a schematic diagram illustrating an example of a map indicating characteristics of slices in a sequence parameter set according to an embodiment of the present application; [Figure 13] 13 shows a schematic diagram illustrating an example of characteristics indicated by a map in the sequence parameter set of FIG. 12 according to an embodiment of the present application. [Figure 14] 1 shows a schematic diagram illustrating an example of association information indicated by utilizing a syntax structure according to an embodiment of the present application; [Figure 15] 10 shows a schematic diagram illustrating another example of a map indicating characteristics of slices in a sequence parameter set according to an embodiment of the present application; [Figure 16] 16 shows a schematic diagram illustrating another example of characteristics shown by a map in the parameter set of FIG. 15 according to an embodiment of the present application. [Figure 17] 10 shows a schematic diagram illustrating a further example of a map indicating slice characteristics in a slice segment header according to an embodiment of the present application; [Figure 18] 18 shows a schematic diagram illustrating further examples of properties indicated by maps in the slice segment header of FIG. 17 according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0026] In the following description, equal or equivalent elements having equal or equivalent functions are designated by equal or equivalent reference numerals.

[0027] In the following description, numerous details are set forth to provide a more thorough explanation of the embodiments of the present application. However, it will be apparent to those skilled in the art that the embodiments of the present application may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the embodiments of the present application. Furthermore, features of the embodiments described below may be combined with each other unless otherwise specified.

[0028] Introduction It should be noted below that the individual aspects described herein may be used individually or in combination, and thus details may be added to each individual aspect without adding details to any other one of the aspects.

[0029] It should also be noted that this disclosure explicitly or implicitly describes features usable in a video decoder (a device for providing a decoded representation of a video signal based on an encoded representation), and thus any of the features described herein can be used in the context of a video decoder.

[0030] Furthermore, features and functions disclosed herein in the context of a method may also be utilized in an apparatus (configured to perform such functions). Furthermore, any feature or function disclosed herein in the context of an apparatus may also be utilized in the corresponding method. In other words, a method disclosed herein may be supplemented by any of the features and functions described in the context of an apparatus.

[0031] To facilitate understanding of the description of the embodiments of the present application with respect to various aspects thereof, Figure 1 illustrates an example environment in which the embodiments described below of the present application may be applied and effectively utilized. In particular, Figure 1 illustrates a system comprising a client 10 and a server 20 interacting via adaptive streaming. For example, Dynamic Adaptive Streaming over HTTP (DASH) may be utilized for communication 22 between the client 10 and the server 20. However, the embodiments outlined below should not be construed as being limited to the use of DASH, and similarly, terms such as Media Presentation Description (MPD) should be understood to be broad enough to include manifest files, which are defined differently from DASH.

[0032] FIG. 1 illustrates a system configured to realize a virtual reality application. That is, the system is configured to present to a user wearing a head-up display 24, i.e., via an internal display 26 of the head-up display 24, a view portion 28 from a time-varying spatial scene 30, the portion 28 corresponding to the orientation of the head-up display 24, as illustratively measured by an internal orientation sensor 32, such as an inertial sensor, of the head-up display 24. That is, the portion 28 presented to the user forms a portion of the spatial scene 30 at a spatial location corresponding to the orientation of the head-up display 24. In the case of FIG. 1, the time-varying spatial scene 30 is depicted as an omnidirectional or spherical video, but the description of FIG. 1 and the embodiments described below can easily be extended to other examples, such as presenting a portion from a video having the spatial location of the portion 28 determined by the intersection of a face access or eye access with a virtual or real projector wall, etc. Furthermore, the sensor 32 and the display 26 may be constituted by different devices, e.g., a remote control and a corresponding television, respectively, or they may be part of a handheld device, such as a mobile device, e.g., a tablet or a mobile phone. Finally, it should be noted that some of the embodiments described below may also be applied to scenarios in which the region 28 presented to the user always covers the entire time-varying spatial scene 30 when presenting a time-varying spatial scene, e.g., with non-uniformity associated with an uneven distribution of quality across the spatial scene.

[0033] Further details regarding the server 20, the client 10, and the manner in which the spatial content 30 is provided at the server 20 are shown in Figure 1 and described below, but these details should also not be treated as limiting the embodiments described below, but rather should serve as examples of how to implement any of the embodiments described below.

[0034] In particular, as shown in Figure 1, the server 20 may comprise a storage 34 and a controller 36, such as a suitably programmed computer, application specific integrated circuit or the like. The storage 34 stores media portions representing a time-varying spatial scene 30. An example is outlined in more detail below with respect to the diagram of Figure 1. The controller 36 responds to requests sent by the client 10 by re-sending the media presentation description for the media portions requested by the client 10, and may also send further information of its own to the client 10. More details in this regard are provided below. The controller 36 may fetch the requested media portions from the storage 34. Within this storage, other information, such as the media presentation description or parts thereof, may be stored in other signals sent from the server 20 to the client 10.

[0035] 1, the server 20 may optionally further comprise a stream modifier 38 for modifying media portions sent from the server 20 to the client 10 in response to a request from the latter, so that, for example, the media portions thus retrieved by the client 10 result in a media data stream at the client 10 that is in fact aggregated from several media streams but forms one single media stream decodable by one associated decoder. However, the presence of such a stream modifier 38 is optional.

[0036] The client 10 of FIG. 1 is exemplarily shown as including a client device or controller 40 or more decoders 42 and a reprojector 44. The client device 40 may be a suitably programmed computer, a microprocessor, a programmed hardware device such as an FPGA, or an application-specific integrated circuit, etc. The client device 40 assumes responsibility for selecting the media portions to be extracted from the server 20 from the plurality of media portions 46 provided therein. To this end, the client device 40 first extracts a manifest or media presentation description from the server 20. From the same, the client device 40 obtains calculation rules for calculating the addresses of the media portions from the plurality of media portions 46 that correspond to a specific required spatial portion of the spatial scene 30. The media portions thus selected are extracted from the server 20 by the client device 40 by sending respective requests to the server 20. These requests include the calculated addresses.

[0037] The media portions extracted by client device 40 in this manner are forwarded by the latter to one or more decoders 42 for decoding. In the example of FIG. 1 , the media portions thus extracted and decoded represent only a spatial portion 48 from the time-varying spatial scene 30 for each time unit in time, but as already mentioned above, this may differ according to other aspects, for example, where the presented view portion 28 always covers the entire scene. A reprojector 44 may optionally reproject and cut out the view portion 28 displayed to the user from the extracted and decoded scene content of the selected, extracted, and decoded media portion. To this end, as shown in FIG. 1 , client device 40 may continuously track and update the spatial location of view portion 28, for example, in response to user orientation data from sensor 32, and may, for example, inform reprojector 44 of the current spatial location of scene portion 28 and the reprojection mapping to be applied to the extracted and decoded media content so as to map it to the region forming view portion 28. In response, reprojector 44 may apply mapping and interpolation onto a regular pixel grid, for example, as displayed on display 26 .

[0038] FIG. 1 illustrates the case where cubic mapping is used to map the spatial scene 30 onto tiles 50. Thus, tiles are depicted as rectangular subregions of a cube onto which the spherically shaped scene 30 is projected. The reprojector 44 reverses this projection. However, other examples may be applied as well. For example, instead of a cubic projection, a projection onto a truncated pyramid or a pyramid without truncation may be used. Furthermore, although the tiles in FIG. 1 are shown as non-overlapping in terms of coverage of the spatial scene 30, the division into tiles may include mutual tile overlap. Also, as outlined in more detail below, the spatial division of the scene 30 into tiles 50, with each tile forming a single representation as described below, is not required.

[0039] Thus, as depicted in FIG. 1, the entire spatial scene 30 is spatially subdivided into tiles 50. In the example of FIG. 1, each of the six faces of a cube is subdivided into four tiles. For illustrative purposes, the tiles are enumerated. For each tile 50, the server 20 provides a video 52 as shown in FIG. 1. More precisely, the server 20 provides multiple videos 52 per tile 50, with these videos having different qualities Q#. Furthermore, the videos 52 are temporally subdivided into time segments 54. The time segments 54 of all the videos 52 of all the tiles T# form one of multiple media portions 46 stored in the storage 34 of the server 20 or are respectively encoded.

[0040] It is again emphasized that even the illustrative example of tile-based streaming shown in Figure 1 merely forms an illustrative example, for which many variations are possible. For example, while Figure 1 seems to suggest that media portions associated with a higher quality representation of scene 30 are associated with tiles that match the tiles to which media portions in which scene 30 is encoded with quality Q1 belong, this correspondence is not necessary, and tiles of different qualities may correspond to tiles of different projections of scene 30. Furthermore, although not discussed thus far, media portions corresponding to different quality levels depicted in Figure 1 may differ in spatial resolution, signal-to-noise ratio, and / or temporal resolution, etc.

[0041] Finally, unlike the tile-based streaming concept in which media portions that may be individually extracted by device 40 from server 20 relate to tiles 50 into which scene 30 is spatially subdivided, media portions provided at server 20 may alternatively have full sampling resolution at different spatial locations in scene 30, e.g., each having scene 30 encoded therein in a spatially complete manner with spatially varying sampling resolution. For example, this can be achieved by providing at server 20 a sequence of segments 54 that relate to the projection of scene 30 onto truncated pyramids, the tips of which are oriented in different directions relative to one another, thereby leading to resolution peaks oriented in different directions.

[0042] Furthermore, it should be noted that, optionally, with respect to the present stream modification device 38, the same may be part of the client 10 or may be located between the client 10 and the server 20 in the network device through which the client 10 and the server 20 exchange signals as described herein.

[0043] There are certain video-based applications in which multiple encoded video bitstreams are jointly decoded, i.e., merged into a joint bitstream and fed to a single decoder as follows: Multi-party conferencing: Encoded video streams from multiple participants are processed on a single endpoint. Or tile-based streaming: e.g., 360-degree tiled video playback in VR applications.

[0044] In the latter, 360-degree video is spatially segmented and each spatial portion is provided to the streaming client in multiple representations at various spatial resolutions, as shown in Figure 2. Figure 2(a) shows a high-resolution tile, and Figure 2(b) shows a low-resolution tile. Figure 2, i.e., Figures 2(a) and (b), show a cube map projecting 360-degree video divided into 6x4 spatial portions at two resolutions. For simplicity, these independently decodable spatial portions are referred to herein as tiles.

[0045] When using a state-of-the-art head-mounted display, a user typically sees only a subset of the tiles that make up the entire 360-degree video through a solid viewport boundary 80 that represents a 90x90 degree field of view, as shown in Figure 3(a). The corresponding tiles, indicated by reference numeral 82 in Figure 3(b), are downloaded at full resolution.

[0046] However, to accommodate sudden changes in the user's orientation, the client application must also download and decode representations of other tiles outside the current viewport, as indicated by reference numeral 84 in FIG. 3(c). Thus, as shown in FIG. 3(c), a client of such an application downloads tiles covering the current viewport at the highest resolution and tiles outside the current viewport at a relatively lower resolution, while the tile resolution selection is constantly adapted to the user's orientation. After client-side downloading, merging the downloaded tiles into a single bitstream processed by a single decoder is a means of addressing the constraints of a typical mobile device with limited computational and power resources. FIG. 4 illustrates a possible tile arrangement in a joint bitstream for the above example. The merging process to generate the joint bitstream must be performed by avoiding compressed-domain processing, i.e., processing on the pixel domain via transcoding.

[0047] While the example from Figure 4 shows a case where all tiles (high-resolution and low-resolution) cover the entire 360-degree space and no tile repeatedly covers the same area, another tile grouping is also available, as shown in Figure 5. That is, the entire low-resolution portion of the video is defined as a "low-resolution fallback" layer, as shown in Figure 5(b), which can be merged with the high-resolution tiles from Figure 5(a) that cover a subset of the 360-degree video. The entire low-resolution fallback video can be coded as a single tile, as shown in Figure 5(c), while the high-resolution tiles are rendered as an overlay on the low-resolution portion of the video in the final stage of the rendering process.

[0048] The client starts a streaming session according to the user's tile selection by downloading all desired tile tracks, as shown in Figure 6, where the client starts the session with tile 0, indicated by reference numeral 90, and tile 1, indicated by reference numeral 92, in Figure 6(a). Whenever a viewport change occurs (i.e., the user turns their head to look in another direction), the tile selection changes in the next available segment, i.e., tile 0 and tile 2, indicated by reference numeral 94 in Figure 6(a), and in the next available segment, the client changes the position of tile 2 and replaces tile 0 with tile 1, as shown in Figure 6(b). It is important to note that every segment needs to start with an IDR (Instantaneous Decoder Refresh) picture, i.e., a prediction chain reset picture; a new tile selection and tile position change would otherwise cause prediction mismatch, artifacts, and drift.

[0049] Encoding each portion with an IDR picture is costly in terms of bitrate. Portions can potentially be very short in duration, e.g., to react quickly to orientation changes, which is why it is desirable to encode multiple variants with varying IDR (or RAP: Random Access Point) durations, as shown in Figure 7. For example, as shown in Figure 7(b), at time t1, there is no reason to break the prediction chain for tile 0, since tile 0 has already been downloaded for time t0 and is located in the same position that the client can select a portion that does not start with a RAP available at the server.

[0050] However, one remaining issue is that slices (tiles) within a coded picture must obey certain constraints. One of them is that a picture may not simultaneously contain Network Abstract Layer (NAL) units of RAP and non-RAP NAL unit types. Therefore, applications have only two less desirable options to address the above issue. First, clients can rewrite the NAL unit type of RAP pictures when they are merged into pictures with non-RAP NAL units. Second, servers can obscure the RAP characteristics of these pictures by using non-RAP from the beginning. However, this prevents the detection of RAP characteristics in systems that must process these coded videos, for example, for file format packaging.

[0051] The present invention is a NAL unit type mapping that allows mapping one NAL unit type to another NAL unit type via an easily rewritable syntax structure.

[0052] In one embodiment of the present invention, a NAL unit type is designated as mappable, and the mapped type is specified in a parameter set, for example, based on Draft 6 V14 of the Versatile Video Coding (VVC) specification with highlighted editing, as follows:

[0053] Figure 8 shows the NAL unit header syntax. The syntax nal_unit_type, i.e., identifier 100, specifies the NAL unit type, i.e., the type of RBSP (Row Byte Sequence Payload) data structure contained in the NAL unit as specified in the table shown in Figure 9.

[0054] The variable NalUnitType is defined as follows:

number

[0055] All references to the syntax element nal_unit_type in the specification are replaced with references to the variable NalUnitType, as well as the following constraints:

[0056] The value of NalUnitType is the same for all coded slice NAL units of a picture. A picture or layer access unit is referred to as having the same NAL unit type as its coded slice NAL unit. That is, as shown in FIG. 9, a first subset of coding unit types 102 indicates "nal_unit_type" 12, i.e., "MAP_NUT" and "VCL" as NAL unit type classes. Accordingly, a second subset 104 of coding unit types indicates "VCL" as a NAL unit type class, i.e., all coded slice NAL units of a picture indicated by numbers 0 to 15 of identifier 100 have the same NAL unit type class of first subset of coding unit types 102, i.e., a coding unit type of VCL.

[0057] FIG. 10 shows a sequence parameter set RBSP syntax including mapped_nut as indicated by reference numeral 106, which indicates a NalUnitType of NAL units with nal_unit_type equal to MAP_NUT.

[0058] In another embodiment, the mapped_nut syntax element is carried in the access unit delimiter AUD.

[0059] In other embodiments, it is a bitstream conformance requirement that the value of mapped_nut must be a VCL NAL unit type.

[0060] In another embodiment, the mapping of NalUnitTypes of NAL units whose nal_unit_type is equal to MAP_NUT is performed by profiling information. Such a mechanism makes it possible to have more than one mappable NAL unit type instead of having a single MAP_NUT and to indicate the required interpretation of the NALUnitTypes of the mappable NAL units within a simple profiling mechanism or a single syntax element mapped_nut_space_idc.

[0061] In other embodiments, the mapping mechanism is used to extend the range of NAL unit types, which is currently limited to 32 (e.g., because it is u(5) as shown in FIG. 10). The mapping mechanism can indicate any unrestricted value, as long as the number of required NAL unit types does not exceed the number of values ​​reserved for mappable NAL units.

[0062] In one embodiment, when a picture simultaneously contains slices of an alternative coding unit type and slices of a normal coding unit type (e.g., existing NAL units of the VCL category), a mapping is performed to result in all slices of the picture having effectively the same coding unit type characteristics, i.e., the alternative coding unit type is equal to the coding unit type of the non-alternative slice of the normal coding type. Furthermore, the above embodiment only holds for pictures with or without the random access property.

[0063] In addition to the described problems regarding NAL unit types and NAL unit type extensibility in merge scenarios and corresponding solutions, there are some video applications where information related to the video and how the video was coded, such as on-the-fly adaptation, is required for system integration and transmission or operation.

[0064] There are some common pieces of information established within the last few years that are widely used in the industry and are clearly designated, and specific bit values ​​are used for such purposes. Temporary ID in NAL unit header NAL unit types, including IDR, CRA, TRAIL, or SPS (Sequence Parameter Set), PPS (Picture Parameter Set), etc. is.

[0065] However, there are some scenarios where additional information can be useful. Although not widely used, some cases have found some utility, e.g., BLA, partial RAP NAL units for subpictures, sub-layer non-reference NAL units, etc. If the extensibility mechanisms described above are used, some of these NAL unit types are feasible. However, as an alternative, some fields in the slice header can also be used.

[0066] Conventionally, additional information is reserved in the slice header that is used to signal specific characteristics of the slice. Discardable flag: indicates that the coded picture can be used as a reference picture for inter prediction and not as a source picture for inter-layer prediction. · cross_layer_bla_flag: Affects the derivation of output pictures for layered coding, and pictures preceding the RAP in higher layers may not be output.

[0067] Similar mechanisms may be envisioned for upcoming video codec standards. However, one constraint of these mechanisms is that the defined flags occupy specific positions in the slice header. In the following, the use of these flags in HEVC is shown in Figure 11.

[0068] As mentioned above, the problem with such a solution is that the positions of the extra slice header bits are allocated in stages, and for applications that utilize rarer information, flags will likely come in later positions, increasing the number of bits that need to be transmitted in the extra bits (e.g., in the case of HEVC, "discardable_flag" and "cross_layer_bla_flag").

[0069] Alternatively, following a similar mechanism described for NAL unit types, the mapping of flags in extra slice header bits in the slice header can be defined in the parameter set. An example is shown as FIG.

[0070] FIG. 12 shows an example of a sequence parameter set including a map indicating association information using "extra_slice_header_bits_mapping_space_idc", i.e., a map indicated by reference numeral 200 indicating the mapping space for extra bits of the slice header.

[0071] Figure 13 shows the mapping of bits to the flags indicated by "extra_slice_header_bits_mapping_space_idc" 200 in Figure 12. As shown in Figure 13, binary characteristics 202 redundantly describe corresponding data in a given video coding unit. In Figure 13, three binary characteristics 202 are shown, namely, "0", "1", and "2". The number of binary characteristics is variable depending on the number of flags.

[0072] In other embodiments, the mapping is performed in a syntax structure (e.g., as shown in FIG. 14) that indicates the presence or absence of syntax elements in the extra bits in the slice header (e.g., as shown in FIG. 11). That is, for example, as shown in FIG. 11, the condition on the presence flag controls the presence of syntax elements in the additional bits in the slice header, i.e., "num_extra_slice_header_bits>i" and "i<num_extra_slice_header_bits i++". In FIG. 11, as described above, each syntax element in the additional bits is placed at a specific position in the slice header, but in this embodiment, it is not necessary for a syntax element, e.g., "discardable_flag", "cross_layer_bla_flag" or "slice_reserved_flag[i]" to occupy a specific position. Instead, when checking the condition regarding the value of a specific presence flag (e.g., "discardable_flag_present_flag" in FIG. 14), when it is shown that the first syntax element (e.g., "discardable_flag" in FIG. 11) does not exist, the following second syntax element, when it exists, takes the position of the first syntax element in the slice header. Also, the syntax elements in the additional bits can be present in the picture header by indicating a flag, e.g., by indicating "sps_extra_ph_bit_present_flag[i]". Also, the syntax structure, e.g., a plurality of syntax elements, i.e., the number of flags presented, indicates the presence / absence of a specific characteristic or the number of syntax elements presented in the additional bits. That is, the number of a specific characteristic or syntax element presented in the additional bits is indicated by counting the number of syntax elements (flags) presented. In FIG. 14, the presence of each syntax element is indicated by flag 210. That is, each flag in 210 indicates the presence of a slice header indication for a specific characteristic of a given video coding unit.Furthermore, the further syntax "[...] / / additional flag" shown in Figure 14 and corresponding to the "slice_reserved_flag[i]" syntax element in the slice header of Figure 11 is used as a placeholder to indicate the presence / absence of a syntax element in the additional bits or as an indication of the presence / absence of an additional flag.

[0073] In another embodiment, the flag type mapping is signaled for each extra slice header bit in the parameter set, as shown in Figure 15. As shown in Figure 15, the syntax "extra_slice_header_bit_mapping_idc", i.e., map 200, is signaled in the sequence parameter set and indicates the location of the mapped feature.

[0074] Figure 16 shows the mapping of bits to the flag indicated by "extra_slice_header_bits_mapping_space_idc" in Figure 15. That is, a binary characteristic 202 corresponding to the map 200 shown in Figure 15 is shown in Figure 16.

[0075] In other embodiments, the slice header extension bit is replaced by idc signaling, which indicates a specific flag value combination, as shown for example in Figure 17. As shown in Figure 17, a map 200, i.e., "extra_slice_header_bit_idc", is indicated in the slice segment header, i.e., the map 200 indicates the presence of a property as shown in Figure 18.

[0076] FIG. 18 shows that the flag value, i.e., binary characteristic 202, represented by a particular value of "extra_slice_header_bit_idc" is either signaled in the parameter set or predefined in the specification (known in advance).

[0077] In one embodiment, the value space of "extra_slice_header_bit_idc", i.e., the value space of map 200, is divided into two ranges: one range representing the combinations of flag values ​​that are known a priori, and one range representing the combinations of flag values ​​that are signaled in the parameter set.

[0078] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent descriptions of corresponding methods, and that blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of a method step also represent descriptions of a corresponding block, item, or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0079] The data stream of the present invention may be stored on a digital storage medium or may be transmitted over a transmission medium, such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0080] Depending on specific implementation requirements, the embodiments of the present application can be realized in hardware or software. The realization can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, on which electronically readable control signals are stored that cooperate (or can cooperate) with a programmable computer system to execute the respective methods. Thus, the digital storage medium may be computer-readable.

[0081] Some embodiments according to the invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system to cause one of the methods described herein to be performed.

[0082] Generally, embodiments of the present application can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer, which program code may for example be stored on a machine-readable carrier.

[0083] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0084] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0085] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer readable medium) comprising, recorded thereon, a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.

[0086] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or the sequence of signals being adapted to be transmitted via a data communication connection, for example the Internet.

[0087] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0088] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0089] Further embodiments according to the present invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.

[0090] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be utilized to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0091] The apparatus described herein may be implemented using a hardware apparatus, a computer, or a combination of a hardware apparatus and a computer.

[0092] The devices described herein, or any components of the devices described herein, may be implemented at least in part in hardware and / or software.

Claims

1. 1. A method of video decoding, comprising: receiving a sequence parameter set (SPS) from a video data stream; Parsing an extra slice header bitmap from the SPS; mapping one or more bits in the extra slice header bitmap to one or more flags indicating the presence or absence of corresponding syntax elements in a slice header; determining a number of extra slice header bits based on the number of the one or more flags indicating presence; parsing syntax elements in the slice header based on the determined number of extra slice header bits from the video data stream; A method comprising:

2. The method of claim 1 , further comprising: determining a position of the syntax element in the slice header based on the determined number of extra slice header bits.

3. The method of claim 1 , wherein the number of the one or more flags indicating presence is determined based on counting the number of the one or more flags indicating presence.

4. The method of claim 1 , wherein each of the one or more flags has a one-bit value.

5. 1. A method of video encoding, comprising: providing a sequence parameter set (SPS) via a video data stream; encoding an extra slice header bitmap into the SPS, the extra slice header bitmap mapping one or more bits in the extra slice header bitmap to one or more flags indicating the presence or absence of corresponding syntax elements in the slice header; determining a number of extra slice header bits based on the number of the one or more flags indicating presence; providing the syntax element in the slice header based on the determined number of extra slice header bits throughout the video data stream; A method comprising:

6. The method of claim 5 , further comprising determining a position of the syntax element in the slice header based on the determined number of extra slice header bits.

7. The method of claim 5 , wherein the number of the one or more flags indicating presence is determined based on counting the number of the one or more flags indicating presence.

8. The method of claim 5 , wherein each of the one or more flags has a one-bit value.

9. A video coding device comprising a processing circuit configured to perform the method according to any one of claims 1 to 8.

10. A non-transitory computer readable medium having instructions that, when executed by a processing circuit, perform the method of any one of claims 1 to 8.