File format suitable for fusion
The proposed file format addresses the inefficiencies in merging and processing spatial video subsets by using source track groups and fusion information, resulting in reduced overhead and latency for tiled streaming applications.
Patent Information
- Application Number
- JP2022519251
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-27
- Filing Date
- 2020-09-28
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2040-09-28
AI Technical Summary
Current file formats for encoded video, such as those using HEVC and VVC, face challenges in efficiently merging and processing spatial subsets of video data, particularly in tiled streaming applications for 360-degree video playback. This results in significant overhead, increased latency, and complexity in handling tile resolution changes and encryption.
A specific file format is proposed that includes a set of source tracks, each representing a spatial portion of a video scene, with group indicators and active source track indicators to manage source track groups. This format also includes fusion information for fusing subsets of source tracks and configurable parameters or SEI messages to adapt to different video stream requirements.
The proposed file format reduces overhead and latency by enabling efficient merging and processing of spatial video subsets, allowing for flexible tile resolution management and improved encryption complexity, thereby enhancing the performance of tiled streaming applications.
Smart Images

Figure 0007695234000013 
Figure 0007695234000014 
Figure 0007695234000015
Abstract
Description
Technical Field
[0001] The present application relates to a file format that enables extraction or fusion of spatial subsets of encoded video using compression domain processing. Specifically, the present application relates to video data for deriving spatially varying portions of a scene, methods and apparatuses for creating video data for deriving spatially varying portions of a scene, and methods and apparatuses for deriving spatially varying portions of a scene from video data formatted in a specific file format. The present application also relates to corresponding computer programs, computer-readable media, and digital storage media.
[0002] 1. Introduction Typically, encoded video data such as data video encoded with AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), or currently under development VVC (Versatile Video Coding) is stored or transmitted in a specific container format such as the ISO base media file format defined in ISO / IEC 14496-12 (Coding of Audio-Visual Objects - Part 12: ISO Base Media File Format), ISO / IEC 14496-15 (Coding of Audio-Visual Objects - Part 12: Transport of Network Abstraction Layer (NAL) Unit Structured Video in the ISO Base Media File Format), ISO / IEC 23008-12 (High Efficiency Coding and Media Delivery in Heterogeneous Environments - Part 12: Image File Format), and various extensions thereof. Such container formats include special provisions targeted at applications that rely on compression domain processing for extraction or combination of spatial subsets of encoded video, for example, for the purpose of using a single decoder on an end device. A non-exhaustive list of such application examples is as follows. · Region of Interest (RoI) streaming that transmits a changing spatial subset of the video · A multi-party conference that decodes encoded video streams from multiple participants together with a single decoder, or, · Tile-based streaming for 360-degree video playback, for example in VR applications.
[0003] 1.1 Tiled Streaming of 360-Degree Video In the latter case, as shown in FIG. 1, the 360-degree video of the scene is spatially segmented, and each spatial segment is provided to the streaming client in multiple representations of different spatial resolutions. The figure shows a cube map-projected 360-degree video divided into 6×4 spatial segments (including left, front, right, rear, bottom, and top) at two resolutions (high resolution and low resolution). For simplicity, in this specification, these independently decodable spatial segments are referred to as tiles. Depending on the selected video encoding technology, structures such as tiles, bricks, and slices can be used to achieve independent encoding of different spatial segments. For example, when encoding each tile in the currently under-development VVC (Versatile Video Coding), this can be achieved by dividing the picture using a suitable tile / brick / slice structure so that, for example, no intra prediction or inter prediction is performed between different tiles / bricks of the same or different pictures. For example, a single tile can be used as a separate slice to encode each independently decodable spatial segment, or the concept of bricks can be further used to perform more flexible tiling.
[0004] As shown at the top of FIG. 2, typically, when using a modern head-mounted display (HMD), the user sees only a subset of the tiles that make up the entire 360-degree video through the solid-line viewport boundary representing a 90×90-degree field of view (FoV). The corresponding tiles shown as shaded areas at the top of FIG. 2 (in this example, four tiles on the right, two tiles on the bottom, one tile on the front, and one tile on the rear) are downloaded at the highest resolution (also shown shaded in the lower left of the figure).
[0005] However, in order for the client application to respond to a sudden change in the user's orientation, it is also necessary to download and decode the representation of other tiles outside the current viewport (not shaded in the upper part of FIG. 2) shown with a different shading in the lower right of FIG. 2. Therefore, the client of such an application downloads the tiles covering the current viewport at the highest resolution and downloads the tiles outside the current viewport at a relatively low resolution, while the selection of the tile resolution is always adapted to the user's orientation. After downloading on the client side, merging the downloaded tiles into a single bitstream so that they are processed by a single decoder is a means to cope with the constraints of a typical mobile device with limited computing resources and power resources. FIG. 3 shows a possible tile arrangement in the joint bitstream of the above example. The merging operation for generating the integrated bitstream needs to be performed at the bitstream level through compression domain processing in order to avoid complex processing in the pixel domain such as transcoding or decoding these tiles independently of each other before synchronously rendering them on a cube.
[0006] The metadata description in the form of so-called supplemental enhancement information (SEI) messages in a symbolized video bitstream describes how the samples of an encoded image are originally related to the positions within the original projection (in this example, a cube map) in order to enable the reconstruction of a cube (or a sphere depending on the projection used) in 3D space. This metadata description, called region-wise-packing (RWP), is essential for a post-decoding renderer that renders the viewport of a media consumption device such as a head-mounted display (HMD). The RWP SEI message defines the displacement / transformation between a rectangular region and its projected video and the packed video, and thus shows the mapping from the projected video (conceptually required for further processing after decoding, as shown on the left side of FIG. 1) and from one specific combination of packed encoded videos (as shown on the right side of FIG. 4 obtained by decoding the integrated bitstream or FIG. 3).
[0007] The examples in FIGS. 1 to 3 show a case where all resolution versions of the content are tiled in the same way, all tiles (high resolution and low resolution) cover the entire 360-degree space, and there are no tiles that repeatedly cover the same region, but other tilings can also be used as shown in FIG. 4. The entire low-resolution version of the video can be merged with high-resolution tiles that cover a subset of the 360-degree video. While the entire low-resolution fallback video can be encoded as one tile, the high-resolution tiles are rendered as an overlay of the low-resolution part of the video in the final stage of the rendering process.
[0008] 1.2 Problems with Tiled Streaming and File Formats Using HEVC In codecs such as HEVC, the need for the merging operation as seen from the video bitstream is related to the tile structure of the picture and the CTU (Coded Tree Unit) addressing signaling of individual tiles (i.e., slices). On the server side, these tiles exist (and thus are downloaded) as individual independent HEVC bitstreams, for example, each of these bitstreams contains a single tile and slice per picture (e.g., first_slice_in_pic_flag equal to 1 in all slice headers, parameter sets that describe bitstreams with only a single tile). In this merging operation, it is necessary to integrate these individual bitstreams into one bitstream by inserting the correct parameter sets and slice headers so as to reflect the tile structure and positions within the integrated picture plane. Not only is the client implementation left to handle the details of the merging (derivation and replacement of parameter sets and slice headers), but MPEG OMAF (Immersive Media Encoded Representation - Part 2: Omnidirectional Media Format; ISO / IEC 23090-2) also defines the latest method to enable the client to merge the bitstream through the following means. · Generating the correct parameter sets and slice headers during the packaging stage, and, · Copying the slice payload using a file format tool called an extractor.
[0009] These extractors are actually NAL (Network Abstraction Layer) units of a special NAL unit type that contain pointers to different tracks, i.e., to other NAL units packaged in different tracks (e.g., containing data for one tile) as defined in ISO / IEC 14496-15, which is an extension of the file format. The extractor itself is stored in a special extractor file format track (the "hvc2" track) that carries only parameter sets and modified slice header data (e.g., reflecting a new tile position, an adjusted quantization step size value relative to the parameter set base value, etc.). While the slice payload (i.e., the entropy-coded data that makes up the actual sample values of the picture during decoding) is referenced by the extractor pointing to a NAL unit (or part of it) within another track, and is copied when such a file format track is read.
[0010] In a 360-degree video tile-based streaming system, this extractor tool results in a design where each tile is provided as an independent HEVC stream in a separate file format track that can be packaged and decoded by a compliant HEVC decoder to yield each spatial subset of the full picture. Further, such a set of extractor tracks is provided, each targeting a specific viewing direction (i.e., a combination of tiles at a specific resolution that concentrates decoding resources such as sample budgets on the tiles within the viewport), which execute a fusion process via the file format tool to yield a single compliant HEVC bitstream containing all the necessary tiles upon reading. The client can select the extractor track most suitable for the current viewport and download the track containing the referenced tiles.
[0011] Each extractor track stores the parameter sets in the HEVCConfigurationBox included in the HEVCSampleEntry. These parameter sets are generated in the file format packaging process and are only available for sample entries, i.e., when the client selects an extractor track, the parameter sets are delivered out-of-band (using the initialization segment), so the parameter sets cannot change over time while playing the same extractor track. The initialization segment of the extractor track includes, in addition to the necessary sample entries, a fixed list of dependent trackIDs within the track reference container (‘tref’). The extractor (included in the media segment of the extractor track) includes an index value that references this ‘tref’ to determine which trackID is being referenced by the extractor.
[0012] However, this design has several drawbacks. · Each view direction (or combination of tiles) needs to be represented through a separate extractor track that clearly references the tiles (i.e., tracks) included, which results in a significant overhead. The client may be able to better select a tile resolution that better suits its needs (i.e., create its own combination), taking into account the client's FoV and latency, etc. Also, the data included in such extractor tracks is often very similar throughout the timeline (the inlines and sample constructors remain the same). · Usually, all slice headers need to be adjusted through the extractor, which results in an even more significant overhead. As a result, there are more pointers to dependent tracks, i.e., a large number of buffer copies need to be performed, which is particularly costly in web applications using JavaScript, etc. · If not all data has been completely downloaded in advance, the file format parser cannot resolve the extractor track. This can add additional latency to the system, for example, when all video data (tiles) has been downloaded and the client is waiting for the fetch of the extractor track. · Since partial encryption needs to be applied (the slice payload must be encrypted independently of the slice header), the complexity of such general encryption of the extractor track increases.
[0013] 1.3 Impact of VVC Design and File Format on Tiled Streaming For the next codec generation such as VVC, two main efforts have been made to simplify the extraction / derivation operations in the compressed region.
[0014] 1.3.1 Tiling Syntax in VVC In HEVC, the subdivision of a picture into slices (NAL units) was finally signaled at the slice header level, i.e., by having multiple slices within one tile or multiple tiles within one slice. In VVC, however, the subdivision of a picture into slices (NAL units) is described only in the parameter set. After the first level of division is signaled through the rows and columns of the tiles, the second level of division is signaled through the so-called brick division of each tile. Tiles that do not include further brick division are also called single bricks. The number of slices per picture and the associated bricks are clearly indicated in the parameter set.
[0015] 1.3.2 Slice Address Signaling in VVC For example, previous codecs such as HEVC relied on slice position signals through the first_slice_in_pic_flag and slice_address having encoded lengths that depend on the slice address in CTU raster scan order within each slice header, particularly the picture size. VVC features an indirection of these addresses instead of these two syntax elements. In this case, the slice header carries, as the slice address, an identifier (e.g., brick_id, tile_id, or subpic_id) that is mapped to a specific picture position by the associated parameter set instead of an explicit CTU position. Thus, if tiles should be rearranged in an extraction or fusion operation, only the indirection of the parameter set needs to be adjusted instead of each slice header.
[0016] 1.3.3 VVC Syntax and Semantics Figure 5 shows an excerpt from the VVC specification (draft 6, version 11) incorporating the relevant parts of the currently assumed picture parameter set and slice header syntax of VVC, with line numbers prefixed to the relevant syntax. The syntax elements from line 5 to line 49 of the picture parameter set syntax are related to the tiling structure, and the syntax elements from line 54 to line 61 of the picture parameter set syntax and the slice_address syntax element of the slice header syntax are related to the slice / tile arrangement.
[0017] The semantics of the syntax elements related to the slice / tile arrangement are as follows. slice_id[i] specifies the slice ID of the i-th slice. The length of the slice_id[i] syntax element is signalled_slice_id_length_minus1 + 1 bits. If it does not exist, the value of slice_id[i] is assumed to be equal to i for each i in the range 0 to num_slices_in_pic_minus1 inclusive. The slice_address specifies the slice address of the slice. If it does not exist, the value of slice_address is assumed to be equal to 0. When rect_slice_flag is 0, · The slice address is the block ID defined by Equation (7 - 59), · The length of slice_address is Ceil(Log2(NumBricksInPic)) bits, · The value of slice_address is in the range of 0 to NumBricksInPic - 1, inclusive. Otherwise (when rect_slice_flag is equal to 1), · The slice address is the slice ID of the slice, · The length of slice_address is signalled_slice_id_length_minus1 + 1 bits, · When signalled_slice_id_flag is 0, the value of slice_address is in the range of 0 to num_slices_in_pic_minus1. Otherwise, the value of slice_address is in the range of 0 to 2 (signalled_slice_id_length_minus1+1) - 1.
[0018] The requirements for bitstream compliance are that the following constraints apply. · The value of slice_address shall not be equal to the value of slice_address of any other coded slice NAL unit of the same coded picture. · When rect_slice_flag is 0, the slices of the picture are in ascending order of their slice_address values. · The shape of the slices of the picture shall be such that at decoding, the entire left border and the entire upper border of each block consist of the picture border or the border of a previously decoded (single or multiple) block.
[0019] For example, in the design of future container format integration such as future file format extensions related to the present invention, it is possible to facilitate the change of the VVC high-level syntax with respect to the HEVC high-level syntax. More specifically, the present invention includes aspects that handle the following. · Basic classification into a set of source tracks (tracks of tiles) that can be merged, · Templates for configurable parameter sets and / or SEI messages, · Extended classification for configurable parameter sets and / or SEI messages, and, · Random access point indication in track combinations.
[0020] According to an aspect of the present invention, video data for deriving spatially varying portions of a scene is provided, the video data being formatted in a file format and including a set of two or more source tracks, each source track including encoded video data representing a spatial portion of a video representing the scene, the set of two or more source tracks including source track groups, and the formatted video data further including one or more group indicators indicating the source tracks belonging to each source track group and one or more active source track indicators indicating the number of two or more active source tracks within the source track group.
[0021] According to another aspect of the present invention, video data for deriving spatially varying portions of a scene is provided, the video data being formatted in a file format and including a set of two or more source tracks, each source track including encoded video data representing a spatial portion of a video representing the scene, and gathering information including fusion information for fusing a subset of the set of two or more source tracks to generate a section-specific video data stream. The formatted video data includes and further includes a set of configurable parameters and / or a template of SEI messages, and the template indicates one or more values of a parameter set or SEI message that needs to be adapted to generate a parameter set or SEI message specific to the section-specific video stream.
[0022] According to another aspect of the present invention, video data for deriving spatially varying portions of a scene is provided, the video data being formatted in a file format and including a set of one or more source tracks including encoded video data representing a spatial portion of a video representing the scene, the encoded video data is encoded using random access points, and the formatted video data further includes one or more random access point alignment indicators indicating whether the random access points in the encoded video data for all spatial portions are aligned.
[0023] According to another aspect of the present invention, a method for creating video data for deriving spatially varying portions of a scene is provided, the video data being formatted in a file format and including a set of two or more source tracks, each source track including encoded video data representing a spatial portion of a video representing the scene, the set of two or more source tracks includes source track groups, and the formatted video data further includes one or more group indicators indicating the source tracks belonging to each source track group and one or more active source track indicators indicating the number of two or more active source tracks within the source track group, The method includes determining the source track groups and the number of two or more active source tracks within the groups, creating one or more group indicators and one or more active source track indicators, and writing them to the formatted video data.
[0024] According to another aspect of the present invention, there is provided a method for creating video data for deriving spatially varying portions of a scene, the video data being formatted in a file format and a set of two or more source tracks, each source track containing encoded video data representing a spatial portion of a video depicting the scene, and collection information including fusion information for fusing a subset of the set of two or more source tracks to generate a section-specific video data stream, wherein the collection information further includes a set of configurable parameters and / or a template for an SEI message, the template indicating one or more values of a set of parameters or an SEI message that need to be adapted to generate a set of parameters or an SEI message specific to the section-specific video stream, the method comprising creating the template and writing it to the collection information of the formatted video data.
[0025] According to another aspect of the present invention, there is provided a method for creating video data for deriving spatially varying portions of a scene, the video data being formatted in a file format and including a set of one or more source tracks containing encoded video data representing a spatial portion of a video depicting the scene, wherein the encoded video data is encoded using random access points, and the formatted video data further includes one or more random access point alignment indicators indicating whether the random access points in the encoded video data for all spatial portions are aligned, the method comprising creating the one or more random access point alignment indicators and writing them to the formatted video data.
[0026] According to another aspect of the present invention, there is provided an apparatus for creating video data for deriving spatially varying portions of a scene, the video data being formatted in a file format, the apparatus being adapted to execute the method according to any one of claims 38 to 55.
[0027] According to another aspect of the present invention, there is provided a method for deriving spatially varying portions of a scene from video data, the video data being formatted in a file format and comprising a set of two or more source tracks, each source track containing encoded video data representing a spatial portion of a video depicting a scene, the set of two or more source tracks including source track groups, the formatted video data further including one or more group indicators indicating the source tracks belonging to each source track group and one or more active source track indicators indicating the number of two or more active source tracks within the source track group, the method comprising reading from the formatted video data one or more group indicators, one or more active source track indicators, and encoded video data from the indicated number of two or more active source track groups, and deriving from this the spatially varying portions of the scene.
[0028] According to another aspect of the present invention, there is provided a method for deriving spatially varying portions of a scene from video data, the video data being formatted in a file format and comprising a set of two or more source tracks, each source track containing encoded video data representing a spatial portion of a video depicting a scene, and collection information including fusion information for fusing a subset of the set of two or more source tracks to generate a section-specific video data stream. including, the collected information further includes a set of configurable parameters and / or a template of the SEI message, and the template indicates one or more values of a parameter set or an SEI message that needs to be adapted to generate a parameter set or an SEI message specific to the section-specific video stream. The method is reading a template from the collected information of the formatted video data, and adapting one or more values of the parameter set or the SEI message indicated by the template to generate a parameter set or an SEI message specific to the section-specific video stream.
[0029] According to another aspect of the present invention, a method for deriving spatially varying portions of a scene from video data is provided, the video data being formatted in a file format and including a set of one or more source tracks including encoded video data representing a spatial portion of a video depicting the scene. The encoded video data is encoded using random access points, and the formatted video data further includes one or more random access point alignment indicators indicating whether the random access points in the encoded video data for all spatial portions are aligned. The method is reading one or more random access point indicators from the formatted video data and accessing the encoded video data based thereon.
[0030] According to another aspect of the present invention, an apparatus for deriving spatially varying portions of a scene from video data is provided, the video data being formatted in a file format, and the apparatus is adapted to execute the method according to any one of claims 57 to 74.
[0031] According to another aspect of the present invention, there is provided a computer program including instructions which, when executed by a computer, cause the computer to execute the method according to any one of claims 38 to 55 or 57 to 74.
[0032] According to another aspect of the present invention, there is provided a computer-readable medium including instructions which, when executed by a computer, cause the computer to execute the method according to any one of claims 38 to 55 or 57 to 74.
[0033] According to another aspect of the present invention, there is provided a digital storage medium storing video data according to any one of claims 1 to 37.
[0034] It should be understood that the video data of claims 1 to 37, the methods of claims 38 to 55, the apparatus of claim 56, the methods of claims 57 to 74, the apparatus of claim 75, the computer program of claim 76, the computer-readable medium of claim 77, and the digital storage medium of claim 78 have similar and / or identical preferred embodiments as specifically defined in the dependent claims.
[0035] It should be understood that the preferred embodiments of the present invention can also be any combination of the dependent claims or any combination of the above embodiments and their respective independent claims.
[0036] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings.
Brief Description of the Drawings
[0037]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
DETAILED DESCRIPTION OF THE INVENTION
[0038] The following description of embodiments of the present invention with respect to the drawings focuses first on embodiments related to the basic classification into a set of fusible source tracks (tile tracks). After that, after describing embodiments related to a set of configurable parameters and / or templates of SEI messages, embodiments related to an extended classification for random access point indications in configurable parameter sets and / or SEI messages and track coupling will be described. For a particular application, all four types of embodiments can be used together to utilize each of these concepts.
[0039] To motivate and facilitate the understanding of embodiments, an example of 360-degree video playback applications based on cube map projections of the scenes shown in FIGS. 1 to 3, tiled into 6×4 spatial segments at two resolutions (high resolution and low resolution), will be described. Such cube map projections constitute video data arranged to derive spatially varying portions of the scene. For example, as shown at the top of FIG. 2, a user can view a 90-degree by 90-degree field of view using a head-mounted display (HMD). In the case of FIG. 2, the subset of tiles necessary to represent the FoV are the four tiles on the right side, the two tiles on the bottom side, the one tile on the front side, and the one tile on the back side of the cube map projection. Of course, depending on the direction of the user's field of view, other subsets of tiles may be required to represent the user's current FoV. In addition to these tiles that can be downloaded and decoded at high resolution, the client application may need to download other tiles outside the viewport to handle sudden changes in the user's orientation. The client application can download and decode these tiles at low resolution. As described above, after downloading on the client side, it is considered desirable to fuse the downloaded tiles into a single bitstream to be processed by a single decoder, for example, to cope with the constraints of a typical mobile device with limited computing resources and power.
[0040] In this example, it is assumed that, using the currently under - development VVC (Versatile Video Coding), each tile is encoded to be independently decodable. This encoding can be achieved, for example, by dividing a picture using a suitable tile / brick / slice structure so that intra or inter prediction is not performed between different tiles / bricks of the same or different pictures. As can be seen from FIG. 5, which shows an excerpt of the syntax of the currently assumed VVC picture parameter set and slice header taken from the VVC specification (Draft 6, Version 11), VVC extends the concepts of tiles and slices known from HEVC by so - called bricks that specify a rectangular region of CTU (Coding Tree Unit) rows within a particular tile in a picture. Thus, a tile can be divided into a plurality of bricks, each consisting of one or more CTU rows within the tile. Using this extended tile / brick / slice structure, it is possible to easily form a tile arrangement as shown in FIG. 3, in which 4×2 spatial segments of high - resolution video and 4×4 spatial segments of low - resolution video are fused into an integrated bitstream through compression domain processing.
[0041] According to the present invention, this fusion process is supported by a specific "fusion - suitable" file format in which the video data is formatted. In this example, this file format is an extension of MPEG OMAF (ISO / IEC 23090 - 2) further based on the ISO - based media file format (ISO / IEC 14496 - 12) that defines the general structure of time - based multimedia files such as video and audio. In this file format, independently decodable video data corresponding to different spatial segments is included in different tracks, which are also referred to herein as source tracks or tile tracks.
[0042] Note that, in this example, VVC is assumed as the basic video codec, but the present invention is not limited to the application of VVC, and different aspects of the present invention can also be realized using other video codecs such as HEVC (High Efficiency Video Coding). Further, in this example, the file format is assumed to be an extension of MPEG OMAF, but the present invention is not limited to such an extension, and different aspects of the present invention can also be realized using other file formats or extensions of other file formats.
[0043] 2. Basic Classification into a Mergeable Set of Source Tracks According to a first aspect of the present invention, a basic classification mechanism enables indicating to a file format parser that some source tracks belong to the same group and a given number of tiles belonging to that group should be played.
[0044] In this regard, the formatted video data includes a set of two or more source tracks, each source track including encoded video data representing a spatial portion of the video indicating a scene. The set of two or more source tracks includes a source track group, and the formatted video data further includes one or more group indicators indicating the source tracks belonging to each source track group and one or more active source track indicators indicating the number of two or more active source tracks within the source track group. In this example, the first source track group includes 6×4 high-resolution tiles of a cube map projection, and the second source track group includes 6×4 low-resolution tiles. This can be indicated by one or more group indicators. Further, as described above, when the user's assumed FoV is 90 degrees × 90 degrees, it is necessary to reproduce 8 out of 24 high-resolution tiles to represent the user's current field of view, while it is also necessary to transmit 16 out of the low-resolution tiles to enable sudden changes in the user's orientation. The 8 source tracks of the first group and the 16 source tracks of the second group can be referred to as "active" source tracks, and the respective numbers thereof can be indicated by one or more active source track indicators.
[0045] In one embodiment, this can be achieved by using a first box of a file format that includes one or more group indicators, for example, a track group type box. The possible syntax and semantics based on the concept of a track group box from the ISO base media file format can be as follows. TIFF0007695234000001.tif87168
[0046] track_group_type indicates a classification type and is set to one of the following values or registered values, or values derived from specifications or registrations. [...] · 'aaaa' indicates that this track belongs to a group of tracks with the same value of track_group_ID and should play a subset of num_active_tracks of them. num_active_tracks must be greater than 1.
[0047] In this case, one or more group indicators are realized by the syntax element track_group_ID, and one or more active source track indicators are realized by the syntax element num_active_tracks. Also, a new track_group_type is defined that indicates that the track group type box includes the syntax element num_active_tracks ('aaaa' is just an example). This type of track group type box can be signaled within each respective source track belonging to the group.
[0048] Since both the source tracks belonging to the first group and the source tracks belonging to the low-resolution group are required to realize 360-degree video playback applications, this application further predicts the possibility of indicating to the file format parser that two or more source track groups are bundled together. In this regard, the formatted video data further includes one or more group bundle indicators indicating such bundling.
[0049] In another embodiment, in combination with the above signaling for each source track, this can be achieved by using another second box, such as a track reference type box, to bundle together multiple groups used in one combination (for example, using one track_group_ID value for high-resolution tiles and one track_group_ID value for low-resolution tiles).
[0050] In a TrackGroupTypeBox of type 'aaaa' that indicates the uniqueness of the track_group_ID, set the value of (flags & 1) to be equal to 1 so that the group can be referenced via 'tref'.
[0051] As implied by the general semantics of track references to track_group_ID, the num_active_tracks tracks of the 'aaaa' source track group are used for the resolution of 'tref'.
[0052] Alternatively, in another embodiment, the source track group does not indicate the number of tracks to be played, and instead this characteristic is represented through an extension of the track reference type box as follows. TIFF0007695234000002.tif100167
[0053] In this case, one or more group indicators indicating the source tracks belonging to each source track group, one or more active source track indicators indicating the number of active source tracks within the source track group, and one or more group bundle indicators indicating that two or more source track groups are bundled together are included in a single box of the file format, in this case the track reference type box.
[0054] The syntax element num_track_group_IDs indicates the number of source track groups bundled within the track reference type box, and the syntax elements track_group_IDs[i] and num_active_tracks_per_track_group_IDs[i] indicate the track group ID and the number of active tracks for each group. In other words, in this embodiment, each source track group is indicated by its respective group ID (e.g., track_group_ID), and two or more source track groups that are bundled together are indicated by an indicator (e.g., num_track_group_IDs) that indicates the number of two or more source track groups that are bundled together and an array of their respective group IDs (e.g., track_group_IDs[i]).
[0055] In the latter two embodiments, the formatted video data can further include a collection track that includes fusion information for fusing a subset of pairs of two or more source tracks to generate a section-specific video data stream, and the track reference box is included in the collection track.
[0056] Alternatively, in yet another embodiment, source track signaling is used to bundle together (sub)groups of source tracks that collect tiles of the same resolution (e.g., high resolution and low resolution). Again, this embodiment can be based on the concept of a track group box from an ISO-based media file, and its possible syntax and semantics are as follows. TIFF0007695234000003.tif133168
[0057] track_group_type indicates a classification type and is set to one of the following values or registered values, or values derived from specifications or registrations. [...] · 'bbbb' means that this track belongs to a track group with the same value of track_group_ID and a subgroup with the same value of track_subgroup_ID, and a subset of num_active_tracks_per_track_subgroup_IDs[i] tracks among them should be played, indicating that track_subgroup_IDs[i] is equal to track_subgroup_ID.
[0058] Thus, in this case, each source track group is indicated to be a subgroup of source tracks by its respective subgroup ID (e.g., track_subgroup_ID), and two or more source track subgroups bundled together are indicated by a common group ID (e.g., track_group_ID), an indicator (e.g., num_track_subgroup_IDs) indicating the number of two or more source track subgroups bundled together, and an array of respective subgroup IDs (e.g., track_subgroup_IDs[i]).
[0059] Alternatively, in yet another embodiment of the present invention, additional group-specific level signaling enables a client to select a group / subgroup combination that matches the level capabilities of the supported decoder. For example, an extension of the last embodiment using a track group type box can be as follows. TIFF0007695234000004.tif143168
[0060] track_group_type indicates a classification type and is set to one of the following values or registered values, or values derived from specifications or registrations. [...] · 'cccc' means that this track belongs to a track group with the same value of track_group_ID and a subgroup with the same value of track_subgroup_ID, and a subset of num_active_tracks_per_track_subgroup_IDs[i] tracks among them should be played. track_subgroup_IDs[i] is equal to track_subgroup_ID, and the playback of the group with track_group_ID corresponds to the level of level_idc of the bitstream corresponding to that group. As a result, the resulting bitstream is shown to accompany the indicated number of num_active_tracks_per_track_subgroup_IDs[i] tracks for each of the num_track_subgroup_IDs subgroups.
[0061] In other words, in this case, the formatted video data further includes a level indicator (e.g., level_idc) indicating the encoding level of the source track group or a bundle of two or more source track groups when the indicated number of tracks are played together.
[0062] Note that the level indicator can also be provided in other embodiments to be described. Further, the two or more source track groups do not necessarily have to differ only in resolution. Rather, in addition to or instead of this, the encoding fidelity can also be different. For example, the first source track group can include source tracks containing encoded video data of the first resolution and / or fidelity, and the second source track group can include source tracks containing encoded video data of a second resolution and / or encoding fidelity different from the first resolution and / or encoding fidelity.
[0063] 3. Template for configurable parameter set and / or SEI message As described above, certain applications require a parameter set or a variant of the SEI message (integrated decoding of tiles in a fused bitstream where tile positions and tile neighbors change) depending on the playout context. Thus, in many cases, it is not easy and not possible to have a single parameter set that applies to multiple combinations.
[0064] One embodiment consists of, for example, signaling a classification mechanism as described above and further indicating that some values of the parameter set template need to be changed. For example, referring to the example of only changing tile selection as described above, the classification mode used indicates that it is necessary to modify the slice_address (HEVC terminology) or slice_id (current VVC terminology used in the picture parameter set syntax table shown in FIG. 5). Another classification mode value indicates that adjustment of the RWP SEI message is required, or that syntax elements related to tiling also need to be adjusted.
[0065] The drawback of such an approach is that for each use case where different syntax elements may need to be changed (sometimes replacing different syntax elements such as slice_id, and in other use cases tiling parameters), it is necessary to signal different group types or similar instructions. A more flexible and general-purpose approach that allows changes to all syntax elements and indicates which syntax elements need to be changed would be beneficial.
[0066] For this purpose, in another embodiment, within the box of the file format, the representation of the parameter set values that are not affected, i.e., the parameter set template, is carried. The client can use this representation to generate the correct parameter set according to its tile / track selection.
[0067] Accordingly, according to this second aspect of the present invention, the formatted video data includes a set of two or more source tracks each including encoded video data representing a spatial portion of the video showing a scene, and collection information including fusion information for fusing a subset of the set of two or more source tracks to generate a section-specific video data stream. The collection information further includes a set of configurable parameters and / or a template of an SEI message, and this template indicates one or more values of a set of parameters or an SEI message that need to be adapted to generate a set of parameters or an SEI message specific to the section-specific video stream. In some embodiments, the formatted video data includes a collection track including the collection information. Different embodiments of this aspect are described below.
[0068] 3.1 XML / JSON Template In one embodiment, the parameter set template and / or the SEI message template is an XML or JSON description of the encoded structure of a parameter set or an SEI message including syntax element names and values, and optionally their encoding. A client (file format parser) can generate a bitstream representation of the parameter set / SEI message by encoding each individual syntax element in its respective form from this XML / JSON description, concatenating the results, and performing anti-emulation. For syntax elements that need to be adjusted by the file format parser, such as the syntax element slice_id or equivalent information for adjusting the position of tiles within a tiled layout, it is desirable that each field be marked in the XML / JSON description as follows. TIFF0007695234000005.tif2689
[0069] In another embodiment, rules for template creation are provided using XML or JSON schemas carried within a file format box. FIG. 6 shows an embodiment of such a schema using XML. The advantage of using an XML / JSON schema is that as long as the syntax element encoding options (e.g., fixed-length vs. variable-length encoding, exponential Golomb coding, etc.) are known, the receiving file format parser can generate a parameter set / SEI message bitstream that conforms without having to pre-recognize the basic codec. A further advantage is that once a single schema is defined, it can be used to easily verify all generated parameter set templates and / or SEI message templates. XML / JSON descriptive metadata with corresponding parameter set templates is preferably stored in the track box (‘trak’) of the collected tracks located within the initialization segment.
[0070] 3.2 Bitstream Templates Without Emulation Prevention In another embodiment, the parameter set template and / or SEI message template is based on the encoded bitstream form of the parameter set / SEI message, i.e., the individual syntax element values are encoded according to a specification (e.g., fixed-length vs. variable-length code, exponential Golomb code, etc.) and concatenated in that specified order. However, this form does not include emulation prevention bytes. Therefore, it is necessary to perform emulation prevention before using such a parameter set within a video bitstream.
[0071] In one embodiment, the parameter set template and / or SEI message template conveys an indication of the gaps where the syntax element values, i.e., their encoded representations such as slice_id, should be inserted.
[0072] Thus, in a general sense, a template can include a parameter set or concatenated coded syntax elements of an SEI message in which values that do not need to be adapted are effectively coded within the template, and further include one or more gap indicators indicating gaps within the template to be filled with effectively coded values that need to be adapted. The one or more gap indicators indicating gaps preferably include the offset and size of the gaps within the template.
[0073] Figure 7 shows the concept of a template gap where a parameter set template is stored within the VVCDecoderConfigurationRecord and gaps are signaled using corresponding offset and size values. The gaps can be signaled by, for example, defining the bitstream blob position (offset) relative to the start of the VVCDecoderConfigurationRecord and the size of the gap, while signaling which element of the parameter set or SEI message is the next element of that blob according to the specification. In one embodiment, a slice_id value (see Figure 5) can be inserted into such a template gap. In another embodiment, tiling structure syntax values (see Figure 5) are inserted into the parameter set template gap.
[0074] The generation of a parameter set or SEI message specific to a section-specific video stream preferably includes performing emulation prevention on the concatenated coded syntax elements to generate a coded bitstream of the parameter set or SEI message after filling the gaps within the template.
[0075] 3.3 Templates with placeholder values In another embodiment, the parameter set template and / or SEI message template stored in the VVCDecoderConfigurationRecord are fully decodable, i.e., they are stored in a bitstream form including anti-emulation similar to a normal non-template parameter set or SEI message, but the fields to be adjusted are filled with valid placeholder values for each encoding. Such a template parameter set fully complies with the specification and can be parsed by a standard corresponding VVC parser. The idea of using such a parameter set template and / or SEI message template is that when the parser processes these parameter sets / SEI messages, the necessary values can be easily overwritten using their instances to complete the definition of the generated parameter sets / SEI messages.
[0076] Thus, in a general sense, a template can include the encoded bitstream of a parameter set or SEI message including anti-emulation bytes filled with valid placeholder values encoded in the encoded bitstream for one or more values that need to be adapted. In the variant of this embodiment described in Section 3.2 above, one or more gap indicators correspond to placeholder value indicators indicating the placeholder values that need to be adapted, and one or more placeholder value indicators indicating the placeholder values are understood to include the offset and size of the placeholder values within the template.
[0077] 3.4 Possible Realizations The following shows a possible realization of the decoder configuration record box within a sample entry that includes the above-described embodiment, i.e., a new sample entry type 'vvcG'. Here, within the loop "for(i = 0; i < numNalus; i++)", the NAL unit can include, for example, a parameter set template or an SEI message template, or a bitstream that forms an XML / JSON base64-encoded representation of the parameter set template or the SEI message template. TIFF0007695234000006.tif240160
[0078] In this realization, the template is included in the decoder configuration record (e.g., VvcDecoderConfigurationRecord), but it can also be included at another location within the initialization segment, such as at another location within the sample description box or at another location within the sample entry box. Furthermore, the presence of the template in the NAL unit is preferably indicated by the NAL unit type (e.g., by defining a specific NAL unit type that indicates the NAL unit containing the template).
[0079] In addition to indicating a parameter set template or an SEI message template within the sample entry of type 'vvcG', the presence of the parameter set template or the SEI message template is preferably also indicated by an additional flag templateNalu within the decoder configuration record of the normal 'vvc1' sample entry. This flag can be provided, for example, for each NAL unit within the loop "for(i = 0; i < numNalus; i++)".
[0080] Thus, in a general sense, the template can be included in a sample entry box, preferably a decoder configuration record, and the presence of the template in the NAL unit is indicated by the sample entry type (e.g., 'vvcG') in the sample entry box and / or one or more template indicators (e.g., templateNalu).
[0081] In these embodiments, additional NAL unit types such as Supplemental Enhancement Information (SEI) messages can be carried in any of the above template forms and modified as appropriate depending on the specific combination selected on the client side. One such SEI message is the RWP SEI message defined by AVC and HEVC.
[0082] To facilitate the replacement of parameters / syntax elements within parameter sets or SEI messages, additional information is present, for example, through a classification mechanism that is signaled in part within collection information such as a collection track and source tracks selected to be combined. This aspect will be further described later in Section 4.
[0083] 3.5 Track-by-track conveyance vs. sample-by-sample conveyance The methods for configurable parameter sets and / or SEI messages to be described can exist, for example, within the decoder configuration record of the initialization segment as in the above embodiments, or within a track in a specific sample. When a parameter set template is included in a track as a media sample, for example, a new sample format can be defined as a parameter set template or an SEI message template in XML / JSON format.
[0084] In another embodiment, a NAL unit having a NAL unit type reserved for external use in VVC is used, and the body of the NAL unit (i.e., the NAL unit payload) needs to be filled with some (distinguishable) parameters and placeholder values that are changed according to some values in the sample group information or the like. For this purpose, either of the methods described (templates in XML / JSON or bitstream format where the "fields to be changed" are identified) can be inserted into the NAL unit payload of this special NAL unit structure.
[0085] Figure 8 shows two types of decoder configuration processes allowed by the file format specification. · An out-of-band parameter set included only in the sample entry in the corresponding decoder configuration record box within the initialization segment. · An in-band parameter set included in the sample entry, which can also be transmitted with the media sample itself and enables the decoder configuration to change over time while playing the same file format track.
[0086] In OMAF version 1, only out-of-band signals are permitted for 360-degree video, and each extractor track includes a predefined parameter set generated by the file format packager for a fixed tiling configuration. Therefore, the client needs to change the collection track and re-initialize the decoder with the corresponding parameter set every time it wants to change this tiling configuration.
[0087] As already explained in the previous section, having a predefined parameter set for such a specific tiling configuration is a major drawback because the client can only act on the predefined extractor track for a specific tiling scheme and cannot flexibly combine the necessary tiles (without extractor NAL units).
[0088] Accordingly, the idea of the present invention is to combine the concept of an in-band parameter set and the concept of an out-of-band parameter set to produce a solution that includes both concepts. FIG. 9 shows the new concept of the generated parameter set. The corresponding collection track includes a parameter set template stored out-of-band (within the sample entry), and this template is used to create a "generated parameter set" that can be in-band when all necessary media segments are selected by the client. The file format track classification mechanism is used to provide information on how to update the parameter set template based on a selected subset of the downloaded tiles.
[0089] In one embodiment, the collection track does not include the media segments themselves such that the media segments are implicitly defined as the sum of the media segments of the selected tiles ('vvcG' in FIG. 9). Accordingly, the initialization segment (such as the sample entry) of the collection track includes all of the metadata necessary to create the generated parameter set.
[0090] In another embodiment, the collection track also includes media segments that provide additional metadata for parameter set generation. This allows the behavior of parameter set generation to be changed over a period of time, rather than relying only on the metadata from the sample entry.
[0091] Thus, in a general sense, the template can be included in the initialization segment of the collection track, preferably in the sample description box, more preferably in the sample entry box, and most preferably in the decoder configuration record. The fusion information includes a media segment that contains references to the encoded video data of a subset of pairs of two or more source tracks. One or more of the media segments further include an indicator indicating that i) the template of the configurable parameter set and / or the SEI message, or ii) the parameter set and / or the SEI message generated together with the template are included in the media segment of the generated section-specific video data stream.
[0092] Note that in all embodiments related to the use of the template of the configurable parameter set and / or the SEI message, the encoded video data included in each source track can be encoded using slices, and the generation of the section-specific video data stream does not require adapting the values of the slice headers of the slices.
[0093] The encoded video data included in each source track is preferably encoded using i) tiles, and the values that need to be adapted are related to the tile structure, and / or ii) bricks, and the values that need to be adapted are related to the brick structure, and / or iii) slices, and the values that need to be adapted are related to the slice structure. In particular, the values that need to be adapted can represent the positions of the tiles and / or bricks and / or slices within the video picture and / or the encoded video data.
[0094] The parameter set is preferably a video parameter set (VPS), a sequence parameter set (SPS), or a picture parameter set (PPS), and / or the SEI message is preferably a regionwise-packing (RWP) SEI message.
[0095] 4. Extended classification for configurable parameter sets and / or SEI messages As described in the preamble, the current state-of-the-art method for indicating that source track groups can be decoded together is by the above-described extractor track that explicitly references each track that carries an appropriate parameter set as shown in FIG. 2 to form one particular valid combination. To reduce the overhead of this state-of-the-art solution (one track per viewport), the present invention more flexibly indicates which tracks can be combined and the rules for combination. Thus, as part of the present invention, a set of two or more source tracks can include one or more boxes in a file format that includes additional information for describing a syntax element that identifies the characteristics of the source tracks, and this additional information enables the generation of a parameter set or SEI message specific to a section-specific video stream without the need to analyze the encoded video data.
[0096] In one embodiment, the additional information describes a syntax element that identifies a slice ID or other information used in a slice header that identifies the slice structure of the associated VCL NAL unit to identify the position of the slice within the integrated bitstream and its combined picture.
[0097] In another embodiment, the additional information describes i) syntax elements that identify the width and height of the encoded video data included in each source track, and / or ii) syntax elements that identify projection mapping, conversion information, and / or guard band information related to the generation of a region-wise packing (RWP) SEI message. For example, the width and height of the encoded video data can be identified in units of encoded samples or in units of maximum encoded blocks. For the RWP SEI message, the syntax element that identifies the projection mapping includes the width and height of the rectangular region within the projection mapping, as well as the upper and left positions. Further, the syntax element that identifies the conversion information can include rotation and mirroring.
[0098] Furthermore, in another embodiment, the additional information further includes the encoding length and / or encoding mode (e.g., u(8), u(v), ue(v)) of each syntax element to facilitate the creation of a configurable parameter set or SEI message.
[0099] In one embodiment, the syntax of the above box is as follows. As described above, each initialization segment of each source track includes a 'trgr' box (track classification indication) inside a 'trak' box (track box) that has an extended track group type box. As a result, new syntax as follows can be carried in the extension of the track group type box. TIFF0007695234000007.tif250166 TIFF0007695234000008.tif90170
[0100] 5. Random Access Point Indication in Track Combining In VVC, NAL unit types may coexist within the same access unit. In this case, IDR NAL units coexist with non-IDR NAL units. That is, some regions can be encoded using inter prediction, while other regions within the picture are intra-encoded to reset the prediction chain for this specific region. In such samples, the client may change its tile selection for a part of the picture. Therefore, it is essential to mark these samples with a file format signaling mechanism so that, for example, even non-IDR NAL units indicate a sub-picture random access point (RAP) with instantaneous decoder refresh (IDR) characteristics at extraction time.
[0101] In this aspect of the present invention, different spatial parts of a video representing a scene can also be provided on a single source track. Accordingly, video data for deriving spatially varying parts of the scene is predicted, and this video data is formatted in a file format and includes a set of one or more source tracks containing encoded video data representing the spatial parts of the video representing the scene. The encoded video data is encoded using random access points, and the formatted video data further includes one or more random access point alignment indicators indicating whether the random access points in the encoded video data for all spatial parts are aligned.
[0102] For example, in one embodiment, different regions of a picture are separated into a plurality of source tracks. In the classification mechanism, it is preferable that whether the RAPs are aligned is signaled. This signaling can be done, for example, by verifying that the RAPs exist within corresponding access units of another source track that includes another spatial part of the picture, regardless of where in the source track the RAPs are located, or by having an additional track (similar to the master track) used for signaling the RAPs. In the second case, only the RAPs signaled within a "master" track, such as a collection track as described above, indicate the RAPs within another source track. If the classification mechanism indicates that the RAPs are not aligned, it is necessary to analyze all the RAP signaling within the other source track. In other words, in this embodiment, the encoded video data representing different spatial parts is included in different source tracks, and the formatted video data further includes a common track that includes one or more random access point indicators indicating the random access points of all the source tracks.
[0103] In another embodiment, all spatial parts are included in the same source track. Nevertheless, in some use cases (e.g., zoom), it may be desirable to extract a part of the entire image (e.g., the central region of interest (RoI)). In such a scenario, the RAPs for the entire picture and the RAPs within the RoI may not always coincide. For example, there may be more RAPs within the RoI than those existing in the entire picture.
[0104] In these embodiments, the formatted video data can further include one or more partial random access point indicators indicating that the access units of the video have random access points for the spatial parts of the video but not for the entire access units. Further, the formatted video data can further include partial random access point information representing the position and / or shape of the spatial parts having random access points.
[0105] In one implementation, this information can be provided using so-called sample groups used in the ISO base media file format to indicate certain characteristics of a picture (e.g., synchronization samples, RAP, etc.). In the present invention, sample groups can be used to indicate that an access unit has partial RAPs, i.e., sub-picture (region-specific) random access points. Further, signaling is added to indicate that each picture can indicate a region without any drift, and the dimensions of the region can be signaled. The syntax of the existing sample to group box is shown below. TIFF0007695234000009.tif96166
[0106] In this embodiment, sample groups are defined for the SampleToGroupBox using a specific classification type 'prap' (partial rap).
[0107] Also, the description of the sample group can be defined, for example, as follows. TIFF0007695234000010.tif32169
[0108] The sample description indicates, for example, a random-accessible region dimension as follows. TIFF0007695234000011.tif49135
[0109] In a further embodiment, different regions are mapped to separate NAL units, i.e., only some of the NAL units of an access unit can be decoded. It is part of the present invention to show that a particular NAL unit can be treated as a RAP when only a subset corresponding to this region is decoded for the bitstream. For this purpose, for example, the concept of an existing subsample information box as follows can be used to derive subsample classification information for sub-pic RAPs. TIFF0007695234000012.tif179137
[0110] The codec_specific_parameters can indicate which subsamples are RAPs and which are not.
[0111] 6. Further embodiments So far, the description of the embodiments of the present invention made below with respect to the drawings has focused on video data for deriving spatially varying parts of a scene and a specific file format in which this part is formatted. However, the present invention also relates to a method and an apparatus for creating video data for deriving spatially varying parts of a scene, and a method and an apparatus for deriving spatially varying parts of a scene from video data formatted in a specific file format. Furthermore, the present invention also relates to a corresponding computer program, computer-readable medium and digital storage medium.
[0112] More specifically, the present invention also relates to the following embodiments.
[0113] A method for creating video data for deriving a spatially varying part of a scene, wherein the video data is formatted in a file format and includes a set of two or more source tracks, each source track including encoded video data representing a spatial part of a video showing the scene The set of two or more source tracks includes a source track group, and the formatted video data further includes one or more group indicators indicating the source tracks belonging to each source track group, and one or more active source track indicators indicating the number of two or more active source tracks within the source track group. The method includes: determining the source track group and the number of two or more active source tracks within the group, creating the one or more group indicators and the one or more active source track indicators, and writing them to the formatted video data.
[0114] In an embodiment of this method, the formatted video data further includes one or more group bundle indicators indicating that two or more source track groups are bundled together, and the method includes: determining the two or more source track groups that are bundled together, creating the one or more bundle indicators, and writing them to the formatted video data.
[0115] In an embodiment of this method, the one or more group indicators indicating the source tracks belonging to each source track group, and the one or more active source track indicators indicating the number of active source tracks within the source track group are included in a first box of a file format different from a second box of the file format that includes the one or more group bundle indicators indicating that the two or more source track groups are bundled together.
[0116] In an embodiment of this method, the first box is a track group type box, and the second box is a track reference type box.
[0117] In an embodiment of this method, the one or more group indicators indicating the source tracks belonging to each of the source track groups, the one or more active source track indicators indicating the number of active source tracks within the source track group, and the one or more group bundle indicators indicating that two or more source track groups are bundled together are included in a single box of the file format.
[0118] In an embodiment of this method, the single box is a track group type box or a track reference type box.
[0119] In an embodiment of this method, the track group type box is included in the source track, and / or the formatted video data further includes a collection track containing fusion information for fusing a subset of the sets of the two or more source tracks to generate a section-specific video data stream. The track reference box is included in the collection track, and the method includes determining the subset of the sets of the two or more source tracks, creating the collection track containing the fusion information, and writing this to the formatted video data.
[0120] In an embodiment of this method, each source track group is indicated by a respective group ID. The two or more source track groups bundled together are indicated by an indicator indicating the number of the two or more source track groups bundled together and an array of the respective group IDs. Alternatively, each source track group is indicated as a subgroup of source tracks by a respective subgroup ID. The two or more subgroups of source tracks bundled together are indicated by a common group ID, an indicator indicating the number of the two or more subgroups of source tracks bundled together, and an array of the respective subgroup IDs.
[0121] In an embodiment of this method, the formatted video data further includes a level indicator indicating the encoding level of the source track group or the encoding level of a bundle of two or more source track groups, and the method includes determining the source track group or the bundle of two or more source track groups, creating the level indicator, and writing it to the formatted video data.
[0122] In an embodiment of this method, the first source track group includes a source track containing encoded video data of a first resolution and / or fidelity, and the second source track group includes a source track containing encoded video data of a second resolution and / or fidelity different from the first resolution and / or encoding fidelity.
[0123] A method of creating video data for deriving spatially varying portions of a scene, the video data being formatted in a file format and including a set of two or more source tracks, each source track containing encoded video data representing a spatial portion of a video depicting the scene, and collection information including fusion information for fusing a subset of the set of two or more source tracks to generate a section-specific video data stream, and wherein the collection information further includes a set of configurable parameters and / or a template of an SEI message, the template indicating one or more values of the set of parameters or the SEI message that need to be adapted to generate a set of parameters or an SEI message specific to the section-specific video stream, the method includes creating the template and writing it to the collection information of the formatted video data.
[0124] In an embodiment of this method, the formatted video data includes a collection track that includes the collection information.
[0125] In an embodiment of this method, the template includes an XML or JSON description of the parameter set or the encoding structure of the SEI message.
[0126] In an embodiment of this method, the formatted video data further includes an XML or JSON schema that provides rules for the creation of the template, and the method includes creating the XLM or JSON schema and writing it to the formatted video data.
[0127] In an embodiment of this method, the template includes concatenated encoding syntax elements of the parameter set or the SEI message. Within the template, values that do not need to be conformed are effectively encoded, and the template further includes one or more gap indicators that indicate gaps within the template that should be filled with effectively encoded values that need to be conformed.
[0128] In an embodiment of this method, the one or more gap indicators indicating the gaps include the offsets and sizes of the gaps within the template.
[0129] In an embodiment of this method, the generation of the parameter set or the SEI message specific to the section-specific video stream includes performing anti-emulation on the concatenated encoding syntax elements to generate an encoded bitstream of the parameter set or the SEI message after filling the gaps within the template.
[0130] In an embodiment of this method, the template includes the parameter set including the anti-emulation byte or the encoded bitstream of the SEI message, and the one or more values that need to be adapted in the encoded bitstream are filled with validly encoded placeholder values.
[0131] In an embodiment of this method, the template is included in the initialization segment of the collection track, preferably in the sample description box, more preferably in the sample entry box, and most preferably in the decoder configuration record.
[0132] In an embodiment of this method, the template is included in the NAL unit, and the presence of the template in the NAL unit is indicated by the NAL unit type.
[0133] In an embodiment of this method, the template is included in the sample entry box, preferably in the decoder configuration record, and the presence of the template in the NAL unit is indicated by the sample entry type and / or one or more template indicators in the sample entry box.
[0134] In an embodiment of this method, it is included in the initialization segment of the collection track, preferably in the sample description box, more preferably in the sample entry box, and most preferably in the decoder configuration record. The fusion information includes a media segment including a reference to the encoded video data of the subset of the set of two or more source tracks, and one or two or more of the media segments include i) a template of a configurable parameter set and / or an SEI message, or ii) an indicator indicating that a parameter set and / or an SEI message generated using the template is included in the media segment of the generated section-specific video data stream.
[0135] In an embodiment of this method, the encoded video data included by each source track is encoded using slices, and the generation of the video data stream specific to the section does not require adapting the values of the slice headers of the slices.
[0136] In an embodiment of this method, the encoded video data included by each source track is encoded using i) tiles, and the values that need to be adapted are related to the tile structure, and / or ii) bricks, and the values that need to be adapted are related to the brick structure, and / or iii) slices, and the values that need to be adapted are related to the slice structure.
[0137] In an embodiment of this method, the values that need to be adapted represent the position of the picture of the video and / or the tiles and / or bricks and / or slices in the encoded video data.
[0138] In an embodiment of this method, the parameter set is a video parameter set (VPS), a sequence parameter set (SPS), or a picture parameter set (PPS), and / or the SEI message is a region-wise packing (RWP) SEI message.
[0139] In an embodiment of this method, the set of two or more source tracks includes one or more boxes of the file format that include additional information for describing a syntax element that identifies the characteristics of the source track, and the additional information enables the generation of the parameter set or the SEI message specific to the video stream specific to the section without the need to analyze the encoded video data.
[0140] In an embodiment of this method, the additional information describes i) a syntax element that identifies the width and height of the encoded video data included by each source track, and / or ii) a syntax element that identifies the projection mapping, conversion information, and / or protection frequency band information related to the generation of the region-wise packing (RWP) SEI message.
[0141] In an embodiment of this method, the encoded video data included by each source track is encoded using slices, and the additional information describes a syntax element that identifies a slice ID, or other information used within the slice header to identify the slice structure.
[0142] In an embodiment of this method, the additional information further includes the encoding length and / or encoding mode of each of the syntax elements.
[0143] In an embodiment of this method, the one or more boxes are an extension of a box of track group type.
[0144] A method of creating video data for deriving spatially varying portions of a scene, the video data being formatted in a file format and including a set of one or more source tracks including encoded video data representing spatial portions of a video depicting the scene, the encoded video data being encoded using random access points, and the formatted video data further including one or more random access point alignment indicators indicating whether the random access points in the encoded video data are aligned for all spatial portions. The method includes creating the one or more random access point alignment indicators and writing them to the formatted video data.
[0145] In an embodiment of this method, the formatted video data further includes one or more partial random access point indicators indicating that the access unit of the video has a random access point for the spatial part of the video but not for the entire access unit, and the method includes creating the one or more partial random access point indicators and writing them into the formatted video data.
[0146] In an embodiment of this method, the formatted video data further includes partial random access point information representing the position and / or shape of the spatial part having the random access point, and the method includes creating the partial random access point information and writing it into the formatted video data.
[0147] In an embodiment of this method, different spatial parts of an access unit are included in different NAL units, and the partial random access point information describes which NAL unit is a random access point for each spatial part, and the partial random access point information is included in a box of the file format, preferably in a subsample information box.
[0148] In an embodiment of this method, the encoded video data representing the different spatial parts is included in different source tracks, and the formatted video data further includes a common track including one or more random access point indicators indicating the random access points of all source tracks.
[0149] An apparatus for creating video data for deriving a spatially varying part of a scene, wherein the video data is formatted in a file format, and the apparatus is adapted to execute the method according to any one of claims 38 to 55 or any of the above embodiments.
[0150] A method for deriving spatially varying portions of a scene from video data, wherein the video data is formatted in a file format and includes a set of two or more source tracks, each source track including encoded video data representing a spatial portion of a video showing the scene, the set of two or more source tracks includes source track groups, and the formatted video data further includes one or more group indicators indicating source tracks belonging to respective source track groups and one or more active source track indicators indicating the number of two or more active source tracks within a source track group, the method includes reading, from the formatted video data, the one or more group indicators, the one or more active source track indicators, and the encoded video data from the number of two or more indicated active source track groups, and deriving the spatially varying portions of the scene based thereon.
[0151] In an embodiment of this method, the formatted video data further includes one or more group bundle indicators indicating that two or more source track groups are bundled together, and the method includes reading, from the formatted video data, the one or more bundle indicators and the two or more source track groups bundled together, and deriving the spatially varying portions of the scene based thereon.
[0152] In an embodiment of this method, the one or more group indicators indicating the source tracks belonging to each of the source track groups, and the one or more active source track indicators indicating the number of active source tracks within the source track group, are included in a first box of a file format different from a second box of a file format that includes the one or more group bundle indicators indicating that the two or more source track groups are bundled together.
[0153] In an embodiment of this method, the first box is a track group type box and the second box is a track reference type box.
[0154] In an embodiment of this method, the one or more group indicators indicating the source tracks belonging to each of the source track groups, the one or more active source track indicators indicating the number of active source tracks within the source track group, and the one or more group bundle indicators indicating that two or more source track groups are bundled together are included in a single box of the file format.
[0155] In an embodiment of this method, the single box is a track group type box or a track reference type box.
[0156] In an embodiment of this method, the track group type box is included in the source track and / or the formatted video data further includes a collection track including fusion information for fusing a subset of pairs of the two or more source tracks to generate a section-specific video data stream, the track reference box is included in the collection track, and the method Reading the fusion information and the subset of the set of two or more source tracks from the formatted video data, and fusing the subset of the set of two or more source tracks to generate the section-specific video data stream based on the fusion information.
[0157] In an embodiment of this method, each source track group is indicated by a respective group ID, and the two or more source track groups bundled together are indicated by an indicator indicating the number of the two or more source track groups bundled together and an array of the respective group IDs, or each source track group is indicated as a subgroup of source tracks by a respective subgroup ID, and the two or more subgroups of source tracks bundled together are indicated by a common group ID, an indicator indicating the number of the two or more subgroups of source tracks bundled together, and an array of the respective subgroup IDs.
[0158] In an embodiment of this method, the formatted video data further includes a level indicator indicating the encoding level of the source track group or the encoding level of a bundle of two or more source track groups, and the method includes Reading the level indicator and the source track group or the bundle of two or more source track groups from the formatted video data, and deriving the spatially varying part of the scene based thereon.
[0159] In an embodiment of this method, the first source track group includes source tracks including encoded video data of a first resolution and / or fidelity, and the second source track group includes source tracks including encoded video data of a second resolution and / or fidelity different from the first resolution and / or encoding fidelity.
[0160] A method for deriving spatially varying portions of a scene from video data, wherein the video data is formatted in a file format and a set of two or more source tracks, each source track containing encoded video data representing a spatial portion of the video depicting the scene, collection information including fusion information for fusing a subset of the set of two or more source tracks to generate a section-specific video data stream, wherein the collection information further includes a set of configurable parameters and / or a template for an SEI message, the template indicating one or more values of the set of parameters or the SEI message that need to be adapted to generate a set of parameters or an SEI message specific to the section-specific video stream, the method comprising: reading the template from the collection information of the formatted video data and adapting the one or more values of the set of parameters or the SEI message indicated by the template to generate the set of parameters or the SEI message specific to the section-specific video stream.
[0161] In an embodiment of this method, the template includes an XML or JSON description of the encoding structure of the set of parameters or the SEI message.
[0162] In an embodiment of this method, the formatted video data further includes an XML or JSON schema providing rules for the creation of the template, and the method comprises: reading the XLM or JSON schema and using it in the generation of the set of parameters or the SEI message.
[0163] In an embodiment of this method, the template includes concatenated encoded syntax elements of the parameter set or the SEI message. Within the template, values that do not need to be adapted are effectively encoded, and the template further includes one or more gap indicators indicating gaps within the template that should be filled with effectively encoded values that need to be adapted.
[0164] In an embodiment of this method, the one or more gap indicators indicating the gaps include the offsets and sizes of the gaps within the template.
[0165] In an embodiment of this method, the generation of the parameter set or the SEI message specific to the section-specific video stream includes performing emulation prevention on the concatenated encoded syntax elements to generate an encoded bitstream of the parameter set or the SEI message after filling the gaps within the template.
[0166] In an embodiment of this method, the template includes an encoded bitstream of the parameter set or the SEI message including emulation prevention bytes, and the one or more values that need to be adapted within the encoded bitstream are filled with effectively encoded placeholder values.
[0167] In an embodiment of this method, the template is included in the initialization segment of the collection track, preferably in the sample description box, more preferably in the sample entry box, and most preferably in the decoder configuration record.
[0168] In an embodiment of this method, the template is included in the NAL unit, and the presence of the template in the NAL unit is indicated by the NAL unit type.
[0169] In an embodiment of this method, the template is preferably included in a decoder configuration record in a sample entry box, and the presence of the template in the NAL unit is indicated by the sample entry type and / or by one or more template indicators within the sample entry box.
[0170] In an embodiment of this method, the template is preferably included in a sample description box, more preferably in a sample entry box, and most preferably in a decoder configuration record in an initialization segment of the collection track. The fusion information includes a media segment that includes a reference to the encoded video data of the subset of the pair of two or more source tracks, and one or two or more of the media segments include i) a template of a configurable parameter set and / or an SEI message, or ii) an indicator indicating that a parameter set and / or an SEI message generated using the template is included in the media segment of the generated section-specific video data stream.
[0171] In an embodiment of this method, the encoded video data included by each source track is encoded using slices, and the generation of the section-specific video data stream does not require adapting the values of the slice headers of the slices.
[0172] In an embodiment of this method, the encoded video data included by each source track is encoded using i) tiles, and the values that need to be adapted are related to the tile structure, and / or ii) bricks, and the values that need to be adapted are related to the brick structure, and / or iii) slices, and the values that need to be adapted are related to the slice structure.
[0173] In an embodiment of this method, the value to be adapted represents the position of the tile and / or block and / or slice within the picture of the video and / or the encoded video data.
[0174] In an embodiment of this method, the parameter set is a video parameter set (VPS), a sequence parameter set (SPS), or a picture parameter set (PPS), and / or the SEI message is a region-wise packing (RWP) SEI message.
[0175] In an embodiment of this method, the set of two or more source tracks includes one or more boxes of the file format that contain additional information for describing a syntax element that identifies the characteristics of the source track, and the additional information enables the generation of the parameter set or the SEI message specific to the video stream of the section without the need to parse the encoded video data.
[0176] In an embodiment of this method, the additional information describes i) a syntax element that identifies the width and height of the encoded video data included by each source track, and / or ii) a syntax element that identifies a projection mapping, conversion information, and / or protected frequency band information related to the generation of a region-wise packing (RWP) SEI message.
[0177] In an embodiment of this method, the encoded video data included by each source track is encoded using slices, and the additional information describes a syntax element that identifies the slice ID or other information for identifying the slice structure used within the slice header.
[0178] In an embodiment of this method, the additional information further includes the encoding length and / or encoding mode of each respective syntax element.
[0179] In an embodiment of this method, the one or more boxes are an extension of a track group type box.
[0180] A method for deriving spatially varying portions of a scene from video data, wherein the video data is formatted in a file format and includes a set of one or more source tracks including encoded video data representing a spatial portion of a video representing the scene, the encoded video data is encoded using random access points, and the formatted video data further includes one or more random access point alignment indicators indicating whether the random access points in the encoded video data are aligned for all spatial portions, the method includes reading the one or more random access point indicators from the formatted video data and accessing the encoded video data based thereon.
[0181] In an embodiment of this method, the formatted video data further includes one or more partial random access point indicators indicating that access units of the video have random access points for spatial portions of the video but not for the entire access unit, and the method includes reading the one or more partial random access point indicators from the formatted video data and accessing the encoded video data based thereon.
[0182] In an embodiment of this method, the formatted video data further includes partial random access point information representing the position and / or shape of the spatial portion having the random access point, and the method includes reading the partial random access point information and accessing the encoded video data based thereon.
[0183] In an embodiment of this method, different spatial portions of the access unit are included in different NAL units, and the partial random access point information describes which NAL unit is a random access point for each spatial portion, and the partial random access point information is included in a box of the file format, preferably in a subsample information box.
[0184] In an embodiment of this method, the encoded video data representing the different spatial portions is included in different source tracks, and the formatted video data further includes a common track including one or more random access point indicators indicating the random access points of all source tracks.
[0185] An apparatus for deriving spatially varying portions of a scene from video data, the video data being formatted in a file format, the apparatus being adapted to execute the method according to any of claims 57 to 74 or any of the above embodiments.
[0186] A computer program comprising instructions which, when executed by a computer, cause the computer to execute the method according to any of claims 38 to 55 or 57 to 74 or any of the above embodiments.
[0187] A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to execute the method according to any of claims 38 to 55 or 57 to 74 or any of the above embodiments.
[0188] A digital storage medium storing video data according to any of claims 1 to 37.
[0189] These methods, apparatuses, computer programs, computer-readable media and digital storage media can have corresponding features as described with respect to the formatted video data.
[0190] Generally, a method of creating video data for deriving spatially varying portions of a scene can include steps of creating information such as different types of indicators, such as one or more group indicators, one or more active source track indicators, one or more group bundle indicators, level indicators, one or more partial random access point indicators, templates such as configurable parameter sets and / or templates of SEI messages, and, for example, i) syntax elements that identify the width and height of the encoded video data included in each source track, and / or ii) projection mappings, conversion information, and / or syntax elements that identify protection frequency band information related to the generation of region-wise packing (RWP) SEI messages, partial random access point information, etc., and writing these to the formatted video data. In this context, it may be necessary to determine specific information signaled in the file format, source track groups, and the number of two or more active source tracks within a group. In some cases, this determination can be made through an interface that allows a user to input the necessary information, or can be derived partially or fully from the encoded video data (e.g., RAP information).
[0191] Similarly, a method of deriving spatially varying portions of a scene from video data can include steps of reading different types of indicators, templates, and information, and using the read data to perform different tasks. This method can include deriving spatially varying portions of the scene based on this, and / or generating a parameter set or SEI message specific to a section-specific video stream, and / or accessing the encoded video data based on the read RAP information.
[0192] Embodiments of the present invention can be implemented in hardware or software according to specific implementation requirements. This implementation can be carried out using a digital storage medium that stores electronically readable control signals, such as a floppy disk, DVD, BluRay, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, which cooperate with (or can cooperate with) a programmable computer system so that each method is executed. Therefore, the digital storage medium can be made computer-readable.
[0193] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to execute some or all of the functions of the methods described herein.
[0194] In some embodiments, a field programmable gate array can cooperate with a microprocessor to execute one of the methods described herein. Generally, these methods are preferably executed by any hardware device.
[0195] The apparatuses described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0196] The apparatuses described herein, or any components of the apparatuses described herein, can be implemented at least partially in hardware and / or software.
[0197] The methods described herein can be executed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0198] Any method described herein, or any component of the apparatus described herein, can be executed at least in part by hardware and / or software.
[0199] The above-described embodiments are merely illustrative of the principles of the present invention. It will be understood by those skilled in the art that various modifications and variations of the configurations and details described herein will become apparent. Therefore, the intention is that the invention is limited only by the appended claims and not by the specific details shown in the description and explanation of the embodiments herein.
Claims
1. A method for creating video data for deriving a spatial subset of a scene, wherein the video data is formatted in a file format and includes a set of two or more source tracks, each source track including encoded video data representing a spatial subset of the video, the set of two or more source tracks includes source track groups, and the formatted video data further includes one or more group indicators indicating the source tracks belonging to each source track group and one or more active source track indicators indicating the number of two or more active source tracks within the source track group, the formatted video data further includes one or more group bundle indicators indicating that two or more source track groups are bundled together, the formatted video data further includes a level indicator indicating the encoding level of the source track group or the encoding level of a bundle of two or more source track groups, the method includes determining the source track group and the number of two or more active source tracks within the group, creating the one or more group indicators and the one or more active source track indicators, and writing them to the formatted video data; determining the two or more source track groups bundled together, creating the one or more bundle indicators, and writing them to the formatted video data; determining the source track group or the bundle of two or more source track groups, creating the level indicator, and writing it to the formatted video data; A method, comprising.
2. The one or more group indicators indicating the source tracks belonging to each of the source track groups, the one or more active source track indicators indicating the number of active source tracks within the source track group, and the one or more group bundle indicators indicating that two or more source track groups are bundled together are included in a single box of the file format, The method according to claim 1.
3. Each source track group is indicated by a respective group ID, and the two or more source track groups that are bundled together are indicated by an indicator indicating the number of the two or more source track groups that are bundled together and by an array of the respective group IDs, or each source track group is indicated as being a subgroup of source tracks by a respective subgroup ID, and the two or more subgroups of source tracks that are bundled together are indicated by a common group ID, an indicator indicating the number of the two or more subgroups of source tracks that are bundled together, and an array of the respective subgroup IDs, The method according to claim 1 or 2.
4. An apparatus for creating video data for deriving a spatial subset, the video data being formatted in a file format, the apparatus being adapted to perform the method according to any one of claims 1 to 3, Apparatus.
5. A method for deriving a spatial subset of a scene from video data, the video data being formatted in a file format and including a set of two or more source tracks in which each source track includes encoded video data representing a spatial subset of the video, The set of two or more source tracks includes a source track group, and the formatted video data further includes one or more group indicators indicating the source tracks belonging to each source track group, and one or more active source track indicators indicating the number of two or more active source tracks within the source track group. The formatted video data further includes one or more group bundle indicators indicating that two or more source track groups are bundled together. The formatted video data further includes a level indicator indicating the encoding level of the source track group or the encoding level of the bundle of two or more source track groups. The method is Reading from the formatted video data the one or more group indicators, the one or more active source track indicators, and the encoded video data from the indicated number of two or more active source track groups, and deriving the spatial subset based thereon. Reading from the formatted video data the one or more bundle indicators and the two or more source track groups bundled together, and deriving the spatial subset based thereon. Reading from the formatted video data the level indicator and the source track group or the bundle of two or more source track groups, and deriving the spatial subset based thereon. A method comprising
6. The one or more group indicators indicating the source tracks belonging to each of the source track groups, the one or more active source track indicators indicating the number of active source tracks within the source track group, and the one or more group bundle indicators indicating that two or more source track groups are bundled together are included in a single box of the file format, The method according to claim 5.
7. Each of the source track groups is indicated by a respective group ID, and the two or more source track groups bundled together are indicated by an indicator indicating the number of the two or more source track groups bundled together and an array of the respective group IDs, or each source track group is indicated as a subgroup of source tracks by a respective subgroup ID, and the two or more subgroups of source tracks bundled together are indicated by a common group ID, an indicator indicating the number of the two or more subgroups of source tracks bundled together, and an array of the respective subgroup IDs, The method according to claim 5 or 6.
8. An apparatus for deriving a spatial subset from video data, the video data being formatted in a file format, the apparatus being adapted to perform the method according to any one of claims 5 to 7, Apparatus.
9. Including instructions that, when executed by a computer, cause the computer to perform the method according to claims 1 to 3 or 5 to 7, A computer program.
Citation Information
Patent Citations
Image processing device and method
WO2015012227A1
Reception device, reception method, transmission device, and transmission method
WO2016006431A1
An apparatus, a method and a computer program for omnidirectional video
WO2019002662A1
Method, device, and computer program for generating timed media data
WO2019072795A1