Apparatus for generating a video data stream and method for generating a video data stream

By providing additional signaling and extraction information in the video data stream to guide the modification of slice addresses and parameter sets, the low efficiency of the video data stream extraction process in the prior art is solved, and more efficient video content processing and adaptability are achieved.

CN116074502BActive Publication Date: 2025-11-28DOLBY VIDEO COMPRESSION LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211353005.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-03-20
Filing Date
2018-03-19
Publication Date
2025-11-28
Estimated Expiration
2038-03-19

AI Technical Summary

Technical Problem

Existing video data stream extraction processes are inefficient and complex when handling video content with different scene resolutions, especially when the receiver is unaware of the video type, making it difficult to effectively handle different types of video content.

Method used

By providing additional signaling and extraction information in the video data stream, the extractor device is guided on how to modify the slice address and parameter set, reducing the processing burden and adapting to different types of video content.

Benefits of technology

It improves the efficiency and adaptability of the video data stream extraction process, enabling more effective processing of video content with different scene resolutions, and reducing the complexity and implementation cost of the extraction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116074502B_ABST
    Figure CN116074502B_ABST
Patent Text Reader

Abstract

A more efficient video data stream extraction technique is presented, e.g., which can more efficiently handle video content of unknown type to the recipient, which has different types of video, e.g., with different aspects such as projection of the view volume to the picture plane, or which reduces the complexity of the extraction process. Furthermore, a technique is described with which collocation of different scene resolution versions of a video scene can be more efficiently provided to the recipient.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application for the invention patent application No. 201880033197.5, entitled "Apparatus for generating a video data stream and method for generating a video data stream", with the filing date of 19 March 2018. TECHNICAL FIELD

[0002] The present application relates to video data stream extraction, i.e. to a way of extracting a downscaled video data stream from a video data stream that has been properly prepared, such that the downscaled video data stream encodes a video of a smaller space, the video of a smaller space corresponding to a spatial portion of the video encoded in the original video data stream. The present application also relates to the transmission of different video versions of a scene, the versions being different in terms of scene resolution or fidelity. BACKGROUND

[0003] The HEVC standard [1] defines a hybrid video codec that allows defining rectangular tile sub-arrays of pictures, for which the video codec follows some encoding constraints, allowing to easily extract a smaller or downscaled video data stream from the whole video data stream, i.e. without re-quantization and without re-doing any motion compensation. As outlined in [2], it is envisaged to add to the HEVC standard syntax to allow a guided extraction process by a video data stream recipient.

[0004] However, there is still a need to make this extraction process more efficient.

[0005] An application field in which video data extraction can be applied relates to the transmission or provision of multiple versions of a video scene, the versions being different in terms of scene resolution. An efficient way to implement such a transmission or provision of versions being different in terms of resolution would be advantageous. SUMMARY

[0006] It is therefore a first object of the present application to provide a more efficient technique for video data stream extraction, e.g. a technique that is able to more efficiently handle video content of unknown type by a recipient, the video content being of different types, e.g. in terms of projection of a view area onto a picture plane, or that reduces the complexity of the extraction process. This object is achieved by the subject-matter of the independent claims according to the first aspect of the present application.

[0007] In particular, according to the first aspect of the present application, the extraction of a video data stream is made more efficient by providing information to the extraction information within the video data stream, the information provided signaling one of a plurality of options for how to modify the slice addresses of the slice portions of each slice within the spatial segment that is extractable in order to indicate within the downscaled video data stream the position where the respective slice is located in the downscaled picture area. In other words, the second information provides the video data stream extraction site with information guiding the extraction process with respect to composing the pictures of the spatially downscaled video from the spatial segment of the original video that the downscaled (extracted) video data stream is based on, thus, either lightening the processing burden of the extraction process or adapting it to the greater variability of the scene types conveyed in the video data stream. With respect to the latter issue, for example, the second information can handle cases where the pictures of the spatially downscaled video advantageously should not just be the result of simply pushing the possibly separated portions of the spatial segment together and maintaining the relative arrangement (or relative order with respect to the encoding order) of these portions of the spatial segment in the original video unchanged. For example, in a spatial segment consisting of zones adjacent to different portions along the periphery of the original picture showing the scene in the interface of a panorama scene to a picture plane projection, the arrangement of the zones of this spatial segment in the smaller picture of the extracted stream should be different than in the case of a picture type being non-panoramic type, but the receiver can not even be aware of this type. Additionally or separately, modifying the slice addresses of the extracted slice portions is a tedious task that can be lightened by explicitly sending information on how to do the modification in the form of, for example, alternative slice addresses.

[0008] It is another object of the present application to provide a technique with which the juxtaposition of versions of a video scene at different scene resolutions can be more efficiently provided to a receiver.

[0009] This object is achieved by the subject matter of the independent claim of the second aspect of the present application.

[0010] In particular, according to the second aspect of the present application, the juxtaposition of versions of a video scene at different scene resolutions is more efficiently provided by assembling these versions in one video that is encoded in one video data stream and providing a signal to the video data stream that indicates that the pictures of the video show the same scene content at different resolutions at different spatial portions of the pictures. Thus, a receiver of the video data stream is able to identify based on the signal whether the video content conveyed by the video data stream belongs to a spatially side-by-side collection of multiple versions of scene content at different scene resolutions. Depending on the capabilities of the receiving site, any attempt to decode the video data stream can be suppressed or the processing of the video data stream can be adapted to the analysis of the signal. BRIEF DESCRIPTION OF DRAWINGS

[0011] Advantageous implementations of embodiments of the above aspects are subject matter of dependent claims. Preferred embodiments of the present application are described below in connection with the accompanying drawings, in which:

[0012] Figure 1 A schematic diagram showing MCTS extraction with adjusted slice addresses is shown;

[0013] Figure 2 A mixed schematic and block diagram showing the way of video data stream extraction processing and involved processes and devices according to embodiments of the first aspect of the present application is shown;

[0014] Figure 3 A syntax example according to the example is shown, which inherits the example of the second information of extraction information, wherein the second information explicitly indicates how to modify the slice address;

[0015] Figure 4 A schematic diagram showing an example of non-adjacent MCTS forming a desired picture sub-segment is shown;

[0016] Figure 5 A specific syntax example according to embodiments is shown, which includes the second information, wherein the second information indicates a certain option among several possible options for modifying the slice address in the extraction process;

[0017] Figure 6 A schematic diagram showing an example of multi-resolution 360° frame packing is shown;

[0018] Figure 7 A schematic diagram showing extraction of an example MCTS containing a mixed resolution representation is shown; and

[0019] Figure 8 A mixed schematic and block diagram showing the way of providing multi-resolution scenes to users and involved devices and video streams and processes according to embodiments of the second aspect of the present application is shown. DETAILED DESCRIPTION

[0020] The description of the present application starts from a description of the first aspect of the present application, and then continues with a description of the second aspect of the present application. More precisely, with respect to the first aspect of the present application, the description of the present application starts from a brief overview of the basic technical problem, to lead to the advantages and the basic technology of the embodiments of the first aspect described below. With respect to the second aspect, the same way of selecting the order of description is chosen.

[0021] In panoramic or 360 video applications, it is often desirable to show only a portion of the picture plane to the user. Certain codec tools such as Motion Constrained Tile Sets (MCTS) allow to extract the encoded data corresponding to the desired picture sub-portion in the compression domain and form a conformant bitstream that can be decoded by legacy decoder devices that do not support decoding MCTS from a full picture bitstream and that can be characterized as lower tier compared to the decoders required for full picture decoding.

[0022] As an example and reference, the signaling involved in the HEVC codec can be found in the following references:

[0023] • Reference [1] specifies in sections D.2.29 and E.2.29 a temporal MCTS SEI message that allows the encoder to signal that a given list of rectangles (each rectangle in the list being defined by the tile indices of its top-left corner and bottom-right corner) belongs to a MCTS.

[0024] • Reference [2] provides additional information such as parameter sets and nested SEI messages that can simplify the work of extracting MCTS as conformant HEVC bitstreams and will be added to the next version of [1].

[0025] As can be seen from [1] and [2], the extraction process includes an adjustment of the slice addresses signaled in the slice headers of the relevant slices, the adjustment being performed in the extractor device.

[0026] Figure 1 An example of MCTS extraction is shown. Figure 1 A picture is shown that has been encoded in a video data stream, i.e. a HEVC video data stream. The picture 100 is subdivided into CTBs, i.e. Coding Tree Blocks, in units of which the picture 100 is encoded. In the example of Figure 1 The picture 100 is subdivided into 16x6 CTBs, but of course the number of rows of CTBs and the number of columns of CTBs is not critical. Reference numeral 102 indicates representatively one such CTB. In units of these CTBs 102, the picture 100 is further subdivided into tiles, i.e. an array of m x n tiles, Figure 1 An exemplary case with m = 8 and n = 4 is shown. In each tile, a reference numeral 104 has been used to indicate representatively one such tile. Each tile is thus a rectangular cluster or sub-array of CTBs 102. For illustration purposes only, Figure 1 It is shown that the tiles 104 can have different sizes, or in other words, the rows of tiles can have different heights from each other and the columns of tiles can have different widths from each other.

[0027] As known in the art, the subdivision of tiles, i.e. the subdivision of the picture 100 into tiles 104, influences the encoding order 106 along which the picture content of the picture 100 is encoded in the video data stream. In particular, the tiles 104 are traversed one after the other along the order of the tiles, i.e. in a raster scan order of the tiles row by row. In other words, first all CTBs 102 within one tile 104 are encoded or traversed according to the encoding order 106, then the encoding order advances to the next tile 104. Within each tile 102, the CTBs are encoded using a raster scan order as well, i.e. using a line-by-line raster scan order. Along the encoding order 106, the encoding of the picture 100 in the video data stream is subdivided to yield so-called slice portions. In other words, a slice of the picture 100, which is traversed by a consecutive portion of the encoding order 106, is encoded into the video data stream as a unit to form a slice portion. In Figure 1 It is assumed in the following that each tile is located within a single slice (or, expressed in terms of HEVC, a slice segment), but this is only an example and different ways can be employed. Typically, one slice 108 (or, expressed in terms of HEVC, one slice segment) is represented by the reference sign 108 in Figure 1 and coincides or corresponds with the corresponding tile 104.

[0028] As far as the encoding of the picture 100 in the video data stream is concerned, it should be noted that this encoding makes use of spatial prediction, temporal prediction, context derivation for entropy coding, motion compensation for temporal prediction, and transform of the prediction residual and / or quantization of the prediction residual. The encoding order 106 not only influences the slices, but also defines the availability of reference bases for spatial prediction and / or context derivation: only those neighboring portions are available which are in front of the encoding order 106. The tiles not only influence the encoding order 106, but also limit the encoding interdependencies within the picture 100: for example, spatial prediction and / or context derivation are limited to refer only to portions within the current tile 104. Portions outside the current tile are not referred to in spatial prediction and / or context derivation.

[0029] Figure 1 It is presently shown that another specific region of the picture 100, namely a so-called MCTS, i.e. a spatial segment 110 within the picture region of the picture 100, can be extracted for which a video to which the picture 100 belongs to can be extracted. In Figure 1 An enlarged illustration of the segment 110 is shown to the right of Figure 1 The MCTS 110 of the picture 100 consists of a set of tiles 104. The tiles located in the segment 110 are each provided with a name, namely a, b, c and d. The fact that the spatial segment 110 is extractable involves a further restriction on the encoding of the video into the video data stream. In particular, Figure 1The illustrated video picture partitioning of picture 100 is adopted by other pictures of the video, and for this picture sequence, the picture content within tiles a, b, c and d is encoded in a way that the coding interdependencies remain within spatial segment 110 even when referring from one picture to another. In other words, for example, temporal prediction and temporal context derivation are restricted in a way that keeps them within the area of spatial segment 110.

[0030] When encoding the video to which picture 100 belongs in a video data stream, one point of interest is the fact that slices 108 are provided with slice addresses indicating their position in the encoded picture area, i.e. their location. The slice addresses are assigned along encoding order 106. For example, the slice addresses indicate the CTB order along encoding order 106 at which the encoding of the respective slice starts. For example, in a data stream encoding the video and picture 100, respectively, the slice portion carrying the slices coinciding with tile a will have a slice address of 7, since the seventh CTB according to encoding order 106 represents the first CTB in tile a according to encoding order 106. In a similar way, the slice addresses within the slice portion carrying the slices pertaining to tiles b, c and d will be 9, 29 and 33, respectively.

[0031] Figure 1 The right side indicates the slice addresses assigned by a receiver of the reduced or extracted video data stream to the two slices corresponding to tiles a, b, c and d. In other words, Figure 1 The numbers 0, 2, 4 and 8 on the right side show the slice addresses assigned by a receiver of the reduced or extracted video data stream, which is obtained from the original video data stream representing the video containing the entire picture 100 by extraction for spatial segment 110, i.e. by MCTS extraction. In the reduced or extracted video data stream, the slice portions of slices 100 of tiles a, b, c and d are encoded in the same order along encoding order 106 as in the original video data stream from which they are taken in the extraction process. Specifically, the receiver places the picture content (i.e. the slices pertaining to tiles a, b, c and d) along encoding order 112 in the form of CTBs reconstructed from the sequence of slice portions in the reduced or extracted video data stream, encoding order 112 traverses spatial segment 110 in the same way as encoding order 106 traverses the entire picture 100, i.e. spatial segment 110 is traversed tile by tile in a raster scan order, and the CTBs within each tile are likewise traversed along the raster scan order before continuing with the next tile. The relative positions of tiles a, b, c and d remain unchanged. That is, as Figure 1As shown, the spatial segment 110 on the right side preserves the relative positions of tiles a, b, c and d as they appear in the picture 100. As a result of determining the addresses by using the encoding order 112, the slice addresses of the slices corresponding to tiles a, b, c and d are 0, 2, 4 and 8, respectively. Thus, a receiver is able to reconstruct a smaller video based on the down-scaled or extracted video data stream, which shows the spatial segment 110 as illustrated on the right side as a standalone picture.

[0032] Summarizing the description of the Figure 1 so far, Figure 1 The adjustments to the slice addresses and CTB units after extraction are shown or explained by using the numbers of the upper left corners of each slice 108. For performing the extraction, the extraction site or extractor device needs to analyze the original video data stream with respect to parameters indicating the CTB size, i.e. the maximum CTB size 114, as well as the number and size of tile columns and tile rows in the picture 100. Furthermore, the nested MCTS-specific sequences and picture parameter sets are inspected in order to derive the output arrangement of tiles therefrom. In Figure 2 , the tiles a, b, c and d within the spatial segment 110 preserve their relative arrangement. In summary, the above-described analysis and inspection to be performed by the extractor device according to the MCTS instructions of the HEVC data stream require dedicated and complex logic to derive the slice addresses of the reconstructed slices 108 from the above-listed parameters. Such dedicated and complex logic in turn leads to additional implementation costs and run-time disadvantages.

[0033] Therefore, the embodiments described in the following use additional signaling in the video data stream as well as respective processing steps on the extraction information generation side and the extraction side, such that the just explained derived overall processing burden of the extractor device can be alleviated by providing readily available information dedicated for extraction purposes. Additionally or alternatively, some of the embodiments described in the following use the additional signaling in order to direct the extraction process in a certain way, thereby enabling a more efficient processing of different types of video content.

[0034] First, the general concept is explained based on Figure 2 . Then, the operation modes of the individual entities participating in the overall process as illustrated in Figure 2 are further described in different ways according to the different embodiments further below. It should be noted that although described together in one figure for the sake of understanding, the entities and blocks shown therein belong to independent devices, each of which individually inherits the features providing the advantages outlined in general in Figure 2 . More precisely, Figure 2A generation process of a video data stream is shown, a process of providing extraction information for such a video data stream, the extraction process itself, then a process of decoding the extracted or down-scaled video data stream, and participating devices, wherein the operating mode of these devices or the execution of individual tasks and steps is in accordance with the presently described embodiments. According to the first reference Figure 2 The described specific implementation examples, and as further outlined below, reduce the processing overhead associated with the extraction task of the extractor device. According to further embodiments, various different types of video content within the original video are additionally or alternatively relieved from processing.

[0035] At the top of Figure 1 , an original video is shown, indicated by reference numeral 120. This video 120 consists of a sequence of pictures, one of which is indicated with reference numeral 100, as it plays the same role as Figure 1 picture 100 in Figure 2 , namely, it shows the picture area from which the spatial segment 110 will be cut out later by the video data stream extraction. However, it is to be understood that the tile subdivision explained above with respect to Figure 2 does not need to be based on the video encoding used by the process shown in Figure 2 , it is to be understood that tiles and CTBs represent semantic entities in video encoding, just as such.

[0036] Figure 2It is shown that video 102 is video encoded in video encoding core 122. The video encoding performed by video encoding core 122 converts video 120 into video data stream 124 using, for example, hybrid video coding. That is, video encoding core 122 uses, for example, block-based predictive coding, which encodes individual picture blocks of pictures of video 120 using one of several supported prediction modes and encodes prediction residuals. The prediction modes can include, for example, spatial prediction and temporal prediction. Temporal prediction can involve motion compensation (i.e., determining a motion field) and transmitting this motion field along with data stream 124 as a motion vector for a block of the temporal prediction. The prediction residuals can be transform coded. That is, some spectral decomposition can be applied to the prediction residuals and the resulting spectral coefficients can be quantized and losslessly encoded into data stream 124 using, for example, entropy coding. The entropy coding can in turn use context adaptivity, i.e., a context can be determined, where this context derivation can depend on spatial and / or temporal neighborhoods. As mentioned above, the encoding can be based on encoding order 106, which has the following restriction on encoding dependencies, namely, only video portions that have already been traversed in encoding order 106 can be used as a basis or reference for encoding a current portion of video 120. Encoding order 106 traverses video 120 picture by picture, but not necessarily in the presentation time order of the pictures. In a picture such as picture 100, video encoding core 122 subdivides the encoding data obtained by the video encoding into slices 108, thereby subdividing picture 100 into slices 108, each of which corresponds to a respective slice portion 126 of video data stream 124. Within data stream 124, the plurality of slice portions 126 form a sequence of slice portions following each other in the order in which the corresponding slices 108 are traversed in picture 100 in encoding order 106.

[0037] Also as Figure 2 shown, video encoding core 122 provides or encodes a slice address into each slice portion 126. For illustration purposes, the slice address is denoted in capital letters in Figure 1 . As described with respect to Figure 2 , the slice address can be determined in some appropriate units, for example, in units of CTBs along encoding order 106 in one dimension, but alternatively, the slice address can be determined differently with respect to some predetermined point within the picture area occupied by the pictures of video 120, for example, the upper left corner of picture 100.

[0038] In this way, video encoding core 122 receives video 120 and outputs video data stream 124.

[0039] As already outlined above, with respect to spatial segment 110, according to Figure 2The generated video data stream will be extractable, accordingly, the video encoding core 122 has adapted the video encoding process appropriately. To this end, the video encoding core 122 restricts the inter-frame coding dependencies in such a way that parts within the spatial segment 110 are encoded into the video data stream 124 in a way that they do not depend on parts outside the spatial segment 110 by, for example, spatial prediction, temporal prediction or context derivation. The slices 108 do not cross the boundaries of the segment 110. Thus, each slice 108 is either completely within the segment 110 or completely outside the segment 110. It should be noted that the video encoding core 122 can adhere to not only one spatial segment 110, but also to several spatial segments. These spatial segments can intersect each other (i.e. they can partially overlap) or one spatial segment can be completely located within another spatial segment. Due to these measures, as will be explained in more detail later on, a downscaled or extracted video data stream of pictures smaller than the pictures of the video 120 (i.e. pictures showing only the content within the spatial segment 110) can be extracted from the video data stream 124 without the need for re-encoding, i.e. without the need to perform complex tasks such as motion compensation, quantization and / or entropy encoding again.

[0040] The video data stream 124 is received by a video data stream generator 128. Specifically, the video data stream 124 is received by the video data stream generator 128 in accordance with the extraction information 132. Figure 2 The video data stream generator 128 comprises, in the shown embodiment, a receiving interface 130 which receives the prepared video data stream 124 from the video encoding core 122. It should be noted that, according to an alternative, the video encoding core 122 can be comprised in the video data stream generator 128, thereby replacing the interface 130.

[0041] The video data stream generator 128 provides extraction information 132 for the video data stream 124. In Figure 2 The resulting video data stream output by the video data stream generator 128 is indicated using reference numeral 124'. For the extraction of the downscaled or extracted video data stream 136 from the video data stream 124', the extraction information 132 indicates how to extract the downscaled or extracted video data stream 136 in which a spatially smaller video 138 corresponding to the spatial segment 110 is encoded. Figure 2 The extractor device shown, the extraction information 132 indicates how to extract the downscaled or extracted video data stream 136 from the video data stream 124' in which a spatially smaller video 138 corresponding to the spatial segment 110 is encoded. The extraction information 132 comprises a first information 140 defining the spatial segment 110 within the picture area transmitted by the picture 100 and a second information 142 signaling one of a plurality of options regarding how to modify the slice addresses of the slice portions 126 falling into each slice 108 within the spatial segment 110 to indicate within the downscaled video data stream 136 where the respective slice is located in the reduced picture area of the picture 144 of the video 138.

[0042] In other words, the video data stream generator 128 only appends, i.e. adds, some information, i.e. the extraction information 132, to the video data stream 124 to derive the video data stream 124'. This extraction information 132 is intended to guide the extractor device 134 that is to receive the video data stream 124' in order to extract from this video data stream 124' a downscaled or extracted video data stream 136 specifically for the section 110. The first information 140 defines the spatial section 110, i.e. its location within the picture area of the video 120 and the picture 100, respectively, and possibly the size and shape of the picture area of the picture 144. As Figure 2 indicated, this section 110 does not necessarily have to be rectangular, convex or not necessarily a connected area. For example, in the example of Figure 2 the section 110 consists of two disjoint partial areas 110a and 110b. In addition, the first information 140 can contain hints as to how the extractor device 134 should modify or replace some encoding parameters or a portion of the data stream 124 or 124', respectively, e.g. an adjustment of a picture size parameter to reflect changes on the picture area that occur when going from the picture 100 to the picture 144 by the extraction operation. In particular, the first information 140 can comprise replacement or modification instructions for a parameter set of the video data stream 124' that is applied by the extractor device 134 in the extraction process in order to modify or replace the corresponding parameter set contained in the video data stream 124' and received into the downscaled or extracted video data stream 136 accordingly.

[0043] In other words, the extractor device 134 receives the video data stream 124', reads the extraction information 132 from the video data stream 124' and obtains the spatial section 110, i.e. its location and position within the picture area of the video 120, from this extraction information based on the first information 140. Thus, based on the first information 140, the extractor device 134 identifies some slice portions 126 in which slices are encoded that fall within the section 110 and that are thus to be received into the downscaled or extracted video data stream 136, while slice portions 126 that relate to slices outside the section 110 are discarded by the extractor device 134. In addition, the extractor device 134 can use the information 140 in order to correctly set, i.e. by modifying or replacing, one or more parameter sets within the data stream 124' before or during the adoption of the one or more parameter sets in the downscaled or extracted video data stream 136 as just outlined. Thus, the one or more parameter sets can relate to a picture size parameter, if the section 110 is as Figure 2The picture size parameter can be set to a size corresponding to the sum of the sizes of the areas of the segments 110, i.e. the sum of the areas of all parts 110a and 110b of the segments 110, according to the information 140, not the connecting areas as exemplarily depicted. This segment 110 sensitive discarding of the slice parts and the parameter set adjustment restricts the video data stream 124' to the segments 110. Furthermore, the extractor device 134 modifies the slice addresses of the slice parts 126 within the downscaled or extracted video data stream 136. In Figure 2 The slices are shown in the middle using hedging. I.e. the hatched slice parts 126 are those parts in which the slices fall into the segments 110 and are thus respectively extracted or received.

[0044] It should be noted that the information 142 can be conceived not only in the case that the information 142 is added to the complete video data stream 124' which comprises a sequence of slice parts comprising slice parts encoding slices located within the spatial segment and slice parts encoding slices located outside the spatial segment. Rather, the data stream containing the information 142 can already be stripped such that the sequence of slice parts comprised by the video data stream comprises slice parts encoding slices located within the spatial segment but does not comprise slice parts encoding slices located outside the spatial segment.

[0045] In the following, different examples for embedding the second information 142 into the data stream 124' and their processing are given. Generally, the second information 142 is signaled within the data stream 124' as a signal which explicitly signals a hint or in the form of one of a plurality of options which is a hint as to how to perform the slice address modification. In other words, the second information 142 is signaled in the form of one or more syntax elements whose possible values can for example explicitly signal a slice address replacement value or can together allow to distinguish a plurality of possibilities to associate the slice address of each slice portion 126 in the video data stream 136 with the setting of a selected one of one or more syntax elements in that data stream. However, it should be noted that the number of meaningful or allowed settings of the one or more syntax elements contained in the second information 142 just mentioned depends on the way the video 120 has been encoded into the video data stream 124 and on the selection of the section 110, respectively. For example, imagine that the section 110 is a rectangular connected area in the picture 100 and that the video encoding core 122 is to perform the encoding on that section 110 without further restricting the encoding with respect to the interior of the section 110. A section 110 consisting of two or more areas 110a and 110b would not be applicable. That is, only dependencies outside the section 110 would be suppressed. In this case, the section 110 has to be mapped onto the picture area of the picture 144 of the video 138 without modification, i.e. without disturbing the position of any sub-area of the section 110, and by placing the interior contents of the section 110 as they are in the picture area of the picture 144, the allocation of the addresses a and b to the slice portions carrying the slices constituting the section 110 can be uniquely determined. In this case, the setting of the information 142 generated by the video data stream generator 128 would be unique, i.e. the video data stream generator 128 has no choice but to set the information 142 in this way, although from an available encoding point of view, the information 142 would have other signal options. However, even in this unvaried case, an explicit indication of the signal 142, for example, of the only slice address modification has the advantage that the extractor device 134 itself does not have to perform the tedious task of determining the slice addresses a and b of the slice portions 126 employed from the stream 124' as described above. Instead, it simply learns from the information 142 how to modify the slice addresses of the slice portions 126.

[0046] Depending on the different embodiments for the nature of the information 142, which are further outlined below, the extractor device 134 either preserves or maintains the order in which the slice portions 126 are received from the stream 124' into the reduced or extracted stream 136, or modifies this order in a way defined by the information 142. In any case, the reduced or extracted data stream 136 output by the extractor device 134 can be decoded by a normal decoder 146. The decoder 146 receives the extracted video data stream 136 and decodes therefrom a video 138, the pictures 134 of which are smaller than the pictures of the video 120 (e.g. the picture 100), and whose picture area is filled by placing the slices 108 decoded from the slice portions 126 within the video data stream 136 in a way defined by the slice addresses a and β contained within the slice portions 126 within the video data stream 136.

[0047] That is, so far, the Figure 1 embodiments have been described in a way that makes their description suitable for various embodiments for the exact nature of the second information 142, which will be described in more detail below.

[0048] The embodiments now described use an explicit signaling of the slice addresses that the extractor device 134 shall use when modifying the slice addresses of the slice portions 126 received from the stream 142 into the stream 136. The embodiments described in the following use a signal 142 that allows signaling to the extractor device 134 one of several allowed options on how to modify the slice addresses. As a result of the already encoded portions 110, the allowed options are for example in a way that limits the encoding interdependencies within a section 110 so as not to cross the spatial boundaries of the section 110, as Figure 1 indicated, the spatial boundaries of the section 110 in turn divide the section 110 into two or more areas, e.g. 110a and 110c or tiles a, b, c, d. The latter embodiments can still involve the tedious task of the extractor device 134 itself performing the computation of the addresses, but allow for an efficient handling of different types of picture content within the original video 120, resulting in a meaningful video 138 at the receiving end from the respective extracted or reduced video data stream 136.

[0049] That is, as outlined above with respect to Figure 1 According to embodiments of the present application, the information on how to modify the addresses in the extraction process is explicitly transmitted by the second information 142, thus relieving the tedious task of determining the slice addresses in the extractor device 134, as outlined above with respect to

[0050] In particular, the information 142 can be used to explicitly signal the new slice addresses to be used in the slice headers of the extracted MCTS by including a list of slice address replacement values in the bitstream 124' in the same order in which the slice portions 126 carrying the slice addresses 124' in the bitstream 124' are included. See, for example, the example in Figure 1 Here, the information 142 would be an explicit signaling of the slice addresses following the order of the slice addresses 124' in the bitstream. Again, the slices 108 and the corresponding slice portions 126 can be received in the extraction process such that the order in the extracted video data stream 136 corresponds to the order in which these slice portions 126 are contained in the video data stream 124'. According to the subsequent syntax example, the information 142 explicitly signals the slice addresses in a way starting with the second slice or slice portion 126. In the case of Figure 3 , this explicit signaling would correspond to the second information 142 indicating or signaling the list {2, 4, 8} with new numbers. In the syntax example of Figure 3 , an exemplary syntax for this embodiment is presented, which shows the corresponding added explicit signaling 142 in addition to the MCTS extraction information SEI learned from [2] by highlighting.

[0051] The semantics are listed below.

[0052] num_associated_slices_minus2[i] plus 2 specifies the number of slices containing MCTSs with mcts identifiers equal to any value in the list mcts_identifier[i][j]. The value of num_extraction_info_sets_minus1[i] shall be in the range of 0 to 2 32 -2, inclusive.

[0053] output_slice_address[i][j] specifies the slice address of the j-th slice belonging to the MCTS with mcts identifiers equal to any value in the list mcts_identifier[i][j] in bitstream order. The value of output_slice_address[i][j] shall be in the range of 0 to 2 32 -2, inclusive.

[0054] It should be noted that the presence of information 142 in the MCTS extraction information SEI or of information 142 in addition to the MTCS related information 140 can be controlled by one flag in the data stream. This flag can be named slice_reordering_enabled_flag or the like. If this flag is set, information 142 such as num_associated_slices_minus2 and output_slice_address is present in addition to information 140, if this flag is not set, information 142 is not present, and the mutual position arrangement of the slices is adhered to or otherwise handled in the extraction process.

[0055] Furthermore, it should be noted that the term "segment" in the syntax element names used in the H.265 / HEVC, Figure 3 The "segment" part in the syntax element names used in the H.265 / HEVC,

[0056] Even more, it should be noted that although num_associated_slices_minus2 suggests that information 142 indicates the number of slices in segment 110 in integer form, this integer represents the number of slices in form of a difference to 2 indicating the number of slices, it is also possible to signal the number of slices within segment 110 directly in the data stream or in form of a difference to 1. For the latter option, for example, num_associated_slices_minus1 can be used as syntax element name. Note that, for example, the number of slices within any segment 110 can also be allowed to be one.

[0057] In addition to the MCTS extraction process as so far expected in [2], other processing steps are associated with the explicit signaling of information 142 by means of, for example, Figures 1 to 3 These other processing steps help the derivation of the slice addresses by the extractor device 134 in the extraction process, and the following outline of this extraction process points out where such help occurs by underlining:

[0058] Let the bitstream inBitstream, the target MCTS identifier mctsIdTarget, the target MCTS extraction information set identifier mctsEISIdTarget, and the target highest Temporalld value mctsTIdTarget be inputs of the sub-bitstream MCTS extraction process.

[0059] The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream.

[0060] The requirement for bitstream consistency of the output bitstream is that any output sub-bitstreams that are the output of the processes specified in this section regarding bitstreams should be consistent bitstreams.

[0061] The output sub-bit stream is obtained as follows:

[0062] - The bitstream outBitstream is set to be the same as the bitstream inBitstream.

[0063] -Remove all NAL cells from outBitstream whose TemporalId is greater than mctsTIdTarget.

[0064] - For each remaining VCL NAL in each access unit of outBitstream, adjust the slice fragment header as follows:

[0065] - For the first VCL NAL unit, set the value of first_slice_segment_in_pic_flag to 1, otherwise set it to 0.

[0066] - Based on the list output_slice_address[i][j], set the value of slice_segment_address for non-first NAL units (i.e., slices) starting from the second one in the bitstream order.

[0067] Just now about Figure 3 The described embodiment variant alleviates the tedious tasks of slice address determination and extraction processing performed by the extractor device 134 by using information 142 as an explicit signal of the slice address. Figure 1 For a specific example, the information includes an alternative slice address 143, which applies only to each second and subsequent slice segment 126 in the slice segment sequence, and maintains the slice order as slice segments 126 relating to slice 108 within segment 110 are received from data stream 124' to data stream 136. The slice address substitution value is associated with a one-dimensional slice address assignment in the picture region of picture 144 in video 138 using sequence 112, and does not conflict with laying out only the picture region of picture 144 along sequence 112, which uses a sequence of slices 108 obtained from the received slice segment sequence. It should be noted, and will be further mentioned below, that an explicit signal can also be applied to the first slice address, i.e., to the slice address of the first slice segment 126. Even for the latter, substitution value 143 may be included in information 142. This signaling of information 142 also enables the placement of slice 108 corresponding to the first slice segment 126, except at the beginning of the encoding sequence 112 (which may be as follows). Figure 3The first slice portion 126 is the first slice portion among the slice portions 126 carrying the slices 108 within the segment 110 in the order within the stream 124' at a position outside the top left corner of the picture shown. The explicit signaling of the slice addresses can also be used to provide greater freedom in rearranging the segment regions 110a and 110b if such a possibility exists or is allowed. For example, in the modification example depicted in Figure 2 , the information 142 also explicitly signals the slice address replacement values 143 (i.e. alpha and beta in the case of Figure 2 ) for the first slice portion 126 within the data stream 136, then the signal 142 will enable distinguishing between the two allowed or available placements of the segment regions 110a and 110b within the output picture 144 of the video 138. Namely, one is that the segment region 110a is placed on the left side of the picture, so that the slices 108 corresponding to the first transmitted slice portion 126 within the data stream 136 remain at the start position in the encoding order 112, and the order between the slices 108 is maintained unchanged compared to the slice order in the picture 100 of the video 120 (in terms of the video data stream 124). The other is that the segment region 120a is placed on the right side of the picture 144, so that the order of the slices 108 within the picture 144 traversed in the encoding order 112 is changed with respect to the order in which these slices are traversed in the original video within the video data stream 124 in the encoding order 106. The signal 142 will explicitly signal the slice addresses used in the modification by the extractor device 134 as a list of slice addresses 143 ordered or assigned to the slice portions 126 within the video data stream 124'. That is, the information 142 will indicate the slice addresses in sequence for each slice within the segment 110 in the order in which the slices 108 would be traversed in the encoding order 106, so that the explicit signaling can result in the following permutation, that the order in which the segment regions 110a and 110b of the segment 110 are traversed by the order 112 is changed compared to the order in which they are traversed by the original encoding order 106. The slice portions 126 received from the stream 124' into the stream 136 will be reordered accordingly by the extractor device 134, i.e. so as to comply with the order in which the slices 108 of the extracted slice portions 126 are sequentially traversed by the order 112. According to the standard conformance of the downscaled or extracted video data stream 136, the slice portions 126 received from the video data stream 124' should strictly follow each other along the encoding order 112, i.e. should have monotonically increasing slice addresses alpha and beta modified by the extractor device 134 in the extraction process, whereby the standard conformance will be maintained. Thus, the extractor device 134 will modify the order between the received slice portions 126 so as to order the received slice portions 126 according to the order of the slices 108 encoded into it along the encoding order 112.

[0068] According to Figure 4 a further variant of the description, the latter aspect is utilized, namely the possibility to rearrange the received slice portions 126 of the slices 108. Here, the second information 142 signals a rearrangement of the order between the slice portions 126 in which any of the slices 108 located within the segment 110 is encoded. One possibility of the presently described embodiment is to signal the slice address replacement values 143 explicitly in a way that leads to a rearrangement of the slice portions. However, the information 142 can signal the rearrangement of the slice portions 126 within the extracted or reduced video data stream 136 in a different way. The present embodiment finally places the reconstructed slices reconstructed from the slice portions 126 strictly along the encoding order 112 in the decoder 146 (which means for example using a tile raster scan order) filling the picture area of the picture 144. The way of rearrangement signaled by the information 142 has been chosen in a way that the order between the extracted or received slice portions 126 is rearranged or changed such that the placement process of the decoder 146 can lead to segment areas such as the segment areas 110a and 110b changing their order compared to the slice portions not being rearranged. If the way of rearrangement signaled by the information 142 preserves the order as originally in the data stream 124', the segment areas 110a and 110b can keep their relative position in the original picture area of the picture 110.

[0069] For explaining the present variant, please refer to Figure 4 . Figure 4 The picture 110 of the original video and the picture area of the picture 144 of the extracted video are shown. Furthermore, Figure 4 an exemplary tile partitioning (i.e. partitioning into tiles 104) is shown and an exemplary extraction segment 110 is shown which comprises two disjoint segment areas or zones 110a and 110b, namely two segment areas 110a and 110b adjacent to opposite edges 150 r and 150 l of the picture 110, here the segment areas are in their position coinciding with the extraction direction of the edges 150 r and 150 l (i.e. along the vertical direction). In particular, Figure 4 it is shown that the picture content of the picture 110 and thus the video to which the picture 110 belongs is of a specific type (i.e. a panoramic video) such that the edges 150 r and 150 l constitute one scene when projecting the 3D scene onto the picture area of the picture 110. Below the picture 110, Figure 4It shows two options on how to place segments 110a and 110b within the output image area of ​​image 144 of the extracted video. Figure 4 The document also uses tile names to describe these two options. These options stem from the fact that video 210 has been encoded into stream 124 in a certain way, such that for each of the two regions 110a and 110b, encoding occurs independently of the external encoding. It should be noted that, in addition to... Figure 4 In addition to the two allowed options shown, two other options are possible in the following cases: each zone in zones 110a and 110b is subdivided into two tiles within each zone, and the video is encoded independently of external encoding into stream 124, i.e. Figure 4 Each tile in the diagram is either an extractable part or an independently coded part (relative to spatial and temporal interdependence). These two options then correspond to shuffling tiles a, b, c, and d in different ways within segment 110.

[0070] Once again, now for reference Figure 4 The described embodiments are intended to alter the order in which the second information 142 is transmitted from the data stream 124' to, adopted, or written into the extracted or scaled video data stream 136 during the extraction process of the extractor device 134, either to the slice 108 or the NAL unit of the data stream 124' carrying the slice 108. Figure 4 This illustrates a scenario where the desired image sub-segment 110 consists of non-adjacent tiles or segment regions 110a and 110b within the image plane spanned by image 110. The complete coded image plane of image 110 is... Figure 4 The top of the image shows a tile boundary and a desired MCTS 110, which consists of two rectangles 110a and 110b comprising tiles a, b, c, and d. Regarding the scene content, or due to... Figure 4 The fact that the video content shown is panoramic video content means that when image 110 covers the 360° environment around the camera by means of the equirectangular projection shown in this example, the desired MCTS 110 surrounds the right and left boundaries 150. r and 150 l In other words, because Figure 4 Given the illustrated scenario (i.e., the image content is panoramic), the second option would actually make more sense than the first (1) option (2) concerning placing segment regions 110a and 110b within the image area of ​​the output image 144. However, the situation might be different if the image content were other types of images (e.g., non-panoramic images).

[0071] In other words, the order of tiles A, B, C and D in the full picture bitstream 124' is {a, b, c, d}. If this order is simply transferred to the encoded order in the placement of the respective tiles in the extracted or downscaled video data stream 136 or output picture 144, then in the above exemplary case, as Figure 4 indicated in the lower left, the extraction process itself would not yield the desired data arrangement within the output bitstream 136. As Figure 4 indicated in the lower right, the preferred arrangement {b, a, d, c} is shown, which results in a video bitstream 136 that produces a continuous picture content on the picture plane of the picture 144 for a legacy device such as a decoder 146. Such a decoder 146 can not have the capability to rearrange the subpicture regions of the output picture 144 in the pixel domain as a post-processing step after decoding (i.e. rendering), and even sophisticated devices prefer to avoid the effort of post-processing.

[0072] Thus, according to the above example with regard to Figure 5 the second information 142 provides the encoding side of the video data stream generator 128 with means to signal a preferred order among several choices or options, respectively, according to which the segment regions 110a and 110b within the video data stream 124', respectively consisting of a set of one or more tiles, shall be arranged in the extracted or downscaled video data stream 136 or the picture area covered by its picture 144. According to Figure 4 the specific syntax example presented in Figure 5 , the second information 142 comprises a list that is encoded into the data stream 124' and indicates the position of each slice 108 falling into a segment 110 within the extracted bitstream in the order of the original or input bitstream, i.e. in the order in which they occur in the bitstream 124'. For example, in the example of Figure 5 , the preferred option 2 would finally become the list read as {1, 0, 3, 2}.

[0073] The semantics are as follows.

[0074] num_associated_slices_minus1[i] plus 1 specifies the number of slices containing MCTSs with mcts identifiers equal to any of the values in the list mcts_identifier[i][j]. The value of num_extraction_info_sets_minus1[i] shall be in the range of 0 to 2 32 -2, inclusive.

[0075] output_slice_order[i][j] identifies the absolute position of the j-th slice in bitstream order, which belongs to the MCTS with mcts identifier equal to the value of the list mcts_identifier[i][j] in the output bitstream

[0076] The value of output_slice_order[i][j] shall be in the range of 0 to 2 23 - 2, inclusive.

[0077] The other processing steps in the extraction process defined in [2] are described next, which help to understand the signaling embodiment of Figure 2 , where the additional content with respect to [2] is highlighted by underlining:

[0078] Let bitstream inBitstream, target MCTS identifier mctsIdTarget, target MCTS extraction information set identifier mctsEISIdTarget, and target highest Temporalld value mctsTIdTarget be the inputs of the sub-bitstream MCTS extraction process.

[0079] The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream.

[0080] The requirement on bitstream conformance for the input bitstream is that any output sub-bitstream as output of the processes specified in this section for the bitstream shall be a conforming bitstream.

[0081] For the i-th extraction information set, derive OutputSliceOrder[j] from the list output_slice_order[i][j].

[0082] The output sub-bitstream is derived as follows:

[0083] - Set the bitstream outBitstream to be identical to the bitstream inBitstream. [...]

[0085] - Remove from outBitstream all NAL units with Temporalld greater than mctsTIdTarget.

[0086] - Order the NAL units of each access unit according to the list OutputSliceOrder[j].

[0087] - For each VCL NAL unit remaining in outBitstream, adjust the slice segment header as follows:

[0088] - for the first VCL NAL unit within each access unit, the value of first_slice_segment_in_pic_flag is set to 1, otherwise it is set to 0.

[0089] - the value of slice_segment_address is set according to the tile setting defined in the PPS with pps_pic_parameter_set_id equal to slice_pic_parameter_set_id.

[0090] Thus, in summary Figure 5 According to the above variants of Figure 3 the above variants discussed with respect to Figure 5 the difference is that the second information 142 does not explicitly signal how to modify the slice addresses. That is, according to the just outlined variants, the second information 142 does not explicitly signal replacement values for the slice addresses of the slice portions extracted from the data stream 124 into the data stream 136. Rather, Figure 5 Embodiments of Figure 5 discussed above, the second information 142 signals reordering information indicating how to reorder the slice portions 126 of the slices 108 located in the segment 110 with respect to the order of the slice portions in the video data stream 124’ when extracting the downsized video data stream 136 from the video data stream 124’. The reordering information 142 can comprise, for example, a set of one or more syntax elements. Among the states possibly signaled by the one or more syntax elements forming the information 142, there can be a state according to which the reordering remains the original order. For example, the information 142 signals an order of the slice portions 126 encoding the slices 108 of the pictures 100 falling into the segment 110 that is identical to the order of the slice portions 126 in the video data stream 124’. Figure 4the extractor device 134 reorders the slice portions 126 within the extracted or downscaled video data stream 136 according to these orders. Then, the extractor device 134 modifies the slice addresses of the slice portions 126 so rearranged within the downscaled or extracted video data stream 136 in the following way: the extractor device 134 knows those slice portions 126 in the video data stream 124' that were extracted from the data stream 124' into the data stream 136. Thus, the extractor device 134 knows the slices 108 within the picture area of the picture 100 that correspond to these received slice portions 126. Based on the rearrangement information provided by the information 142, the extractor device 134 is able to determine how the segment areas 110a and 110b are mutually displaced in a translational manner so as to form the rectangular picture area corresponding to the picture 144 of the video 138. For example, in the exemplary case of option 2 of Figure 4 In the exemplary case of option 2 of the above, the extractor device 134 assigns the slice address 0 to the slice corresponding to tile b, since the slice address 0 appears in the second position in the order list provided by the second information 142. Thus, the extractor device 134 is able to place one or more slices related to tile b and then continues with the next slice, which according to the reordering information is associated with a slice address pointing to a position in the picture area following tile b according to the encoding order 112. In the exemplary case of option 2 of the above, this is the slice related to tile a, since the next order position indicates the first slice a in the segment 110 of the picture 100. In other words, the reordering is limited to yield any possible rearrangement of the segment areas 110a and 110b. Separately for each segment area, the respective segment area is still traversed in the same path in the encoding order 106 and 112. However, due to the tile partitioning, it is possible that, in the case of option 2, the areas corresponding to the segment areas 110a and 110b in the picture area of the picture 144, i.e. the areas of the combination of tiles b and d and the combination of tiles A and C, are traversed in an interleaved manner according to the encoding order 112 and, accordingly, the related slice portions 126 encoding the slices located in the respective tiles are interleaved in the extracted or downscaled video stream 136. Figure 5 In the exemplary case of option 2 of the above, the extractor device 134 assigns the slice address 0 to the slice corresponding to tile b, since the slice address 0 appears in the second position in the order list provided by the second information 142. Thus, the extractor device 134 is able to place one or more slices related to tile b and then continues with the next slice, which according to the reordering information is associated with a slice address pointing to a position in the picture area following tile b according to the encoding order 112. In the exemplary case of option 2 of the above, this is the slice related to tile a, since the next order position indicates the first slice a in the segment 110 of the picture 100. In other words, the reordering is limited to yield any possible rearrangement of the segment areas 110a and 110b. Separately for each segment area, the respective segment area is still traversed in the same path in the encoding order 106 and 112. However, due to the tile partitioning, it is possible that, in the case of option 2, the areas corresponding to the segment areas 110a and 110b in the picture area of the picture 144, i.e. the areas of the combination of tiles b and d and the combination of tiles A and C, are traversed in an interleaved manner according to the encoding order 112 and, accordingly, the related slice portions 126 encoding the slices located in the respective tiles are interleaved in the extracted or downscaled video stream 136.

[0091] Another embodiment is to signal the guarantee that another order signaled using existing syntax reflects the preferred output slice order. More specifically, this embodiment can be implemented by interpreting the occurrence of the MCTS extraction SEI message [2] as the guarantee that the order of the rectangles constituting the MCTS in the MCTS SEI message from sections D.2.29 and E.2.29 in [1] represents the preferred output order of tiles / NAL units. In Figure 1In a specific example of this, this would result in using the rectangles in the order {b, a, d, c} for each included tile. The example of this embodiment is identical to the example described above, e.g., except for the derivation of OutputSliceOrder[j].

[0092] OutputSliceOrder[j] is derived from the order of rectangles signaled in the MCTS SEI message.

[0093] Summarizing the above examples, the second information 142 can signal to the extractor 134 the following: how to reorder the slice portions 126 of slices falling into the spatial segments 110, relative to their ordering in the sequence of slice portions of the video data stream 124', along the first encoding scan order 106 that traverses the picture area, the slice addresses of each slice portion 126 in the sequence of slice portions of the video data stream 124', the one-dimensionally indexed encoding start positions of the slices 108 encoded into the respective slice portions 126, and wherein the picture 100 has been encoded into the sequence of slice portions of the video data stream along the first encoding scan order. Thereby, the slice addresses of the slice portions in the sequence of slice portions within the video data stream 124' monotonically increase, the modification of the slice addresses when extracting the downscaled video data stream 136 from the video data stream 124' is defined by sequentially placing the slices encoded into the slice portions along the second encoding scan order 112 that traverses the downscaled picture area, the slice portions being the slice portions the downscaled video data stream 136 is defined in and being the slice portions reordered as notified by the second information 142, and setting the slice addresses of the slice portions 126 to index the encoding start positions of the slices along the second encoding scan order 112. The first encoding scan order 106 traverses the picture area within each of the set of at least two segment areas in the same way as the second encoding scan order 112 traverses the respective spatial area. Each of the set of at least two segment areas is indicated by the first information 140 as a subarray of rectangular tiles (the picture 100 is subdivided into rows and columns of which the rectangular tiles) where the first and second encoding scan orders use a row-by-row tile raster scan to completely traverse the current tile before continuing with the next tile.

[0094] As mentioned above, the output slice order can be derived from another syntax element, e.g. output_slice_address[i][j] as mentioned above. In this case, an important supplement to the above example syntax for output_slice_address[i][j] is that the slice address is signaled for all associated slices, including the first associated slice, in order to enable ordering, i.e. num_associated_slices_minus2[i] becomes num_associated_slices_minus1[i]. The example of this embodiment is identical to the above example except for the derivation of OutputSliceOrder[j], e.g.

[0095] For the i-th extraction information set, OutputSliceOrder[j] is derived from the list output_slice_address[i][j].

[0096] Even further embodiments consist of a single flag on information 142 which indicates that the video content wraps around at a set of picture boundaries, e.g. the vertical picture boundaries. Accordingly, the output order is derived in extractor 134 which applies to the picture sub-portion including tiles on two picture boundaries as outlined previously. In other words, information 142 can signal one of two options: a first one of the multiple options indicates that the video is a panoramic video which shows the scene in a way that different edge portions of the picture abut on each other on the scene, and a second one of the multiple options indicates that the different edge portions do not abut on each other on the scene. The at least two segment areas a, b, c, d included by segment 110 form first and second zones 110a, 110b which abut on different edge portions, i.e. right edge 150 r and left edge 150 l , of the picture 144, and accordingly, in case the second information 142 signals the first option, the reduced picture area is constituted by putting the set of at least two segment areas together in a way that the first and second zones abut along the different edge portions, and in case the second information 142 signals the second option, the reduced picture area is constituted by putting the set of at least two segment areas together in a way that the different edge portions of the first and second zones face away from each other.

[0097] For the sake of completeness only, it should be noted that the shape of the picture area of picture 144 is not limited to conforming to the following shape: putting the individual areas, e.g. tiles a, b, c and d of segment 110, together while keeping the relative arrangement of any connected clusters unchanged, e.g. Figure 2(a, c, b, d) in (a, c, b, d); or piecing together any such clusters along the shortest possible interconnection direction, e.g. Figure 1 horizontal piecing zones (a, c) and (b, d) in (a, c, b, d). Conversely, e.g. in Figure 2 the picture region of a picture of the extracted data stream can be a column of four regions, in Figure 6 it can be a column of a row of four regions. Generally, the size and shape of the picture region of a picture 144 of the video 138 can be present in different parts of the data stream 124': e.g. this information can be given in the form of one or more parameters within the information 140, in order to guide the extraction process in the extractor 134 with respect to the adaptation of the parameter set when extracting a particular sub-stream 136 from the stream 124'. For example, a nested parameter set for replacing a parameter set in the stream 124' can be contained in the information 140, wherein the parameter set e.g. contains parameters related to the picture size, e.g. indicating the size and shape of the picture 144 in pixels, such that during extraction in the extractor 134, the replacement of the parameter set of the stream 124' covers the old parameters in the parameter set indicating the size of the picture 100. However, the picture size of the picture 144 can additionally or alternatively be indicated as part of the information 142. Explicitly signalling the shape of the picture 144 in an easily readable high-level syntax element such as the information 142 can be particularly advantageous if the slice addresses are not explicitly provided in the information 142. Then nested parameter sets would need to be parsed in order to derive the addresses.

[0098] It is further noted that in more complex system setups, a cubical projection can be used. This projection avoids the known drawbacks of equirectangular projection, e.g. the large variation in sampling density. However, when using a cubical projection, it is necessary to have a rendering stage to recreate a continuous view volume from the content (or a sub-portion thereof). Such a rendering stage can be traded off between complexity / capability, i.e. some feasible and off-the-shelf rendering modules can expect a given arrangement of the content (or a sub-portion thereof). In such a case, it is essential to enable the possibility of the arrangement by the disclosure below.

[0099] In the following, embodiments relating to the second aspect of the present application are described. The description of embodiments of the second aspect of the present application again starts with a brief introduction to the general problem that is envisaged and solved by these embodiments.

[0100] In the case of 360° video, but not limited to this case, one relevant use case for MCTS extraction is a composite video containing multiple resolution variants of the content next to each other on the picture plane, as Figure 6 shown in Fig. 1. Figure 6The bottom of Figure 3 depicts a synthetic video 300 with multiple resolutions, and the line 302 shows the tile boundaries of tiles 304 into which the synthesis of a high resolution video 306 and a low resolution video 308 is subdivided for encoding into respective data streams.

[0101] More precisely, Figure 6 A picture in the high resolution video is shown at 306, and a picture of the low resolution video, such as a co-temporal picture, is shown at 308. For example, both videos 306 and 308 show exactly the same scene, i.e. with the same graphics or view, but with different resolutions. However, the views can alternatively only partially overlap each other, and in the overlapping zone, the spatial resolution is manifested in that the number of samples of the picture of video 306 is different compared to video 308. In fact, the fidelity of videos 306 and 308 is different, i.e. the number of samples or pixels of the equal scene portion is different. The juxtaposed pictures of videos 306 and 308 are composed side by side into a larger picture, thereby generating a picture of the synthetic video 300. For example, Figure 6 The picture of video 308 is shown split in two halves in the horizontal direction, and one of the two halves is on top of the other, which are appended to the right of the co-temporal picture of video 306, thereby yielding the corresponding video 300. The tile partitioning is performed in such a way that no tile 304 crosses the border between the high resolution picture of video 306 and the picture content originating from the low resolution video 308. In the example of Figure 3, the picture 300 is partitioned into 8x4 tiles of the high resolution picture content of the picture of video 300, and 2x2 tiles of each of the two halves from the low resolution video 308. The picture of video 300 is thus composed of 10x4 tiles in total. Figure 7

[0102] When such a multi-resolution synthetic video 300 is encoded with MCTS in an appropriate way, MCTS extraction can yield a variant 312 of the content. As Figure 8 shown, such a variant 312 can for example be designed to depict a predefined sub-picture 310a at high resolution, and the rest of the scene or another sub-portion at low resolution, with the MCTS 310 in the compound picture bitstream compared to three aspects or three separate areas 310a, 310b, 310c.

[0103] That is, the picture of the extracted video, i.e. picture 312, has three image domains 314a, 314b and 314c, each corresponding to one of the MCTS areas 310a, 310b, 310c, with area 310a being a sub-area of the high resolution picture area of picture 300, and the other two areas 310b and 310c being sub-areas of the low resolution video content of picture 308. ​

[0104] As already mentioned, Figure 8 Embodiments of this application relating to the second aspect of this application are described. In other words, Figure 8 This illustrates a scheme for presenting multi-resolution content to a recipient site, specifically the various sites involved in generating multiple versions at different resolutions for the recipient site. It should be noted that along... Figure 8 The processing paths described herein, with each device and process located at each site, represent individual devices and methods; therefore, Figure 2 This should not be interpreted as merely illustrating a systemic or holistic approach. Regarding Figure 2 Similar statements are also correct. Figure 8 Individual devices and methods are also shown. All these sites are shown together in one figure solely for ease of understanding their interrelationships and the advantages of the embodiments described with respect to these figures.

[0105] Figure 8 Video data stream 330 is shown, containing video 332 encoded with image 334. Image 334 is the result of stitching together simultaneous image 336 of high-resolution video 338 and image 340 of low-resolution video 342. More precisely, images 336 and 340 of videos 338 and 342 either correspond to the exact same view area 344, or at least partially overlap to show the same scene in the overlapping area. However, for the same scene content, the number of sampled samples in high-resolution video image 336 is higher than the number of samples in the corresponding low-resolution image 340; therefore, the scene resolution fidelity of image 336 is higher than that of image 340. The synthesis of image 334 of synthesized video 332 is performed by compositor 346 based on videos 338 and 342. Compositor 346 stitches images 336 and 340 together. In doing so, the compositor 346 can subdivide the images of the high-resolution video 338 and the images 340 of the low-resolution video, or both, to obtain favorable filling and patching of the image regions of the images 334 of the composite video 332. The video encoder 348 then encodes the composite video 332 into a video data stream 330. The video data stream generation device 350 can be included in the video encoder 348, or can be connected to the output of the video encoder 348, to provide a signal 352 to the video data stream 330 indicating that the same scene content is displayed multiple times (i.e., in different spatial portions at different resolutions) for each image 334 or each image or video 332 in a sequence of images in the video 332 encoded into the video data stream 330. These portions are in... Figure 8H and L are shown in order to indicate their origin by the composition done by the compositor 346, and are indicated using reference numerals 354 and 356. It should be noted that more than two different resolution versions can be combined to form the content of the picture 332, and the use of two versions as shown in Figure 6 , Figure 7 and Figure 8 is only for illustrative purposes.

[0106] For the sake of illustration, Figure 6 a video stream processor 358 is shown receiving the video stream 330. The video stream processor 358 can for example be a video decoder. In any case, the video stream processor 358 is able to inspect the signal 352 in order to determine, based on this signal, whether the video stream processor 358 should start processing such as decoding the video stream 330, which further depends on certain functionalities of the video stream processor 358 or of a device connected downstream thereof. For example, the video stream processor 358 can only be able to render pictures 332 that are fully encoded in the video stream 330, then the video stream processor 358 can refuse to process a video stream 330 whose signal 352 indicates that its respective pictures show the same scene content at different spatial resolutions at different spatial portions of the respective pictures, i.e. the signal 352 indicates that the video stream 330 is a multi-resolution video stream.

[0107] The signal 352 can for example comprise a flag contained within the data stream 330, which flag can be switched between a first state and a second state. The first state can for example indicate the fact just outlined, i.e. that the respective pictures of the video 332 show the same scene content at different resolutions. The second state indicates that this is not the case, i.e. that the pictures only show one scene content at one resolution. The video stream processor 358 will therefore respond to the flag 352 being in the first state, so as to refuse to perform certain processing tasks.

[0108] The signal 352 such as the flag described above can be conveyed within the data stream 330, within its sequence parameter set or video parameter set. In the following description, as a possible candidate, a possible syntax element reserved for future use in HEVC is exemplarily identified.

[0109] As shown above in relation to Figure 7 and Figure 2 it is possible, although not necessary, that the composed video 332 has been encoded into the video stream 330 in such a way that, for each of a set of spatial regions 360 (such as tiles or tile sub-arrays) that do not overlap each other, the encoding is independent of the outside of the respective spatial region 360. As shown above in relation to Figure 7The coding dependency can limit the spatial and temporal prediction and / or context derivation such that no boundary between spatial regions 360 is crossed. Thus, for example, the coding dependency can limit the coding of a certain spatial region 360 of a certain picture 334 of the video 332 to only refer to spatial regions of the same position within another picture 334 of the video 332 that is subdivided into spatial regions 360 in the same way as the picture 334. The video encoder 348 can provide extraction information such as the information 140, or a combination of the information 140 and 142, or a corresponding device such as the video data stream generator 128 can be connected to its output, to the video data stream 330. The extraction information can relate to certain or all possible combinations of spatial regions 360 as extraction portions, for example Figure 2 the extraction sections 310 in the information 140. The signal 352 can in turn comprise information on a spatial subdivision of the pictures 334 of the video 332 into sections 354 and 356 with different scene resolutions, i.e. information on the size and position of each spatial section 354 and 356 within a picture region of the picture 332. Based on such information within the signal 352, the video data stream processor 358 can exclude certain extraction sections from a list of possible extractable sections of the video stream 330, for example, extraction sections that mix spatial regions 360 falling into different sections 354 and 356, thus avoiding performing video extraction for extraction sections that mix different resolutions. In this case, the video stream processor 358 can comprise an extractor device such as the extractor device 128, for example. ​

[0110] Additionally or alternatively, the signal 330 can comprise information on different resolutions at which the pictures 334 of the video 332 show the same scene content. Furthermore, the signal 352 can also only indicate a count of different resolutions at which the pictures 334 of the video 332 show the same scene content at different picture positions multiple times.

[0111] As already mentioned, the video data stream 330 can comprise extraction information on a list of possible extraction regions with respect to which the video data stream 330 can be extracted. The signal 352 can then comprise a further signal that indicates, for each of the at least one or more of these extraction regions, a view direction of a section region of the respective extraction region at which the same scene content is shown at the highest resolution within the respective extraction region, and / or an area share of a section region of the respective extraction region at which the same scene content is shown at the highest resolution within the respective extraction region in the total region of the respective extraction region, and / or a spatial subdivision of the respective extraction region from the entire region of the respective extraction region that subdivides the respective extraction region into section regions that show the same scene content at mutually different resolutions, respectively.​

[0112] Thus, such a signal 352 can be exposed at a high level in the bitstream so that it can be easily pushed into streaming systems.

[0113] One option is to use one of the generally reserved zero X bits flags in the profile tier level syntax. This flag can be named a general non multi-resolution flag:

[0114] A general non multi-resolution flag equal to 1 specifies that the decoded output pictures do not contain multiple versions of the same content with different resolutions (i.e. the corresponding syntax such as region packing is constrained). A general non multi-resolution flag equal to 0 specifies that the bitstream can contain such content (i.e. unconstrained).

[0115] In addition, the application thus comprises a signal informing of the nature of the complete bitstream content characteristics, i.e. the number of variants and resolutions in the composition. Furthermore, other signals provide information on the following features of each MCTS in the encoded bitstream in an easily accessible form:

[0116] • Main view orientation:

[0117] Orientation of the MCTS high resolution view area center, e.g. delta yaw, delta pitch and / or delta roll with respect to a predefined initial view area center.

[0118] • Overall coverage

[0119] Percentage of the total content represented in the MCTS.

[0120] • Ratio high / low resolution

[0121] Ratio between high resolution area and low resolution area in the MCTS, i.e. how much of the overall coverage is represented in high resolution / fidelity.

[0122] Signaling information has been proposed for the view direction or the overall coverage of a full omnidirectional video. Similar signaling needs to be added for the sub-areas that can be extracted. This information is in the form of SEI [2], so it can be included in the Motion Constrained Tile Set Extraction Information NESTED SEI. However, this kind of information is needed to select the MCTS to be extracted. Including the information in the Motion Constrained Tile Set Extraction Information NESTED SEI adds an additional level of indirection and requires deeper parsing (the Motion Constrained Tile Set Extraction Information NESTED SEI contains other information that is not needed to select the extraction set) to select a given MCTS. From a design point of view, a more straightforward approach is to signal this information or part of it in the center point that contains only the important information to select the extraction set. In addition, the mentioned signaling contains information about the whole bitstream, in the proposed case, it is desirable to signal the coverage of the high resolution and the coverage of the low resolution or, if more resolutions are mixed, the coverage of each resolution and the view direction of the mixed resolution extracted video.

[0123] One embodiment is to add the coverage of each resolution and add it to the 360 ERP SEI from [2]. Thus, this SEI can be included in the Motion Constrained Tile Set Extraction Information NESTED SEI and the aforementioned cumbersome task needs to be performed.

[0124] In another embodiment, a flag is added to the MCTS Extraction Information Set SEI, for example omnidirectional information indicating the presence of the signal in question, so that only the MCTS Extraction Information Set SEI is needed for selecting the set to be extracted.

[0125] Although some aspects have been described in the context of an apparatus, it is clear that separate aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps can be performed by (or using) a hardware apparatus, like, for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, some one or more of the most important method steps can be performed by such an apparatus.

[0126] The inventive data stream can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0127] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or in software. The implementation can be carried out using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium can be computer readable.

[0128] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0129] Generally, embodiments of the present application can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code can for example be stored on a machine readable carrier.

[0130] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0131] In other words, an embodiment of the inventive methods is, therefore, a computer program for performing one of the methods described herein, when the computer program runs on a computer.

[0132] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary.

[0133] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be configured to be transferred via a data communication connection, for example via the Internet.

[0134] A further embodiment comprises processing means, for example a computer, or a programmable logic device, configured to or adapted for performing one of the methods described herein.

[0135] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0136] Another embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0137] In some embodiments, a programmable logic device (for example a field programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0138] The apparatuses described herein can be implemented using a hardware apparatus, or using a computer, or using a combination of hardware and computer.

[0139] The apparatuses described herein, or any components of the apparatuses described herein, can be implemented at least partially in hardware and / or in software.

[0140] The methods described herein can be performed using a hardware apparatus, or using a computer, or using a combination of hardware and computer.

[0141] The methods described herein, or any components of the apparatuses described herein, can be performed at least partially in hardware and / or in software.

[0142] The embodiments described above are merely illustrative of the principles of this application. Modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended that the scope of the application be limited only by the scope of the appended patent claims, and not by the specific details presented by way of description and explanation of the embodiments herein.

[0143] Example Embodiment 1, a video data stream, comprising:

[0144] a sequence of slice portions (126), a respective slice (108) of a plurality of slices of a picture (100) of a video (120) being encoded in each slice portion (126), wherein each slice portion comprises a slice address indicating a position in a picture area of the video at which the slice encoded in the respective slice portion is located;

[0145] extracting information (132) indicative of how to extract a downsized video data stream (136) from the video data stream (130), wherein the downsized video data stream (136) has encoded therein a spatially smaller video (138) corresponding to a spatial segment (110) of the video, by restricting the video data stream to slice portions having encoded therein any slices within the spatial segment (110), and modifying the slice addresses to relate to a downsized picture area of the spatially smaller video (138), wherein the extracting information (132) comprises:

[0146] first information (140) defining the spatial segment (110) within the picture area, wherein no slice spans a boundary of the spatial segment (110); and

[0147] second information (142) signaling one of a plurality of options, or explicitly signaling, how to modify the slice addresses of the slice portions of each slice within the spatial segment (110) in order to indicate within the downsized video data stream (136) where the respective slice is located in the downsized picture area.

[0148] Example embodiment 2, the video data stream according to example embodiment 1, wherein the first information (140) defines the spatial segment (110) within the picture area as being composed of a set of at least two segment areas (110a, 110b; a, b, c, d), in each of which the video is encoded in the video data stream independently from outside the respective segment area, wherein no slice (108) spans a boundary of at least any one of the at least two segment areas.

[0149] Example embodiment 3, the video data stream according to example embodiment 2, wherein

[0150] a first one of the plurality of options indicates that the video is a panoramic video showing a scene in a way that different edge portions of the picture adjoin each other on the scene,

[0151] a second one of the plurality of options indicates that the different edge portions do not adjoin each other on the scene,

[0152] wherein the at least two segment areas form a first and a second zone (110a, 110b) adjacent to different ones of the different edge portions,

[0153] in order to,

[0154] In case the second information (142) signals the first option, the reduced picture region is composed by putting together the set of at least two segment regions such that the first zone and the second zone abut along the different edge portions, and

[0155] In case the second information (142) signals the second option, the reduced picture region is composed by putting together the set of at least two segment regions such that the different edge portions of the first zone and the second zone face away from each other.

[0156] Example embodiment 4, the video data stream according to example embodiment 2, wherein

[0157] The second information (142) signals how slice parts (126) of slices within the spatial segment (110) are reordered relative to an ordering of the slice parts (126) in a sequence of slice parts of the video data stream when extracting the reduced video data stream (136) from the video data stream, and

[0158] A slice address of each slice part in the sequence of slice parts of the video data stream indexes a coding start position of a slice encoded into the respective slice part along a first coding scan order (106) traversing the picture region, and wherein the picture (100) has been encoded in the sequence of slice parts of the video data stream along the first coding scan order such that

[0159] The slice addresses of the slice parts in the sequence of slice parts within the video data stream are monotonically increasing, and

[0160] The modification of the slice addresses when extracting the reduced video data stream (136) from the video data stream is defined by sequentially placing slices encoded into slice parts along a second coding scan order (112) traversing the reduced picture region, the slice parts being slice parts to which the reduced video data stream (136) is confined and being slice parts reordered as signaled by the second information (142), and setting a slice address of the slice part (126) to index a coding start position of a slice determined along the second coding scan order (112).

[0161] Example embodiment 5, the video data stream according to example embodiment 4, wherein the first coding scan order (106) traverses the picture region within each of the set of at least two segment regions in the same way as the second coding scan order (112) traverses the respective spatial region.

[0162] Example embodiment 6, the video data stream according to example embodiment 4 or 5, wherein each of the set of at least two segment regions is indicated by the first information (140) as a rectangular tile subarray, the picture (100) being subdivided into rows and columns of the rectangular tile subarrays, wherein the first and the second coding scan order use a line-by-line tile raster scan to completely traverse a current tile before proceeding to a next tile.

[0163] Example embodiment 7, the video data stream according to example embodiment 1, wherein the second information (142) signals a replacement value for a slice address for each of at least a subset of slice portions in which any slice within the spatial segment is encoded.

[0164] Example embodiment 8, the video data stream according to example embodiment 7, wherein the slice address of each slice portion in the sequence of slice portions of the video data stream relates to the picture region of the video, and the replacement value for the slice address relates to the reduced picture region.

[0165] Example embodiment 9, the video data stream according to example embodiment 6 or 7, wherein the slice address of each slice portion in the sequence of slice portions of the video data stream one-dimensionally indexes a coded start position of a slice encoded into the respective slice portion along a first coding scan order traversing the picture region, and wherein the picture has been encoded in the sequence of slice portions of the video data stream along the first coding scan order, and the replacement value for the slice address one-dimensionally indexes the coded start position along a second coding scan order traversing the reduced picture region.

[0166] Example embodiment 10, the video data stream according to any one of example embodiments 7 to 9, wherein the subset of slice portions does not include a first one of the slice portions of the sequence of slice portions in which any slice within the spatial segment (110) is encoded, and wherein the second information (142) provides the replacement value for each of the subset of slice portions in turn in the order of occurrence of the slice portions in the video data stream.

[0167] Example embodiment 11, the video data stream according to any one of example embodiments 7 to 9, wherein the first information (140) defines a spatial segment (110) within the picture region as consisting of a set of at least two segment regions (110a, 110b) in each of which the video is encoded in the video data stream independently from outside the respective segment region, wherein none of the plurality of slices crosses a boundary of any one of the first and the second segment region.

[0168] Example embodiment 12, the video data stream according to example embodiment 11, wherein the subset of slice portions comprises all slice portions of the sequence of slice portions in which any slice within the spatial segment is encoded, and wherein the second information provides a replacement value for each slice portion in the order in which the slice portion occurs in the video data stream.

[0169] Example embodiment 13, the video data stream according to example embodiment 12, wherein,

[0170] the second information (142) further signals how slice portions of slices within the spatial segment are reordered relative to their ordering in the sequence of slice portions of the video data stream when extracting the reduced video data stream from the video data stream, and

[0171] the second information (142) provides a replacement value for each of the subsets of slice portions in the order in which the reordered slice portions occur in the reduced video data stream.

[0172] Example embodiment 14, the video data stream according to any one of example embodiments 1 to 13, wherein,

[0173] the sequence of slice portions comprised by the video data stream comprises slice portions in which slices located within the spatial segment are encoded and slice portions in which slices located outside the spatial segment are encoded, or

[0174] the sequence of slice portions comprised by the video data stream comprises slice portions in which slices located within the spatial segment are encoded but does not comprise slice portions in which slices located outside the spatial segment are encoded.

[0175] Example embodiment 15, the video data stream according to any one of example embodiments 1 to 14, wherein the second information (142) defines a shape and a size of the reduced picture area.

[0176] Example embodiment 16, the video data stream according to any one of example embodiments 1 to 15, wherein the first information (140) defines a shape and a size of the reduced picture area by means of a size parameter in parameter set replacement, which parameter set replacement shall be used to replace a respective parameter set in the video data stream when extracting the reduced video data stream (136) from the video data stream.

[0177] Example embodiment 17, an apparatus for generating a video data stream, configured to:

[0178] providing a sequence of slice parts to the video data stream, each slice part encoding a respective slice of a plurality of slices of a picture of the video, wherein each slice part comprises a slice address indicating a position of the slice encoded in the respective slice part within a picture area of the video;

[0179] providing extraction information to the video data stream, the extraction information indicating how to extract a reduced video data stream (136) from the video data stream, wherein the reduced video data stream (136) encodes a spatially smaller video corresponding to a spatial segment of the video, by restricting the video data stream to slice parts encoding any slice within the spatial segment and modifying the slice addresses to relate to a reduced picture area of the spatially smaller video, wherein the extraction information comprises:

[0180] first information defining the spatial segment within the picture area, within which the video is encoded in the video data stream independently from outside the spatial segment, wherein none of the plurality of slices crosses a boundary of the spatial segment; and

[0181] second information signaling one of a plurality of options, or explicitly signaling, how to modify the slice address of the slice part of each slice within the spatial segment in order to indicate within the reduced video data stream a position of the respective slice within the reduced picture area.

[0182] Example Embodiment 18, an apparatus for extracting a reduced video data stream from a video data stream in which a video is encoded, the reduced video data stream encoding a spatially smaller video, the video data stream comprising a sequence of slice parts, each slice part encoding a respective slice of a plurality of slices of a picture of the video, wherein each slice part comprises a slice address indicating a position of the slice encoded in the respective slice part within a picture area of the video, wherein the apparatus is configured to:

[0183] reading extraction information from the video data stream,

[0184] deriving a spatial segment within the picture area from the extraction information, wherein none of the plurality of slices crosses a boundary of the spatial segment, and wherein the reduced video data stream is restricted to slice parts encoding any slice within the spatial segment,

[0185] modifying the slice address of the slice part of each slice within the spatial segment using one of a plurality of options determined from the plurality of options by explicit signaling of the extraction information,

[0186] to indicate within the down-scaled video data stream a position at which a respective slice is located in a down-scaled picture area of the spatially smaller video.

[0187] Example embodiment 19, the apparatus according to example embodiment 17 or 18, wherein the video data stream is a video data stream according to any one of example embodiments 2 to 16.

[0188] Example embodiment 20, a video data stream in which a video (332) is encoded, wherein the video data stream comprises a signal (352) indicating that a picture (334) of the video displays a same scene content (304) at different spatial portions (354, 356) of the picture at different resolutions.

[0189] Example embodiment 21, the video data stream according to example embodiment 20, wherein the signal (352) comprises a flag switchable between a first state and a second state, wherein the first state indicates that a picture (334) of the video comprises different spatial portions at which the picture displays a same scene content at different resolutions, and wherein the second state indicates that the picture of the video displays each portion of a scene at only one resolution, wherein the flag is set to the first state.

[0190] Example embodiment 22, the video data stream according to example embodiment 21, wherein the flag is contained in a sequence parameter set or a video parameter set of the video data stream.

[0191] Example embodiment 23, the video data stream according to any one of example embodiments 20 to 22, wherein the video data stream is encoded independently of outside a respective spatial region (360) for each spatial region of a set of spatial regions (360) that do not overlap with each other.

[0192] Example embodiment 24, the video data stream according to any one of example embodiments 20 to 23, wherein the signal (352) comprises information on a count of different resolutions at which a picture of the video displays a same scene content at different spatial portions.

[0193] Example embodiment 25, the video data stream according to any one of example embodiments 20 to 24, wherein the signal (352) comprises information on a spatial subdivision of a picture of the video into segments (354, 356) each displaying a same scene content and at least two of the segments displaying the same scene content at different resolutions.

[0194] Example embodiment 26, the video data stream according to any one of the example embodiments 20 to 25, wherein the signal (352) comprises different information on resolutions at which pictures of the video display identical scene content at the different spatial portions (354, 356).

[0195] Example embodiment 27, the video data stream according to any one of the example embodiments 20 to 26, wherein the video data stream comprises extraction information indicating a set of extraction regions for which the video data stream allows extraction of independent video data streams from the video data without re-encoding, the independent video data streams representing videos confined within picture content of the video by respective extraction regions.

[0196] Example embodiment 28, the video data stream according to example embodiment 27, wherein the video data stream is encoded for each of a set of spatial regions (360) that do not overlap with each other in a manner independent of outside the respective spatial region (360), and the extraction information indicates each extraction region of the set of extraction regions as a set of one or more of the spatial regions (360).

[0197] Example embodiment 29, the video data stream according to example embodiment 27 or 28, wherein pictures of the video are spatially subdivided into an array of tiles and are sequentially encoded in the data stream, wherein the spatial regions (360) are clusters of one or more tiles.

[0198] Example embodiment 30, the video data stream according to any one of the example embodiments 27 to 29, wherein the video data stream comprises a further signal indicating for each of at least one or more extraction regions:

[0199] whether pictures of the video display identical scene content within the respective extraction region at different resolutions.

[0200] Example embodiment 31, the video data stream according to example embodiment 30, wherein the further signal indicates for each of at least one or more extraction regions further:

[0201] a view direction of a segment region of the respective extraction region at which identical scene content is displayed within the respective extraction region at the highest resolution, and / or

[0202] a share of area of a segment region of the respective extraction region at which identical scene content is displayed within the respective extraction region at the highest resolution in a total region of the respective extraction region, and / or

[0203] The respective extraction region is spatially subdivided from the entire area of the respective extraction region into segment areas that each display the same scene content at mutually different resolutions.

[0204] Example embodiment 32, the video data stream according to example embodiment 30, wherein the further signal is contained in a SEI message separate from the one or more SEI messages containing the extraction information.

[0205] Example embodiment 33, an apparatus for processing a video data stream according to any one of example embodiments 20 to 32, wherein the apparatus supports a predetermined processing task and is configured to inspect the signal to decide whether or not to perform the predetermined processing task on the video data stream.

[0206] Example embodiment 34, the apparatus according to example embodiment 33, wherein the processing task comprises:

[0207] decoding and / or

[0208] extracting a reduced video data stream from the video data stream without re-encoding.

[0209] Example embodiment 35, an apparatus for generating a video data stream according to any one of example embodiments 20 to 32.

[0210] Example embodiment 36, a method for generating a video data stream, the method comprising:

[0211] providing the video data stream with a sequence of slice parts, each slice part having encoded therein a respective slice of a plurality of slices of a picture of a video, wherein each slice part comprises a slice address indicating a position at which the slice encoded in the respective slice part is located in a picture area of the video;

[0212] providing the video data stream with extraction information indicating how to extract a reduced video data stream from the video data stream, wherein the reduced video data stream has encoded therein a spatially smaller video corresponding to a spatial segment of the video, by restricting the video data stream to slice parts having encoded therein any slices within the spatial segment and modifying the slice addresses to relate to a reduced picture area of the spatially smaller video, wherein the extraction information comprises:

[0213] first information defining a spatial segment within a picture area, within which the video is encoded in the video data stream independently from outside the spatial segment, wherein no slice of the plurality of slices crosses a boundary of the spatial segment; and

[0214] a second information signaling one of a plurality of options or explicitly signaling the following: how to modify the slice address of the slice portion of each slice within the spatial segment in order to indicate within the reduced video data stream the position of the respective slice within the reduced picture area of the spatially reduced video.

[0215] Example embodiment 37, a method for extracting a reduced video data stream from a video data stream in which a video is encoded, the reduced video data stream encoding a spatially reduced video, the video data stream comprising a sequence of slice portions, each slice portion encoding a respective slice of a plurality of slices of a picture of the video, wherein each slice portion comprises a slice address indicating the position of the slice encoded in the respective slice portion within a picture area of the video, wherein the method comprises:

[0216] reading extraction information from the video data stream,

[0217] deriving a spatial segment within the picture area from the extraction information, wherein none of the plurality of slices crosses a boundary of the spatial segment, and wherein the reduced video data stream is restricted to slice portions encoding any slice within the spatial segment,

[0218] modifying the slice address of the slice portion of each slice within the spatial segment using one of a plurality of options determined from a plurality of options by explicit signaling of the extraction information,

[0219] in order to indicate within the reduced video data stream the position of the respective slice within the reduced picture area of the spatially reduced video.

[0220] Example embodiment 38, a method for processing a video data stream according to any one of example embodiments 20 to 32, wherein the processing comprises a predetermined processing task, and the method comprises inspecting the signal to decide whether or not to perform the predetermined processing task on the video data stream.

[0221] Example embodiment 39, a method for generating a video data stream according to any one of example embodiments 20 to 32.

[0222] Example embodiment 40, a computer readable medium storing a computer program having a program code for performing the method according to example embodiments 36, 37, 38 or 39, when the computer program is run on a computer.

[0223] List of references

[0224] [1] ITU-T, Recommendation H.265 (04 / 13), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video, High efficiency video coding, Online: http: / / www.itu.int / rec / T-REC-H.265-201304-I;

[0225] [2] Jill Boyce et al., JCTVC-Z1005, HEVC Additional Supplemental Enhancement Information (Draft 1), Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 26th Meeting: Geneva, CH, 12-20 January 2017.

Claims

1. A method of storing a video, comprising: storing a video data stream on a digital storage medium, wherein the video data stream comprises a sequence of slice portions, each slice portion encoding a respective slice of a plurality of slices of a picture of the video, wherein each slice portion comprises a slice address indicating a position at which the slice encoded in the respective slice portion is located in a picture area of the video; wherein the video data stream further comprises extraction information for extracting a reduced video data stream from the video data stream, wherein the reduced video data stream encodes a spatially reduced video corresponding to a spatial segment of the video of the video data stream, the extraction involving restricting the video data stream to slice portions encoding any slices within the spatial segment and modifying the slice addresses to relate to a reduced picture area of the spatially reduced video, wherein the extraction information comprises: first information defining the spatial segment within the picture area; and second information explicitly signaling replacement slice addresses for replacing the slice addresses of the slice portions of the slices within the spatial segment, the replacement slice addresses indicating the positions at which the slices are located in the reduced picture area within the reduced video data stream, wherein the first information defines the spatial segment within the picture area as consisting of a plurality of segment areas, in each of which the video is encoded in the video data stream independently from outside the respective segment area by: a syntax element indicating a number of segment areas, and an identifier of each segment area for identifying the respective segment area, wherein none of the plurality of slices crosses a boundary of any of the plurality of segment areas.

2. The method of claim 1, wherein, the slice addresses of each slice portion of the sequence of slice portions of the video data stream are one-dimensionally indexed along a first coding scan order traversing the picture area, and the replacement values for replacing the slice addresses of the respective slice portions are one-dimensionally indexed along a second coding scan order traversing the reduced picture area, to the coding start positions of the slices encoded into the respective slice portions.

3. The method of claim 1, wherein, the second information provides the replacement slice addresses for the slice portions in a sequence in which the slice portions occur in the video data stream.

4. The method of claim 1, wherein, the sequence of slice portions comprised by the video data stream comprises slice portions encoding slices located within the spatial segment and slice portions encoding slices located outside the spatial segment.

5. The method of claim 1, wherein, the extraction information further defines a shape and a size of the reduced picture area.

6. The method of claim 5, wherein, the extraction information defines the shape and the size of the reduced picture area by a size parameter in a parameter set replacement value, which shall be used to replace a corresponding parameter set in the video data stream when extracting the reduced video data stream from the video data stream.

7. An apparatus for generating a video data stream, comprising: an apparatus for providing a sequence of slice parts to the video data stream, each slice part encoding a respective slice of a plurality of slices of a picture of the video, wherein each slice part comprises a slice address indicating a position of the slice encoded in the respective slice part within a picture area of the video; an apparatus for providing extraction information to the video data stream, the extraction information being used for extracting a reduced video data stream from the video data stream, wherein the reduced video data stream encodes a spatially reduced video corresponding to a spatial segment of the video of the video data stream, the extraction involving restricting the video data stream to slice parts encoding any slice within the spatial segment and modifying the slice addresses to relate to a reduced picture area of the spatially reduced video, wherein the extraction information comprises: first information defining the spatial segment within the picture area, within which the video is encoded in the video data stream independently from outside the spatial segment, and second information explicitly signaling replacement slice addresses for replacing the slice addresses of the slice parts of the slices within the spatial segment, the replacement slice addresses indicating a position of the slices within the reduced picture area within the reduced video data stream, wherein the first information defines the spatial segment within the picture area as consisting of a plurality of segment areas, within each segment area the video is encoded in the video data stream independently from outside the respective segment area by: a syntax element indicating a number of segment areas, and an identifier of each segment area for identifying the respective segment area, wherein none of the plurality of slices crosses a boundary of any of the plurality of segment areas.

8. An apparatus for extracting a reduced video data stream from a video data stream in which a video is encoded, the reduced video data stream encoding a spatially reduced video, the video data stream comprising a sequence of slice parts, each slice part encoding a respective slice of a plurality of slices of a picture of the video, wherein each slice part comprises a slice address indicating a position of the slice encoded in the respective slice part within a picture area of the video, wherein the apparatus comprises: an apparatus for reading extraction information from the video data stream, an apparatus for deriving a spatial segment within the picture area from first information in the extraction information, wherein the reduced video data stream is restricted to slice parts encoding any slice within the spatial segment, an apparatus for deriving replacement slice addresses from second information in the extraction information, and an apparatus for replacing the slice addresses of the slice parts of the slices within the spatial segment by the replacement slice addresses. means for replacing, using the replacement slice address explicitly signaled using the extraction information, the slice address of a slice portion of a slice within the spatial segment, the replacement slice address indicating within the downscaled video data stream a position of the slice within a downscaled picture area of the spatially smaller video, such that the slice portion of a slice within the spatial segment is reordered relative to an ordering of the slice portion in a sequence of slice portions of the video data stream, wherein the first information defines the spatial segment within the picture area as consisting of a plurality of segment areas, in each of which the video is encoded in the video data stream independently from outside the respective segment area by: a syntax element indicating a number of segment areas, and an identifier of each segment area for identifying the respective segment area, wherein none of the plurality of slices crosses a boundary of any of the plurality of segment areas.

9. A method for generating a video data stream, comprising: providing a sequence of slice portions to the video data stream, each slice portion encoding a respective slice of a plurality of slices of a picture of a video, wherein each slice portion comprises a slice address indicating a position of the slice encoded in the respective slice portion within a picture area of the video; providing extraction information to the video data stream, the extraction information being used for extracting a downscaled video data stream from the video data stream, wherein the downscaled video data stream encodes a spatially smaller video corresponding to a spatial segment of the video of the video data stream, the extraction involving restricting the video data stream to slice portions encoding any slices within the spatial segment and modifying the slice addresses to relate to a downscaled picture area of the spatially smaller video, wherein the extraction information comprises: first information defining the spatial segment within the picture area; second information explicitly signaling replacement slice addresses for replacing slice addresses of slice portions of slices within the spatial segment, the replacement slice addresses indicating within the downscaled video data stream a position of the slice within the downscaled picture area, wherein the first information defines the spatial segment within the picture area as consisting of a plurality of segment areas, in each of which the video is encoded in the video data stream independently from outside the respective segment area by: a syntax element indicating a number of segment areas, and an identifier of each segment area for identifying the respective segment area, wherein none of the plurality of slices crosses a boundary of any of the plurality of segment areas.

10. A method for extracting a reduced video data stream from a video data stream in which a video is encoded, the reduced video data stream having a spatially reduced video encoded therein, the video data stream comprising a sequence of slice portions, each slice portion having a respective slice of a plurality of slices of a picture of the video encoded therein, wherein each slice portion comprises a slice address, the slice address indicating a position at which the slice encoded in the respective slice portion is located in a picture area of the video, wherein the method comprises: reading extraction information from the video data stream, deriving a spatial segment within the picture area from first information within the extraction information, wherein the reduced video data stream is restricted to slice portions in which any slices within the spatial segment are encoded, deriving a replacement slice address from second information within the extraction information, and replacing the slice address of the slice portion of a slice within the spatial segment with the replacement slice address explicitly signaled using the extraction information, the replacement slice address indicating a position at which the slice is located in a reduced picture area of the spatially reduced video within the reduced video data stream, such that the slice portion of the slice within the spatial segment is reordered relative to an ordering of the slice portion in the sequence of slice portions of the video data stream, wherein the first information defines the spatial segment within the picture area as consisting of a plurality of segment areas, in each of which the video is encoded in the video data stream independently from outside the respective segment area by: a syntax element indicating a number of segment areas, and an identifier of each segment area for identifying the respective segment area, wherein none of the plurality of slices crosses a boundary of any of the plurality of segment areas.

11. A non-transitory digital storage medium having stored thereon a computer program which, when executed by a computer, performs the method for generating a video data stream according to claim 9.

12. A non-transitory digital storage medium having stored thereon a computer program which, when executed by a computer, performs the method for extracting a reduced video data stream from a video data stream in which a video is encoded according to claim 10.

Citation Information

Patent Citations

  • Padding of segments in coded slice NAL units

    CN103959781A

  • Image encoding method, image decoding method, image encoding device, image decoding device, and image encoding / decoding device

    CN104737541A