Advanced video data stream extraction and multi-resolution video transmission
By incorporating additional signaling to modify slice addresses and encoding parameters, the video data stream extraction process becomes more efficient, addressing inefficiencies in handling diverse video content types and enabling effective extraction and decoding of spatially smaller video streams.
Patent Information
- Application Number
- JP2025103557
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-03-20
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-15
AI Technical Summary
The existing video data stream extraction process is inefficient, particularly when handling different recipients with varying video content types and requires complex adjustments to slice addresses during the extraction process.
The process is enhanced by providing additional signaling in the video data stream to guide the extraction process, specifically by modifying slice addresses and encoding parameters, allowing for efficient handling of different video content types and enabling the extraction of spatially smaller video streams without re-encoding.
This approach reduces processing overhead and complexity, enabling efficient extraction and decoding of video streams with varying resolutions, suitable for multiple recipients and different video content types.
Smart Images

Figure 2025157235000001_ABST
Abstract
Description
[Technical Field]
[0001] The present application relates to the concept of extracting a video data stream, i.e. extracting a reduced video data stream from a suitably prepared video data stream, such that the reduced video data stream has encoded therein a spatially smaller video that corresponds to a spatial section of the video encoded in the original video data stream, and more particularly to the transmission of different video versions of a scene, the versions differing in resolution or fidelity of the scene. [Background technology]
[0002] The HEVC standard [1] defines a hybrid video codec that allows the definition of a sub-array of rectangular tiles of an image, such that a smaller or scaled-down video data stream can be easily extracted from the entire video data stream, i.e., without requantization or the need to redo motion compensation, provided the video codec adheres to some coding constraints. It is envisioned that additions to the HEVC standard syntax, as outlined in [2], will guide the extraction process for the receiver of the video data stream.
[0003] However, the extraction process needs to be made more efficient.
[0004] An application area in which video data extraction may be applied concerns the transmission or provision of multiple versions of a video scene, where the resolution of the scene differs from one another. It would be advantageous to have an efficient way of installing such transmission or providing different resolution versions. Summary of the Invention [Problem to be solved by the invention]
[0005] It is therefore a first object of the present invention to provide a video data stream extraction concept that is more efficient, i.e. allows for more efficient handling of video content of unknown types for different recipients, e.g. from a viewport to a screen, etc., or to reduce the complexity of the extraction process. This object is achieved by the subject matter of the independent claims of the present application according to a first aspect. [Means for solving the problem]
[0006] In particular, according to a first aspect of the present application, the extraction of a video data stream becomes more efficient by providing information in the video data stream that signals one of several options for modifying the slice addresses of each slice in the extractable spatial section to indicate where the respective slice is located in the reduced (extracted) image area within the reduced video data stream. In other words, the second information provides information to the extraction site of the video data stream, guiding the extraction process regarding the composition of the spatially smaller video images in the reduced (extracted) video data stream based on the spatial section of the original video, thereby mitigating the extraction process or enabling it to adapt to greater variations in scene types conveyed within the video data stream. Regarding the latter issue, for example, the second information can address various opportunities to take advantage of the results of potentially discontinuous portions of the spatial section jostling each other while maintaining the relative positioning of these portions of the spatial section in the original video, or their relative order in terms of encoding order. For example, in a spatial section consisting of zones that abut different parts along the periphery of an original image showing a scene at the seam interface of a projection from a panoramic scene to a screen, the arrangement of the zones of the spatial section in the smaller image of the extracted stream will be different from that of a non-panoramic image type, which the receiver may not even know. Correcting the slice addresses of the extracted slice parts, either additionally or individually, is a tedious task that may be alleviated by explicitly sending information on how to correct them, for example in the form of alternative slice addresses.
[0007] Another object of the present invention is to provide a concept for more efficiently presenting different versions of a video scene, with different resolution versions of the scene, side-by-side to a receiver.
[0008] This object is achieved by the subject matter of the pending independent claims of the second aspect of the present application.
[0009] In particular, according to a second aspect of the present application, the juxtaposition of several versions of a video scene with different scene resolutions is more efficiently displayed by combining these versions into a single video encoded in a single video data stream, and providing this video data stream with signaling indicating that the images of the video display common scene content in different spatial portions of the image at different resolutions. A recipient of the video data stream can therefore recognize, based on the signaling, whether the video content conveyed by the video data stream relates to a spatially parallel collection of several versions of scene content at different scene resolutions. Depending on the capabilities of the receiving site, attempts to decode the video data stream may be suppressed, or processing of the video data stream may be adapted to an analysis of the signaling.
[0010] Advantageous implementations of the embodiments of the above-mentioned aspects are the subject of the dependent claims.Preferred embodiments of the present application are described below with reference to the drawings. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a schematic diagram illustrating MCTS extraction using adjusted slice addresses. [Figure 2] FIG. 2 is a combined schematic and block diagram illustrating the concept of a video data stream extraction process according to an embodiment of the first aspect of the present application and examples of participating processes and devices. [Figure 3] FIG. 3 is a diagram illustrating an example syntax that inherits an example of second information of extraction information, followed by an example that explicitly shows how the second information modifies a slice address. [Figure 4] FIG. 4 is a schematic diagram illustrating an example of non-adjacent MCTSs forming a desired image subsection. [Figure 5]FIG. 5 illustrates a particular syntax example including second information according to an embodiment in which the second information indicates a particular option for modifying a slice address in an extraction process among several possible options. [Figure 6] FIG. 6 is a schematic diagram showing an example of packaging a multi-resolution 360° frame. [Figure 7] FIG. 7 is a schematic diagram illustrating the extraction of an exemplary MCTS including a mixed resolution representation. [Figure 8] FIG. 8 is a mixed schematic and block diagram illustrating an efficient method for providing multi-resolution scenes to users and participating devices and video streams and processes according to an embodiment of the second aspect of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0012] The following description begins with a description of the first aspect of the present application, followed by a description of the second aspect of the present application. More precisely, for the first aspect of the present application, the description begins with an overview of the underlying technical problem in order to motivate the advantages and underlying concepts of the embodiments of the first aspect described below. For the second aspect, the order of description is chosen in the same way.
[0013] In panoramic or 360 video applications, it is common for only a subsection of the screen to need to be presented to the user. Specific codec tools, such as Motion Constrained Tile Sets (MCTS), allow for the extraction of coded data corresponding to the desired image subsection in the compressed domain, forming a conforming bitstream that can be decoded by legacy decoder devices that do not support MCTS decoding from the full image bitstream, and are characterized as being lower tier compared to the decoder required for full image decoding.
[0014] As an example and for reference, the signaling included in the HEVC codec can be found here: Reference [1] specifies a temporary MCTS SEI message in sections D2.29 and E.2.29, which allows an encoder to signal a list of designated rectangles, defined by their top-left and bottom-right tile indices, that belong to the MCTS. Reference [2], which will be added to the next version of [1], provides additional information such as parameter sets and nested SEI messages that will facilitate the extraction of MCTS as a conforming HEVC bitstream.
[0015] As can be seen from [1] and [2], the extraction procedure involves adjusting the slice address indicated in the slice header of the relevant slice, which is performed on the extraction device.
[0016] FIG. 1 illustrates an example of MCTS extraction. FIG. 1 illustrates a video data stream, i.e., an image encoded into an HEVC video data stream. The image 100 is subdivided into CTBs, i.e., coding tree blocks, into which the image 100 is encoded. In the example of FIG. 1, the image 100 is subdivided into 16×6 CTBs, although the number of CTB rows and the number of CTB columns is of course not important. Reference numeral 102 representatively denotes such a CTB. In units of these CTBs 102, the image 100 is further subdivided into tiles, i.e., an array of m×n tiles, and FIG. 1 illustrates an exemplary case where m=8 and n=4. Reference numeral 104 is used in each tile to representatively denote one of such tiles. Each tile is thus a rectangular cluster or subarray of CTBs 102. For illustrative purposes only, FIG. 1 shows that tiles 104 may be of different sizes, or alternatively, rows of tiles may be columns of tiles of different heights and different widths, respectively.
[0017] As is known in the art, the tile subdivision, i.e., the subdivision of the image 100 into tiles 104, affects the encoding order 106 by which the image content of the image 100 is encoded into a video data stream. In particular, the tiles 104 are traversed one after another along the tile order, i.e., in a tile-row-by-tile raster scan order. In other words, all CTBs 102 within one tile 104 are first coded or traversed according to the coding order 106 before the coding order proceeds to the next tile 104. Within each tile 102, the CTBs are also coded using a raster scan order, i.e., a row-wise raster scan order. Along the coding order 106, the coding of the image 100 into a video data stream is subdivided to yield so-called slice portions. In other words, slices of the image 100 traversed by successive portions of the coding order 106 are coded into the video data stream as a unit to form a slice portion. 1, each tile is assumed to be within a single slice, or in HEVC terminology, a slice segment, but this is merely an example and can be created in different ways. Typically, one slice 108, or in HEVC terms, one slice segment, is typically designated by reference numeral 108 in FIG. 1, and matches or fits into a corresponding tile 104.
[0018] As far as encoding of the image 100 into a video data stream is concerned, it should be noted that this encoding utilizes spatial prediction, temporal prediction, context derivation of entropy coding, motion compensation of the temporal prediction, transformation of the prediction residual, and / or quantization of the prediction residual. The coding order 106 not only affects slicing but also defines the availability of reference bases for spatial prediction and / or context derivation. Only those neighboring portions preceding it in the coding order 106 are available. Tiling not only affects the coding order 106 but also limits coding interdependencies within the image 100. For example, spatial prediction and / or context derivation are limited to referencing only portions within the current tile 104. Portions outside the current tile are not referenced in spatial prediction and / or context derivation.
[0019] FIG. 1 illustrates a further specific region of image 100, the so-called MCTS, i.e., a spatial section 110 within the image region of image 100 from which video of which image 100 is a part can be extracted. A close-up of section 110 is shown on the right side of FIG. 1. The MCTS 110 of FIG. 1 is composed of a set of tiles 104. The tiles in section 110 are labeled a, b, c, and d, respectively. The fact that spatial section 110 is extractable imposes further constraints on the encoding of the video into a video data stream. In particular, the image division of the video of image 100 shown in FIG. 1 is adopted by other images in the video, and for this image sequence, the image content within tiles a, b, c, and d is coded such that coding interdependencies remain within spatial section 110 even when referencing one image to another. In other words, for example, temporal prediction and temporal context derivation are constrained to fit within the region of spatial section 110.
[0020] An interesting aspect of encoding video that is part of image 100 into a video data stream is the fact that slices 108 are provided with slice addresses that indicate their encoding start, i.e., their location in the encoded image area. Slice addresses are assigned along the encoding order 106. For example, the slice address indicates the CTB rank along the encoding order 106 at which the encoding of each slice begins. For example, within the data streams encoding video and image 100, respectively, the slice portion carrying a slice corresponding to tile a has slice address 7 as the seventh CTB of the encoding order 106, indicating the first CTB of the encoding order 106 within tile a. Similarly, the slice addresses within the slice portions carrying slices associated with tiles b, c, and d will be 9, 29, and 33, respectively.
[0021] The right side of Figure 1 shows slice addresses assigned by the receiver of the reduced or extracted video data stream to two slices corresponding to tiles a, b, c, and d. In other words, Figure 1 uses the numbers 0, 2, 4, and 8 on the right side to indicate slice addresses assigned by the receiver of the reduced or extracted video data stream, which are obtained from the original video data stream representing the video including the entire image 100 by extraction with respect to spatial section 110, i.e., by MCTS extraction. In the reduced or extracted video data stream, the slice portions encoded in slice 100 for tiles a, b, c, and d are arranged in coding order 106 exactly as they were in the original video data stream retrieved during the extraction process. In particular, the receiver arranges the image content in coding order in the form of a CTB reconstructed from the sequence of slice portions in the reduced or extracted video data stream, i.e., the slices for tiles a, b, c, and d. The coding order 112 traverses spatial section 110 in the same manner as coding order 106 for the entire image 100. That is, spatial section 110 is traversed tile by tile in raster scan order, and the CTB within each tile is also traversed in raster scan order before proceeding to the next tile. The relative positions of tiles a, b, c, and d are maintained. That is, spatial section 110 as shown on the right side of FIG. 1 maintains the relative positions among tiles a, b, c, and d that occurred in image 100. The slice addresses of the slices resulting from determining tiles a, b, c, and d using encoding order 112 are 0, 2, 4, and 8, respectively. Thus, a receiver can reconstruct a smaller video based on the reduced or extracted video data stream illustrating spatial section 110, shown as a self-contained image on the right side.
[0022] To summarize the discussion of FIG. 1 so far, FIG. 1 illustrates or describes the adjustment of slice addresses and CTB units after extraction by using the numbers in the upper left corner of each slice 108. To perform the extraction, the extraction site or extraction device must analyze the original video data stream for parameters indicating the size of the CTB, i.e., maximum CTB size 114, and the number and size of tile columns and tile rows in the image 100. Additionally, nested MCTS-specific sequence and image parameter sets are examined to derive the output arrangement of tiles therefrom. In FIG. 1, tiles a, b, c, and d in spatial section 110 retain their relative arrangement. In summary, the analysis and examination performed by the extraction device in accordance with the above and the MCTS instructions of the HEVC data stream as a whole requires dedicated and sophisticated logic to derive slice addresses of the reconstructed slices 108 from the parameters listed above. This dedicated and sophisticated logic incurs additional implementation costs as well as runtime drawbacks.
[0023] The embodiments described below can therefore reduce the overall processing burden of the extraction device described above by using additional signaling in the video data stream and corresponding processing steps at the extraction information generation side and extraction side to provide readily available information for the specific purpose of extraction. Additionally or alternatively, some embodiments described below use additional signaling to guide the extraction process in such a way that more effective processing of different types of video content is achieved.
[0024] The general concept is first described based on FIG. 2 . Later, this general description of the operation modes of the individual entities participating in the overall process depicted in FIG. 2 will be further clarified in different ways according to different embodiments below. It should be noted that while the entities and blocks depicted therein are collectively described in one diagram for ease of understanding, each relates to a self-contained device that individually inherits the functionality that provides the overview of the benefits of FIG. 2 as a whole. More precisely, FIG. 2 illustrates a generation process that generates a video data stream and provides extraction information for such video data stream. The extraction process itself is followed by decoding of the extracted or reduced video data stream and the participating devices. The operation modes of these devices or the performance of the individual tasks and steps are in accordance with the embodiments described herein. According to the specific implementation example initially described with respect to FIG. 2 and as further outlined below, the processing overhead associated with the extraction task of the extraction device is reduced. According to further embodiments, the processing of various different types of video content in the original video is additionally or alternatively reduced.
[0025] At the top of Figure 2, an original video is shown, designated by reference numeral 120. This video 120 is composed of a series of images, one of which is designated using reference numeral 100, since it serves the same role as image 100 in Figure 1, i.e., it is an image that indicates an image region from which spatial section 110 will subsequently be cropped by extraction of the video data stream. However, it should be understood that the tile subdivision described above with respect to Figure 1 need not be used in the video coding underlying the process shown in Figure 2; rather, tiles and CTBs indicate semantic entities in video coding that are merely optional as far as the embodiment of Figure 2 is concerned.
[0026] FIG. 2 illustrates that video 102 is the subject of video coding by video coding core 122. The video coding performed by video coding core 122 transforms video 120 into video data stream 124, for example, using hybrid video coding. That is, video coding core 122 uses block-based predictive coding, for example, to encode individual image blocks of images of video 120 using one of several supported prediction mode encodings of the prediction residual. Prediction modes may include, for example, spatial and temporal prediction. Temporal prediction may include motion compensation, i.e., determining a motion field representing the motion vector of the temporally predicted block and transmitting it via data stream 124. The prediction residual may be transform coded. That is, several spectral decompositions may be applied to the prediction residual, and the resulting spectral coefficients may be subjected to quantization and losslessly coded into data stream 124, for example, using entropy coding. The entropy coding may then use context adaptivity, i.e., determine a context whose derivation depends on spatial and / or temporal neighborhoods. As mentioned above, the encoding may be based on an encoding order 106 that restricts the encoding dependencies to the extent that only portions of the video that the encoding order 106 has already traversed can be used as a basis or reference for encoding the current portion of the video 120. The encoding order 106 traverses the video 120 image by image, but not necessarily in the presentation time order of the images. Within an image, such as image 100, the video encoding core 120 subdivides the encoded data obtained by video encoding, thereby dividing the image 100 into slices 108 that each correspond to a corresponding slice portion 126 in the video data stream 124. Within the data stream 124, the slice portions 126 form a sequence of slice portions that follow each other in the order in which the corresponding slices 108 are traversed by the encoding order 106 of the image 100.
[0027] As also shown in Figure 2, video encoding core 122 provides or encodes a slice address for each slice portion 126. Slice addresses are shown in uppercase in Figure 2 for illustrative purposes. As described with respect to Figure 1, slice addresses may be determined in any suitable units, such as units of CTBs, one-dimensionally along encoding order 106, or alternatively, they may be determined differently for some predetermined point within the image area used by an image of video 120, such as the upper left corner of image 100.
[0028] Thus, the video encoding core 122 receives the video 120 and outputs a video data stream 124 .
[0029] As already outlined above, the video data stream generated according to FIG. 2 can be extracted as far as the spatial section 110 is concerned, and the video encoding core 120 adapts the video encoding process accordingly. To this end, the video encoding core 122 restricts inter-coding dependencies so that parts within the spatial section 110 are coded in a certain way in the video data stream 124, so as not to depend on parts outside the section 110, for example, by spatial prediction, temporal prediction, or context derivation. Slices 108 do not cross the boundaries of the section 110. Thus, each slice 108 is either entirely within the section 110 or entirely outside the section 110. It should be noted that the video encoding core 122 can follow several spatial sections, not just one spatial section 110. These spatial sections may intersect with each other, i.e., the same spatial sections may partially overlap, or one spatial section may be entirely within another spatial section. By these means, as will be explained in more detail later, it is possible to extract from the video data stream 124 a downscaled or extracted video data stream of images smaller than the images of the video 120, i.e., images that show the content in the spatial section 110 simply without re-encoding, i.e., without having to perform complex tasks such as motion compensation, quantization and / or entropy coding again.
[0030] The video data stream 124 is received by a video data stream generator 128. In particular, according to the embodiment shown in Figure 2, the video data stream generator 128 includes a receiving interface 130 that receives the prepared video data stream 124 from the video encoding core 122. It should be noted that, according to an alternative example, the video encoding core 122 can be included in the video data stream generator 128, thereby replacing the interface 130.
[0031] The video data stream generator 128 provides extraction information 132 to the video data stream 124. In FIG. 2, the resulting video data stream output by the video data stream generator 128 is indicated using the reference numeral 124′. The extraction information 132 indicates to an extractor, indicated in FIG. 2 by the reference numeral 134, how to extract from the video data stream 124′ an encoded reduced or extracted video data stream 136, which contains a spatially smaller video 138 corresponding to the spatial section 110. The extraction information 132 includes first information 140 defining the spatial section 110 within the image area transmitted by the image 100, and second information 142 indicating one of several options on how to modify the slice address of the slice portion 126 of each slice 108 falling within the spatial section 110 to indicate where, within the reduced video data stream 136, the respective slice is located within the reduced image area of the image 144 of the video 138.
[0032] In other words, the video data stream generator 128 simply adds to the video data stream 124 to produce the video data stream 124', i.e., the extracted information 132. This extracted information 132 is intended to guide the extractor 134 receiving the video data stream 124' in extracting a reduced or extracted video data stream 136, particularly with respect to the section 110, from the video data stream 124'. The first information 140 defines the spatial section 110, i.e., its location within the image area of the video 120 and the image 100, respectively, and possibly the size and shape of the image area of the image 144. As shown in FIG. 2, the section 110 does not necessarily have to be rectangular, convex, or even connected. For example, in the example of FIG. 2, the section 110 is composed of two separate sub-areas 110a and 110b. Furthermore, the first information 140 may include hints on how the extractor 134 should modify or replace some of the encoding parameters or parts of the data stream 124 or 124′, respectively, such as image size parameters adapted to reflect the change in image area due to the extraction from image 100 to image 144. In particular, the first information 140 may include substitution or modification instructions for parameter sets of the video data stream 124′ to be applied by the extractor 134 in the extraction process in order to correspondingly change or replace the corresponding parameter sets included in the video data stream 124′ and carry them over to the reduced or extracted video data stream 136.
[0033] In other words, the extractor 134 receives the video data stream 124′, reads the extracted information 132 from the video data stream 124′, and derives from the extracted information its position and arrangement within the spatial section 110, i.e., the image area of the video 120, i.e., based on the first information 140. Therefore, the extractor 130 identifies, based on the first information 140, those slice portions 126 encoded in the slices that fall within the section 110, and thus carries them over to the reduced or extracted video data stream 136, while the slice portions 126 relating to slices outside the section 110 are dropped by the extractor 134. Furthermore, the extractor 134 may use the information 140 to correctly set one or more parameter sets previously in the data stream 124′, as outlined above, or adopt them in the reduced or extracted video data stream 136, i.e., by modification or replacement. Thus, one or more parameter sets may relate to a picture size parameter that is set according to the information 140 to a size corresponding to the sum of the area sizes of the section 110, i.e., the sum of the areas of all portions 110a and 110b of the section 110 if the section 110 is not a connected area as exemplarily shown in FIG. 2. The section 110-sensitive dropping of slice portions due to the parameter set adaptation limits the video data stream 124' to the section 110. Furthermore, the extractor 134 modifies the slice addresses of the slice portions 126 in the reduced or extracted video data stream 136. These slices are indicated using hedges in FIG. 2. That is, the hatched slice portions 126 are slices that fall into the section 110 and are therefore extracted or taken over, respectively.
[0034] It should be noted that information 142 can be imagined not only in a situation where it is added to the complete video data stream 124', but also where the sequence of slice portions contained therein includes slice portions encoded for slices within the spatial section as well as slice portions encoded for slices outside the spatial section. Rather, the data stream including information 142 may already be stripped out so that the sequence of slice portions contained in the video data stream includes slice portions encoded for slices within the spatial section, but no slice portions encoded for slices outside the spatial section.
[0035] In the following, different examples for embedding the second information 142 in the data stream 124′ and the processing thereof are presented. Generally, the second information 142 is conveyed as signaling that indicates hints about how to perform slice address modification within the data stream 124′, either explicitly or by selecting one of several options. In other words, the second information 142 is conveyed in the form of one or more syntax elements, the possible values of which may, for example, explicitly signal slice address alternatives or, together, distinguish a number of possible signaling possibilities for associating slice addresses for each slice portion of the video data stream 136 by selecting the setting of one or more syntax elements of the data stream. However, it should be noted that the number of meaningful or possible settings of the one or more syntax elements embodying the second information 142 depends on how the video 120 was encoded into the video data stream 124 and the selection of the section 110, respectively. For example, assume that section 110 is a rectangular, connected region within image 100, and that video encoding core 120 is concerned with performing encoding on this section without further restricting the encoding as far as the interior of section 110 is concerned. The organization of section 110 by more than one region 110a and 110b would not apply. That is, only dependencies on the exterior of section 110 would be suppressed. In this case, section 110 would have to be mapped unmodified to the image area of image 144 of video 138, i.e., without scrambling the position of any subregion of section 110, and the assignment of addresses α and β to the slice portions carrying the slices that make up section 110 would be uniquely determined by aligning the interior of section 110 as is to the image area of image 144. In this case, the configuration of information 142 generated by video data stream generator 128 would be unique. That is, video data stream generator 128 would have no choice but to configure information 142 in this way. However, there would be other signaling options available for information 142 from an encoding perspective.However, even in the absence of this alternative, for example, signaling 142 explicitly indicating the inherent slice address modification has the advantage that extractor 134 does not have to perform the aforementioned onerous task of determining by itself the slice addresses α and β for slice portions 126 employed from stream 124′, but rather simply derives from information 142 how to modify the slice addresses of slice portions 126.
[0036] Depending on different embodiments regarding the nature of the information 142, as outlined further below, the extractor 134 either preserves or maintains the order in which the slice portions 126 are passed from the stream 124' to the reduced or extracted stream 136, or modifies the order in a manner defined by the information 142. In either case, the reduced or extracted data stream 136 output by the extractor 134 can be decoded by a conventional decoder 146. The decoder 146 receives the extractor video data stream 136 and decodes a video 138 therefrom. The image 134 is smaller than an image of the video 120, such as image 100. The image area is then filled by placing the decoded slices 108 from the slice portion 126 in the video data stream 136 in a manner defined by the slice addresses α and β conveyed in the slice portion 126 within the video data stream 136.
[0037] That is, up to this point, FIG. 2 has been described in such a way that the description is tailored to various embodiments regarding the exact nature of second information 142, which are described in more detail below.
[0038] The embodiments described herein use explicit signaling of slice addresses to be used by extractor 134 when modifying slice addresses of slice portions 126 carried over from stream 142 to stream 136. The embodiments described below use signaling 142 that allows extractor 134 to signal one of several permitted options for how to modify slice addresses. For example, optional allowances resulting from section 110 being coded in a manner that limits coding interdependencies within section 110 so as not to cross spatial boundaries of section 110, which in turn divides section 110 into two or more regions, such as 110a and 110c or tiles a, b, c, and d, as shown in FIG. 1. The latter embodiment may further include extractor 134 to perform the tedious task of calculating addresses by itself, but still allows for efficient processing of different types of image content within original video 120 to result in meaningful video 138 at the receiving end based on the corresponding extracted or downscaled video data stream 136.
[0039] 1, the tedious task of determining slice addresses in the extractor 134 is alleviated in accordance with an embodiment of the present application by explicitly transmitting how to modify the addresses in the extraction process via second information 142. An example of a syntax that can be used for this purpose is given below:
[0040] In particular, information 142 can be used to explicitly signal new slice addresses to be used in the slice header of the extracted MCTS by including a list of slice address replacements included in stream 124′ in the same order as the slice portions 126 are carried in bitstream 124′. For example, see the example of FIG. 1. Here, information 142 is an explicit signaling of slice addresses that follows the order of slice addresses 124′ in the bitstream. Again, as slices 108 and corresponding slice portions 126 are carried over in the extraction process, the order of extracted video data stream 136 corresponds to the order in which these slice portions 126 were included in video data stream 124′. According to the following syntax example, information 142 explicitly signals slice addresses in a manner that starts with the second slice or slice portion 126 onward. In the case of FIG. 1, this explicit signaling would correspond to second information 142 indicating or signaling the list {2, 4, 8}. An exemplary syntax for this embodiment is shown in the syntax example of FIG. 3, highlighting which shows the corresponding addition of explicit signaling 142 in addition to the MCTS extraction information SEI known from [2].
[0041] The semantics are shown below.
[0042] num_associated_slices_minus2[i] + 2 indicates the number of slices containing MCTSs with mcts identifiers equal to any value in the list mcts_identifier[i][j]. The value of num_extraction_info_sets_minus1[i] ranges from 0 to 2. 32 Must be in the range -2.
[0043] output_slice_address[i][j] identifies the slice address of the j-th slice in bitstream order belonging to an MCTS whose mcts identifier is equal to any value in the list mcts_identifier[i][j]. The value of output_slice_address[i][j] ranges from 0 to 2. 32 Must be in the range -2.
[0044] Note that this can be controlled by the presence of information 142 in the MCTS extraction information SEI or by a flag in the data stream in addition to the MTCS related information 140. This flag can have a name such as slice_reordering_enabled_flag. If set, information 142 such as num_associated_slices_minus2 or output_slice_address will be present in addition to information 140; otherwise, information 142 will not be present and the inter-positional arrangement of slices will be respected in the extraction process or otherwise handled.
[0045] Furthermore, it should be noted that using the H.265 / HEVC nomenclature, the "_segment_" part of the syntax element names used in Figure 3 can be replaced with "_segment_address_" instead, but the technical content remains the same.
[0046] And further, it should be noted that although num_associated_slices_minus2 suggests that information 142 indicates the number of slices in section 110 in the form of an integer indicating this number in the form of a difference of 2, the number of slices in section 110 can also be signaled in the data stream directly or as a difference of 1. In the latter case, for example, num_associated_slices_minus1 would be used instead as the syntax element name. It should be noted that the number of slices in any section 110 can also be, for example, one.
[0047] In addition to the MCTS extraction process previously envisaged in [2], additional processing steps are associated with explicit signaling by information 142, as embodied in Figure 3. These additional processing steps facilitate the derivation of slice addresses within the extraction process performed by extractor 134, and the following overview of this extraction process underlines where this facilitation takes place.
[0048] Let the bitstream inBitstream, the target MCTS identifier mctsIdTarget, the target MCTS Extraction Information Set identifier mctsEISIdTarget, and the target highest TemporalId value mctsTIdTarget be the inputs to the sub-bitstream MCTS extraction process. The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream. It is a bitstream conformance requirement of the output bitstream that any output sub-bitstream that is the output of the process specified with the bitstream in this section be a conforming bitstream.
[0049] The output sub-bitstreams are derived as follows. - The bitstream outBitstream is set to be equal to the bitstream inBitstream. - Remove all NAL units with TemporalId greater than mBitsTIdTarget from outBitstream. - For each remaining VCL NAL in each access unit in outBitstream, adjust the slice segment header as follows: - For the first VCL NAL unit, set the value of first_slice_segment_in_pic_flag to 1, otherwise set it to 0. - Set the values of slice_segment_address of NAL units (i.e., slices) other than the first, starting from the second in bitstream order, according to the list output_slice_address[i][j].
[0050] A variation of the embodiment described with reference to FIGS. 1 to 3 eases the burdensome task of the slice address determination and extraction process performed by the extractor 134 by using information 142 as explicit slice address signaling. According to the specific example of FIG. 3 , the information simply includes a substitute slice address 143 for each slice portion 126 subsequent to the first in the slice portion order, while maintaining the slice order when transferring slice portions 126 related to slices 108 in section 110 from data stream 124′ to data stream 136. The slice address substitution relates to the assignment of one-dimensional slice addresses to the image area of each image 144 of video 138. Using order 112 does not conflict with simply covering the image area of image 144 in accordance with order 112 using the sequence of slices 108 obtained from the sequence of transferred slice portions. It should be noted, and will be further discussed below, that explicit signaling can also be applied to the first slice address, i.e., the slice address of the first slice portion 126. Even in the latter case, information 142 may include a substitute 143. Such signaling of information 142 may also enable the placement of the corresponding slice 108 first in the order within stream 124′. The slice portion 126 of the slice portions 126 carrying the slice 108 within section 110 may be located somewhere other than the start of encoding order 112, which may be the upper-left corner of the image, as shown in FIG. 1 . Explicit signaling of slice addresses may be used to provide additional flexibility in rearranging cross-sectional regions 110a and 110b if such a possibility exists or is permitted. For example, in modifying the example shown in FIG. 3 , information 142 also explicitly signals slice address permutation 143 for the first slice portion 126 within data stream 136, i.e., α and β in the case of FIG. 2 . Signaling 142 then enables distinguishing between two permitted or available placements of section regions 110a and 110b within the output image of video 138, namely, placement of section region 110a on the left side of the image.This leaves the slice 108 corresponding to the slice portion 126 transmitted first in the data stream 136 at the start of the encoding order 112, and maintains the order among the slices 108 as far as the video data stream 124 is concerned, compared to the slice order of the image 100 of the video 120. The section region 120a is then placed to the right of the image 144, which changes the order of the slices 108 in the image 144 traversed by the encoding order 112, compared to the order in which the same slices were traversed in the original video in the video data stream 124 by the encoding order 106. The signaling 142 explicitly indicates the slice addresses used for correction by the extractor 134 as a list of slice addresses 143 ordered or assigned to the slice portions 126 in the video data stream 124'. That is, the information 142 indicates the slice addresses of each slice in the section 110 consecutively in the order in which these slices 108 are traversed by the encoding order 106, and this explicit signaling may result in a permutation that changes the order in which the section regions 110a and 110b of the section 110 are traversed by the order 112 compared to the order in which they were traversed by the original encoding order 106. The slice portions 126 carried over from the stream 124' to the stream 136 are reordered by the extractor 134 accordingly, i.e., so as to conform to the order in which the slices 108 of the extracted slice portions 126 are traversed consecutively by the order 112. The standard conformance of the reduced or extracted video data stream 136 according to the slice portions 126 carried over from the video data stream 124' must strictly follow each other along the encoding order 112, i.e., be maintained with monotonically increasing slice addresses α and β as modified by the extractor 134 in the extraction process. Therefore, the extractor 134 modifies the order between the inherited slice portions 126 so as to order the inherited slice portions 126 according to the order of the slices 108 coded therein along the coding order 112 .
[0051] The latter embodiment, i.e., the possibility of rearranging slices 108 in inherited slice portions 126, is exploited according to a further variation of the description of FIG. 2. Here, second information 142 signals a rearrangement of the order between slice portions 126 coded in any slice 108 within section 110. Explicitly signaling slice address substitution 143 in a manner that results in the rearrangement of slice portions is one possibility for this currently described embodiment. However, the rearrangement of slice portions 126 within the extracted or downsized video data stream 136 may be signaled by information 142 in a different manner. This embodiment ultimately arranges reconstructed slices reconstructed from slice portions 126 in a decoder 146 strictly according to the coding order 112, which would, for example, use a tile raster scan order, thereby filling the image area of image 144. The rearrangement signaled by signaling 142 may be selected to rearrange or change the order among the extracted or inherited slice portions 126, and the arrangement by decoder 146 may cause partitioned regions, such as regions 110a and 110b, to change their order compared to a case where the order of the slice portions is not changed. If the rearrangement signaled by information 142 leaves the order that was in the original data stream 124', then sub-regions 110a and 110b may maintain the relative positions that they had in the original image regions of image 110.
[0052] To illustrate the present variation, reference is made to Figure 4, which shows an image region of an image 110 of an original video and an image 144 of an extracted video. Furthermore, Figure 4 shows an exemplary tile division into tiles 104 and two disjoint cross-sectional regions or zones 110a and 110b, i.e., opposite sides 150 of image 110. r and 150 l 1 shows an exemplary extraction section 110 consisting of cross-sectional areas 110a and 110b abutting side 150. r and 150 lAs far as the extraction direction of the two images is concerned, i.e., along the vertical direction, their positions coincide. In particular, FIG. 4 shows that the image content of the image 110, and therefore the video to which the image 110 belongs, is of a particular type, i.e., a panoramic video. Therefore, the side 150 r and 150 l forms a scene when projecting a 3D scene onto the image area of image 110. Under image 110, FIG. 4 shows two options for how sections 110a and 110b are positioned within the output image area of image 144 of the extracted video. Again, tile names are used in FIG. 4 to describe the two options. The two options result from the video 210 being coded into stream 124 in such a way that each of the two zones 110a and 110b is coded independently from the outside. Note that in addition to the two permissible choices shown in FIG. 4, two more options are possible if each of zones 110a and 110b is subdivided into two tiles, and within each tile, in turn, the video is coded independently from the outside into stream 124, i.e., if each tile in FIG. 4 is a part that can be extracted by itself or is coded independently (in terms of spatial and temporal interdependence). The two options then correspond to differently scrambling tiles a, b, c, and d of section 110.
[0053] Again, the embodiment now described with respect to FIG. 4 aims at slice 108 or the NAL units of data stream 124′ carrying slice 108 having second information 142 that changes the order and is transferred, adopted, or written from data stream 124′ to extracted or reduced video data stream 136 during the extraction process by extractor 134. FIG. 4 illustrates the case where desired image subsection 110 consists of non-abutting tiles or section areas 110a and 110b in the image plane that image 110 spans. This complete encoded image plane of image 110 is shown in the upper part of FIG. 4 with a tiled boundary with desired MCTS 110 consisting of two rectangles 110a and 110b containing tiles a, b, c, and d. As far as scene content is concerned, or due to the fact that the video content shown in FIG. 4 is panoramic video content, when image 110 covers the camera's surroundings over 360° in an isometric projection as exemplarily shown here, desired MCTS 110 is bounded by left and right boundaries 150. r and 150 l 4, where the image content is panoramic content, among options 1 and 2 for arranging the section regions 110a and 110b within the image area of the output image 144, the second option actually makes more sense. However, if the image content is of another type other than panoramic, the situation may be different.
[0054] In other words, the order of tiles A, B, C, and D in the complete image bitstream 124' is {a, b, c, d}. If this order were simply transferred to the encoding order of the extracted or downscaled video data stream 136 or the arrangement of the corresponding tiles in the output image 144, as shown in the lower left of FIG. 4, the extraction process itself would not result in the desired data arrangement in the output bitstream 136 in the above example case. As shown in the lower right of FIG. 4, a preferred arrangement {b, a, d, c} is shown, which results in a video bitstream 136 that results in continuous image content on the image plane of the picture 144 for a legacy device such as a decoder 146. Such a legacy device 146 may not have the ability to rearrange sub-image regions of the output image 144 as a post-processing step in the pixel domain after decoding. That is, rendering, and even sophisticated devices, may prefer to avoid the effort of post-processing.
[0055] Thus, according to the motivating example described above with respect to FIG. 4, the second information 142 provides a means for signaling a preferred ordering among several choices or options to the encoding side of the video data stream generator 128. For example, the section regions 110a, 110b, each consisting of a set of one or more tiles in the video data stream 124′, should be positioned in the extracted or downsized video data stream 136 or its image region spanned by the image 144. According to the specific syntax example presented in FIG. 5, the second information 142 includes a list encoded in the data stream 124′ and indicates the position of each slice 108 that falls within the section 110 in the extracted bitstream in its original or input bitstream order, i.e., the order of occurrence in the bitstream 124′. For example, in the example of FIG. 4, preferred option 2 would be a list reading {1, 0, 3, 2}. FIG. 5 illustrates a specific example of syntax including the second information 142 in an MCTS Extraction Information Set SEI message.
[0056] The semantics are as follows:
[0057] num_associated_slices_minus1[i] plus 1 indicates the number of slices containing MCTSs with mcts identifiers equal to any value in the list mcts_identifier[i][j]. The value of num_extraction_info_sets_minus1[i] ranges from 0 to 2. 32 Must be in the range -2.
[0058] output_slice_address[i][j] identifies the absolute position of the jth slice in bitstream order belonging to the MCTS whose mcts identifier in the output bitstream is equal to any value in the list mcts_identifier[i][j]. The value of output_slice_address[i][j] ranges from 0 to 2. 23 in the range of -2.
[0059] Additional processing steps of the extraction process defined in [2] are now described to facilitate understanding of the signaling embodiment of Figure 5. Additions related to [2] are highlighted with underlines.
[0060] The bitstream inBitstream, the target MCTS identifier mctsldTarget, the target MCTS Extraction Information Set identifier mctsEISIdTarget, and the target highest TemporalId value mctsTldTarget are the inputs to the sub-bitstream MCTS extraction process. The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream. It is a bitstream conformance requirement of the input bitstream that the output sub-bitstream, which is the output of the process specified in this section along with the bitstream, be a conforming bitstream.
[0061] OutputSliceOrder[j] is derived from the list of the ith extracted information set, output_slice_order[i][j].
[0062] The output sub-bitstreams are derived as follows. The bitstream outBitstream is set to be the same as the bitstream inBitstream. [...] - Remove all NAL units with TemporalId greater than mBitsTIdTarget from outBitstream. - Sort the NAL units of each access unit according to the list OutputSliceOrder[j]. - For each remaining VCL NAL unit in outBitstream, adjust the slice segment header as follows: - For the first VCL NAL unit in each access unit, set the value of first_slice_segment_in_pic_flag to 1, otherwise set it to 0. - Set the value of slice_segment_address according to the tile settings defined in the PPS where pps_pic_parameter_set_id is equal to slice_pic_parameter_set_id.
[0063] Therefore, the above-described variation of the embodiment of FIG. 2 is summarized in FIG. 5. This variation differs from that discussed above with respect to FIG. 3 in that the second information 142 does not explicitly signal how the slice addresses should be modified. That is, the second information 142 does not explicitly signal the substitution of slice addresses for the slice portions extracted from the data stream 124 to the data stream 136 according to the variation outlined above. Rather, the embodiment of FIG. 5 relates to the case where the first information 140 defines a spatial section 110 within the image region of the image 100 as consisting of at least a first subregion 110a, in which video independent from outside the first subregion 110a is encoded into a video data stream 124′, and a second subregion 110b, in which video 120 is encoded into a video data stream 124′ independent from outside the second subregion 110b. None of the slices 108 intersects the boundary of either the first or second subregion 110a or 110b. At least for these regions 110a and 110b, the images 144 of the output video 138 of the extracted video data stream 136 may be configured differently. Thus, according to the two options and the variants discussed above with respect to FIG. 5, the second information 142 signals re-sorting information that indicates the order of the slice portions and how to re-sort the slice portions 126 of the slices 108 located in the region 110 when extracting the reduced video data stream 136 from the associated video data stream 124′. The re-sorting information 142 may, for example, include a set of one or more syntax elements. Among the possible conditions signaled by the one or more syntax elements forming the information 142, there may be a condition in which the re-sorting maintains the original order. For example, information 142 indicates the rank (compared to 141 in FIG. 5) of each slice portion 126 that encodes one of the slices 108 of image 100 that fall within region 110, and extraction device 134 then sorts the slice portions 126 in extracted or reduced video data stream 136 according to these ranks. Next, the extractor 134 modifies the slice addresses of the reconstructed slice portions 126 in the reduced or extracted video data stream 136 in the following manner: The extractor 134 knows about these slice portions 126 in the video data stream 124′ that were extracted from the data stream 124′ into the data stream 136. Thus, the extractor 134 knows about the slices 108 that correspond to these inherited slice portions 126 in the image area of the image 100. Based on the rearrangement information provided by the information 142, the extractor 134 can determine how the sub-areas 110a and 110b have been translated and shifted relative to each other to result in a rectangular image area corresponding to the image 144 of the video 138. For example, in the example of FIG. 4 , option 2, the extractor 134 assigns slice address 0 to the slice corresponding to tile b, since slice address 0 occurs at the second position in the list of ranks provided by the second information 142. Thus, the extractor 134 can locate one or more slices associated with tile b, and then, according to the re-sorting information, track the next slice associated with the slice address pointing to the location in the coding order 112 immediately following tile b in the image region. In the example of FIG. 4, this is the slice associated with tile a, since it is the next ranked location indicated for the first slice a in region 110 of image 100. In other words, the re-sorting is limited to leading to any of the possible rearrangements of sub-regions 110 a and 110 b. It is true that, for each sub-region individually, the coding orders 106 and 112 traverse the respective sub-regions in the same path. However, due to tiling, the areas corresponding to sub-areas 110a and 110b of the image area of image 144, i.e., in the case of option 2, the area of the combination of tiles b and d on the one hand and the area of the combination of tiles A and C on the other hand, are traversed in an interleaved manner by the encoding order 112, and the associated slice portions 126 encoding the slices in the corresponding tiles are interleaved accordingly within the extracted or downsized video stream 136.
[0064] A further embodiment is to signal a guarantee that the further ordering signaled using existing syntax reflects the preferred output slice order. More specifically, this embodiment can be implemented by interpreting the occurrence of the MCTS Extraction SEI message [2] as guaranteeing the ordering of the rectangles forming the MCTS in the MCTS SEI message from sections D.2.29 and E.2.29 of [1], which indicates the preferred output order of tiles / NAL units. In the example of Figure 5, rectangles are used in the order {b, a, d, c} for each contained tile. An example of this embodiment would be identical to the one above, except for the derivation of OutputSliceOrder[j].
[0065] OutputSliceOrder[j] is derived from the rectangle order signaled in the MCTS SEI message.
[0066] To summarize the above example, the second information 142 can signal to the extractor 134 how to re-sort the slice portions 126 of the slices entering the spatial section 110 when extracting the downsized video data stream 136 from the video data stream, with respect to how the slice portions 126 are ordered in the sequence of slice portions 126 in the video data stream 124'. The slice address of each slice portion 126 in the sequence of slice portions in the video data stream 124' one-dimensionally indexes the encoding start position of the slice 108 encoded in each slice portion 126 along the first encoding scan order 106, which in turn traverses the image domain along which the image 100 is encoded in the sequence of slice portions in the video data stream. This results in the slice addresses of a series of slice portions in the video data stream 124' increasing monotonically, and the modification of the slice addresses in extracting the reduced video data stream 136 from the video data stream 124' sets the slice addresses of the slice portions 126 to index the encoding start positions of the slices measured along the second encoding scanning order 112 across the reduced picture area defined by sequentially arranging the slices encoded in the slice portions to which the reduced video data stream 136 is confined, and reordered as signaled by the second information 142. The first coding scan order 106 traverses the image regions within each of the set of at least two partitioned regions in a manner that matches the manner in which each spatial region is traversed by the second coding scan order 112. Each of the set of at least two partitioned regions is indicated by the first information 140 as a subarray of rectangular tiles into rows and columns into which the image 100 is subdivided, where the first and second coding scan orders use a row-wise tile raster scan that completely traverses the current tile before proceeding to the next tile.
[0067] As already explained above, the order of the output slices can be derived from another syntax element, such as output_slice_address[i][j] above. The important addition to the above syntax example for output_slice_address[i][j] in this case is that the slice addresses of all associated slices, including the first one that enables sorting, are signaled. That is, num_associated_slices_minus2[i] becomes num_associated_slices_minus1[i]. This example embodiment would be identical to the one above, except for the derivation of OutputSliceOrder[j].
[0068] OutputSliceOrder[j] is derived from the list of the ith extracted information set, output_slice_address[i][j].
[0069] Yet another embodiment consists of a single flag in information 142 indicating that the video content wraps around a set of image boundaries, such as vertical image boundaries. Thus, an output order is derived in extractor 134 that corresponds to an image subsection that includes tiles on both image boundaries, as described above. In other words, information 142 can signal one of two options: a first option in the plurality of options indicates that the video is a panoramic video that shows a scene such that different edge portions of the image abut each other scene-wise, and a second option in the plurality of options indicates that the different edge portions are not adjacent to each other scene-wise. The at least two section regions a, b, c, d of which section 110 is composed are composed of different portions of different edge portions, i.e., first and second zones 110a, 110b adjacent to left and right edges 150r and 150l, so that when second information 142 signals a first option, the reduced image region is composed by combining a set of at least two partial regions, the first zone and the second zone abutting along different edge portions, and when second information 142 signals a second option, the reduced image region is composed by combining a set of at least two partial regions, the first and second zones having different edge portions facing opposite each other.
[0070] For the sake of completeness, it should be noted that the shape of the image region of image 144 is not limited to being compatible with stitching various regions, such as tiles a, b, c, and d, of section 110 together in a manner that maintains the relative placement of connected clusters, such as (a, c, b, d) in FIG. 1, or stitching such clusters along the shortest possible interconnecting direction, such as stitching zones (a, c) and (b, d) in FIG. 2 horizontally. Rather, for example, in FIG. 1, the image region of the extracted data stream image could be a column of all four regions, and in FIG. 2, it could be a column of all four regions. In general, the size and shape of the image region of image 144 of video 138 can be present in different parts of data stream 124′. For example, this information can be provided in the form of one or more parameters in information 140 to guide the extraction process of extractor 134 regarding the adaptation of a parameter set when extracting section-specific substreams 136 from stream 124′. The nested parameter set that replaces the parameter set for stream 124′ may be included in information 140, for example, and may include the following: For example, parameters related to the size of a picture may indicate the size and shape of picture 144, e.g., in pixels, and replacing the parameter set for stream 124′ during extraction in extractor 134 overwrites the old parameters in the parameter set that indicate the size of picture 100. However, additionally or alternatively, the picture size of picture 144 may also be indicated as part of information 142. Explicitly signaling the shape of picture 144 in an easily readable high-level syntax element such as information 142 may be particularly advantageous when slice addresses are not explicitly provided in information 142. To derive the addresses, parsing the nested parameter set may be necessary.
[0071] It is also worth noting that in more sophisticated system setups, a cubic projection can be used. This projection avoids known weaknesses of the equirectangular projection, such as the large variation in sampling density. However, when using a cubic projection, a rendering stage is required to recreate a continuous viewport from the content (or a subsection thereof). Such rendering stages may offer different trade-offs between complexity and functionality, i.e., some viable off-the-shelf rendering modules may expect a predetermined placement of the content (or a subsection thereof). In such scenarios, the possibility to manipulate the placement is essential, as will be possible in the next invention.
[0072] In the following, embodiments relating to the second aspect of the present application will be described. The description of the embodiments of the second aspect of the present application will again begin with a brief introduction to the general problem or problem envisioned and addressed by these embodiments.
[0073] An interesting use case for MCTS extraction in a context not limited to 360° video is a composite video containing multiple resolution variations of adjacent content on the screen, as shown in Figure 6. The bottom part of Figure 6 shows a composite video 300 with multiple resolutions, where lines 302 indicate the tile boundaries of tiles 304 into which a composition of high resolution video 306 and low resolution video 308 has been subdivided for encoding into corresponding data streams.
[0074] More precisely, FIG. 6 illustrates images such as an image of a high-resolution video at 306 and a simultaneous time image of a low-resolution video at 308. For example, both videos 306 and 308 show the exact same scene, i.e., have the same view or field of view, but at different resolutions. However, the fields of view may only partially overlap each other, and in the overlap zone, the spatial resolution indicates a different number of samples in the photograph of video 306 compared to video 308. In practice, videos 306 and 308 differ in fidelity, i.e., the number of samples or pixels of the same scene section. Photos of the same location in videos 306 and 308 are composited side-by-side into a larger photograph, resulting in the composite video 300. For example, FIG. 6 illustrates an image of video 308 being horizontally halved, with the two halves overlapping one above the other and attached to the right of the simultaneous time image of video 306 to obtain the corresponding video 300. The tiles 304 are executed so as not to cross the junction between the high resolution image of the video 306 on the one hand and the image content coming from the low resolution video 308 on the other hand. In the example of Figure 6, the tiling of the image 300 results in 8x4 tiles into which the high resolution image content of the image of the video 300 is divided and 2x2 tiles into which each half coming from the low resolution video 308 is divided. Overall, the width of the image of the video 300 is 10x4.
[0075] When such a multiple resolution composite video 300 is MCTS encoded in an appropriate manner, the MCTS extraction can produce a content variant 312. Such a variant 312 can be designed, for example, to depict a given subpicture 310a in high resolution and the rest or another subsection of the scene in lower resolution, as shown in Figure 7, where the MCTS 310 in the composite image bitstream is compared to three sides or distinct regions 310a, b, c.
[0076] That is, the extracted video image, i.e., image 312, has three fields 314a, 314b, and 314c, each corresponding to one of the MCTS regions 310a, b, and c, where region 310a is a sub-area of the high-resolution image area of image 300, and the other two regions 310b and 310c are sub-regions of the low-resolution video content of image 308.
[0077] With this being said, embodiments of the present application relating to the second aspect of the present application will now be described with respect to FIG. 8. In other words, FIG. 8 illustrates a scenario in which multi-resolution content is presented to a recipient site, particularly to individual sites participating in the generation of multiple versions at different resolutions up to the receiving site. It should be noted that the individual devices and processes located at various sites along the process path illustrated in FIG. 8 represent individual devices and methods, and therefore, FIG. 8 should not be interpreted as illustrating only the entire system or method. A similar statement applies with respect to FIG. 2, which also illustrates individual devices and methods. The reason for displaying all of these sites together in one diagram is solely to facilitate understanding of the interrelationships and advantages resulting from the embodiments described with respect to these figures.
[0078] FIG. 8 shows a video data stream 330 encoded into a video 332 of picture 334. Image 334, in turn, is the result of stitching together a contemporaneous image 336 from a high-resolution video 338 and an image 340 from a low-resolution video 342. More precisely, images 336 and 340 from videos 338 and 342 correspond to the exact same viewport 344 or at least partially overlap to display the same scene in the overlapping region. However, the high-resolution video image 336 samples the same scene content as the corresponding low-resolution image 340 with a higher number of samples, and therefore the scene resolution fidelity of image 336 is higher than that of image 340. The composition of image 334 of composition video 332 based on videos 338 and 342 is performed by a composer 346. Composer 346 stitches together images 336 and 340. In doing so, the composer 346 can subdivide the images 340 of the low-resolution video and / or the images of the high-resolution video 338 to provide advantageous filling and patching of image regions of the images 334 of the composition video 332. The video encoder 348 then encodes the composite video 332 into the video data stream 330. A video data stream generator 350 may be included in the video encoder 348 or connected to the output of the video encoder 348 to provide signaling 352 indicative of the images 334, or each image or video 332, in the video data stream 330, or each picture or video 332 within a particular sequence of pictures in the video 332, once encoded into the video data stream 330, to display common scene content multiple times, i.e., in different spatial portions at different resolutions from each other. These portions are shown in FIG. 8 using H and L to indicate their origin due to the composition performed by the composer 346, and are indicated using reference numerals 354 and 356. Two or more different resolution versions may be put together to form the content of image 332, and care should be taken to use two versions as shown in Figures 8 and 6 and 7, respectively, which are for illustrative purposes only.
[0079] For illustrative purposes, FIG. 8 shows a video data stream processor 358 receiving a video data stream 330. The video data stream processor 358 may be, for example, a video decoder. In either case, the video data stream processor 358 may examine the signaling 352 to determine whether the video data stream processor 358 should begin processing, such as decoding the video data stream 330, based on this signaling, which further depends on the specific capabilities of the video data stream processor 358 or a device connected downstream thereof. For example, the video data stream processor 358 may only present fully encoded images 332 in the video data stream 330, and then the video data stream processor 358 may refuse to process the video data stream 330, with the signaling 352 indicating that the individual images show common scene content at different spatial resolutions in different spatial portions of these individual images. That is, the signaling 352 indicates that the video data stream 330 is a multi-resolution video data stream.
[0080] Signaling 352 may include, for example, a flag conveyed within data stream 330, which is switchable between a first state and a second state. The first state may indicate, for example, the fact just outlined, i.e., that individual images of video 332 show multiple versions of the same scene content at different resolutions. The second state indicates that this situation does not exist, i.e., the images display only one scene content at one resolution. Thus, video data stream processor 358 would respond to flag 352 being in the first state by refusing to perform certain processing tasks.
[0081] Signaling 352, such as the aforementioned flags, may be conveyed within the data stream 330 within the sequence parameter set or video parameter set. Possible syntax elements reserved for future use in HEVC are exemplarily identified as possible candidates in the following description.
[0082] As illustrated above with respect to Figures 6 and 7, it is possible, but not necessary, that the composite video 332 be encoded in the video data stream 330 in such a way that, for each of a set of non-overlapping spatial regions 360, such as tiles or tile sub-arrays, the encoding is independent from outside the respective spatial regions 360. Coding independence, as explained above with respect to Figure 2, restricts spatial and temporal prediction and / or context derivation to not cross boundaries between spatial regions 360. Thus, for example, coding dependence may restrict the encoding of a particular spatial region 360 of a particular image 334 of the video 332 to only reference co-located spatial regions in another image 334 of the video 332 that is an image similarly subdivided into spatial regions 360 as the image 334. The video encoder 348 may provide the extracted information, such as information 140 or a combination of information 140 and 142, to the video data stream 330, or may have a respective device, such as device 128, connected to its output. The extraction information may relate to specific or all possible combinations of spatial regions 360 as extraction sections, such as extraction section 310 of FIG. 7. Signaling 352 may then include information regarding the spatial subdivision of image 334 of video 332 into sections 354 and 356 of different scene resolutions, i.e., the size and location of each spatial section 354 and 356 within the image region of image 332. Based on such information in signaling 352, video data stream processor 358 may, for example, exclude certain extraction sections from the list of possibly extractable sections of video stream 330. For example, these extraction sections that combine spatial regions 360 divided into different sections 354 and 356 may avoid the performance of video extraction for extraction sections that combine different resolutions. In this case, video stream processor 358 may include an extraction device, such as the extraction device of FIG. 2.
[0083] Additionally or alternatively, signaling 330 may include information regarding the different resolutions at which images 334 of video 332 show common scene content with respect to one another. Further, signaling 352 may simply indicate a count of the different resolutions at which images 334 of video 332 show common scene content multiple times at different image locations.
[0084] As already mentioned, the video data stream 330 may include extraction information regarding a list of potential extraction regions from which the video data stream 330 may be extracted. The signaling 352 may then include, for each of at least one or more of these extraction regions, further signaling indicating a viewport direction of a subregion of each extraction region in which the common scene content is shown at the highest resolution within each extraction region, and / or an areal share of the cross-sectional area of each extraction region in which the common scene content is shown at the highest resolution within each extraction region, and / or a spatial subdivision of each extraction region into subregions in which the common scene content is shown at mutually different resolutions within the overall area of each extraction region.
[0085] Therefore, such signaling 352 may be exposed at a high level in the bitstream so that it can be easily pushed up to a streaming system.
[0086] One option is to use one of the commonly reserved zero X-bit flags in the profile layer level syntax, which can be named as the general non-multiresolution flag.
[0087] A general non-multiresolution flag equal to 1 specifies that the decoded output picture will not contain multiple versions of the same content at different resolutions (i.e., each construct such as region packing is constrained). A general non-multiresolution flag equal to 0 (zero) specifies that the bitstream may contain such content (i.e., no constraints).
[0088] In addition, the present invention therefore consists in signaling that informs about the nature of the complete bitstream content characteristics, i.e., the number and resolution of variants in the configuration, and additional signaling that provides information about the coded bitstream in an easily accessible format regarding the following characteristics of each MCTS: Main Viewpoint Orientation: What is the orientation of the MCTS high-resolution viewport center, e.g., in terms of delta-yaw, pitch, and / or roll from a predefined initial viewport center? Overall coverage Percentage of complete content represented in MCTS. High to low resolution ratio The ratio between high-resolution and low-resolution areas of the MCTS, i.e., how much of the total covered content is represented in high resolution / fidelity.
[0089] Proposed signaling information already exists for viewport orientation or overall coverage of full omnidirectional video. Similar signaling would need to be added for subregions that could potentially be extracted. This information is in the form of an SEI [2], so it could be included in the motion-constrained tileset extraction information nesting the SEI. However, such information is required to select the MCTS to extract. Obtaining motion-constrained tileset extraction information nesting the SEI would add additional indirection and require deeper analysis to select a specific MCTS (the motion-constrained tileset extraction information nesting the SEI contains additional information not necessary for selecting the extracted set). From a design perspective, a cleaner approach would be to signal this information, or a subset of it, at a central point containing only the key information for selecting the extracted set. Furthermore, the mentioned signaling includes information about the entire bitstream; in the proposed case, it would be preferable to signal which is the high-resolution coverage and which is the low-resolution coverage. Also, if more resolutions are mixed, the coverage of each resolution and the viewport orientation of the video extracted at the mixed resolution.
[0090] One implementation is to add coverage for each resolution to the 360 ERP SEI from [2]. This SEI can then be included in the motion-constrained tileset extraction information that nests the SEI, which requires the tedious task described above.
[0091] In another embodiment, a flag is added to the MCTS extraction information set SEI, e.g., omnidirectional information, indicating the presence of the discussed signaling, so that only the MCTS extraction information set SEI can be needed to select the set to be extracted.
[0092] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, with blocks or devices corresponding to method steps or features of method steps. Similarly, aspects described in the context of a method step also represent a description of the corresponding block or item or function of the corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0093] The data stream of the present invention can be stored on a digital storage medium or can be transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0094] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, ERROM, EEPROM, or flash memory, on which electronically readable control signals are stored, which cooperates (or can cooperate) with a programmable computer system so that the respective methods are performed. Thus, the digital storage medium may be computer-readable.
[0095] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0096] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code may be stored on, for example, a machine-readable carrier.
[0097] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0098] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0099] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.
[0100] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or sequence of signals being adapted to be transmitted via a data communication connection, for example the Internet.
[0101] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0102] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0103] Further embodiments according to the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0104] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0105] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0106] The devices described herein, or any components of the devices described herein, may be implemented at least in part in hardware and / or software.
[0107] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0108] The methods described herein, or any components of the apparatus described herein, may be at least partially implemented by hardware and / or software.
[0109] The above-described embodiments are merely illustrative of the principles of the present invention. It is to be understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented as descriptions and explanations of the embodiments herein.
Claims
1. a sequence of slice portions (126), each of which encodes a respective slice (108) of a plurality of slices of an image (100) of a video (120), each of which includes a slice address indicating where the slice it encodes is located within an image area of the video; extraction information (132) indicating how to extract from the video data stream a reduced video data stream encoding the spatially smaller video (138) corresponding to the spatial section (110) of the video by limiting the video data stream to slice portions encoding any slice within the spatial section (110) and modifying the slice addresses to relate to reduced image areas of the spatially smaller video (138), the extraction information (132) comprising: first information (140) defining a spatial section (110) within the image region such that none of the plurality of slices crosses a boundary of the spatial section (110); and second information (142) indicating in the downscaled video data stream (136) where each of the slices is located within the downscaled region, signaling one of a plurality of options or explicitly signaling how to modify the slice address of the slice portion of each slice within the spatial section (110). A video data stream, including:
2. 2. The video data stream of claim 1, wherein the first information (140) defines the spatial section (110) within the image domain as consisting of a set of at least two sub-regions (110a, 110b, a, b, c, d) in which the video is encoded into the video data stream independently of outside each sub-region, and wherein none of the plurality of slices (108) crosses the boundaries of the at least two sub-regions.
3. a first option of the plurality of options indicating that the video is a panoramic video that displays a scene such that different edge portions of the images abut one another in a scene-like manner; a second option of the plurality of options indicating that the different edge portions are not scene-adjacent to one another; The at least two partial regions form first and second zones (110a, 110b) adjacent to different ones of the different edge portions, if the second information (142) signals the first option, the reduced image area is configured by combining the set of at least two sub-areas such that the first and second zones abut along the different edge portions. and 3. The video data stream of claim 2, wherein if the second information (142) signals the second option, the reduced image area is configured by combining a set of at least two sub-areas having the first and second zones with the different edge portions facing each other.
4. the second information (142) indicating how to re-sort the slice portions (126) of slices within the spatial section (110) relative to how the slice portions (126) are ordered in the sequence of slice portions of the video data stream when extracting the downsized video data stream (136) from the video data stream; and the slice addresses of each slice portion of the sequence of slice portions of the video data stream index linearly the coding start positions of the slices coded into each slice portion along a first coding scan order (106) that traverses the image area and codes the image (100) into the sequence of slice portions of the video data stream; the slice addresses of the slice portions of the sequence of slice portions in the video data stream are monotonically increasing, and 3. The video data stream of claim 2, wherein the modification of the slice addresses in the extraction of the reduced video data stream (136) from the video data stream is determined by sequentially arranging the slices encoded in the slice portions, to which the reduced video data stream (136) is limited, in a second encoding scanning order (112) across the reduced image area, the slices being reordered as signaled by the second information (142), and setting the slice addresses of the slice portions (126) to index the encoding start positions of the slices measured along the second encoding scanning order (112).
5. 5. The video data stream of claim 4, wherein the first coding scanning order (106) traverses the image regions in each of the sets of at least two sub-regions in a manner consistent with how the second coding scanning order (112) traverses each of the spatial regions.
6. 6. A video data stream as described in claim 4 or claim 5, wherein each of the sets of at least two sub-regions is represented by the first information (140) as a sub-array of rectangular tiles into which the image (100) is subdivided by rows and columns, and wherein the first and second encoding scan orders use a row-wise tile raster scan and completely traverse a current tile before proceeding to the next tile.
7. 2. The video data stream of claim 1, wherein the second information (142) signals an alternative to the slice address for each of a subset of the slice portions in which at least any of the slices in the spatial section are coded.
8. 8. The video data stream of claim 7, wherein the slice address of each slice portion of the sequence of slice portions of the video data stream relates to an image area of the video, and the alternative of the slice address relates to the reduced image area.
9. 8. A video data stream as claimed in claim 6 or claim 7, wherein the slice address of each slice portion of the sequence of slice portions of the video data stream indexes, one-dimensionally, the position of the encoding start of the slice encoded in each slice portion along a first encoding scan order across the image area in which the image is encoded into the sequence of slice portions of the video data stream, and the alternative of the slice address indexes, one-dimensionally, the position of the encoding start along a second encoding scan order across the reduced image area.
10. A video data stream as described in any of claims 7 to 9, wherein the subset of slice portions excludes a first slice portion of the sequence of slice portions in which any of the slices in the spatial section (110) is coded, and the second information (142) provides an alternative for each of the subset of slice portions sequentially according to the order in which the slice portions occur in the video data stream.
11. 10. A video data stream as described in any one of claims 7 to 9, wherein the first information (140) identifies the spatial section (110) within the image domain as consisting of a set of at least two sub-regions (110a, 110b) each having video encoded into the video data stream independently from outside each sub-region, and wherein none of the plurality of slices intersects a boundary of either the first or second sub-region.
12. 12. The video data stream of claim 11, wherein the subset of slice portions includes all slice portions of the sequence of slice portions in which any of the slices in the spatial section are encoded, and the second information provides alternatives for each of the slice portions sequentially according to the order in which the slice portions occur in the video data stream.
13. the second information (142) further signals how, upon extraction of the downsized video data stream from the video data stream, to re-sort the slice portions of the slices within the spatial section relative to how the slice portions are ordered in a sequence of slice portions of the video data stream; and 13. The video data stream of claim 12, wherein the second information (142) provides an alternative for each subset of the slice portions sequentially according to the order in which the re-sorted slice portions occur in the reduced video data stream.
14. the sequence of slice portions included in the video data stream includes slice portions encoding slices within the spatial section and slice portions encoding slices outside the spatial section, or A video data stream as described in any one of claims 1 to 13, wherein the sequence of slice portions included in the video data stream includes slice portions that encode slices within the spatial section, but does not include slice portions that encode slices outside the spatial section.
15. A video data stream according to any of claims 1 to 14, wherein said second information (142) defines the shape and size of said reduced image area.
16. 16. A video data stream as described in any one of claims 1 to 15, wherein the first information (140) defines the shape and size of the reduced image region by size parameters of parameter set alternatives to be used to substitute corresponding parameter sets in the video data stream when extracting the reduced video data stream (136) from the video data stream.
17. 1. An apparatus for generating a video data stream, comprising: providing a video data stream with a sequence of slice portions encoded in respective slices of a plurality of slices of an image of the video, each slice portion including a slice address indicating where, within an image area of the video, the slice in which the respective slice portion is encoded is located; configured to provide extraction information in the video data stream indicating how to extract from the video data stream a reduced video data stream encoded in spatially small video corresponding to the spatial section of video by limiting the video data stream to slice portions encoded in any slice within the spatial section and modifying the slice addresses to relate to a reduced image area of the spatially small video; The extracted information is first information defining the spatial section within the image region in which the video is encoded into the video data stream independently from outside the spatial section, First information that none of the slices crosses a boundary of the spatial section; and extraction information including second information indicating in the reduced video data stream where each of the slices is located within the reduced image area, signaling one of a plurality of options or explicitly signaling how to modify the slice addresses of the slice portions of each slice within the spatial section.
20. An apparatus configured to provide video data.
18. 1. An apparatus for extracting a reduced video data stream encoding a spatially smaller video from a video data stream encoding a video, the video data stream comprising a sequence of slice portions, each slice portion encoding a respective slice of a plurality of slices of an image of a video, each slice portion comprising a slice address indicating where the slice encoded by the respective slice portion is located within an image area of the video; The device comprises: configured to read extracted information from the video data stream; configured to derive a spatial section within the image region from the extracted information, wherein none of the plurality of slices crosses a boundary of the spatial section, and the reduced video data stream is limited to slice portions that encode any slice within the spatial section; An apparatus configured to modify the slice address of the slice portion of each slice in the spatial section using one of a plurality of options determined from the plurality of options using explicit signaling by the extracted information to indicate where, within the reduced video data stream, the respective slice is located in a reduced image area of the spatially smaller video.
19. The apparatus according to claim 17 or claim 18, wherein the video data stream is one according to any one of claims 2 to 16.
20. 1. A video data stream (332) encoding video, the video data stream including signaling (352) indicating that images (334) of the video display common scene content (304) at different resolutions in different spatial portions (354, 356) of the image.
21. 21. The video data stream of claim 20, wherein the signaling (352) includes a flag switchable between a first state and a second state, the first state indicating that the images (334) of the video include different spatial portions where the images display common scene content at different resolutions, and the second state indicating that the images of the video display each portion of the scene at only one resolution, and the flag is set to the first state.
22. 22. The video data stream of claim 21, wherein the flag is included in a sequence parameter set or a video parameter set of the video data stream.
23. 23. A video data stream according to any one of claims 20 to 22, wherein the video data stream is coded for each of a set of mutually non-overlapping spatial regions (360) independently from outside the respective spatial region (360).
24. A video data stream according to any of claims 20 to 23, wherein the signalling (352) comprises information about the counts of the different resolutions at which the images of the video display the common scene content at different spatial portions.
25. A video data stream as claimed in any one of claims 20 to 24, wherein the signalling (352) includes information about a spatial subdivision of the images of the video into sections (354, 356) each displaying the common scene content, at least two of which display the common scene content at different resolutions.
26. A video data stream as described in any one of claims 20 to 25, wherein the signaling (352) includes information regarding differences in resolution at which the images of the video display the common scene content in the different spatial portions (354, 356).
27. 27. A video data stream as claimed in any one of claims 20 to 26, wherein the video data stream includes extraction information indicating a set of extraction regions that allow extraction of a self-contained video data stream from the video data stream without re-encoding, the self-contained video data stream representing video limited to the image content of the video within each extraction region.
28. 28. The video data stream of claim 27, wherein the video data stream is encoded independently for each set of non-overlapping spatial regions (360) outside the respective spatial regions (360), and the extraction information indicates each set of extraction regions as one or more of the sets of spatial regions (360).
29. 29. A video data stream as described in claim 27 or claim 28, wherein the images of the video are spatially subdivided into an array of tiles and encoded continuously into the data stream, and wherein the spatial region (360) is a cluster of one or more tiles.
30. 30. A video data stream as described in any one of claims 27 to 29, wherein the video data stream includes, for each of at least one or more of the extraction regions, further signaling indicating whether the images of the video display common scene content at different resolutions within each of the extraction regions.
31. The further signaling may include, for each of the at least one or more extracted regions: a viewport orientation of a subregion within each of the extracted regions such that the common scene content is displayed at the highest resolution within each of the extracted regions; and / or the area share of the sub-area of each extracted area in which the common scene content is displayed at the highest resolution within each extracted area, relative to the total area of each extracted area; and / or a spatial subdivision of each of the extracted regions from its entire area into subregions in which the common scene content is displayed at different resolutions; 31. The video data stream of claim 30, further comprising:
32. 31. The video data stream of claim 30, wherein the further signaling is limited to an SEI message separate from the one or more SEI messages containing the extracted information.
33. Apparatus for processing a video data stream according to any of claims 20 to 32, said apparatus being configured to support predetermined processing tasks and to examine signalling to determine the performance or avoidance of performance of said predetermined processing tasks in said video data stream.
34. 34. The apparatus of claim 33, wherein the processing tasks include subjecting the video data stream to decoding and / or subjecting the video data stream to extraction of a reduced video data stream from the video data stream without re-encoding.
35. An apparatus for generating a video data stream according to any one of claims 20 to 32.
36. 1. A method for generating a video data stream, said method comprising: providing a sequence of slice portions in the video data stream, each slice portion encoding a respective slice of a plurality of slices of an image of the video, each slice portion including a slice address indicating where the slice encoded by each slice portion is located within an image area of the video; providing extraction information into the video data stream indicating how to extract from the video data stream a reduced video data stream encoding a spatially smaller video corresponding to the spatial section of video by limiting the video data stream to slice portions encoding any slices within the spatial section and modifying slice addresses to relate to reduced image areas of the spatially smaller video; The extracted information is first information defining the spatial section within the image region in which the video is encoded into the video data stream independently from outside the spatial section, wherein none of the plurality of slices crosses a boundary of the spatial section; and second information indicating in the downsized video data stream where the respective slices are located in the downsized image area, signaling one of a plurality of options or explicitly signaling how to modify the slice addresses of the slice portions of each slice in the spatial section; providing extracted information,
37. 1. A method for extracting a spatially smaller video-encoded downsized video data stream from a video-encoded video data stream, the video data stream comprising a sequence of slice portions, each slice portion encoding a respective slice of a plurality of slices of an image of a video, each slice portion comprising a slice address indicating where the slice encoded by the respective slice portion is located within an image area of the video; The method comprises: reading extracted information from the video data stream; deriving a spatial section within the image region from the extracted information, wherein none of the slices crosses a boundary of the spatial section, and the reduced video data stream is limited to slice portions that encode any slice within the spatial section; The method is configured to include a step of modifying the slice address of the slice portion of each slice in the spatial section to indicate in the reduced video data stream where the respective slice is located within the reduced image area of the spatially smaller video, using one of a plurality of options determined from the plurality of options using explicit signaling by the extracted information.
38. A method of processing a video data stream according to any one of claims 20 to 32, wherein the processing comprises predetermined processing tasks, the method comprising examining signalling to determine whether to perform or avoid performing the predetermined processing tasks on the video data stream.
39. A method for generating a video data stream according to any one of claims 20 to 32.
40. A computer program having a program code for performing the method according to any one of claims 36 to 39 when the computer program runs on a computer.
Citation Information
Patent Citations
Object data processor, object data recording device, data storage medium and data transmission structure
JP1998304353A
Parallelization of video decoder tiles
JP2014525151A