High - level Video Data Stream Extraction and Multiresolution Video Transmission

By modifying slice addresses and using additional signaling, the method enhances the efficiency of video data stream extraction and multi-resolution video scene display, addressing inefficiencies in existing technologies and enabling seamless adaptation and decoding.

JP7701417B2Active Publication Date: 2025-07-01DOLBY VIDEO COMPRESSION LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023120144
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-03-20
Filing Date
2023-07-24
Publication Date
2025-07-01
Estimated Expiration
2038-03-19

AI Technical Summary

Technical Problem

Existing video data stream extraction methods, such as those defined by the HEVC standard, are inefficient and cumbersome, particularly when extracting a spatial section of a video for different recipients or combining video scenes with varying resolutions, as they require complex reprocessing and lack efficient signaling for slice address modification.

Method used

The method involves modifying slice addresses within the video data stream to indicate the location of slices in a reduced image area, using additional signaling to guide the extraction process, allowing for efficient adaptation to different video content types and combining multiple resolution versions into a single stream.

Benefits of technology

This approach reduces processing complexity and enables efficient extraction and display of spatial sections and multi-resolution video scenes, allowing for seamless adaptation and decoding by legacy devices without the need for complex reprocessing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701417000001
    Figure 0007701417000001
  • Figure 0007701417000002
    Figure 0007701417000002
  • Figure 0007701417000003
    Figure 0007701417000003
Patent Text Reader

Abstract

To provide a device and a method for extracting a video data stream more efficiently.SOLUTION: In video data stream extraction processing, each slice portion 126 of a video data stream 124 includes: a sequence of slice portions, each slice portion including a slice address indicating a location where, in a picture area of video 102, a slice 108 is located which each slice portion has encoded thereinto; and extraction information used in extracting a video data stream 136 reduced from the video data stream. The extraction information includes many extraction information sets for different versions of the reduced video data stream. Each extraction information includes first information 140 which identifies a spatial section 110 within the picture area, and second information 142 which explicitly signals an alternative slice address for replacing the slice address in the slice portion of the slice in the spatial section.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the concept of extracting a video data stream, i.e., extracting a reduced video data stream from a properly prepared video data stream such that the reduced video data stream can encode a spatially smaller video corresponding to a spatial section of the video encoded in the original video data stream. Further, it relates to the transmission of different video versions of one scene, where the versions differ in the resolution or fidelity of the scene.

Background Art

[0002] The HEVC standard [1] defines a hybrid video codec that enables the definition of a sub-array of rectangular tiles of an image when the video codec conforms to some encoding constraints so that a smaller or reduced video data stream can be easily extracted from the entire video data stream, i.e., without requantization and without having to redo motion compensation. As outlined in [2], it is assumed that the HEVC standard syntax that can guide the extraction process of the receiver of the video data stream will be added.

[0003] However, the extraction process needs to be made more efficient.

[0004] Application areas where video data extraction may be applied relate to the transmission or provision of multiple versions of one video scene with different resolutions of the scene. An efficient way to install such transmissions or to provide different resolution versions is advantageous.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Accordingly, a first object of the present invention is to provide a more efficient concept of video data stream extraction, i.e., for example, different types of video can more efficiently process unknown type of video content for different recipients, such as, for example, projection from a viewport to a screen, or to reduce the complexity of the extraction process. This object is achieved by the subject matter of the independent claims of the present application according to the first aspect.

Means for Solving the Problems

[0006] In particular, according to a first aspect of the present application, regarding a method of modifying the slice address of the slice portion of each slice in an extractable spatial section so as to indicate the location where each slice is located in a reduced (extracted) image area within a reduced video data stream, information for signaling one of a plurality of options to the extraction information in the video data stream, or by providing a method of explicitly signaling, the extraction of the video data stream becomes more efficient. In other words, the second information provides information to the extraction site of the video data stream that guides the extraction process regarding the composition of the spatially small video image of the reduced (extracted) video data stream based on the spatial section of the original video. And thus, it reduces the extraction process or enables adaptation to greater variations in the scene type transmitted within the video data stream. Regarding the latter problem, for example, the second information is the relative arrangement of these portions of the spatial section within the original video, or relative from the perspective of the coding order, of the pictures of the spatially small video in orderAs a result of squeezing together potentially discontinuous portions of a spatial section while maintaining the order, it is possible to address various opportunities that should be advantageous. Regarding the latter problem, for example, the second information is not only the result of squeezing together potentially discontinuous portions of a spatial section, but also the relative arrangement of these portions of a spatial section within the original video, or the maintenance of the relative order from the perspective of the encoding order, and it is possible to address various opportunities that should be advantageous. For example, in a spatial section composed of zones that abut different portions along the periphery of an original image showing a scene of a seam interface of a projection from a panoramic machine to a screen, the arrangement of the zones of the spatial section in the small images of the extracted stream should be different from the case of a non-panoramic type of image type, but the recipient may not know about that type either. Additionally or separately, it is a cumbersome task to modify the slice addresses of the extracted slice portions, and it may be alleviated, for example, by explicitly transmitting information regarding the modification method in the form of alternative slice addresses.

[0007] Another object of the present invention is to provide a concept of juxtaposing different versions of a video scene with different resolutions of a scene to more efficiently provide it to a recipient.

[0008] This object is achieved by the subject matter of the independent claims pending in the second aspect of the present application.

[0009] In particular, according to the second aspect of the present application, the juxtaposition of several versions of video scenes with different scene resolutions is more efficiently displayed by combining these versions into one video encoded in one video data stream, and in this video data stream, a signaling is provided that indicates that the video images of the video display common scene content in different spatial parts of the image at different resolutions. A recipient of the video data stream can thus, based on the signaling, recognize whether the video content conveyed by the video data stream relates to a spatial parallel collection of several versions of the scene content at different scene resolutions. Depending on the function of the receiving site, attempts to decode the video data stream can be suppressed, or the processing of the video data stream can be adapted to the analysis of the signaling.

[0010] Advantageous implementations of the embodiments of the above-described aspects are the subject matter of the dependent claims. Preferred embodiments of the present application are described below with respect to the drawings.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

[0012] The following description begins with an explanation of the first aspect of the present application and is then followed by an explanation of the second aspect of the present application. More precisely, with respect to the first aspect of the present application, the explanation begins with an overview of the underlying technical problem in order to motivate the advantages and underlying concepts of the embodiments of the first aspect described below. For the second aspect, the order of explanation is selected in the same way.

[0013] In panorama or 360 - video applications, it is common to need to present only a sub - section of the screen to the user. Using certain codec tools such as Motion Constrained Tile Sets (MCTS), it is possible to extract the encoded data corresponding to the desired image sub - section within the compression domain and form a compliant bitstream. And it can be decoded by a legacy decoder device that does not support MCTS decoding from the complete image bitstream. And it is characterized as being at a lower level compared to the decoder required for complete image decoding.

[0014] As an example and for reference, the signaling included in the HEVC codec is in the following locations. · Refer to [1]. Specify the temporary MCTS SEI messages in Sections D2.29 and E.2.29, whereby the encoder can signal the specified list of rectangles. Each is defined by the top-left and bottom-right tile indices and belongs to MCTS. · Refer to [2]. Provide additional information such as parameter sets and nested SEI messages that facilitate the task of extracting MCTS as a compliant HEVC bitstream, to be added to the next version of [1].

[0015] As can be seen from [1] and [2], the extraction procedure includes the adjustment of the slice address indicated in the slice header of the relevant slice executed on the extraction device.

[0016] Figure 1 shows an example of MCTS extraction. Figure 1 shows a video data stream, i.e., an image encoded in an HEVC video data stream. Image 100 is subdivided into CTBs, i.e., coding tree blocks which are the units in which Image 100 is encoded. In the example of Figure 1, Image 100 is subdivided into 16×6 CTBs, although of course the number of rows and columns of CTBs is not important. Reference numeral 102 typically represents such a CTB. In units of these CTBs 102, Image 100 is further subdivided into tiles, i.e., an array of m×n tiles, and Figure 1 shows an exemplary case where m = 8 and n = 4. In each tile, reference numeral 104 is used to typically represent one such tile. Thus, each tile is a rectangular cluster or subarray of CTBs 102. For illustrative purposes only, Figure 1 shows that the tiles 104 can be of different sizes, or alternatively, that the rows of tiles can be columns of tiles with different heights and different widths from each other.

[0017] As is known in the art, the tiling of an image, i.e., the subdivision of the image 100 into tiles 104, affects the encoding order 106 in which the image content of the image 100 is encoded into a video data stream. In particular, the tiles 104 are traversed one after another along the order of the tiles, i.e., in a raster scan order for each row of tiles. In other words, all CTBs 102 within one tile 104 are first encoded or traversed by the encoding order 106 before the encoding order proceeds to the next tile 104. Within each tile 102, the CTBs are also encoded using a raster scan order, i.e., a raster scan order in the row direction. Along the encoding order 106, the encoding of the image 100 into the video data stream is subdivided to yield so-called slice parts. In other words, the slices of the image 100 that the successive parts of the encoding order 106 traverse are encoded as units into the video data stream to form slice parts. In FIG. 1, each tile is assumed to be within a single slice, or in the terms of HEVC, within a slice segment, but this is merely an example and can be created in different ways. Typically, one slice 108, or for HEVC, one slice segment is typically denoted by reference numeral 108 in FIG. 1 and conforms to or fits the corresponding tile 104.

[0018] As far as the encoding of the image 100 into the video data stream is concerned, it should be noted that this encoding utilizes spatial prediction, temporal prediction, context derivation for entropy encoding, motion compensation for temporal prediction, transformation of the prediction residue and / or quantization of the prediction residue. The encoding order 106 not only affects slicing but also defines the availability of the reference base for spatial prediction and / or context derivation. For example, only adjacent parts preceding in the encoding order 106 are available. Tiling not only affects the encoding order 106 but also limits the encoding interdependencies within the image 100. For example, spatial prediction and / or context derivation are limited to referring only to parts within the current tile 104. Parts outside the current tile are not referred to in spatial prediction and / or context derivation.

[0019] FIG. 1 shows a further specific region of the image 100, namely the so-called MCTS, i.e., a spatial section 110 within the image region of the image 100 from which a video of which the image 100 is a part can be extracted. An enlarged view of the section 110 is shown on the right side of FIG. 10. The MCTS 110 in FIG. 1 is composed of a set of tiles 104. The tiles in the section 110 are named a, b, c, and d, respectively. The fact that the spatial section 110 is extractable involves further restrictions regarding the encoding into the video data stream of the video. In particular, the division of the video image of the image 100 shown in FIG. 1 is adopted by other images of the video, and for this image sequence, the image content within tiles a, b, c, and d is encoded such that encoding interdependencies remain within the spatial section 110 even when referring from one image to another. In other words, for example, temporal prediction and temporal context derivation are restricted to fit within the region of the spatial section 110.

[0020] An interesting point when encoding a video that is part of the image 100 into the video data stream is the fact that a slice address is provided to the slice 108, and this slice address indicates the start of its encoding, i.e., its position, in the encoded image region. The slice address is assigned along the encoding order 106. For example, the slice address indicates the CTB rank along the encoding order 106 at which the encoding of each slice starts. For example, within the data streams encoding the video and the image 100 respectively, the slice part carrying the slice that coincides with tile a has a slice address of 7 as the 7th CTB of the encoding order 106 indicating the first CTB of the encoding order 106 within tile a. Similarly, the slice addresses within the slice parts carrying the slices related to tiles b, c, and d are 9, 29, and 33, respectively.

[0021] On the right side of FIG. 1, it shows the slice addresses where the recipient of the reduced or extracted video data stream assigns two slices corresponding to tiles a, b, c, and d. In other words, FIG. 1 uses the numbers 0, 2, 4, and 8 on the right side to show the slice addresses assigned by the recipient of the reduced or extracted video data stream, which are obtained from the original video data stream representing the video including the entire image 100 by extraction regarding the spatial section 110, i.e., by MCTS extraction. In the reduced or extracted video data stream, the slice portions encoded in slice 100 of tiles a, b, c, and d are arranged in the encoding order 106 as if they exactly existed in the original video data stream taken out during the extraction process. In particular, this recipient arranges the image content in the form of CTBs reconstructed from the sequence of slice portions of the reduced or extracted video data stream. That is, the slices regarding tiles a, b, c, and d along the encoding order 112 that traverses the spatial section 110 in the same way as the encoding order 106 for the entire image 100, i.e., by traversing the spatial section 110 in raster scan order for each tile within the tile. Also, the scanning of CTBs within each tile is performed along the raster scan order and then proceeds to the next tile. The relative positions of tiles a, b, c, and d are maintained. That is, the spatial section 110 as shown on the right side of FIG. 1 maintains the relative positions among tiles a, b, c, and d that occurred in image 100. The slice addresses of the slices resulting from determining tiles a, b, c, and d using the encoding order 112 are 0, 2, 4, and 8 respectively. Therefore, the recipient can reconstruct a smaller video based on the reduced or extracted video data stream showing the spatial section 110 presented as a self - contained image on the right side.

[0022] Summarizing the description of FIG. 1 thus far, FIG. 1 shows or shows the adjusted slice address and CTB units after extraction by using the numbers in the upper left corner of each slice 108. To perform the extraction, the extraction site or device needs to analyze the original video data stream with respect to the size of the CTB, i.e., the maximum CTB size 114, and the parameters indicating the number and size of tile columns and tile rows in the image 100. Further, the nested MCTS-specific sequences and sets of image parameters are examined to derive the output placement of the tiles therefrom. In FIG. 1, tiles a, b, c, and d within the spatial section 110 maintain their relative placement. In short, overall, the analysis and examination performed by the extraction device in accordance with the above and the MCTS instructions of the HEVC data stream require dedicated sophisticated logic for deriving the slice address of the reconstructed slice 108 from the parameters listed above. This dedicated and sophisticated logic not only has runtime drawbacks but also incurs additional implementation costs.

[0023] The embodiments described below thus use additional signaling in the video data stream and corresponding processing steps on the extraction information generation side and the extraction side to provide information that is readily available for a particular purpose of extraction, thereby reducing the overall processing burden of the extraction device described above. Additionally or alternatively, some of the embodiments described below use additional signaling to guide the extraction process in such a way that more effective processing of different types of video content is achieved.

[0024] Based on FIG. 2, a general concept will be described first. Later, this general description of the operating modes of the individual entities participating in the overall process shown in FIG. 2 will be further clarified in different ways according to different embodiments hereinafter. For the sake of easy understanding of the entities and blocks shown therein, they are described together in one figure, but it should be noted that each is related to a self - contained device that individually inherits the functions providing an overview of the advantages of FIG. 2. More precisely, FIG. 2 shows a generation process that generates a video data stream and provides extraction information to such a video data stream. Following the extraction process itself, decoding of the extracted or reduced video data stream and participating devices is performed. The operating modes of these devices or the performance of individual tasks and steps are in accordance with the embodiments described herein. According to a specific implementation example first described with respect to FIG. 2 and as further outlined hereinafter, the processing overhead associated with the extraction tasks of the extraction device is reduced. According to further embodiments, the processing of various different types of video content within the original video is additionally or alternatively reduced.

[0025] At the top of FIG. 2, an original video indicated by reference numeral 120 is shown. This video 120 serves the same role as the image 100 in FIG. 1, so it is composed of a series of images, one of which is shown using reference numeral 100. That is, it is an image showing an image area from which the spatial section 110 will later be cut out by extraction of the video data stream. However, it should be understood that the tile sub - division described above with respect to FIG. 1 need not be used in the video encoding that forms the basis of the process shown in FIG. 2. Rather, tiles and CTBs simply represent semantic entities in video encoding that are optional as far as the embodiments of FIG. 2 are concerned.

[0026] FIG. 2 shows that video 102 is the object of video encoding by video encoding core 122. The video encoding executed by video encoding core 122 changes video 120 into video data stream 124, for example, using hybrid video encoding. That is, video encoding core 122 uses block-based predictive encoding that encodes individual image blocks of the image of video 120 using, for example, one of several supported predictive mode encodings of prediction residuals. The predictive mode may include, for example, spatial and temporal prediction. Temporal prediction may include motion compensation, that is, determination of a motion field representing motion vectors of temporally predicted blocks, and its transmission by data stream 124. The prediction residuals may be transform-encoded. That is, several spectral decompositions can be applied to the prediction residuals, and the resulting spectral coefficients can be quantized and reversibly encoded into data stream 124, for example, using entropy encoding. Next, the entropy encoding may use context adaptability. That is, this context derivation may determine a context that depends on the spatial neighborhood and / or the temporal neighborhood. As described above, the encoding may be based on encoding order 106 that limits the encoding dependencies such that only the parts of the video that the encoding order 106 has already passed through can be used as a basis or reference for encoding the current part of video 120. The encoding order 106 traverses video 120 for each image, but is not necessarily in the presentation time order of the images. Within an image such as image 100, video encoding core 120 subdivides the encoded data obtained by video encoding, whereby image 100 is subdivided into slices 108 that respectively correspond to corresponding slice portions 126 of video data stream 124. Within data stream 124, slice portions 126 form a sequence of slice portions that follow each other in the order in which the corresponding slices 108 are traversed by the encoding order 106 of image 100.

[0027] As also shown in FIG. 2, the video encoding core 122 provides or encodes a slice address for each slice portion 126. The slice addresses are shown in capital letters in FIG. 2 for illustration purposes. As described with respect to FIG. 1, the slice addresses may be determined in some suitable unit, such as units of CTBs, linearly along the encoding order 106. Alternatively, they may be determined differently for some predetermined points within the image area already used by the image of the video 120, such as the upper left corner of the image 100, for example.

[0028] Thus, the video encoding core 122 receives the video 120 and outputs a video data stream 124.

[0029] As already outlined above, the video data stream generated according to FIG. 2 is extractable as far as the spatial section 110 is concerned, and thus the video encoding core 120 appropriately adapts the video encoding process. For this purpose, the video encoding core 122, for example, by means of spatial prediction, temporal prediction, or context derivation, etc., restricts the inter-encoding dependencies so that the parts within the spatial section 110 are encoded in such a way that they do not depend on the parts outside the section 110. The slice 108 does not cross the boundary of the section 110. Thus, each slice 108 is either completely within the section 110 or completely outside the section 110. It should be noted that the video encoding core 122 can follow not only one spatial section 110 but also several spatial sections. These spatial sections may intersect each other, i.e., the same spatial section may partially overlap, or one spatial section may be completely present within another spatial section. By means of these measures, as will be explained in detail later, it is possible to extract from the video data stream 124 a reduced or extracted video data stream of an image smaller than the image of the video 120. That is, there is no need to re-execute complex tasks such as motion compensation, quantization, and / or entropy encoding for an image showing the content within the spatial section 110 without simply re-encoding it.

[0030] The video data stream 124 is received by the video data stream generator 128. In particular, according to the embodiment shown in FIG. 2, the video data stream generator 128 includes a reception interface 130 that receives the video data stream 124 prepared by the video encoding core 122. According to an alternative, it should be noted that the video encoding core 122 can be included in the video data stream generator 128, thereby replacing the interface 130.

[0031] The video data stream generator 128 provides the extraction information 132 to the video data stream 124. In FIG. 2, the resulting video data stream output by the video data stream generator 128 is indicated using the reference numeral 124'. The extraction information 132 indicates how to extract, from the video data stream 124', an encoded reduced or extracted video data stream 136 for the extraction device indicated by the reference numeral 134 in FIG. 2, which includes a spatially small video 138 corresponding to the spatial section 110. The extraction information 132 includes a first information 140 that defines the spatial section 110 within the image area transmitted by the image 100, and a second information 142 that indicates one of a plurality of options regarding how to modify the slice address of the slice portion 126 of each slice 108 entering the spatial section 110 so as to indicate where each slice is located within the reduced image area of the image 144 of the video 138 within the reduced video data stream 136.

[0032] In other words, the video data stream generator 128 simply accompanies, i.e., adds something to, the video data stream 124 in order to produce the video data stream 124', i.e., the extraction information 132. This extraction information 132 is for the purpose of guiding the extraction device 134 that receives the video data stream 124' when extracting, in particular, the video data stream 136 that has been reduced or extracted with respect to the section 110 from the video data stream 124'. The first information 140 defines the spatial section 110, i.e., its location within the image regions of the video 120 and the image 100 respectively, and optionally the size and shape of the image region of the image 144. As shown in FIG. 2, this section 110 does not necessarily have to be rectangular, convex, or even a connected region. For example, in the example of FIG. 2, the section 110 is composed of two separate sub-regions 110a and 110b. Furthermore, the first information 140 can include hints regarding how the extraction device 134 modifies or replaces part of the encoding parameters or part of the data stream 124 or 124' respectively. For example, it could be an image size parameter adapted to reflect the change in the image region from the image 100 to the image 144 upon extraction. In particular, the first information 140 can include alternative or modification instructions regarding the parameter set of the video data stream 124' that are applied by the extraction device 134 in the extraction process in order to correspondingly change or replace the corresponding parameter set included in the video data stream 124' and pass it on to the reduced or extracted video data stream 136.

[0033] In other words, the extraction device 134 receives the video data stream 124', reads the extraction information 132 from the video data stream 124', and derives the spatial section 110, i.e., its position and arrangement within the image area of the video 120, from the extraction information. That is, it is based on the first information 140. Accordingly, the extraction device 130 identifies the slice portion 126 encoded in the slice entering the section 110 based on the first information 140. And thus, while the slice portions 126 related to the slices outside the section 110 are dropped by the extraction device 134, they are passed on to the reduced or extracted video data stream 136. Further, as outlined above, the extraction device 134 may use the information 140 to correctly set one or more parameter sets in the data stream 124' previously, or adopt it in the reduced or extracted video data stream 136, i.e., by modification or replacement. Accordingly, one or more parameter sets may relate to the image size parameter set to a size corresponding to the total area of the section 110 region according to the information 140, i.e., if the section 110 is not the connected region as exemplarily shown in FIG. 2, the total area of all parts 110a and 110b of the section 110. By dropping the slice portions dependent on the section 110 and adapting the parameter sets, the video data stream 124' is limited to the section 110. Further, the extraction device 134 modifies the slice addresses of the slice portions 126 in the reduced or extracted video data stream 136. These slices are shown using the header of FIG. 2. That is, the hatched slice portions 126 are the slices that enter the section 110 and are thus the slices extracted or passed on respectively.

[0034] The information 142 can be imagined not only in the situation of being added to the complete video data stream 124', but it should be noted that the sequence of slice parts included therein also includes slice parts encoded for slices outside the spatial section, just as it includes slice parts encoded for slices within the spatial section. Rather, the data stream containing the information 142 may already have been removed such that the sequence of slice parts included in the video data stream includes slice parts encoded for slices within the spatial section. However, there are no slice parts encoded for slices outside the spatial section.

[0035] In the following, different examples for embedding the second information 142 into the data stream 124' and their processing are presented. Generally, the second information 142 is conveyed as signaling within the data stream 124' that indicates a hint regarding how to perform slice address modification, either explicitly or in the form of a selection of one of several options. In other words, the second information 142 is conveyed in the form of one or more syntax elements, and its possible values can distinguish, for example, between signaling that explicitly notifies an alternative of the slice address or, together, a number of possibilities of associating the slice address for each slice part of the video data stream 136 by selecting settings of one or more syntax elements of the data stream. Note, however, that the number of meaningful or possible settings of the one or more syntax elements that embody the second information 142 depends respectively on the way the video 120 is encoded into the video data stream 124 and the selection of the section 110. For example, assume that the section 110 is a rectangular connection area within the image 100 and the video encoding core 120 performs encoding with respect to this section without further restricting the encoding as far as the inside of the section 110 is concerned. A configuration of the section 110 by two or more regions 110a and 110b will not be applicable. That is, only dependencies external to the section 110 are suppressed. In this case, the section 110 must be mapped without being modified to the image area of the image 144 of the video 138, i.e., without scrambling positions of any sub-regions of the section 110. Also, the assignment of addresses α and β to the slice parts of the section 110 that make up the slice 110 will be uniquely determined by placing the inside of the section 110 as it is into the image area of the image 144. In this case, the setting of the information 142 generated by the video data stream generator 128 is unique. That is, there will be no other options for the video data stream generator 128 to set the information 142 in this way. However, there may be other signaling options available for the information 142 from the perspective of encoding.However, even in this no-alternative case, for example, the signaling 142 that explicitly indicates the native slice address modification has the advantage that the extraction device 134 does not have to perform the aforementioned cumbersome task of itself determining the slice addresses α and β for the slice portion 126 adopted from the stream 124'. Rather, it simply derives how to modify the slice address of the slice portion 126 from the information 142.

[0036] Depending on different embodiments regarding the nature of the information 142 to be further outlined below, the extraction device 134 either preserves or maintains the order in which the slice portion 126 is passed from the stream 124' to the reduced or extracted stream 136, or modifies the order in the manner defined by the information 142. In either case, the reduced or extracted data stream 136 output by the extraction device 134 can be decoded by a normal decoder 146. The decoder 146 receives the extracted video data stream 136 and decodes the video 138 therefrom. The image 134 is smaller than the image of the video 120 such as the image 100. And the image area is filled by arranging the decoded slice 108 from the slice portion 126 in the video data stream 136 in a manner defined by the slice addresses α and β transmitted within the slice portion 126 in the video data stream 136.

[0037] That is, so far, FIG. 2 has been described in a manner that is adapted to various embodiments regarding the exact nature of the second information 142, which will be described in more detail below.

[0038] The embodiments described herein use an explicit signaling of the slice address to be used by the extraction device 134 when modifying the slice address of the slice portion 126 inherited from stream 142 to stream 136. The embodiments described below use signaling 142 that enables signaling to the extraction device 134 one of several permitted options of how to modify the slice address. For example, the resulting option allowance of section 110 encoded in a way that limits the coding interdependencies inside section 110 so as not to cross the spatial boundaries of section 110, in turn, as shown in FIG. 1, divides section 110 into two or more regions such as 110a and 110c or tiles a, b, c, d. The latter embodiment may further include the extraction device 134 to perform the cumbersome task of calculating the address by itself. However, based on the corresponding extracted or reduced video data stream 136, it enables efficient processing of different types of image content in the original video 120 so that meaningful video 138 is obtained on the receiving side.

[0039] That is, as outlined above with respect to FIG. 1, the cumbersome task of slice address determination in the extraction device 134 is alleviated according to the embodiments of the present application by the explicit transmission of the method of modifying the address in the extraction process by the second information 142. A specific example of the syntax that can be used for this purpose is shown below.

[0040] In particular, by using the information 142 to include a list of slice address substitutions in the stream 124' in the same order as the slice portion 126 is carried in the bitstream 124', the new slice address used in the slice header of the extracted MCTS can be explicitly signaled. For example, refer to the example of FIG. 1. Here, the information 142 is an explicit signaling of slice addresses following the order of the slice addresses 124' in the bitstream. Again, since the slice 108 and the corresponding slice portion 126 are carried over in the extraction process, the order of the extracted video data stream 136 corresponds to the order in which these slice portions 126 were included in the video data stream 124'. According to an example of the subsequent syntax, the information 142 explicitly notifies the slice address in a way that starts after the second slice or slice portion 126. In the case of FIG. 1, this explicit signaling would correspond to the second information 142 that indicates or signals the list {2, 4, 8}. An exemplary syntax of this embodiment is shown in the syntax example of FIG. 3, and the highlight shows the corresponding addition of the explicit signaling 142 in addition to the MCTS extraction information SEI known from [2].

[0041] The semantics are shown below.

[0042] num_associated_slices_minus2 [i] + 2 indicates the number of slices containing an MCTS with an mcts identifier equal to any value in the list mcts_identifier [i] [j]. The value of num_extraction_info_sets_minus1 [i] must be in the range from 0 to 32 -2.

[0043] output_slice_address [i] [j] identifies the slice address of the j-th slice in the bitstream order of the MCTS to which the mcts identifier is equal to any value within the list mcts_identifier [i] [j]. The value of output_slice_address [i] [j] must be in the range of 0 to 2 32 -2.

[0044] Note that in addition to the presence of information 142 in the MCTS extraction information SEI or MTCS-related information 140, it can be controlled by a flag in the data stream. This flag can be named, for example, slice_reordering_enabled_flag. When set, in addition to information 140, information 142 such as num_associated_slices_minus2 and output_slice_address exists, and if not, information 142 does not exist, and the relative positioning of the slices is observed or processed in the extraction process.

[0045] Furthermore, using the H.265 / HEVC nomenclature, it should be noted that the "_segment_" part of the syntax element name used in Figure 3 can be replaced with "_segment_address_", but the technical content is the same.

[0046] And further, num_associated_slices_minus2 implies that information 142 indicates the number of slices in section 110 in the form of an integer showing this number as the difference from 2, but it should be noted that the number of slices in section 110 can also be signaled directly in the data stream or as the difference from 1. In the latter case, for example, num_associated_slices_minus1 would be used as the syntax element name. It should be noted that the number of slices in any section 110 can be, for example, one.

[0047] In addition to the MCTS extraction process previously anticipated in [2], the additional processing steps are associated with explicit signaling by information 142, as embodied in FIG. 3. These additional processing steps facilitate the derivation of slice addresses within the extraction process executed by extraction device 134, and the following outline of this extraction process indicates where facilitation occurs by underlining.

[0048] Input to the sub-bitstream MCTS extraction process is the bitstream inBitstream, the target MCTS identifier mctsIdTarget, the target MCTS extraction information set identifier mctsEISIdTarget, and the target maximum TemporalId value mctsTIdTarget. The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream. It is a requirement for bitstream conformance of the output bitstream that any output sub-bitstream that is the output of the process specified with the bitstream in this section is a conforming bitstream.

[0049] The output sub-bitstream is derived as follows. - The bitstream outBitstream is set identical to the bitstream inBitstream. - Remove from outBitstream all NAL units having a TemporalId greater than mBitsTIdTarget. - For each remaining VCL NAL of each access unit unit of outBitstream, adjust the slice segment header as follows. - For the first VCL NAL unit, set the value of first_slice_segment_in_pic_flag to 1. Otherwise, set it to 0. - Set the value of slice_segment_address of the NAL unit (slice) other than the first one starting from the second in bitstream order according to the list output_slice_address[i][j].

[0050] The variations of the embodiments described with respect to FIGS. 1 to 3 reduce the cumbersome tasks of the slice address determination and extraction process executed by the extraction device 134 by using the information 142 as an explicit signaling of the slice address. According to a specific example of FIG. 3, the information includes an alternative slice address 143 for each of the second and subsequent slice portions 126 in the slice portion order while maintaining the slice order when taking over the slice portion 126 related to the slice 108 in the section 110 from the data stream 124' to the data stream 136. The replacement of the slice address is related to the assignment of the one-dimensional slice address in the image area of the image 144 of the video 138, respectively. Using the order 112 and using the sequence of slices 108 obtained from the sequence of the taken-over slice portions does not conflict with simply paving the image area of the image 144 along the order 112. It should be noted that the explicit signaling can also be applied to the first slice address, i.e., the slice address of the first slice portion 126, and will be further mentioned below. Even in the latter case, there may be an alternative 143 included in the information 142. Such signaling for the information 142 may also enable the placement of the corresponding slice 108 first in the order within the stream 124'. The slice portion 126 of the slice portion(s) 126 carrying the slice 108 in the section 110 may be anywhere other than the start position of the coding order 112, which may be the upper left image corner as shown in FIG. 1. In the rearrangement of the cross-section areas 110a and 110b, if such a possibility exists or is permitted, the explicit signaling of the slice address can be used to increase the degree of freedom. For example, when changing the example shown in FIG. 3, the information 142 also explicitly notifies the slice address replacement 143 for the first slice portion 126 in the data stream 136, i.e., for α and β in the case of FIG. 2. Next, the signaling 142 enables the distinction between the two permitted or available arrangements of the section areas 110a and 110b in the output image of the video 138, i.e., the one where the section area 110a is arranged on the left side of the image.As a result, the slice 108 corresponding to the first transmitted slice portion 126 in the data stream 136 at the start position of the encoding order 112 remains, and as far as the video data stream 124 is concerned, the order among the slices 108 is maintained as compared with the slice order of the image 100 of the video 120. And the section area 120a is arranged on the right side of the image 144, whereby the order of the slices 108 in the image 144 traversed by the encoding order 112 is changed as compared with the order in which the same slice is traversed in the original video in the video data stream 124 by the encoding order 106. The signaling 142 explicitly indicates the slice addresses used for correction by the extraction device 134 as a list of slice addresses 143 ordered or assigned to the slice portions 126 in the video data stream 124'. That is, the information 142 successively indicates the slice addresses of each slice in the section 110 in the order in which these slices 108 are traversed by the encoding order 106, and this explicit signaling may result in a replacement such that the section areas 110a and 110b of the section 110 change the order traversed by the order 112 as compared with the order traversed by the original encoding order 106. The slice portions 126 inherited from the stream 124' to the stream 136 are accordingly reordered by the extraction device 134, that is, made to conform to the order in which the slices 108 of the extracted slice portions 126 are successively traversed by the order 112. In accordance with the slice portions 126 inherited from the video data stream 124', the compliance of the reduced or extracted video data stream 136 needs to strictly follow each other along the encoding order 112. That is, it is maintained with slice addresses α and β that monotonically increase as corrected by the extraction device 134 in the extraction process. Therefore, the extraction device 134 corrects the order between the inherited slice portions 126 so as to order the inherited slice portions 126 according to the order of the slices 108 encoded therein along the encoding order 112.

[0051] The latter embodiment, i.e., the possibility of repositioning slice 108 of the inherited slice portion 126, is utilized according to a further variant of the description of FIG. 2. Here, the second information 142 signals a reordering of the order between the slice portions 126 encoded in any slice 108 within section 110. An explicit signaling slice address substitution 143 in a way that results in repositioning of the slice portions is one possibility of the presently described embodiment. However, the repositioning of the slice portions 126 within the extracted or reduced video data stream 136 may be notified by the information 142 in a different way. This embodiment finally arranges the reconstructed slices reconstructed from the slice portions 126 along the encoding order 112 exactly at the decoder 146, which will use, for example, a tile raster scan order, thereby filling the image area of the image 144. The repositioning signaled by the signaling 142 is selected to reorder or change the order between the extracted or inherited slice portions 126, and the arrangement by the decoder 146 may change the order compared to the case where the section areas such as areas 110a and 110b do not change the order of the slice portions. If the repositioning signaled by the information 142 leaves the order that was in the original data stream 124', the sub-areas 110a and 110b can maintain their relative positions that were in the original image area of the image 110.

[0052] To explain the present variant, refer to FIG. 4. FIG. 4 shows the image 110 of the original video and the image area of the extracted video's image 144. Further, FIG. 4 shows an exemplary tile division into tiles 104, and two separate cross-section areas or zones 110a and 110b, i.e., opposite sides 150 of the image 110 r and 150 l and an exemplary extraction section 1110 consisting of the cross-section areas 110a and 110b that abut on the sides 150 r and 150 lAs far as the extraction direction is concerned, i.e., their positions coincide along the vertical direction. In particular, FIG. 4 shows that the image content of image 110, and thus the video to which image 110 belongs, is of a specific type, i.e., a panoramic video. Thus, side surfaces 150 r and 150 l form the scene when projecting the 3D scene onto the image area of image 110. Below image 110, FIG. 4 shows two options for how to place sections 110a and 110b within the output image area of the extracted video's image 144. Again, tile names are used in FIG. 4 to explain the two options. The two options are the result of video 210 being encoded in stream 124 in such a way that it is done independently from the outside for each of the two regions 110a and 110b. In addition to the two allowed choices shown in FIG. 4, each of zones 11a and 110b is subdivided into two tiles, and within each of those tiles, in turn, if the video is encoded independently from the outside in stream 124, i.e., if each tile in FIG. 4 is a separately extractable part or a part encoded independently (with respect to spatial and temporal interdependencies), it should be noted that there could be two more options. Next, the two options correspond to scrambling the tiles a, b, c, d of section 110 differently.

[0053] Again, with respect to FIG. 4, the embodiment being described now is for the purpose that the NAL unit of slice 108 or data stream 124' carrying slice 108 has second information 142 for changing the order, and during the extraction process by extraction device 134, it is extracted from data stream 124' or transferred or adopted or written into reduced video data stream 136. FIG. 4 shows a case where the desired image subsection 110 consists of non-adjacent tile or section regions 110a and 110b within the image plane in which image 110 spreads. This fully encoded image plane of image 110 is shown at the top of FIG. 4 having a tiled boundary with a desired MCTS 110 consisting of two rectangles 110a and 110b including tiles a, b, c, and d. As far as the scene content is concerned, or due to the fact that the video content shown in FIG. 4 is panoramic video content, when image 110 covers the surroundings of the camera over 360° by the equirectangular projection method exemplarily shown here, the desired MCTS 110 surrounds the left and right boundaries 150 r and 150 l . In other words, due to the example of FIG. 4 where the image content is panoramic content, of options 1 and 2 for the arrangement of section regions 110a and 110b within the image area of output image 144, the second option is actually more reasonable. However, if the image content is of another type other than panoramic, the situation may be different.

[0054] In other words, the order of tiles A, B, C, and D in the complete image bitstream 124' is {a, b, c, d}. If this order is simply transferred to the encoding order of the extracted or reduced video data stream 136, or the arrangement of the corresponding tiles in the output image 144, as shown in the lower left of FIG. 4, the extraction process does not result in the desired data arrangement in the output bitstream 136 in the above exemplary case itself. As shown in the lower right of FIG. 4, the preferred arrangement {b, a, d, c} is shown, which results in a video bitstream 136 that provides continuous image content on the image plane of the picture 144 for legacy devices such as the decoder 146. Such a legacy device 146 may not have the ability to rearrange sub-image regions of the output image 144 as a post-processing step in the decoded pixel domain. That is, rendering, and even sophisticated devices, may prefer to avoid the effort of post-processing.

[0055] Therefore, according to the above motivated example with respect to FIG. 4, the second information 142 respectively provides means for signaling a preferred order among several choices or options to the encoding side of the video data stream generator 128. For example, the section regions 110a, 110b each composed of a set of one or more tiles within the video data stream 124' should be arranged in the image region that the extracted or reduced video data stream 136 or its image 144 spans. According to the specific syntax example presented in FIG. 5, the second information 142 includes a list encoded in the data stream 124', and indicates the position of each slice 108 that enters the section 110 in the extracted bitstream in the original bitstream order or the input bitstream order, that is, the order of occurrence in the bitstream 124'. For example, in the example of FIG. 4, the priority option 2 becomes a list that reads {1, 0, 3, 2}. FIG. 5 shows a specific example of the syntax including the second information 142 in the MCTS extraction information set SEI message.

[0056] The semantics are as follows.

[0057] num_associated_slices_minus1[i] plus 1 indicates the number of slices containing MCTS with an MCTS identifier equal to any value in the list mcts_identifier[i][j]. The value of num_extraction_info_sets_minus1[i] must be in the range from 0 to 232 - 2.

[0058] output_slice_address[i][j] identifies the absolute position of the j-th slice in bitstream order belonging to the MCTS whose MCTS identifier in the output bitstream is equal to any value within the list mcts_identifier[i][j]. The value of output_slice_address[i][j] is in the range from 0 to 2 23 - 2.

[0059] Next, additional processing steps of the extraction process defined in [2] are described to facilitate understanding of the signaling embodiment of FIG. 5. Additions related to [2] are highlighted in underline.

[0060] Use the bitstream inBitstream, the target MCTS identifier mctsldTarget, the target MCTS extraction information set identifier mctsEISIdTarget, and the target maximum TemporalId value mctsTldTarget as inputs to the sub-bitstream MCTS extraction process. The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream. It is a requirement of the bitstream conformity of the input bitstream that the output sub-bitstream, which is the output of the process specified with the bitstream in this section, is a conforming bitstream.

[0061] OutputSliceOrder[j] is derived from the list output_slice_order[i][j] of the i-th extraction information set.

[0062] The output sub-bitstream is derived as follows. - The bitstream outBitstream is set to be the same as the bitstream inBitstream. [...] - Remove all NAL units in outBitstream that have a TemporalId greater than mBitsTIdTarget. - Sort the NAL units of each access unit according to the list OutputSliceOrder [j]. - For each remaining VCL NAL unit in outBitstream, adjust the slice segment header as follows. - For the first VCL NAL unit in each access unit, set the value of first_slice_segment_in_pic_flag to 1. Otherwise, set it to 0. - Set the value of slice_segment_address according to the tile setting defined by the PPS where pps_pic_parameter_set_id is equal to slice_pic_parameter_set_id.

[0063] Thus, the above-described variant of the embodiment of FIG. 2 according to FIG. 5 is summarized. This variant differs from what was discussed above with respect to FIG. 3 in that the second information 142 does not explicitly signal how the slice address should be modified. That is, the second information 142 does not explicitly signal an alternative to the slice address of the slice portion extracted from the data stream 124 to the data stream 136 according to the variant outlined above. Rather, the embodiment of FIG. 5, in the case where the first information 140 defines the spatial section 110 within the image area of the image 100, consists of at least a first sub-region in which a video independent from the outside of the first sub-region 100a is encoded into the video data stream 124', and a second sub-region 110b, in which the video 120 is encoded into the video data stream 124' independent from the outside of the second sub-region 110b. None of the plurality of slices 108 intersects any boundary of either the first and second sub-regions 110a and 110b. The image 144 of the output video 138 of the extracted video data stream 136 can be configured differently, at least in units of these regions 110a and 110b. Thus, according to two options and according to the variant discussed above with respect to FIG. 5, the second information 142 notifies reordering information indicating how to reorder the slice portions 126 of the slices 108 located in the region 110 when extracting the reduced video data stream 136 from the video data stream 124' associated with the slice order and the video data stream 124'. The reordering information 142 can include, for example, a set of one or more syntax elements. Among the states that can be signaled by the one or more syntax elements forming the information 142, there can be a state in which the reordering maintains the original order. For example, the information 142 notifies the rank (compared to 141 in FIG. 5) of each slice portion 126 encoding one of the slices 108 of the image 100 entering the region 110, and the extraction device 134 rearranges the slice portions 126 within the extracted or reduced video data stream 136 according to these ranks. Next, the extraction device 134 modifies the slice addresses of the thus reconstructed slice portions 126 within the reduced or extracted video data stream 136 in the following manner. The extraction device 134 knows about these slice portions 126 within the video data stream 124' that are extracted from the data stream 124' to the data stream 136. Thus, the extraction device 134 knows about the slices 108 corresponding to these inherited slice portions 126 within the image area of the image 100. Based on the rearrangement information provided by the information 142, the extraction device 134 can determine how the sub-regions 110a and 110b have been translated relative to each other and shifted relative to each other so as to result in a rectangular image area corresponding to the image 144 of the video 138. For example, in the example of FIG. 4, Option 2, the extraction device 134 assigns the slice address 0 to the slice corresponding to the tile address b when the slice address 0 occurs at the second position in the list of ranks provided by the second information 142. Thus, the extraction device 134 can place one or more slices associated with tile b and then track the next slice associated with the slice address indicating the position according to the resort information, and the coding order 112 follows immediately after tile b of the image area. In the example of FIG. 4, this is the slice regarding tile a. This is because the next ranked position relative to the first slice a of the area 110 of the image 100 is shown. In other words, the resort is limited so as to lead to any possible rearrangement of the sub-regions 110a and 110b. For each sub-region, individually, it is true that the coding orders 106 and 112 cross their respective sub-regions within the same path. However, due to the tiling, the regions corresponding to the sub-regions 110a and 110b of the image area of the image 144, i.e., in the case of Option 2, the region of the combination of one tile b and d and the region of the combination of the other tile A and C, are traversed in an interleaved manner by the coding order 112, and accordingly, the associated slice portions 126 encoding the slices in the corresponding tiles are interleaved within the extracted or reduced video stream 136.

[0064] A further embodiment is to signal the guarantee that a further order signaled using the existing syntax reflects the preferred output slice order. More specifically, this embodiment can be implemented by interpreting the occurrence of the MCTS extraction SEI message [2] as guaranteeing the order of the rectangles that form the MCTS with the MCTS SEI messages from Sections D.2.29 and E.2.29 of [1] that indicate the priority output order of the tiles / NAL units. In the specific example of FIG. 5, the rectangles are used in the order {b, a, d, c} for each tile included. An example of this embodiment would be the same as the above except for the derivation of OutputSliceOrder [j].

[0065] OutputSliceOrder [j] is derived from the order of the rectangles notified by the MCTS SEI message.

[0066] Summarizing the above example, the second information 142 can signal to the extraction device 134 how to resort the slice portions 126 of the slices that enter the spatial section 110 when extracting the reduced video data stream 136 from the video data stream, regarding how the slice portions 126 of the video data stream 124' are ordered in the sequence of the slice portions 126 of the video data stream. The slice address of each slice portion 126 of the sequence of slice portions of the video data stream 124' indexes one-dimensionally the coding start position of the slice 108 encoded in each slice portion 126 along the first coding scan order 106, and it traverses the image region along which the image 100 is encoded in the sequence of slice portions of the video data stream. Thereby, the slice addresses of a series of slice portions in the video data stream 124' increase monotonically, and the correction of the slice addresses in the extraction of the reduced video data stream 136 from the video data stream 124' is to sequentially arrange the slices encoded in the slice portions in which the reduced video data stream 136 is confined and rearranged as signaled by the second information 142. Along the second coding scan order 112 that traverses the reduced picture area, the slice address of the slice portion 126 is set to index the coding start position of the slice measured along the second coding scan order 112. The first coding scan order 106 traverses the image region within each of at least two sets of partitioned regions in a way that matches the way each spatial region is traversed by the second coding scan order 112. Each of the at least two sets of partitioned regions is shown by the first information 140 as a sub-array of rectangular tiles into the rows and columns into which the image 100 is subdivided, where the first and second coding scan orders use a raster scan of tiles in the row direction that completely traverses the current tile before proceeding to the next tile.

[0067] As already explained above, the order of the output slices can be derived from other syntactic elements such as the above output_slice_address[i][j]. An important addition to the above syntax example for output_slice_address[i][j] in this case is that the slice addresses of all relevant slices including the first one enabling sorting are notified. An example of this embodiment would be the same as the above except for the derivation of OutputSliceOrder[j].

[0068] OutputSliceOrder[j] / OutputSliceOrder[j] is derived from the list output_slice_address[i][j] of the i-th extraction information set.

[0069] Yet another embodiment consists of a single flag on information 142 indicating that the video content wraps around an image boundary set, such as the boundaries of an image in the vertical direction. Thus, the output order is derived by the extraction device 134 corresponding to the image subsections including the tiles on both image boundaries as described above. In other words, the information 142 can notify one of two options. The first option of the plurality of options indicates that the video is a panoramic video showing a scene such that the edge portions of different images of the scene abut each other, and the second option of the plurality of options indicates that different edge portions do not abut each other scene-wise. At least two section regions a, b, c, d that make up section 110 are composed of different portions of different edge portions, that is, the first and second zones 110a, 110b abutting the left and right edges 150r and 150l. Therefore, when the second information 142 notifies the first option, the reduced image region is composed by combining a set of at least two partial regions, the first zone and the second zone abut along different edge portions, and when the second information 142 notifies the second option, the reduced image region is composed by combining a set of at least two partial regions, the first and second zones have different edge portions facing opposite each other.

[0070] For the sake of completeness, it should be noted that the shape of the image area of the image 144 is not restricted to conform to joining together various regions such as tiles a, b, c, and d of section 110 in a way that maintains the relative arrangement of connected clusters such as (a, c, b, d) in FIG. 1, or joining such clusters along the direction of possible interconnection in the shortest distance, such as horizontally joining zones (a, c) and (b, d) in FIG. 2. Rather, for example, in FIG. 1, the image area of the image of the extracted data stream could be columns of all four regions, and in FIG. 2 it could be columns of rows of all four regions. In general, the size and shape of the image area of the image 144 of the video 138 can exist in different parts of the data stream 124'. For example, this information can be given in the form of one or more parameters within the information 140 so as to aim to guide the extraction process of the extraction device 134 regarding the adaptation of the parameter set when extracting the section-specific sub-stream 136 from the stream 124'. The nested parameter set that replaces the parameter set of the stream 124' is included, for example, in the information 140, and this parameter set includes the following. For example, the parameter related to the size of the image indicates the size and shape of the image 144, for example, in pixel units, and when replacing the parameter set of the stream 124' during extraction by the extraction device 134, the old parameters within the parameter set indicating the size of the image 100 are overwritten. However, additionally or alternatively, the image size of the image 144 may be shown as part of the information 142. Explicitly signaling the shape of the picture 144 with a high-level syntax element such as the information 142 that is readily readable may be particularly advantageous when the slice address is not explicitly provided in the information 142. To derive the address, it is necessary to analyze the nested parameter set.

[0071] Also, in a more sophisticated system configuration, it is necessary to note that third-order projection can be used. This projection method avoids the known weaknesses of the orthographic cylindrical projection method, such as a large change in sampling density. However, when using third-order projection, a rendering stage is required to recreate a continuous viewport from the content (or its subsections). In such a rendering stage, the trade-off between complexity and functionality may vary. That is, some executable off-the-shelf rendering modules may expect a predetermined arrangement of the content (or its subsections). In such a scenario, it is essential to manipulate the arrangement to enable the following inventions.

[0072] In the following, embodiments related to the second aspect of the present application will be described. The description of the embodiments of the second aspect of this application starts again from a brief introduction to the general problems or issues assumed and addressed by these embodiments.

[0073] An interesting use case of MCTS extraction in a context not limited to 360° video is a composite video that includes variations in multiple resolutions of adjacent content on the screen, as shown in FIG. 6. The lower part of FIG. 6 shows a composited video 300 with multiple resolutions, and the line 302 indicates the tile boundaries of the tiles 304 where the composition of the high-resolution video 306 and the low-resolution video 308 is subdivided to encode the corresponding data stream.

[0074] More precisely, FIG. 6 shows an image of a high-resolution video at 306 and images such as a simultaneous time image of a low-resolution video at 308. For example, both videos 306 and 308 show exactly the same scene, i.e., have the same view or field of view, but differ in resolution. However, the fields of view may only partially overlap alternatively with each other, and in the overlap zone, the spatial resolution indicates that the number of samples of the photo of video 306 is different compared to video 308. In practice, videos 306 and 308 differ in fidelity, i.e., the number of samples or pixels of the same scene section. Photos at the same location in videos 306 and 308 are arranged side by side and combined into a large photo, which becomes the photo of the composite video 300. For example, FIG. 6 shows that the image of video 308 is halved horizontally, and the two halves overlap vertically and are attached to the right side of the simultaneous time image of video 306 so that the corresponding video 300 is obtained. Tile 304 is performed so as not to cross the junction between the high-resolution image of video 306 on the one hand and the image content resulting from the low-resolution video 308 on the other hand. In the example of FIG. 6, the tiling of image 300 results in an 8×4 tile in which the high-resolution image content of the image of video 300 is divided, and a 2×2 tile in which each half resulting from the low-resolution video 308 is divided. Overall, the width of the image of video 300 is 10×4.

[0075] When such a composite video 300 of multiple resolutions is encoded by MCTS in an appropriate manner, MCTS extraction can generate a deformation 312 of the content. Such a deformation example 312 can be designed, for example, as shown in FIG. 7, to depict a predetermined sub-picture 310a at a high resolution and the rest of the scene or another sub-section at a low resolution, where the MCTS 310 in the composite image bitstream is compared with three sides or separate regions 310a, b, c.

[0076] That is, the image of the extracted video, i.e., Image 312, has three fields 314a, 314b, and 314c corresponding to one of the MCTS regions 310a, b, c, respectively. Region 310a is a sub-region of the high-resolution image area of Image 300, and the other two regions 310b and 310c are sub-regions of the low-resolution video content of Image 308.

[0077] Having stated this, with reference to FIG. 8, embodiments of the present application relating to the second aspect of the present application will be described. In other words, FIG. 8 shows a scenario of presenting multi-resolution content to a recipient site, in particular, to individual sites participating in the generation of multiple versions of different resolutions up to the receiving site. It should be noted that the individual devices and processes arranged at various sites along the process path shown in FIG. 8 represent individual devices and methods, and thus FIG. 8 should not be construed as showing only the overall system or method. A similar description also applies to FIG. 2, which similarly shows individual devices and methods. The reason for presenting all these sites together in one figure is only to facilitate the understanding of the interrelationships and advantages arising from the embodiments described with respect to these figures.

[0078] Figure 8 shows a video data stream 330 encoded in video 332 of picture 334. On the other hand, image 334 is the result of stitching together the simultaneous time image 336 of high-resolution video 338 and the image 340 of low-resolution video 342. More precisely, the images 336 and 340 of videos 338 and 342 correspond to exactly the same viewport 344 or at least partially overlap to display the same scene in the overlapping area. However, the number of samples of the high-resolution video image 336 that samples the same scene content as the corresponding low-resolution image 340 is larger, and thus the scene resolution of the fidelity of image 336 is higher than that of image 340. The composition of image 334 of the composition video 332 based on videos 338 and 342 is performed by composer 346. Composer 346 stitches together images 336 and 340. By doing so, composer 346 can re-divide either the image 340 of the low-resolution video or the image of high-resolution video 338 or both in order to provide advantageous filling and patching of the image area of image 334 of the composition video 332. Next, video encoder 348 encodes the composite video 332 into video data stream 330. The video data stream generation device 350 may be included in the video encoder 348 or connected to the output of the video encoder 348 to provide the image 334, or a signaling 352 indicating each image or video 332, to the video data stream 330. Alternatively, each picture or video 332 within a particular sequence of pictures in video 332 is encoded into video data stream 330 and displays the common scene content multiple times, i.e., in different spatial portions at different resolutions. These portions are shown in Figure 8 using H and L to indicate their origin due to the composition performed by composer 346 and are shown using reference numerals 354 and 356. There may be two or more different resolution versions combined to form the content of image 332, and as shown in Figures 8, 6, and 7, which are for illustrative purposes only, attention needs to be paid to the use of the two versions.

[0079] FIG. 8 shows, for purposes of explanation, a video data stream processor 358 that receives a video data stream 330. The video data stream processor 358 can be, for example, a video decoder. In any case, the video data stream processor 358 can examine the signaling 352 to determine whether the video data stream processor 358 should begin processing, such as decoding the video data stream 330, which further depends on the specific capabilities of, for example, the video data stream processor 358 or a device connected downstream thereof, based on this signaling. For example, the video data stream processor 358 can simply present the fully encoded images 332 in the video data stream 330, and then the video data stream processor 358 may reject the processing of the video data stream 330, where the signaling 352 means that the individual images of the video 332 represent the same scene content at different spatial resolutions in different spatial portions of these individual images. That is, the signaling 352 means that the video data stream 330 is a multi-resolution video data stream.

[0080] The signaling 352 can include, for example, a flag transmitted within the data stream 330, and the flag is switchable between a first state and a second state. The first state can indicate, for example, the fact just outlined, that is, that the individual images of the video 332 represent multiple versions of the same scene content at different scene resolutions. The second state indicates that such a situation does not exist. That is, the image displays only one scene content at one resolution. Thus, the video data stream processor 358 will respond to the flag 352 being in the first state by rejecting the execution of a particular processing task.

[0081] The signaling 352, such as the aforementioned flag, can be transmitted within the data stream 330 in its sequence parameter set or video parameter set. Possible syntax elements reserved for future use in HEVC are illustratively identified as possible candidates in the following description.

[0082] As noted above with respect to FIGS. 6 and 7, while it is not necessary, it is possible for the composite video 332 to be encoded in the video data stream 330 in such a way that it is independently encoded from outside each of a set of non-overlapping spatial regions 360, such as tiles or tile sub-arrays. Encoding independence restricts spatial and temporal prediction and / or context derivation so as not to cross the boundaries between the spatial regions 360, as explained above with respect to FIG. 2. Thus, for example, encoding dependence can refer the encoding of a particular spatial region 360 of a particular image 334 of the video 332 only to spatial regions in the same location within another image 334 of the image 332, which is also subdivided into spatial regions 360 like the image 334. The video encoder 348 can provide the video data stream 330 with extracted information, such as the information 140 or a combination of the information 140 and 142, or can connect each device, such as the device 128, to its output. The extracted information can be related to a particular or all possible combinations of the spatial regions 360, as an extraction section, such as the extraction section 310 of FIG. 7. Next, the signaling 352 can include information regarding the spatial subdivision of the video 332 into sections 354 and 356 of different scene resolutions of the image 334, i.e., the size and position of each spatial section 354 and 356 within the image area of the image 332. Based on such information in the signaling 352, the video data stream processor 358 can, for example, exclude a particular extraction section from a list of possibly extractable sections of the video stream 330. For example, these extraction sections that mix spatial regions 360 divided into different ones of the sections 354 and 356 avoid the performance of video extraction for extraction sections that mix different resolutions. In this case, the video stream processor 358 can include an extraction device, such as the extraction device of FIG. 2, for example.

[0083] Additionally or alternatively, signaling 330 can include information regarding different resolutions where the images 334 of video 332 show common scene content with each other. Further, signaling 352 can simply indicate the count of different resolutions where the images 334 of video 332 show common scene content multiple times at different image positions.

[0084] As already described, video data stream 330 can include extraction information regarding a list of potential extraction regions regarding which video data streams 330 are extractable. Next, signaling 352 can include, for at least one or more of these extraction regions, further signaling indicating the viewport direction of the sub-regions of each extraction region where common scene content is shown at the highest resolution within each extraction region, and / or the area share of the cross-section regions of each extraction region where common scene content is displayed at the highest resolution within the overall region of each extraction region, and / or the spatial subdivision of each extraction region into sub-regions where common scene content is shown at different resolutions with respect to each other within the overall region of each extraction region.

[0085] Thus, such signaling 352 may be exposed at a high level of the bitstream so as to be easily pushed up to the streaming system.

[0086] One option is to use one of the zero X bit flags that are generally reserved in the profile layer level syntax. The flag can be named as a general non-multiple resolution flag.

[0087] A general non-multiple resolution flag equal to 1 specifies that the decoded output image does not contain multiple versions of the same content at different resolutions (i.e., each syntax such as region packing is restricted). A general non-multiple resolution flag equal to 0 (zero) specifies that such content may be included in the bitstream (i.e., no restrictions).

[0088] In addition, and thus, the present invention is composed of signaling that notifies about the nature of the complete bitstream content characteristics, i.e., the number of variants and the resolution in the configuration. Further, additional signaling that provides information on the coded bitstream regarding the following characteristics of each MCTS in a form that can be easily accessed: · Direction of the main view point: What is the orientation of the MCTS high-resolution viewport center? For example, from the perspective of pitch (delta-yaw) and / or roll from a pre-defined initial viewport center. · Overall coverage The proportion of the complete content represented by the MCTS. · High-to-low resolution ratio The ratio between the high-resolution and low-resolution regions of the MCTS, i.e., what proportion of the total content covered is represented in high resolution / fidelity.

[0089] Signaling information has already been proposed for the orientation of the viewport or the overall coverage of a complete omnidirectional video. Similar signaling needs to be added to potentially extractable sub-regions. Since this information is in the form of SEI[2], it can be included in the motion-constrained tile set extraction information that nests the SEI. However, such information is required to select the MCTS to extract. Obtaining the extraction information of the Motion-constrained tile set that nests the SEI adds additional indirection and requires deeper analysis (the extraction information of the motion-constrained tile set that nests the SEI contains additional information that is not necessary to select the extracted set) to select a specific MCTS. From a design perspective, it is a cleaner approach to signal this information or a subset of it at a central point that only contains important information for selecting the extracted set. Furthermore, the signaling mentioned contains information about the entire bitstream, and in the proposed case, it would be desirable to signal which is the high-resolution coverage and which is the low-resolution coverage. Also, when more resolutions are mixed, it is the coverage of each resolution and the orientation of the viewport of the video extracted at the mixed resolution.

[0090] One embodiment is to add the coverage of each resolution and add it to the 360 ERP SEI from [2]. Next, this SEI may be included in the motion-constrained tile set extraction information that nests the SEI, and the above-mentioned cumbersome tasks need to be performed.

[0091] In another embodiment, a flag is added to the omnidirectional information indicating the presence of the signaling discussed, such as the MCTS extraction information set SEI, so that only the MCTS extraction information set SEI is required to select the set to be extracted.

[0092] Although some aspects have been described in the context of an apparatus, these aspects also represent corresponding method descriptions, and it is clear that a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent corresponding block or item or corresponding apparatus functionality descriptions. Some or all of the method steps may be performed (or used) by a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.

[0093] The data stream of the present invention can be stored in a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0094] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be carried out using a digital storage medium storing electronically readable control signals, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, ERROM, EEPROM, or flash memory, on which a programmable computer system can cooperate (or be capable of cooperating) so that each method is executed. Thus, the digital storage medium may be computer-readable.

[0095] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is executed.

[0096] In general, embodiments of the present invention can be implemented as a computer program product having program code, which operates to execute one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, on a machine-readable carrier.

[0097] Other embodiments include a computer program for executing one of the methods described herein, stored on a machine-readable carrier.

[0098] In other words, thus, an embodiment of the method of the present application is a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.

[0099] Thus, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) recording a computer program for executing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.

[0100] Thus, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for executing one of the methods described herein. The data stream or sequence of signals may be configured to be transferred via a data communication connection such as the Internet, for example.

[0101] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to execute one of the methods described herein.

[0102] Further embodiments include a computer installed with a computer program for executing one of the methods described herein.

[0103] Further embodiments according to the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0104] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0105] The apparatus described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0106] The apparatus described herein, or any component of the apparatus described herein, may be implemented at least partially in hardware and / or software.

[0107] The methods described herein can be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0108] The methods described herein, or any component of the apparatus described herein, can be performed at least partially by hardware and / or software.

[0109] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and changes to the arrangements and details described herein will be apparent to other skilled persons. Accordingly, it is intended to be limited only by the claims as they now exist, rather than by the specific details presented as descriptions and explanations of the embodiments herein.

Claims

Claim 1 An apparatus for generating a video data stream, configured to provide to the video data stream a sequence of slice parts, each slice part encoding a respective slice of a plurality of slices of a video picture, each slice part including a slice address indicating a location within the picture area of the video where the slice encoded by the respective slice part is located, configured to provide to the video data stream extraction information used when extracting a reduced video data stream from the video data stream, the reduced video data stream encoding a spatially smaller video corresponding to a spatial section of the video of the video data stream, the extraction involving restricting the video data stream to slice parts encoding any slice within the spatial section and modifying the slice address to relate to a reduced picture area of the spatially smaller video, wherein the extraction information includes a number of extraction information sets for different versions of the reduced video data stream, each extraction information set including first information defining the spatial section within the picture area, where the video is encoded independently of outside the spatial section in the video data stream, and second information explicitly signaling an alternative slice address for replacement of the slice address of the slice part of the slice within the spatial section, the alternative slice address indicating a location within the reduced picture area of the reduced video data stream where the slice is located, the extraction information configured to provide to the video data stream, wherein the apparatus is configured to provide to the extraction information a syntax element indicating the number of extraction information sets, apparatus. Claim 2 An apparatus for extracting a reduced video data stream encoding a spatially smaller video from a video data stream encoding a video, wherein the video data stream includes a sequence of slice portions each encoding a respective one of a plurality of slices of an image of the video, and each slice portion includes a slice address indicating a location where the slice encoded by the respective slice portion is located within an image region of the video. The apparatus is configured to read extraction information from the video data stream, where the extraction information includes a number of extraction information sets for different versions of the reduced video data stream. configured to derive a spatial section within the image region from first information of the extraction information, where the reduced video data stream is limited to slice portions encoding any slice within the spatial section. configured to derive alternative slice addresses from second information of the extraction information. configured to replace the slice address of the slice portion of the slice within the spatial section with the alternative slice address explicitly signaled by the extraction information, where the alternative slice address indicates a location within a reduced image region of the spatially smaller video where the slice is located in the reduced video data stream, and the order of the slice portions of the slices within the spatial section is reordered with respect to how the slice portions are ordered in the sequence of slice portions of the video data stream. The apparatus is configured to derive the number of extraction information sets from syntax elements included in the extraction information. Apparatus. **Claim 3** A method for generating a video data stream, the method comprising: providing, to the video data stream, a sequence of slice portions each encoding a respective one of a plurality of slices of an image of the video, wherein each slice portion includes a slice address indicating a location where the slice encoded by the respective slice portion is located within an image region of the video. A step of providing extraction information used when extracting a reduced video data stream from the video data stream, wherein the reduced video data stream encodes a spatially smaller video corresponding to a spatial section of the video of the video data stream, and the extraction involves limiting the video data stream to a slice portion encoding any slice within the spatial section, and modifying the slice address so as to be related to a reduced image region of the spatially smaller video, The extraction information includes a number of extraction information sets for different versions of the reduced video data stream, and each extraction information set respectively, First information defining the spatial section within the image region, Second information explicitly signaling an alternative slice address for replacement of the slice address of the slice portion of the slice within the spatial section, wherein the alternative slice address indicates a location within the reduced image region of the reduced video data stream where each respective slice is located, second information Including the step of providing extraction information, Including, A method in which a syntax element indicating the number of extraction information sets is supplied to the extraction information.

4. A method of extracting a reduced video data stream encoding a spatially smaller video from a video data stream encoding a video, wherein the video data stream includes a sequence of slice portions each encoding a respective slice of a plurality of slices of an image of the video, and each slice portion includes a slice address indicating a location within an image region of the video where the slice encoded by each slice portion is located, The method includes, A step of reading extraction information from the video data stream, wherein the extraction information includes a number of extraction information sets for different versions of the reduced video data stream, the step of reading extraction information, A step of deriving a spatial section within the image region from the first information of the extraction information, wherein the reduced video data stream is limited to a slice portion encoding any slice within the spatial section, the step of deriving, a step of deriving an alternative slice address from the second information of the extracted information; a step of replacing the slice address of the slice portion of the slice in the spatial section with the alternative slice address explicitly signaled by the extracted information, wherein the alternative slice address indicates a location within a reduced image region of the spatially smaller video in which the slice is located in the reduced video data stream, and the order of the slice portions of the slices in the spatial section is reordered with respect to how the slice portions are ordered in the sequence of the slice portions of the video data stream; comprising; the number of the extracted information sets is derived from syntax elements included in the extracted information; method. **Claim 5** A non-transitory digital storage medium storing a computer program for executing the method of generating a video data stream according to claim 3 when executed on a computer. **Claim 6** A non-transitory digital storage medium storing a computer program for executing the method of extracting according to claim 4 when executed on a computer.

Citation Information

Patent Citations

  • Video composition

    WO2016026526A2