Concept about image / video data stream for enabling effective reduction or effective random access

The described method allows for efficient encoding and random access of non-rectangular video content by using subregion-specific encoding and dual random access points, addressing computational complexity and bitrate peak issues in existing codecs.

JP2025170281APending Publication Date: 2025-11-18FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025132413
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-02-09
Filing Date
2025-08-07
Publication Date
2025-11-18

Smart Images

  • Figure 2025170281000001_ABST
    Figure 2025170281000001_ABST
Patent Text Reader

Abstract

To enable effective random access.SOLUTION: A video stream 10 undergoes rendering in a reduceable manner by reducing such that an image of a reduced video data stream is limited to only a prescribed sub area 22 of an image of an original video data stream and also that transcoding such as re-quantization is avoided to maintain the adaptability of the reduced video data stream for a codec to be a base of the original video data stream. This is achieved by providing the video data stream with information 50 including an instruction of the prescribed sub area and a replacement parameter for adjusting a first set 20a of coding parameter setting so as to cause a second set 20b of a replacement index and / or coding parameter setting 20 to be redirected so as to refer to an index 48 in a payload part 18.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to video / image coding, and in particular to concepts that allow for efficient reduction of such data streams, concepts that allow for easier handling of such data streams, and / or concepts that allow for more efficient random access to video data streams. [Background technology]

[0002] Many video codecs exist that enable scalability of video data streams without transcoding, i.e., without the need to perform decoding and encoding consecutively. Examples of this kind of scalable video data streams are data streams that are scalable, for example, with respect to temporal resolution, spatial resolution, or signal-to-noise ratio, by simply omitting some of the enhancement layers of each scalable video data stream. However, to date, no video codec has enabled scalability with low computational complexity regarding scene delimitation. In HEVC, concepts exist or have been proposed for restricting HEVC data streams to image subregions, but these are still computationally complex.

[0003] Furthermore, depending on the application, the image content to be encoded into the data stream may be in a form that cannot be efficiently encoded within the rectangular image area typically provided. For example, panoramic image content may be projected onto a two-dimensional plane, and the projection target, i.e., the footprint of the panoramic scene onto the image area, forms an image area that may be non-rectangular or even non-convex. In that case, more efficient encoding of the image / video data would be advantageous.

[0004] Furthermore, random access points are provided to existing video data streams in such a way that they cause significant bitrate peaks. To reduce the negative impact caused by these bitrate peaks, one can consider reducing the temporal granularity of the occurrence of these random access points. However, this increases the average time duration for randomly accessing such a video data stream, and therefore it would be advantageous to have a concept that solves this problem at hand in a more efficient way. Summary of the Invention [Problem to be solved by the invention]

[0005] The object of the present invention is therefore to solve the above-mentioned problems. According to the present application, this is achieved by the subject matter of the independent claims. [Means for solving the problem]

[0006] According to a first aspect of the present application, a video data stream is reducibly rendered such that image restrictions of the reduced video data stream are limited only to a predetermined subregion of the image of the original video data stream, transcoding such as requantization is avoided, and compatibility of the reduced video data stream with the underlying codec of the original video data stream is maintained. This is achieved by providing a video data stream containing information consisting of an indication of the predetermined subregion, a substitution index for redirecting an index configured by a payload portion for reference, and / or a substitution parameter for adjusting a first set of encoding parameter settings to produce a second set of encoding parameter settings. The payload portion of the original video data stream has been coded with an image of the parameterized video using a first set of encoding parameter settings indexed by an index included in the payload portion. Additionally or alternatively, similar measures can be implemented for auxiliary enhancement information. It is therefore possible to reduce the video data stream to a reduced video data stream by performing redirection and / or adjustment, such that the second set of coding parameter settings is indexed by the payload portion's index and thus becomes the valid coding parameter setting set, remove parts of the payload portion that refer to areas of the image outside the predetermined sub-region, and modify location indications such as slice addresses of the payload portions to indicate locations measured from the periphery of the predetermined sub-region instead of the periphery of the image. Alternatively, a data stream that has already been reduced so that it does not include parts of the payload portion that refer outside the predetermined sub-region may be modified by fly adjustment of parameters and / or supplemental enhancement information.

[0007] According to a further aspect of the present invention, transmission of image content is made more efficient in that the image content does not need to be formed or ordered in a predetermined manner, e.g., the way typically rectangular image regions supported by an underlying codec are written. Rather, a data stream having images encoded therein is provided that includes, for a set of at least one predetermined subregion of the image, displacement information indicating the displacement within the region of the target image relative to an undistorted or one-to-one corresponding or congruent copy of the region of the target image of the set. Providing such displacement information is useful, for example, when conveying the projection of a panoramic scene within an image when the projection is non-rectangular. This displacement information is also useful when, due to data stream reduction, the image content conveyed within a smaller image of the reduced video data stream loses its suitability, for example, in the case of an interesting panoramic view section to be transmitted within the reduced video data stream crossing a transition boundary, such as a prediction.

[0008] According to a further aspect of the present invention, the negative impact of bit rate peaks in a video data stream caused by random access points is reduced by providing a video data stream with two sets of random access points: encoding a first set of one or more images into the video data stream with temporal prediction suspended within at least a first image sub-region to form a set of one or more first sets of random access points, and encoding a second set of one or more images into the video data stream with temporal prediction suspended within a second image sub-region different from the first image sub-region to form a set of one or more second random access points. In this way, a decoder wishing to randomly access or resume decoding the video data stream can select one of the first and second random access points that are distributed in time and enable random access at least with respect to the second image sub-region in the case of the second random access point and with respect to at least the first image sub-region in the case of the first random access point.

[0009] The above concepts can be used advantageously together. Further advantageous embodiments are the subject of the dependent claims. Preferred embodiments of the present application are described below with reference to the drawings. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 shows a schematic diagram of a video data stream according to an embodiment of the present application according to a first aspect, whereby the video data stream can be reduced into a reduced video data stream relating to a sub-region of an image of the reducible video data stream. [Figure 2] FIG. 2 shows a schematic diagram illustrating the interdependence between the payload portion and the parameter set portion of the reducible video data stream of FIG. 1 according to an embodiment to illustrate the parameterization with which images are coded into the reduced video data stream. [Figure 3]FIG. 3 shows a schematic diagram for explaining a possible content of information provided by an embodiment of the video data stream of FIG. 1 allowing reduction. [Figure 4] FIG. 4 shows a schematic diagram illustrating a network device that receives a reducible video data stream and derives a reduced video data stream therefrom. [Figure 5] FIG. 5 shows a schematic diagram illustrating the operation mode when reducing the video data stream according to an embodiment using parameter set redirection. [Figure 6] FIG. 6 shows a schematic diagram illustrating a video decoder 82 that receives a reduced video data stream and reconstructs images of the reduced video data stream that instead merely show sub-regions of the images of the original video data stream. [Figure 7] FIG. 7 shows a schematic diagram of an alternative mode of operation when reducing a reducible video data stream at this point using parameter set adjustment using substitutions in the information provided by the reducible video data stream. [Figure 8] FIG. 8 shows an example syntax for information that a reducible video data stream can provide. [Figure 9] FIG. 9 shows an alternative example of a syntax for the information that a reducible video data stream can provide. [Figure 10] FIG. 10 shows a further example of the syntax of information that a reducible video data stream can provide. [Figure 11] FIG. 11 shows a further example of syntax for the information now to replace the SEI message. [Figure 12] FIG. 12 shows an example of a syntax table that can be used to form information relating to a multi-layer video data stream. [Figure 13]Figure 13 shows a schematic diagram illustrating the relationship between tiles in a sub-region of an image in one reducible video data stream and corresponding tiles in an image in the other reduced video data stream, in accordance with one embodiment, to illustrate the possibility of spatial rearrangement of these tiles. [Figure 14] FIG. 14 shows an example of an image obtained by linearly predicting a panoramic scene. [Figure 15] FIG. 15 shows an example of an image with image content corresponding to a three-dimensional projection of a panoramic scene. [Figure 16] FIG. 16 shows an image efficiently filled using the three-dimensional projection content of FIG. 15 by rearranging it. [Figure 17] FIG. 17 shows an example of a syntax table for displacement information using a data stream that can be provided by an embodiment according to the second aspect of the present invention. [Figure 18] FIG. 18 shows a schematic diagram illustrating the structure of a video data stream according to an embodiment related to the second aspect of the present invention. [Figure 19] FIG. 19 shows a schematic diagram illustrating possible contents for replacing information according to an embodiment. [Figure 20] FIG. 20 shows a schematic diagram illustrating an encoder configured to form displacement information and simultaneously a reducible data stream. [Figure 21] FIG. 21 shows a schematic diagram illustrating a decoder configured to receive a data stream including displacement information to illustrate a possibility of how the displacement information can be used advantageously. [Figure 22] FIG. 22 shows a schematic diagram illustrating a video data stream including sub-region specific random access images according to an embodiment related to a further aspect of the present application. [Figure 23A] FIG. 23A shows a schematic diagram illustrating a possible arrangement of sub-regions used according to different alternatives. [Figure 23B] FIG. 23B shows a schematic diagram illustrating a possible arrangement of sub-regions used according to different alternatives. [Figure 23C]FIG. 23C shows a schematic diagram illustrating a possible arrangement of sub-regions used according to different alternatives. [Figure 23D] FIG. 23D shows a schematic diagram illustrating possible arrangements of sub-regions used according to different alternatives. [Figure 23E] FIG. 23E shows a schematic diagram illustrating a possible arrangement of sub-regions used according to different alternatives. [Figure 24] FIG. 24 shows a schematic diagram illustrating a video decoder configured to receive a video data stream having sub-region specific random access images interspersed therein, according to an embodiment. [Figure 25] Figure 25 shows a schematic diagram illustrating the situation of Figure 24, but illustrates an alternative mode of operation of the video decoder in that when the video decoder randomly accesses the video data stream, the video decoder waits until the image region of the input video data stream is completely covered by the subregions of the subregion-specific random access image before outputting or presenting the video. [Figure 26] FIG. 26 shows a schematic diagram illustrating a network device receiving a video data stream containing sub-region specific random access images, the sub-regions simultaneously forming sub-regions relative to which the video data stream can be reduced. [Figure 27] FIG. 27 shows a schematic diagram illustrating a network device 231 receiving a data stream that can be reduced and provided with displacement information, to illustrate how the network device 231 can potentially provide sub-region specific displacement information for the reduced video data stream. [Figure 28] FIG. 28 shows an example of an isolated region of interest sub-region of an image, illustratively a cylindrical panorama. [Figure 29] FIG. 29 shows a syntax table of the TMCTS SEI message of HEVC. DETAILED DESCRIPTION OF THE INVENTION

[0011] The present description relates to the above-identified aspects of the present application. To provide background for the first aspect of the present application, which relates to sub-region extraction / reduction of video data streams, an example application in which such a need may arise and the challenges of meeting this need are described and overcome are motivated as follows using HEVC as an example.

[0012] A spatial subset, i.e., a set of tiles, can be signaled in HEVC using a Temporal Motion Constrained Tile Set (TMCTS) SEI message. The tile set defined in such a message has the following characteristics: "The inter-prediction process is constrained to have no sample values ​​outside each identified tile set, and no sample values ​​at fractional sample positions derived using one or more sample values ​​outside the identified tile set are used for inter-prediction of any sample within the identified tile set." In other words, samples in a TMCTS can be decoded independently from samples not associated with the same TMCTS in the same layer. A TMCTS encompasses one or more rectangular unions of one or more tiles, as shown in Figure A using rectangle 900. In the figure, the region of interest 900 seen by the user encompasses two discontinuous image patches.

[0013] The detailed syntax of the TMCTS SEI message is shown in Figure B for reference.

[0014] There are many applications where it is beneficial to create independently decodable rectangular spatial subsets of a video bitstream, i.e., regions of interest (RoIs), without the need for processing-intensive video transcoding. These applications include, but are not limited to:

[0015] Panoramic Video Streaming: Only the unique spatial region of the wide-angle video, e.g., a 360° field of view, is displayed to the end user via a head-mounted display. Aspect Ratio Adjustment Streaming: The aspect ratio of the encoded video is adjusted live on the server side according to the client's display characteristics. · Decoding complexity adjustment: Low-cost / low-tech devices that cannot decode a given encoded video bitstream due to level limitations may accommodate a spatial subset of the video.

[0016] Considering the state of the art described thus far for the above list of exemplary applications, a number of issues arise.

[0017] · There is no means to make the HRD parameters, i.e., buffering / timing information, of a spatial subset of a video bitstream available to the system layer. · No conformance points exist in the video, so that a spatial subset of a given video bitstream can be easily transformed into a conforming video bitstream. · There is no way for an encoder to convey a guarantee that a tile set with a given identifier can be easily converted into a conforming video bitstream.

[0018] Given solutions to the problems listed above, all of the above example applications can be realized in a standards-compliant manner. Defining this capability within the video coding layer is expected to be a key fit point for the application and system layers.

[0019] The HEVC specification already includes a process for the extraction of sub-bitstreams that may reduce the temporal resolution or amount of layers of the coded video bitstream, i.e., the spatial resolution, signal fidelity or number of views.

[0020] The present invention provides solutions to the identified problems, in particular: 1. A means for extracting a spatial subset, i.e., a single TMCTS-based video bitstream, from a coded video sequence via the definition of a TMCTS-based sub-image extraction process. 2. Means for conveying and identifying the correct parameter set values ​​and (optionally) SEI information for the extracted sub-image video sequence. 3. A means for the encoder to convey specific subregion extraction guarantees that allow for bitstream constraints on the video bitstream and TMCTS.

[0021] The embodiments described below overcome the outlined problem by providing a video data stream with information not required for the reconstruction of video images from the payload portion of the video data stream, including an index of a given sub-region and a replacement index and / or replacement parameter, the significance and function of which are described in more detail below. The following description is not limited to HEVC or improvements to HEVC only. Rather, the embodiments described next may be implemented in any video codec technology to provide such video coding technology with additional adaptation points for providing reduced sub-region-specific video data streams. Details are then presented of how the embodiments described next may specifically be implemented to form extensions to HEVC.

[0022] 1 illustrates a video data stream 10 according to one embodiment of the present application, i.e., the video data stream can be reduced in a conformation-preserving manner into a reduced video data stream, whose images simply represent predetermined sub-regions of images 12 of a video 14 that are encoded into the video data stream 10 without the need for transcoding, or more precisely, for time-consuming and computationally complex operations such as requantization, spatial-to-spectral conversion and vice versa, and / or re-performing motion estimation.

[0023] Video data stream 10 in Figure 1 is shown to include a parameter setting portion 16 that indicates encoding parameter settings 80, and a payload portion 18 in which images 12 of video 14 are encoded. In Figure 1, portions 16 and 18 are illustratively distinguished from one another by using hatching for payload portion 18 and indicating that parameter setting portion 16 is not hatched. Furthermore, portions 16 and 18 are illustratively interleaved with one another within data stream 10, although this is not necessarily the case.

[0024] Payload portion 18 encodes images 12 of video 14 in a special manner. In particular, FIG. 1 shows exemplary predetermined sub-regions 22 for which video data stream 10 should be capable of being reduced to a reduced video data stream. Payload portion 18 has images 12 encoded therein such that, as far as a predetermined sub-region 22 is concerned, any coding dependency is constrained to not cross the boundaries of the sub-region 22. That is, an image 12 is encoded in payload portion 18 such that coding within a sub-region 22 is not dependent on the spatial neighborhood of such region 22 within this image. If images 12 are encoded in payload portion 18 using temporal prediction, the temporal prediction may be constrained within sub-region 22 such that any portion within sub-region 22 of a first image of video 14 is coded to depend on regions of reference (other images) of video 14 outside of sub-region 22. That is, the corresponding encoder generating the video data stream 14 restricts the set of available motion vectors for encoding the sub-region 22 in such a way that they do not point to parts of the reference image, so that the formation of the motion compensated prediction signal requires or involves samples of the reference image outside the sub-region 22. As far as spatial dependencies are concerned, it should be noted that the spatial dependency restrictions may relate to sample-by-sample spatial prediction, for example by continuing the arithmetic coding spatially across the boundaries of the sub-region 22, spatial prediction of the coding parameters, and spatial prediction of the coding dependencies.

[0025] Payload portion 18 thus encodes outlined image 12 according to restricting coding dependencies so as not to reach portions outside of predetermined sub-region 22, and thus may include a syntactically ordered sequence 24 of syntax elements including, for example, motion vectors, image reference indices, partition information, coding modes, transform coefficients, or residual sample values ​​indicating quantized prediction residuals, or one or any combination thereof. Most importantly, however, payload portion 18 has images 12 of video 14 encoded therein in a parameterized manner using a first set 20a of coding parameter settings 20. For example, the coding parameter settings in set 20a may specify, for example, the image size of image 12, e.g., the vertical height and horizontal width of image 12. To explain how image size "parameterizes" the encoding of image 12 into payload portion 18, reference is made briefly to FIG. 2, which illustrates a picture size coding parameter 26 as one example of the coding parameter settings of set 20a. Obviously, the image size 26 indicates the size of the image area that must be "encoded" by the payload portion 18, indicating that each sub-block of an image 12 remains unencoded, e.g., filled with a predetermined sample value, such as zero, which may correspond to black. Thus, the image size 26 influences the quantity or size 30 of the syntax description 24 of the payload portion 18. Furthermore, the image size 26 influences the position indicators 32 in the syntax description 24 of the payload portion 18, e.g., regarding the value range of the position indicators 32 and the order in which the position indicators 32 appear in the syntax description 24. For example, the position indicators 32 may include slice addresses within the payload portion 18. A slice 34 is a portion of the data stream 10 in units in which the data stream 10 can be transmitted to a decoder, as shown in FIG. 1 . Each image 12 may be encoded into the data stream 10 in units of such slices 34 and subdivided into slices 34 according to the decoding order in which the image 12 is encoded into the data stream 10.Each slice 34 corresponds to and is encoded within a corresponding region 36 of image 12, but the region 36 is either inside or outside of subregion 22, i.e., does not cross a subregion boundary. In such cases, each slice 34 may be given a slice address that indicates the location of the corresponding region 36 within the image region of image 12, i.e., relative to the perimeter of image 12. To take a specific example, the slice address may be measured relative to the upper left corner of image 12. Obviously, such slice address should not exceed the value of the slice address within an image of image size 26.

[0026] Similar to image size 26, set of encoding parameter settings 20a can also define a tile structure 38 into which image 12 can be subdivided. Using dash-dotted lines 40, FIG. 1 illustrates an example of a subdivision of image 12 into tiles 42, such that the tiles are arranged in a tile array of columns and rows. In any case where image 12 is encoded into payload portion 18 using tile subdivision into tiles 42, this implies, for example, that 1) spatial interdependencies across tile boundaries are not permitted and therefore not used, and 2) the decoding order in which image 12 is encoded into data stream 10 traverses image 12 in raster-scan tile order, i.e., each tile is traversed before visiting the next tile in tile order. Thus, tile structure 38 influences 28 the decoding order 44 in which image 12 is encoded into payload portion 18 and accordingly influences syntax description 24. In a similar manner to image size 26, tile structure 38 also has an effect 28 on location indicators 32 within payload portion 18, i.e., the order in which different instantiations of location indicators 32 are allowed to occur within syntax description 24.

[0027] The encoding parameter settings of set 20a may also include buffer timing 46. Buffer timing 46 may indicate, for example, coded image buffer removal times. Because certain portions of data stream 10, such as individual slices 34 or portions of data stream 10 referencing one image 12, should be removed from the decoder's coded image buffer, and these temporal values ​​affect or relate to the size 28 of the corresponding portions in data stream 10, buffer timing 46 also affects 28 the quantity / size 30 of payload portions 18.

[0028] That is, as illustrated in the description of FIG. 2, the encoding of image 12 into payload portion 18 is "parameterized" or "described" using set 20a of encoding parameter settings 20, on the one hand, and payload portion 18 and its syntax description 24, on the other hand, are identified as conforming, in the sense that any discrepancy between set 20a of encoding parameter settings 20 and payload portion 18 and its syntax description 24 is identified as inconsistent with the conformance requirements required to comply with the data stream identified as conforming.

[0029] The first set of coding parameter settings 20a is included in the payload portion 18 and is referenced or indexed by an index 48 interspersed or contained within the syntax description 24. For example, the index 48 may be included in the slice header of the slice 34.

[0030] The indexed set 20a of encoding parameter settings is cancelled simultaneously with or together with the payload portion 18, for portions of the payload portion 18 that do not belong to the sub-region 22, and the resulting reduced data stream is modified to maintain conformance. This approach is not followed by the embodiment of Figure 1. While such correlation modification of both the encoding parameter settings in the indexed set 20a on the one hand and the encoding parameter settings of the payload portion 18 on the other hand does not require a full decoding and encoding detour, the computational overhead to perform this correlation modification would nevertheless require a significant amount of analysis steps, etc.

[0031] Thus, the embodiment of FIG. 1 follows another approach in which the video data stream 10 includes, i.e., is provided with, information 50 that is not required to reconstruct the image 12 of the video from the payload portion 18. The information may include an indication of a given subregion and a replacement index and / or replacement parameters. For example, the information 50 may indicate a given subregion 22 with respect to its position within the image 12. The information 50 may indicate the location of the subregion 22, for example, in terms of tiles. Thus, the information 50 may identify a set of tiles 42 within each image 12 to form the subregion 22. The set of tiles 42 within each image 12 may be fixed between images 12. That is, the tiles formed within each image 12, the subregions 22 may be aligned with each other, and the tile boundaries of these tiles forming the subregions 22 may spatially vary between different images 12. It should be noted that the set of tiles is not limited to forming a contiguous rectangular tile subarray of the image 12. However, there may be an overlay-free, gap-free abutment of tiles within each of these images 12 forming a subregion 22, with this gapless and overlay-free abutment or juxtaposition forming a rectangular region. However, it should be appreciated that the representation 50 is not limited to showing the subregions 22 in tiles. It should be recalled that the use of tile subdivisions of the images 12 is, in any case, merely optional. The representation 50 may, for example, show the subregions 22 in units of samples or by some other means. In yet another embodiment, the location of the subregions 22 may even form default information known to participating network devices and decoders likely to handle the video data stream 10, with the information 50 simply indicating the reduction possibilities associated with or present in the subregions 22. As already mentioned above and as shown in FIG. 3 , the information 50 includes, in addition to the representation 52 of a given subregion, a replacement index 54 and / or replacement parameters 56.The replacement index and / or replacement parameter is for modifying the indexed set of encoding parameter settings, i.e. the set of encoding parameter settings indexed by an index within the payload portion 18, so that the indexed set of encoding parameter settings fits the payload portion of the reduced video data stream in which the payload portion 18 is modified on the one hand by removing those parts relating to parts of the image 12 outside the sub-region 22 and modifying the position indication 32 to relate to the periphery of the sub-region 22 rather than the periphery of the image 12.

[0032] To clarify the latter situation, reference is made to Figure 4, which shows a network device 60 configured, according to Figure 1, to receive and process video data stream 10 and derive therefrom a reduced video data stream 62. The term "reduced" in "reduced video data stream" 62 refers to two things: first, the fact that reduced video data stream 62 corresponds to a lower bit rate compared to video data stream 10, and second, the images into which reduced video data stream 62 is encoded are smaller than images 12 of video data stream 10, in that the smaller images of reduced video data stream 62 simply represent sub-regions 22 of images 12.

[0033] As will be explained in more detail below, to perform its task, method apparatus 60 includes a reader 64 configured to read information 50 from data stream 10, and a reducer 66 that performs a reduction or expansion process based on information 50, as will be explained in more detail below.

[0034] 5 illustrates the functionality of network device 60 in an exemplary case using permutation index 54 in information 50. Specifically, as shown in FIG. 5, network device 60 uses information 50 to remove 68 from payload portion 18 references to portions 70 of data stream 10 relating to regions of image 12 that are not related to subregion 22, i.e., outside subregion 22. Removal 68 can be performed, for example, on a slice basis, with reducer 66 indicating slices 34 in payload portion 18 that are not related to subregion 22 based on location indications or slice addresses in the slice headers of slices 34, on the one hand, and indication 52 in information 50, on the other hand.

[0035] In the example of FIG. 5 , where information 50 includes replacement indexes 54, parameter setting portion 16 of video data stream 10 includes, in addition to indexed set 20a of encoding parameter settings, an unindexed set 20b of encoding parameter settings that is not referenced or indexed by index 48 in payload portion 18. In performing the reduction, reducer 66 replaces index 48 in data stream 10 with replacement index 54 according to the replacement shown in FIG. 5 using curve 72, as exemplarily shown in FIG. 5 . By replacing index 48 with replacement index 54, redirection 72 occurs according to which indexes included in the payload portion of reduced video data stream 62 reference or index second set 20b of encoding parameter settings, resulting in first set 20a of encoding parameter settings no longer being indexed. Thus, redirection 72 can include reducer 66 removing 74 from parameter setting portion 16 the no longer indexed set 20a of encoding parameter settings.

[0036] Reducer 66 also modifies the position indications 32 within payload portion 18 as measured relative to the perimeter of predetermined sub-region 22. This modification is illustrated in FIG. 5 , with the change of merely one illustratively represented position indication 32 from data stream 10 to reduced video data stream 62 being shown schematically by curved arrows 78 indicating the position indications 32 in data stream 10 without hatching, while showing the position indications 32 in reduced video data stream 62 in a cross-hatched manner.

[0037] 5, network device 60 can obtain reduced video data stream 62 in a relatively low-complexity manner. The laborious task of properly adapting set of encoding parameter settings 20b to properly parameterize or adapt quantity / size 30, position indication 32, and decoding order 44 of payload portion 18 of reduced video data stream 62 can be performed elsewhere, such as in encoder 80, representatively shown using the dashed box in FIG. 1. An alternative is to change the order between evaluation of information 50 and reduction by reducer 66, as will be explained further below.

[0038] 6 illustrates a situation in which reduced video data stream 62 is provided to video decoder 82 to show that reduced video data stream 62 is encoded into video 84 of smaller image 86, i.e., image 86 of a smaller size than image 12 and showing only a sub-region 22 thereof. Reconstruction of video 84 therefore occurs by video decoder 82 decoding reduced video data stream 62. As described with reference to FIG. 5, reduced video data stream 62 has reduced payload portion 18 encoded into smaller image 86 as parameterized by or correspondingly described second set of encoding parameter settings 20b.

[0039] Video encoder 80 may encode image 12 into video data stream 10, for example, according to the encoding constraints described above with respect to FIG. 1 in relation to subregion 22. Encoder 80 may perform this encoding using, for example, an appropriate rate-distortion optimization function. As a result of this encoding, payload portion 18 indexes set 20a. Furthermore, encoder 80 generates set 20b. To this end, encoder 80 may, for example, adapt image size 26 and tile structure 38 from their values ​​in set 20a to correspond to the size and occupied tile set of subregion 22. Beyond that, encoder 80 essentially performs the reduction process as described above with respect to FIG. 5 itself and calculates buffer timing 46 so that a decoder, such as video decoder 82, can correctly manage the encoded image buffer using the thus calculated buffer timing 46 in second set 20b of encoding parameter settings.

[0040] 7 illustrates an alternative mode of operation of the network device, namely, the use of substitution parameters 56 within information 50. According to this alternative, as shown in FIG. 7, parameter setting portion 16 simply includes indexed set of encoding parameter settings 20a, and re-indexing or redirection 72 and set remover 74 do not need to be performed by reducer 66. However, instead, reducer 66 uses substitution parameters 56 obtained from information 50 to adjust 88 indexed set of encoding parameter settings 20a to become set 20b of encoding parameter settings. According to this alternative, reducer 66 performing steps 68, 78, and 88 to derive reduced video data stream 62 from original video data stream 10 is a relatively uncomplicated operation.

[0041] In other words, in FIG. 7, substitution parameters 56 may include, for example, one or more of image size 26, tile structure 38 and / or buffer timing 46.

[0042] 5 and 7, it should be noted that there may also be a mixture of both alternatives, with information 50 including both substitution indexes and substitution parameters. For example, the coding parameter settings subject to change from set 20a to 20b may be distributed over or comprised of different parameter set slices, such as SPS, PPS, VPS, etc. Different ones of these slices may therefore be subjected to different processing, e.g., according to FIG. 5 or FIG. 7.

[0043] With regard to the task of modifying 78 the position indication, it should be noted that this task must be performed relatively frequently, e.g., for each payload slice of slices 34 in payload portion 18, but the calculation of the new replacement value of position indication 32 is relatively uncomplicated. For example, the position indication may indicate a position in terms of horizontal and vertical coordinates, and modification 78 may calculate the new coordinates of the position indication by, e.g., forming a subtraction between the corresponding coordinates of the original position indication 32 and the offset of subregion 22 relative to the upper left corner of data stream 10 and image 12. Alternatively, position indication 32 may indicate a position in some suitable units, e.g., in units of coding blocks such as tree root blocks in which image 12 is regularly divided into rows and columns, e.g., using some linear scale that follows the decoding order described above. In such a case, the position indication is newly calculated in step 78 taking into account only the coding order of these code blocks within subregion 22. In this regard, it should also be noted that the reduction / extraction process just outlined for forming a reduced video data stream 62 from a video data stream 10 is also suitable for forming a reduced video data stream 62 in such a manner that the smaller images 86 of the video 84 encoded in the reduced video data stream 62 represent the sections 22 in a spatially stitched manner, and the same image content of the images 84 may be arranged within the images 12 of the sub-regions 22 in spatially different ways.

[0044] 6, it should be noted that the video decoder 82 shown in FIG. 6 may or may not be capable of decoding video data stream 10 to reconstruct images 12 of video 14. A reason that video decoder 82 may not be able to decode video data stream 10 may be, for example, that the profile level of video decoder 82 is sufficient to handle the size and complexity of reduced video data stream 62, but insufficient to decode original video data stream 10. However, in principle, both data streams 62 and 10 can be fitted to one video codec by the above-mentioned appropriate adaptation of the indexed sets of coding parameter settings by re-indexing and / or parameter adjustment.

[0045] After describing a fairly general embodiment for video stream reduction / extraction for a particular sub-region of an image of a video data stream to be reduced, the above motivation and problem description for such extraction for HEVC is resumed below and specific examples for implementing the above-described embodiment are provided.

[0046] 1. Signaling Aspects for Single Layer Subregions

[0047] 1.1. Parameter Set: The following aspects of the parameter set need to be adjusted when extracting spatial subsets: VPS: No normative information on single-layer coding SPS: Level Information Image size Cut or fit window information Buffering and timing information (i.e., HRD information) Potentially adding more VUI (Video Usability Information) items such as motion_vectors_over_pic_boundaries_flag, min_spatial_segmentation_idc, etc. PPS: Spatial segmentation information, i.e., tile information about the amount and size of tiles in horizontal and vertical directions

[0048] Signaling Embodiments Signaling 1A: For each TMCTS, the encoder can transmit additional unused (i.e., not activated at all) VPS, SPS, and PPS in-band (i.e., as respective NAL units) and provide their mapping to the TMCTS in a Supplemental Enhancement Information (SEI) message.

[0049] An example syntax / semantics for the Signaling 1A SEI is shown in FIG.

[0050] Syntax element 90 is optional as it can be derived from an image parameter set identifier.

[0051] The semantics are as follows: num_extraction_information_sets_minus1 indicates the number of information sets contained in a given Signaling1A SEI that should be applied in the sub-image extraction process.

[0052] num_applicable_tile_set_identifiers_minus1 indicates the number of mcts_id values ​​of tile sets for which the following i-th information set applies to the subimage extraction process.

[0053] mcts_identifier[i][k] indicates the value of mcts_id of the tile set for which all num_applicable_tile_set_identifers_minus1 apply the next i-th information set to the sub-image extraction process + 1.

[0054] num_mcts_pps_replacements[i] indicates the number of pps identifier replacements to signal in the Signaling1A SEI for the tile set with mcts_id equal to mcts_id_map[i].

[0055] mcts_vps_id[i] indicates that the mcts_vps_idx[i]th video parameter set is used in the sub-image extraction process for the tile set with mcts_id equal to mcts_id_map[i].

[0056] mcts_sps_id[i] indicates that the mcts_sps_idx[i]th sequence parameter set is used in the sub-image extraction process for the tile set with mcts_id equal to mcts_id_map[i].

[0057] mcts_pps_id_in[i][j] indicates the jth value of the num_mcts_pps_replacements[i] pps identifier in the slice header syntax structure of the tile set with mcts_id equal to mcts_id_map[i] that should be replaced in the sub-image extraction process.

[0058] mcts_pps_id_out[i][j] indicates the jth value of the num_mcts_pps_replacements pps identifier in the slice header syntax structure of the tile set with mcts_id equal to mcts_id_map[i] to replace the pps identifier equal to the value mcts_pps_id_in[i][j] in the sub-image extraction process.

[0059] Signaling 1B: The encoder can send VPS, SPS, and PPS for each TMCTS, and can map them to all TMCTSs included in the container-style SEI.

[0060] An example syntax / semantics for the Signaling 1B SEI is shown in FIG.

[0061] The yellow syntax elements 92 are optional as they can be derived from the image parameter set identifier.

[0062] The semantics are as follows: num_vps_in_message_minus1 indicates the number of vps syntax structures in a given Signaling1B SEI that should be used in the subimage extraction process.

[0063] num_sps_in_message_minus1 indicates the number of sps syntax structures in a given Signaling1B SEI that should be used in the subimage extraction process.

[0064] num_pps_in_message_minus1 indicates the number of pps syntax structures in a given Signaling1B SEI that should be used in the subimage extraction process.

[0065] num_extraction_information_sets_minus1 indicates the number of information sets contained in a given Signaling1B SEI that should be applied in the sub-image extraction process.

[0066] num_applicable_tile_set_identifiers_minus1 indicates the number of mcts_id values ​​of tile sets to which the subsequent i-th information set applies for the subimage extraction process.

[0067] mcts_identifier[i][k] indicates all num_applicable_tile_set_identifers_minus1 plus 1 of the mcts_id value of the tile set for which the next i-th information set applies to the sub-image extraction process.

[0068] mcts_vps_idx[i] indicates that for the tile set with mcts_id equal to mcts_id_map[i] in the sub-image extraction process, the mcts_vps_idx[i]th video parameter set signaled in the Signaling1B SEI is used.

[0069] mcts_sps_idx[i] indicates that the mcts_sps_idx[i]th sequence parameter set signaled in the Signaling1B SEI in the sub-image extraction process is used for the tile set with mcts_id equal to mcts_id_map[i].

[0070] num_mcts_pps_replacements[i] indicates the number of pps identifier replacements to signal in the Signaling1B SEI for the tile set with mcts_id equal to mcts_id_map[i].

[0071] mcts_pps_id_in[i][j] indicates the jth value of the num_mcts_pps_replacements[i] pps identifier in the slice header syntax structure of the tile set with mcts_id equal to mcts_id_map[i] to be replaced in the sub-image extraction process.

[0072] mcts_pps_idx_out[i][j] indicates that in the sub-image extraction process, the picture parameter set with a pps identifier equal to mcts_pps_id_in[i][j] should be replaced with the mcts_pps_idx_out[i][j]th signaled picture parameter set in the Singalling1C SEI.

[0073] Signaling 1C: The encoder may provide parameter set information related to non-derivable TMCTS (essentially additional buffering / timing (HRD) parameters and mapping to applicable TMCTS in the SEI).

[0074] An exemplary syntax / semantics for the Signaling 1C SEI is shown in FIG. The HRD information in the next SEI is configured to allow the extraction process to replace consecutive blocks of syntax elements in the original VPS with respective consecutive blocks of syntax elements from the SEI.

[0075] num_extraction_information_sets_minus1 indicates the number of information sets contained in a given signaling 1C SEI that are applied in the sub-image extraction process. num_applicable_tile_set_identifiers_minus1 indicates the number of mcts_id values ​​of tile sets to which the following i-th information set applies for the sub-image extraction process.

[0076] mcts_identifier[i][k] denotes all num_applicable_tile_set_identifers_minus1 plus one of the mcts_id values ​​of the tilesets for which the next i-th set of information applies to the sub-image extraction process.

[0077] mcts_vps_timing_info_present_flag[i] equal to 1 indicates that mcts_vps_num_units_in_tick[i], mcts_vps_time_scale[i], mcts_vps_poc_proportional_to_timing_flag[i] and mcts_vps_num_hrd_parameters[i] are present in the VPS. mcts_vps_timing_info_present_flag[i] equal to 0 indicates that mcts_vps_num_units_in_tick[i], mcts_vps_time_scale[i], mcts_vps_poc_proportional_to_timing_flag[i] and mcts_vps_num_hrd_parameters[i] are not present in the Signaling1C SEI.

[0078] mcts_vps_num_units_in_tick[i] is the number of time units of the clock running at a frequency of mcts_vps_time_scale Hz, which corresponds to one increment (called a clock tick) of the clock tick counter. The value of mcts_vps_num_units_in_tick[i] must be greater than 0. A clock tick (in seconds) is equal to mcts_vps_num_units_in_tick divided by mcts_vps_time_scale. For example, if the picture rate of a video signal is 25 Hz, mcts_vps_time_scale is equal to 27,000,000 and mcts_vps_num_units_in_tick is equal to 1,080,000, so a clock tick may be 0.04 seconds.

[0079] mcts_vps_time_scale[i] is the number of time units passing in a second. For example, a time coordinate system that measures time using a 27MHz clock has a vps_time_scale of 27000000. The value of vps_time_scale must be greater than 0.

[0080] mcts_vps_poc_proportional_to_timing_flag[i] equal to 1 indicates that the image order count value of each image in the CVS that is not the first image in the CVS, in decoding order, is proportional to the image output time relative to the output time of the first image in the CVS. mcts_vps_poc_proportional_to_timing_flag[i] equal to 0 indicates that the image order count value of each image in the CVS that is not the first image in the CVS, in decoding order, may or may not be proportional to the image output time relative to the output time of the first image in the CVS.

[0081] mcts_vps_num_ticks_poc_diff_one_minus1[i] plus 1 specifies the number of clock ticks that correspond to a difference in picture order count values ​​equal to 1. The value of mcts_vps_num_ticks_poc_diff_one_minus1[i] must be in the range 0 to 2-2.

[0082] mcts_vps_num_hrd_parameters[i] specifies the number of hrd_parameters() syntax structures present in the i-th entry of the Signaling1C SEI. The value of mcts_vps_num_hrd_parameters must be in the range of 0 to vps_num_layer_sets_minus1 + 1.

[0083] mcts_hrd_layer_set_idx[i][j] specifies the index into the list of layer sets specified by the VPS of the ith entry in the Signaling1C SEI for the layer set to which the jth hrd_parameters() syntax structure in the Signaling1C SEI applies and which is used in the sub-image extraction process. The value of mcts_hrd_layer_set_idx[i][j] must be in the range of (vps_base_layer_internal_flag?0:1) to vps_num_layer_sets_minus1. It is a bitstream conformance requirement that the value of mcts_hrd_layer_set_idx[i][j] must not equal the value of hrd_layer_set_idx[i][k] for any value of j not equal to k.

[0084] mcts_cprms_present_flag[i][j] equal to 1 specifies that the HRD parameters common to all sublayers are present in the jth hrd_parameters() syntax structure of the ith entry of the Signaling1C SEI. mcts_cprms_present_flag[i][j] equal to 0 specifies that the HRD parameters common to all sublayers are not present in the ith hrd_parameters() syntax structure of the ith entry of the Signaling1C SEI and are derived to be the same as the (i-1)th hrd_parameters() syntax structure of the ith entry of the Signaling1C SEI. mcts_cprms_present_flag[i][0] is inferred to be equal to 1.

[0085] Since the above HRD information is VPS related, signaling of similar information for SPS VUI HRD parameters can be implemented in the same way, for example by extending the above SEI or as a separate SEI message.

[0086] It is worth noting that further embodiments of the present invention may use the mechanisms implemented by signaling 1A, 1B and 1C as an extension of other bitstream syntax structures or parameter sets, such as VUI.

[0087] SEI Message The occurrence of any of the following SEI messages in the original video bitstream may require a reconciliation mechanism to avoid inconsistencies after TMCTS extraction: HRD related buffering period, picture timing and decode unit information SEI Pan Scan SEI * FramePackingArrangement * SEI DecodedPictureHash SEI (Decoded image hash SEI) TMCTS SEI

[0088] Signaling embodiment: Signaling 2A: The encoder can provide appropriate permutations of the SEIs associated with the TMCTS within the container-style SEI for all TMCTSs. Such signaling can be combined with the embodiment of Signaling 1C and is shown in Figure 11. In other words, in addition to or instead of the above, a video data stream 10 representing a video 14 can include a payload portion 18 in which an image 12 of the video is encoded, and a supplemental enhancement information message indicating supplemental enhancement information corresponding to the payload portion 18, or more precisely, how the image 12 of the video is encoded in the payload portion 18, and further including a replacement supplemental enhancement information message for replacing the information 50 including a representation 52 of a predetermined sub-region 22 of the image 12 and the supplemental enhancement information message, the replacement supplemental enhancement information message selecting to remove 68 a portion 70 of the payload portion 18 and to modify 78 a position indication 32 within the payload portion 18 to reference an area of ​​the image 12 outside the predetermined sub-region 22 and to indicate a position in place of the image 12 as measured from the periphery of the predetermined sub-region 22, such that the replacement supplemental enhancement information message indicates replacement supplemental enhancement information corresponding to the reduced payload portion, i.e., how the sub-area-specific picture 86 is encoded in the reduced payload portion 18. In addition to or as an alternative to the parameter generation described above, parameter setter 80a generates supplemental enhancement information messages that are subject to potential replacement by replacement SEI messages, which replacement is performed by network device 60 in addition to or as an alternative to the redirection and / or adjustment described above.

[0089] all_tile_sets_flag equal to 0 specifies that the applicable_mcts_id[0] list is specified by wapplicable_mcts_id[i] for all tile sets defined in the bitstream. all_tile_sets_flag equal to 1 specifies that the list applicable_mcts_id[0] consists of all values ​​of nuh_layer_id present in the current access unit that are greater than or equal to the nuh_layer_id of the current SEI NAL unit in ascending order of value. tile_sets_max_temporal_id_plus1 minus1 indicates the maximum temporal level that should be extracted in the sub-image extraction process for tilesets with mcts_id equal to an element of the array applicable_mcts_id[i].

[0090] num_applicable_tile_set_identifiers_minus1 plus 1 indicates the number of next applicable mcts ids that should use the next SEI message in the subimage extraction process.

[0091] mcts_identifier[i] indicates all num_applicable_tile_set_identifiers_minus1 values ​​of mcts_id for which the next SEI message should be inserted when extracting the respective tileset with mcts_id equal to applicable_mcts_id[i] using the tileset sub-image extraction process.

[0092] 2. Sub-image extraction process:

[0093] The details of the extraction process will obviously depend on the signaling scheme applied.

[0094] Constraints on the tile setup and TMCTS SEI, especially the extracted TMCTS, must be formulated to guarantee conforming output. The presence of any of the above signaling embodiments in the bitstream indicates a guarantee that the encoder will follow the constraints formulated below while creating the video bitstream.

[0095] input: Bitstream ·Target MCTS identifier MCTSIdTarget. Target layer identifier list layerIdListTarget.

[0096] Constraints or bitstream requirements: tiles_enabled_flag equals 1. ·num_tile_columns_minus1> 0 || num_rows_minus1> 0. · A TMCTS SEI message with mcts_id[i] equal to MCTSIdTarget exists and is associated with every image that should be output. · A TMCTS with mcts_id [i] equal to MCTSIdTarget exists in the TMCTS SEI. · The appropriate level for a TMCTS with mcts_id[i] equal to MCTSIdTarget shall be indicated either by the TMCTS SEI syntax elements mcts_tier_level_idc_present_flag[i], mcts_tier_idc[i], mcts_level_idc[i] or by one of the signaling variables 1A or 1B above. · The HRD information of the TMCTS is present in the bitstream via one of the signaling variables 1A, 1B or 1C. ·All rectangles in the TMCTS with mcts_id[i] equal to MCTSIdTarget have the same height and / or width in luma samples.

[0097] process: Remove all tile NALUs that do not exist in the tileset associated with mcts_id[i] equal to MCTSIdTarget. · Replace / adjust parameter sets according to signaling 1X. Adjust the remaining NALU slice headers as follows:

[0098] ○ Adjust slice_segment_address and first_slice_segment_in_pic_flag to create a common image plane from all rectangles in the tileset. Adjust pps_id as needed. · In the presence of signaling 2A, it removes or replaces the SEI.

[0099] As an alternative embodiment, the constraints or bitstream requirements described above as part of the extraction process can take the form of dedicated signaling in the bitstream, for example separate SEI messages or VUI indications, the presence of which is a prerequisite for them on the extraction process.

[0100] 2. Multi-layer

[0101] In some scenarios, there may be interest in layered codecs, for example to provide different qualities per region. It may be interesting to provide a larger spatial region at a lower layer quality, so that if required by the user and some specific region of wide-angle video is not available at the higher layer, the content can be upsampled at the lower layer and presented together with the content at the higher layer. The degree to which the video region of the lower layer extends the video region of the higher layer should be allowed to vary depending on the usage situation.

[0102] In addition to the TMCTS SEI described, in the layered extension of the HEVC specification (i.e., Annex F), an ILCTS (Inter-layer Constrained Tile Sets) SEI message is specified, which indicates similar constraints of nature for inter-layer prediction. For reference, a syntax table is shown in Figure 12. Therefore, as a further part of the present invention, a similar extraction process is implemented for layered coded video bitstreams that takes additional information into account.

[0103] The main difference from the signaling and processing disclosed above when considering the signaling aspects of multi-layer sub-pictures is that the target data portion of the bitstream is no longer identified by a single value of the mcts_id identifier, but instead the identifier of the layer set and, if applicable, multiple identifiers of the TMCTS within each containing layer and, if applicable, each ILCTS identifier between the contained layers form a multi-dimensional vector that identifies the target portion of the bitstream.

[0104] The present invention embodiments are variations of the single layer signaling 1A, 1B, 1C and 2A disclosed above, in which the syntax element mcts_identifier[i][k] is replaced by the multidimensional identifier vector described.

[0105] Furthermore, the encoder constraints or bitstream requirements are extended as follows:

[0106] 2.2. Extraction Process:

[0107] input: A multidimensional identifier vector consisting of: Target layer set layerSetIdTarget Target layer TMCTS identifier MCTSIdTarget_Lx for at least the highest layer in the set of layers with identifier layerSetIdTarget Target ILCTS identifier LTCTSIdTarget_Lx_refLy corresponding to Bitstream

[0108] Bitstream Request In addition to those defined for the single layer case:

[0109] A TMCTS SEI message with mcts_id[i] equal to MCTSIdTarget_Lx exists for each layer Lx, and an ILCTS SEI message with ilcts_id[i] equal to ILTCTSIdTarget_Lx_refLy exists for each layer Lx for any used reference layer Ly included in the bitstream and layer set layerSetIdTarget.

[0110] To exclude the presence of reference samples that are missing from the extracted bitstream portion, the TMCTS and ILCTS that define the bitstream portion must further satisfy the following constraints:

[0111] For each reference layer A with tile set tsA associated with mcts_id[i] equal to MCTSIdTarget_LA: The tiles in layer A that make up tsA are the same tiles associated with the tile set with ilcts_id[i] equal to ILTCTSIdTarget_LA_refLy. · For each referenced layer B having a tileset tsB associated with mcts_id[i] equal to MCTSIdTarget_LB: The tiles of layer B that make up tsB are entirely contained in the associated reference tileset, denoted ilcts_id[i] ILCTSIdTarget_Lx_refLB.

[0112] process For each layer x: mcts id [i] remove all tile NAL units that are not in the tile set with identifier MCTS Id Target_Lx.

[0113] Before moving on to the next aspect of the present application, a brief note should be made regarding the above-mentioned possibility that the sub-region 22 may be composed of a set of tiles, the relative position of which within the image 12 may differ from the relative position of these tiles to each other within the smaller image 86 of the video 84 represented by the reduced video data stream 62.

[0114] Figure 13 shows an image 12 subdivided into an array of tiles 42, listed in decoding order with capital letters A through I. For illustrative purposes only, Figure 13 shows an exemplary image and its subdivision into 3x3 tiles 42. Assume that image 12 is encoded into video data stream 10 such that coding dependencies do not cross tile boundaries 40, and that coding dependencies include not only intra-picture spatial interdependencies but are also limited, for example, by temporal independence. Thus, tiles 42 in the current image depend only on themselves or on co-located tiles in previously coded / decoded images, i.e., temporal reference images.

[0115] In this situation, subregion 22 may be composed of a set of discontinuous tiles 42, such as the set of tiles [D, F, G, I]. Because of their mutual independence, image 86 of video 84 may depict subregion 22 such that the participating tiles are spatially positioned differently within image 86. This is illustrated in FIG. 13: image 86 is also encoded into reduced video data stream 62 in units of tiles 42, but the tiles forming subregion 22 of image 12 are spatially positioned differently relative to one another within image 86. In the example of FIG. 13, tiles 42 forming subregion 22 occupy tiles on opposite sides of image 12, as if image 12 were depicting a horizontal panoramic view; these opposing side tiles actually represent adjacent portions of the panoramic scene. However, in image 86, tiles within each tile row switch positions relative to their relative positions in image 12. That is, for example, tile F appears to the left of tile D compared to the mutual horizontal positions of tiles D and F in image 12.

[0116] Before proceeding to the next aspect of the present application, it should be noted that neither the tiles 42 nor the sections 22 need be coded into the image 12 in the manner described above in which coding dependencies are limited so as not to cross their boundaries. Of course, this restriction relaxes the above-described concept of reducing / extracting the video data stream, but such coding dependencies tend to only affect small edge portions along the boundaries of the sub-regions 22 / tiles 42, depending on the application in which distortions at these edge portions are acceptable.

[0117] Furthermore, it should be noted that the above-described embodiment shows the possibility of extending current video codecs to newly include the described compliance points, i.e., to include the possibility of reducing a video stream to a reduced video stream related to a sub-region 22 of the original image 12 while only maintaining conformance, and to this end, the information 50 is exemplarily hidden in a SEI message, a VUI, or a parameter set extension, i.e., in a portion of the original video data stream that can be skipped by the decoder according to likes and dislikes. However, the information 50 can alternatively be conveyed within the video data stream in a standard portion. That is, new video codecs can be configured from the beginning to include the described compliance points.

[0118] Furthermore, for completeness, a further specific example of the above-described embodiment is described, which shows the possibility of extending the HEVC standard to implement the above-described embodiment. For this purpose, a new SEI message is provided. In other words, a modification to the HEVC specification is described that allows for the extraction of motion constrained tile sets (MCTS) as individual HEVC-compliant bitstreams. Two SEI messages are used and are described below.

[0119] The first SEI message, the MCTS Extraction Information Set SEI message, provides the syntax for the carriage of MCTS-specific replacement parameter sets and defines the extraction process in semantics. The second SEI message, the MCTS Extraction Information Nest SEI message, provides the syntax for MCTS-specific nested SEI messages.

[0120] Therefore, to include these SEI messages in the HEVC framework, the general SEI message syntax of HEVC is modified to include the new types of SEI messages.

[0121] JPEG2025170281000002.jpg101163

[0122] Thus, the list SingleLayerSeiList is set to include payload type values ​​3, 6, 9, 15, 16, 17, 19, 22, 23, 45, 47, 56, 128, 129, 131, 132 and 134 through 153 inclusive. Similarly, the lists VclAssociatedSeiList and PicUnitRepConSeiList are extended with new SEI message type numbers 152 and 153, which, of course, have been chosen for illustrative purposes only.

[0123] Table D.1 of HEVC further includes a hint to the new tape in the SEI message in the lifetime of the SEI message.

[0124] JPEG2025170281000003.jpg30161

[0125] Its syntax is as follows: The MCTS Extraction Information Set SEI message syntax can be designed as follows:

[0126] JPEG2025170281000004.jpg229162

[0127] As far as semantics are concerned, the MCTS Extract Information Set SEI message is an example of information 50 that uses substitution parameters 56 . The MCTS Extraction Information Set SEI message provides auxiliary information for performing sub-bitstream MCTS extraction as specified below to derive an HEVC compliant bitstream from a motion-constrained tile set, i.e., a set of tiles that form a fragment 84 of the whole image region. This information consists of a number of extraction information sets, each containing an identifier of the motion-constrained tile set to which the extraction information set applies. Each extraction information set contains the RBSP bytes of the replacement video parameter set, sequence parameter set, and picture parameter set used during the sub-bitstream MCTS extraction process. Let the set of pictures associatedPicSet be the pictures from the access unit containing the MCTS Extraction Information Set SEI message (inclusive), including the first of the following, in decoding order, up to but not including the first: - The next access unit in decoding order that contains an MCTS Extraction Information Set SEI message. - Next IRAP picture in decoding order with NoRaslOutputFlag equal to 1. - The next IRAP access unit in decoding order with NoClrasOutputFlag equal to 1. The scope of the MCTS Extraction Information Set SEI message is the set of images associatedPicSet. If an MCT Extraction Information Set tile set SEI message is present for any image in associatedPicSet, then a temporally motion constrained tile set SEI message shall be present for the first image in associatedPicSet in decoding order and MAY be present for other images in associatedPicSet. The temporally motion constrained tile set SEI message MUST have mcts_id[] equal to mcts_identifer[] for all images in associatedPicSet. If an MCTS Extraction Information Set Tile Set SEI message is present for any image in associatedPicSet, then an MCTS Extraction Information Set SEI message MUST be present for the first image in associatedPicSet in decoding order, and MAY also be present for other images in associatedPicSet. For any PPS that is active for any picture in the associatedPicSet, if tiles_enabled_flag is equal to 0, then the MCTS Extraction Information Set SEI message shall not be present for any picture in the associatedPicSet. The MCTS Extraction Information Set SEI message shall not be present for any image in the associatedPicSet unless all PPSs that are active for any image in the associatedPicSet have the same values ​​for the syntax elements num_tile_columns_minus1, num_tile_rows_minus1, uniform_spacing_flag, column_width_minus1[i] and row_height_minus1[i]. NOTE 1 - This constraint is similar to the constraint related to tiles_fixed_structure_flag being equal to 1, and it is desirable for tiles_fixed_structure_flag to be equal to 1 if an MCTS Extraction Information Set SEI message is present (although this is not required). If there are multiple MCTS Extraction Information Set SEI messages for the images in the associatedPicSet, they shall contain the same content. A NAL unit containing tiles belonging to tile set tileSetA must not contain tiles that do not belong to tile set tileSetA. The number of MCTS Extraction Information Set SEI messages in each access unit must not exceed five. num_extraction_info_sets_minus1 plus 1 indicates the number of extraction info sets included in the MCTS Extraction Info Sets SEI message that are applied to the mcts extraction process. The value of num_extraction_info_sets_minus1 is between 0 and 2. 32 Must be in the range -2. The i-th extracted information set is assigned an MCTS extracted information set identifier value equal to i. num_associated_tile_set_identifiers_minus1[i] plus 1 indicates the number of mcts_id values ​​for tile sets in the i-th extraction info set. The value of num_extraction_info_sets_minus1[i] ranges from 0 to 2. 32 Must be in the range -2. mcts_identifier[i][j] identifies the j-th tile set with mcts_id equal to mcts_identifier[i][j] associated with the ith extracted information set. The value of mcts_identifier[i][j] is between 0 and 2. 32 Must be in the range -2. num_vps_in_extraction_info_set_minus1[i] plus 1 indicates the number of replacement video parameter sets in the i-th extraction information set. The value of num_vps_in_extraction_info_set_minus1[i] must be in the range of 0 to 15. vps_rbsp_data_length[i][j] indicates the number of bytes vps_rbsp_data_bytes[i][j][k] of the next j-th replacement video parameter set in the i-th extraction information set. num_sps_in_extraction_info_set_minus1[i] plus 1 indicates the number of permutation sequence parameter sets in the i-th extraction information set. The value of num_sps_in_extraction_info_set_minus1[i] must be in the range 0 to 15. sps_rbsp_data_length[i][j] indicates the number of bytes of sps_rbsp_data_bytes[i][j][k], which is the next j-th replacement sequence parameter set in the i-th extraction information set. num_pps_in_extraction_info_set_minus1[i] plus 1 indicates the number of replaced image parameter sets in the i-th extraction information set. The value of num_pps_in_extraction_info_set_minus1[i] must be in the range of 0 to 63. pps_nuh_temporal_id_plus1[i][j] specifies the temporal identifier for generating the PPS NAL unit associated with the PPS data specified in the PPS RBSP specified in pps_rbsp_data_bytes[i][j][] for the j-th replaced picture parameter set for the i-th extraction information set. pps_rbsp_data_length[i][j] indicates the number of bytes pps_rbsp_data_bytes[i][j][k] of the next j-th replaced image parameter set set in the i-th extraction information set. mcts_alignment_bit_equal_to_zero must be equal to 0. vps_rbsp_data_bytes[i][j][k] contains the k-th byte of the RBSP of the next j-th replacement video parameter set in the i-th extraction information set. sps_rbsp_data_bytes[i][j][k] contains the k-th byte of the RBSP of the next j-th substitution sequence parameter set in the i-th extraction information set. pps_rbsp_data_bytes[i][j][k] contains the k-th byte of the RBSP of the next j-th replacement image parameter set in the i-th extraction information set. The sub-bitstream MCTS extraction process is applied as follows. The bitstream inBitstream, the target MCTS identifier mctsIdTarget, the target MCTS extraction information set identifier mctsEISIdTarget, and the target highest TemporalId value mctsTIdTarget are taken as inputs to the sub-bitstream MCTS extraction process. The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream. The bitstream conformance requirement for an input bitstream is that any output sub-bitstreams that are the output of the process specified in this section along with the bitstream are conforming bitstreams. The output sub-bitstreams are derived as follows. The bitstream outBitstream is set to be the same as the bitstream inBitstream. The lists ausWithVPS, ausWithSPS and ausWithPPS are set to consist of all access units in outBitstream that contain VCL NAL units of type VPS_NUT, SPS_NUT and PPS_NUT. Remove all SEI NAL units that have nuh_layer_id equal to 0 and that contain non-nested SEI messages. NOTE 2 - A "smart" bitstream extractor may include appropriate non-nested SEI messages in the extracted sub-bitstream, provided that the SEI messages applicable to the sub-bitstream are present as nested SEI messages in the original bitstream's mcts_extraction_info_nesting(). -remove all types of NAL units from outBitstream: - A VCL NAL unit containing a tile that does not belong to a tileset with mcts_id[i] equal to mctsIdTarget. - Non-VCL NAL units with type VPS_NUT, SPS_NUT, or PPS_NUT. - Insert into all access units in the list ausWithVPS in the outBitstream num_vps_in_extraction_info_minus1 [mctsEISIdTarget] plus 1 NAL unit with type VPS_NUT generated from the VPS RBSP data in the mctsElSldTarget-th MCTS Extraction Info Set, i.e., vps_rbsp_data_bytes [mctsEISIdTarget] [j] [] for all values ​​of j in the range from 0 to num_vps_in_extraction_info_minus1 [mctsEISIdTarget]. For each generated VPS_NUT, nuh_layer_id is set equal to 0 and nuh_temporal_id_plus1 is set equal to 1. - Insert into all access units in the list ausWithSPS in the outBitstream num_sps_in_extraction_info_minus1[mctsEISIdTarget] plus 1 NAL unit with type SPS_NUT generated from the SPS RBSP data in the mctsEISIdTarget-th MCTS Extraction Info Set, i.e., sps_rbsp_data_bytes[mctsEISIdTarget][j][] for all values ​​of j in the range from 0 to num_sps_in_extraction_info_minus1[mctsEISIdTarget]. For each generated SPS_NUT, nuh_layer_id is set equal to 0 and nuh_temporal_id_plus1 is set equal to 1. - Insert into all access units in the list ausWithPPS in outBitstream NAL units with type PPS_NUT generated from the mctsEISIdTarget-th MCTS extraction information set, i.e., the PPS RBSP data in pps_rbsp_data_bytes[mctsEISIdTarget][j][] for all values ​​of j in the range 0 to num_pps_in_extraction_info_minus1[mctsEISIdTarget], where pps_nuh_temporal_id_plus1[mctsEISIdTarget][j] is less than or equal to mctsTIdTarget. For each generated PPS_NUT, nuh_layer_id is set equal to 0 and nuh_temporal_id_plus1 is set equal to pps_nuh_temporal_id_plus1[mctsEISIdTarget][j] for all values ​​of j in the range from 0 to num_pps_in_extraction_info_minus1[mctsEISIdTarget] and pps_nuh_temporal_id_plus1[mctsEISIdTarget][j] is less than or equal to mctsTIdTarget. - Remove all NAL units with TemporalId greater than mctsTIdTarget from outBitstream. - For each remaining VCL NAL unit in outBitstream, adjust the slice segment headers as follows: - For the first VCL NAL unit in each access unit, set the value of first_slice_segment_in_pic_flag equal to 1, and set it equal to 0 otherwise. - Set the value of slice_segment_address according to the tile settings defined in the PPS where pps_pic_parameter_set_id is equal to slice_pic_parameter_set_id.

[0128] The MCTS Extract Information Nested SEI message syntax can be designed as follows:

[0129] JPEG2025170281000005.jpg127165

[0130] With regard to semantics, it should be noted that MCTS Extraction Information Nested SEI messages may be present in addition to or instead of MCTS Extraction Information Set SEI messages to form information 50 .

[0131] The MCTS Extract Information Nested SEI message carries nested SEI messages and provides a mechanism for associating the nested SEI messages with bitstream subsets corresponding to one or more motion-constrained tile sets. In the sub-bitstream MCTS extraction process specified in the semantics of the MCTS extraction information set SEI message, nested SEI messages contained in the MCTS extraction information nested SEI message may be used to replace non-nested SEI messages in the access unit containing the MCTS extraction information nested SEI message. all_tile_sets_flag equal to 0 sets the mcts_identifier list to consist of mcts_identifier[i]. all_tile_sets_flag equal to 1 indicates that the list mcts_identifier[i] consists of all values ​​of mcts_id[] in temporal_motion_constrained_tile_sets SEI messages present in the current access unit.

[0132] num_associated_mcts_identifiers_minus1 plus 1 indicates the number of subsequent mcts_identifiers. The value of num_associated_mcts_identifiers_minus1 [i] is 0 to 2. 32 Must be in the range -2. mcts_identifier[i] indicates the tileset with mcts_id equal to mcts_identifier[i] associated with the next nested SEI message. The value of mcts_identifier[i] is between 0 and 2. 32 Must be in the range -2. num_seis_in_mcts_extraction_seis_minus1 plus 1 indicates the number of next nested SEI messages. mcts_nesting_zero_bit must be equal to 0.

[0133] It has already been indicated above that the evaluation or generation of information 50, i.e., the information guiding the parameter and / or SEI adaptation, can alternatively take place at outer encoder 80, i.e., outside the location where the actual encoding of image 12 into stream 10 takes place. According to such an alternative, data stream 10 may be transmitted accompanied by original parameters 20a and / or original SEI messages relating only to unreduced data stream 10. Optionally, information regarding one or more supported subregions 22 of image 12 may be present in video stream 10, although this is not required, as evaluation of information 50 may be based solely on evaluating the tile structure of stream 12 to determine one or more subregions. In doing so, the burden of evaluating information 50 is shifted from the encoder site to a location closer to the client, or even to a user site, such as immediately upstream of final decoder 82, while avoiding the need to transmit the complete, i.e., unreduced, data stream 10 by omitting transmission of portions 70 of payload portion 18 referencing regions of image 12 outside the desired subregions 22. The original encoding parameter set 20a and / or SEI message for the unreduced data stream 12 would of course be transmitted. The network entity 60 that performs the actual reduction or removal 68 of the portion 70 can be located immediately upstream of the entity that also performs the evaluation of the information 50. For example, a streaming device specifically downloads only the portion of the payload portion 18 of the data stream 10 that does not belong to the portion 70. For this purpose, some download methods, such as a manifest file, can be used. The DASH protocol can be used for this purpose. The evaluation of the information 50 can actually be performed in such a network device located before the decoder, simply as a preparation for the actual parameter adjustment and / or SEI replacement according to FIG. 3 or FIG. 5, respectively.Overall, the network device may include an interface for receiving a reduced data stream 10, according to the alternative described above, that includes fewer payload portions 18 (70) but still has parameters 20a that parameterize the complete encoding of the image 12 in the payload portion 18 including portion 70. The position indication 32 may still be the original position indication. That is, the received data stream is actually erroneously encoded. The received data stream corresponds to the data stream 62 shown in FIG. 5 with parameters 20b, but still has the original parameters 20a, and the position indication 32 is still incorrect. The index 54 within that stream references the parameters 20a within that stream and may be unchanged relative to the original encoding of the unreduced data stream. In fact, the unreduced original data stream may simply differ from the received original data stream due to the omission of portion 70. Optionally, one or more SEI messages are included, which reference the original encoding, e.g., the size of the original image 12 or other characteristics of the complete encoding. At the output of such a network device, a data stream 62 is output that is decoded by a decoder 82. Between this input and output are connected modules that adapt the SEI messages and / or parameters to fit the sub-area 12 that the inbound data stream 10 has already been reduced to, i.e., such a module must perform task 74, i.e., adjusting parameter 20a to become parameter 20b, and task 78, i.e., adjusting position indication 32 to refer precisely to the periphery of sub-area 22. Knowledge of the sub-area 22 for which reduction has been performed may be internal to such a network device, in the latter case specifically limiting the download of stream 10 to that part of payload portion 18 with reference to sub-area 22, or may be provided externally in the case of another network device that may be upstream of the network device, assuming the task of reduction 68.For a network device including a module between the input and the output, the same statements as those made above with respect to the network device of FIG. 4 apply, e.g., in terms of implementation in software, firmware, or hardware. To summarize the just-outlined embodiment, the embodiment relates to a network device configured to receive a data stream including a portion of a payload portion in which a video image is encoded. This portion corresponds to the result when a portion 70 is excluded from a payload portion 18 that refers to an area of ​​the image outside a predetermined subregion 22 of the image. The video image 12 is encoded in the payload portion 18 in a parameterized manner without exclusion using the encoding parameter settings of the parameter-configured portion of the data stream. That is, the parameter-configured portion of the received data stream correctly parameters the encoding of the image 12 in the payload portion 18 if no portion 70 is left. Additionally or alternatively, the video image 12 is encoded in the payload portion 18 in a manner that does not exclude and that matches the supplemental enhancement information indicated by the supplemental enhancement message of the data stream. That is, an SEI message optionally included by the received data stream actually coincides with the unreduced payload portion 18. The network device modifies 78 the position indication 32 in the payload portion to indicate a position measured from the periphery of the predetermined sub-region 22 instead of the image 12, adjusts the encoding parameter settings of the parameter setting portion and / or adjusts the supplemental enhancement information message so that the modified data stream has a sub-region-specific image 84 indicating part of the payload portion 18, i.e., all but 70, or in other words the predetermined sub-region of the image, coded to be correctly parameterized using the thus adjusted encoding parameter settings and / or coded to match the supplemental enhancement information indicated by the adjusted supplemental enhancement information message after adjustment. The generation of parameters 20b in addition to parameters 20a has been performed by encoder 80 according to the previous embodiment to result in a data stream carrying both parameter settings 20a and 20b.Here, in the above alternative, parameter 20a is changed to corresponding parameter 20b on the fly. Adjusting parameter 20a to parameter 20b requires modifying parameter 20a using knowledge of subregion 22. For example, image size 26 in setting 20a corresponds to the size of complete image 12, but after adjustment of setting 20b, image size must indicate the size of subregion 22 or image 86, respectively. Similarly, tile structure 38 in setting 20a corresponds to the tile structure of complete image 12, but after adjustment of setting 20b, tile structure 38 must indicate the tile structure of subregion 22 or image 86, respectively. Similar statements apply, for example, but not exclusively, with respect to buffer size and timing 46. If there are no retransmitted SEI messages in the received data stream, no adjustment is necessary. Alternatively, the SEI messages can simply be aborted instead of being adjusted.

[0134] With respect to the above embodiment, it should be noted that the adaptation of the auxiliary enhancement information may be related to buffer size and / or buffer timing data. In other words, the type of information in any present SEI that is matched or different between the original SEI and the replacement SEI to fit the decomposed or reduced video stream may be related, at least in part, to buffer size and / or buffer timing data. That is, the SEI data in stream 10 may have buffer size and / or buffer timing data related to the full encoding, while the replacement SEI data, which is conveyed in addition to the former as described with respect to FIG. 1 or generated on the fly as described in the previous paragraph, has buffer size and / or buffer timing data related to reduced stream 62 and / or reduced image 86.

[0135] The following description relates to a second aspect of the present application, namely, a concept for enabling more efficient transport of video data that does not fit the typical rectangular image shape of a video codec. As before, with respect to the first aspect, the following description begins with a kind of introduction, i.e., an exemplary description of an application in which such a problem may arise, in order to motivate the advantages resulting from the embodiments described below. However, it should also be noted that this preliminary description should not be understood as limiting the breadth of the embodiments described below. Furthermore, it should be noted that the aspects of the present application described next can also be advantageously combined with the embodiments described above. More details on this are also provided below.

[0136] The problem explained next arises from the different projections used for panoramic video, especially when processes such as the subregion extraction mentioned above are applied.

[0137] For example, in the following description we will use the so-called cubic projection. Cubic projection is a special case of rectilinear projection, also called gnomonic projection. This projection describes the transformation approximated for most conventional camera systems / lenses when an image representation of a scene is acquired. Straight lines in the scene are mapped to straight lines in the resulting image, as shown in Figure 14.

[0138] A cubic projection applies a rectangular parallelepiped projection to map the perimeter of a cube onto six faces, each with a 90° x 90° viewing angle from the center of the cube. The result of this type of cube projection is shown as image A in Figure 15. Other arrangements of the six faces in a common image are possible as well.

[0139] To now derive a more coding-friendly representation of the image A thus obtained (i.e., with fewer unused image regions 130 and rectangular shapes), the image patches can be displaced within the image, resulting in image B, as shown, for example, in FIG. 16.

[0140] From a system perspective, it is important to have an understanding of how various image patches in image B (Figure 16) are spatially related to the original (world-view) continuous representation of image A, i.e., additional information to derive image A (Figure 15) given the representation of image B. In particular, under the above-described circumstances, processing such as the above-described subregion extraction in the coding domain of Figures 1-13 is performed on the server side or on a network element. In this case, only the portion of image B shown via the isolated ROI 900 in Figure A is available to the end device. The end device must be able to remap the relevant region of a given (sub)video to the correct location and area expected by the end display / rendering device. A server or network device that modifies the encoded video bitstream according to the above-described extraction process can generate, add, or adjust the respective displacement signaling according to the modified video bitstream.

[0141] Therefore, the embodiments described later provide signaling indicating within a video bitstream (rectangle) group of samples of image B. Furthermore, the displacement of each group of samples with respect to the samples of image B in the horizontal and vertical directions. In further embodiments, the bitstream signaling includes explicit information about the resulting image size of image A. Furthermore, default luma and chroma values ​​for samples not covered by the displaced sample group or samples initially covered by the displaced sample group. Furthermore, some of the samples of image A may be initialized with the sample values ​​of the corresponding samples of image B.

[0142] An exemplary embodiment is shown in the syntax table of FIG.

[0143] A further embodiment utilizes tile structure signaling for indication of samples belonging to a group of displaced samples.

[0144] With reference to Figure 18, an embodiment of a data stream 200 according to a second aspect is shown. The data stream 200 has video 202 encoded therein. While the embodiment of Figure 18 may be modified to refer to one image encoded in the data stream 200 alone, to facilitate understanding of examples in which an embodiment according to the second aspect is combined with any embodiment of the other aspects of the present application, Figure 18 shows an example in which the data stream 200 is a video data stream rather than an image data stream. As just mentioned, the data stream 200 has an image 204 of the video 202 encoded therein.

[0145] However, the video data stream 200 also includes displacement information 206. The displacement information has the following meaning: In practice, the data stream 200 must convey image content within an image 204 with a non-rectangular perimeter. FIG. 18 shows an example of such image content at 200. That is, image content 208 indicates the image content that the data stream 200 should convey within one timestamp, i.e., one image 204. However, the perimeter 210 of the actual image content is non-rectangular. In the example of FIG. 18, the actual image content corresponds rather to a six-sided projection of a cube 212, whose sides are distinguished from one another by the normal distribution of the numbers 1 to 6 in FIG. 18, six on each of the six faces of the cube 212, i.e., the number of opposite faces equals seven in total. Each face thus represents one sub-region 214 of the actual image content 208, and may represent, for example, the appropriate projection of the sixth of a complete 3D panoramic scene onto the respective sub-region 214, which, according to this example, is square-shaped. However, as already mentioned above, the subregions 214 may be different shapes and may be arranged within the image content 208 in a manner other than a regular arrangement of rows and columns. In any event, the actual image content shown in 208 is non-rectangular in shape, and therefore the smallest possible rectangular target image area 216 that completely encompasses the cropped panoramic content 208 will have an unused portion 130, i.e., a portion that is not occupied by actual image content associated with the panoramic scene.

[0146] Therefore, in order to not "waste" image area within the image 204 of the video 202 carried within the data stream 200, the image 204 carries the full real-world image content 208 such that the spatial relative positions of the sub-regions 214 are changed relative to their positions within the target image region 216.

[0147] As shown in FIG. 18 , FIG. 18 illustrates an example in which four subregions, namely, subregions 1, 2, 4, and 5, should not be displaced when undistorting or congruently copying image 204 to target image region 216, while subregions 3 and 6 must be displaced. Illustratively, in FIG. 18 , the displacement is a pure translational displacement that can be described by a two-dimensional vector 218; however, according to alternative embodiments, more complex displacements can be selected, such as, for example, displacements that further include scaling and / or reflection (mirroring) and / or rotation of each subregion when transitioning between image 204 on the one hand and target image region 216 on the other. Displacement information 206 can indicate, for each of a set of at least one predetermined subregion of image 204, the displacement of the respective subregion within target image region 216 relative to an undistorted or undistorted copy of image 204 to target image region 216. For example, in the example of FIG. 18 , displacement information 206 can indicate the displacement of a set of subregions 214 that encompasses only subregions 3 and 6. Alternatively, the displacement information 206 can indicate a displacement for a subregion 214 or a set of displaced subregions 214 relative to some default point of the target region 216, such as its upper left corner. By default, the remaining subregions in the example image 204 of Figures 18, 1, 2, 4, and 5 can be treated as remaining intact upon undistorted or congruent copying onto the target image 216 relative to the default point, for example.

[0148] 19 shows that the displacement information can include a count 220 of subregions 214 to be displaced, a target image region size parameter indicating the size of target image regions 216, 222, and, for each of the n displacement subregions, coordinates 224 describing the displacement 218 of each subregion 214 to be displaced when mapping the subregions 214 in their original locations in the image 204 onto the target region 216, e.g., by measuring the displacement relative to or its relative position relative to the aforementioned default point. In addition to the coordinates 224, the information 206 can include a scaling 226 for each subregion 214 to be displaced, i.e., an indication of how each displaced subregion 214 should be scaled according to its respective coordinates 224 within the target image region 216 when mapping the respective subregion 214. The scaling 226 can result in an enlargement or reduction relative to the non-displaced subregions 214. Alternatively, horizontal and / or vertical reflections and / or rotations can be signaled to each subregion 214. Additionally, the displacement information 206 may include the coordinates of the upper left and lower right corners of each subregion, again measured relative to a default point or respective corner if the subregion 214 were to be mapped onto the target region 216 without displacement, so that each subregion 214 is displaced or positioned within the target region 216. This allows for positioning, scaling, and reflection to be signaled. Thus, the subregions 214 of the transmitted image 204 may be freely displaced relative to their original relative positioning within the image 204.

[0149] The displacement information 206 may have a scope, i.e., validity, for a time interval of the video 202 larger than one timestamp or one image 204, as in the case of, for example, a sequence of images 204 or the entire video 202. Furthermore, Figure 18 shows that the data stream 200 may optionally also include default filling information 228, indicating the default filling with which portions 130 of the target image region should be filled at the decoding side, i.e., that these portions 130 are not covered by any of the sub-regions 214, i.e., either the displaced or non-displaced sub-regions 214 of the image 204. In the example of FIG. 18, for example, subregions 1, 2, 4, and 5 form the non-displaced portion of image 204, which is shown unhatched in FIG. 18, while the remainder of image 204, i.e., subregions 3 and 6, which are shown hatched in FIG. 18, and all six of these subregions, do not cover the shaded portion 130 of target image region 216 after subregions 3 and 6 are displaced according to information 206 so that portion 130 is filled according to default fill information 228.

[0150] An encoder 230 suitable for generating data stream 200 is shown in FIG. 20. Encoder 230 simply accompanies or provides information 206 to data stream 200 indicating the displacement required to fill the target image region with image 204 to be encoded into data stream 200. FIG. 20 further illustrates that encoder 230 can generate data stream 200 such that it can be reduced in a compliance-preserving manner, such as data stream 10 of the embodiments of FIGS. 1-13. In other words, encoder 230 may be implemented to implement encoder 80 of FIG. 1. For illustrative purposes, FIG. 20 illustrates the case where image 204 is subdivided into an array of subregions 214, i.e., a 2×3 array of subregions according to the example of FIG. 18. Image 204 can correspond to image 12, i.e., an image of unreduced data stream 200. FIG. 20 exemplarily illustrates subregions 22 with respect to which encoder 230 can reduce data stream 200. As described with respect to FIG. 1, there may be several such subregions 22. By this means, a network device can reduce data stream 200 such that it simply extracts a portion of data stream 200 to obtain a reduced video data stream in which image 86 simply represents sub-region 22. Data stream 200 can be reduced, and the reduced video data stream associated with sub-region 22 is shown in FIG. 20 using reference numeral 232. While image 204 in unreduced video data stream 200 represents image 204, the image in reduced video data stream 232 simply represents sub-region 22. Encoder 230 can additionally provide video data stream 200 with displacement information 206' specific to sub-region 22 to provide a receiver, such as a decoder receiving reduced video data stream 232, with the ability to fill the target image region with the image content of reduced video data stream 232. A network device, such as network device 60 in FIG. 4, can remove displacement information 206 and simply carry over displacement information 206' to reduced video data stream 232 as a result of reducing video data stream 200 to produce reduced video data stream 232.By this means, a receiver such as a decoder can fill the image content of the reduced video data stream 232 related to the content of the sub-region 22 onto the image region of interest 216 shown in FIG.

[0151] Again, it should be emphasized that Figure 20 should not be understood as being limited to reducible video data streams: if video data stream 200 is reducible, different concepts may be used than those presented above with respect to Figures 1 to 13.

[0152] 21 illustrates a possible reception. The decoder 234 receives the video data stream 200 or the reduced video data stream 232 and reconstructs its image, i.e., an image simply showing the sub-region 22 based on the image 204 or the reduced video data stream 232, respectively. The receiver is a decoder, and in addition to the decoding core 234, it also comprises a displacer 236 that uses the displacement information of the video data stream 206 for the video data stream 200 or the reduced video data stream 206' for the reduced video data stream to fill the target image region 216 based on the image content. The output of the displacer 236 is thus the filled target image region 216 for each image of the respective data stream 200 or 232. As mentioned above, some parts of the target image region remain unfilled by the image content. Optionally, a renderer 238 may be connected after the displacer 236 or to the output of the displacer 236. The renderer 238 applies an injective projection to the target image formed by the target image region 216—or at least a subregion thereof within the fill region of the target image region—to form an output image or output scene 240 that corresponds to the currently viewed scene section. The injective projection performed by the renderer 238 may be the inverse of a cubic projection.

[0153] Thus, the above embodiments enable rectangular region-by-region packing of image data, such as for panoramic or semi-panoramic scenes. A specific syntax example is given below: Provided below is a syntax example in the form of pseudocode called RectRegionPacking(i), which specifies how a source rectangular region of the projected frame, i.e., 216, is packed into a destination rectangular region of the packed frame, i.e., 204. Horizontal mirroring and rotation of 90, 180, or 270 degrees is indicated, and vertical and horizontal resampling is inferred from the region width and height.

[0154] aligned(8)class RectRegionPacking(i)[ unsigned int(32)proj_reg_width [i]; unsigned int(32)proj_reg_height [i]; unsigned int(32)proj_reg_top [i]; unsigned int(32)proj_reg_left [i]; unsigned int(8)transform_type [i]; unsigned int(32)packed_reg_width [i]; unsigned int(32)packed_reg_height [i]; unsigned int(32)packed_reg_top [i]; unsigned int(32)packed_reg_left [i]; ]

[0155] The semantics are as follows:

[0156] proj_reg_width[i], proj_reg_height[i], proj_reg_top[i], and proj_reg_left[i] are in pixels, with width and height equal to proj_frame_width and proj_frame_height, respectively, in the projected frame, i.e., 216. i is the index for the respective region, i.e., tile 214, as compared to Figure 18. proj_reg_width[i] specifies the width of the i-th region in the projected frame. proj_reg_width[i] must be greater than 0. proj_reg_height[i] specifies the height of the i-th region in the projected frame. proj_reg_height[i] must be greater than 0. proj_reg_top[i] and proj_reg_left[i] specify the topmost sample row and leftmost sample column in the projected frame. Values ​​range from 0 or greater, inclusively indicating the top-left corner of the projection frame, to proj_frame_height and proj_frame_width, respectively, exclusive. proj_reg_width[i] and proj_reg_left[i] must be constrained such that proj_reg_width[i] + proj_reg_left[i] is less than proj_frame_width. proj_reg_height[i] and proj_reg_top[i] must be constrained such that proj_reg_height[i] + proj_reg_top[i] is less than proj_frame_height. If the projected frame 216 is stereoscopic, proj_reg_width[i], proj_reg_height[i], proj_reg_top[i], and proj_reg_left[i] are such that the area identified by these fields on the projected frame lies within a single constituent frame of the projected frame.transform_type[i] specifies the rotation and mirroring applied to the i-th region of the projection frame to map to the packed frame, and therefore the mapping must be reversed to map the respective region 214 back to region 216. Naturally, it can indicate a mapping from image 204 to region of interest 216. If transform_type[i] specifies both rotation and mirroring, the rotation is applied after the mirroring. Naturally, the reverse is also possible. According to the embodiment, the following values ​​are specified, although other values ​​may be reserved:

[0157] 1: no transformation, 2: horizontal mirror, 3: 180 degree rotation (counterclockwise), 4: horizontal mirror followed by 180 degree rotation (counterclockwise), 5: horizontal mirror followed by 90 degree rotation (counterclockwise), 6: 90 degree rotation (counterclockwise), 7: horizontal mirror followed by 270 degree rotation (counterclockwise), 8: 270 degree rotation (counterclockwise). Note that the values ​​correspond to the EXIF ​​image orientation tag.

[0158] packed_reg_width[i], packed_reg_height[i], packed_reg_top[i], and packed_reg_left[i] specify the width, height, topmost sample row, and leftmost sample column of the region within the packed frame, i.e., the region covered by tiles 214 in image 204. The rectangle specified by packed_reg_width[i], packed_reg_height[i], packed_reg_top[i], and packed_reg_left[i] shall not overlap with the rectangle specified by packed_reg_width[j], packed_reg_height[j], packed_reg_top[j], and packed_reg_left[j] for any value of j ranging from 0 to i-1, inclusive.

[0159] To summarize and generalize the above example, the embodiments further described above may differ in that for each region or tile 214 of image 214, two rectangular regions are shown, i.e., the region that each region or tile 214 covers within the target region 216 and the rectangular region that each region or tile 214 covers within the image region 204, along with mapping rules for mapping the image content, i.e., reflection and / or rotation, of each region or tile 214 between these two regions. Scaling can be signaled by signaling a pair of regions of different sizes.

[0160] The third aspect of the present application will now be described. The third aspect relates to an advantageous concept of distributing access points to a video data stream. In particular, access points associated with one or more sub-regions of an image encoded in the video data stream are introduced. Advantages resulting therefrom are described below. As with the other aspects of the present application, the description of the third aspect is accompanied by an introduction that explains the problems that arise. As with the description of the first aspect, this introduction exemplarily refers to HEVC, but this context should also not be interpreted as limiting the embodiments described subsequently to refer only to HEVC and its extensions.

[0161] In the context of the TMCTS system described above, tile-specific random access points may offer clear benefits. Random access in tiles at different time instances allows for a more even distribution of bitrate peaks across images in a video sequence. All or a subset of the mechanisms for image-specific random access in HEVC can be transferred to tiles.

[0162] One picture-specific random access mechanism is the indication of an intra-coded picture or access unit in which a) the picture in presentation order or b) the following picture in coding and presentation order does not inherit prediction dependencies on the picture samples preceding the intra-coded picture. In other words, a reference picture buffer reset is indicated either immediately in case b) or from the first subsequent picture on in case a). In HEVC, such access units are indicated on the Network Abstraction Layer (NAL) via a specific NAL unit type, namely, a so-called Intra Random Access Point (IRAP) access unit such as BLA, CRA (both above category a), or IDR (above category b). Embodiments described further below can use a NAL unit header level indication, e.g., via a new NAL unit type or an SEI message for backward compatibility, that indicates to a decoder or network intermediate box / device that a given access unit meets condition a) or b), i.e., contains at least one intra-coded slice / tile for which some form of reference picture buffer reset is applied on a slice / tile-by-slice / tile basis. Furthermore, slices / tiles can be identified by slice header level indication for the picture at the encoder side in addition to or instead of NAL unit type signaling. The pre-decoding operation allows to reduce the DPB size required for post-decoding.

[0163] For this to happen, the constraints expressed with fixed_tile_structure enabled must be satisfied: samples from the previous tile of the specified access unit should not be referenced by the same tile (and other tiles) of the current image.

[0164] According to some embodiments, the encoder can constrain coding dependencies via inter-subregion temporal prediction so that, for each subregion where RA occurs, the image region used as a basis for temporal prediction in the reference image is extended by the image region covered by further subregions if these further subregions also undergo RA. These slices / tiles / subregions are indicated in the bitstream, for example at the NAL unit or slice level or in an SEI message. Such a structure prevents subregion extraction but mitigates the penalty of constrained temporal prediction. The type of subregion random access (whether it allows extraction or not) must be distinguishable from the bitstream indication.

[0165] Another embodiment takes advantage of the above signaling opportunity by adopting a specific coding-dependent structure, in which image-wise random access points exist at specific points in time and with a coarse temporal granularity that allows instantaneous random access without drift with existing state-of-the-art signaling.

[0166] However, at finer temporal granularity, the coding structure allows tile-wise random access that distributes the bitrate load of intra-coded image samples over time towards less fluctuating bitrate behavior. For backward compatibility, this tile-wise random access can be signaled via SEI messages, keeping the respective slices as non-RAP pictures.

[0167] In the sub-picture bitstream extraction process, the NAL unit type indicated by the above SEI message indicating tile-based random access within such a stream structure is changed to picture-wise random access, signaling instantaneous random access opportunities in each picture of the extracted sub-bitstream, if necessary.

[0168] A video data stream 300 according to an embodiment of the third aspect of the present application will be described with reference to Figure 22. The video data stream 300 has encoded therein a sequence of images 302, i.e., video 304. As with the other embodiments described above, the temporal order in which the images 302 are shown may correspond to a presentation time order which may or may not coincide with the decoding order in which the images 302 are encoded into the data stream 300. That is, although not described with reference to the other previous figures, the video data stream 300 may be subdivided into a sequence of access units 306, each access unit 306 being associated with a respective one of the images 302, and the order in which the images 302 are associated with the sequence of access units 306 corresponds to the decoding order.

[0169] The images 302 are encoded into the video data stream 300 using temporal prediction, i.e., predictively coded images between the images 302 are coded using temporal prediction based on one or more temporal reference images that precede each image in decoding order.

[0170] Instead of having just one type of random access picture, the video data stream 300 includes at least two different types, as will be described below. In particular, normal random access pictures are pictures in which no temporal prediction is used; that is, each picture is coded in a manner independent of any previous pictures in the decoding order. For such normal random access pictures, the stopping of temporal prediction concerns the entire picture area. According to the embodiments described below, the video data stream 300 may or may not include such normal picture-by-picture random access pictures.

[0171] As just explained, random access pictures do not depend on previous pictures in the decoding order. Therefore, they allow random access for decoding of the video data stream 300. However, encoding pictures without temporal prediction implies encoding penalties in terms of compression efficiency. Therefore, typical video data streams experience bit rate peaks, i.e., bit rate maxima, at random access pictures. These problems can be solved by the above-described embodiments.

[0172] According to the embodiment of Figure 22, the video data stream 300 includes a first set of one or more images of type A encoded into the video data stream 300 forming a first set of one or more first random access points that suspend temporal prediction within at least a first image sub-region A, and a second set of one or more images of type B encoded into the video data stream 300 forming a second set of one or more second random access points of the video data stream 300 while suspending temporal prediction within a second image sub-region B different from the first image sub-region A.

[0173] In FIG. 22, the first and second image subregions A and B are indicated using hatching and do not overlap each other, as shown in FIG. 23a, but rather are adjacent to each other along a common boundary 308 such that subregions A and B cover the entire image area of ​​image 302. However, this is not necessarily the case. As shown in FIG. 23b, subregions A and B may partially overlap, or as shown in FIG. 23c, the first image region A may actually cover the entire image area of ​​image 302. In the case of FIG. 23c, type A images are picture-wise random access points where temporal prediction is completely switched off, i.e., they are encoded into data stream 300 without temporal prediction throughout their respective images. For completeness, FIG. 23d shows that subregion B does not have to be located inside image area 302, but can also be adjacent to the outer image boundary 310 of the two images 302. FIG. 23e shows that in addition to type A and B images, there is a type C image with an associated subregion C, which together can completely cover the image area of ​​image 302.

[0174] The consequence of restricting the regions in images B and A of FIG. 22 where temporal prediction is suspended to sub-regions A and B is as follows: Typically, the bit rate for encoding image 302 into video data stream 300 is large for images forming random access points, since temporal prediction is refrained from being used throughout the entire respective image region, and prediction from previous images (in decoding order) is broken for subsequent images (at least in presentation order). For images (pictures) of types A and B of FIG. 22, the avoidance of the use of temporal prediction is only used in sub-regions A and B, respectively, resulting in relatively low bit rate peaks 312 for these images A and B compared to image-wise random access point images. However, as explained below, the reduction in bit rate peaks 312 occurs at a relatively low cost, at least with respect to the random access rate of the complete image, excluding coding-dependent constraints at the boundaries of the sub-regions. This is illustrated in FIG. 22, where curve 314 represents a function showing the temporal evolution of the bit rate depending on time t. As explained, the peaks 312 at the time points of images A and B are lower than the peaks resulting from image-by-image random access images. Compared to simply using image-by-image random access images, the image-related random access rate of video data stream 300 corresponds to the rate at which sub-region-related random access images covering the entire image region are traversed, which in the case of FIG. 22 is the rate at which at least one image of type B and at least one image of type A are encountered. Even the presence of a regular image-by-image random access image in video data stream 300, i.e., image A, as in the example of FIG. 23c, is advantageous over a regular video data stream in which such images are simply dispersed in time throughout the video data stream. In particular, in such cases, the presence of the sub-region-related random access image in the case of FIG. 23c, i.e., image B, can be utilized to reduce the image rate of image A, which comes with a high bitrate peak, as will be explained in more detail below.However, compensating for the increased random access latency by interspersing type B images in the video data stream 300 between type A images allows for sub-region restricted random access to the video data stream, filling the time until the next image unit random access point is encountered, i.e., the time until the next image A.

[0175] Before proceeding to a description of a decoder that utilizes a special type of random access image within video data stream 300, some remarks are made regarding sub-region B and / or sub-region A, and way image 302 is encoded into video data stream 300 by taking into account sub-regions beyond the interruption of temporal prediction within sub-regions A and B while applying temporal prediction within the same image outside sub-regions A and B.

[0176] FIG. 22 uses a dashed box 316 to illustrate a video encoder configured to encode image 302 into video data stream 300 to include just-outlined images A and B, respectively. As already outlined above, video encoder 316 may be a hybrid video encoder that uses motion-compensated prediction for encoding image 302 into video data stream 300. Video encoder 316 may use any GOP (Group of Pictures) structure for encoding image 302 into video data stream 300, such as an open GOP structure or a closed GOP structure. With regard to sub-region-related random access images A and B, this means that video encoder 316 interposes between images 302 an image whose sub-region, i.e., A or B, is independent of any previous image in decoding order. It will be explained later that such sub-regions B and / or A may correspond to sub-regions 22 according to the embodiments of FIGS. 1 to 13, i.e., sub-regions into which video data stream 300 can be reduced. However, it should be noted that this is merely an example and that reducibility is not a necessary property of the video data stream 300, although it would be advantageous if the image 312 were encoded in the video data stream 300 so as to conform to the boundaries of sub-regions B and / or A, similar to the discussion above regarding the boundaries of sub-region 22 in relation to Figures 1-13.

[0177] In particular, while the reach of spatial coding dependency mechanisms when encoding images 302 within video data stream 300 is typically short, it is advantageous if sub-region-related random access images, i.e., images A and B of FIG. 22, are encoded into video data stream 300 such that the coding dependency for encoding each sub-region B / A does not cross the boundaries of each sub-region, so as not to introduce coding dependencies outside or spatially neighboring each sub-region. That is, within each sub-region B / A, sub-region-related random access images A and B are coded without temporal prediction and spatial coding dependency on portions of each image outside their respective sub-regions A / B. In addition to this, it is advantageous if images between random access images A and B are also coded into video data stream 300, taking into account the section boundaries of the section in which the immediately preceding section-specific random access image forms a sub-region-specific random access point.

[0178] For example, in FIG. 22 , image B forms a sub-region-specific random access point with respect to sub-region B, and therefore it is advantageous if image 302 follows this image B and precedes the first occurrence of the next sub-region random access image, i.e., image A. This sequence is exemplarily indicated using curly brackets 317. Image 302 is coded taking into account the boundaries of section B. In particular, it is advantageous if the spatial and temporal prediction and coding dependencies for coding these images into video data stream 300 are constrained so that section B does not depend on any part of these images or image B itself that is outside section B. With regard to motion vectors, for example, video encoder 316 constrains the motion vectors available for coding section B of the images between images B and A so that they do not point to image B and any part of the images between image B and image A that extends beyond the sub-region of reference image B. Beyond that, image B forms a random access point with respect to subregion B, so that no temporal reference image for temporally predicting subregion B of image 317 should be upstream relative to image B. Spatial dependencies for coding subregion B of intermediate images between images B and A are similarly constrained, i.e., so as not to introduce dependencies in neighboring portions outside of subregion B. Again, this constraint may be relaxed depending on the application; furthermore, see the discussion of FIGS. 1-13 for possible countermeasures against drift errors. Similarly, the constraint just described with respect to images between random access images whose predecessor is a subregion-specific random access image such as image B may only apply with respect to temporal prediction, where spatial dependencies have less significant impact with respect to drift errors.

[0179] The argument raised in the immediately preceding paragraph concerns the restriction of the coding dependency for coding the immediate successor (in terms of decoding order) of a sub-region-specific random access image B only with respect to the coding of sub-region B, i.e., the image 317 within the sub-region for which image B forms a sub-region-specific random access point. A question that should be addressed separately is whether the coding dependency of a coded image 317 outside section B, i.e., sub-region A in the case of FIG. 22, should be restricted so as to render the coding of an outer part of image 317 dependent on sub-region B. That is, the question is whether sub-region A of image 317 should be coded such that its spatial coding dependency is restricted so as not to reach, for example, section B of the same image, and whether the temporal coding dependency for coding sub-region A of image 317 should be restricted so as not to reach sub-region B of a reference image, which is either one of the preceding images B or image 317 in the coding / decoding order. More precisely, it should be noted that the reference picture used for coding sub-region A of picture 317 can belong on the one hand to one of the pictures coded / decoded before picture 317, and on the other hand to the leading random access picture B or be located upstream (in decoding order) with respect to picture B. In the case of a reference picture located upstream with respect to picture B, the temporal coding dependency on coded sub-region A of picture 317 is still restricted so as not to reach sub-region B. Rather, the question addressed here is whether the coding dependency for coding sub-region A of picture 317 is or is not restricted so as not to reach section B, insofar as the reference picture is one of pictures B and is a temporal coding dependency with respect to any picture previously coded / decoded among pictures 317. Both options have merits.If subregion B of image B and the subregion boundaries of image 317 are also followed when encoding subregion A of image 317, i.e., if the coding dependency for coding subregion A is limited and does not reach subregion B, then subregion A continues to be coded in a manner independent of section B, thus forming a subregion from which data stream 300 can be extracted or reduced as described above. The same applies to subregion B when considering the same situation with respect to the coding of subregion B of an image immediately following subregion-specific random access image A. If the reducibility of specific subregions such as subregion A and / or subregion B is not so important, it can be beneficial in terms of coding efficiency if the immediately preceding coding dependency reaches the subregion with respect to which the immediately preceding section-wise random access image can form a subregion-specific random access point. In that case, other subregions, such as subregion A in the preceding discussion, would no longer be able to be reduced, but coding efficiency would increase because the video encoder 316 would have fewer restrictions on exploiting redundancy by selecting motion vectors across the boundaries of subregion B, such that it temporally predicts part of subregion A of image 317 based on subregion B of either of these images or image B. For example, data stream 300 would no longer be able to be reduced with respect to region A if the following were done: image 317 is encoded into video data stream 300 in an image region A outside second image subregion B using temporal prediction that at least partially references a second image subregion B of a reference image among images 317 and B. That is, in the case of FIG. 22, the entire image region would be available for referencing temporally predictively coded region A of image 317, i.e., not only A but also B. In terms of decoding order, the image 317' following image A, i.e., the image 317' between A and a subsequent image of type B not shown in Figure 22, is encoded in the video data stream 300 and restricts temporal prediction within image area A outside the second image sub-area B so as not to refer to the second image sub-area B of the preceding reference image in terms of the decoding order between image A and image 317'.That is, in the case of Figure 22, region A is available simply because it is referenced by temporal predictive coding region A of picture 317'. The same is true for sub-region B: picture B attaches to sub-region B within sub-region B, and the coding of B of picture B can use sub-region A as well as A. In that case, stream 300 would no longer be reducible to sub-region B.

[0180] Before proceeding with a description of a decoder configured to decode the video data stream 300 of Figure 22, it should be noted that the video encoder 316 may be configured to provide signal processing 319 to the video data stream 300 indicating a spatial subdivision of the corresponding images into sub-regions A and B, to the extent that it relates to the entire video 304 or encompasses the sequence of images 302, or according to any of the alternatives as illustrated in Figures 23a-23e. The encoder 316 also provides the data stream 300 with signaling 318 indicating images A and B, i.e., signaling 318 marking particular images that are sub-region-specific random access points for any of the sub-regions indicated by signaling 319. That is, the signaling 319 signals a constant spatial subdivision among images for which signaling 319 is valid, and the signaling 318 distinguishes sub-region-specific random access images from other images and associates these sub-region random access images with one of the sub-regions A and B.

[0181] On the one hand, it is possible to make no particular distinction between images B and A in video data stream 300, while still distinguishing among other images as far as picture type is concerned. In the example of FIG. 22, images B and A are "merely" sub-region-specific random access images, and therefore not actual image-based random access images such as IDRs. Thus, as far as picture type signaling 320 in video data stream 300 is concerned, video data stream 300 does not distinguish between images B and A, on the one hand, and other temporally predicted images, on the other hand. For example, signaling 318 may be included in slices or slice headers of these slices to indicate that image 302 is encoded in stream 300 in that unit, thereby indicating that the image regions corresponding to each slice form section B. Alternatively, a combination of signaling 320 and 318 may be used to indicate to a decoder that a certain image is a sub-region-specific random access image and that the same image belongs to a specific sub-region. For example, signaling 320 may be used to indicate that a specific image is a sub-region-specific random access image, but does not reveal the sub-regions for which each image represents a sub-region-specific random access point. The latter indication is performed by signaling 318, which indicates which of the sub-regions of the picture subdivision is signaled by signaling 319, and associating a picture signal with a sub-region specific random access picture by signaling 320. However, signaling 320, which may be a syntax element of NAL unit type, can instead distinguish or differentiate not only between picture-wise random access pictures such as B-pictures and P-pictures, IDR-pictures, and sub-region-wise random access points as between pictures A and B, but also between temporally predicted pictures such as B-pictures and P-pictures, picture-wise random access pictures such as IDR-pictures, and sub-region-wise random access points as between pictures A and B, by using different values ​​for different sub-regions, i.e., pictures A and B, respectively.

[0182] Further signaling 321 may be inserted by video encoder 316 to signal for a particular subregion whether data stream 300 is reducible with respect to the respective subregion. Signaling 321 may be signaled in a manner that allows one of the subregions within data stream 300 to be signaled as a subregion in which data stream 300 is reducible, while other subregions do not form such subregions with respect to which data stream 300 is reducible. Alternatively, signaling 321 may simply allow binary signaling of reducibility with respect to all subregions, i.e., signaling 321 may signal that all subregions are subregions in which data stream 300 is reducible, or signal that data stream 300 is not reducible with respect to any of these subregions. However, the signaling 321 leaves the effect that sub-regions such as sub-regions A and B in the example of FIG. 22 are treated as completely independently coded sub-regions with respect to whether data stream 300 can be reduced, or not, in which case the asymmetric coding dependencies described above are used across sub-region boundaries, as previously described.

[0183] It should be noted that although sub-regions B and A have been shown to be contiguous areas, sub-regions B and A may alternatively be non-contiguous areas, such as sets of tiles of image 302. With regard to the specific processing with respect to tiles into which image 302 may be encoded into data stream 300 according to this example, reference is made to the descriptions of Figures 1 through 13.

[0184] With respect to Figure 23c, note that since image A is a picture-by-picture random access image, the picture type of image A will be different from that of the other temporally predicted images, and therefore image A will be signaled in picture type signaling 320.

[0185] FIG. 24 illustrates a video decoder 330 configured to decode a video data stream 300 using images B and A. The video data stream 300 in FIG. 24 is illustratively shown corresponding to the description of FIG. 22. That is, sub-region-related random access images B and A are interspersed between images 302 of a video 304 encoded in the video data stream 300. The video decoder 330 is configured to wait for the next random access image to occur, i.e., a sub-region-related image or an image-by-image random access image, when randomly accessing the video data stream 300. For video data streams 300 that do not include image-by-image random access images, the video decoder 330 may not even respond to such images. In any case, the video decoder 330 resumes decoding the video data stream 300 as soon as it encounters the first sub-region-related random access image, which in the example of FIG. 24 is image B. Starting from this image B, the video decoder 330 begins reconstructing, decoding, and outputting an image 322 showing only sub-region B. Alternatively, the video decoder 330 will decode, reconstruct, and output the portions of these images 322 outside of sub-region B, with or without signaling accompanying these images that indicates to a presentation device, such as a display device, that these portions of these images 322 outside of sub-region B are missing a reference image for this outer sub-region and therefore will experience drift error.

[0186] The video decoder 330 continues decoding the video data stream 300 in this manner until it encounters the next random access picture, which is picture A in the example of FIG. 24. Because sub-region A represents a random access point for the remaining image area, i.e., the area of ​​image 302 outside of sub-region B, the video decoder 330 fully decodes, reconstructs, and outputs pictures 302 from picture A onward. That is, the functionality described in connection with FIG. 24 provides the user with the opportunity to gain the advantage of viewing the video 304 earlier with respect to at least sub-region B, i.e., the sub-region of random access picture B associated with the first encountered sub-region. Then, after encountering the next sub-region-associated random access picture for the sub-region covering the remaining image 302, the video decoder 330 can provide the complete picture 302 without drift error. In the example of FIG. 23e, this is the case after encountering the first of each of the sub-region-specific random access pictures for sub-regions A, B, and C, respectively.

[0187] 25 illustrates another mode of operation of the video decoder 330. Here, the video decoder 330 begins decoding and reconstructing sub-region B upon encountering the first sub-region-specific random access image, here image B, but waits until it encounters enough random access images so that the sub-region covers the entire image area of ​​image 302, before the video decoder 330 actually outputs image 302 in its entirety. In this example, this is the case when image A and image B are encountered. That is, from image A, the video decoder 330 outputs image 302, but sub-region B would have been available from image B.

[0188] FIG. 26 illustrates a network device 322 receiving a video data stream 300 that includes a sub-region-specific random access image. However, this time, the video data stream 300 is reducible with respect to one or more sub-regions of the sub-region-specific random access image. For example, FIG. 26 illustrates a case where the video data stream 300 is reducible with respect to sub-region B. With respect to the reducibility and the corresponding functionality of the network device 322, it is noted that this functionality may or may not be configured to correspond to the description of FIGS. 1-13. In either case, image B is indicated in the video data stream 300 using the signaling 318 described above. When the network device 322 reduces the video data stream 300 to relate only to image B, it extracts it from the video data stream 300 or reduces the video data stream 300 to a reduced video data stream 324, whose image 326 forms a video 328 that simply shows the content of sub-region B of the video 304, i.e., the content of the unreduced video data stream 300. However, as measured in this manner, since subregion B no longer represents a subregion relative to picture 326, reduced video data stream 324 no longer includes signaling 318. Rather, the picture of video 328 corresponding to subregion B of subregion-specific random access picture B of the video data stream is signaled in reduced video data stream 324 by picture type signaling 320 to be a picture-wise random access picture, such as an IDR picture. There are various ways to do this. For example, network device 322, if configured to correspond to network device 60, can change picture type signaling 320 in the NAL unit headers of corresponding NAL units when reducing video data stream 300 to reduced video data stream 324, in addition to the redirection and / or parameter set revisions described above with respect to Figures 5 and 7.

[0189] For the sake of completeness, Fig. 27 shows a network device 231 configured to process the data stream 20 of Fig. 20. However, Fig. 20 shows that the information 206' may or may not already be present in the reducible video data stream 200, such as in the information 50. As already mentioned above, this network device 231 can be configured to discard the displacement information 206 when reducing the video data stream 200 to the reduced video data stream 232, by simply carrying over the sub-region-specific displacement information 206' from the video data stream 200 to the reduced video data stream 232, or the network device 231 can form a readjustment of the displacement information 206 to become sub-region-specific displacement information 206, based on knowledge of the position of the sub-region 22 with respect to the images of the reducible video data stream 200.

[0190] Thus, the above description has revealed processes and signaling for, for example, extraction of tile sets with temporal motion and inter-layer prediction constraints. Extraction or spatial subsets of coded video bitstreams using single or multi-layer video coding has also been described.

[0191] With regard to the above description, it should be noted that any encoder, decoder, or network device shown may be embodied or implemented in hardware, firmware, or software. If implemented in hardware, the respective encoder, decoder, or network device may be implemented, for example, in the form of an application-specific integrated circuit (ASIC). If implemented in firmware, the respective device may be implemented as a field-programmable array, and if implemented in software, the respective device may be a processor or computer programmed to perform the described functions.

[0192] While some aspects are described in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, and that a block or apparatus corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or used in) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most significant method steps may be performed by such an apparatus.

[0193] The encoded data stream or signal of the present invention can be stored on a digital storage medium or transmitted over a transmission medium, such as a wireless transmission medium or a wired transmission medium such as the Internet. Although the insertion or encoding of some information into a data stream has been described above, this description is also to be understood as a disclosure that the resulting data stream includes the respective information, flag syntax elements, etc.

[0194] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementation can be performed using a digital storage medium, such as a floppy disk (floppy is a registered trademark), DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM or flash memory, on which electronically readable control signals are stored, which cooperates (or can cooperate) with a programmable computer system to execute the respective method. Thus, the digital storage medium may be computer-readable.

[0195] Some embodiments according to the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0196] Generally, embodiments of the present invention may be implemented as a computer program product having program code operative to perform one of the methods when the computer program product is run on a computer, which program code may for example be stored on a machine readable carrier.

[0197] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0198] In other words, an embodiment of the inventive method is a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0199] A further embodiment of the inventive methods is, therefore, a data carrier (or digital storage medium or computer readable medium) having recorded thereon the computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.

[0200] A further embodiment of the inventive methods is, therefore, a data stream or sequence of signals representing the computer program for performing one of the methods described herein, The data stream or sequence of signals can for example be adapted to be transmitted via a data communications connection, for example the Internet.

[0201] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0202] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0203] Further embodiments according to the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0204] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0205] The apparatus described herein can be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0206] The devices described herein, or any components of the devices described herein, may be implemented at least in part in hardware and / or software.

[0207] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0208] The methods described herein, or any components of the apparatus described herein, may be performed at least in part by hardware and / or software.

[0209] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims, and not by the specific details shown by the description and explanation of the embodiments herein.

Claims

1. a video data stream representing a video, a parameter setting portion indicating encoding parameter settings; a payload portion in which the images of the video are parameterized and encoded using the first set of encoding parameter settings, the first set being indexed by an index included in the payload portion; Including, The video data stream displaying a predetermined sub-region of the image; substitution parameters for adjusting the first set of encoding parameter settings to produce a second set of encoding parameter settings, the encoding parameter settings defining buffer sizes and timing; information including The second set of encoding parameters comprises: removing a portion of the payload portion that falls in an area outside the predetermined sub-area of ​​the image; and modifying the location indication within the payload portion such that a location measured from a perimeter of the predetermined sub-region rather than the image indicates a location including a reduced payload portion encoding a sub-region-specific image that represents the predetermined sub-region of the image parameterized with the second set of encoding parameter settings. selecting a reduced video data stream modified when compared to said video data stream by The information is included in an SEI message, a VUI, or a parameter set extension. Video data stream.

2. 1. An encoder for encoding video into a video data stream, comprising: a parameter setter configured to determine encoding parameter settings and to generate a parameter setting portion of the video data stream indicative of the encoding parameter settings; an encoding core configured to encode images of the video parameterized using the first set of encoding parameter settings into a payload portion of the video data stream, the first set being indexed by an index included in the payload portion; Equipped with The encoder comprises: displaying a predetermined sub-region of the image; substitution parameters for adjusting the first set of encoding parameter settings to produce a second set of encoding parameter settings, the encoding parameter settings defining buffer sizes and timing; configured to provide information in the video data stream including The second set of encoding parameters comprises: removing a portion of the payload portion that falls in an area outside the predetermined sub-area of ​​the image; and modifying a position indication within the payload portion to indicate a position measured from the periphery of the predetermined sub-region rather than from the image; a reduced video data stream modified compared to the video data stream by the encoder is configured to insert the information into an SEI message, a VUI, or a parameter set extension. Encoder.

3. 1. A decoder for decoding a video data stream, the video data stream comprising: a parameter setting portion indicating encoding parameter settings; a payload portion in which an image of a video is parameterized and encoded using said first set of encoding parameter settings, said first set being indexed by an index included in said payload portion; Equipped with The decoder displaying a predetermined sub-region of the image; substitution parameters for adjusting the first set of encoding parameters to produce a second set of encoding parameter settings, the encoding parameter settings defining buffer sizes and timing; configured to read information from the video data stream including performing redirection and / or adjustment such that the second set of encoding parameter settings is indexed by an index of the payload portion; removing a portion of the payload portion that falls in an area outside the predetermined sub-area of ​​the image; and Modifying the position indication on the payload portion to indicate a position measured from the periphery of the predetermined sub-region rather than from the image. reducing the video data stream to a reduced video data stream modified by the decoder is configured to read the information from an SEI message, a VUI, or a parameter set extension of the video data stream. decoder.

Citation Information

Patent Citations

  • Object data processor, object data recording device, data storage medium and data transmission structure

    JP1998304353A