Apparatus for generating video data streams and methods for generating video data streams.
By providing additional signaling and extraction information in the video data stream to guide the modification of slice addresses, the problem of low video data stream extraction efficiency in existing technologies is solved, achieving more efficient video data stream processing and adaptability to different scene resolutions.
Patent Information
- Application Number
- CN202211352845.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-03-20
- Filing Date
- 2018-03-19
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2038-03-19
AI Technical Summary
Existing video data stream extraction processes are inefficient, especially when handling the transmission of video versions with different scene resolutions. They are complex and difficult to adapt to unknown types of video content.
By providing additional signaling and extraction information in the video data stream, the extractor device is guided on how to modify slice addresses and parameters, reducing processing burden and adapting to different types of video content.
It improves the efficiency and adaptability of video data stream extraction, reduces the complexity of the extraction process, and can more effectively handle the transmission of video versions with different scene resolutions.
Smart Images

Figure CN115955559B_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application No. 201880033197.5 entitled "Apparatus for generating video data stream and method for generating video data stream", filed on March 19, 2018. Technical Field
[0002] This application relates to video data stream extraction, specifically, the method of extracting a reduced video data stream from a properly prepared video data stream, such that the reduced video data stream encodes a smaller video portion corresponding to the spatial portion of the video encoded in the original video data stream. This application also relates to the transmission of different video versions of a scene, which differ in scene resolution or fidelity. Background Technology
[0003] The HEVC standard [1] defines a hybrid video codec that allows the definition of rectangular tile subarrays of images for which the video codec follows certain coding constraints, thereby allowing for easy extraction of smaller or reduced video streams from the entire video data stream, i.e., without requantization or any motion compensation. As outlined in [2], it is envisioned that this be added to the HEVC standard syntax to allow guidance for the extraction process by the video data stream receiver.
[0004] However, there remains a need to make the extraction process more efficient.
[0005] Applications involving video data extraction may involve the transmission or provision of multiple versions of a video scene with varying resolutions. An efficient approach would be advantageous for achieving this transmission or provision of versions with different resolutions. Summary of the Invention
[0006] Therefore, a first object of the present invention is to provide a more efficient technique for video data stream extraction, for example, a technique capable of more effectively processing video content of unknown type to the receiver, video content having different types of videos, such as differing in projection from the viewport to the image plane, or a technique that reduces the complexity of the extraction process. This object is achieved by the subject matter of the independent claims of this application according to the first aspect.
[0007] Specifically, according to the first aspect of this application, by providing information to the extraction information within the video data stream, the extraction of the video data stream becomes more efficient. This information signals one of several options, or explicitly signals how to modify the slice address of the slice portion of each slice within the extractable spatial segment, so as to indicate the position of the corresponding slice in the reduced image region within the reduced video data stream. In other words, the second information provides the video data stream extraction site with guidance for the extraction process, which aims to synthesize images of the reduced (extracted) spatial video from the original video based on spatial segments. Therefore, this reduces the processing burden of the extraction process or adapts it to greater variability in the types of scenes transmitted in the video data stream. Regarding the latter issue, for example, the second information can handle various situations where the images of the reduced spatial video should advantageously not be merely the result of simply piling together possible separate portions of the spatial segments while maintaining the relative arrangement (or relative order in terms of encoding order) of these portions of the spatial segments in the original video. For example, in a spatial segment consisting of zones adjacent to different parts along the perimeter of the original image (the original image displays the scene in the seam interface of the panoramic scene projected onto the image plane), the arrangement of these zones in the smaller images of the extracted stream should differ from that in the case of a non-panoramic image type, even though the receiver may not be aware of this type. Additionally, or separately, modifying the slice address of the extracted slice is a tedious task, which can be alleviated by explicitly sending information about how to make the modification, for example, in the form of an alternative slice address.
[0008] Another object of the present invention is to provide a technique that can more effectively provide the receiver with the juxtaposition of different scene resolution versions of a video scene.
[0009] This objective is achieved through the subject matter of the independent claims of the second aspect of this application.
[0010] Specifically, according to a second aspect of this application, the juxtaposition of versions of a video scene at different scene resolutions is provided more effectively by: summarizing and encoding these versions within a single video stream and providing a signal to the video stream indicating that images of the video display the same scene content at different resolutions in different spatial portions of the images. Therefore, a receiver of the video stream can identify, based on this signal, whether the video content transmitted by the video stream belongs to a spatially juxtaposed set of multiple versions of scene content at different scene resolutions. Depending on the capabilities of the receiving site, any attempts to decode the video stream can be suppressed, or the processing of the video stream can be adapted to the analysis of this signal. Attached Figure Description
[0011] Advantageous implementations of the embodiments described above are the subject matter of the dependent claims. Preferred embodiments of this application are described below with reference to the accompanying drawings, in which:
[0012] Figure 1 A schematic diagram of MCTS extraction with adjusted slice addresses is shown;
[0013] Figure 2 A mixed schematic diagram and block diagram illustrating a video data stream extraction and processing method according to an embodiment of the first aspect of this application, as well as the participating processes and devices, are shown.
[0014] Figure 3 A syntax example based on the example is shown, which inherits from the example of extracting second information, where the second information explicitly indicates how to modify the slice address.
[0015] Figure 4 A schematic diagram illustrating an example of non-adjacent MCTS forming the desired image sub-segments is shown;
[0016] Figure 5 An example of a specific syntax example including second information according to an embodiment is shown, wherein the second information indicates one of several possible options for modifying the slice address during the extraction process;
[0017] Figure 6 A schematic diagram illustrating an example of multi-resolution 360° frame packaging is shown;
[0018] Figure 7 A schematic diagram illustrating the extraction of an exemplary MCTS containing a mixed resolution representation is shown; and
[0019] Figure 8 The illustration shows a schematic diagram and block diagram illustrating an efficient way of providing multi-resolution scenes to users and participating devices according to an embodiment of the second aspect of this application, as well as a mixture of video streams and processes. Detailed Implementation
[0020] The description of this application begins with a description of the first aspect, and then continues with a description of the second aspect. More precisely, regarding the first aspect, the description begins with a brief overview of the basic technical problem to introduce the advantages and basic techniques of the embodiments of the first aspect described below. The second aspect follows the same order of description.
[0021] In panoramic or 360-degree video applications, it is often only necessary to show the user a portion of the image plane. Certain codec tools, such as Motion Constrained Tile Sets (MCTS), allow the extraction of encoded data corresponding to the desired sub-part of the image in the compressed domain, forming a standard-compliant bitstream that can be decoded by conventional decoder devices that do not support decoding the MCTS from the complete image bitstream and can be characterized at a lower level compared to the decoder required for decoding the complete image.
[0022] As an example and reference, the signaling involved in the HEVC codec can be found in the following references:
[0023] Reference [1] specifies the time MCTS SEI message in sections D.2.29 and E.2.29, which allows the encoder to signal that a given list of rectangles (each rectangle in the list is defined by the tile index of its upper left and lower right corners) belongs to the MCTS.
[0024] • Reference [2] provides additional information such as parameter sets and nested SEI messages, which can simplify the work of extracting MCTS into a standard-compliant HEVC bitstream, and will be added to the next version of [1].
[0025] As can be seen from [1] and [2], the extraction process includes adjusting the slice address signaled in the slice header of the relevant slice, which is performed in the extractor device.
[0026] Figure 1 An example of MCTS extraction is shown. Figure 1 The image shown is already encoded in a video data stream (i.e., a HEVC video data stream). Image 100 is subdivided into CTBs, or Code Tree Blocks, and image 100 is encoded in units of these CTBs. Figure 1 In the example, image 100 is subdivided into 16×6 CTBs; however, the number of rows and columns of the CTBs are not critical. Reference numeral 102 representatively indicates such a CTB. Using these CTBs 102 as units, image 100 is further subdivided into tiles, i.e., an array of m×n tiles. Figure 1 An exemplary case is shown with m=8 and n=4. In each block, reference symbol 104 has been used to representatively indicate one such block. Therefore, each block is a rectangular cluster or subarray of CTB 102. For illustrative purposes only, Figure 1 It is shown that the tile 104 can have different sizes, or in other words, the rows of the tile can each have different heights and the columns of the tile can each have different widths.
[0027] As is known in the art, the subdivision of a tile (i.e., subdividing image 100 into tiles 104) affects the encoding order 106 along which the image content of image 100 is encoded in the video data stream. Specifically, tiles 104 are traversed one after another in the order of the tiles, i.e., in a raster scan order row by row. In other words, all CTBs 102 within a tile 104 are first encoded or traversed according to the encoding order 106, and then the encoding order moves to the next tile 104. Within each tile 102, the CTBs are also encoded using the raster scan order (i.e., using a row-by-row raster scan order). Along the encoding order 106, the encoding of image 100 in the video data stream is subdivided to produce so-called slice portions. In other words, slices of image 100 traversed by consecutive portions of the encoding order 106 are encoded as units into the video data stream to form slice portions. Figure 1 This assumes that each tile is located within a single slice (or, in HEVC terminology, a slice fragment), but this is just an example and different approaches can be taken. Typically, a slice 108 (or, in HEVC terminology, a slice fragment) is located within... Figure 1 The reference numeral 108 is representative of the reference numeral 108, and it coincides with or corresponds to the corresponding block 104.
[0028] Regarding the encoding of image 100 in the video data stream, it should be noted that this encoding utilizes spatial prediction, temporal prediction, contextual derivation for entropy coding, motion compensation for temporal prediction, and transformation and / or quantization of the prediction residuals. The encoding order 106 not only affects the tiles but also defines the availability of the reference base for spatial prediction and / or contextual derivation: only those adjacent portions preceding the current tile in the encoding order 106 are available. The tiles not only affect the encoding order 106 but also restrict the encoding interdependencies within image 100: for example, spatial prediction and / or contextual derivation are restricted to referencing only portions within the current tile 104. Portions outside the current tile are not referenced in spatial prediction and / or contextual derivation.
[0029] Figure 1 Currently, another specific region of image 100 is shown, namely the so-called MCTS, which is the spatial segment 110 within the image area of image 100. It is possible to extract the video to which image 100 belongs by targeting this specific region. Figure 1 The right side shows an enlarged view of section 110. Figure 1 The MCTS110 consists of a set of tiles 104. The tiles located within segment 110 are all given names, namely a, b, c, and d. The fact that spatial segment 110 can be extracted involves further constraints on the encoding of the video-to-video data stream. Specifically, Figure 1The video image segmentation of image 100 shown is adopted by other images in the video, and for this image sequence, the image content within tiles a, b, c, and d is encoded in a way that the encoded interdependencies remain within spatial segment 110 even when referencing from one image to another. In other words, temporal prediction and temporal context derivation are constrained, for example, in a way that remains within the region of spatial segment 110.
[0030] When encoding the video to which image 100 belongs in the video data stream, a point of interest is the fact that slice 108 is provided with a slice address, which indicates the start of encoding of the slice within the encoded image region, i.e., its position. Slice addresses are assigned along the encoding order 106. For example, the slice address indicates the CTB order along the encoding order 106 at the start of encoding for each slice. For instance, in the data stream encoding both video and image 100 separately, the slice portion carrying the slice corresponding to patch a will have a slice address of 7, because the seventh CTB according to the encoding order 106 represents the first CTB in patch a according to the encoding order 106. Similarly, the slice addresses within the slice portions carrying slices related to patches b, c, and d will be 9, 29, and 33, respectively.
[0031] Figure 1 The right side indicates the slice addresses assigned by the receiver of the scaled-down or extracted video data stream to the two slices corresponding to tiles a, b, c, and d. In other words, Figure 1 The numbers 0, 2, 4, and 8 on the right indicate the slice addresses assigned by the receiver of the scaled-down or extracted video data stream, which is obtained from the original video data stream representing the video containing the entire picture 100 through extraction of spatial segment 110 (i.e., extraction via MCTS). In the scaled-down or extracted video data stream, the slice portions of slice 100 encoded with tiles a, b, c, and d are arranged in the same encoding order 106 as in the original video data stream (from which these slice portions are extracted during the extraction process). Specifically, the receiver places the image content (i.e., slices involving tiles a, b, c, and d) along the encoding order 112 in the form of a CTB reconstructed from a sequence of slice portions in the scaled-down or extracted video data stream. The encoding order 112 traverses the spatial segment 110 in the same way as the encoding order 106 traverses the entire image 100, that is, traversing the spatial segment 110 tile by tile in raster scan order, and also traversing the CTB within each tile in raster scan order before continuing to traverse the next tile. The relative positions of tiles a, b, c, and d remain unchanged. That is, as Figure 1As shown, the spatial segment 110 on the right preserves the relative positions of tiles a, b, c, and d when they appear in image 100. As a result of determining addresses using encoding order 112, the slice addresses corresponding to tiles a, b, c, and d are 0, 2, 4, and 8, respectively. Therefore, the receiver can reconstruct a smaller video based on a scaled-down or extracted video data stream, which is displayed as an independent image in the spatial segment 110 shown on the right.
[0032] Summarizing the results so far Figure 1 The description, Figure 1 The adjustments made to the slice address and CTB unit after extraction are shown or illustrated using the numbers in the top left corner of each slice 108. For extraction to be performed, the extraction station or extractor device needs to analyze the raw video data stream against parameters indicating the CTB size (i.e., the maximum CTB size) 114, and the number and size of tile columns and rows in picture 100. Furthermore, nested MCTS-specific sequences and picture parameter sets are examined to deduce the output tile arrangement. Figure 1 In this context, tiles a, b, c, and d within spatial segment 110 maintain their relative arrangement. In summary, the aforementioned analysis and checks, to be performed by the extractor device according to the MCTS instructions of the HEVC data stream, require specialized and complex logic to derive the slice address of the reconstructed slice 108 from the parameters listed above. This specialized and complex logic incurs additional implementation costs and runtime disadvantages.
[0033] Therefore, the embodiments described below utilize additional signaling in the video data stream and corresponding processing steps on the information generation and extraction sides to reduce the overall processing burden derived just described by the extractor device by providing readily available information specifically for extraction purposes. Additionally or alternatively, some embodiments described below use additional signaling to guide the extraction process in a certain way, thereby enabling more efficient processing of different types of video content.
[0034] First, based on Figure 2 Explain the general concepts. Then, participate. Figure 2 This general description of the operating modes of the various entities in the entire process illustrated is further presented in different ways according to the following various embodiments. It should be noted that although described together in one figure for ease of understanding, the entities and frames shown therein belong to independent devices, each of which is individually provided. Figure 2 The overall advantages and characteristics summarized in the text. More precisely, Figure 2The diagram illustrates: the process of generating a video data stream, the process of extracting information for such a video data stream, the extraction process itself, then the process of decoding the extracted or downscaled video data stream, and the participating devices, wherein the operating modes of these devices or the execution of various tasks and steps are according to the embodiments described herein. (Referring to the first reference...) Figure 2 The specific implementation examples described, and further outlined below, reduce the processing overhead associated with the extraction tasks of the extractor device. According to another embodiment, the processing of various types of video content within the original video is additionally or alternatively mitigated.
[0035] exist Figure 2 At the top, the original video, indicated by reference numeral 120, is shown. This video 120 consists of a sequence of images, one of which is indicated by reference numeral 100 because it serves a similar function to... Figure 1 The same function is shown in image 100, that is, it illustrates the image region from which spatial segments 110 will be extracted later via video data stream. However, it should be understood that the above regarding Figure 1 The explained tile subdivision does not need to be determined by Figure 2 The process illustrated is based on video coding. It should be understood that tiles and CTBs represent semantic entities in video coding. Figure 2 For the purposes of this embodiment, these are merely optional.
[0036] Figure 2Video encoding of video 102 is illustrated in video coding core 122. The video encoding performed by video coding core 122 converts video 120 into video data stream 124 using, for example, hybrid video coding. Specifically, video coding core 122 uses, for example, block-based predictive coding, which encodes individual picture blocks of images in video 120 using one of several supported prediction modes, and encodes prediction residuals. Prediction modes may include, for example, spatial prediction and temporal prediction. Temporal prediction may involve motion compensation (i.e., determining a motion field) and transmits this motion field, represented as motion vectors for the temporally predicted blocks, along with data stream 124. Prediction residuals may be transform-coded. That is, some spectral decomposition may be applied to the prediction residuals, and the resulting spectral coefficients may be quantized using, for example, entropy coding and losslessly encoded into data stream 124. Entropy coding may further utilize context adaptability, i.e., determining the context, where the context derivation may depend on the spatial and / or temporal neighborhoods. As described above, encoding can be based on encoding order 106, which imposes the following restriction on encoding dependencies: only video portions already traversed according to encoding order 106 can be used as the basis or reference for encoding the current portion of video 120. Encoding order 106 traverses video 120 frame by frame, but not necessarily in the order of the frames' presentation time. In a frame such as frame 100, video encoding kernel 122 subdivides the encoded data obtained through video encoding, thereby subdividing frame 100 into slices 108, each slice corresponding to a corresponding slice portion 126 in video data stream 124. Within data stream 124, multiple slice portions 126 form a sequence of slice portions that follow each other in the order in which the corresponding slices 108 in frame 100 are traversed according to encoding order 106.
[0037] Similarly, Figure 2 As shown, the video encoding core 122 provides a slice address to each slice portion 126, or encodes the slice address into each slice portion 126. For illustrative purposes, in Figure 2 Slice addresses are represented by uppercase letters. For example, regarding... Figure 1 As described, the slice address can be determined in some appropriate unit (e.g., in CTBs one-dimensionally along the encoding order 106), but alternatively, the slice address can be determined differently relative to a predetermined point in the picture area occupied by the picture of video 120 (e.g., the upper left corner of picture 100).
[0038] In this way, the video encoding core 122 receives the video 120 and outputs the video data stream 124.
[0039] As outlined above, regarding spatial segment 110, according to Figure 2The generated video data stream will be extractable, and accordingly, the video coding kernel 122 appropriately adapts to the video coding process. To this end, the video coding kernel 122 restricts inter-frame coding dependencies so that portions within spatial segment 110 are encoded into video data stream 124 in such a way that portions within spatial segment 110 do not depend on portions outside segment 110 through, for example, spatial prediction, temporal prediction, or contextual derivation. Slices 108 do not cross the boundaries of segment 110. Therefore, each slice 108 is either entirely within or entirely outside segment 110. It should be noted that the video coding kernel 122 can not only adhere to one spatial segment 110, but also to several spatial segments. These spatial segments can intersect each other (i.e., they can partially overlap), or one spatial segment can be entirely within another spatial segment. Because of these measures, as will be explained in more detail later, a downsized or extracted video data stream of images smaller than those in video 120 (i.e., images showing only the content within spatial segment 110) can be extracted from video data stream 124 without re-encoding, i.e., without having to perform complex tasks such as motion compensation, quantization, and / or entropy coding again.
[0040] Video data stream 124 is received by video data stream generator 128. Specifically, according to Figure 2 In the illustrated embodiment, the video stream generator 128 includes a receiving interface 130 that receives a prepared video stream 124 from the video encoding core 122. It should be noted that, according to alternatives, the video encoding core 122 may be included within the video stream generator 128, thereby replacing the interface 130.
[0041] Video data stream generator 128 provides extraction information 132 for video data stream 124. Figure 2 In this context, the resulting video data stream output by the video data stream generator 128 is indicated by reference numeral 124'. For the data stream indicated by reference numeral 134... Figure 2 The extractor device shown includes extraction information 132 instructing how to extract a scaled-down or extracted video data stream 136 from the video data stream 124', in which a smaller video 138 corresponding to spatial segment 110 is encoded. Extraction information 132 includes first information 140 and second information 142. First information 140 defines spatial segment 110 within the picture area sent by picture 100, and second information 142 signals one of several options regarding how to modify the slice address of slice portion 126 of each slice 108 falling within spatial segment 110, to indicate the location of each slice within the scaled-down picture area of picture 144 of video 138 within the scaled-down video data stream 136.
[0042] In other words, the video stream generator 128 only appends (i.e. adds) some information (i.e., extraction information 132) to the video stream 124 to derive the video stream 124'. This extraction information 132 is intended to guide the extractor device 134, which is to receive the video stream 124', so that it can extract a reduced or extracted video stream 136 from the video stream 124' specifically for segment 110. The first information 140 defines the spatial segment 110, i.e., its position within the image regions of video 120 and image 100, respectively, and may define the size and shape of the image region of image 144. Figure 2 As shown, this segment 110 does not necessarily have to be rectangular, raised, or a connected region. For example, in Figure 2 In the example, segment 110 consists of two non-overlapping partial regions 110a and 110b. Additionally, the first information 140 may contain hints about how the extractor device 134 should modify or replace some encoded parameters or portions of data streams 124 or 124', for example, adjusting image size parameters to reflect changes occurring in the image area when the image is transformed from image 100 to image 144 through the extraction operation. Specifically, the first information 140 may include replacement or modification instructions for the parameter set of the video data stream 124' applied by the extractor device 134 during the extraction process, so as to correspondingly modify or replace the corresponding parameter set contained in the video data stream 124' and received in the scaled-down or extracted video data stream 136.
[0043] In other words, the extractor device 134 receives video data stream 124', reads extraction information 132 from video data stream 124', and obtains spatial segment 110 (i.e., its position and location within the picture area of video 120) from the extraction information based on first information 140. Therefore, based on the first information 140, the extractor device 134 identifies some slice portions 126 that encode slices falling within segment 110, and these slice portions 126 will therefore be received into the scaled-down or extracted video data stream 136, while slice portions 126 relating to slices outside segment 110 will be discarded by the extractor device 134. Additionally, the extractor device 134 can use information 140 to correctly set (i.e., modify or replace) one or more parameter sets, as just outlined, before or during the adoption of one or more parameter sets within data stream 124' in the scaled-down or extracted video data stream 136. Therefore, one or more parameter sets may relate to picture size parameters, if segment 110 is such as Figure 2As exemplarily depicted, this is not a connected region; the image size parameter can be set according to information 140 to a size corresponding to the sum of the areas of segment 110 (i.e., the sum of the areas of all portions 110a and 110b of segment 110). This segment-sensitive discarding of the slice portion and parameter set adjustment confines the video data stream 124' to segment 110. Furthermore, the extractor device 134 modifies the slice address of the slice portion 126 within the scaled-down or extracted video data stream 136. Figure 2 These slices are shown using hedging. That is, the shaded slice portion 126 is the portion of the slice that falls into section 110 and is therefore extracted or received separately.
[0044] It should be noted that information 142 can be envisioned not only when it is added to the complete video data stream 124', which includes a sequence of slice portions encoding slices located within spatial segments and slice portions encoding slices located outside spatial segments, but also when the data stream containing information 142 has been stripped so that the sequence of slice portions included in the video data stream includes slice portions encoding slices located within spatial segments, but does not include slice portions encoding slices located outside spatial segments.
[0045] The following provides various examples of embedding the second information 142 into the data stream 124' and their processing. Typically, the second information 142 is communicated within the data stream 124' as a signal, either explicitly signaled or signaled as one of several options regarding how to perform a slice address modification. In other words, the second information 142 is communicated as one or more syntax elements whose possible values may, for example, be explicitly signaled as slice address replacement values, or may together allow for the differentiation of multiple possibilities to associate the slice address of each slice portion 126 in the video data stream 136 with a setting of one of the selected syntax elements in that data stream. However, it should be noted that the number of meaningful or permissible settings of the one or more syntax elements contained in the second information 142 mentioned above depends on how the video 120 has been encoded into the video data stream 124 and the selection of segment 110. For example, imagine that segment 110 is a rectangular connected region in picture 100, and video encoding core 122 will perform encoding on this segment, without further restricting the encoding in terms of the interior of segment 110. Segment 110 consisting of two or more regions 110a and 110b will not be applicable. That is, only the dependence on the exterior of segment 110 will be suppressed. In this case, segment 110 must be mapped onto the picture region of picture 144 of video 138 without modification, that is, without disturbing the position of any sub-region of segment 110, and by placing the interior content of segment 110 as is in the picture region of picture 144, the allocation of addresses α and β to the slice portions carrying the slices constituting segment 110 can be uniquely determined. In this case, the settings of information 142 generated by video stream generator 128 will be unique, that is, video stream generator 128 has no choice but to set information 142 in this way, although information 142 will have other signal options from the perspective of available encoding. However, even in this unchanged case, the signal 142, which explicitly indicates, for example, a unique slice address modification, has advantages because the extractor device 134 itself does not have to perform the tedious task of determining the slice addresses α and β of the slice portion 126 adopted from the stream 124'. Instead, it simply learns from the information 142 how to modify the slice address of the slice portion 126.
[0046] Depending on the different embodiments of the nature of information 142 further outlined below, extractor device 134 may either retain or maintain the order in which slice portions 126 are received from stream 124' into the reduced or extracted stream 136, or modify that order in a manner defined by information 142. In any case, the reduced or extracted data stream 136 output by extractor device 134 can be decoded by a common decoder 146. Decoder 146 receives the extracted video data stream 136 and decodes video 138 from it, wherein the image 134 of video 138 is smaller than the image of video 120 (e.g., image 100), and its image area is filled by placing slices 108 decoded from slice portions 126 within video data stream 136, such placement being in a manner defined by slice addresses α and β contained within slice portions 126 within video data stream 136.
[0047] In other words, so far, it has been described in some way. Figure 2 This description is adapted to various embodiments of the exact nature of the second information 142, which will be described in more detail below.
[0048] The embodiments described here use an explicit signal for a slice address, which is the address that the extractor device 134 should use when modifying the slice address of slice portion 126 received from stream 142 in stream 136. The embodiments described below use a signal 142 that allows signaling to the extractor device 134 of one of several permitted options regarding how to modify the slice address. As a result of the already encoded portion 110, the permission for the option is, for example, in a way that restricts the interdependencies of the encoding within segment 110 so as not to cross the spatial boundaries of segment 110, such as... Figure 1 As shown, the spatial boundary of segment 110 further divides segment 110 into two or more regions, such as 110a and 110c, or tiles a, b, c, and d. Subsequent embodiments may still involve the extractor device 134 performing the tedious task of calculating addresses itself, but allow for efficient processing of different types of image content within the original video 120, thereby generating a meaningful video 138 at the receiving end based on the corresponding extracted or reduced video data stream 136.
[0049] That is, as mentioned above regarding Figure 1 As outlined, according to embodiments of this application, information about how addresses are modified during the extraction process is explicitly transmitted via second information 142, thereby alleviating the tedious task of determining slice addresses in the extractor device 134. Specific examples of syntax that can be used for this purpose are listed below.
[0050] Specifically, information 142 can be used to explicitly signal the new slice address to be used in the slice header of the extracted MCTS by including a list of slice address substitution values contained in stream 124' in the same order as the slice portion 126 carried in bit stream 124'. See, for example, [link to relevant documentation]. Figure 1 Example from the example. Here, information 142 will be explicit signaling of the slice addresses following the order of slice addresses 124' in the bitstream. Again, slice 108 and the corresponding slice portion 126 can be received during the extraction process such that the order in the extracted video data stream 136 corresponds to the order in which these slice portions 126 are contained in the video data stream 124'. According to the following syntax example, information 142 explicitly signals the slice address by starting with the second slice or slice portion 126. Figure 1 In this case, the explicit signaling will correspond to the second message 142 indicating or notifying list {2, 4, 8} with a new number. Figure 3 The syntax example presents an exemplary syntax for this embodiment, which highlights the explicit signaling 142 added in addition to the MCTS extraction information SEI obtained from [2].
[0051] The semantics are listed below.
[0052] `num_associated_slices_minus2[i]` plus 2 represents the number of slices of the MCTS that contain an MCTS identifier equal to any value in the list `mcts_identifier[i][j]`. The value of `num_extraction_info_sets_minus1[i]` should be between 0 and 2. 32 The range is within -2 (inclusive).
[0053] `output_slice_address[i][j]` identifies the slice address of the j-th slice in bitstream order, which belongs to an MCTS whose MCTS identifier is any value in the list `mcts_identifier[i][j]`. The value of `output_slice_address[i][j]` should be between 0 and 2. 32 The range is within -2 (inclusive).
[0054] It should be noted that the presence of information 142 in the MCTS extraction information SEI, or information 142 other than MTCS-related information 140, can be controlled by a flag in the data stream. This flag can be named slice_reordering_enabled_flag, etc. If this flag is set, information 142 such as num_associated_slices_minus2 and output_slice_address exists in addition to information 140; if this flag is not set, information 142 does not exist, and the relative positions of the slices are respected or otherwise handled during the extraction process.
[0055] In addition, it should be noted that the terminology of H.265 / HEVC is used. Figure 3 The “_segment_” part in the syntax element name used can be replaced with “_segment_address_”, but the technical content remains unchanged.
[0056] It should also be noted that although the suggestion message 142 num_associated_slices_minus2 indicates the number of slices in segment 110 as an integer, representing the number of slices as a difference of 2, the number of slices in segment 110 can also be signaled directly in the data stream, or indicated as a difference of 1. For the latter option, num_associated_slices_minus1 can be used as a syntax element name, for example. Note that, for example, the number of slices in any segment 110 can also be one.
[0057] In addition to the MCTS extraction process anticipated so far in [2], other processing steps are also similar to those using, for example... Figure 3 The explicit signal notification associated with information 142 shown is also present. These additional processing steps facilitate the derivation of the slice address by the extractor device 134 during the extraction process, and the following summary of the extraction process underlines where this assistance occurs:
[0058] Let the bitstream inBitstream, the target MCTS identifier mctsIdTarget, the target MCTS extract information set identifier mctsEISIdTarget, and the target highest TemporalId value mctsTIdTarget be the inputs to the sub-bitstream MCTS extraction process.
[0059] The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream.
[0060] The requirement for bitstream consistency of the output bitstream is that any output sub-bitstreams that are the output of the processes specified in this section regarding bitstreams should be consistent bitstreams.
[0061] The output sub-bit stream is obtained as follows:
[0062] - The bitstream outBitstream is set to be the same as the bitstream inBitstream.
[0063] -Remove all NAL cells from outBitstream whose TemporalId is greater than mctsTIdTarget.
[0064] - For each remaining VCL NAL in each access unit of outBitstream, adjust the slice fragment header as follows:
[0065] - For the first VCL NAL unit, set the value of first_slice_segment_in_pic_flag to 1, otherwise set it to 0.
[0066] - Based on the list output_slice_address[i][j], set the value of slice_segment_address for non-first NAL units (i.e., slices) starting from the second one in the bitstream order.
[0067] Just now about Figures 1 to 3 The described embodiment variant alleviates the tedious tasks of slice address determination and extraction processing performed by the extractor device 134 by using information 142 as an explicit signal of the slice address. Figure 3 For a specific example, the information includes an alternative slice address 143, which applies only to each second and subsequent slice segment 126 in the slice segment sequence, and maintains the slice order as slice segments 126 relating to slice 108 within segment 110 are received from data stream 124' to data stream 136. The slice address substitution value is associated with a one-dimensional slice address assignment in the picture region of picture 144 in video 138 using sequence 112, and does not conflict with laying out only the picture region of picture 144 along sequence 112, which uses a sequence of slices 108 obtained from the received slice segment sequence. It should be noted, and will be further mentioned below, that an explicit signal can also be applied to the first slice address, i.e., to the slice address of the first slice segment 126. Even for the latter, substitution value 143 may be included in information 142. This signaling of information 142 also enables the placement of slice 108 corresponding to the first slice segment 126, except at the beginning of the encoding sequence 112 (which may be as follows). Figure 1Located outside the top left corner of the image shown, the first slice portion 126 is the first slice portion in the slice portion 126 carrying slice 108 within segment 110, following the order within stream 124'. If this possibility exists or is permitted, explicit signals of the slice address can also be used to provide greater flexibility in rearranging segment regions 110a and 110b. For example, in Figure 3 In the modified example depicted, information 142 also explicitly signals the replacement value 143 of the slice address of the first slice portion 126 within data stream 136 (i.e., in...). Figure 2 In the case of α and β), signal 142 will then enable the differentiation of two permitted or available placement methods of segment regions 110a and 110b within the output picture of video 138. One is that segment region 110a is placed on the left side of the picture, thus keeping slice 108 corresponding to the first slice portion 126 transmitted in data stream 136 at the beginning of encoding order 112, and maintaining the order between slices 108 compared to the slice order in picture 100 of video 120 (with respect to video data stream 124). The other is that segment region 120a is placed on the right side of picture 144, thereby changing the order of slices 108 within picture 144 traversed according to encoding order 112, relative to the order in which these slices are traversed in the original video within video data stream 124 according to encoding order 106. Signal 142 will explicitly notify the extractor device 134 of the slice addresses used in the modification as a list of slice addresses 143, which are sorted or assigned to slice portions 126 within the video data stream 124'. That is, information 142 will sequentially indicate the slice addresses for each slice within segment 110 in the order in which these slices 108 will be traversed in encoding order 106. Therefore, this explicit signal can produce the following permutation: the order in which segment regions 110a and 110b of segment 110 are traversed in sequence 112 is changed compared to the order in which they are traversed in the original encoding order 106. The slice portions 126 received from stream 124' into stream 136 will be reordered accordingly by the extractor device 134, i.e., to conform to the order in which the slices 108 of the extracted slice portion 126 are traversed in sequence 112. According to the standard conformance of the reduced or extracted video data stream 136, the slices 126 received from the video data stream 124' should strictly follow each other along the encoding order 112, that is, they should have monotonically increasing slice addresses α and β modified by the extractor device 134 during the extraction process, thereby maintaining standard conformance. Therefore, the extractor device 134 will modify the order among the received slices 126 to sort them according to the order in which the slices 108 are encoded along the encoding order 112.
[0068] according to Figure 2 Another variation of the description utilizes the latter aspect, namely the possibility of rearranging the slices 108 of the received slice portion 126. Here, the second information 142 signals the rearrangement of the order among the slice portions 126, in which any slice 108 located within segment 110 is encoded. One possibility of the embodiment currently described is to explicitly signal the slice address replacement value 143 in a manner that results in the rearrangement of the slice portions. However, information 142 may signal the rearrangement of the slice portions 126 within the extracted or reduced video data stream 136 in different ways. This embodiment ultimately places the reconstructed slices rebuilt from the slice portions 126 in the decoder 146 strictly along the encoding order 112 (meaning, for example, using a tile raster scan order) to fill the picture area of picture 144. A rearrangement method, as notified by signal 142, has been selected in a certain way to rearrange or change the order between the extracted or received slice portions 126, such that the placement process of decoder 146 may cause segment regions such as regions 110a and 110b to change their order compared to slice portions that are not rearranged. If the rearrangement method notified by information 142 preserves the original order in data stream 124', then segment regions 110a and 110b can maintain their relative positions in the original image region of image 110.
[0069] To explain the current variant, please refer to [reference needed]. Figure 4 . Figure 4 Image regions of image 110 from the original video and image 144 from the extracted video are shown. Furthermore, Figure 4 An exemplary tile division (i.e., division into tile 104) is shown, and an exemplary extraction segment 110 is shown, which includes two non-intersecting segment regions or zones 110a and 110b, namely the two opposite edges 150 of the image 110. r and 150 l Adjacent segments 110a and 110b, where these segments are located relative to edge 150. r and 150 l The extraction direction is consistent (i.e., along the vertical direction). Specifically, Figure 4 As shown, the image content of image 110 and therefore the video to which image 110 belongs are of a specific type (i.e., panoramic video), such that when the 3D scene is projected onto the image area of image 110, the edge 150... r and 150 l This creates a scene. Below image 110, Figure 4It shows two options on how to place segments 110a and 110b within the output image area of image 144 of the extracted video. Figure 4 The document also uses tile names to describe these two options. These options stem from the fact that video 210 has been encoded into stream 124 in a certain way, such that for each of the two regions 110a and 110b, encoding occurs independently of the external encoding. It should be noted that, in addition to... Figure 4 In addition to the two allowed options shown, two other options are possible in the following cases: each zone in zones 110a and 110b is subdivided into two tiles within each zone, and the video is encoded independently of external encoding into stream 124, i.e. Figure 4 Each tile in the diagram is either an extractable part or an independently coded part (relative to spatial and temporal interdependence). These two options then correspond to shuffling tiles a, b, c, and d in different ways within segment 110.
[0070] Once again, now for reference Figure 4 The described embodiments are intended to alter the order in which the second information 142 is transmitted from the data stream 124' to, adopted, or written into the extracted or scaled video data stream 136 during the extraction process of the extractor device 134, either to the slice 108 or the NAL unit of the data stream 124' carrying the slice 108. Figure 4 This illustrates a scenario where the desired image sub-segment 110 consists of non-adjacent tiles or segment regions 110a and 110b within the image plane spanned by image 110. The complete coded image plane of image 110 is... Figure 4 The top of the image shows a tile boundary and a desired MCTS110, which consists of two rectangles 110a and 110b comprising tiles a, b, c, and d. In terms of scene content, or due to... Figure 4 The fact that the video content shown is panoramic video content means that when image 110 covers the 360° environment around the camera by means of the equirectangular projection shown in this example, the desired MCTS 110 surrounds the right and left boundaries 150. r and 150 l In other words, because Figure 4 Given the illustrated scenario (i.e., the image content is panoramic), the second option would actually make more sense than the first (1) option (2) concerning placing segment regions 110a and 110b within the image area of the output image 144. However, the situation might be different if the image content were other types of images (e.g., non-panoramic images).
[0071] In other words, the order of tiles A, B, C, and D in the complete image bitstream 124' is {a, b, c, d}. If this order is simply transferred to the encoded order of the placement of the corresponding tiles in the extracted or downsized video data stream 136 or the output image 144, then in the above exemplary case, as... Figure 4 As shown in the lower left, the extraction process itself will not produce the desired data arrangement within the output bitstream 136. (As...) Figure 4 As shown in the lower right corner, a preferred arrangement {b, a, d, c} is illustrated, which yields a video bitstream 136 that produces continuous image content on the image plane of image 144 for conventional devices such as decoder 146. Such a decoder 146 may not have the capability (i.e., rendering) to rearrange sub-image regions of the output image 144 in the pixel domain as a post-processing step after decoding, and even sophisticated devices tend to avoid post-processing.
[0072] Therefore, based on the above regarding Figure 4 The resulting example provides the encoding side of the video stream generator 128 with means for signaling a preferred order among several selections or options, wherein segment regions 110a and 110b within the video stream 124', each consisting of, for example, sets of one or more tiles, should be arranged in the extracted or scaled-down video stream 136 or in the picture region covered by its picture 144, according to this preferred order. Figure 5 The specific syntax example presented in the second information 142 includes a list encoded into the data stream 124', indicating the position of each slice 108 falling within segment 110 within the extracted bitstream, in the order of the original or input bitstream, i.e., the order in which they appear in the bitstream 124'. For example, in Figure 4 In the example, the preferred option 2 will eventually become a list read as {1, 0, 3, 2}. Figure 5 A specific example of the syntax is shown, which includes second information 142 within the MCTS extract information set SEI message.
[0073] The semantics are as follows.
[0074] `num_associated_slices_minus1[i]` plus 1 indicates the number of slices of the MCTS that contain any value in the list `mcts_identifier[i][j]` whose mcts identifier is equal to any value in the list `mcts_identifier[i][j]`. The value of `num_extraction_info_sets_minus1[i]` should be between 0 and 2. 32 The range is within -2 (inclusive).
[0075] `output_slice_order[i][j]` identifies the absolute position of the j-th slice in the bitstream order, where the mcts identifier is equal to the list `mcts_identifier[i][j]` in the output bitstream.
[0076] Any value of MCTS. The value of output_slice_order[i][j] should be between 0 and 2. 23 The range is within -2 (inclusive).
[0077] The following describes the other processing steps in the extraction process defined in [2], which will help in understanding Figure 5 The signal implementation, wherein the additional content relative to [2] is highlighted by underline:
[0078] Let the bitstream inBitstream, the target MCTS identifier mctsIdTarget, the target MCTS extract information set identifier mctsEISIdTarget, and the target highest TemporalId value mctsTIdTarget be the inputs to the sub-bitstream MCTS extraction process.
[0079] The output of the sub-bitstream MCTS extraction process is the sub-bitstream outBitstream.
[0080] The requirement for bitstream consistency of the input bitstream is that any output sub-bitstreams that are the output of the processes specified in this section regarding bitstreams should be consistent bitstreams.
[0081] For the i-th extracted information set, derive OutputSliceOrder[j] from the list output_slice_order[i][j].
[0082] The output sub-bit stream is obtained as follows:
[0083] - Set the bitstream outBitstream to be the same as the bitstream inBitstream. [...]
[0085] -Remove all NAL cells from outBitstream whose TemporalId is greater than mctsTIdTarget.
[0086] - Sort the NAL cells of each access cell according to the list OutputSliceOrder[j].
[0087] - For each remaining VCL NAL unit in outBitstream, adjust the slice fragment header as follows:
[0088] - For the first VCL NAL unit within each access unit, set the value of first_slice_segment_in_pic_flag to 1, otherwise set it to 0.
[0089] - Set the value of slice_segment_address based on the tile settings defined in PPS where pps_pic_parameter_set_id equals slice_pic_parameter_set_id.
[0090] Therefore, in conclusion Figure 2 According to the embodiments Figure 5 The above variations, which are related to the above regarding Figure 3 The difference in the variant discussed is that the second information 142 does not explicitly signal how the slice address should be modified. That is, according to the variant just outlined, the second information 142 does not explicitly signal the replacement value for the slice address of the slice portion extracted from data stream 124 into data stream 136. Instead, Figure 5 The embodiments involve the following situation: First information 140 defines a spatial segment 110 within the image area of image 100 as including at least a first segment region 110a and a second segment region 110b, where video within the first segment region is encoded into video data stream 124' independently of the first segment region 110a, and video 120 within the second segment region is encoded into video data stream 124' independently of the second segment region 110b, wherein multiple slices 108 do not cross the boundary of either the first segment region or the second segment region 110a and 110b; at least on a unit basis of these regions 110a and 110b, the images 144 of the output video 138 of the extracted video data stream 136 can be constructed differently, thereby generating two options, and according to the above-mentioned... Figure 5 In a variation of the discussion, the second information 142 signals reordering information indicating how the slice portions 126 of slice 108 located in segment 110 should be reordered relative to the order of slice portions in video data stream 124' when extracting the reduced video data stream 136 from video data stream 124'. The reordering information 142 may include, for example, a set of one or more syntax elements. Among the states that may be notified by one or more syntax elements forming information 142, there may be a state according to which the reordering maintains the original order. For example, information 142 signals the order in which each slice portion 126 of one of the slices 108 of picture 100 falling into segment 110 is encoded. Figure 5Reference numeral 141 in the accompanying drawings, and the extractor device 134 reorders the slice portions 126 within the extracted or reduced video data stream 136 according to these orders. Then, the extractor device 134 modifies the slice addresses of the thus rearranged slice portions 126 within the reduced or extracted video data stream 136 in such a way that the extractor device 134 knows which slice portions 126 in the video data stream 124' were extracted from the data stream 124' into the data stream 136. Therefore, the extractor device 134 knows the slices 108 within the image area of image 100 corresponding to these received slice portions 126. Based on the rearrangement information provided by information 142, the extractor device 134 is able to determine how the segment regions 110a and 110b are shifted relative to each other in a translational manner, thereby forming a rectangular image area corresponding to image 144 of video 138. For example, in Figure 4 In an exemplary case of Option 2, the extractor device 134 assigns slice address 0 to the slice corresponding to tile b because slice address 0 appears in the second position of the sequence list provided by the second information 142. Therefore, the extractor device 134 is able to place one or more slices associated with tile b and then proceed to the next slice, which, according to the reordering information, is associated with a slice address pointing to a position in the image region immediately following tile b according to the encoding order 112. Figure 4 In the example, this is the slice associated with tile a, because the position of the next order indicates the first slice a in segment 110 of image 100. In other words, the reordering is restricted to produce any possible rearrangement of segment regions 110a and 110b. Individually for each segment region, the corresponding segment region is still traversed along the same path in encoding order 106 and 112. However, due to the tile division, it is possible that, in the case of option 2, the regions corresponding to segment regions 110a and 110b in the image region of image 144 (i.e., the region of the combination of tiles b and d and the region of the combination of tiles A and C) are traversed in an interleaved manner according to encoding order 112, and accordingly, the relevant slice portions 126 that encode the slices located in the corresponding tiles are interleaved in the extracted or downsized video stream 136.
[0091] Another embodiment uses signaling to ensure that the order of signaling, using existing syntax, reflects the preferred output slice order. More specifically, this embodiment can be achieved by interpreting the appearance of the MCTS extract SEI message [2] as a guarantee that the order in which the rectangles constituting the MCTS in the MCTS SEI messages from sections D.2.29 and E.2.29 of [1] represent the preferred output order of the tile / NAL units. Figure 5In a specific example, this would result in rectangles being used for each contained tile in the order {b, a, d, c}. Except for the derivation of OutputSliceOrder[j], this example is the same as the example above, for example,
[0092] OutputSliceOrder[j] is derived from the order of the rectangles signaled in the MCTS SEI message.
[0093] Summarizing the above examples, the second information 142 can signal to the extractor 134 the following: when extracting the reduced video data stream 136 from the video data stream, how to reorder the slice portions 126 falling into the spatial segment 110 relative to the order of the slice portions 126 in the slice portion sequence of the video data stream 124', the slice address of each slice portion 126 in the slice portion sequence of the video data stream 124' is one-dimensionally indexed to the encoding start position of the slice 108 in the corresponding slice portion 126 along the first encoding scan order 106 that traverses the image region, and wherein the image 100 has been encoded into the slice portion sequence of the video data stream along the first encoding scan order. Therefore, the slice addresses of the slice portions in the slice portion sequence within the video data stream 124' monotonically increase. The modification of the slice addresses when extracting the reduced video data stream 136 from the video data stream 124' is defined as follows: slices encoded into the slice portion are sequentially placed along the second encoding scan order 112 that traverses the reduced image region. These slice portions are those defined by the reduced video data stream 136 and are reordered as notified by the second information 142; and the slice address of the slice portion 126 is set to the index of the encoding start position of the slice determined along the second encoding scan order 112. The first encoding scan order 106 traverses the image regions within each segment region of the set of at least two segment regions in the same way as the second encoding scan order 112 traverses the corresponding spatial regions. Each segment region in the set of at least two segment regions is indicated by the first information 140 as a subarray of rectangular tiles (the image 100 is subdivided into rows and columns of the rectangular tiles), wherein the first and second coded scanning sequences use row-by-row tile raster scanning to completely traverse the current tile before continuing to scan the next tile.
[0094] As described above, the output slice order can be derived from another syntax element (e.g., output_slice_address[i][j] as described above). In this case, an important addition to the example syntax above regarding output_slice_address[i][j] is that the slice addresses are signaled for all associated slices (including the first associated slice) to enable sorting; that is, num_associated_slices_minus2[i] becomes num_associated_slices_minus1[i]. Except for the derivation of OutputSliceOrder[j], this example is the same as the example above, for example,
[0095] For the i-th extracted information set, derive OutputSliceOrder[j] from the list output_slice_address[i][j].
[0096] Even further embodiments consist of a single flag on information 142 indicating that the video content wraps around a set of image boundaries (e.g., vertical image boundaries). Therefore, an output order is derived in extractor 134 that applies to image sub-parts including tiles on both image boundaries, as previously outlined. In other words, information 142 can signal one of two options: a first option among multiple options indicates that the video is a panoramic video, showing the scene in a way that different edge portions of the images are adjacent to each other in the scene, and a second option among multiple options indicates that the different edge portions are not adjacent to each other in the scene. The at least two segment regions a, b, c, d included in segment 110 form first and second bands 110a, 110b, which are associated with different edge portions (i.e., the right edge 150). r and left edge 150 l In the case where different edge portions of the first and second zones are adjacent, the reduced image area is formed by placing a set of at least two segment regions together such that the first zone and the second zone are adjacent along different edge portions. In the case where the second information 142 signals the second option, the reduced image area is formed by placing a set of at least two segment regions together such that the different edge portions of the first and second zones are opposite to each other.
[0097] For the sake of completeness only, it should be noted that the shape of the image region in Figure 144 is not limited to conforming to the following shape: combining individual regions (e.g., tiles a, b, c, and d of segment 110) together while maintaining the relative arrangement of any connected clusters during the combination, for example... Figure 1(a, c, b, d); or join any such clusters together along the shortest possible interconnection direction, for example... Figure 2 The horizontally joined zones (a, c) and (b, d) in the image. Conversely, for example, in... Figure 1 In the extracted data stream, the image region can be a column consisting of all four regions. Figure 2 In this context, it can be a column consisting of a row of all four regions. Typically, the size and shape of the image region of image 144 in video 138 can exist in different parts of the data stream 124': for example, this information can be given in the form of one or more parameters within information 140 to guide the extraction process in extractor 134 by adaptive adjustments to the parameter set when extracting a specific substream 136 from stream 124'. For example, a nested parameter set used to replace the parameter set in stream 124' can be included in information 140, where the parameter set includes, for example, parameters related to image size, indicating the size and shape of image 144 in pixels, such that during extraction in extractor 134, the replacement of the parameter set in stream 124' overrides the old parameters in the parameter set indicating the size of image 100. However, the image size of image 144 can be indicated as part of information 142, either alternatively or as a separate feature. If the slice address is not explicitly provided in Message 142, it may be particularly advantageous to explicitly signal the shape of Image 144 in an easily readable high-level syntax element such as Message 142. The nested set of parameters will then need to be parsed to derive the address.
[0098] It should also be noted that in more complex system setups, cubic projection can be used. This projection avoids the known drawbacks of isorectangular projection, such as large variations in sampling density. However, when using cubic projection, a rendering stage is required to recreate a continuous viewport from the content (or its sub-parts). Such a rendering stage may involve a trade-off between complexity and capability; that is, some feasible and readily available rendering modules may expect a given arrangement of the content (or its sub-parts). In such cases, realizing the possibility of arrangement through the following disclosures is crucial.
[0099] Hereinafter, embodiments relating to the second aspect of this application are described. The description of the embodiments of the second aspect of this application again begins with a brief introduction to the general problems conceived and solved by these embodiments.
[0100] In the case of 360° video (but not limited to this), a relevant use case for MCTS extraction is composite video, which contains multiple resolution variations of content that are adjacent to each other on the image plane, such as... Figure 6 As shown. Figure 6The bottom depicts a composite video 300 with multiple resolutions, and line 302 shows the tile boundaries of tile 304. The composite of high-resolution video 306 and low-resolution video 308 is subdivided into such tiles to be encoded into corresponding data streams.
[0101] More precisely, Figure 6 A picture from a high-resolution video is shown at position 306, and a picture from a low-resolution video (such as a picture taken at the same time) is shown at position 308. For example, videos 306 and 308 both show the exact same scene, i.e., with the same graphics or field of view, but at different resolutions. However, the field of view can alternatively overlap only partially with each other, and in the overlapping band, the spatial resolution is manifested in the fact that the number of samples in the picture of video 306 is different compared to that of video 308. In effect, videos 306 and 308 have different fidelity, i.e., the number of samples or pixels in the equal scene portions are different. The juxtaposed pictures of videos 306 and 308 are arranged side by side to form a larger picture, thereby generating the picture of the composite video 300. For example, Figure 6 The image shown in video 308 is horizontally divided into two halves, with one half on top of the other. These halves are appended to the right of the simultaneous image in video 306, thus producing the corresponding video 300. Tile partitioning is performed in such a way that no tile 304 crosses the boundary between the high-resolution image of video 306 and the image content originating from the low-resolution video 308. Figure 6 In the example, the tile division of image 300 is as follows: an 8x4 tile formed by dividing the high-resolution image content of video 300, and a 2x2 tile formed by dividing each of the two halves from low-resolution video 308. The total number of images in video 300 is 10x4 tiles.
[0102] When such a multi-resolution composite video 300 is encoded using MCTS in an appropriate manner, MCTS extraction can produce a variation 312 of the content. For example... Figure 7 As shown, such a variant 312 can, for example, be designed to depict a predefined sub-picture 310a at high resolution and the rest of the scene or another sub-picture at low resolution, wherein the MCTS 310 in the composite picture bitstream is compared with three aspects or three separate regions 310a, 310b, 310c.
[0103] That is, the extracted video image (i.e., image 312) has three regions 314a, 314b and 314c, each region corresponding to one of the MCTS regions 310a, 310b and 310c, where region 310a is a sub-region of the high-resolution image region of image 300, and the other two regions 310b and 310c are sub-regions of the low-resolution video content of image 308.
[0104] As already mentioned, Figure 8 Embodiments of this application relating to the second aspect of this application are described. In other words, Figure 8 This illustrates a scheme for presenting multi-resolution content to a recipient site, specifically the various sites involved in generating multiple versions at different resolutions for the recipient site. It should be noted that along... Figure 8 The processing paths described herein, with each device and process located at each site, represent individual devices and methods; therefore, Figure 8 This should not be interpreted as merely illustrating a systemic or holistic approach. Regarding Figure 2 Similar statements are also correct. Figure 2 Individual devices and methods are also shown. All these sites are shown together in one figure solely for ease of understanding their interrelationships and the advantages of the embodiments described with respect to these figures.
[0105] Figure 8 Video data stream 330 is shown, containing video 332 encoded with image 334. Image 334 is the result of stitching together simultaneous image 336 of high-resolution video 338 and image 340 of low-resolution video 342. More precisely, images 336 and 340 of videos 338 and 342 either correspond to the exact same view area 344, or at least partially overlap to show the same scene in the overlapping area. However, for the same scene content, the number of sampled samples in high-resolution video image 336 is higher than the number of samples in the corresponding low-resolution image 340; therefore, the scene resolution fidelity of image 336 is higher than that of image 340. The synthesis of image 334 of synthesized video 332 is performed by compositor 346 based on videos 338 and 342. Compositor 346 stitches images 336 and 340 together. In doing so, the compositor 346 can subdivide the images of the high-resolution video 338 and the images 340 of the low-resolution video, or both, to obtain favorable filling and patching of the image regions of the images 334 of the composite video 332. The video encoder 348 then encodes the composite video 332 into a video data stream 330. The video data stream generation device 350 can be included in the video encoder 348, or can be connected to the output of the video encoder 348, to provide a signal 352 to the video data stream 330 indicating that the same scene content is displayed multiple times (i.e., in different spatial portions at different resolutions) for each image 334 or each image or video 332 in a sequence of images in the video 332 encoded into the video data stream 330. These portions are in... Figure 8H and L are used to indicate their origin through the composition completed by synthesizer 346, and are indicated by reference numerals 354 and 356. It should be noted that more than two versions at different resolutions can be combined to form the content of image 332, and respectively as shown in the figure. Figure 8 , Figure 6 and Figure 7 The use of the two versions shown is for illustrative purposes only.
[0106] For the purpose of explanation, Figure 8 A video stream processor 358 is shown that receives video data stream 330. The video stream processor 358 may be, for example, a video decoder. In any case, the video stream processor 358 is able to check signal 352 to determine, based on this signal, whether the video stream processor 358 should begin processing such as decoding the video data stream 330, which further depends on certain functions of, for example, the video stream processor 358 or the device connected downstream thereto. For example, if the video stream processor 358 is only capable of presenting the images 332 fully encoded in the video data stream 330, then the video stream processor 358 may refuse to process the following video data stream 330: i.e., its signal 352 indicates that its individual images display the same scene content at different spatial resolutions in different spatial portions of those individual images; that is, signal 352 indicates that the video data stream 330 is a multi-resolution video data stream.
[0107] Signal 352 may include, for example, a flag contained within data stream 330 that can switch between a first state and a second state. The first state may, for example, indicate the fact just outlined, that the individual images of video 332 display multiple versions of the same scene content at different resolutions. The second state indicates that this is not the case, that is, the images only show one scene content at one resolution. Video data stream processor 358 will therefore refuse to perform certain processing tasks in response to flag 352 being in the first state.
[0108] Signals 352, such as the flags described above, can be transmitted within the data stream 330 within its sequence parameter set or video parameter set. In the following description, possible syntax elements reserved for future use in HEVC are exemplarily identified as possible candidates.
[0109] As mentioned above Figure 6 and Figure 7 As shown, although not mandatory but possible, it is possible that the synthesized video 332 has been encoded into the video data stream 330 in such a way that for each of a set of non-overlapping spatial regions 360 (e.g., tiles or tile subarrays), the encoding is independent of the exterior of the corresponding spatial region 360. As mentioned above regarding... Figure 2The encoding independence restricts spatial and temporal prediction and / or contextual derivation to avoid crossing boundaries between spatial regions 360. Therefore, for example, encoding dependency can restrict the encoding of a spatial region 360 of a certain image 334 of video 332 to refer only to the spatial region at the same location within another image 334 of video 332, which is subdivided into spatial regions 360 in the same manner as image 334. Video encoder 348 can provide extraction information such as information 140, or a combination of information 140 and 142, to video data stream 330, or a corresponding device such as video data stream generator 128 can be connected to its output. The extraction information can involve some or all possible combinations of spatial regions 360 as extraction components, for example... Figure 7 The extraction segment 310 in the signal 352 may further include information such as the spatial subdivision of the image 334 of the video 332 into segments 354 and 356 with different scene resolutions, i.e., information on the size and position of each spatial segment 354 and 356 within the image area of the image 332. Based on such information within the signal 352, the video stream processor 358 may, for example, exclude certain extraction segments from the list of potentially extractable segments of the video stream 330. These excluded extraction segments may, for example, be a mixture of spatial regions 360 of different segments falling into segments 354 and 356, thus avoiding video extraction for extraction segments with mixed resolutions. In this case, the video stream processor 358 may, for example, include an extractor device, such as... Figure 2 Extractor equipment.
[0110] Additionally or alternatively, signal 330 may include information about different resolutions, wherein images 334 of video 332 show the same scene content at these different resolutions. Furthermore, signal 352 may also simply indicate a count of different resolutions, wherein images 334 of video 332 show the same scene content multiple times at different image locations at these different resolutions.
[0111] As already mentioned, video data stream 330 may include extraction information on a list of possible extraction regions that video data stream 330 may extract relative to it. Then, signal 352 may include additional signals indicating, for each of at least one or more of these extraction regions: the viewport direction of a segment of the respective extraction region displaying the same scene content at the highest resolution within the respective extraction region, and / or the area share of the segment of the respective extraction region displaying the same scene content at the highest resolution within the respective extraction region in the total region of the respective extraction region, and / or a spatial subdivision of the respective extraction region from its entire region, said spatial subdivision dividing the respective extraction region into segment regions that display the same scene content at mutually different resolutions.
[0112] Therefore, such a signal 352 can be exposed at a high level in the bitstream, thus making it easy to push into the streaming system.
[0113] One option is to use one of the generally reserved zero-X-bit flags in the configuration file hierarchy syntax. This flag can be named a general non-multi-resolution flag:
[0114] A general non-multi-resolution flag equal to 1 indicates that the decoded output image does not contain multiple versions of the same content at different resolutions (i.e., corresponding syntax such as region packing is constrained). A general non-multi-resolution flag equal to 0 indicates that the bitstream may contain such content (i.e., no constraints).
[0115] Additionally, the present invention therefore includes signals that inform the following: the nature of the complete bitstream content characteristics, i.e., the number and resolution of variants in the composition. Furthermore, other signals provide information in an easily accessible form within the coded bitstream regarding the following characteristics of each MCTS:
[0116] • Main viewpoint direction:
[0117] The orientation of the MCTS high-resolution view center, such as δ yaw, δ pitch, and / or δ roll relative to a predefined initial view center.
[0118] Overall coverage
[0119] The percentage of all content represented in MCTS.
[0120] • Ratio of high to low resolution
[0121] The ratio between high-resolution and low-resolution areas in MCTS, i.e., how much of the total content covered is represented in high resolution / fidelity.
[0122] The proposed signal information for the viewport orientation or overall coverage of the complete omnidirectional video already exists. Similar signals need to be added for the sub-regions that may be extracted. This information takes the form of an SEI[2] and can therefore be included in a motion-limited tile set extraction information nested SEI. However, such information is needed to select the MCTS to be extracted. Including the information in the motion-limited tile set extraction information nested SEI adds additional indirection and requires deeper parsing (the motion-limited tile set extraction information nested SEI contains additional information not needed for selecting the extraction set) in order to select a given MCTS. From a design perspective, a more straightforward approach is to signal the information or a portion thereof at a central point containing only the essential information for selecting the extraction set. Additionally, the signals mentioned include information about the entire bitstream, and in the proposed case, it is desirable to signal the coverage of high-resolution and low-resolution areas, or, if more resolutions are mixed, the coverage of each resolution, as well as the viewport orientation of the video extracted from the mixed resolutions.
[0123] One embodiment is to add the coverage for each resolution and add it to the 360ERP SEI from [2]. Thus, the SEI may be included in a nested SEI for motion-limited tile set extraction information and require the aforementioned tedious task to be performed.
[0124] In another embodiment, a flag is added to the MCTS Extraction Information Set SEI, such as omnidirectional information indicating the presence of the signal in question, so that only the MCTS Extraction Information Set SEI is needed to select the set to be extracted.
[0125] Although some aspects have been described in the context of the apparatus, it is clear that these aspects also represent a description of the corresponding method, where a box or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of method steps also represent a description of a corresponding box or item or feature of the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices (e.g., microprocessors, programmable computers, or electronic circuits). In some embodiments, one or more of the most important method steps may be performed by such devices.
[0126] The data stream of this invention can be stored on a digital storage medium, or it can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium (such as the Internet).
[0127] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementation can be carried out using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) on which electronically readable control signals are stored, which cooperate with (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0128] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0129] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable medium.
[0130] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0131] In other words, embodiments of the method of the present invention are therefore computer programs having program code for performing one of the methods described herein when the computer program is run on a computer.
[0132] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded, the computer program being used to perform one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.
[0133] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence of a computer program used to perform one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0134] Another embodiment includes a processing means, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.
[0135] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.
[0136] Another embodiment of the invention includes an apparatus or system configured to transmit a computer program to a receiver (e.g., electronically or optically) for performing one of the methods described herein. The receiver may be, for example, a computer, mobile device, storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.
[0137] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0138] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.
[0139] The apparatus described herein, or any component thereof, may be implemented, at least in part, in hardware and / or software.
[0140] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.
[0141] The methods or any components of the apparatus described herein may be performed, at least in part, by hardware and / or by software.
[0142] The embodiments described above are merely illustrative of the principles of the invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the invention is intended to be limited only by the scope of the appended claims and not by the specific details given by way of the description and explanation of the embodiments herein.
[0143] Example 1: A video data stream, comprising:
[0144] A sequence of slice portions (126), each slice portion (126) containing a corresponding slice (108) of a plurality of slices of an image (100) of a video (120), wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion in the image region of the video;
[0145] Extraction information (132) instructs how to extract a reduced video stream (136) from the video stream by: restricting the video stream to slice portions thereof, wherein the reduced video stream (136) encodes a smaller spatial video (138) corresponding to a spatial segment (110) of the video; and modifying the slice addresses to relate to a reduced image region of the smaller spatial video (138), wherein the extraction information (132) includes:
[0146] First information (140), the first information defines the spatial segment (110) within the image area, wherein multiple slices do not cross the boundary of the spatial segment (110); and
[0147] The second information (142) signals one of several options regarding how to modify the slice address of the slice portion of each slice within the spatial segment (110) to indicate the position of the corresponding slice in the reduced image region within the reduced video data stream (136).
[0148] Example 2, according to the video data stream described in Example 1, wherein the first information (140) defines a spatial segment (110) within the image area as a set consisting of at least two segment regions (110a, 110b; a, b, c, d), wherein the video is encoded in the video data stream independently of the corresponding segment region, wherein the plurality of slices (108) do not cross the boundary of either of the at least two segment regions.
[0149] Example 3, based on the video data stream described in Example 2, wherein,
[0150] The first of the multiple options indicates that the video is a panoramic video, which shows the scene in a way that different edge portions of the image are adjacent to each other on the scene.
[0151] The second option among the multiple options indicates that the different edge portions are not adjacent to each other in the scene.
[0152] Wherein, the at least two segment regions form a first zone and a second zone (110a, 110b) adjacent to different portions of the different edge portions.
[0153] so that,
[0154] When the second information (142) signals the first option, the reduced image region is formed by placing the set of the at least two segment regions together so that the first band and the second band are adjacent along the different edge portions, and
[0155] When the second information (142) signals the second option, the reduced image area is formed by combining the set of at least two segment regions in such a way that the different edge portions of the first and second zones are opposite to each other.
[0156] Example 4, based on the video data stream described in Example 2, wherein,
[0157] The second information (142) signals how the slice portions (126) within the spatial segment (110) should be reordered relative to the order of the slice portions (126) in the sequence of slice portions in the video data stream when extracting the reduced video data stream (136) from the video data stream, and
[0158] The slice address of each slice in the video data stream slice sequence is one-dimensionally indexed to the encoding start position of the slice in the corresponding slice, along the first encoding scan order (106) that traverses the image region, and wherein the image (100) has been encoded in the video data stream slice sequence along the first encoding scan order, thereby
[0159] The slice addresses of the slice portions in the slice portion sequence within the video data stream monotonically increase, and
[0160] The modification of the slice address when extracting the reduced video data stream (136) from the video data stream is defined as follows: slices encoded into a slice portion are placed sequentially along a second encoding scan order (112) that traverses the reduced image region, the slice portion being the slice portion confined to the reduced video data stream (136) and being the slice portion that has been reordered as signaled by the second information (142); and the slice address of the slice portion (126) is set to the starting position of the encoding of the slice determined along the second encoding scan order (112).
[0161] Example 5, based on the video data stream described in Example 4, wherein the first encoding scanning order (106) traverses the image regions in each segment region of the set of at least two segment regions in the same way as the second encoding scanning order (112) traverses the corresponding spatial regions.
[0162] Example 6, a video data stream according to Example 4 or 5, wherein each segment region in the set of at least two segment regions is indicated by the first information (140) as a rectangular tile subarray, the image (100) is subdivided into rows and columns of the rectangular tile subarray, wherein the first coded scan order and the second coded scan order use row-by-row tile raster scanning to completely traverse the current tile before proceeding to the next tile.
[0163] Example 7, according to the video data stream of Example 1, wherein, for each subset of at least one subset of the slice portion, the second information (142) signals an alternative value for the slice address, wherein any slice within the spatial segment is encoded in the slice portion.
[0164] Example embodiment 8, according to the video data stream described in Example embodiment 7, wherein the slice address of each slice portion in the slice portion sequence of the video data stream is related to the image region of the video, and the substitute value of the slice address is related to the reduced image region.
[0165] Example embodiment 9, according to the video data stream described in Example embodiment 6 or 7, wherein the slice address of each slice portion of the slice portion sequence of the video data stream is one-dimensionally indexed along a first coded scan order traversing the image region to index the encoding start position of the slice encoded into the corresponding slice portion, and wherein the image has been encoded in the slice portion sequence of the video data stream along the first coded scan order, and the substitute value of the slice address is one-dimensionally indexed along a second coded scan order traversing the reduced image region to index the encoding start position.
[0166] Example embodiment 10, a video data stream according to any one of example embodiments 7 to 9, wherein the subset of the slice portion does not include a first slice portion of the slice portion sequence in which any slice within the spatial segment (110) is encoded, and wherein the second information (142) provides alternative values for each subset of the slice portion in the order in which the slice portions appear in the video data stream.
[0167] Example 11, a video data stream according to any one of Example 7 to 9, wherein the first information (140) defines a spatial segment (110) within the image area as a set consisting of at least two segment regions (110a, 110b), wherein the video is encoded in the video data stream independently of the corresponding segment region outside the segment region in each segment region, wherein multiple slices do not cross the boundary of either the first segment region or the second segment region.
[0168] Example 12, a video data stream according to Example 11, wherein the subset of the slice portions includes all slice portions of the slice portion sequence wherein any slice within the spatial segment is encoded, and wherein the second information provides alternative values for each slice portion in the order in which the slice portions appear in the video data stream.
[0169] Example 13, according to the video data stream described in Example 12, wherein,
[0170] The second information (142) also signals how to reorder the slice portions of the segments within the spatial segment relative to the order of the slice portions in the sequence of slice portions in the video data stream when extracting the reduced video data stream from the video data stream, and
[0171] The second information (142) provides alternative values for each subset of the slice portion in the order in which the reordered slice portions appear in the reduced video data stream.
[0172] Example 14: A video data stream according to any one of Example Examples 1 to 13, wherein...
[0173] The video data stream includes a sequence of slice portions, comprising slice portions encoding slices located within the spatial segment, and slice portions encoding slices located outside the spatial segment, or...
[0174] The video data stream includes a sequence of slice portions that encode slices located within the spatial segment, but excludes slice portions that encode slices located outside the spatial segment.
[0175] Example embodiment 15, a video data stream according to any one of example embodiments 1 to 14, wherein the second information (142) defines the shape and size of the reduced image region.
[0176] Example embodiment 16, a video data stream according to any one of example embodiments 1 to 15, wherein the first information (140) defines the shape and size of the reduced image region by a size parameter in a parameter set substitution, wherein the parameter set substitution shall be used to substitute the corresponding parameter set in the video data stream when the reduced video data stream (136) is extracted from the video data stream.
[0177] Example embodiment 17, an apparatus for generating video data streams, is configured as follows:
[0178] The video data stream is provided with a sequence of slice portions, each slice portion containing a corresponding slice from a plurality of slices of images encoded in the video, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within the image region of the video;
[0179] Extraction information is provided to the video data stream, indicating how to extract a reduced video data stream (136) from the video data stream by: restricting the video data stream to slice portions where any slice within the spatial segment is encoded, and modifying the slice addresses to relate to the reduced image region of the smaller video, wherein the extraction information includes:
[0180] First information defines a spatial segment within an image area, within which the video is encoded in the video data stream independently of the spatial segment, wherein multiple slices do not cross the boundary of the spatial segment; and
[0181] The second information, which signals one of several options regarding how to modify the slice address of the slice portion of each slice within the spatial segment to indicate the position of the corresponding slice in the reduced image region within the reduced video data stream, or explicitly signals how to modify the slice address of the slice portion of each slice within the spatial segment to indicate the position of the corresponding slice in the reduced image region within the reduced video data stream.
[0182] Example 18: An apparatus for extracting a reduced video data stream from a video data stream in which video is encoded, the reduced video data stream encoding a smaller video, the video data stream comprising a sequence of slice portions, each slice portion encoding a corresponding slice from a plurality of slices of video images, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within the image region of the video, wherein the apparatus is configured to:
[0183] Information is read and extracted from the video data stream.
[0184] Spatial segments within the image region are derived from the extracted information, wherein none of the plurality of slices cross the boundaries of the spatial segments, and wherein the reduced video data stream is restricted to slice portions therein that encode any slice within the spatial segments.
[0185] The slice address of the slice portion of each slice within the spatial segment is modified by using one of several options determined from multiple choices through explicit signaling of extracted information.
[0186] This is to indicate the location of the corresponding slice within the reduced image area of the smaller video within the reduced video data stream.
[0187] Example embodiment 19, the apparatus according to example embodiment 17 or 18, wherein the video data stream is the video data stream according to any one of example embodiments 2 to 16.
[0188] Example embodiment 20, a video data stream wherein video (332) is encoded, wherein the video data stream includes a signal (352) indicating that pictures (334) of the video display the same scene content (304) at different resolutions in different spatial portions (354, 356) of the pictures.
[0189] Example embodiment 21, according to the video data stream of example embodiment 20, wherein the signal (352) includes a flag that is switchable between a first state and a second state, wherein the first state indicates that the images (334) of the video include different spatial portions, the images displaying the same scene content at different resolutions in the different spatial portions, wherein the second state indicates that the images of the video display each portion of the scene at only one resolution, wherein the flag is set to the first state.
[0190] Example embodiment 22, a video data stream according to example embodiment 21, wherein the flag is included in the sequence parameter set or video parameter set of the video data stream.
[0191] Example embodiment 23, a video data stream according to any one of example embodiments 20 to 22, wherein the video data stream is encoded in a manner independent of the outside of the respective spatial region (360) for each of the set of non-overlapping spatial regions (360).
[0192] Example embodiment 24, a video data stream according to any one of example embodiments 20 to 23, wherein the signal (352) includes information about counts of different resolutions, wherein the images of the video display the same scene content at the different spatial portions at the different resolutions.
[0193] Example embodiment 25, a video data stream according to any one of example embodiments 20 to 24, wherein the signal (352) includes information about spatially subdividing images of the video into segments (354, 356), each segment displaying the same scene content, and at least two segments displaying the same scene content at different resolutions.
[0194] Example embodiment 26, a video data stream according to any one of example embodiments 20 to 25, wherein the signal (352) includes different information about resolution, wherein the images of the video display the same scene content at the different spatial portions (354, 356) at the resolution.
[0195] Example embodiment 27, a video data stream according to any one of example embodiments 20 to 26, wherein the video data stream includes extraction information indicating a set of extraction regions, for which the video data stream allows the extraction of independent video data streams from the video data without re-encoding, the independent video data streams representing videos restricted to the picture content of the video by the corresponding extraction regions.
[0196] Example embodiment 28, a video data stream according to example embodiment 27, wherein the video data stream is encoded in a manner independent of the outside of the corresponding spatial region (360) for each spatial region in the set of non-overlapping spatial regions (360), and the extraction information indicates each extraction region in the set of extraction regions as a set of one or more of the spatial regions (360).
[0197] Example embodiment 29, a video data stream according to example embodiment 27 or 28, wherein the images of the video are spatially subdivided into an array of tiles and are sequentially encoded in the data stream, wherein the spatial region (360) is a cluster of one or more tiles.
[0198] Example 30, a video data stream according to any one of Example Examples 27 to 29, wherein the video data stream includes an additional signal indicating, for each of at least one or more extraction regions:
[0199] Whether the images in the video are displayed at different resolutions within the corresponding extraction area, showing the same scene content.
[0200] Example 31, the video data stream according to Example 30, wherein the additional signal further indicates for each of at least one or more extraction regions:
[0201] The viewport direction of the corresponding extracted area segment of the same scene content within the corresponding extracted area, displayed at the highest resolution, and / or
[0202] The area share of a segment of the corresponding extracted region that displays the same scene content at the highest resolution within the corresponding extracted region, and / or the area of the segment within the total area of the corresponding extracted region.
[0203] The corresponding extracted region is further subdivided into segments that display the same scene content at different resolutions.
[0204] Example embodiment 32, a video data stream according to example embodiment 30, wherein the additional signal is contained in an SEI message separate from one or more SEI messages containing the extracted information.
[0205] Example embodiment 33, an apparatus for processing a video data stream according to any one of example embodiments 20 to 32, wherein the apparatus supports a predetermined processing task and is configured to examine the signal to determine whether to perform or not perform the predetermined processing task on the video data stream.
[0206] Example embodiment 34, according to the apparatus described in example embodiment 33, wherein the processing task includes:
[0207] Decode and / or
[0208] Extract the reduced video data stream from the video data stream without re-encoding.
[0209] Example embodiment 35, an apparatus for generating a video data stream according to any one of example embodiments 20 to 32.
[0210] Example 36: A method for generating a video data stream, the method comprising:
[0211] The video data stream is provided with a sequence of slice portions, each slice portion containing a corresponding slice from a plurality of slices of images encoded in the video, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within the image region of the video;
[0212] Extraction information is provided to the video data stream, indicating how to extract a reduced video data stream from the video data stream by: restricting the video data stream to slice portions thereof, where any slice within the spatial segment is encoded, and modifying the slice addresses to relate to the reduced image region of the smaller video, wherein the extraction information includes:
[0213] First information defines a spatial segment within an image area, within which the video is encoded in the video data stream independently of the spatial segment, wherein multiple slices do not cross the boundary of the spatial segment; and
[0214] The second information, which signals one of several options regarding how to modify the slice address of the slice portion of each slice within the spatial segment to indicate the position of the corresponding slice in the reduced image region within the reduced video data stream, or explicitly signals how to modify the slice address of the slice portion of each slice within the spatial segment to indicate the position of the corresponding slice in the reduced image region within the reduced video data stream.
[0215] Example 37: A method for extracting a reduced video data stream from a video data stream in which video is encoded, the reduced video data stream encoding a smaller video, the video data stream comprising a sequence of slice portions, each slice portion encoding a corresponding slice from a plurality of slices of images of video, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within the image region of the video, wherein the method comprises:
[0216] Information is read and extracted from the video data stream.
[0217] Spatial segments within the image region are derived from the extracted information, wherein none of the plurality of slices cross the boundaries of the spatial segments, and wherein the reduced video data stream is restricted to slice portions therein that encode any slice within the spatial segments.
[0218] The slice address of the slice portion of each slice within the spatial segment is modified by using one of several options determined from multiple choices through explicit signaling of extracted information.
[0219] This is to indicate the location of the corresponding slice within the reduced image area of the smaller video within the reduced video data stream.
[0220] Example 38: A method for processing a video data stream according to any one of Example 20 to 32, wherein the processing includes a predetermined processing task, and the method includes checking the signal to determine whether to perform or not perform the predetermined processing task on the video data stream.
[0221] Example 39, a method for generating a video data stream according to any one of Example 20 to 32.
[0222] Example embodiment 40: A computer-readable medium storing a computer program having program code that, when run on a computer, performs the method described according to example embodiments 36, 37, 38, or 39.
[0223] Reference List
[0224] [1]ITU-T,Recommendation H.265(04 / 13),SERIES H:AUDIOVISUAL ANDMULTIMEDIA SYSTEMS,Infrastructure of audiovisual services-Coding of movingvideo,High efficiency video coding,Online:http: / / www.itu.int / rec / T-REC-H.265-201304-I;
[0225] [2]Jill Boyce et al.,JCTVC-Z1005,HEVC Additional SupplementalEnhancement Information(Draft 1),Joint Collaborative Team on Video Coding(JCT-VC)of ITU-T SG 16WP 3and ISO / IEC JTC 1 / SC 29 / WG 11,26th Meeting:Geneva,CH,12-20January 2017。
Claims
1. A method for storing video, comprising: Storing video data streams on digital storage media The video data stream includes a sequence of slice portions, each slice portion containing a corresponding slice from a plurality of slices of video images, wherein each slice portion includes a slice address, the slice address indicating the position of the slice encoded in the corresponding slice portion within the image region of the video; The video data stream further includes extraction information for extracting a reduced video data stream from the video data stream, wherein the reduced video data stream encodes a smaller video corresponding to a spatial segment of the video in the video data stream. The extraction involves: restricting the video data stream to slice portions where any slice within the spatial segment is encoded, and modifying the slice address to relate to a reduced image region of the smaller video. The extraction information includes: First information, the first information defines the spatial segment within the image area; and The second information explicitly signals the following: a substitute slice address for the slice portion of the slice within the spatial segment, the substitute slice address indicating the position of the slice within the reduced image region in the reduced video data stream.
2. The method according to claim 1, wherein, The first information defines the spatial segment within the image area as a set of at least two segment regions, wherein the video is encoded in the video data stream independently of the corresponding segment region, and wherein none of the multiple slices crosses the boundary of any of the at least two segment regions.
3. The method according to claim 1, wherein, The slice address of each slice portion of the video data stream is one-dimensionally indexed to the encoding start position of the slice encoded in the corresponding slice portion along a first coded scan order that traverses the image region, and the substitute value used to replace the slice address of the corresponding slice portion is one-dimensionally indexed to the encoding start position of the slice encoded in the corresponding slice portion along a second coded scan order that traverses the reduced image region.
4. The method according to claim 1, wherein, The second information provides the alternative slice addresses for the slices in the order in which they appear in the video data stream.
5. The method according to claim 1, wherein, The video data stream includes a sequence of slice portions, which encode slices located within the spatial segment and slice portions that encode slices located outside the spatial segment.
6. The method according to claim 1, wherein, The extracted information also defines the shape and size of the reduced image region.
7. The method according to claim 6, wherein, The extracted information defines the shape and size of the reduced image region using the size parameter in the parameter set substitution value, wherein the parameter set substitution value should be used to replace the corresponding parameter set in the video data stream when extracting the reduced video data stream from the video data stream.
8. An apparatus for generating a video data stream, comprising: A means for providing a sequence of slice portions to the video data stream, each slice portion containing a corresponding slice of a plurality of slices of a picture of the video, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within the picture region of the video; A means for providing extraction information to the video data stream, the extraction information being used to extract a reduced video data stream from the video data stream, wherein the reduced video data stream encodes a spatially smaller video corresponding to a spatial segment of the video in the video data stream, the extraction involving: restricting the video data stream to slice portions wherein any slice within the spatial segment is encoded, and modifying the slice address to relate to a reduced image region of the spatially smaller video, wherein the extraction information includes: First information, the first information defines the spatial segment within the image area, within which the video is encoded in the video data stream independently of the spatial segment outside the spatial segment; as well as The second information explicitly signals the following: a substitute slice address for the slice portion of the slice within the spatial segment, the substitute slice address indicating the position of the slice in the reduced image region within the reduced video data stream.
9. An apparatus for extracting a reduced video data stream from a video data stream in which video is encoded, the reduced video data stream encoding a smaller video, the video data stream comprising a sequence of slice portions, each slice portion encoding a corresponding slice of a plurality of slices of images of video, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within an image region of the video, wherein the apparatus comprises: A means for reading and extracting information from the video data stream. A means for deriving spatial segments within the image region from the extracted information, wherein the reduced video data stream is restricted to slice portions encoded with any slices within the spatial segment. A means for explicitly replacing the slice address of a slice portion of a slice within the spatial segment with a signaled alternative slice address using the extracted information, the alternative slice address indicating the position of the slice within the reduced video data stream in the reduced image region of the smaller video space.
10. A method for generating a video data stream, the method comprising: The video data stream is provided with a sequence of slice portions, each slice portion containing a corresponding slice from a plurality of slices of images encoded in the video, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within the image region of the video; Extraction information is provided to the video data stream for extracting a reduced video data stream from the video data stream, wherein the reduced video data stream encodes a smaller video corresponding to a spatial segment of the video. The extraction restricts the video data stream to slice portions where any slice within the spatial segment is encoded, and modifies the slice addresses to be associated with the reduced image region of the smaller video. The extraction information includes: First information, the first information defines the spatial segment within the image area; as well as The second information explicitly signals the following: a substitute slice address for the slice portion of the slice within the spatial segment, the substitute slice address indicating the position of the slice within the reduced image region in the reduced video data stream.
11. A method for extracting a reduced video data stream from a video data stream in which video is encoded, the reduced video data stream encoding a smaller video, the video data stream comprising a sequence of slice portions, each slice portion encoding a corresponding slice of a plurality of slices of images of video, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within an image region of the video, wherein the method comprises: Information is read and extracted from the video data stream. Spatial segments within the image region are derived from the extracted information, wherein the reduced video data stream is restricted to slice portions encoding any slices within the spatial segments. The extracted information is used to explicitly replace the slice address of the slice portion within the spatial segment with a signal-notified alternative slice address, the alternative slice address indicating the position of the slice within the reduced video data stream in the reduced image region of the smaller video.
12. A non-transitory digital storage medium storing thereon a computer program for performing a method for generating a video data stream, said computer program, when run by a computer, the method comprising: The video data stream is provided with a sequence of slice portions, each slice portion containing a corresponding slice from a plurality of slices of images encoded in the video, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within the image region of the video; Extraction information is provided to the video data stream for extracting a reduced video data stream from the video data stream, wherein the reduced video data stream encodes a spatially smaller video corresponding to a spatial segment of the video in the video data stream. The extraction involves: restricting the video data stream to slice portions where any slice within the spatial segment is encoded, and modifying the slice addresses to relate to a reduced image region of the spatially smaller video. The extraction information includes: First information, the first information defines the spatial segment within the image area, within the spatial segment the video is encoded in the video data stream independently of the space segment outside the spatial segment; as well as The second information explicitly signals the following: a substitute slice address for the slice portion of the slice within the spatial segment, the substitute slice address indicating the position of the slice in the reduced image region within the reduced video data stream.
13. A non-transitory digital storage medium storing thereon a computer program for performing a method for extracting a reduced video data stream from a video data stream in which video is encoded, the reduced video data stream encoding a smaller video, the video data stream comprising a sequence of slice portions, each slice portion encoding a corresponding slice of a plurality of slices of video images, wherein each slice portion includes a slice address indicating the position of the slice encoded in the corresponding slice portion within an image region of the video, wherein the method, when executed by a computer, comprises: Information is read and extracted from the video data stream. Spatial segments within the image region are derived from the extracted information, wherein the reduced video data stream is restricted to slice portions encoding any slices within the spatial segments. The extracted information is used to explicitly replace the slice address of the slice portion within the spatial segment with a signal-notified alternative slice address, the alternative slice address indicating the position of the slice within the reduced video data stream in the reduced image region of the smaller video.
Citation Information
Patent Citations
Frame splitting in video coding
US20120170648A1
Video composition
WO2016026526A2