Broadcast streaming of panoramic video for interactive clients
By packetizing the encoded data of different spatial segments or sequence spatial segment groups of panoramic video into a separate substream and selectively extracting and combining the substreams on the receiver side, the problem of difficulty in sending panoramic videos higher than the decoder resolution in the prior art is solved, and efficient video transmission is achieved.
Patent Information
- Application Number
- CN202111445789.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-05-26
- Filing Date
- 2017-03-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2037-03-28
AI Technical Summary
The prior art is difficult to efficiently send panoramic videos to the decoder at resolutions higher than the decoder can decode.
By packetizing encoded data of different spatial segments or sequence spatial segment groups of the panoramic video on the transmitter side into a separate substream, and selectively extracting and combining appropriate substreams on the receiver side, to form a data stream that can be decoded by the decoder.
It realizes sending high-resolution panoramic video to the decoder, solving the problem that the decoder cannot handle high-resolution video, and improving the efficiency and feasibility of video transmission.
Smart Images

Figure CN114125504B_ABST
Abstract
Description
[0001] This application is a divisional application whose applicant is Fraunhofer-Gesellschaft, whose application date is March 28, 2017, whose application number is 201780046872.3, and whose invention name is “Broadcast stream of panoramic video for interactive clients”. Technical Field
[0002] Embodiments relate to stream multiplexers. Additional embodiments relate to stream demultiplexers. Additional embodiments relate to video streams. Some embodiments relate to broadcast streams of panoramic video for interactive clients. Background Art
[0003] Region of Interest (RoI) streaming, i.e., interactive panoramic streaming, is becoming increasingly popular. The idea behind this streaming service is to enable navigation within a very wide angle and high resolution video, while showing only a portion of the entire video, i.e., the RoI, on the receiver side.
[0004] Typically, the entire panoramic video is encoded at a very high resolution, e.g., 16K, and cannot be sent to the receiver as is because it cannot be decoded by existing hardware, e.g., 4K. Summary of the invention
[0005] Therefore, according to a first aspect of the present application, the object of the invention is to provide a concept that allows to send a panoramic video to a decoder even if the panoramic video comprises a higher resolution than the decoder is able to decode.
[0006] This object is solved by the independent claims of the first aspect.
[0007] Embodiments of the first aspect provide a stream multiplexer comprising a receiving interface and a data stream former. The receiving interface is configured to receive coded data of each of at least two different spatial segments or different sequence spatial segment groups of a video image of a video stream. The data stream former is configured to packetize the coded data of each of at least two different spatial segments or different sequence spatial segment groups into a separate substream and provide the separate substream at an output.
[0008] An embodiment provides a stream demultiplexer comprising a data stream former and an output interface. The data stream former is configured to selectively extract at least two separate substreams from a separate substream group, the at least two separate substreams containing coded data of different spatial segments or different sequential spatial segment groups of video images of a coded video stream, wherein the data stream former is configured to combine the at least two separate substreams into a data stream containing coded data of different spatial segments or different sequential spatial segment groups of video images of the coded video stream.
[0009] In an embodiment, in order to transmit a panoramic video including a resolution higher than that which can be decoded by a decoder, at the transmitter side, the coded data of different spatial segments or different groups of spatial segments of a video image of a coded video stream are packetized into separate substreams to obtain separate groups of substreams. At the receiver side, from the separate groups of substreams, a suitable subset (i.e., only a part) of the separate substreams is extracted and combined into a data stream containing coded data of a suitable subset (i.e., only a part) of the spatial segments or groups of sequential spatial segments of the video images of the respective coded video streams. Thus, a decoder decoding the data stream can decode only a sub-area of the video images of the video stream, which sub-area is defined by the spatial segments or groups of spatial segments coded in the coded data contained in the data stream.
[0010] Another embodiment provides a method for stream multiplexing, the method comprising:
[0011] - receiving coded data for each of at least two different spatial segments or different sequential groups of spatial segments of a video picture of a video stream; and
[0012] - Packetizing the coded data for each of the at least two different spatial segments into separate sub-streams.
[0013] Another embodiment provides a method for stream demultiplexing, the method comprising:
[0014] - selectively extracting from the group of individual substreams at least two individual substreams, said at least two individual substreams containing coded data of different spatial segments or different sequential groups of spatial segments of a video picture of the coded video stream; and
[0015] - combining the at least two separate substreams into a data stream containing coded data of different spatial segments or different sequential groups of spatial segments of a video picture of the coded video stream.
[0016] Another embodiment provides an encoder configured to encode video images of a video stream by encoding at least two different spatial segments or different sequential spatial segment groups of the video images of the video stream, so that the encoded data includes at least two slices. The encoder can be configured to provide signaling information indicating whether the coding constraint is satisfied. The coding constraint is satisfied when at least one slice header of the at least two slices is removed while maintaining the slice header of the first slice of the at least two slices with respect to the coding order and concatenating the at least two slices or an appropriate subset of the at least two slices using the first slice header to produce a data stream that complies with the standard.
[0017] Further embodiments provide a set of individual substreams, wherein each of the individual substreams contains coded data for a different spatial segment or a different sequential group of spatial segments of a video picture of the coded video stream.
[0018] Advantageous embodiments are addressed in the dependent claims.
[0019] An embodiment provides a transmitter for encoding a video stream. The transmitter includes an encoding stage and a stream multiplexer. The encoding stage is configured to encode at least two different spatial segments or different sequence spatial segment groups of a video image of the video stream to obtain encoded data for each of the at least two different spatial segments or different sequence spatial segment groups. The stream multiplexer includes a receiving interface, which is configured to receive encoded data for each of the at least two different spatial segments or different sequence spatial segment groups of the video image of the video stream. In addition, the stream multiplexer includes a data stream former, which is configured to packetize the encoded data of each of the at least two different spatial segments into a separate sub-stream and provide the separate sub-stream at the output.
[0020] In an embodiment, the encoding stage of the transmitter may be configured to construct the video images of the video stream in spatial segments and encode the spatial segments individually or sequential (ordered with respect to encoding order) groups of spatial segments to obtain coded data for each of the spatial segments or sequential groups of spatial segments. The data stream former of the transmitter may be configured to packetize the coded data for each of the spatial segments or sequential groups of spatial segments into separate sub-streams and provide the separate sub-streams at the transmitter output.
[0021] For example, the encoding stage of the transmitter may be configured to encode a first spatial segment or a first set of sequential spatial segments to obtain first coded data, and to encode a second spatial segment or a second set of sequential spatial segments to obtain second coded data. The data stream former of the transmitter may be configured to packetize the first coded data in a first substream or a first set of substreams, and to packetize the second coded data in a second substream or a second set of substreams.
[0022] In an embodiment, the encoding stage of the transmitter may be configured to receive a plurality of video streams, each video stream comprising image segments of a video image, i.e., the image segments contained in the plurality of video streams together form a video image. Thus, the encoding stage may be configured to construct the video images of the video streams in spatial segments, such that each spatial segment or group of spatial segments corresponds to one of the video streams.
[0023] In an embodiment, individual sub-streams or groups of individual sub-streams may be transmitted, broadcast or multicast.
[0024] In an embodiment, the encoding stage may be configured to encode at least two different spatial segments separately such that the encoded data of each of the at least two different spatial segments or different sequential spatial segment groups is decodable by itself. For example, the encoding stage may be configured such that no information is shared between different spatial segments. In other words, the encoding stage may be configured to encode the video stream such that inter-frame prediction is constrained in such a way that a spatial segment of a video image is not predicted from a different spatial segment of a previous video image.
[0025] In an embodiment, the encoding stage of the transmitter can be configured to encode at least two different spatial segments or different sequence spatial segment groups of a video image of a video stream, such that the encoded data of each of the at least two different spatial segments or different sequence spatial segment groups comprises at least one slice, wherein the data stream former of the transmitter can be configured to pack each slice in a separate sub-stream.
[0026] For example, the encoding stage of the transmitter may be configured to encode each spatial segment so that the encoded data of each spatial segment includes one slice, i.e., one slice per spatial segment. In addition, the encoding stage of the transmitter may be configured to encode each spatial segment so that the encoded data of each spatial segment includes two or more slices, i.e., two or more slices per spatial segment. In addition, the encoding stage of the transmitter may be configured to encode each sequence spatial segment group so that the encoded data of each sequence spatial segment group includes one slice, i.e., one slice per group of sequence spatial segments (e.g., two or more spatial segments). In addition, the encoding stage of the transmitter may be configured to encode each group of sequence spatial segments so that the encoded data of each group of sequence spatial segments includes two or more slices, i.e., two or more slices per group of sequence spatial segments (e.g., two or more spatial segments), such as one slice per spatial segment, two or more slices per spatial segment, etc.
[0027] In an embodiment, the data stream former of the transmitter may be configured to pack each slice in a separate sub-stream without a slice header. For example, the data stream former of the transmitter may be configured to remove the slice header from the slice before packing the slice in the separate sub-stream.
[0028] In an embodiment, the data stream former may be configured to provide a further separate stream (or more than one, e.g. different image sizes) comprising a suitable slice header and the required parameter sets. For example, a suitable slice header may contain a parameter set referencing a new image size.
[0029] In an embodiment, the data stream former of the transmitter may be configured to generate a sub-stream descriptor which assigns a unique sub-stream identification to each individual sub-stream.
[0030] For example, the data stream former may be configured to provide a separate substream group, the separate substream group comprising a separate substream and a substream descriptor, for example, a program map table comprising substream descriptors (e.g., in the program map table, a substream descriptor may be assigned to each substream, i.e., one substream descriptor per substream). The substream descriptor may be used to find a separate substream in the separate substream group based on a unique substream identifier.
[0031] In an embodiment, the data stream former of the transmitter may be configured to generate a sub-region descriptor signaling a pattern of sub-stream identifications of sub-regions of the video image belonging to the video stream, or in other words, for each of at least one spatial subdivision of the video image of the video stream into sub-regions, signaling a set of sub-stream identifications for each sub-region.
[0032] For example, the data stream former may be configured to provide a set of individual substreams, the set of individual substreams comprising individual substreams and subregion descriptors, e.g., a program map table comprising subregion descriptors. The subregion descriptors describe which substreams contain coded data for coding spatial segments or groups of sequential spatial segments that together form a valid subregion (e.g., a contiguous subregion) of a video image of the video stream. The subregion descriptors may be used to identify an appropriate subset of substreams belonging to a suitable subregion of a video image of the video stream.
[0033] In an embodiment, the encoding stage of the transmitter may be configured to combine the coded data of at least two different spatial segments or different sequence spatial segment groups into one slice. The data stream former of the transmitter may be configured to split a slice or a bit stream of a slice at the spatial segment boundaries in the slice portion and pack each slice portion in a substream.
[0034] In an embodiment, at least two different spatial segments or different sequential spatial segment groups may be encoded into a video stream with one slice per video picture. Thus, the data stream former may be configured to, for each video picture, pack a portion of a slice in which a corresponding one of the tiles or tile groups is encoded into a separate substream, while for at least one of the at least two tiles or tile groups there is no slice header of the slice.
[0035] For example, at least two different spatial segments or groups of different sequence spatial segments may be entropy encoded by the encoding stage. Thus, the encoding stage may be configured to entropy encode the at least two different spatial segments or groups of different sequence spatial segments such that each of the at least two different spatial segments or groups of different sequence spatial segments is decodable by itself, i.e., such that no encoding information is shared between the at least two different spatial segments or groups of different sequence spatial segments and other at least two different spatial segments or groups of different sequence spatial segments of previous video images of the video stream. In other words, the encoding stage may be configured to reinitialize the entropy encoding of the different spatial segments or groups of different spatial segments after encoding each of the different spatial segments or groups of different spatial segments.
[0036] In an embodiment, the transmitter may be configured to signal the stream type.
[0037] For example, a first stream type may signal that individual substreams may be aggregated according to information found in the sub-region descriptors to produce a data stream that complies with the standard.
[0038] For example, the second stream type may signal that aggregation of individual sub-streams according to information found in the sub-region descriptors produces a data stream that needs to be modified or further processed to obtain a version of the data stream that complies with the standard.
[0039] In an embodiment, the data stream former of the transmitter may be configured to provide a transport stream comprising the individual sub-streams. The transport stream may be, for example, an MPEG-2 transport stream.
[0040] In an embodiment, the sub-stream may be a base stream. According to an alternative embodiment, multiple sub-streams may be transmitted together via one base stream.
[0041] In an embodiment, the spatial segment of the video image of the video stream may be a tile. For example, the spatial segment of the video image of the video stream may be a HEVC tile.
[0042] In an embodiment, the encoding stage of the transmitter may be an encoding stage compliant with a standard. For example, the encoding stage of the transmitter may be an encoding stage compliant with the HEVC (HEVC=High Efficiency Video Coding) standard.
[0043] In an embodiment, the stream multiplexer may be configured to signal at least one of two stream types. A first stream type may signal that a combination of an appropriate subset of individual substreams corresponding to at least one spatial subdivision of the video image of the video stream to the sub-region produces a data stream that complies with the standard. A second stream type may signal that a combination of an appropriate subset of individual substreams corresponding to at least one spatial subdivision of the video image of the video stream to the sub-region produces a data stream that needs to be further processed (e.g., whether header information must be added or modified, or whether a parameter set must be adjusted) to obtain a version of the data stream that complies with the standard.
[0044] In an embodiment, at least two different spatial segments or different sequence spatial segment groups of a video image of a video stream are encoded so that the encoded data includes at least two slices. Therefore, the encoded data may include signaling information indicating whether the encoding constraint is met, or wherein the data stream former is configured to determine whether the encoding constraint is met. When at least one slice header of at least two slices is removed, while the slice header of the first slice of at least two slices is maintained with respect to the coding order and the first slice header is used to cascade at least two slices or an appropriate subset of at least two slices to produce a data stream that conforms to the standard, the encoding constraint is met. The stream multiplexer may be configured to signal at least one of the two stream types according to the encoding constraint (e.g., if the encoding constraint is met, the first stream type is signaled and / or otherwise the second stream type is signaled). In addition, the multiplexer may be configured to signal the second stream type even if the encoding constraint is met, but the parameter set must be adjusted to obtain a data stream that conforms to the standard.
[0045] Other embodiments provide a receiver. The receiver includes a stream demultiplexer and a decoding stage. The stream demultiplexer includes a data stream former and an output interface. The data stream former is configured to selectively extract at least two separate substreams from the separate substream groups, the at least two separate substreams containing coded data of different spatial segments or different sequence spatial segment groups of a video image of the coded video stream, wherein the data stream former is configured to combine the at least two separate substreams into a data stream containing coded data of different spatial segments or different sequence spatial segment groups of a video image of the coded video stream. The output interface is configured to provide the data stream. The decoding stage is configured to decode the coded data contained in the data stream to obtain at least two different spatial segments of the video image of the video stream.
[0046] In an embodiment, the decoding stage of decoding the data stream may decode only a sub-region of a video image of the video stream, the sub-region being defined by a spatial segment or a group of spatial segments encoded in the encoded data contained in the data stream.
[0047] In an embodiment, the receiver may be configured to receive, for example from the above-mentioned transmitter, a group of individual substreams, the group of individual substreams comprising individual substreams. Each of the individual substreams may include coded data of a spatial segment or a group of sequential spatial segments of a plurality of spatial segments in which a video image of the coded video stream is constructed.
[0048] In an embodiment, the data stream former of the receiver may be configured to selectively extract a suitable subset of the individual substreams from the group of individual substreams, the suitable subset of the individual substreams containing coded data of a suitable subset of the spatial segments of the video images or of the sequential spatial segments of the coded video stream. The data stream former of the receiver may be configured to combine the individual substreams extracted from the group of individual substreams into a new data stream containing coded data of a suitable subset of the spatial segments of the video images or of the different sequential spatial segments of the coded video stream. The decoding stage of the receiver may be configured to decode the coded data contained in the data stream to obtain a suitable subset of the spatial segments of the video images or of the sequential spatial segments of the video stream.
[0049] For example, the data stream former of the receiver may be configured to selectively extract only a part of the plurality of individual substreams contained in the individual substream group and combine the individual substreams extracted therefrom into a new data stream. Since the individual substreams extracted from the individual substream group contain only a part of the spatial segment in which the video image of the video stream is constructed, the new data stream combined from the individual substreams extracted from the individual substream group also contains coded data of only a part of the spatial segment in which the video image of the coded video stream is constructed.
[0050] In an embodiment, the decoding stage of the receiver can be configured to decode a data stream containing only the encoded data contained in the extracted subset of the multiple individual sub-streams, thereby decoding a sub-region of a video image of the video stream, the sub-region of the video image being smaller than the video image and the sub-region being defined by spatial segments encoded in the encoded data contained in the extracted subset of the multiple individual elementary streams.
[0051] In an embodiment, a data stream former of the receiver may be configured to extract sub-region descriptors from the group of separate sub-streams, the sub-region descriptors signaling a pattern of sub-stream identifications of sub-regions of a video image belonging to the video stream, wherein the data stream former may be configured to select a subset of sub-streams to be extracted from the group of separate sub-streams using the sub-region descriptors.
[0052] For example, the data stream former of the receiver may be configured to extract sub-region descriptors from the individual sub-stream groups, e.g., a program map table comprising sub-region descriptors. The sub-region descriptors describe which sub-streams contain coded data for coding spatial segments or groups of sequential spatial segments, which together form a valid sub-region (e.g., a continuous sub-region) of a video image of the video stream. Based on the sub-region descriptors, the data stream former of the receiver may select an appropriate subset of sub-streams for extraction, which belong to a suitable sub-region of the video image of the video stream.
[0053] In an embodiment, a data stream former of the receiver may be configured to extract substream identifiers from the group of individual substreams, the substream identifiers assigning a unique substream identification to each individual substream, wherein the data stream former may be configured to localize a subset of substreams to be extracted from the group of individual substreams using the substream descriptors.
[0054] For example, the data stream former of the receiver may be configured to extract substream descriptors, e.g., a program map table including substream descriptors, from the group of individual substreams. Based on the substream descriptors that assign a unique substream identifier to each individual substream, the data stream former of the receiver may identify (e.g., find or locate) the individual substreams in the group of individual substreams. According to an alternative embodiment, the descriptor in the adaptation field is used to identify a data packet belonging to a certain substream.
[0055] In some embodiments, the receiver may include a data stream processor 127. The data stream processor may be configured to further process the data stream 126 if the data stream 126 provided by the data stream former 121 does not conform to the standard, i.e. cannot be decoded by the decoding stage 124, to obtain a processed version 126' (i.e. a version conforming to the standard) of the data stream 126. If further processing is required, the transmitter 100 may signal this in the stream type.
[0056] For example, a first stream type may signal that aggregating the individual substreams according to the information found in the sub-region descriptors produces a data stream that complies with the standard. Thus, if the first stream type is signaled, no further processing of the data stream 126 is required, i.e., the data stream processor 127 may be bypassed and the decoding stage 124 may directly decode the data stream 126 provided by the data stream former 122. A second stream type may signal that aggregating the individual substreams according to the information found in the sub-region descriptors produces a data stream that needs to be modified or further processed to obtain a version of the data stream that complies with the standard. Thus, if the first stream type is signaled, further processing of the data stream 126 is required, i.e., in this case, the data stream processor 127 may further process the data stream 126 to obtain a processed version 126' (i.e., a version that complies with the standard) of the data stream 126.
[0057] In an embodiment, the stream demultiplexer may be configured to identify the stream type of the signaled separate substream group from at least two stream types. The first stream type may indicate that a combination of an appropriate subset of separate substreams corresponding to one of at least one spatial subdivision of the video image of the video stream to the sub-area produces a data stream that complies with the standard. The second stream type indicates that a combination of an appropriate subset of separate substreams corresponding to one of at least one spatial subdivision of the video image of the video stream to the sub-area produces a data stream that requires further processing to obtain a version of the data stream that complies with the standard. The stream demultiplexer may include a data stream processor configured to further process the data stream using processing information contained in at least one substream in the separate substream group to obtain a version of the data stream that complies with the standard.
[0058] In an embodiment, the individual substream groups may be transport streams. The transport stream may be, for example, an MPEG-2 transport stream.
[0059] In an embodiment, a sub-stream may be an elementary stream.
[0060] In an embodiment, the spatial segment of the video image of the video stream may be a tile. For example, the spatial segment of the video image of the video stream may be a HEVC tile.
[0061] In an embodiment, the data stream former of the receiver may be configured to combine the individual substreams extracted from the individual substream groups into a data stream that complies with the standard. The data stream may be, for example, a data stream that complies with the HEVC standard.
[0062] In an embodiment, the decoding stage of the receiver may be a decoding stage that complies with a standard. For example, the decoding stage of the receiver may be a decoding stage that complies with the HEVC standard.
[0063] In an embodiment, the encoded data provided by the encoding stage of the transmitter forms an encoded video stream.
[0064] In an embodiment, a coded video stream comprises coded data of at least two different spatial segments or different sequential groups of spatial segments of a video image of the coded video stream using at least two slices. The coded video stream comprises signaling information indicating whether a coding constraint is satisfied, wherein the coding constraint is satisfied when at least one slice header of the at least two slices is removed while a slice header of a first slice of the at least two slices is maintained with respect to a coding order and concatenating the at least two slices or a proper subset of the at least two slices using the first slice header produces a data stream that complies with the standard.
[0065] For example, the signaling information may indicate that the values of syntax elements of the slice headers of the first slice and all other slices of the video picture are similar to the extent that removing all slice headers except the slice header of the first slice and concatenating the data produces a coded video stream compliant with the standard.
[0066] In an embodiment, the signaling information is present in the encoded video data stream in the form of a flag in the video usability information or the supplemental enhancement information.
[0067] In an embodiment, a coded video stream (e.g., provided by an encoding stage) may include more than one slice per video picture (e.g., one slice per spatial segment with a corresponding slice header) and signaling information indicating that different spatial segments satisfy coding constraints, such as using the same reference structure or slice type, i.e., the syntax elements of the corresponding slice headers are similar (e.g., slice type), in a way that only the first slice header is required and if this single slice header is concatenated with any number of slice payload data corresponding to more than one spatial segment; the resulting stream is a compliant video stream, provided that as long as the appropriate parameter sets are prepended to the stream (i.e., as long as the parameter sets are, for example, modified). In other words, if the signaling information is present, all slice headers except the first slice may be stripped, and concatenating the first slice and all other slices without their slice headers produces a coded video stream that conforms to the standard. In the case of HEVC, the signaling information of the above constraints may be present in the coded video stream in the form of markers in the Video Usability Information (VUI) or the Supplementary Enhancement Information (SEI).
[0068] According to a second aspect of the application, the invention aims to provide a concept which allows sending sub-streams or partial data streams relating only to segments of the original image in a way which allows easier processing on the receiving side.
[0069] This object is solved by the independent claims of the second aspect.
[0070] According to a second aspect, sending a substream or a partial data stream related only to a segment of an original image in a manner that allows easier processing at the receiving side is achieved by adding one or more NAL units of a second set to one or more NAL units of a first set forming a self-contained data stream parameterized to encode a first image, the NAL units of the first set being selected from one or more NAL unit types of the first set, each of the one or more NAL units of the second set being one of one or more predetermined NAL unit types of the second set, disjoint from the first set, and determined to cause a legacy decoder to ignore the corresponding NAL unit. Thus, by transparently presenting portions of the original version of the partial data stream to a legacy decoder, portions that would interfere with a legacy decoder when reconstructing a particular segment from the partial data stream can be preserved, but processing devices interested in these original portions can deduce them anyway. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Embodiments of the present invention are described herein with reference to the accompanying drawings.
[0072] Figure 1 shows a schematic block diagram of a transmitter according to an embodiment;
[0073] Figure 2 shows a schematic block diagram of a receiver according to an embodiment;
[0074] Figure 3 An illustrative view of a video image of a video stream structured in a plurality of tiles, and an illustrative view of a sub-region of the video image defined by a suitable subset of the plurality of tiles;
[0075] Figure 4 shows a schematic block diagram of a transport stream demultiplexer;
[0076] Figure 5 illustrative views of video images of a video stream structured in a plurality of tiles, and illustrative views of individual tiles and combinations of tiles that form a conforming bitstream when transmitted as separate elementary streams;
[0077] Figure 6 An illustrative view of a video image of a video stream structured in a plurality of tiles is shown, as well as an illustrative view of a sub-region of the video image defined by a suitable subset of the plurality of tiles and an offset pattern of tiles for the sub-region;
[0078] Figure 7 A flow chart of a method for stream multiplexing according to an embodiment is shown;
[0079] Figure 8 A flow chart of a method for stream demultiplexing according to an embodiment is shown;
[0080] Fig. 9 A schematic diagram illustrating an image or video being encoded into a data stream, the image or video being spatially divided into two segments so that two fragments are generated, one for each segment, ready to be multiplexed onto a segment-specific partial data stream;
[0081] Fig.10 The diagram shows a Fig. 9 to illustrate the situation that prevents one of the fragments from being cut off directly as a self-contained segment-specific data stream;
[0082] Fig.11a and 11b The diagram illustrates the modification of parameter set NAL units and / or slice NAL units from Fig.10 A self-contained segment-specific data stream derived from the corresponding fragment in ;
[0083] Fig.12 A schematic block diagram of a stream formatter and its operating mode is shown, in which the stream formatter Fig. 9 The data stream of the MPEG-4 decoder is distributed into two partial data streams, one for each segment, where one partial data stream is ready to be processed by a conventional decoder for reconstructing the corresponding segment into a self-contained picture by adding / inserting specific NAL units into Fig.11a and 11b This is achieved in the modified partial data flow shown in FIG.
[0084] Fig.13 A schematic block diagram of a processing device and its operating mode in which the processing device again uses special NAL units inserted transparently to conventional decoders from Fig.12 The partial data stream is reconstructed from the contained data stream 302';
[0085] Fig.14 Shows Fig.12 An example of a special type of added NAL unit, i.e., a hidden NAL unit that carries information hidden from legacy decoders;
[0086] Fig.15 The diagram shows the Fig.12 A schematic diagram of the mechanism of adding a second type of NAL units that form a hidden instruction to ignore subsequent NAL units in the corresponding data stream, which hidden instruction is ignored by a legacy decoder; and
[0087] Fig.16 A schematic diagram illustrating a scenario of a stream formatter is shown, which is similar to Fig.12 distributing the data stream into partial data streams in such a manner that one of them is converted into a self-contained data stream ready for processing by a legacy decoder for reconstructing corresponding segments therefrom as a self-contained picture; additionally preparing the partial data streams in such a manner that the partial data streams are concatenated as indicated by instructions signaled by specific NAL units inserted and assumed to be ignored by the legacy decoder as being of a disallowed NAL unit type; also generating a legacy-compliant data stream that may be decoded by a legacy decoder to reconstruct the entire picture; here, using NAL units forming hidden instructions to ignore a portion of subsequent NAL units; and
[0088] Fig.17 Shows the ratio Fig.16 A more detailed illustration is provided by skipping the Fig.16 Schematic diagram of the NAL unit signaling modification and instruction mechanism used in the embodiments.
[0089] One or more identical or equivalent elements having the same or equivalent function are denoted by the same or equivalent reference numerals in the following description. DETAILED DESCRIPTION
[0090] In the following description, many details are set forth to provide a more thorough explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other cases, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present invention. In addition, unless otherwise specifically stated, the features of the different embodiments described below may be combined with each other.
[0091] Although in the following description and corresponding figures, a transmitter and a receiver, a transmitter including an encoding stage and a stream multiplexer, and a receiver including a stream demultiplexer and a decoding stage are discussed and shown as examples and for illustrative purposes only, it should be noted that embodiments of the present invention relate to a stream multiplexer and a stream demultiplexer, respectively. That is, when practicing embodiments of the present invention, the encoding stage and the decoding stage may be omitted.
[0092] Figure 1 A schematic block diagram of a transmitter 100 according to an embodiment of the present invention is shown. The transmitter 100 comprises a stream multiplexer 103 and an encoding stage 104. The encoding stage 104 is configured to encode at least two different spatial segments 108 or different sequential (subsequent) spatial segments 108 groups of video images 110 of a video stream 102 to obtain coded data 112 for each of at least two different spatial segments 108 or different groups of sequential spatial segments 108. The stream multiplexer 103 comprises a receiving interface 105 and a data stream former 106. The receiving interface 105 is configured to receive the coded data for each of at least two different spatial segments or different groups of sequential spatial segments of the video images of the video stream encoded in the coded data. The data stream former 106 is configured to packetize the coded data 112 for each of at least two different spatial segments 108 or different groups of sequential spatial segments into separate substreams 114 and provide the separate substreams at the transmitter output.
[0093] The encoding stage 104 of the transmitter 100 may be configured to construct the video images 110 of the video stream in spatial segments 108. For example, the encoding stage 104 may be configured to construct the video images 110 of the video stream in N×M spatial segments 108. N may be a natural number describing the number of columns in which the video images 110 of the video stream 102 are constructed. M may be a natural number describing the number of rows in which the video images 110 of the video stream 102 are constructed. Thus, one of N and M may be greater than or equal to 2, wherein the other of N and M may be greater than or equal to 1.
[0094] like Figure 1As shown in the example of , the encoding stage 104 may be configured to construct the video images 110 of the video stream 102 in four columns (N=4) and three rows (M=3), ie, in twelve spatial segments 108.
[0095] The encoding stage 104 of the transmitter 100 may be configured to encode a first spatial segment (e.g., spatial segment 108_1,1) or a first group of consecutive spatial segments (e.g., spatial segments 108_1,1 and 108_1,2) to obtain first coded data (e.g., coded data 112_1), and to encode a second spatial segment (e.g., spatial segment 108_1,2) or a second group of consecutive spatial segments (e.g., spatial segments 108_1,3 and 108_1,4) to obtain second coded data (e.g., coded data 112_2). The data stream former of the transmitter may be configured to packetize the first coded data in a first substream (e.g., substream 114_1) or a first group of substreams (e.g., substreams 114_1 and 114_2), and to packetize the second coded data (e.g., coded data 112_2) in a second substream (e.g., second substream 114_2) or a second group of substreams.
[0096] like Figure 1 As shown in the example in , the encoding stage 104 of the transmitter 100 can be configured to encode each spatial segment 108 separately to obtain encoded data for each spatial segment 108, wherein the data stream former 106 of the transmitter 100 can be configured to packetize each encoded data 112 in a separate substream 114, i.e., one encoded data per spatial segment and one substream per encoded data. However, the encoding stage 104 of the transmitter 100 can also be configured to encode sequential spatial segment groups separately to obtain encoded data for each sequential spatial segment group, i.e., one encoded data per sequential spatial segment group (i.e., two or more sequential spatial segments). In addition, the data stream former 106 of the transmitter 100 can also be configured to packetize each encoded data 112 in more than one separate substream, i.e., one separate substream per encoded data (i.e., two or more separate substreams).
[0097] The data stream former 106 of the transmitter 100 may be configured to provide (at its output) an individual substream group 116 comprising the individual substreams 114 .
[0098] Individual sub-streams or groups of individual sub-streams may be transmitted, broadcast or multicast.
[0099] Figure 2A schematic block diagram of a receiver 120 according to an embodiment is shown. The receiver 120 comprises a stream demultiplexer 121 and a decoding stage. The stream demultiplexer comprises a data stream former 122 and an output interface 123. The data stream former 122 is configured to selectively extract at least two individual substreams 114 from the group of individual substreams 116, the at least two individual substreams 114 containing coded data of different spatial segments 108 or different groups of sequential spatial segments 108 of video images 110 of the video stream 102, wherein the data stream former 122 is configured to combine the at least two individual substreams 114 into a data stream 126 containing coded data of different spatial segments 108 or different groups of sequential spatial segments 108 of video images 110 of the coded video stream 102. The output interface 123 is configured to provide the data stream 126. The decoding stage 124 is configured to decode the encoded data contained in the data stream 126 to obtain at least two different spatial segments 108 or different groups of sequential spatial segments 108 of the video images 110 of the video stream 102 .
[0100] By decoding the data stream 126, the decoding stage 124 may decode only a sub-region 109 of a video image of the video stream, the sub-region 109 being defined by a spatial segment or a group of spatial segments encoded in the encoded data contained in the data stream, for example 108_1,1 and 108_1,2.
[0101] The individual substream group 116 may include a plurality of individual substreams 114, each of which encodes a different spatial segment of the plurality of spatial segments in which the video images of the video stream are constructed or a different group of sequential spatial segments. For example, the video images 110 of the video stream may be constructed in N×M spatial segments 108. N may be a natural number describing the number of columns in which the video images 110 of the video stream 102 are constructed. M may be a natural number describing the number of rows in which the video images 110 of the video stream 102 are constructed. Thus, one of N and M may be greater than or equal to 2, wherein the other of N and M may be greater than or equal to 1.
[0102] The data stream former 122 of the receiver 120 can be configured to selectively extract appropriate subsets of individual substreams (e.g., substreams 114_1 and 114_2) from the individual substream group 116, the appropriate subsets of individual substreams containing encoded data of appropriate subsets of spatial segments of the video image 110 of the encoded video stream 102 (e.g., spatial segments 108_1,1 and 108_1,2) or a group of sequential spatial segments (e.g., a first group of sequential spatial segments 108_1,1 and 108_1,2 and a second group of sequential spatial segments 108_1,3 and 108_1,4). The data stream former 122 of the receiver 120 can be configured to combine the individual substreams (e.g., substreams 114_1 and 114_2) extracted from the individual substream groups 116 into a new data stream 126, which contains encoded data of an appropriate subset of spatial segments of the video image 110 of the encoded video stream 102 (e.g., spatial segments 108_1,1 and 108_1,2) or a group of sequential spatial segments (e.g., a first group of sequential spatial segments 108_1,1 and 108_1,2 and a second group of sequential spatial segments 108_1,3 and 108_1,4).
[0103] The decoding stage 124 of the receiver 120 can be configured to decode the encoded data contained in the data stream 126 to obtain an appropriate subset of spatial segments (e.g., spatial segments 108_1,1 and 108_1,2) or a group of sequential spatial segments (e.g., a first group of sequential spatial segments 108_1,1 and 108_1,2 and a second group of sequential spatial segments 108_1,3 and 108_1,4) of the video image 110 of the video stream 102, i.e., a sub-area 109 of the video image 110 of the video stream defined by the spatial segments or groups of spatial segments encoded in the encoded data contained in the data stream 126.
[0104] like Figure 2 As shown in the example in Figure 1The transmitter 100 shown in the figure provides a group of individual substreams 116 comprising twelve individual substreams 114_1 to 114_12. Each of the individual substreams 114_1 to 114_12 contains coded data of one of the twelve spatial segments 108 in which the video image 110 of the coded video stream 102 is constructed, i.e., one spatial segment per coded data and one coded data per substream. The data stream former 122 can be configured to extract only the first individual substream 114_1 and the second individual substream 114_2 of the twelve individual substreams 114_1 to 114_12. The first individual substream 114_1 comprises coded data of a first spatial segment 108_1,1 of the video image 110 of the coded video stream, wherein the second individual substream 114_2 comprises coded data of a second spatial segment 108_1,2 of the video image 110 of the coded video stream. Furthermore, the data stream former 122 may be configured to combine the first individual substream 114_1 and the second individual substream 114_2 to obtain a new data stream 126 containing coded data of the first spatial segment 108_1,1 and the second spatial segment 108_1,2 of the video image of the coded video stream. Thus, by decoding the new data stream 126, the decoding stage obtains the first spatial segment 108_1,1 and the second spatial segment 108_1,2 of the video image of the video stream, i.e. the sub-area 109 of the video image 110 of the video stream which is defined only by the first spatial segment 108_1,1 and the second spatial segment 108_1,2 coded in the coded data contained in the data stream 126.
[0105] In some embodiments, the receiver 120 may include a data stream processor 127. The data stream processor may be configured to further process the data stream 126 to obtain a processed version 126' (i.e., a standard-compliant version) of the data stream 126 if the data stream 126 provided by the data stream former 121 is not compliant with the standard, i.e., is not decodable by the decoding stage 124. If further processing is required, this may be signaled by the transmitter 100 in the stream type. A first stream type may represent or indicate that the aggregation of the individual substreams according to the information found in the sub-region descriptors produces a data stream that complies with the standard. Thus, if the first stream type is signaled, no further processing of the data stream 126 is required, i.e., the data stream processor 127 may be bypassed, and the decoding stage 124 may directly decode the data stream 126 provided by the data stream former 122. A second stream type may represent or indicate that the data stream generated by the aggregation of the individual substreams according to the information found in the sub-region descriptors needs to be modified or further processed to obtain a standard-compliant version of the data stream. Therefore, if the second stream type is signaled, further processing of the data stream 126 is required, i.e., in this case, the data stream processor 127 may further process the data stream 126 to obtain a processed version 126' (i.e., a version that complies with the standard) of the data stream 126. The data stream processor 127 may use additional information contained in the data stream 126, for example, additional information contained in one of the substreams 114s and 114p of the substream group, to perform additional processing. The substream 114s may contain one or more slice headers, wherein the substream 114p may contain one or more parameter sets. If the substream group contains the substreams 114s and 114p, the data stream former 122 may also extract these substreams.
[0106] In other words, the data stream 126 may require additional processing in order to be formed into a standard-compliant data stream that can be correctly decoded by the decoding stage 124 as indicated by the stream type, i.e., whether further processing is required is indicated in the stream type (e.g., the first new stream type and the second new stream type as described below). The processing includes the use of additional information that is either placed into the encoded data 112 by the encoding stage 104 or into one of the substreams 114 by the data stream former 106 and is subsequently included in the data stream 126. By using the additional information, the data stream processor 127 specifically adjusts the encoding parameters (e.g., parameter sets) and slice headers in the data stream 126 to reflect the actual subset of 116 to be output by 123, i.e., a data stream that is different from 112 in terms of image size, for example.
[0107] In the following description, it is exemplarily assumed that the encoding stage is the HEVC encoding stage and the decoding stage is the HEVC decoding stage. However, the following description is also applicable to other encoding and decoding stages, respectively.
[0108] Furthermore, in the following description, it is exemplarily assumed that the individual substream group is a transport stream (eg, an MPEG-2 transport stream), wherein the individual substreams of the individual substream group are elementary streams.
[0109] HEVC bitstreams can be generated using a "tile" concept that destroys intra-image prediction dependencies (including entropy decoding dependencies). The data generated by the encoder for each such tile can be processed separately, for example, by one processor / core. If tiles are used, the entire video is constructed as a rectangular pattern of N×M tiles. Optionally, each tile can be included in a different slice, or many tiles can be included in the same slice. The encoder can be configured in such a way that information is not shared between different tiles. For certain use cases, such as presenting a smaller window (i.e., a region of interest (RoI)) taken from a large panorama, only a subset of the tiles need to be decoded. In particular, the HEVC bitstream can be encoded in such a way that inter-frame prediction is constrained in a way that a block of an image is not predicted from a different block of a previous image.
[0110] In this document, a portion of a bitstream that allows decoding of a tile or a subset of a tile is referred to as a substream. The substream may include a slice header that indicates the original position of the tile within the full panorama. In order to use existing hardware decoders, such a substream may be converted into a bitstream compliant with the HEVC standard by adjusting the data indicating the position of the tile before decoding. In addition, when converting the substream into a new bitstream 126 or 126', the reference to the picture parameter set (pps_id) in the slice header may also be adjusted. Thus, by indirectly referencing the sequence parameter set, parameters such as the image size may be adjusted and a bitstream compliant with the HEVC standard may be generated.
[0111] If the entire bitstream including the encoded data of all tiles is sent to the receiver via a broadcast channel, a receiver capable of decoding a smaller RoI may not be able to process the large amount of data corresponding to the complete panorama. There are different transport protocols for broadcasting, among which the MPEG-2 Transport Stream (TS) is widely used. In the MPEG-2 Systems standard, the TS is specified as a sequence of packets with a fixed length, which carries a PID (packet identifier) for identifying different ESs in the multiplexed stream. PID 0 is used to carry the PAT (Program Association Table), which points to one or more PMTs (Program Map Tables) by indicating the PID of each PMT. Within the PMT, the program map section is used to represent the attributes of the ES belonging to the program. However, these sections are limited to 1021 bytes for the description of all ESs, which typically include video and may include multiple audio streams or subtitle information, so the substream and sub-region information must be very compact.
[0112] MPEG-2 TS currently provides signaling for HEVC coded video bitstreams sent in elementary streams (ES) containing the full panorama. However, the signaling included in the TS indicates the profile / layer / level required to decode the entire bitstream and if the decoder's capabilities are insufficient to decode a bitstream with such a high level, it is very likely that the receiver will not start decoding if the target display resolution is much smaller than the full panorama.
[0113] In an embodiment, the bitstream can be split into separate ESs, and the client can select the subset required to decode the RoI from the separate ESs. An alternative option of using a descriptor in the adaptation field is described later, in which one ES transmits more than one substream. In any case, such a subset of substreams is called a subregion. In this case, the current MPEG-2TS standard does not provide signaling to tell the decoder which level the subregion conforms to. The receiver cannot also find out which ES sets need to be combined in order to decode a specific subregion. Different subregion sizes can be used, that is, a subregion set can consist of a certain number of rows and columns, while another subregion set can consist of a different number of rows and columns. They are referred to as different subregion layouts below.
[0114] Figure 3 An illustrative view of a video image 110 of a video stream constructed with a plurality of tiles 108 and an illustrative view of a sub-region 109 of the video image 110 defined by a suitable subset of the plurality of tiles 108 are shown. From the plurality of tiles 108, suitable subsets may be aggregated to obtain the sub-region 109 of the video image. For example, Figure 3 The video image 110 shown in FIG. 1 is structured into N=6 columns and M=5 rows (or rows), i.e., NxM=30 blocks. The position of each block 108 can be described by the indexes h and v, where h describes the row and v describes the column. From these 30 blocks, 9 blocks can be gathered and form a sub-region 109 of the video image.
[0115] In an embodiment, the stream sender may generate the substream 114 included in the TS 116 as a separate ES. Inside the ES, each coded picture may be encapsulated in a PES (Packetized Elementary Stream) packet. There are several options for generating the substream.
[0116] For example, according to a first option, the transmitter 100 may generate one slice per substream (ie, one block 108 or a fixed set of sequenced blocks 108) and packetize the slice data of each slice into a PES packet, thereby creating a separate ES for each substream.
[0117] According to the second option, the transmitter 100 may generate one slice per substream, and the stream multiplexer 103 may strip all slice headers before packetizing the slice data of each slice into a PES packet, thereby establishing a separate ES for each substream 114. In addition, the transmitter 100 may generate additional separate substreams 114s, for example, separate ESs, which provide a suitable slice header that, when combined with the slice data, produces a bitstream compliant with HEVC.
[0118] According to a third option, the transmitter 100 may generate only one slice containing all the blocks and split the bitstream at the block boundaries. The data portions constituting the substreams may be packed into PES packets, thereby establishing a separate ES for each substream. In addition, the transmitter 100 may generate additional separate substreams 114s, e.g., separate ESs, which provide a suitable slice header that, when combined with the slice data, produces a bitstream compliant with HEVC.
[0119] According to a fourth option, the transmitter 100 may generate one slice per substream and introduce signaling information (e.g., in the form of a flag in VUI (Video Usability Information) or SEI (Supplemental Enhancement Information)), which indicates a constraint that allows the removal of slice headers of all slices except the first slice, and the stream multiplexer 103 strips all slice headers before packetizing the slice data of each slice into a PES packet based on parsing the signaling information, thereby establishing a separate ES for each substream 114. In addition, the transmitter 100 may generate another separate substream 114s, e.g., a separate ES, which provides a suitable slice header, which, when combined with the slice data, produces a bitstream compliant with HEVC.
[0120] In the second and fourth options, the stream multiplexer (103) may add a single slice header for each video picture (each DTS (decoding time stamp)) to the further stream 114s, i.e., there may be a constraint that the PES packets in the further stream contain a single slice header, so that the demultiplexer can easily rearrange the PES packets without having to detect video picture boundaries.
[0121] In a first option, the transmitter 100 may also generate a separate sub-stream 114s, e.g. a separate ES, providing additional data consisting of one or more parameter sets or appropriate information, such as a syntactic construct containing parameter sets and supplementary information and information about their association with sub-regions, to derive parameter sets which, when combined with the slice data, allow extraction processing to be performed by the data stream processor 127, which then produces a conforming bitstream.
[0122] In the second, third and fourth options, the transmitter 100 may also include one or more additional parameter sets in the same individual substream 114, e.g. individual ES, or generate additional individual substreams 114p, e.g. individual ES including (only) these parameter sets.
[0123] In the first case, the substream 114 consisting of the backward compatible (upper left) tile 108_1,1 and optionally sequence tiles that together form a rectangular region may use the HEVC stream type and the legacy descriptors for HEVC specified in the HEVC standard.
[0124] The first new stream type may indicate that an ES contains sub-streams. The first new stream type indicates that a conforming bitstream is generated from the aggregation of the ES according to the information found in the sub-region descriptor as described below.
[0125] Furthermore, a second new stream type may indicate that an ES contains sub-streams. This second new stream type indicates that aggregation of ESs according to information found in sub-region descriptors as described below produces a bitstream that needs to be modified by processing as specified below before being decoded.
[0126] This information may be sufficient to allow aggregation of sub-regions from the appropriate set of sub-streams in the TS Demux, as will be seen below Figure 4 This is more clearly seen in the discussion.
[0127] Figure 4 A schematic block diagram of a transport stream demultiplexer 121 according to an embodiment is shown. The transport stream demultiplexer 121 can be configured to reconstruct each elementary stream 114 using a chain of three buffers 140, 142, and 144. In each chain, a first buffer 140 "TB" (i.e., a transport buffer) can be configured to store a transport packet if its PID matches a value found in the PMT for a certain ES. A second buffer 142 "MB" (i.e., a multiplexing buffer) that exists only for a video stream can be configured to aggregate the payload of a sequence of TS packets by stripping off the TS packet headers, thereby establishing a PES packet. A third buffer 144 "SB" (substream buffer) can be configured to aggregate ES data. In an embodiment, the data in the third buffer 144 can belong to a certain substream 114, and the data of multiple substream buffers can be aggregated to form an encoded representation of a subregion. Which group of ES needs to be aggregated in which order is signaled by information transmitted in the descriptor specified below, which is found in the PMT of that program.
[0128] In detail, the descriptors specified below are the set of descriptors specified in the extended MPEG-2 Systems standard. The type and length of the descriptors are provided by header bytes, which are not shown in the following syntax table.
[0129] Sub-stream signaling is described subsequently.
[0130] For each ES containing a substream (i.e., a block or a fixed set of sequence blocks), the newly defined substream descriptor assigns a SubstreamID to the substream. It optionally contains additional SubstreamIDs needed to form subregions or an index indicating the pattern of these additional SubstreamIDs through an offset array found in the subregion descriptor.
[0131] The following syntax (Syntax No. 1) can be used for substream descriptors:
[0132]
[0133] The substream descriptor can be used in three different versions, each signaling a SubstreamID:
[0134] According to the first version, if its size is only one byte (excluding the leading header byte), it represents a value of 0 for PatternReference (refer to SubstreamOffset[k][0][i] in the sub-region descriptor described below).
[0135] According to the second version, if ReferenceFlag is set to "1", it specifies the index (except index 0) of the pattern used to calculate the additional SubstreamID.
[0136] According to the third version, if ReferenceFlag is set to "0", it directly specifies the additional SubstreamID.
[0137] The value SubstreamCountMinus1 can be found in the subregion descriptor.
[0138] Figure 5 An illustrative view of a video image 110 of a video stream constructed from a plurality of tiles 108 is shown, as well as an illustrative view of individual tiles and combinations of tiles that form a conforming bitstream when sent as separate elementary streams. Figure 5 As shown, a single block or a fixed set of sequences of blocks (i.e., sequences with respect to coding order) can be sent as separate elementary streams, where i is the substreamID. Figure 5 As shown, when sent as a separate basic stream, the first block can form a conforming bit stream, the entire panorama can form a conforming bit stream, the first and second blocks can form a conforming bit stream, blocks 1 to N can form a conforming bit stream, blocks 1 to 2N can form a conforming bit stream, and blocks 1 to Z=NxM can form a conforming bit stream.
[0139] Thus, N may be signaled in the sub-region descriptor in the field SubstreamIDsPerLine, where Z may be signaled in the sub-region descriptor in the field TotalSubstreamIDs.
[0140] Subsequently, sub-region signaling is described.
[0141] A newly defined sub-region descriptor may be associated with the entire program. The sub-region descriptor may signal the pattern of SubstreamIDs belonging to the sub-region 109. It may signal different layouts, consisting of, for example, different numbers of sub-streams 114, and indicate the level of each pattern. The value of LevelFullPanorama may indicate the level of the full panorama.
[0142] The following syntax can be used for subregion descriptors:
[0143]
[0144] This syntax can be extended in the following way by the flag SubstreamMarkingFlag which signals one of two options for substream marking:
[0145] a) Each substream is associated with an individual elementary stream, and the substream is identified by mapping via SubstreamDescriptors in the PMT, as already discussed above;
[0146] b) Multiple substreams are transported in a common elementary stream and the substreams are identified by the af_substream_descript found in the adaptation field of the transport packet carrying the start of the PES packet. This alternative is discussed in more detail below.
[0147] The following syntax can then be used for subregion descriptors:
[0148]
[0149] Thus, N1 may be the number of different sub-region layouts indexed by l that may be selected from the entire panorama. Its value may be implicitly given by the descriptor size. PictureSizeHor[l] and PictureSizeVert[l] may indicate the horizontal and vertical sub-region dimensions measured in pixels.
[0150] Figure 6An illustrative view of a video image 110 of a video stream constructed with a plurality of tiles 108 is shown, as well as an illustrative view of a sub-region 109 of the video image 110 defined by a suitable subset of the plurality of tiles 108 and an offset pattern of tiles of the sub-region. Figure 6 The video image 110 shown in is structured into N=6 columns and M=5 rows (or lines), ie into N×M=30 blocks, wherein 9 blocks can be gathered from the 30 blocks and form a sub-region 109 of the video image 110 .
[0151] for Figure 6 Taking the example shown in , the offset pattern of the sub-regions of the 3x3 sub-stream can be indicated by the following array:
[0152] SubstreamOffset[0]:1
[0153] SubstreamOffset[1]:2
[0154] SubstreamOffset[2]:N
[0155] SubstreamOffset[3]:N+1
[0156] SubstreamOffset[4]:N+2
[0157] SubstreamOffset[5]:2N
[0158] SubstreamOffset[6]:2N+1
[0159] SubstreamOffset[7]:2N+2
[0160] Similarly, the offset pattern of the sub-regions of a 2x2 sub-stream may be indicated by the following array:
[0161] SubstreamOffset[0]:1
[0162] SubstreamOffset[1]:N
[0163] SubstreamOffset[2]:N+1
[0164] Subregion assembly is described subsequently.
[0165] The process or method of accessing a sub-region may include a first step of selecting a suitable sub-region size from a sub-region descriptor at the receiver side based on a level indication or a sub-region dimension. This selection implicitly generates a value l.
[0166] Furthermore, the process or method of accessing a sub-region may include a step of selecting an ES containing an upper left sub-stream (reading all sub-stream descriptors) of the region to be displayed based on the SubstreamID.
[0167] Furthermore, the process or method of accessing a sub-region may include the step of checking whether the applicable substream descriptor provides a PatternReference. Using this PatternReference, it selects the applicable SubstreamOffset value:
[0168] SubstreamOffset[k][PatternReference][l]
[0169] Among them, 0 <k<SubstreamCountMinus1[l]
[0170] Furthermore, the process or method of accessing a sub-region may include the step of defaulting the reference to the index to 0 if no PatternReference is present, meaning that the descriptor size is equal to 1.
[0171] There may be an ES that is not suitable for forming the top left substream of a subregion, for example, because it is located at the right edge or bottom edge of the panorama. This can be signaled by a PatternReference value greater than PatternCount[l]-1, which means that no SubstreamOffset value is assigned.
[0172] Furthermore, the process or method of accessing a sub-area may include the following steps: if the condition is met, for each PES packet of ESx having a substream descriptor indicating that SubstreamID is equal to SubstreamIDx, performing the following operations:
[0173] - If PreambleCount[l]>0: then the PES packets of the DTS with the same ES are preceded by a substream descriptor indicating that SubstreamID is equal to PreambleSubstreamID[j][l]. The order of the PES packets in the assembled bitstream is given by increasing the value of the index j.
[0174] - if SubstreamCountMinus1[l]>0: then before decoding, the PES packets with the DTS of the same ES are preceded by the substream descriptor indicating SubstreamID equal to AdditionalSubstreamID[l] given in the substream descriptor of ESx resp.SubstreamIDx+SubstreamOffset[k][j] (the SubstreamOffset array is found in the subregion descriptor, where j is given by the value of PatternReference in the substream descriptor of ESx and k ranges from 0 to SubstreamCountMinus1[l]). The order of the PES packets in the assembled bitstream is given by increasing values of SubstreamID, which also corresponds to increasing values of index k.
[0175] Figure 7 A flow chart of a method 200 for stream multiplexing according to an embodiment is shown. The method 200 comprises a step 202 of receiving coded data of each of at least two different spatial segments or different groups of sequential spatial segments of a video image of a video stream. Furthermore, the method 200 comprises a step 204 of packing the coded data of each of the at least two different spatial segments into separate sub-streams.
[0176] Figure 8 A flow chart of a method 220 for stream demultiplexing is shown. The method 220 comprises a step 222 of selectively extracting at least two individual substreams from the group of individual substreams, the at least two individual substreams comprising coded data of different spatial segments or different groups of sequential spatial segments of a video image of the coded video stream. Furthermore, the method comprises a step 224 of combining the at least two individual substreams into a data stream comprising coded data of different spatial segments or different groups of sequential spatial segments of a video image of the coded video stream.
[0177] To briefly summarize the above, a stream demultiplexer 121 has been described, which comprises a data stream former 122, which is configured to selectively extract at least two individual substreams from the group of individual substreams 116, the at least two individual substreams 114 containing coded data of different spatial segments 108 or different groups of sequential spatial segments 108 of coded images 110. The coded data originate from or are of the video stream 102. They are obtained therefrom by stream multiplexing in the stream multiplexer 100. The data stream former 122 is configured to combine the at least two individual substreams 114 into a data stream 126, which contains coded data of different spatial segments 108 or different groups of sequential spatial segments 108 of the video images 110 of the video stream 102 of the coded extracted at least two individual substreams 114. The data stream 126 is provided at an output interface 123 of the stream demultiplexer 121.
[0178] As described above, the individual substream group 116 consists of a broadcast transport stream including TS packets. The individual substream group 116 includes a plurality of individual substreams 114, which contain coded data of different spatial segments 108 or different groups of sequential spatial segments 108 of the coded video stream 102. In the above example, each individual substream is associated with an image block. The program map table also consists of the individual substream group 116. The stream demultiplexer 121 can be configured to derive a stream identifier from the program map table for each of the plurality of individual substreams 114 and to distinguish each of the plurality of individual substreams 114 in the broadcast transport stream using the corresponding stream identifier referred to as SubstreamID in the above example. For example, the stream demultiplexer 121 derives a predetermined packet identifier from a program association table transmitted in a packet of packet identifier zero in the broadcast transport stream, and derives the program map table from a packet of the broadcast transport stream having the predetermined packet identifier. That is, the PMT can be transmitted in a TS packet whose packet ID is equal to the predetermined packet ID indicated in the PAT of the program of interest (i.e., the panoramic content). According to the above-described embodiment, each substream 104 or even each substream 106 is contained in a separate elementary stream, i.e., within TS packets of mutually different packet IDs. In this case, the program map table uniquely associates, for example, each stream identifier with a corresponding packet identifier, and the stream demultiplexer 121 is configured to depacketize each of the plurality of individual substreams 104 from a packet of the broadcast transport stream having a packet identifier associated with the stream identifier of the corresponding individual substream. In this case, the substream identifier and the packet identifier of the substream 104 are quasi-synonymous, as long as there is a bijective mapping therebetween in the PMT. In the following, an alternative is described in which the substream 104 is multiplexed into an elementary stream using the concept of marking the NAL unit of the substream 104 via an adaptation field in a packet header of a TS packet of one elementary stream. In particular, the packet header of a TS packet of an elementary stream falling into its payload portion at the beginning of any PES packet containing one or more NAL units of the substream 104 is provided with an adaptation field, which in turn is provided with the substream ID of the substream to which the corresponding one or more NAL units belong. Later, it will be shown that the adaptation field also includes information related to the substream descriptor. According to this alternative, the stream demultiplexer 121 is configured to unpack a sequence of NAL units from packets of a broadcast transport stream having a packet identifier indicated in a program map table, and to associate each NAL unit of the sequence of NAL units with one of a plurality of separate substreams depending on the substream ID indicated in the adaptation field of the packets of the broadcast transport stream having the packet identifier indicated in the program map table.
[0179] Furthermore, as described above, the stream demultiplexer 121 may be configured to read from a program map table information about the spatial subdivision of the video and video images 110 into segments 108, respectively, and to derive stream identifiers for the plurality of individual substreams 114 inherently from the spatial subdivision by using a mapping from the spatially subdivided segments to stream identifiers. Figure 5 An example is shown in which the stream demultiplexer 121 is configured to read information about the spatial subdivision of the video into segments from a program map table and derive stream identifiers of the plurality of individual substreams 114 such that the stream identifiers of the plurality of individual substreams 114 order the segments within the spatial subdivision along a raster scan. Figure 5 While a row-by-row raster scan is shown, any other favorable assignment of segments may be used, such as a column-by-column raster scan for association between segments, and the stream identifiers and / or packet IDs may be implemented in other ways.
[0180] The stream demultiplexer 121 may be configured to read from a program map table or, according to the just mentioned alternative described further below, from an adaptation field of a data packet carrying the group of individual substreams 116, examples of which are listed above and given below, the substream descriptors. Each substream descriptor may index one of the plurality of individual substreams 104 by a substream ID associated with one of the individual substreams and comprises information about which one or more of the plurality of individual substreams 104 together with the indexed individual substream form an encoded representation of a subregion 109 extractable from the group of individual substreams 116 as at least two individual substreams 114, the subregion 109 consisting of a spatial segment 108 or a sequence of spatial segments 108 of the one or more individual substreams that together with the indexed individual substream form an encoded representation 126. The stream demultiplexer 121 may also be configured to read from the program map table information about subregion descriptors indicating one or more spatial subdivisions of the video into subregions 109 indexed in the above example using index 1. For each such subdivision, each sub-region 109 is a set of spatial segments 108 or a sequential set of spatial segments 108 of one or more individual sub-streams 104. For each spatial subdivision of the video into a sub-region 109, the sub-region descriptor may indicate the size of the sub-region 109, such as using parameters PictureSizeHor[l] and PictureSizeVert[l]. In addition, the encoding level may be signaled. Thus, at least two individual sub-streams selectively extracted from the group of individual sub-streams 116 may together contain coded data encoding different spatial segments 108 or different groups of sequential spatial segments 108 of one of the sub-regions of one of the one or more spatial subdivisions of the video.
[0181] One or more of the substream descriptors (in the example above, those with ReferenceFlag=1) may contain information about which one or more of the plurality of individual substreams are to be extracted together with the indexed individual substream as at least two individual substreams 114 from the group of individual substreams 116, in the form of a reference index, such as a PatternReference in a list of sets of stream identifier offsets (such as a list SubstreamOffset[...][j][l] pointed to therein by using PatternReference as j=PatternReference), and there is one such list indexed by l for each subdivision. Each stream identifier offset SubstreamOffset[k][j][l] indicates an offset relative to the stream identifier SubstreamID of the indexed individual substream, i.e. the referenced substream has a substream ID equal to the SubstreamID of the substream to which the substream descriptor belongs plus SubstreamOffset[k][j][l].
[0182] Alternatively or additionally, one or more of the substream descriptors (in the example above, those with ReferenceFlag=0) may contain information about which one or more of the multiple individual substreams are to be extracted together with the indexed individual substream as at least two individual substreams 114 from the individual substream group 116, in the form of a set of stream identifier offsets, e.g. AdditionalSubstreamID[i], each indicating an offset relative to the stream identifier of the indexed individual substream, i.e., by an offset explicitly signaled in the substream descriptor.
[0183] One or more substreams within the group 106 may include a slice header and / or parameter set stripped from any of the plurality of individual substreams 114 or dedicated to modifying or replacing a slice header and / or parameter set of any of the plurality of individual substreams 114. Whether or not the slice header and / or parameter set of the substream 104 is included in the additional substream 106, modifying or replacing the slice header and / or parameter set to achieve a standard-compliant data stream 126′ for decoding by the decoder 124 may be performed by the data stream processor 127 of the stream demultiplexer 121.
[0184] The just mentioned alternative to spending a separate ES for each chunk is now described in more detail. The described alternative may be advantageous in the case of existing implementations that rely on a demultiplexer structure at the receiver, which pre-allocates buffers for all ESs that are potentially decoded. Such implementations would therefore overestimate the buffer requirements, thereby compromising some of the benefits of the solution embodiments presented above. In this case, it may be beneficial to send multiple chunks within the same ES and assign substream identifiers to the data portions within that ES so that the advanced demultiplexer 121 can remove the unneeded data portions from the ES before the ES is stored in the elementary stream buffer. In this case, the TS Demux 121 still uses a buffer such as Figure 4 The three buffer chains shown reconstruct each elementary stream. In each chain, the first buffer "TB" (i.e., transport buffer) stores a transport (TS) packet if its PID matches the value found in the PMT of a certain ES. In addition to evaluating the PID included in each TS packet header, it also parses the adaptation field, which optionally follows the TS packet header.
[0185] The syntax of the adaptation field from ISO / IEC 13818-1 is as follows:
[0186]
[0187]
[0188] The new syntax for carrying a substream descriptor in the adaptation field might look like this:
[0189]
[0190] The new flag identifies the af_substream_descriptor that carries the substream descriptor. Within the adaptation field, a Substream_descriptor according to syntax No.1 is sent whenever the TS packet payload contains the start of a PES packet. The multiplexing buffer "MB" aggregates the payloads of sequence TS packets with the same PID by stripping the TS packet headers and the adaptation field, thereby building up a PES packet. If the Substream_descriptor indicates a SubstreamID that is not required for decoding a subregion, the demultiplexer 121 discards the entire PES packet, and the PES packet with the SubstreamID matching the subregion is stored in the substream buffer "SB".
[0191] In addition to the substream identification using Substream_descriptor, as described above, the subregion descriptor is sent in the PMT associated with the program. The optional information in Substream_descriptor is used according to the example above:
[0192] A pattern may be signaled that indicates a set of offset values that are added to the value of the SubstreamID present in the descriptor, resulting in additional Substream IDs that complement the desired sub-region;
[0193] - Alternatively, an array of additional SubstreamIDs may be indicated directly or explicitly in a for loop extending to the end of the descriptor, the length of the descriptor being indicated by the length value af_descr_length included in the adaptation field preceding the Substream_descriptor.
[0194] The entire extracted bitstream 126 may be forwarded to a data stream processor 127 which removes unwanted data and further processes the bitstream 126 before forwarding the output bitstream 126' to the decoding stage 124 in order to store a standard-compliant bitstream for the sub-region in the decoder's coded picture buffer. In this case, some fields of the Subregion_descriptor syntax may be omitted, or the following simplified syntax may be used to indicate the desired level of decoding the sub-region and the resulting horizontal and vertical sub-region dimensions:
[0195]
[0196] N1 can be the number of different sub-region layouts indexed by l that can be selected from the entire panorama. Its value can be implicitly given by the descriptor size. If this descriptor is present in the PMT, it is mandatory to have the MCTS extraction information set SEI message in all random access points. The value of SubstreamMarkingFlag (whose presence in the above syntax example is optional) or the presence of Subregion_level_descriptor indicates that af_substream_descriptor is used to identify the substream. In this case, the client can adjust the buffer size of the SB to the CPB buffer size indicated by Level[l].
[0197] The following description of the present application relates to embodiments for tasks of stream multiplexing, stream demultiplexing, image and / or video encoding and decoding and corresponding data streams, which tasks do not necessarily involve providing the receiver side with an opportunity to select or change a sub-area within an image area of the video regarding which stream extraction is performed from a broadcast transport stream. However, the embodiments described below relate to an aspect of the present application that can be combined or used in conjunction with the above-described embodiments. Therefore, at the end of the description of the embodiments described subsequently, there is an overview of how the embodiments described subsequently can be advantageously used to implement the above-described embodiments.
[0198] In particular, the embodiments described below attempt to address Fig. 9 The issues outlined in . Fig. 9 An image 300 is shown. The image 300 is encoded into a data stream 302. However, the encoding has been done in such a way that different segments 304 of the image are encoded into the data stream 302 individually or independently of each other, i.e. wherein inter-coding dependencies between different segments are suppressed.
[0199] Fig. 9 Also illustrated is the possibility that the segments 304 of the image are encoded sequentially into the data stream 302, i.e. such that each segment 304 is encoded into a respective consecutive fragment 306 of the data stream 302. For example, the image 300 is encoded into the data stream 302 using a certain coding order 308 which traverses the image 300 in such a way that the segments 304 are traversed one after the other, thereby completely traversing one segment before continuing with the next segment in the coding order 308.
[0200] exist Fig. 9 , only two segments 3041 and 3042 and the corresponding two fragments 3061 and 3062 are shown, but the number of segments can naturally be higher than two.
[0201] As previously described, the image 300 may be an image of a video and in the same manner just described with respect to the image 300, further images of the video 310 may be encoded into the data stream 302, wherein these further images are subdivided into segments 304 in the same manner as described with respect to the image 300, and suppression of inter-segment coding dependencies may also follow with respect to coding dependencies between different images, so that each segment, such as the segment 3042 of the image 300, may be encoded into its corresponding fragment 3062 in a manner that depends on a corresponding, i.e. concatenated, segment of another image, rather than on another segment of another image. The fragments 306 belonging to one image will form a continuous portion 312 of the data stream 302, which may be referred to as an access unit, and they are not interleaved with parts of fragments belonging to other images of the video 310.
[0202] Fig.10The encoding of an image 300 into a data stream 302 is illustrated in more detail. In particular, Fig.10 The corresponding part of the data stream 302 into which the image 300 is encoded is shown, namely an access unit 312 consisting of a concatenation of fragments 3061 and 3062 . Fig.10 The diagram shows the data flow from Fig.10 The NAL unit is composed of a series of NAL units represented by rectangle 314 in FIG. There may be different types of NAL units. For example, Fig.10 The NAL units 314 in the two slices have been encoded into the corresponding slices in the segment 304 corresponding to the fragment 306 of which the corresponding NAL unit 314 is a part. The number of such slice NAL units per segment 304 can be any number including one or greater than one. The sub-divisions in the two slices follow the coding order 308 described above. In addition, Fig.10 A NAL unit is illustrated by a shaded box 314 that contains a parameter set 316 in its payload portion instead of having slice data of a corresponding slice of a corresponding segment 304 encoded therein. Parameter sets such as parameter set 316 are used to mitigate or reduce the amount of transmission of parameters that parameterize the encoding of the image 300 into a data stream. In particular, a slice NAL unit 314 may reference or even point to such a parameter set 316, such as Fig.10 , so as to adopt the parameters in such parameter set 316. That is, for example, there may be one or more pointers in the slice NAL unit pointing to the parameter set 316 contained in the parameter set NAL unit.
[0203] like Fig.10 As shown, parameter set NAL units may be redundantly or repeatedly interspersed between slice NAL units, for example, to increase the resistivity of the data stream against NAL unit loss, etc., but naturally, the parameter set 316 referenced by the slice NAL units of fragments 3061 and 3062 should consist of at least one parameter set NAL unit that precedes the referenced slice NAL unit, and thus, Fig.10 Such a parameter set NAL unit is depicted at the outermost left side of the NAL unit sequence of access unit 312 .
[0204] It should be mentioned here that the previous and subsequent Fig. 9 and 10 Many of the details described are only for the purpose of illustrating the problems that may occur when trying to feed only a portion of the data stream 302 to a decoder in order for the decoder to recover a certain segment of the video 310 or image 300. Although these details are helpful in understanding the basic ideas of the embodiments described herein, it should be clear that these details are only illustrative, as the embodiments described below should not be limited by these details.
[0205] In any case, one example of a situation that may prevent the successful feeding of only one fragment 306 to the decoder in order for the decoder to successfully decode the segment 304 corresponding to the fragment 306 may be a parameter within the parameter set 316 related to the size of the image 300. For example, such a parameter may explicitly indicate the size of the image 300, and since the size of each segment 304 is smaller than the size of the image 300 as a whole, a decoder that receives only fragments of a certain image 300 in which a plurality of fragments 306 are encoded without any modification will be corrupted by such a parameter within the parameter set 316. It should be mentioned that a certain fragment 304 that has just been cut out of the original data stream 302 without any modification may even lack any parameter set NAL unit and thus lacks an essential component of a valid data stream, i.e. a parameter set, at all. In particular, as a minimum, the data stream 302 only needs a parameter set NAL unit at the beginning of an access unit, or more precisely, before a slice NAL unit that refers to it. Therefore, while the first fragment 3061 needs to have a parameter set NAL unit, this is not the case for all subsequent fragments 3062 in the coding order.
[0206] There may additionally or alternatively be other examples of situations that may prevent the successful excision of the fragment 306 from the data stream 302 and result in the successful decoding of the corresponding segment of the fragment. Only one such additional example is now discussed. In particular, each slice NAL unit may include a slice address, which indicates its location within the image area, where the data stream to which the corresponding slice NAL unit belongs is located. This parameter is also discussed in the following description.
[0207] If we consider the above situation, then Fig. 9 The data stream 302 can be transformed into a data stream associated with one of the segments 304 only in the following cases: 1) a segment 306 corresponding to the segment is cut off from the data stream 302, 2) at least one of the image-sensitive parameters just mentioned in the parameter set 316 carried in the parameter set NAL unit is modified, or even such an adaptation parameter set NAL unit is newly inserted, and 3) if necessary, the slice address stored in the slice NAL unit of the cut-off segment is modified, wherein the modification of the slice address is necessary whenever the segment 304 of interest is not the first segment in the coding order among all segments 304, for example, the slice addresses are measured along the above-mentioned coding order 308, so that for it, the slice address associated with the entire image 300 and the slice address associated only with the segment of the corresponding slice are consistent.
[0208] therefore, Fig.11a Illustration of extraction and modification Fig.103061 of the entire data stream 312, i.e. the access units affecting only the parameter set NAL units related to the picture 300 in which the segment coded into the segment 3061 is located. Fig.11b The result of the corresponding extraction and modification is illustrated with respect to segment 3042. Here, the self-contained segmented data stream 320 includes a segment 3062 of the data stream 302 as an access unit related to the picture 300, but additional modifications have been made. For example, in addition to modifying at least one parameter of the parameter set 316, the parameter set NAL units within the segment 3062 of the data stream 302 have been shifted towards the beginning of the segment 3062 with respect to the self-contained segmented data stream 320, and the slice NAL units have been modified with respect to the slice address. In contrast, Fig.11a In this case, with respect to the first segment 3041 of the coding order 308, the slice NAL unit remains unchanged because the slice address is the same.
[0209] About Fig.11a and 11b Modifying the segments 306 of the data stream 302 in the manner shown results in a loss of information. In particular, the modified parameters and / or slice addresses are no longer visible or known to the receiver of the self-contained data stream 320. Therefore, it would be advantageous if this information loss could be avoided without interfering with the receiving decoder's successful decoding of the corresponding self-contained data stream 320.
[0210] Fig.12 The stream formatter according to an embodiment of the present application is shown, and the stream formatter can achieve two tasks: first, Fig.12 The stream formatter 322 of FIG. 304 is capable of demultiplexing the inbound data stream 302 into a plurality of partial data streams 324, each of which is associated with a corresponding segment 304. That is, the number of partial data streams 324 corresponds to the number of segments 304, and corresponding indices are used for these reference numbers, as is also the case with respect to the segments 306. For example, the partial data streams 324 may be identified according to the above description of Figures 1 to 8 One of the concepts discussed, namely, for example, by using one ES for each partial data stream 324 for broadcasting.
[0211] The stream formatter 322 distributes the segments 306 within the data stream 302 onto corresponding partial data streams 324. However, secondly, the stream formatter 322 converts at least one of the partial data streams 324 into a properly parameterized self-contained data stream so as to be successfully decodable by a conventional decoder with respect to the segment 304 associated with the corresponding partial data stream. Fig.123061, this is illustrated with respect to the first segment, i.e., with respect to the first partial data stream 3241. Thus, the stream formatter 322 subjects the NAL units 314 within the fragment 3061 to the above Fig.11a The modification discussed thereby obtains NAL unit 314' within data stream 3241 from NAL unit 314 within segment 3061 of data stream 302. Note that not all NAL units must be modified. The same type of modification can be made with respect to portions of data stream 3242 associated with any other segment such as segment 3042, but will not be discussed further here.
[0212] However, as an additional task, the stream formatter 322 adds NAL units 326 of a specific NAL unit type to the NAL units 314' of the partial data stream 3241. The latter newly added NAL unit 326 has a NAL unit type selected from a set of NAL unit types that is disjoint from the set of NAL unit types to which the NAL units 314 and 314' belong. In particular, while the NAL unit types of the NAL units 314 and 314' are types that are understood and processed by legacy decoders, the NAL units 326 are types that are expected to be ignored or discarded by legacy decoders because they are, for example, reserved types reserved for future use. Therefore, legacy decoders ignore the NAL units 326. However, through these NAL units 326, the stream formatter 322 is able to signal to a decoder of the second type (i.e., a non-legacy decoder) to perform certain modifications to the NAL unit 314'. In particular, as discussed further below, one of the NAL units 326, such as having Fig.12 The B NAL unit shown in , may indicate that the immediately subsequent NAL unit 314' of the parameter set NAL unit parameterized for encoding the segment 3041 as an independent picture is to be discarded or ignored along with the B NAL unit 326. Additionally or alternatively, an A NAL unit 326 newly added by the stream formatter 322 may carry in its payload portion the original parameter set NAL unit associated with the picture 300 of the data stream 302, and may signal to a non-legacy decoder that the NAL unit carried in its payload portion is to be inserted in place of itself or as a replacement for itself. For example, such an A NAL unit 326 may follow or be followed by a modified parameter set NAL unit 314' at the beginning of a fragment 3061 within the partial data stream 3241 so as to "update" the parameter set 316 before the slice NAL unit within the fragment 3061 begins, so that a non-legacy decoder only preliminarily uses the NAL unit 314 to set the parameters in the parameter set 316 before the A NAL unit 326, and then correctly sets the parameters based on the correct parameter set NAL unit stored in the payload portion of the A NAL unit 326.
[0213] In summary, a legacy decoder receiving the partial data stream 3241 receives a self-contained data stream into which the segment 3041 is encoded as a self-contained picture, and the legacy decoder is not disturbed by the newly added NAL unit 326. As shown in the apparatus 340 for processing an inbound data stream Fig.13 A more complex reception of the two partial data streams shown in 300 enables reconstruction of the entire video or image in one step, or by reconstruction of the self-contained entire data stream 302', which again forms a coding equivalent data stream 302 from the partial data stream 324 and feeds the partial data stream 324 to a conventional decoder capable of handling the image size of the image 300. The minimum task may be to convert the self-contained partial data stream 3241 into the inverse parameterized partial data stream 334 by following the instructions indicated by the inserted NAL units 326. Furthermore, only a subset of the original partial data stream 324 may be output by the device 340, which subset relates to the extracted portion in the original image 300. The device may or may not perform any slice address inverse adaptation of the slice addresses of the NAL units 314' within the partial data stream 3241, so as to relate again to the overall image 300.
[0214] It should be noted that the ability of the stream formatter 322 to extract the partial data stream 3241 from the data stream 302 and present it as a self-contained data stream decodable by a conventional decoder, but still carrying the original unmodified parameter data, may be useful without deriving (one or more) other partial data streams 3242. Therefore, it would be the task of the data stream generator to generate only the partial data stream 3241 from the data stream 302, which also forms an embodiment of the present application. Again, all statements are also valid in the other direction, i.e. if only the partial data stream 3242 is generated, all statements are also valid.
[0215] In the above-described embodiment, the stream multiplexer 103 may be configured to communicate with one or more sub-streams 114. Fig.12 The stream formatter 322 operates similarly. Fig.123242. In the embodiment of the present invention, the stream formatter 322 may alternatively or additionally perform the modification plus addition of the NAL unit 326 with respect to any other partial data stream, such as the partial data stream 3242. Again, the number of partial data streams may be greater than 2. On the receiver side, any stream former, such as a non-legacy decoder or a stream demultiplexer, such as the stream demultiplexer 121, which receives the partial data stream 3241 thus processed, may perform, among other tasks, the task of following the instructions defined by the NAL unit 326, thereby producing a partial data stream that is correct at least with respect to the parameter set 316 when concatenated with other fragments of the original data stream 302. For example, if the concatenation of fragments 3241 and 3242 has to be decoded as a whole, then only the slice addresses within the slice NAL units may have to be modified.
[0216] It goes without saying that the partial data streams 324 may be transmitted, for example, within separate elementary streams. That is, they may be packetized into transport stream packets, i.e., packets of one packet ID for one of the partial data streams 324, and packets of another packet ID for another partial data stream 324. Other ways of multiplexing the partial data streams 324 within a transport stream have been discussed above and may be reused with respect to transmission of partial data streams in accordance with conventional procedures.
[0217] In the following, a NAL unit 326 of type B is presented exemplarily and is referred to as jump over NAL units, and type A NAL units 326 use what is called hide The syntax of the NAL unit is illustrated.
[0218] therefore, Figures 9 to 12 An exemplary case involving at least two ESs for transmission of a complete panorama is involved. One of the ESs is a substream such as 3241, which represents the coded video of the subregion 3041, and the substream itself is a standard-compliant stream and can be decoded by a legacy decoder if the legacy decoder is capable of decoding the level of the substream.
[0219] The combination with at least one other substream 3242 transmitted in a different ES is sent to a data stream processor such as 127 for further processing, which may result in the extraction of conforming bitstreams for different sub-regions. In this case, some data portions at the input of the data stream processor 127 necessary for its normal function, such as the parameter set 316 providing information for the entire panorama, prevent a legacy decoder from decoding a sub-region intended to be decoded by such a legacy device. To solve this problem, these data portions are made invisible to legacy decoders, while advanced devices can process them.
[0220] The decoder processes the video bitstream as a sequence of data units, which in case of HEVC encoding are represented by so-called "NAL units". The size of the NAL unit is implicitly indicated by a start code indicating the start of each NAL unit. After the start code, each NAL unit starts with a header containing information about the NAL unit type. If the decoder does not recognize the type indicated in the NAL unit header, it ignores the NAL unit. Some NAL unit type values are reserved, and NAL units indicating reserved types will be ignored by all conforming decoders.
[0221] To make the NAL units invisible to legacy decoders, the header with the original NAL unit type field is prepended with a header with a reserved value. The advanced processor recognizes this type value and implements different processing, i.e., the insertion is restored and the original NAL unit is processed. The two reserved NAL unit type values are used to form two different pseudo-NAL units. If the advanced processor encounters the first type, the insertion is restored. If the second type is encountered, the pseudo-NAL unit and the NAL unit that follows it are removed from the bitstream.
[0222] NAL unit syntax in ISO / IEC 23008-2:
[0223]
[0224] New syntax for hiding NAL unit headers:
[0225]
[0226] hiding_nal_unit_type is set to 50, which is the NAL unit type marked as "unspecified" in ISO / IEC 23008-2.
[0227] nuh_layer_id and nuh_temporal_id_plus1 are copied from the original NAL unit header.
[0228] New syntax for skipping NAL units:
[0229]
[0230] skipping_nal_unit_type is set to 51, which is the NAL unit type marked as "unspecified" in ISO / IEC 23008-2.
[0231] nuh_layer_id and nuh_temporal_id_plus1 are copied from the original NAL unit header.
[0232] exist Fig.14 In the concealment process shown by the transition from top to bottom, a single NAL unit is formed by the concealed NAL unit header together with the original NAL unit. There is only one start code before the combined NAL unit.
[0233] Fig.14 The un-hiding process shown by the transition from bottom to top in FIG. 1 removes only the hidden NAL unit header, which is located between the start code and the original NAL unit.
[0234] The process of allowing advanced decoders to skip the original NAL unit requires that the skip NAL unit be inserted into the bitstream before the original NAL unit, which means that there are two NAL units, each prepended with a start code. Fig.15 The skip is illustrated by a transition from top to bottom in Fig.15 The preparation (insert) is illustrated in FIG. 1 by the transition from bottom to top.
[0235] According to the advantageous alternative scheme now described, it is allowed to signal by the inserted NAL unit 326 that the NAL unit is not skipped as a whole, but only a part of it is skipped. In principle, there are two options: or the following NAL unit is skipped in its entirety. This instruction is only signaled by the skip NAL unit 326 according to the above example. According to the alternative scheme, only a part of any subsequent NAL unit is skipped. The latter is useful in the case where multiple elementary streams are combined and the skipped part of the NAL unit is replaced by information in the elementary stream that contains the partially skipped NAL unit placed in front of the stream. This will be explained in more detail below. In this case, the above grammar example can be adapted in the following way:
[0236]
[0237]
[0238] Here, bytes_to_skip indicates the number of bytes to skip of the following NAL unit. If this value is set to zero, the entire following NAL unit is skipped.
[0239] about Fig.16 This alternative is described in more detail. Fig.16 A stream formatter 322 and a scenario in which the stream formatter 322 receives a data stream 302 having an image 300 or video encoded therein, wherein the image 300 is segmented into spatial segments 3041 and 3042 as described above, i.e. in a manner that enables extraction of parts of the data stream with respect to the spatial segments. As with all the above embodiments, the number of segments is not critical and may even be increased.
[0240] In order to facilitate understanding of Fig.16 The variant discussed presupposes that the encoding of the image 300 into the data stream 302 is performed in the following manner, wherein for each segment 3041 and 3042 only one slice NAL unit is spanned, and the access unit associated with the image 300 and consisting of the fragments 3061 and 3062 contains only one parameter set NAL unit 314 at the beginning. The three generated NAL units in the data stream 300 of the image 300 are numbered #1, #2 and #3, wherein NAL unit #1 leads and is a parameter set NAL unit. Therefore, it is distinguished from NAL units #2 and #3 by shading. NAL units #2 and #3 are slice NAL units. NAL unit #2 is NAL unit 314 within the fragment 3061 associated with the segment 3041, and NAL unit #3 is NAL unit 314 within the fragment 3062 associated with the spatial segment 3042. It is the only NAL unit within the fragment 3062.
[0241] The stream formatter 322 seeks to split the data stream 302 into two partial data streams 3241 and 3242, the former being associated with the spatial segment 3041 and the latter being associated with the spatial segment 3042, wherein the partial data stream 3242 appears as a self-contained data stream which can be decoded by a conventional decoder as a self-contained picture region with respect to the spatial segment 3042, whereas the partial data stream 3241 is not, as in Fig.12 As in the discussion of Fig.12As can be clearly seen in the discussion of FIG. 3 , stream formatter 322 handles two problems associated with converting fragment 3062 into a self-contained data stream; namely, 1) fragment 3062 does not yet have any parameter set NAL units, and 2) fragment 3062 contains a slice NAL unit; namely, NAL unit 314#3, whose slice header 400 contains an erroneous slice address, i.e., it contains a slice address that addresses the start of segment 3042 relative to the entire picture area of picture 300, such as, for example, with respect to the upper left corner of picture 300, rather than a slice address relative to spatial segment 3042 representing a self-contained picture area of partial data stream 3242 for low-budget legacy decoders. Therefore, stream formatter 322 adds a suitably designed parameter set NAL unit 314′ to fragment 3062 when transmitting fragment 3062 onto partial data stream 3242. The stream formatter 322 places this parameter set NAL unit 314' in front of the slice NAL unit 314' of the fragment 3062; i.e., in front of the slice NAL unit #3, which differs from the corresponding slice NAL unit 314#3 in the data stream 302 with respect to its slice header. To illustrate the difference in the slice header, the slice header of the slice NAL unit #3 is shown unhatched and indicated using the reference numeral 400 with respect to its version contained in the data stream 302, and is indicated using cross-hatching and using the reference numeral 402 with respect to its modified version in the partial data stream 3242. In addition, however, the stream formatter 322 precedes the parameter set NAL unit 314' with a skip NAL unit 326 of type B so that the parameter set NAL unit 314' will be skipped by a processing device or data stream former 340 fed with the partial data stream 3042. However, in addition to this, the stream formatter 322 places the slice NAL unit 326 of type C before the slice NAL unit 314'#3. That is, this partially skipped NAL unit 326 of type C is also a NAL unit type that is not allowed with respect to conventional decoders. However, the data stream former or processing device 340, which is designed to correctly interpret the NAL unit 326 of type C, interprets the instruction conveyed by this NAL unit as to the extent that the portion 402 of the subsequent slice NAL unit 314' is to be cut off or ignored. The composition of this "lost" header will be clear from the following description. However, first, Fig.16 3242. In particular, the low-budget legacy decoder 500 will ignore or discard NAL units 326 of type B and C since these NAL units are not allowed NAL unit types, so that only NAL unit 314' remains, as described above with respect to Fig.11b As discussed, this results in a correct and self-contained decoding of an image whose shape and content are consistent with segment 3042.
[0242] However, in addition, the stream formatter 322 adds the previous version 400 of the slice header of the slice NAL unit #3 at the end of the stream fragment 3061, which has already been distributed by the stream formatter 322 to the partial data stream 3241, so as to be adopted therein without any modification. This in turn means that the partial data stream 3241 has in it an "incomplete fragment" as an access unit associated with the timestamp of the picture 300; i.e., an exact copy of the NAL unit 314 within the corresponding fragment 3061, followed by an incomplete leading part of the only NAL unit of the subsequent fragment 3062; i.e., its previous slice header 400. The significance of this is as follows.
[0243] In particular, when the data stream former 340 receives both partial data streams 3241 and 3242, the data stream former 340 performs the action indicated by the special NAL unit 326 on the self-contained data stream 3242, and then concatenates the partial data streams 3241 and 3242'. However, the result of this concatenation plus instruction execution is a data stream 302', in which the NAL unit associated with the moment of the image 300 is an exact copy of the three NAL units 314 of the original data stream 302. Therefore, the high-end legacy decoder 502 receiving the data stream 302' will decode the entire image 300 including both the spatial segments 3041 and 3042 therein.
[0244] about Fig.16 The code modification instructions caused by NAL unit type C 326 are discussed with respect to Fig.17 In particular, it is shown here that a NAL unit 326 of type C may indicate a portion to be cut from a subsequent NAL unit 314'; that is, a portion 510 may be indicated using a length indication indicating the portion to be cut from the subsequent NAL unit 314' within the corresponding partial data stream 3242, such as measured from the beginning of the corresponding subsequent NAL unit 314', rather than such as for easier understanding as described above with respect to Fig.16 The portion to be cut off is semantically determined by indicating that the NAL unit header 402 is to be cut off as described. Naturally, indicating which portion is to be cut off may alternatively be implemented in any other way. For example, the length indication indicated in the NAL unit 326 of type C may indicate the portion to be cut off, i.e., the portion 510, by a length indication, the length indicated by the length indication being measured from the end of the previous NAL unit relative to the NAL unit 326 of type C, or any other indication may be used.
[0245] Therefore, about Figures 9 to 17 The description reveals a data flow that has been made self-contained, i.e., Fig.12 3241 and Fig.163242 in . It consists of a NAL unit sequence. The NAL unit sequence includes one or more NAL units 314' of a first set, which form a self-contained data stream parameterized to encode a first image, and the NAL units 314' of the first set are selected from one or more NAL unit types of the first set. In addition, the NAL unit sequence includes one or more NAL units 326 of a second set, each NAL unit 326 is one of one or more predetermined NAL unit types of the second set, which is disjoint from the first set and is determined to cause a legacy decoder to ignore the corresponding NAL unit. The one or more NAL units 326 of the second set may be interspersed in the NAL unit sequence or arranged within the NAL unit sequence such that, for each NAL unit 326 of the second set, the NAL unit sequence is converted into a converted sequence 334 of NAL units that, in addition to the slice addresses contained in one or more of the one or more NAL units of the first set, forms a fragment 3061 of a larger self-contained data stream having a plurality of spatial segments 304 encoded therein by discarding from the NAL unit sequence an immediately following NAL unit 330 or a portion 402 thereof of the first set and the corresponding NAL unit 326 of the second set and / or inserting into the NAL unit sequence a NAL unit 332 of one of the first NAL unit types carried in a payload portion of the corresponding NAL unit 326 of the second set in place of the corresponding NAL unit of the second set. 1,2 Composed image 300, multiple spatial segments 304 1,2 The predetermined one 3041 of the first set is the first picture. The fragment 3061 is modifiable for a larger self-contained data stream by concatenating the transformed sequence of NAL units 334 with one or more other NAL unit sequences 3242, wherein each spatial segment 3042 of the picture other than the predetermined spatial segment is encoded therein, and the transformed sequence of NAL units 334 is parameterized to encode the predetermined spatial segment 3041 within the picture when concatenated with the one or more other NAL unit sequences. As described with respect to Fig.16As shown, the predetermined spatial segment may be a spatial segment other than the first spatial segment in the coding order 308. The one or more NAL units 326 of the second set may include a predetermined NAL unit of the second set that precedes a first NAL unit 330 of the first set, the first NAL unit 330 of the first set including a payload portion carrying a first parameter set referenced by at least one second NAL unit of the first set, the predetermined NAL unit 326 indicating that the first NAL unit 330 or a portion 402 thereof is to be discarded from the NAL unit sequence along with the predetermined NAL unit. The one or more NAL units of the second set may include a predetermined NAL unit 326 of the second set that follows a first NAL unit of the first set, the first NAL unit of the first set including a payload portion carrying a first parameter set referenced by at least one second NAL unit of the first set, the predetermined NAL unit 326 having a payload portion carrying a third NAL unit 332 of one of the first NAL unit types, the third NAL unit including a payload portion carrying a second parameter set that replaces the first parameter set when the third NAL unit 332 is inserted into the NAL unit sequence in place of the predetermined NAL unit 326. In case a predetermined one of the one or more NAL units 326 of the second set indicates that only a portion 402 of the subsequent NAL unit 314' is to be removed, the length of the portion 402 may also be optionally indicated. The substitute, i.e. the incomplete slice fragment, may be appended by the stream formatter to another fragment 306, but alternatively a concealed NAL unit carrying the partial substitute may be used and arranged in a manner preceding the partially skipped NAL unit. The partially skipped subsequent NAL unit may be a slice NAL unit comprising a payload portion carrying slice data, the slice data comprising a slice header at least partially covered by the portion 402 to be skipped, followed by slice payload data of the corresponding slice in which the slice NAL unit is encoded. The incomplete slice fragment and the portion to be discarded comprise at least a slice address of the slice header, the slice address comprised by the incomplete slice fragment being defined relative to the perimeter of the picture, and the slice address comprised by the portion to be discarded being defined relative to the perimeter of a predetermined space segment 3042, such as in both cases relative to the upper left corner from which the coding order 308 starts.
[0246] Already in the above Fig.12 3241 from the second data stream 302. In the case where more than one partial data stream 3241 is output, it may be a stream separator. The second data stream has a plurality of spatial segments 304 encoded therein. 1,2An image 300 composed of a plurality of NAL units, wherein the second data stream consists of a sequence of NAL units, the sequence of NAL units comprising one or more NAL units 314 of a first set parameterized to encode a predetermined spatial segment 3041, the NAL units of the first set being selected from one or more NAL unit types of the first set, wherein the apparatus is configured to cut out the one or more NAL units 314 of the first set from the second data stream 302 to adopt them into the first data stream 3241; re-parameterize the one or more NAL units of the first set to encode the predetermined spatial segment as a self-contained image; and insert one or more NAL units 326 of the second set into the first data stream, each NAL unit 326 of the second set being one of the one or more predetermined NAL unit types of the second set, being disjoint from the first set, and being determined to cause a legacy decoder to ignore the corresponding NAL unit 326. The apparatus 322 may be configured to arrange one or more NAL units of the second set relative to the one or more NAL units of the first set, i.e., to intersperse the second set into the one or more NAL units of the first set, and / or to prepend the second set to and / or to append to the one or more NAL units of the first set, such that each NAL unit of the second set specifies that an immediately following NAL unit 330 of the first set and a corresponding NAL unit of the second set are discarded from the sequence of NAL units, and / or specifies that a hidden NAL unit 332 of one of the first NAL unit types carried in a payload portion of a corresponding NAL unit of the second set is inserted into the sequence of NAL units in place of the corresponding NAL unit of the second set. The immediately following NAL unit is, for example, a parameter set NAL unit, and / or the hidden NAL unit is a parameter set NAL unit as contained in the second data stream.
[0247] Already about Fig.13An apparatus 340 for processing a data stream such as 3241 is described. The apparatus 340 may be extended to a decoder if, instead of a data stream 302', image content such as an image 300 is output, or a set of one or more fragments thereof, such as, for example, a fragment 334 derived only from a portion of the data stream 3241. The apparatus 340 receives a data stream 3241 consisting of a sequence of NAL units, the NAL unit sequence comprising one or more NAL units 314' of a first set, which form a self-contained data stream parameterized to encode a first image, the NAL units of the first set being selected from one or more NAL unit types of the first set, and the NAL unit sequence comprising one or more NAL units 326 of a second set, each of which is one of one or more predetermined NAL unit types of the second set, disjoint from the first set, wherein the one or more NAL units of the second set are arranged in the NAL unit sequence or interspersed in the NAL unit sequence. The apparatus distinguishes each NAL unit of the second type from the set of one or more NAL units of the first type by checking a NAL unit type syntax element in each NAL unit indicating the NAL unit type. If the syntax element has a value within the first set of values, it is designated as a NAL unit of the first NAL unit type, and if the syntax element has a value within the second set of values, it is designated as a NAL unit of the second NAL unit type. Different types of NAL units 326 (second type) can be distinguished. This can be done by the NAL unit type syntax element, that is, by assigning different NAL unit types (the second set of values of the NAL unit type syntax element) to these different NAL units 326 such as types A, B and C as described above, reserved NAL unit types. However, the distinction between the latter can also be done by a syntax element in the payload portion of each inserted NAL unit 326. The NAL unit type syntax element also distinguishes slice NAL units from parameter set NAL units, that is, they have NAL unit type syntax elements with different values in the first set of values. For the predetermined NAL unit 326 of the second set, the device 340 discards the immediately following NAL unit 330 of the first set or its portion 402 and the corresponding NAL unit of the second set from the NAL unit sequence. For one NAL unit 326, the device 340 may discard the subsequent NAL unit slice in its entirety, and for another, only discard the portion 402. However, the device may also process only one of these options. The portion 402 may be indicated by a length indication in the payload portion of the corresponding NAL unit 326, or the portion 402 may be determined by 340, or otherwise such as by parsing. The portion 402 is, for example, a slice header of a subsequent NAL unit.Additionally or alternatively, the device may, for another predetermined NAL unit 326 of the second set, discard the immediately following NAL unit 330 of the first set or a portion thereof 402 and the corresponding NAL unit of the second set from the NAL unit sequence, and insert a NAL unit 332 of one of the first NAL unit types carried in the payload portion of the corresponding NAL unit of the second set into the NAL unit sequence in place of the corresponding NAL unit of the second set. Additionally or alternatively, the device may, for even another predetermined NAL unit 326 of the second set, insert a NAL unit 332 of one of the first NAL unit types carried in the payload portion of the corresponding NAL unit of the second set into the NAL unit sequence in place of the corresponding NAL unit of the second set. The device 340 may be configured to concatenate the transformed sequence of NAL units obtained by discarding and / or inserting (wherein the slice address contained in one or more of the one or more NAL units of the first set is potentially modified) with one or more other NAL unit sequences 3242, wherein each spatial segment of a larger image other than the predetermined spatial segment is encoded. The concatenation orders the sequence of NAL units along the coding order such that the slice addresses in the slice NAL units within the sequence of NAL units increase monotonically. The concatenation together with the execution of the instructions indicated by the NAL unit 326 produces a data stream 302' which is equal to the data stream 302 or is an equivalent thereof, i.e., its decoding produces an image 300 having a predetermined spatial portion as one of its spatial segments. If any self-contained partial data stream would be fed to a legacy decoder, the decoding of the corresponding spatial segment would result due to the discarding of the NAL unit 326.
[0248] about Fig.12 and 16 It should be noted that the stream formatter 322 shown therein can be split into two consecutive entities, i.e. one that performs only the task of distributing the NAL units of all partial data streams from a common stream to the respective partial data streams, i.e. the task of demultiplexing or stream distribution, and another connected upstream of the former that performs the remaining tasks mentioned above. At least some of the tasks can even be performed by the encoder that generates the original NAL unit 314.
[0249] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent descriptions of corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding blocks or items or features of corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices, such as microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps may be performed by such devices.
[0250] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. This implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM or FLASH memory, which has electronically readable control signals stored thereon that cooperate (or can cooperate) with a programmable computer system to perform the corresponding method. Therefore, the digital storage medium may be computer readable.
[0251] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0252] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.The program code may, for example, be stored on a machine readable carrier.
[0253] Other embodiments comprise the computer program for performing one of the methods described herein stored on a machine readable carrier.
[0254] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0255] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory.
[0256] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.The data stream or the sequence of signals may, for example, be configured to be transmitted via a data communication connection, for example via the Internet.
[0257] A further embodiment comprises a processing means, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0258] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0259] Another embodiment according to the invention comprises an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transmitting the computer program to the receiver.
[0260] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may collaborate with a microprocessor to perform one of the methods described herein. Typically, the method is preferably performed by any hardware device.
[0261] The above embodiments are only used to illustrate the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the upcoming patent claims, rather than by the specific details presented by the description and explanation of the embodiments herein.
Claims
1. A stream multiplexer (103), comprising: A receiving interface (105) for receiving coded data in which each of at least two different spatial segments (108) or groups of different sequence spatial segments (108) of a video image (110) of a video stream (102) is coded; and a data stream former (106) configured to packetize the coded data of each of the at least two groups of different spatial segments (108) or different sequence spatial segments (108) into a separate elementary stream (114) and to provide the separate elementary stream (114) at an output; wherein at least two different spatial segments or different sequential groups of spatial segments of a video picture of the video stream are encoded such that for each video picture the encoded data comprises at least one slice; wherein the data stream former is configured to pack one or more slices of each spatial segment or group of spatial segments into a single elementary stream; The data stream former is configured to, for a slice of at least one of a spatial segment or a sequential spatial segment group, Packing slices of at least one of a spatial segment or a sequence of spatial segment groups into a single elementary stream; as well as One or more elementary streams are provided, the one or more elementary streams comprising slice headers and / or parameter sets dedicated for modifying or replacing slice headers and / or parameter sets of a plurality of individual elementary streams.
2. The stream multiplexer (103) of claim 1, wherein the at least two different spatial segments (108) or different sequence spatial segments (108) groups are coded such that the coded data of each of the at least two different spatial segments (108) or different sequence spatial segments (108) groups contained in the respective individual elementary streams (114) only references coded data contained in the same individual elementary stream.
3. The stream multiplexer (103) of claim 1, wherein the data stream former (106) is configured to packetize slices of at least one of a spatial segment (108) or a group of sequential spatial segments (108) into separate elementary streams (114) in a manner removing their slice headers; The data stream former (106) is configured to provide another separate stream, the other separate stream comprising a suitable slice header modified with respect to its image position or image size relative to the removed slice header.
4. The stream multiplexer (103) of claim 1, wherein at least two different spatial segments (108) or groups of different sequence spatial segments (108) are encoded into the video stream with one slice per video picture; The data stream former (106) is configured to pack, for each video picture (110), a portion of a slice in which a corresponding one of the spatial segments (108) or groups of spatial segments (108) is encoded into a separate elementary stream (114), without a slice header for the slice for at least one of the at least two spatial segments (108) or groups of spatial segments (108).
5. The stream multiplexer (103) of claim 4, wherein the data stream former (106) is configured to provide a further individual stream comprising a slice header or a modified version thereof modified with respect to the removed slice header with respect to its picture position or picture size.
6. The stream multiplexer (103) of claim 4, wherein the data stream former (106) is configured to provide a further separate stream with parameter sets for each of at least one spatial subdivision of the video images (110) of the video stream (102) into sub-regions (109).
7. The stream multiplexer (103) of claim 3, wherein the data stream former (106) is configured to provide the parameter sets and the slice header or a modified version thereof in the same further separate elementary stream.
8. The stream multiplexer (103) of claim 3, wherein the data stream former (106) is configured to signal the presence of the further elementary streams in the descriptor.
9. The stream multiplexer (103) of claim 1, wherein the data stream former (106) is configured to generate an elementary stream descriptor which assigns a unique elementary stream identification to each of the individual elementary streams (114).
10. The stream multiplexer (103) of claim 9, wherein the data stream former (106) is configured to generate a sub-region descriptor which, for each of at least one spatial subdivision of a video image (110) of a video stream (102) to a sub-region (109), signals a set of elementary stream identifications for each sub-region (109).
11. The stream multiplexer (103) of claim 1, wherein the stream multiplexer (103) is configured to signal at least one of two stream types; wherein the first stream type signals that the combination of an appropriate subset of the individual elementary streams (114) corresponding to one of at least one spatial subdivision of the video image (110) of the video stream (102) to the sub-region (109) produces a data stream compliant with the standard; The second stream type signals that the combination of an appropriate subset of the individual elementary streams (114) corresponding to at least one spatial subdivision of the video image (110) of the video stream (102) to the sub-area (109) produces a data stream that needs to be further processed to obtain a version of the data stream that complies with the standard.
12. The stream multiplexer (103) of claim 11, wherein at least two different spatial segments (108) or groups of different sequence spatial segments (108) of video images (110) of the video stream (102) are encoded such that the encoded data comprises at least two slices; wherein the coded data comprises signaling information indicating whether the coding constraint is satisfied, or wherein the data stream former is configured to determine whether the coding constraint is satisfied; wherein the coding constraint is satisfied when removing at least one slice header of the at least two slices while maintaining a slice header of a first slice of the at least two slices with respect to coding order and concatenating the at least two slices or a suitable subset of the at least two slices using the first slice header produces a data stream conforming to the standard; Wherein the stream multiplexer is configured to signal at least one of the two stream types depending on the satisfaction of the coding constraint.
13. A stream demultiplexer (121), comprising: A data stream former (122) configured to selectively extract at least two separate elementary streams (114) from the group of separate elementary streams (116), the at least two separate elementary streams (114) containing coded data of different spatial segments (108) or different groups of sequential spatial segments (108) of video pictures (110) of the coded video stream (102), wherein the data stream former (122) is configured to combine the at least two separate elementary streams (114) into a data stream (126) containing coded data of different spatial segments (108) or different groups of sequential spatial segments (108) of video pictures (110) of the coded video stream (102), wherein the at least two different spatial segments or different groups of sequential spatial segments of the video pictures of the video stream are coded such that for each video picture the coded data comprises at least one slice; wherein one or more slices of each spatial segment or group of spatial segments are packed into one separate elementary stream; in Packing slices of at least one of a spatial segment or a sequence of spatial segments into separate elementary streams; wherein one or more elementary streams include slice headers and / or parameter sets dedicated to modifying or replacing slice headers and / or parameter sets of multiple separate elementary streams; and an output interface (123) configured to provide a data stream (126); The data stream former is configured to combine at least two separate elementary streams into a data stream that complies with the HEVC standard.
14. The stream demultiplexer (121) of claim 13, wherein the group of individual elementary streams (116) comprises a plurality of individual elementary streams (114), the individual elementary streams (114) containing coded data of different spatial segments (108) or different sets of sequential spatial segments (108) of video pictures (110) of the coded video stream (102); The data stream former (122) is configured to selectively extract a subset of the plurality of individual elementary streams (114) of the individual elementary stream group (116).
15. A stream demultiplexer (121) as claimed in claim 14, wherein the data stream former (122) is configured to extract from the separate elementary stream group (116) a sub-region descriptor signalling a set of elementary stream identifications for each of at least one spatial subdivision of the video image (110) of the video stream to the sub-region (109), wherein the data stream former (122) is configured to use the sub-region descriptors to select a subset of elementary streams (114) to be extracted from the coded video stream (112).
16. The stream demultiplexer (121) of claim 15, wherein the data stream former (122) is configured to extract an elementary stream identifier from the group of individual elementary streams (116), the elementary stream identifier assigning a unique elementary stream identification to each of the individual elementary streams (114), wherein the data stream former (122) is configured to use the elementary stream descriptors to locate a subset of elementary streams (114) to be extracted from the group of individual elementary streams (116).
17. The stream demultiplexer (121) of claim 13, wherein the stream demultiplexer (121) is configured to identify at least one of at least two stream types; wherein the combination of the first stream type indication with an appropriate subset of the individual elementary streams (114) corresponding to one of the at least one spatial subdivisions of the video image (110) of the video stream (102) to the sub-region (109) produces a data stream compliant with the standard; wherein the combination of the second stream type indication with an appropriate subset of the individual elementary streams (114) corresponding to one of the at least one spatial subdivision of the video image (110) of the video stream (102) to the sub-region (109) produces a data stream that needs to be further processed to obtain a version of the data stream that complies with the standard; The stream demultiplexer (121) comprises a data stream processor (127) configured to further process the data stream (126) according to the identified stream type using processing information contained in at least one elementary stream (114s, 114p) of the separate elementary stream group to obtain a version (126') of the data stream that complies with the standard.
18. The stream demultiplexer (121) of claim 13, wherein A separate elementary stream group (116) is included in the broadcast transport stream and comprises: a plurality of separate elementary streams (114), the separate elementary streams (114) containing coded data of different spatial segments (108) or different sets of sequential spatial segments (108) of the coded video stream (102); and Program map table; as well as The stream demultiplexer (121) is configured to derive a stream identifier from the program map table for each of the plurality of individual elementary streams (114) and to distinguish each of the plurality of individual elementary streams (114) in a broadcast transport stream using the corresponding stream identifier.
19. The stream demultiplexer (121) of claim 18, wherein the stream demultiplexer (121) is configured to derive the predetermined packet identifier from a program association table transmitted in a packet of packet identifier zero in the broadcast transport stream, and to derive the program map table from a packet of the broadcast transport stream having the predetermined packet identifier.
20. The stream demultiplexer (121) of claim 18, wherein the program map table uniquely associates each stream identifier with a corresponding packet identifier, and the stream demultiplexer is configured to depacketize each of the plurality of individual elementary streams from packets of the broadcast transport stream having a packet identifier associated with the stream identifier of the corresponding individual elementary stream, or The stream demultiplexer is configured to depacketize a sequence of NAL units from packets of a broadcast transport stream having packet identifiers indicated in a program map table and to associate each NAL unit with one of a plurality of separate elementary streams according to a stream identifier indicated in an adaptation field of the packets of the broadcast transport stream.
21. The stream demultiplexer (121) of claim 18, wherein the stream demultiplexer (121) is configured to read information about the spatial subdivision of the video into segments from a program map table and to derive stream identifiers of the plurality of individual elementary streams (114) inherently from the spatial subdivision by using a mapping from segments of the spatial subdivision to stream identifiers.
22. A stream demultiplexer (121) as claimed in claim 18, wherein the stream demultiplexer (121) is configured to read information about the spatial subdivision of the video into segments from a program map table and derive stream identifiers of the plurality of individual elementary streams (114) such that the stream identifiers of the plurality of individual elementary streams (114) order the segments within the spatial subdivision along a raster scan.
23. A stream demultiplexer (121) as claimed in claim 18, wherein the stream demultiplexer (121) is configured to read elementary stream descriptors from a program map table or an adaptation field of a data packet carrying a group of individual elementary streams (116), each elementary stream descriptor indexing one of a plurality of individual elementary streams by a stream identifier of the individual elementary stream and comprising information about which one or more individual elementary streams of the plurality of individual elementary streams together with the indexed individual elementary stream form a coded representation of a sub-region extractable from the group of individual elementary streams (116) as at least two individual elementary streams (114), the sub-region consisting of a spatial segment (108) or a group of sequence spatial segments of the one or more individual elementary streams together with the indexed individual elementary stream forming a coded representation.
24. A stream demultiplexer (121) as claimed in claim 23, wherein the stream demultiplexer (121) is configured to read information about a sub-region descriptor from a program map table, the sub-region descriptor indicating one or more spatial subdivisions of the video to the sub-region such that each sub-region is a set of spatial segments (108) of one or more separate elementary streams or a sequential set of spatial segments, and a size of the sub-region under each spatial subdivision of the video to the sub-region (109), wherein at least two separate elementary streams selectively extracted from the group of separate elementary streams (116) together contain coded data for encoding one of the different spatial segments (108) or different sequential sets of spatial segments (108) of the sub-region forming one of the one or more spatial subdivisions of the video.
25. A stream demultiplexer (121) as claimed in claim 22, wherein one or more of the elementary stream descriptors contain information about which one or more individual elementary streams of a plurality of individual elementary streams together with the indexed individual elementary streams are to be extracted as at least two individual elementary streams (114) from the group of individual elementary streams (116) in the form of a reference index to a list of sets of stream identifier offsets, each stream identifier offset indicating an offset relative to a stream identifier of the indexed individual elementary stream.
26. The stream demultiplexer (121) of claim 22, wherein one or more of the elementary stream descriptors contain information about which one or more individual elementary streams of the plurality of individual elementary streams together with the indexed individual elementary streams are to be extracted as at least two individual elementary streams (114) from the group of individual elementary streams (116) in the form of a set of stream identifier offsets, each stream identifier offset indicating an offset relative to a stream identifier of the indexed individual elementary stream.
27. The stream demultiplexer (121) of claim 13, wherein The individual elementary stream group (116) comprises one or more subgroups of individual elementary streams, Each encoding a corresponding space segment (108) or a corresponding sequence of space segments (108) groups, Each consists of a sequence of NAL units, and The NAL unit sequence consists of the following: a first set of one or more NAL units forming a conforming version of a data stream representing a corresponding spatial segment (108) or a corresponding set of sequence spatial segments (108) of a corresponding elementary stream, and The one or more NAL units of the second set are one of a set of one or more predetermined NAL unit types for ignoring the corresponding NAL units by a legacy decoder.
28. The stream demultiplexer (121) of claim 27, wherein the one or more NAL units of the second set are arranged in a NAL unit sequence, each NAL unit of the second set: indicating to a non-legacy decoder that an immediately following NAL unit of the first set, or a portion thereof, is to be dropped from the NAL unit sequence along with the corresponding NAL unit of the second set, and / or Contains a payload portion carrying NAL units to be inserted into the NAL unit sequence in place of corresponding NAL units of the second set.
29. The stream demultiplexer (121) of claim 27, configured to, if any one of the subgroup of individual elementary streams is in the extracted at least two individual elementary streams (114), for each NAL unit of the second set: discarding from the NAL unit sequence the immediately following NAL unit of the first set or a portion thereof and the corresponding NAL unit of the second set, and / or The NAL unit carried in the payload portion of the corresponding NAL unit of the second set is inserted into the NAL unit sequence to replace the corresponding NAL unit of the second set.
30. A method (200) for stream multiplexing, the method comprising: Receiving (202) coded data for each of at least two different spatial segments (108) or groups of different sequence spatial segments (108) of a video picture (110) of a video stream (102); and Packing (204) coded data of each of at least two different spatial segments (108) or groups of different spatial segments (108) into a single elementary stream, wherein at least two different spatial segments or groups of different sequences of spatial segments of a video picture of the video stream are coded such that for each video picture the coded data comprises at least one slice; wherein one or more slices of each spatial segment or group of spatial segments are packed into a single elementary stream; wherein slices of at least one of a spatial segment or a sequence of spatial segment groups are packed into a single elementary stream; and One or more elementary streams are provided, the one or more elementary streams comprising slice headers and / or parameter sets dedicated for modifying or replacing slice headers and / or parameter sets of a plurality of individual elementary streams.
31. A method (220) for stream demultiplexing, the method comprising: Selectively extracting (222) at least two separate elementary streams (114) from the group of separate elementary streams (116), the at least two separate elementary streams (114) containing coded data of different spatial segments (108) or different groups of sequential spatial segments (108) of video pictures (110) of the coded video stream (102), wherein the at least two different spatial segments or different groups of sequential spatial segments of the video pictures of the video stream are coded such that for each video picture the coded data comprises at least one slice; wherein one or more slices of each spatial segment or group of spatial segments are packed into one separate elementary stream; wherein slices of at least one of the spatial segments or groups of sequential spatial segments are packed into one separate elementary stream; wherein one or more elementary streams comprise slice headers and / or parameter sets dedicated to modifying or replacing slice headers and / or parameter sets of the plurality of separate elementary streams; and At least two separate elementary streams (114) are combined (224) into a data stream (126) containing coded data of different spatial segments (108) or different sequence groups of spatial segments (108) of video images (110) of a coded video stream (102), wherein the at least two separate elementary streams are combined into a data stream compliant with the HEVC standard.
32. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it is used to perform the method according to claim 30 or 31.
Citation Information
Patent Citations
Sub-slices in video coding
US20120189049A1
Spatially-Segmented Content Delivery
US20140089990A1
Method and apparatus for transreceiving broadcast signal for panorama service
WO2015126144A1