Apparatus, method and computer program for processing random access pictures in video coding
By repeatedly reusing random access pictures in video bitstreams, the problem of failure to utilize time-on-interleaved lens similarity in the prior art is solved, and encoding efficiency and quality are improved.
Patent Information
- Application Number
- CN202080050322.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-16
- Filing Date
- 2020-05-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-05-14
AI Technical Summary
Existing video encoding technologies fail to effectively utilize the similarity between time-interleaved lenses, resulting in inefficient encoding.
Reusable IRAP pictures are achieved by reusing random access pictures (IRAP pictures) multiple times in the video bitstream and controlling their display or not to display through the bitstream or encoder.
Improve the efficiency of video encoding, reduce redundant data, and improve the encoding quality.
Smart Images

Figure CN114270868B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to apparatuses, methods, and computer programs for video encoding and decoding. Background Art
[0002] This section is intended to provide background or context for the present invention as recited in the claims. The description herein may include concepts that could be pursued but that were not necessarily previously envisioned or pursued. Thus, unless otherwise indicated herein, the content described in this section is not prior art to the specification and claims in this application and will not be admitted to be prior art by inclusion in this section.
[0003] A video coding system may include an encoder that transforms an input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a visual form. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form, e.g., to enable storage / transmission of video information at a lower bitrate than might otherwise be required.
[0004] Many types of content may be captured using several cameras, which may be temporally interleaved to produce video to be encoded. Such temporally interleaved shots from different cameras may be present in, for example, movies, sitcoms, late-night talk shows, conferences, and surveillance. Shots that are not temporally adjacent and that originate from the same camera may have substantial similarity or correlation. However, since each shot conventionally starts with a random access picture, the similarity between shots has not been exploited in traditional video compression.
[0005] Some problems have been identified in some methods that utilize cross-shot correlation. Summary of the Invention
[0006] Now, in order to at least mitigate the above problems, an enhanced coding method is introduced herein. Embodiments are provided for reusing random access pictures multiple times in video encoding and decoding. In some embodiments, methods, apparatuses, and computer program products are provided for reusing random access pictures multiple times.
[0007] A Random Access Segment (RAS) may be defined as a sequence of coded pictures that starts with a random access picture in decoding order. It may be specified, but in some cases need not be specified, that no picture other than the first picture in the RAS is a random access picture. For example, in some cases, the duration of a random access segment is constant and the random access picture may also appear in the middle of the RAS.
[0008] A random access picture can be defined as a picture that does not refer to any picture other than itself as a reference for prediction during its decoding process, and constraints can be imposed on subsequent pictures in decoding and output order such that those pictures do not refer to any picture that is earlier in decoding order than the random access picture as a reference for prediction during their decoding processes.
[0009] Although concepts specific to HEVC and / or H.266 / VVC are used in describing several embodiments, similar aspects can be used with any video codec. For example, instead of using the term Intra Random Access Point (IRAP) picture, the term random access picture can also be used.
[0010] Encoding according to an embodiment is performed in a manner in which a selected IRAP picture (hereinafter "reusable IRAP picture") is repeated in the encoded bitstream. Similarly, decoding according to an embodiment can be arranged in a manner in which the reusable IRAP picture is repeated in the bitstream to be decoded. Thus, the reusable IRAP picture can occur repeatedly in a conventional video bitstream (similarly, DASH representation or the like). However, multiple transmissions of the reusable IRAP picture can be avoided by system components as described in this specification.
[0011] The reusable IRAP can be used only as a reference for prediction and not for display. This can be controlled by the encoder in the bitstream or along with the bitstream, for example, using the pic_output_flag of HEVC and H.266 / VVC. For example, when the value of the pic_output_flag assigned to the reusable IRAP picture is 0, the reusable IRAP picture will not be displayed, and when the value of the pic_output_flag assigned to the reusable IRAP picture is 1, the reusable IRAP picture can be displayed.
[0012] Some embodiments can utilize Figure 1a to illustrate. Pictures that start a RAS and are indicated with the same symbol (circle / square / triangle) are the same reusable IRAP picture. In Figure 1aAmong them, there are three pairs of identical IRAP pictures, the first pair (circular, circular with hatching) in the first RAS and the sixth RAS, the second pair (square, square with hatching) in the second RAS and the fourth RAS, and the third pair (triangle, triangle with hatching) in the third RAS and the fifth RAS. In this example, the reusable IRAP pictures from the first RAS to the third RAS are assigned a pic_output_flag with a value of 1 indicating that these pictures can be output, and the reusable IRAP pictures from the fourth to the sixth RAS are assigned a withpic_output_flag with a value of 0 indicating that these pictures do not need to be output but can be used as references.
[0013] The method according to the first aspect includes:
[0014] Obtaining an intra random access point picture from a first position of an encoded video bitstream;
[0015] Determining whether the intra random access point picture is reusable in at least a second position of the encoded video bitstream, the at least second position being different from the first position;
[0016] In the case of determining that the intra random access point picture is reusable, providing an identification assigned to the intra random access point picture that the intra random access point picture is a reusable intra random access point picture.
[0017] The method according to the second aspect includes:
[0018] Encoding pictures into multiple layers, where the encoding includes:
[0019] Selecting a layer for each picture among the pictures;
[0020] Encoding each of the pictures into an encoded picture within the layer; and
[0021] Including temporally related pictures into the same independent layer, where the pictures within the independent layer are not predicted from the pictures of other layers.
[0022] The apparatus according to the third aspect includes at least one processor and at least one memory, the at least one memory including computer program code, the memory and the computer program code being configured to, together with the at least one processor, cause the apparatus to at least perform the following operations:
[0023] Obtaining an intra random access point picture from a first position of an encoded video bitstream;
[0024] Determining whether the intra random access point picture is reusable in at least a second position of the encoded video bitstream, the at least second position being different from the first position;
[0025] When it is determined that the picture of the internal random access point is reusable, provide an identification assigned to the picture of the internal random access point indicating that the picture of the internal random access point is a reusable picture of the internal random access point.
[0026] The non-transitory computer program product according to the fourth aspect includes computer program code, and the computer program code is configured to cause a device or system when running on at least one processor:
[0027] Obtain a picture of an internal random access point from a first position of an encoded video bitstream;
[0028] Determine whether the picture of the internal random access point is reusable in at least a second position of the encoded video bitstream, where the at least second position is different from the first position;
[0029] When it is determined that the picture of the internal random access point is reusable, provide an identification assigned to the picture of the internal random access point indicating that the picture of the internal random access point is a reusable picture of the internal random access point.
[0030] The device according to the fifth aspect includes:
[0031] A first circuit system configured to obtain a picture of an internal random access point from a first position of an encoded video bitstream;
[0032] A second circuit system configured to determine whether the picture of the internal random access point is reusable in at least a second position of the encoded video bitstream, where the at least second position is different from the first position;
[0033] A third circuit system configured to provide an identification assigned to the picture of the internal random access point indicating that the picture of the internal random access point is a reusable picture of the internal random access point when it is determined that the picture of the internal random access point is reusable.
[0034] The method according to the sixth aspect includes:
[0035] Receive an identification associated with a picture of an internal random access point in a first position of an encoded video bitstream, the identification indicating whether the picture of the internal random access point is reusable in at least a second position of the encoded video bitstream, where the at least second position is different from the first position;
[0036] When the identification indicates that the picture of the internal random access point is reusable, check whether the corresponding previously received picture of the internal random access point is available;
[0037] When it is checked that the corresponding previously received picture of the internal random access point is available, omit processing the picture of the internal random access point.
[0038] The apparatus according to the seventh aspect comprises at least one processor and at least one memory, the at least one memory comprising computer program code, the memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least perform the following operations:
[0039] Receiving an identification associated with an intra random access point picture at a first position of an encoded video bitstream, the identification indicating whether the intra random access point picture is reusable at at least a second position of the encoded video bitstream, the at least second position being different from the first position;
[0040] Checking whether a corresponding previously received intra random access point picture is available in case the identification indicates that the intra random access point picture is reusable;
[0041] Omitting processing of the intra random access point picture in case it is checked that the corresponding previously received intra random access point picture is available.
[0042] The apparatus according to the eighth aspect comprises at least one processor and at least one memory, the at least one memory comprising computer program code, the memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least perform the following operations:
[0043] A first circuitry configured to receive an identification associated with an intra random access point picture at a first position of an encoded video bitstream, the identification indicating whether the intra random access point picture is reusable at at least a second position of the encoded video bitstream, the at least second position being different from the first position;
[0044] A second circuitry configured to check whether a corresponding previously received intra random access point picture is available in case the identification indicates that the intra random access point picture is reusable;
[0045] A third circuitry configured to omit processing of the intra random access point picture in case it is checked that the corresponding previously received intra random access point picture is available.
[0046] The apparatus according to the ninth aspect comprises:
[0047] Means for obtaining an intra random access point picture from a first position of an encoded video bitstream;
[0048] Means for determining whether the intra random access point picture is reusable at at least a second position of the encoded video bitstream, the at least second position being different from the first position;
[0049] A component for providing, when it is determined that an intra random access point picture is reusable, an identification that the intra random access point picture assigned to the intra random access point picture is a reusable intra random access point picture.
[0050] The apparatus according to the tenth aspect includes:
[0051] A component for receiving an identification associated with an intra random access point picture in a first position of an encoded video bitstream, the identification indicating whether the intra random access point picture is reusable in at least a second position of the encoded video bitstream, the at least second position being different from the first position;
[0052] A component for checking whether a corresponding previously received intra random access point picture is available when the identification indicates that the intra random access point picture is reusable;
[0053] A component for omitting processing of the intra random access point picture when it is checked that the corresponding previously received intra random access point picture is available.
[0054] The method according to the eleventh aspect includes:
[0055] Decoding pictures encoded into multiple layers from a bitstream, wherein the decoding includes:
[0056] Selecting a layer for each picture among the pictures;
[0057] Decoding each picture according to the encoded pictures within the layer, wherein the decoding is independent of other layers; and
[0058] Including temporally related decoded pictures into the same independent layer, wherein the pictures within the independent layer are not predicted from the pictures of other layers.
[0059] Another aspect relates to an apparatus and a computer-readable storage medium having code stored thereon, the code being arranged to execute one or more of the above methods and their related embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] To better understand the present invention, reference will now be made by way of example to the accompanying drawings, in which:
[0061] In the drawings:
[0062] Figure 1a An example sequence of some reusable pictures is shown;
[0063] Figure 1b An example of including pictures into several layers is shown;
[0064] Figure 2ais a flowchart illustrating an encoding method according to an embodiment;
[0065] Figure 2b is a flowchart illustrating a decoding method according to an embodiment;
[0066] Figure 2c is a flowchart illustrating a decoding method according to another embodiment;
[0067] Figure 3a shows a schematic diagram of an encoder suitable for implementing an embodiment of the present invention;
[0068] Figure 3b shows a schematic diagram of a decoder suitable for implementing an embodiment of the present invention;
[0069] Figure 4a shows some elements of a video coding section according to an embodiment;
[0070] Figure 4b shows some elements of a video decoding section according to an embodiment;
[0071] Figure 5 shows a device according to an embodiment. Detailed Description
[0072] In the following, several embodiments will be described in the context of a video coding arrangement. However, it should be noted that the embodiments are not limited to this particular device. For example, some embodiments may be applicable to video coding systems such as streaming systems, DVD (Digital Versatile Disc) players, digital television receivers, personal video recorders, systems and computer programs on personal computers, handheld computers and communication devices, and network elements such as transcoders and cloud computing devices where video data is processed.
[0073] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some but not all embodiments are shown. In fact, the various embodiments may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms "data", "content", "information" and similar terms may be used interchangeably to refer to data that can be sent, received and / or stored according to an embodiment. Thus, the use of any such term should not be construed as limiting the spirit and scope of the embodiments.
[0074] Additionally, as used herein, the term "circuitry" refers to (a) only hardware circuit implementations (e.g., implementations in analog and / or digital circuitry); (b) a combination of circuitry and one or more computer programs including software and / or firmware instructions stored on one or more computer-readable memories that work together to cause an apparatus to perform one or more functions described herein; and (c) circuitry, such as, for example, a portion of a (multi-)microprocessor or (multi-)microprocessors that requires software or firmware for operation even if the software or firmware is not physically present. This definition of "circuitry" applies to all uses of the term herein (including in any claims). As another example, as used herein, the term "circuitry" also includes implementations that include one or more processors and / or portions thereof and accompanying software and / or firmware. As another example, as used herein the term "circuitry" also includes, for example, baseband integrated circuits or application processor integrated circuits for mobile phones, or similar integrated circuits in servers, cellular network devices, other network devices, and / or other computing devices.
[0075] As defined herein, a "computing-readable storage medium" (which refers to a non-transitory physical storage medium (e.g., a volatile or non-volatile memory device)) can be distinguished from a "computer-readable transmission medium", which refers to an electromagnetic signal.
[0076] In the following, several embodiments are described using a convention related to (de)coding, which indicates that the embodiments can be applied to decoding and / or encoding.
[0077] The Advanced Video Coding standard (which may be abbreviated as AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunication Standardization Sector (ITU-T) of the International Telecommunication Union and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard was published by two parent standardization organizations and it is known as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, and is also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, each integrating new extensions or features into the specification. These extensions include Scalable Video Coding (SVC) and Multi-View Video Coding (MVC).
[0078] The High Efficiency Video Coding standard, which may be abbreviated as HEVC or H.265 / HEVC, was developed by the Joint Collaborative Team on Video Coding (JCT-VC) of VCEG and MPEG. This standard is published by two parent standardization organizations and is known as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, and is also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Extensions of H.265 / HEVC include scalable, multi-viewpoint, three-dimensional, and Fidelity Range Extensions, which may be referred to as SHVC, MV-HEVC, 3D-HEVC, and REXT, respectively. References to H.265 / HEVC, SHVC, MV-HEVC, 3D-HEVC, and REXT in this specification have been made for the purpose of understanding the definitions, and unless otherwise indicated, the structures or concepts of these standard specifications should be understood as references to the latest versions of these standards available prior to the date of this application.
[0079] The Versatile Video Coding standard (VVC, H.266, or H.266 / VVC) is currently under development by the Joint Video Exploration Team (JVET), which is a collaboration between ISO / IEC MPEG and ITU-T VCEG.
[0080] Some key definitions, bitstreams, and coding structures, as well as some of the concepts of H.264 / AVC, HEVC, VVC, and their extensions, are described in this section as examples of video encoders, decoders, coding methods, decoding methods, and bitstream structures, in which embodiments may be implemented. Aspects of various embodiments are not limited to H.264 / AVC or HEVC or VVC or their extensions, but rather this specification is given on a possible basis on which the embodiments herein may be partially or fully implemented. Whenever reference is made hereinafter to any of VVC or its draft versions, it is to be understood that this specification matches the VVC draft specification, and changes may exist in later draft versions and (multiple) final versions of VCC, and the specification and embodiments may be adjusted to match the (multiple) final versions of VVC.
[0081] A video codec may include an encoder that transforms an input video into a compressed representation suitable for storage / transmission and a decoder that may decompress the compressed video representation back into a visual form. The compressed representation may be referred to as a bitstream or a video bitstream. The video encoder and / or video decoder may also be separated from each other, i.e., it is not necessary to form a codec. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bitrate).
[0082] A hybrid video codec, such as ITU-T H.264, can encode video information in two stages. First, pixel values in a particular picture region (or “block”) are predicted, for example, by a motion compensation component (which finds and indicates a region in a previously encoded video frame that closely corresponds to the block being encoded) or by a spatial component (which uses pixel values surrounding the block to be encoded in a specified manner). Then, the prediction error, i.e., the difference between the predicted pixel block and the original pixel block, is encoded. This can be done by transforming the difference using a specified transform (e.g., the discrete cosine transform (DCT) or a variant thereof), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size or transmission bit rate).
[0083] In temporal prediction, the source of prediction is a previously decoded picture (i.e., a reference picture). In Intra Block Copy (IBC; i.e., Intra Block Copy prediction or current picture reference), prediction is similarly applied to temporal prediction, but the reference picture is the current picture and only previously decoded samples can be involved in the prediction process. Inter-layer or inter-view prediction can be similarly applied to temporal prediction, but correspondingly, the reference picture is a decoded picture from another scalable layer or from another view. In some cases, inter-frame prediction may only refer to temporal prediction, while in other cases, inter-frame prediction may collectively refer to temporal prediction and any of Intra Block Copy, inter-layer prediction, and inter-view prediction, provided that they are performed using the same or a similar process as temporal prediction. Inter-frame prediction or temporal prediction can sometimes be referred to as motion compensation or motion-compensated prediction.
[0084] Intra prediction takes advantage of the fact that neighboring pixels within the same picture are likely to be correlated. Intra prediction can be performed in the spatial or transform domain, i.e., sample values or transform coefficients can be predicted. Intra prediction is typically utilized in intra coding, where no inter-frame prediction is applied.
[0085] One result of the encoding process is a set of encoded parameters, such as motion vectors and quantized transform coefficients. Many parameters can be more efficiently entropy encoded if they are first predicted from spatially or temporally neighboring parameters. For example, a motion vector can be predicted from a temporally neighboring motion vector and only the difference relative to the motion vector predictor can be encoded. The prediction of encoded parameters and intra prediction can be collectively referred to as intra-picture prediction.
[0086] Figure 3a A block diagram of a video encoder suitable for adopting an embodiment of the present invention is shown. Figure 3aAn encoder for two layers is presented, but it will be appreciated that the presented encoder can be similarly simplified to encode only one layer or extended to encode more than two layers. Figure 3a An embodiment of a video encoder is illustrated that includes a first encoder section 500 for a base layer and a second encoder section 502 for an enhancement layer. Each of the first encoder section 500 and the second encoder section 502 can include similar elements for encoding incoming pictures. The encoder sections 500, 502 can include pixel predictors 302, 402, prediction error encoders 303, 403, and prediction error decoders 304, 404. Figure 3a Embodiments of the pixel predictors 302, 402 are also shown to include an inter-frame predictor 306, 406, an intra-frame predictor 308, 408, a mode selector 310, 410, filters 316, 416, and reference frame memories 318, 418. The pixel predictor 302 of the first encoder section 500 receives the base layer image of the video stream 300 to be encoded at both the inter-frame predictor 306 (which determines the difference between the image and the motion-compensated reference frame 318) and the intra-frame predictor 308 (which determines the prediction for an image block based only on the already processed portions of the current frame or picture). The outputs of both the inter-frame predictor and the intra-frame predictor are passed to the mode selector 310. The intra-frame predictor 308 can have more than one intra-frame prediction mode. Thus, each mode can perform intra-frame prediction and provide the predicted signal to the mode selector 310. The mode selector 310 also receives a copy of the base layer picture 300. Correspondingly, the pixel predictor 402 of the second encoder section 502 receives the enhancement layer image of the video stream 400 to be encoded at both the inter-frame predictor 406 (which determines the difference between the image and the motion-compensated reference frame 418) and the intra-frame predictor 408 (which determines the prediction for an image block based only on the already processed portions of the current frame or picture). The outputs of both the inter-frame predictor and the intra-frame predictor are passed to the mode selector 410. The intra-frame predictor 408 can have more than one intra-frame prediction mode. Thus, each mode can perform intra-frame prediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the enhancement layer picture 400.
[0087] Depending on which encoding mode is selected to encode the current block, the output of the inter-frame predictors 306, 406 or the output of one of the optional intra-frame predictor modes or the output of the surface encoder within the mode selector is passed to the outputs of the mode selectors 310, 410. The output of the mode selector is passed to first summing devices 321, 421. The first summing devices can subtract the outputs of the pixel predictors 302, 402 from the base layer picture 300 / enhancement layer picture 400 to produce first prediction error signals 320, 420, which are input to the prediction error encoders 303, 403.
[0088] The pixel predictors 302, 402 also receive a combination of the predicted representations of the image blocks 312, 412 from the preliminary reconstructors 339, 439 and the outputs 338, 438 of the prediction error decoders 304, 404. The preliminarily reconstructed images 314, 414 may be passed to the intra predictors 308, 408 and filters 316, 416. The filters 316, 416 that receive the preliminary representations may filter the preliminary representations and output the finally reconstructed images 340, 440, which may be stored in the reference frame memories 318, 418. The reference frame memory 318 may be connected to the inter predictor 306 to be used as a reference image for comparison in the inter prediction operation of the future base layer picture 300. Under the condition that the base layer is selected according to some embodiments and indicated as a source for inter-layer sample prediction and / or inter-layer motion information prediction for the enhancement layer, the reference frame memory 318 may also be connected to the inter predictor 406 to be used as a reference image for comparison in the inter prediction operation of the future enhancement layer picture 400. In addition, the reference frame memory 418 may be connected to the inter predictor 406 to be used as a reference image for comparison in the inter prediction operation of the future enhancement layer picture 400.
[0089] Under the condition that the base layer is selected according to some embodiments and indicated as a source for predicting the filtering parameters of the enhancement layer, the filtering parameters of the filter 316 from the first encoder portion 500 may be provided to the second encoder portion 502.
[0090] The prediction error encoders 303, 403 include transform units 342, 442 and quantizers 344, 444. The transform units 342, 442 transform the first prediction error signals 320, 420 into the transform domain. This transform is, for example, a DCT transform. The quantizers 344, 444 quantize the transform domain signals (e.g., DCT coefficients) to form quantized coefficients.
[0091] The prediction error decoders 304, 404 receive the outputs from the prediction error encoders 303, 403 and perform the opposite process of the prediction error encoders 303, 403 to generate the decoded prediction error signals 338, 438, which, when combined with the predicted representations of the image blocks 312, 412 at the second summing devices 339, 439, generate the preliminarily reconstructed images 314, 414. The prediction error decoder may be considered to include: an inverse quantizer 361, 461, which inverse quantizes the quantized coefficient values (e.g., DCT coefficients) to reconstruct the transform signal; and an inverse transform unit 363, 463, which performs an inverse transform on the reconstructed transform signal, where the output of the inverse transform unit 363, 463 contains the (multiple) reconstructed blocks. The prediction error decoder may also include a block filter, which may filter the (multiple) reconstructed blocks according to additional decoded information and filtering parameters.
[0092] Entropy encoders 330, 430 receive the outputs of prediction error encoders 303, 403 and may perform appropriate entropy coding / variable length coding on the signals to provide error detection and correction capabilities. The outputs of entropy encoders 330, 430 may be inserted into the bitstream, for example, via multiplexer 508.
[0093] Figure 3b A block diagram of a video decoder suitable for implementing an embodiment of the present invention is shown. Figure 3b The structure of a two-layer decoder is depicted, but it will be appreciated that the decoding operations may be similarly employed in a single-layer decoder.
[0094] Video decoder 550 includes a first decoder portion 552 for base layer pictures and a second decoder portion 554 for enhancement layer pictures. Block 556 illustrates a demultiplexer for delivering information about the base layer pictures to the first decoder portion 552 and for delivering information about the enhancement layer pictures to the second decoder portion 554. Reference numeral P'n represents a predicted representation of an image block. Reference numeral D'n represents a reconstructed prediction error signal. Blocks 704, 804 illustrate a preliminarily reconstructed image (I'n). Reference numeral R'n represents a final reconstructed image. Blocks 703, 803 illustrate an inverse transform (T-1). Blocks 702, 802 illustrate inverse quantization (Q-1). Blocks 700, 800 illustrate entropy decoding (E-1). Blocks 706, 806 illustrate a reference frame memory (RFM). Blocks 707, 807 illustrate prediction (P) (inter-frame prediction or intra-frame prediction). Blocks 708, 808 illustrate filtering (F). Blocks 709, 809 may be used to combine the decoded prediction error information with the predicted base or enhancement layer picture to obtain a preliminarily reconstructed image (I'n). The preliminarily reconstructed and filtered base layer picture may be output 710 from the first decoder portion 522, and the preliminarily reconstructed and filtered enhancement layer picture may be output 810 from the second decoder portion 554.
[0095] Herein, a decoder may be construed to cover any operating unit capable of performing decoding operations, such as a player, receiver, gateway, demultiplexer, and / or decoder.
[0096] The decoder reconstructs the output video by applying a prediction component similar to the encoder to form a predicted representation of the pixel block (using the motion or spatial information created by the encoder and stored in the compressed representation) and decoding the prediction error (the inverse operation of prediction error encoding to recover the quantized prediction error signal in the spatial pixel domain). After applying the prediction and prediction error decoding components, the decoder adds the prediction and prediction error signals (pixel values) together to form the output video frame. The decoder (and encoder) may also apply additional filtering components to improve the quality of the output video before delivering the output video for display and / or storing the output video as a prediction reference for upcoming frames in the video sequence.
[0097] Entropy encoding / decoding can be performed in many ways. For example, context-based encoding / decoding can be applied, where both the encoder and decoder modify the context state of the encoding parameters based on previously encoded / decoded encoding parameters. Context-based encoding can be, for example, context-adaptive binary arithmetic coding (CABAC) or context-based variable length coding (CAVLC) or any similar entropy coding. Entropy encoding / decoding can alternatively or additionally be performed using variable length coding schemes such as Huffman coding / decoding or Exp-Golomb coding / decoding. The decoding of the encoding parameters from the entropy-encoded bitstream or codewords can be referred to as parsing.
[0098] Video coding standards can specify the bitstream syntax and semantics and the decoding process for an error-free bitstream, while the encoding process may not be specified, but the encoder may only be required to generate a consistent bitstream. Bitstream and decoder compliance can be verified using a hypothetical reference decoder (HRD). The standard can include coding tools to help handle transmission errors and losses, but the use of the tools in encoding can be optional and the decoding process for an erroneous bitstream may not be specified.
[0099] Syntax elements can be defined as elements of the data represented in the bitstream. A syntax structure can be defined as zero or more syntax elements presented together in a specified order in the bitstream.
[0100] Accordingly, the basic unit for the input to the encoder and the output of the decoder is typically a picture. A picture given as the input to the encoder can also be referred to as a source picture, and a picture decoded by the decoder can be referred to as a decoded picture or a reconstructed picture.
[0101] Both the source picture and the decoded picture can include one or more sample arrays, such as one of the following set of collections of sample arrays:
[0102] - Only luminance (Y) (monochrome).
[0103] - Luminance and two chrominances (YCbCr or YCgCo).
[0104] - Green, blue, and red (GBR, also known as RGB).
[0105] - An array representing other unspecified monochromatic or trichromatic samplings (e.g., YZX, also known as XYZ).
[0106] Hereinafter, these arrays may be referred to as luminance (or L or Y) and chrominance, where the two chrominance arrays may be referred to as Cb and Cr; regardless of the actual color representation method in use. The actual color representation method in use may be indicated, for example, in an encoded bitstream using the video usability information (VUI) syntax of HEVC or the like. A component may be defined as an array or a single sample from one of the three sample arrays (luminance and two chrominances) or an array or a single sample of an array that constitutes a picture in a monochromatic format.
[0107] A picture may be defined as a frame or a field. A frame includes a matrix of luminance samples and possibly corresponding chrominance samples. When the source signal is interlaced, a field is a set of alternative sample lines of a frame and may be used as an encoder input. The chrominance sample array may be missing (and thus monochromatic sampling may be in use) or the chrominance sample array may be undersampled when compared to the luminance sample array.
[0108] Some chrominance formats may be summarized as follows:
[0109] - In monochromatic sampling, there is only one sample array, which may be nominally considered as the luminance array.
[0110] - In 4:2:0 sampling, each of the two luminance arrays has half the height and half the width of the luminance array.
[0111] - In 4:2:2 sampling, each of the two luminance arrays has the same height and half the width of the luminance array.
[0112] - In 4:4:4 sampling, when no separate color plane is in use, each of the two luminance arrays has the same height and width as the luminance array.
[0113] An encoding format or standard may allow encoding the sample arrays as separate color planes into the bitstream and decoding the separately encoded color planes from the bitstream accordingly. When separate color planes are in use, each of them is processed separately (by the encoder and / or decoder) as a picture with monochromatic sampling.
[0114] When chrominance undersampling is in use (e.g., 4:2:0 or 4:2:2 chrominance sampling), the position of chrominance samples relative to luma samples can be determined at the encoder side (e.g., as a preprocessing step or as part of encoding). The position of chrominance samples relative to the luma sample positions can be predefined, for example, in an encoding standard such as H.264 / AVC or HEVC, or can be indicated in the bitstream as part of the VUI of, for example, H.264 / AVC or HEVC.
[0115] In general, the (multiple) source video sequences provided as input for encoding may represent interlaced source content or progressive source content. Fields of opposite parity have been captured at different times for interlaced source content. Progressive source content contains captured frames. The encoder can encode fields of interlaced source content in two ways: a pair of interlaced fields can be encoded into an encoded frame or a field can be encoded as an encoded field. Similarly, the encoder can encode frames of progressive source content in two ways: a frame of progressive content can be encoded into an encoded frame or a pair of encoded fields. A field pair or complementary field pair can be defined as two fields that are close to each other in decoding and / or output order, have opposite parity (i.e., one is a top field and the other is a bottom field) and never belong to any other complementary field pair. Some video coding standards or schemes allow mixing of encoded frames and encoded fields in the same encoded video sequence. Additionally, predicting an encoded field from fields in an encoded frame and / or predicting an encoded frame for a complementary field pair (encoded as fields) can be enabled in encoding and / or decoding.
[0116] Partitioning can be defined as dividing a set into subsets such that each element of the set is in only one of the subsets.
[0117] In H.264 / AVC, a macroblock is a 16×16 luma sample block and corresponding chrominance sample blocks. For example, in the 4:2:0 sampling mode, a macroblock contains an 8×8 chrominance sample block for each chrominance component. In H.264 / AVC, a picture is partitioned into one or more slice groups, and a slice group contains one or more slices. In H.264 / AVC, a slice consists of an integer number of macroblocks that are consecutively ordered in raster scan within a particular slice group.
[0118] When describing the operations of HEVC encoding and / or decoding, the following terms may be used. An encoding block may be defined as an N×N sample block for some value of N such that the partitioning of an encoding tree block into encoding blocks is a split. A coding tree block (CTB) may be defined as an N×N sample block for some value of N such that the partitioning of a component into coding tree blocks is a split. A coding tree unit (CTU) may be defined as a coding tree block of luma samples, the corresponding coding tree blocks of chroma samples of two pictures having three sample arrays, or a coding tree block of samples of a monochrome picture or a picture encoded using three separate color planes and a syntax structure for encoding samples. A coding unit (CU) may be defined as a coding block of luma samples, the corresponding coding blocks of chroma samples of two pictures having three sample arrays, or a coding block of samples of a monochrome picture or a picture encoded using three separate color planes and a syntax structure for encoding samples.
[0119] In some video codecs, such as the High Efficiency Video Coding (HEVC) codec, a video picture may be partitioned into coding units (CUs) that cover regions of the picture. A CU includes one or more prediction units (PUs) that define a prediction process for samples within the CU and one or more transform units (TUs) that define a prediction error coding process for the samples in the CU. A CU may include a square sample block having a size that can be selected from a predefined set of possible CU sizes. A CU having the maximum allowed size may be named an LCU (largest coding unit) or a coding tree unit (CTU) and the video picture is partitioned into non-overlapping LCUs. An LCU may be further split into a combination of smaller CUs, e.g., by recursively splitting the LCU and the resulting CUs. Each resulting CU may have at least one PU and at least one TU associated therewith. Each PU and TU may be further split into smaller PUs and TUs in order to increase the granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated therewith that defines what kind of prediction is to be applied to the pixels within the PU (e.g., motion vector information for a PU for inter prediction and intra prediction directional information for a PU for intra prediction).
[0120] Each TU may be associated with information that describes a prediction error decoding process for the samples within the TU (including, e.g., DCT coefficient information). It may be signaled at the CU level whether prediction error coding is applied to each CU. If there is no prediction error residue associated with a CU, then it may be considered that there is no TU for the CU. The partitioning of the picture into CUs and the partitioning of CUs into PUs and TUs may be signaled in the bitstream, thereby allowing the decoder to reproduce the expected structure of these units.
[0121] In the draft version of H.266 / VVC, the following partitioning is applied. It should be noted that the content described herein may still evolve in later draft versions of H.266 / VVC until the standard is finalized. Similarly to HEVC, pictures are partitioned into CTUs, although the maximum CTU size has been increased to 128×128. Coding tree units (CTUs) are first partitioned by a quadtree (i.e., a four-way tree) structure. Then, the quadtree leaf nodes can be further partitioned by a multi-type tree structure. There are four splitting types in the multi-type tree structure: vertical binary split, horizontal binary split, vertical ternary split, and horizontal ternary split. The multi-type tree leaf nodes are called coding units (CUs). CUs, PUs, and TUs have the same block size, unless the CU is too large for the maximum transform length. The partitioning structure for CTUs is a quadtree with nested multi-type trees using binary and ternary splits, i.e., there is no separate CU, PU, and TU concept in use, except when needed for a CU with a size too large for the maximum transform length. A CU can have a square or rectangular shape.
[0122] The decoder reconstructs the output video by applying prediction components similar to those of the encoder to form a predicted representation of the pixel block (using motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (the inverse operation of prediction error encoding to recover the quantized prediction error signal in the spatial pixel domain). After applying the prediction and prediction error decoding components, the decoder adds the prediction and prediction error signals (pixel values) together to form the output video frame. The decoder (and encoder) can also apply additional filtering components to improve the quality of the output video before passing the output video for display and / or storing the output video as a prediction reference for upcoming frames in the video sequence.
[0123] Filtering can include, for example, one or more of the following: deblocking, sample adaptive offset (SAO), and / or adaptive loop filtering (ALF).
[0124] The deblocking loop filter can include multiple filtering modes or strengths, which can be adaptively selected based on characteristics of the blocks adjacent to the boundary (such as quantization parameter values) and / or signaling included by the encoder in the bitstream. For example, the deblocking loop filter can include a normal filtering mode and a strong filtering mode, which can differ in the number of filter taps (i.e., the number of samples being filtered on both sides of the boundary) and / or the filter tap values. For example, when omitting the potential impact of the clipping operation, filtering of two samples on both sides of the boundary can be performed using a filter with an impulse response of (3 7 9 -3) / 16.
[0125] Motion information can be indicated using motion vectors associated with each motion-compensated image block in a video codec. Each of these motion vectors represents a shift of an image block in a picture to be encoded (on the encoder side) or decoded (on the decoder side) relative to a predicted source block in one of the previously encoded or decoded pictures. To represent motion vectors efficiently, those that are different from a block-specific prediction can be encoded differently. The predicted motion vectors can be created in a predefined manner, such as by computing the median of the encoded or decoded motion vectors of neighboring blocks. Another way to create a motion vector prediction is to generate a list of candidate predictions based on neighboring blocks and / or co-located blocks in a temporally referenced picture and signal the selected candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of the previously encoded / decoded picture can be predicted. The reference index can be predicted from neighboring blocks and / or co-located blocks in a temporally referenced picture. Furthermore, an efficient video codec can employ an additional motion information encoding / decoding mechanism, often referred to as the merge mode, in which all motion field information including the motion vectors and the corresponding reference picture indices for each available reference picture list is predicted and used without any modification / correction. Similarly, the predicted motion field information is performed using the motion field information of neighboring blocks and / or co-located blocks in a temporally referenced picture, and the motion field information used is signaled among a list of motion field candidates populated with the motion field information of available neighboring / co-located blocks.
[0126] A video codec can support motion-compensated prediction from one source image (single prediction) and two sources (dual prediction). In the case of single prediction, a single motion vector is applied, while in the case of dual prediction, two motion vectors are signaled and the motion-compensated prediction from the two sources is averaged to create the final sample prediction. In the case of weighted prediction, the relative weights of the two predictions can be adjusted, or a signaled offset can be added to the prediction signal.
[0127] In addition to applying motion compensation to inter-picture prediction, a similar method can be applied to intra-picture prediction. In this case, the shift vector indicates that from the same picture, a sample block can be copied to form the prediction of the block to be encoded or decoded. This in-block copy method can significantly improve the encoding efficiency in cases where there are repetitive structures within a frame (such as text or other graphics).
[0128] The prediction residual after motion compensation or intra-picture prediction can first be transformed using a transform kernel (such as DCT) and then encoded. The reason for this is that there is usually still some correlation remaining in the residuals, and the transform can in many cases help reduce this correlation and provide more efficient encoding.
[0129] A video encoder can utilize a Lagrangian cost function to find an optimal coding mode, such as a desired macroblock mode and associated motion vectors. This cost function uses a weight factor λ to relate the (exact or estimated) image distortion due to the lossy coding method to the (exact or estimated) amount of information required to represent the pixel values in the image region:
[0130] C = D + λR (Equation 1)
[0131] where C is the Lagrangian cost to be minimized, D is the image distortion (e.g., mean squared error) for the considered mode and motion vectors, and R is the number of bits required to represent the data needed to reconstruct the image block at the decoder (including the amount of data representing the candidate motion vectors).
[0132] Some codecs use the concept of a picture order count (POC). The value of the POC is obtained for each picture and increases with increasing picture position in the output order rather than decreasing. The POC thus indicates the output order of the pictures. The POC can be used during the decoding process, for example, for implicit scaling of motion vectors and for reference picture list initialization. Additionally, the POC can be used in the verification of output order consistency.
[0133] In video coding standards, a compliant bitstream must be decodable by a hypothetical reference decoder, which can conceptually be connected to the output of the encoder and consists of at least a pre-decoder buffer, a decoder, and an output / display unit. This virtual decoder can be referred to as a hypothetical reference decoder (HRD) or a video buffer verifier (VBV). Video coding standards use variable bitrate coding, which is caused, for example, by the flexibility of the encoder to adaptively select between intra-coding techniques and inter-coding techniques for compressing video frames. To handle the fluctuations in the bitrate variations of the compressed video, buffering can be used on both the encoder and decoder sides. The HRD can be considered a hypothetical decoder model that specifies the constraints on the variability within a compliant bitstream, a compliant NAL unit stream, or a compliant byte stream that the coding process can produce. If a stream can be decoded by the HRD without buffer overflow or, in some cases, underflow, then the stream is compliant. Buffer overflow occurs when more bits are to be placed in the buffer when the buffer is full. Buffer underflow occurs when some bits are to be retrieved from the buffer for decoding / playback and those bits are not in the buffer. One of the motivations for the HRD is to avoid so-called evil bitstreams that would consume such a large amount of resources that an actual decoder implementation would not be able to handle.
[0134] If a bitstream can be decoded by an HRD without buffer overflow or, in some cases, underflow, then the bitstream can be considered compliant. Buffer overflow occurs when more bits are to be placed in the buffer when the buffer is full. Buffer underflow occurs when some bits are to be taken from the buffer for decoding / playback and the bits are not in the buffer.
[0135] The HRD can be part of the encoder or operatively connected to the output of the encoder. The buffer occupancy of the HRD and possibly other information can be used to control the encoding process. For example, if the encoded data buffer in the HRD is about to overflow, then the encoding bitrate can be decreased, for example, by increasing the quantizer step size.
[0136] The operation of the HRD can be controlled by HRD parameters such as (a) buffer size(s) and (a) initial delay(s). The HRD parameters can be classified into: CPB parameters such as (a) CPB buffer size(s) and (a) CPB initial delay(s); and DPB parameters such as (a) DPB buffer size(s) and (a) DPB initial delay(s). The HRD parameter values can be created to include or be operatively connected to part of the encoded HRD process. Alternatively, the HRD parameters can be generated separately from the encoding, for example, in an HRD validator that processes an input bitstream with a specified HRD process and generates such HRD parameter values in accordance with the bitstream being consistent therewith. Another use for the HRD validator is to verify that a given bitstream and given HRD parameters actually result in consistent HRD operation and output.
[0137] The HRD model typically includes instantaneous decoding, and the input bitrate to the HRD's coded picture buffer (CPB) can be considered a constraint on the encoder and the bitstream in terms of the decoding rate of the encoded data and a requirement on the decoder in terms of the processing rate.
[0138] On the encoder side, the CPB as specified in the HRD can be used to verify and control compliance with buffer constraints during encoding. The decoder implementation can also have a CPB that can, but does not necessarily, operate similarly or identically to the CPB specified for the HRD.
[0139] The CPB can operate on the basis of decoding units. The decoding unit can be an access unit (AU) or it can be a subset of an access unit, such as an integral number of NAL units. The choice of the decoding unit can be indicated by the encoder in the bitstream.
[0140] The decoded picture buffer (DPB) can be used in the encoder and / or the decoder. There can be two reasons for buffering decoded pictures, for reference in inter prediction and for reordering the decoded pictures into output order. Some coding formats such as HEVC provide a great deal of flexibility for both reference picture marking and output reordering, and separate buffers for reference picture buffering and output picture buffering can waste memory resources. Thus, the DPB can include a unified decoded picture buffering process for reference pictures and output reordering. Decoded pictures can be removed from the DPB when they are no longer used as references and are not needed for output. The HRD can also include the DPB. The DPB of the HRD and the decoder implementation can but need not operate the same way.
[0141] The output order can be defined as the order in which decoded pictures are output from the decoded picture buffer (for decoded pictures to be output from the decoded picture buffer).
[0142] The decoder and / or the HRD can include a picture output process. The output process can be considered as the process in which the decoder provides the decoded and cropped pictures as the output of the decoding process. The output process is typically part of a video coding standard and typically part of the hypothetical reference decoder specification. In output cropping, rows and / or columns of samples can be removed from the decoded picture according to a crop rectangle to form the output picture. The cropped decoded picture can be defined as the result of cropping the decoded picture based on, for example, a conforming crop window specified in the sequence parameter set referred to by the corresponding coded picture.
[0143] In some video coding specifications, the encoder can control the output of decoded (and cropped) pictures through the picture output process or a similar decoder-side process. The encoder can include one or more syntax elements in the bitstream or accompanying the bitstream for controlling picture output. For example, the bitstream syntax can include pic_output_flag in, for example, the picture header and / or the slice header (e.g., the slice header). The semantics of pic_output_flag can be specified in the following way: when pic_output_flag is equal to 0, the corresponding decoded picture is not output (through the picture output process or the like), and when pic_output_flag is equal to 1, the corresponding decoded picture is output, unless otherwise concluded during the decoding process.
[0144] The decoder can control whether the decoded (and cropped) picture is output through the picture output process or a similar decoder-side process. The decoder can decode one or more syntax elements for controlling picture output from the bitstream or along the bitstream. For example, the decoder can decode the pic_output_flag syntax element from, for example, the picture header and / or the slice header (e.g., slice header) of the bitstream. The decoder can obtain a variable PicOutputFlag for each picture. The obtaining of the value of PicOutputFlag can involve but is not limited to the following: If the current picture is a RASL picture (or the like) associated with an IRAP picture from which the decoding process starts or starts encoding the video sequence, then PicOutputFlag can be set to be equal to 0. Otherwise, PicOutputFlag is set to be equal to pic_output_flag. If PicOutputFlag is equal to 0 for the decoded picture, then the picture is not output through the picture output process or the like, otherwise the picture is output through the picture output process or the like.
[0145] The HRD can operate as follows, for example. The data associated with the decoded units flowing into the CPB according to the specified arrival schedule can be delivered by the Hypothetical Stream Scheduler (HSS). The arrival schedule can be determined by the encoder and indicated, for example, by the picture timing SEI message, and / or the arrival schedule can be obtained, for example, based on the bitrate that can be indicated as part of the HRD parameters in the Video Availability Information. The HRD parameters in the Video Availability Information can contain a number of parameter sets, each parameter set for a different bitrate or delivery schedule. The data associated with each decoded unit can be instantaneously removed and decoded at the CPB removal time through an instantaneous decoding process. The CPB removal time can be determined, for example, using the initial CPB buffer delay and different removal delays indicated for each picture, for example, by the picture timing SEI message. The initial CPB buffer delay can be determined by the encoder and indicated, for example, by the buffering period SEI message. The initial arrival time of the first decoded unit (i.e., the arrival time of the first bit) can be determined to be 0. The initial arrival time of any subsequent decoded unit can be determined to be equal to the final arrival time of the previous decoded unit. Each decoded picture is placed in the DPB. The decoded picture can be removed from the DPB at a later time of the DPB output time or at a time when it is no longer needed for inter-frame prediction reference. Thus, the operation of the CPB of the HRD can include the timing of the initial arrival of the decoded units (when the first bit of the decoded unit enters the CPB), the timing of the removal of the decoded units, and the decoding of the decoded units, while the operation of the DPB of the HRD can include the removal of pictures from the DPB, picture output, and the marking and storage of the decoded pictures.
[0146] The operation of AU-based coded picture buffering in HRD can be described in a simplified manner as follows. Assume that bits arrive at the CPB at a constant arrival bit rate (when the so-called low-delay mode is not in use). Thus, a coded picture or access unit is associated with an initial arrival time, which indicates when the first bit of the coded picture or access unit enters the CPB. Additionally, in the low-delay mode, a coded picture or access unit is assumed to be instantaneously removed when the last bit of the coded picture or access unit is inserted into the CPB and the corresponding decoded picture is then inserted into the DPB, thus stimulating instantaneous decoding. This time is referred to as the removal time of the coded picture or access unit. The removal time of the first coded picture of an encoded video sequence is typically controlled, for example, by buffering Supplemental Enhancement Information (SEI) messages for the period. This so-called initial coded picture removal delay ensures that any variation in the encoded bit rate relative to the constant bit rate used to fill the CPB does not cause starvation or overflow of the CPB. It should be understood that the operation of the CPB is slightly more complex than described herein, having, for example, low-delay operation modes and capabilities to operate at many different constant bit rates. Additionally, the operation of the CPB can be specified differently in different standards.
[0147] Video coding standards can specify profiles and levels. A profile can be considered a subset of the algorithmic features of the standard. Alternatively, a profile can be defined as a specified subset of the syntax of the coding standard. A level can be defined as a set of restrictions on coding parameters, which impose a set of constraints on decoder resource consumption. Alternatively, a level can be defined as a set of definitions of constraints on the values that can be taken by the syntax elements and variables of the coding standard. The same set of levels can be defined for all profiles, where most aspects of the definition of each level are common across different profiles, although aspects can also differ between profiles. Profiles and levels can be used to signal the signal attributes of a media stream and to signal the capabilities of a media decoder. Each pair of profile and level can be considered to form an "operational point".
[0148] Through the combination of profiles and levels, a decoder can declare its ability to decode a stream without actually attempting the decoding process. If a decoder is unable to decode a stream, then it can cause the decoder to crash, operate slower than real-time, and / or discard data due to buffer overflow.
[0149] The concept of tiers has been specified and used in HEVC and can be similarly specified and used in other codecs. A tier can be defined as a specified class of level constraints imposed on the values of syntax elements in a bitstream, where the level constraints are nested within the tier and a decoder conforming to a particular tier and level will be able to decode all bitstreams of the same tier or a lower tier that conform to that level or any level below it.
[0150] One or more syntax structures for (decoded) reference picture marking may be present in a video coding system. The encoder generates an instance of the syntax structure, for example, in each coded picture, and the decoder decodes an instance of the syntax structure, for example, from each coded picture. For example, decoding of the syntax structure may cause a picture to be adaptively marked as "for reference" or "not for reference".
[0151] The reference picture set (RPS) syntax structure of HEVC is an example of a syntax structure for reference picture marking. The reference picture set that is valid or active for a picture includes all reference pictures that can be used as references for that picture and all reference pictures that are kept marked as "for reference" for any subsequent pictures in decoding order. A reference picture that is kept marked as "for reference" for any subsequent pictures in decoding order but is not used as a reference for the current picture or picture segment may be considered inactive. For example, they may not be included in the (multiple) initial reference picture list.
[0152] In some coding formats and codecs, a distinction is made between so-called short-term reference pictures and long-term reference pictures. This distinction may affect some decoding processes, such as motion vector scaling. The (multiple) syntax structures for marking reference pictures may indicate to mark a picture as "for long-term reference" or "for short-term reference".
[0153] A reference picture list may be defined as a list of reference pictures for inter prediction of P or B slices. In some coding formats, the reference pictures for inter prediction may be indicated by an index to the reference picture list. In some codecs, two reference picture lists (reference picture list 0 and reference picture list 1) are generated for each bi-predicted (B) slice, and one reference picture list (reference picture list 0) is formed for each inter-coded (P) slice.
[0154] Reference picture lists such as reference picture list 0 and reference picture list 1 can be constructed in two steps: First, an initial reference picture list is generated. The initial reference picture list can be generated using an algorithm predefined in the standard. Such an algorithm can use, for example, POC and / or temporal sublayers as a basis. The algorithm can process reference pictures with (multiple) specific markers (such as "for reference") and ignore other reference pictures, i.e., avoid inserting other reference pictures into the initial reference picture list. An example of such other reference pictures is a reference picture marked as "not for reference" but still residing in the decoded picture buffer waiting to be output from the decoder. Second, the initial reference picture list can be recorded by a specific syntax structure, such as the reference picture list reordering (PRLR) command for H.264 / AVC or the reference picture list modification syntax structure for HEVC or any similar one. Additionally, the number of active reference pictures can be indicated for each list, and the use of pictures other than the active reference pictures in the list as references for inter prediction is disabled. One or both of reference picture list initialization and reference picture list modification can only process the active reference pictures among those reference pictures marked as "for reference" or the like.
[0155] In VVC, the reference picture list is directly indicated in the reference picture list syntax structure instead of indicating a reference picture set as described above and using an initialization and optional reordering process. When a picture is present in any reference picture list of the current picture (within the active or inactive entries of any reference picture list), it is marked as "for long-term reference" or "for short-term reference". When a picture is not present in the reference picture list of the current picture, it is marked as "not for reference". The abbreviation RPL can be used to refer to the reference picture list syntax structure and / or one or more reference picture lists.
[0156] Scalable video coding refers to an encoding structure in which one bitstream can contain multiple representations of the content at different bitrates, resolutions, or frame rates. In these cases, depending on its characteristics (e.g., best matching the resolution of the display device), the receiver can extract the desired representation. Alternatively, depending on, for example, the network characteristics or processing capabilities of the receiver, the server or network element can extract a portion of the bitstream to be sent to the receiver. A scalable bitstream can include a "base layer" providing the lowest available quality video and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layer. To improve the encoding efficiency for the enhancement layer, the encoded representation of the layer can depend on the lower layer. For example, the motion and mode information of the enhancement layer can be predicted from the lower layer. Similarly, the pixel data of the lower layer can be used to create a prediction for the enhancement layer.
[0157] Scalable video codecs for quality scalability (also known as signal-to-noise ratio or SNR) and / or spatial scalability can be implemented as follows. For the base layer, traditional non-scalable video encoders and decoders are used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer for the enhancement layer. In H.264 / AVC, HEVC, and similar codecs that use (multiple) reference picture lists for inter-frame prediction, similar to the decoded reference pictures of the enhancement layer, the decoded pictures of the base layer can be inserted into the (multiple) reference picture lists for the encoding / decoding of the enhancement layer pictures. Thus, the encoder can select the base layer reference picture as the inter-frame prediction reference and indicate its use, for example, using a reference picture index in the encoded bitstream. The decoder decodes from the bitstream (e.g., from the reference picture index) that the base layer picture is used as the inter-frame prediction reference for the enhancement layer. When the decoded base layer picture is used as the prediction reference for the enhancement layer, it is called an inter-layer reference picture.
[0158] Scalability modes or scalability dimensions can include but are not limited to the following:
[0159] ● Quality scalability: The base layer pictures are encoded at a lower quality than the enhancement layer pictures, which can be achieved, for example, by using a larger quantization parameter value (i.e., a larger quantization step for the transformation coefficient quantization) in the base layer than in the enhancement layer.
[0160] ● Spatial scalability: The base layer pictures are encoded at a lower resolution than the enhancement layer pictures (i.e., with fewer samples). Spatial scalability and quality scalability can sometimes be considered the same type of scalability.
[0161] ● Bit-depth scalability: The base layer pictures are encoded at a lower bit depth (e.g., 8 bits) than the enhancement layer pictures (e.g., 10 or 12 bits).
[0162] ● Dynamic range scalability: The scalable layers represent different dynamic ranges and / or images obtained using different tone mapping functions and / or different optical transfer functions.
[0163] ● Chroma format scalability: In the chroma sample array (e.g., encoded in 4:2:0 chroma format), the base layer pictures provide a lower spatial resolution than the enhancement layer pictures (e.g., 4:4:4 format).
[0164] ● Gamut scalability: The enhancement layer pictures have a richer / wider color representation range than the color representation range of the base layer pictures. For example, the enhancement layer can have a UHDTV (ITU-R BT.2020) gamut, and the base layer pictures can have an ITU-R BT.709 gamut.
[0165] ● Region of Interest (ROI) scalability: The enhancement layer represents a spatial subset of the base layer. ROI scalability can be used in conjunction with other types of scalability (e.g., quality or spatial scalability) such that the enhancement layer provides higher subjective quality for the spatial subset.
[0166] ● Viewpoint scalability, which can also be referred to as multi-view coding. The base layer represents the first viewpoint, while the enhancement layer represents the second viewpoint.
[0167] ● Depth scalability, which can also be referred to as depth-enhanced coding. The layer or some layers of the bitstream can represent the (multiple) texture viewpoints, while other layers or multiple layers can represent the (multiple) depth viewpoints.
[0168] In all of the above scalability cases, the base layer information can be used to encode the enhancement layer to minimize the additional bitrate overhead.
[0169] Scalability can be achieved in two basic ways. Either by introducing new coding modes for performing prediction of pixel values or syntax from the lower layers of the scalable representation, or by placing the lower layer pictures into the higher layer reference picture buffer (decoded picture buffer DPB). The first method is more flexible and thus can provide better coding efficiency in most cases. However, the second reference-frame-based scalability method can be implemented very efficiently with minimal changes to a single-layer codec while still achieving most of the available coding efficiency gains. Basically, a reference-frame-based scalability codec can be implemented by using the same hardware or software implementation for all layers, just responsible for DPB management by an external module.
[0170] A transmitter, gateway or the like can select the sending layer and / or sublayer of a scalable video bitstream, or similarly, a receiver, client, player or the like can request the transmission of a selected layer and / or sublayer of the scalable video bitstream. The terms layer extraction, extraction of layers or layer down-switching can refer to sending fewer layers than are available in the bitstream. Layer up-switching can refer to sending additional (multiple) layers compared to the layers sent before the layer up-switching, i.e., restarting the transmission of one or more layers that were earlier stopped in a layer down-switching. Similar to layer down-switching and / or up-switching, down-switching and / or up-switching of temporal sublayers can be performed. Both layer and sublayer down-switching and / or up-switching can be performed similarly. Layer and sublayer down-switching and / or up-switching can be performed in the same access unit or the like (i.e., basically simultaneously) or can be performed in different access units or the like (i.e., basically at different times). Layer up-switching can occur at a random access picture (e.g., an IRAP picture in HEVC). Sublayer up-switching can occur at a specific type of picture (e.g., an STSA or TSA picture in HEVC).
[0171] The following definitions can be made with respect to the High Efficiency Video Coding standard but can also apply to other codecs. An independent layer is a layer that does not have a direct reference layer, i.e., is not inter-layer predicted. A non-base layer is a layer in which all VCL NAL units have the same nuh_layer_id value greater than 0. An independent non-base layer is both an independent layer and a non-base layer.
[0172] In the following, an example of the sub-bitstream extraction process will be briefly explained. As follows, the bitstream outBitstream can be generated according to the independent non-base layer of the bitstream inBitstream. The bitstream outBitstream is set to be the same as the bitstream inBitstream. NAL units with a nal_unit_type not equal to SPS_NUT, PPS_NUT, and EOB_NUT and with a nuh_layer_id not equal to the assignedBaseLayerId are removed from outBitstream. NAL units with a nal_unit_type equal to SPS_NUT or PPS_NUT and with a nuh_layer_id not equal to 0 or the assignedBaseLayerId are removed from outBitstream. NAL units with a nal_unit_type equal to VPS_NUT are removed from outBitstream. All NAL units with a TemporalId greater than tIdTarget are removed from outBitstream. In each NAL unit of outBitstream, the nuh_layer_id is set to be equal to 0. The bitstream outBitstream can be decoded using the HEVC decoding process.
[0173] In the following, an example of the Video Parameter Set (VPS) of HEVC for indicating layer attributes will be briefly explained. The Video Parameter Set contains an extension part, and a part of it is presented below:
[0174]
[0175]
[0176] The Video Parameter Set of HEVC specifies a scalability mask that indicates the (multiple) types of scalability in use for the layer:
[0177] scalability_mask_flag[i] being equal to 1 indicates the existence of the dimension_id syntax element corresponding to the i-th scalability dimension in Table F.1. scalability_mask_flag[i] being equal to 0 indicates the non-existence of the dimension_id syntax element corresponding to the i-th scalability dimension.
[0178] Table F.1 – Mapping of ScalabiltyId to Scalability Dimensions
[0179]
[0180]
[0181] layer_id_in_nuh[i] specifies the value of the nuh_layer_id syntax element in the VCL NAL unit of the i-th layer. When i is greater than 0, layer_id_in_nuh[i] shall be greater than layer_id_in_nuh[i - 1]. For any value of i in the inclusive range of 0 to MaxLayersMinus1, when it does not exist, the value of layer_id_in_nuh[i] is presumed to be equal to i.
[0182] For i in the inclusive range from 0 to MaxLayersMinus1, the variable LayerIdxInVps[layer_id_in_nuh[i]] is set to be equal to i.
[0183] dimension_id[i][j] specifies the identifier of the j-th existing scalability dimension type in the i-th layer. The number of bits used for the representation of dimension_id[i][j] is dimension_id_len_minus1[j] + 1 bits.
[0184] Depending on splitting_flag, the following applies. If splitting_flag is equal to 1, for inclusive i from 0 to MaxLayersMinus1 and inclusive j from 0 to NumScalabilityTypes-1, dimension_id[i][j] is presumed to be equal to ((layer_id_in_nuh[i] & ((1<<dimBitOffset[j+1]) - 1)) >> dimBitOffset[j]). If splitting_flag is not equal to 1, (splitting_flag is equal to 0), then for inclusive j from 0 to NumScalabilityTypes-1, dimension_id[0][j] is presumed to be equal to 0.
[0185] The variable ScalabilityId[i][smIdx] that specifies the identifier of the smIdx-th scalability dimension type of the i-th layer and the variables DepthLayerFlag[lId], ViewOrderIdx[lId], DependencyId[lId], and AuxId[lId] that correspondingly specify the depth flag, view order index, spatial / quality scalability identifier, and auxiliary identifier of the layer with the nuh_layer_id equal to lId can be obtained as follows:
[0186]
[0187] The output layer set (OLS) can be defined as a set of layers, where one or more layers in the set are indicated as output layers. Similar to determining whether a picture is output in a single-layer bitstream as described earlier, the pictures of the output layers can be determined to be output by the decoder. The pictures that are among the layers of the OLS but not among the output layers are not output by the decoder. If multiple OLSs are indicated for a bitstream, then which OLS is used in decoding and output can be indicated to the decoder through an interface, for example. The OLS can be indicated in, for example, the VPS.
[0188] The basic unit for the output of an encoder for some coding formats (such as HEVC) and the input of a decoder for some coding formats (such as HEVC) is the network abstraction layer (NAL) unit. For transmission over a packet-oriented network or storage in a structured file, the NAL unit can be encapsulated into packets or similar structures.
[0189] For a transmission or storage environment that does not provide a frame structure, a byte stream format can be specified for the NAL unit stream. The byte stream format separates NAL units from each other by appending a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, the encoder runs a byte-oriented start code emulation prevention algorithm that adds emulation prevention bytes to the NAL unit payload in cases where a start code would otherwise have occurred. To enable simple gateway operation between packet-oriented and stream-oriented systems, start code emulation prevention can always be performed, regardless of whether the byte stream format is in use.
[0190] A NAL unit can be defined as a syntax structure that contains an indication of the type of data that follows and the bytes in the form of a raw byte sequence payload (RBSP) that contains that data, optionally interspersed with emulation prevention bytes as needed. An RBSP can be defined as a syntax structure that contains an integral number of bytes encapsulated within a NAL unit. An RBSP is either empty or in the form of a string of data bits that contains syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0191] A NAL unit consists of a header and a payload. In HEVC, a two-byte NAL unit header is used for all specified NAL unit types, while in other codecs, the NAL unit header can be similar to the NAL unit header in HEVC.
[0192] In HEVC, the NAL unit header contains one reserved bit, a six-bit NAL unit type indication, a three-bit temporal_id_plus1 indication for the temporal level or sublayer (which can be required to be greater than or equal to 1), and a six-bit nuh_layer_id syntax element.
[0193] In VVC Draft 5, the NAL unit header contains a five-bit NAL unit type indication separated into two syntax elements that are not adjacent within the NAL unit header, a three-bit temporal_id_plus1 indication for the temporal level or sublayer (which can be required to be greater than or equal to 1), a seven-bit nuh_layer_id syntax element, and one reserved bit.
[0194] In HEVC and VVC Draft 5, the syntax element temporal_id_plus1 can be considered as a temporal identifier for NAL units, and the zero-based TemporalId variable can be obtained as follows: TemporalId = temporal_id_plus1 - 1. The abbreviation TID can be used interchangeably with the TemporalId variable. A TemporalId equal to 0 corresponds to the lowest temporal level. The value of temporal_id_plus1 is required to be non-zero to avoid start code emulation involving two NAL unit header bytes. A bitstream created by excluding all VCL NAL units with a TemporalId greater than or equal to a selected value and including all other VCL NAL units remains compliant. Thus, a picture with a TemporalId equal to tid_value does not use any picture with a TemporalId greater than tid_value as an inter-prediction reference. A sub-layer or temporal sub-layer can be defined as a temporal scalability layer (or temporal layer TL) of a temporally scalable bitstream. Such a temporal scalability layer can include VCL NAL units with a specific value of the TemporalId variable and associated non-VCL NAL units.
[0195] In HEVC and VVC Draft 5, nuh_layer_id can be understood as a scalability layer identifier. In VVC Draft 5, inter-layer prediction is not enabled, i.e., all layers are independent layers.
[0196] NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units. In HEVC, VCL NAL units contain syntax elements representing one or more CUs. In HEVC, a range of NAL unit types indicates VCL NAL units, and the VCL NAL unit type indicates the picture type.
[0197] An image can be split into independently encodable image segments and independently decodable image segments (e.g., slices or tiles or groups of tiles). Such image segments can enable parallel processing. In this specification, a "slice" can refer to an image segment composed of a specific number of basic coding units processed in the default encoding or decoding order, and a "tile" can refer to an image segment that has been defined as a rectangular image region. A group of tiles can be defined as a group of one or more tiles. The image segments can be encoded as separate units in a bitstream, such as VCL NAL units in H.264 / AVC, HEVC, and VVC. The encoded image segments can include a header and a payload, where the header contains parameter values required to decode the payload. The payload of a slice can be referred to as slice data.
[0198] In the HEVC standard, a picture can be partitioned into tiles, where a tile is rectangular and contains an integer number of CTUs. In the HEVC standard, the partitioning into tiles forms a grid that can be characterized by a list of tile column widths (in CTUs) and a list of tile row heights (in CTUs). The tiles are sequentially ordered in the raster scan order of the tile grid in the bitstream. A tile can contain an integer number of slices.
[0199] In HEVC, a slice includes an integer number of CTUs. In the case where tiles are not in use, the CTUs are scanned in the raster scan order of the CTUs within a tile or within a picture. A slice can contain an integer number of slices or a slice can be contained within a tile. Within a CTU, the CUs can have a specific scan order.
[0200] In HEVC, a slice is defined as an integer number of coding tree units contained in an independent slice segment and all subsequent dependent slice segments (if any) that come before the next independent slice segment (if any) within the same access unit. In HEVC, a slice segment is defined as an integer number of coding tree units that are consecutively ordered in a tile scan and are contained within a single NAL (Network Abstraction Layer) unit. The partitioning of each picture into slice segments is a segmentation. In HEVC, an independent slice segment is defined as a slice segment for which the values of the syntax elements of the slice segment header are not inferred from the values for the previous slice segment, and a dependent slice segment is defined as a slice segment for which the values of some of the syntax elements of the slice segment header are inferred from the values for the previous independent slice segment in decoding order. In HEVC, a slice header is defined as the slice segment header of an independent slice segment that is the current slice segment or the independent slice segment that comes before the current dependent slice segment, and a slice segment header is defined as the part of the slice segment that contains the data elements encoded in relation to the first or all of the coding tree units represented in the slice segment. In the case where tiles are not in use, CUs are scanned in the raster scan order of the LCUs within the tile or picture. Within an LTU, a CU can have a specific scan order.
[0201] In VVC Draft 5, pictures are partitioned into tiles (similar to HEVC) along a tile grid. A tile is a sequence of coding tree units (CTUs) that covers a "cell" (i.e., a rectangular region of the picture) in the tile grid. A tile is partitioned into one or more bricks, each of which includes a number of CTU rows within the tile. A tile that is not partitioned into multiple bricks is also referred to as a brick. However, a brick that is a proper subset of a tile is not referred to as a tile. A slice contains a number of tiles of a picture or a number of bricks of a tile. A slice is a VCL NAL unit. Two slice modes are supported, namely the raster scan slice mode and the rectangular slice mode. In the raster scan slice mode, a slice contains a sequence of tiles in the tile raster scan of the picture. In the rectangular slice mode, a slice contains a number of bricks that together form a rectangular region of the picture. The bricks within a rectangular slice are in the order of the brick raster scan of the slice. It should be noted that the content described in this paragraph may still evolve in later drafts of H.266 / VVC until the standard is finalized.
[0202] A Motion Constrained Tile Set (MCTS) causes the inter - frame prediction process to be constrained in encoding such that sample values outside the tile set without motion constraints and sample values at fractional sample positions obtained without one or more sample values outside the tile set with motion constraints are not used for the inter - frame prediction of any sample within the motion constrained tile set. Additionally, the encoding of the MCTS is constrained in such a way that variables obtained from blocks outside the MCTS and any decoding results are not used in any decoding process within the MCTS. For example, the encoding of the MCTS is constrained in such a way that motion vector candidates are not obtained from blocks outside the MCTS. This can be enforced by turning off the temporal motion vector prediction of HEVC or by not allowing the encoder to use TMVP candidates or any motion vector prediction candidates after the TMVP candidates in the merge or AMVP candidate lists for PUs immediately to the left of the right - hand tile boundary of the MCTS, except for the last one at the bottom - right of the MCTS. Generally, an MCTS can be defined as a tile set independent of any sample values and encoded data (such as motion vectors) outside the MCTS. An MCTS sequence can be defined as a sequence of corresponding MCTSs in one or more encoded video sequences or the like. In some cases, an MCTS may be required to form a rectangular region. It should be understood that, depending on the context, an MCTS can refer to a tile set within a picture or the corresponding tile sets in a sequence of pictures. The corresponding tile sets can be but generally do not need to be co - located in the sequence of pictures. A motion constrained tile set can be considered an independently encoded tile set since it can be decoded without other tile sets.
[0203] It should be recognized that sample positions used in inter - frame prediction can be saturated such that positions that would otherwise be outside the picture are saturated to point to the corresponding boundary samples of the picture. Thus, in some usage cases, if the tile boundary is also a picture boundary, then a motion vector can actually cross the boundary or a motion vector can actually cause fractional sample interpolation that would reference positions outside the boundary, since the sample positions are saturated to the boundary. In other usage cases, specifically, if an encoded tile can be extracted from a bitstream where it is located in a position adjacent to the picture boundary of another bitstream and the tile is positioned in a location not adjacent to the picture boundary, then similar to any MCTS boundary, the encoder can constrain the motion vector to the picture boundary.
[0204] The HEVC Temporal Motion Constrained Tile Set SEI (Supplemental Enhancement Information) message can be used to indicate the presence of a motion constrained tile set in the bitstream.
[0205] Non-VCL NAL units can be, for example, one of the following types: sequence parameter set, picture parameter set, supplementary enhancement information (SEI) NAL unit, access unit delimiter, end of sequence NAL unit, end of bitstream NAL unit, or filler data NAL unit. Parameter sets may be required for the reconstruction of decoded pictures, while many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.
[0206] Some coding formats specify parameter sets that can carry parameter values and are required for the decoding or reconstruction of decoded pictures. Parameters that remain unchanged throughout the coded video sequence can be included in the sequence parameter set (SPS). In addition to parameters that may be required by the decoding process, the sequence parameter set may optionally contain video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. The picture parameter set (PPS) contains parameters that are very likely to remain unchanged across several coded pictures. The picture parameter set may include parameters that can be referenced by the coded picture segments of one or more coded pictures. The header parameter set (HPS) has been proposed to contain such parameters that can vary on a picture-by-picture basis.
[0207] The video parameter set (VPS) can be defined as a syntax structure that contains syntax elements applied to zero or more entire coded video sequences. The VPS can provide information about the layer dependencies in the bitstream, as well as many other information that can be applied to all slices across all layers in the entire coded video sequence. In HEVC, the VPS can be considered to consist of two parts, the base VPS and the VPS extension, where the VPS extension may optionally be present. The video parameter set RBSP can include parameters that can be referenced by one or more sequence parameter set RBSPs.
[0208] A parameter set can be activated when it is referenced, for example, by its identifier. For example, the header of a picture segment (such as a slice header) can contain the identifier of the PPS that is activated for decoding the coded picture containing the picture segment. The PPS can contain the identifier of the SPS that is activated when the PPS is activated. The activation of a particular type of parameter set can cause the deactivation of the previously active parameter set of the same type.
[0209] The relationships and hierarchies among VPS, SPS, and PPS can be described as follows. VPS resides at a level above SPS in the parameter set hierarchy. VPS can include parameters common to all picture segments across all layers in the entire coded video sequence. SPS includes parameters common to all picture segments in a specific layer across the entire coded video sequence and can be shared by multiple layers. PPS includes parameters common to all picture segments in the coded picture and may be shared by all picture segments in multiple coded pictures.
[0210] Instead of or in addition to parameter sets at different hierarchical levels (e.g., sequence and picture), the video coding format can include header syntax structures such as a sequence header or a picture header. In bitstream order, the sequence header can be before any other data in the coded video sequence. In bitstream order, the picture header can be before any coded video data for the picture.
[0211] Phrases such as "associated with the bitstream" (e.g., indicating associated with the bitstream) or "associated with the coded units of the bitstream" (e.g., indicating associated with the coded tiles) can be used in the claims and described embodiments to refer to the transmission, signaling, or storage in such a manner that "out-of-band" data is correspondingly associated with the bitstream or coded units but not included within the bitstream or coded units. Phrases such as "associated with the bitstream" or "associated with the coded units of the bitstream" or the like can refer to decoding the involved out-of-band data associated with the bitstream or coded units respectively (which can be obtained from out-of-band transmission, signaling, or storage). For example, the phrase "associated with the bitstream" can be used when the bitstream is contained in a container file (such as a file conforming to the ISO base media file format), and specific file metadata is stored in the file in a manner that associates the metadata with the bitstream (such as boxes in the sample entry for the track containing the bitstream, sample groups for the track containing the bitstream, or timing metadata associated with the track containing the bitstream).
[0212] A coded picture is the coded representation of a picture.
[0213] A random access point (RAP) picture (which may also be referred to as an intra random access point (IRAP) picture) can include only intra-coded picture segments. Additionally, a RAP picture can constrain subsequent pictures in the output order such that they can be correctly decoded without performing decoding processing on any pictures before the RAP picture in the decoding order. There may be pictures in the bitstream that contain only intra-coded slices of non-IRAP pictures.
[0214] In some contexts, the term random access picture may be used interchangeably with the terms RAP picture or IRAP picture. In some contexts, a RAP picture or an IRAP picture may be defined as a category of random access pictures characterized by the fact that they contain only intra-coded picture segments, while other categories of random access pictures may allow for in-picture prediction, such as intra-block copy.
[0215] In HEVC, the abbreviations of picture types may be defined as follows: Trail (TRAIL) picture, Temporal Sub-layer Access (TSA), Stepwise Temporal Sub-layer Access (STSA), Random Access Decodable Leading (RADL) picture, Random Access Skip Leading (RASL) picture, Broken Link Access (BLA) picture, Instantaneous Decoding Refresh (IDR) picture, Clean Random Access (CRA) picture. Some picture types are more fine-grained than those indicated in the table above. For example, three types of BLA pictures are specified: BLA without leading picture, BLA with decodable leading picture (i.e., without RASL picture), and BLA with any leading picture.
[0216] In HEVC, an IRAP picture may be a BLA picture, a CRA picture, or an IDR picture. In VVC Draft 5, an IRAP picture may be a CRA picture or an IDR picture.
[0217] VVC Draft 5 includes the following non-IRAP picture types similar to the corresponding picture type specifications in HEVC: TRAIL, STSA, RADL, and RASL. Additionally, a Gradual Random Access (GRA) picture is a Trail picture from which decoding can start and which guarantees the correct content of the decoded pictures at and after the indicated position (i.e., recovery point).
[0218] In HEVC and VVC Draft 5, if the necessary parameter sets are available when they need to be activated, an IRAP picture on an independent layer and all subsequent non-RASL pictures on the independent layer in decoding order can be correctly decoded without performing the decoding process of any picture before the IRAP picture in decoding order.
[0219] In HEVC and VVC Draft 5, a CRA picture can be the first picture in decoding order in the bitstream or can appear later in the bitstream. A CRA picture in HEVC allows so-called leading pictures to follow it in decoding order but before it in output order. Some leading pictures (so-called RASL pictures) can use pictures decoded before the CRA picture as references. If random access is performed on a CRA picture, the pictures following the CRA picture in decoding and output order are decodable and thus a clean random access similar to the clean random access function of an IDR picture is achieved.
[0220] A CRA picture can have associated RADL or RASL pictures. When a CRA picture is the first picture in decoding order in the bitstream, the CRA picture is the first picture of the coded video sequence in decoding order and any associated RASL pictures are not output by the decoder and may be undecodable as they can contain references to pictures not present in the bitstream.
[0221] A leading picture is a picture that is before an associated RAP picture in output order and after the associated RAP picture in decoding order. The associated RAP picture is the previous RAP picture (if any) in decoding order. In some coding specifications (such as HEVC and VVC Draft 5), a leading picture is either a RADL picture or a RASL picture.
[0222] All RASL pictures are leading pictures of an associated BLA or CRA picture. When the associated RAP picture is a BLA picture or the first coded picture in the bitstream, the RASL pictures are not output and may not be correctly decodable as they can contain references to pictures not present in the bitstream. However, if decoding starts from the RAP picture before the RASL picture's associated RAP picture, the RASL pictures can be correctly decoded. RASL pictures are not used as reference pictures for the decoding process of non-RASL pictures. When present, all RASL pictures are before all trailing pictures of the same associated RAP picture in decoding order.
[0223] All RADL pictures are leading pictures. RADL pictures are not used as reference pictures for the decoding process of trailing pictures of the same associated RAP picture. When present, all RADL pictures are before all trailing pictures of the same associated RAP picture in decoding order. In decoding order, a RADL picture does not refer to any picture before the associated RAP picture and thus can be correctly decoded when decoding starts from the associated RAP picture.
[0224] When a portion of the bitstream starting from a CRA picture is included in another bitstream, the RASL pictures associated with the CRA picture may not be decoded correctly because some of their reference pictures may not exist in the combined bitstream. To simplify the splicing operation in the HEVC bitstream, the NAL unit type of the CRA picture can be changed to indicate that it is a BLA picture. The RASL pictures associated with the BLA picture may not be decoded correctly and thus cannot be output / displayed. Additionally, the RASL pictures associated with the BLA picture can be omitted from decoding. Another way of such a splicing operation applicable to HEVC bitstreams and VVC Draft5 bitstreams is to include an end-of-sequence (EOS) NAL unit just before the CRA access unit. This causes the CRA picture to be processed as if it were the first picture of the bitstream, and the RASL pictures associated with the CRA picture are thus not output by the decoder.
[0225] In HEVC, a BLA picture can be the first picture in decoding order in the bitstream or can appear later in the bitstream. Each BLA picture starts a new coded video sequence and has a similar effect on the decoding process as an IDR picture. However, a BLA picture contains syntax elements that specify a non-empty reference picture set.
[0226] Two IDR picture types can be defined and indicated: an IDR picture without a leading picture and an IDR picture that can have an associated decodable leading picture (i.e., a RADL picture).
[0227] A trailing picture can be defined as a picture that follows an associated RAP picture in output order (and also in decoding order). Additionally, a trailing picture may not be required to be classified as any other picture type, such as a TSA or STSA picture.
[0228] In HEVC, there are two picture types, the TSA and STSA picture types that can be used to indicate temporal sub-layer switching points. VVC Draft 5 specifies an STSA picture type similar to the STSA picture type in HEVC. If temporal sub-layers with a TemporalId not exceeding N have been decoded up to, but not including, a TSA or STSA picture and the TSA or STSA picture has a TemporalId equal to N+1, then the TSA or STSA picture enables the decoding of all subsequent pictures (in decoding order) with a TemporalId equal to N+1. The TSA picture type can impose constraints on the TSA picture itself and all pictures in the same sub-layer that follow the TSA picture in decoding order. None of these pictures are allowed to use inter-prediction from any picture in the same sub-layer that precedes the TSA picture in decoding order. The TSA definition can further impose constraints on pictures in higher sub-layers that follow the TSA picture in decoding order. If a TSA picture belongs to the same or a higher sub-layer as the TSA picture, then none of these pictures are allowed to reference pictures that precede the TSA picture in decoding order. The TSA picture has a TemporalId greater than 0. The STSA is similar to the TSA picture, but does not impose constraints on pictures in higher sub-layers that follow the STSA picture in decoding order and thus only enables switching up to the sub-layer in which the STSA picture resides. In nested temporal scalability, all (trailing) pictures with a TemporalId greater than 0 can be marked as TSA pictures.
[0229] An access unit may include encoded video data for a single temporal instance and associated other data. In HEVC, an access unit (AU) can be defined as a set of NAL units that are associated with each other according to specified classification rules, are consecutive in decoding order, and contain at most one picture with any particular nuh_layer_id value. In addition to the VCL NAL units containing encoded pictures, an access unit may also contain non-VCL NAL units. The specified classification rules may, for example, associate pictures with the same output time or picture output count value to the same access unit. An access unit delimiter NAL unit may indicate the start of an access unit.
[0230] It may be required that encoded pictures appear in a specific order within an access unit. For example, within the same access unit, an encoded picture with a nuh_layer_id equal to nuhLayerIdA may be required to be before, in decoding order, an encoded picture with a nuh_layer_id greater than nuhLayerIdA.
[0231] A bitstream can be defined as a sequence of bits, which can be in the form of a NAL unit stream or a byte stream in some coding formats or standards, which forms the representation of coded pictures and the associated data forming one or more coded video sequences. The first bitstream can then be the second bitstream in the same logical channel (such as in the same file or the same connection of a communication protocol). An elementary stream (in the context of video coding) can be defined as a sequence of one or more bitstreams. In some coding formats or standards, the end of the first bitstream can be indicated by a specific NAL unit, which can be called the end-of-bitstream (EOB) NAL unit and which is the last NAL unit of the bitstream.
[0232] A coded video sequence (CVS) can be defined as a sequence of coded pictures in decoding order, which is independently decodable and then followed by another coded video sequence or the end of the bitstream. When a specific NAL unit (which can be called the end-of-sequence (EOS) NAL unit) appears in the bitstream, the coded video sequence can be additionally or alternatively specified as ended. In HEVC, the EOB NAL unit with nuh_layer_id equal to 0 ends the coded video sequence.
[0233] A bitstream or a coded video sequence can be coded to be temporally scalable as follows. Each picture can be assigned to a specific temporal sublayer. The temporal sublayers can be enumerated, for example, from 0 upwards. The lowest temporal sublayer (sublayer 0) can be decoded independently. Pictures at temporal sublayer 1 can be predicted from the reconstructed pictures at temporal sublayers 0 and 1. Pictures at temporal sublayer 2 can be predicted from the reconstructed pictures at temporal sublayers 0, 1, and 2, and so on. In other words, pictures at temporal sublayer N do not use any pictures at temporal sublayers greater than N as references for inter-frame prediction. The bitstream created by excluding all pictures with a sublayer value greater than or equal to the selected sublayer value and including the pictures still conforms.
[0234] Sub-layer access pictures can be defined as pictures from which the decoding of a sub-layer can be correctly started, i.e., all pictures of a sub-layer can be correctly decoded starting from it. In HEVC, there are two picture types, the Temporal Sub-layer Access (TSA) and the Stepwise Temporal Sub-layer Access (STSA) picture types that can be used to indicate the temporal sub-layer switching points. If temporal sub-layers with TemporalId not exceeding N have been decoded up to but not including the TSA or STSA picture and the TSA or STSA picture has a TemporalId equal to N + 1, then the TSA or STSA picture enables the decoding of all subsequent pictures (in decoding order) with TemporalId equal to N + 1. The TSA picture type can impose constraints on the TSA picture itself and all pictures in the same sub-layer that follow the TSA picture in decoding order. None of these pictures are allowed to use inter-prediction from any picture in the same sub-layer that precedes the TSA picture in decoding order. The TSA definition can further impose constraints on pictures in higher sub-layers that follow the TSA picture in decoding order. If a TSA picture belongs to the same or a higher sub-layer as the TSA picture, none of these pictures are allowed to reference pictures that precede the TSA picture in decoding order. The TSA picture has a TemporalId greater than 0. The STSA is similar to the TSA picture, but does not impose constraints on pictures in higher sub-layers that follow the STSA picture in decoding order and thus only enables switching up to the sub-layer in which the STSA picture resides.
[0235] A video shot can be defined as a set of temporally consecutive pictures captured using a camera or otherwise related (e.g., computer-generated animated scenes). Metadata on the shot can accompany the captured content, or the shot can be detected from a video sequence that includes several shots in a time-interleaved manner. Shot detection methods operating on uncompressed video content can include, for example, pairwise pixel comparison, block-based comparison, histogram comparison, and segmentation followed by segment-based comparison.
[0236] When describing H.264 / AVC, HEVC, VCC, and example embodiments, the following description can be used to specify the parsing process of each syntax element.
[0237] -u(n): Unsigned integer using n bits. When n is "v" in the syntax table, the number of bits varies in a manner depending on the value of other syntax elements. The pairing process for this descriptor is specified by the n next bits of the binary representation of an unsigned integer read from the bitstream and written first as the most significant bit.
[0238] -ue(v): An unsigned integer Exponential-Golomb coded (i.e., exp-Golomb coded) syntax element having the leftmost bit first.
[0239] An Exponential-Golomb bit string can be converted to a code number (codeNum) using, for example, the following table:
[0240] bit string codeNum 1 0 0 1 0 1 0 1 1 2 0 0 1 0 0 3 0 0 1 0 1 4 0 0 1 1 0 5 0 0 1 1 1 6 0 0 0 1 0 0 0 7 0 0 0 1 0 0 1 8 0 0 0 1 0 1 0 9 … …
[0241] Available media file format standards include the ISO Base Media File Format (ISO / IEC 14496-12, which can be abbreviated as ISOMFF), the MPEG-4 file format (ISO / IEC 14496-14, also known as the MP4 format), the file format for NAL unit structured video (ISO / IEC 14496-15), and the 3GPP file format (3GPP TS26.244, also known as the 3GP format). The ISO file format is the derived basis for all of the above file formats (excluding the ISO file format itself). These file formats (including the ISO file format itself) are generally referred to as the ISO family of file formats.
[0242] Some concepts, structures, and specifications of ISOBMFF are described below as an example of a container file format on which embodiments can be implemented. Aspects of the present invention are not limited to ISOBMFF, but rather a possible basis is described on which the present invention can be implemented in part or in whole.
[0243] The basic building blocks of the ISO Base Media File Format are called boxes. Each box has a header and a payload. The box header indicates the type and size of the box in bytes. A box can enclose other boxes, and the ISO file format specifies which box types are allowed within a certain type of box. Additionally, in each file, the presence of some boxes can be mandatory, while the presence of other boxes can be optional. Additionally, for some box types, more than one box can be allowed to exist in the file. Thus, the ISO Base Media File Format can be considered to specify a hierarchy of boxes.
[0244] According to the ISO family of file formats, a file includes media data and metadata encapsulated in boxes. Each box is identified by a four-character code (4CC) and starts with a header that tells the type and size of the box.
[0245] In a file conforming to the ISO base media file format, media data can be provided in one or more instances of a MediaDataBox (“mdat”), and a MovieBox (“moov”) can be used to enclose metadata for timed media. In some cases, for the file to be operable, both the “mdat” box and the “moov” box may be required to be present. The “moov” box can include one or more tracks, and each track can reside in a corresponding track box (“trak”). Each track is associated with a handle of a specified track type identified by a four-character code. Video, audio, and image sequence tracks can be collectively referred to as media tracks, and they contain elementary media streams. Other track types include hint tracks and timed metadata tracks. A track includes samples, such as audio or video frames. For a video track, a media sample can correspond to an encoded picture or access unit. A media track refers to samples (which can also be referred to as media samples) formatted according to a media compression format (and its encapsulation for the ISO base media file format). A hint track refers to hint samples that contain menu instructions for constructing packets for transmission over an indicated communication protocol. A timed metadata track can refer to samples and / or hint samples that describe the media involved.
[0246] The “trak” box includes, in its box hierarchy, a SampleDescriptionBox that gives detailed information about the encoding type used and any initialization information required by that encoding. The SampleDescriptionBox contains a number of entries and as many sample entries as indicated by the number of entries. The format of a sample entry is track type-specific but derived from a common class (e.g., VisualSampleEntry, AudioSampleEntry). Which type of sample entry form is used for deriving the track type-specific sample entry format is determined by the media handle of the track.
[0247] Movie fragments can be used, for example, when recording content to an ISO file, for example, to avoid data loss if the recording application crashes, the memory space is exhausted, or some other event occurs. In the absence of movie fragments, data loss may occur because the file format may require all metadata (e.g., movie boxes) to be written to a contiguous region of the file. Additionally, when recording a file, there may not be enough memory space (e.g., random access memory RAM) to buffer the movie boxes for the size of the available storage device, and it may be too slow to recalculate the content of the movie boxes when the movie is closed. Additionally, movie fragments can enable simultaneous recording and playback of the file using a conventional ISO file parser. Additionally, a shorter initial buffering duration may be required for progressive download, e.g., when movie fragments are used, simultaneous reception and playback of the file, and the initial movie box is smaller compared to a file with the same media content but not structured with movie fragments.
[0248] The movie fragment feature can enable metadata that might otherwise reside in a movie box to be split into multiple chunks. Each chunk can correspond to a certain time period of the track. In other words, the movie fragment feature enables file metadata and media data to intersect. As a result, the size of the movie box can be restricted, and the above use cases can be implemented.
[0249] In some examples, if the media samples for a movie fragment are in the same file as the moov box, they can reside in the mdat box. However, for the metadata of the movie fragment, a moof box can be provided. The moof box can include information for a certain duration of playback time that should have previously been in the moov box. The moov box itself may still represent a valid movie, but additionally, it can include an mvex box that indicates that movie fragments will follow in the same file. Movie fragments can extend the presentation associated with the moov box in real time.
[0250] Within a movie fragment, there can be a set of track fragments, including any position from zero to multiple for each track. A track fragment in turn can include any position from zero to multiple track runs (i.e., track fragment runs), where each track run file of a track run is a consecutive run of samples for that track. Within these structures, many fields are optional and can be default. The metadata that can be included in the moof box can be limited to a subset of the metadata that can be included in the moov box and can be encoded differently in some cases. Details about the boxes that can be included in the moof box can be found in the ISO Base Media File Format specification. An independent movie fragment can be defined as containing moof boxes and mdat boxes in file order continuity, where the mdat box contains the samples of the movie fragment (for which the moof box provides metadata) and does not contain samples of any other movie fragment (i.e., any other moof box).
[0251] Track reference mechanisms can be used to correlate tracks with each other. The TrackReferenceBox includes one or more boxes, and each box of the boxes provides a reference from the containing track to other sets of tracks. These references are labeled by the box type (i.e., the four-character code of the box) of the one or more contained boxes.
[0252] The TrackGroupBox contained within the TrackBox can indicate track groups, where each group shares a specific characteristic or the tracks within the group have a specific relationship. The box contains zero or more boxes, and the specific characteristic or relationship is represented by the box type of the contained boxes. The contained boxes include identifiers that can be used to infer the tracks belonging to the same track group. Tracks that contain the same type of contained boxes within the TrackGroupBox and have the same identifier value within these contained boxes belong to the same track group.
[0253] The ISOBMFF format includes three mechanisms for timing metadata that can be associated with specific samples: sample groups, timing metadata tracks, and sample auxiliary information. The resulting specification can use one or more of these three mechanisms to provide similar functionality.
[0254] Sample grouping in the ISO base media file format and its derivatives (such as ISO / IEC 14496-15 (Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format)) can be defined as assigning each sample in a track to a member of a sample group based on a grouping criterion. The sample groups in sample grouping are not limited to being consecutive samples and can include non-adjacent samples. Since there can be more than one sample grouping for the samples in a track, each sample grouping can have a type field to indicate the type of the grouping. Sample grouping can be represented by two linked data structures: (1) a SampleToGroupBox (sbgp box) that represents the assignment of samples to sample groups; and (2) a SampleGroupDescriptionBox (sgpd box) that contains sample group entries for each sample group, describing the attributes of the group. Based on different grouping criteria, there can be multiple instances of SampleToGroupBox and SampleGroupDescriptionBox. These can be distinguished by a type field used to indicate the grouping type. The SampleToGroupBox can include a grouping_type_parameter field, which can be used, for example, to indicate the subtype of the grouping.
[0255] A Uniform Resource Identifier (URI) can be defined as a string used to identify a resource name. This identification enables interaction with a representation of the resource over a network using a specific protocol. URIs are defined by schemes that specify a specific syntax and an associated protocol. A Uniform Resource Locator (URL) and a Uniform Resource Name (URN) are forms of URIs. A URL can be defined as a URI that identifies a web resource and specifies how to act on the resource or obtain a representation of the resource, specifying both its primary access mechanism and its network location. A URN can be defined as a URI that identifies a resource by name within a particular namespace. A URN can be used to identify a resource without implying its location or access method.
[0256] Recently, the Hypertext Transfer Protocol (HTTP) has been widely used to deliver real-time multimedia content (such as in video streaming applications) over the Internet. Different from using the Real-Time Transport Protocol (RTP) over the User Datagram Protocol (UDP), HTTP is easy to configure and is typically authorized to traverse firewalls and Network Address Translation (NAT) programs, which makes it attractive for multimedia streaming applications.
[0257] Such as Smooth Streaming, Adaptive HTTP live streaming and Several commercial solutions for adaptive streaming over HTTP for both adaptive HTTP live streaming and dynamic streaming have come to the market, and standardization projects have also been launched. Adaptive HTTP streaming (AHS) was first standardized in Release 9 of the 3rd Generation Partnership Project (3GPP) Packet Switched Streaming (PSS) service (3GPP TS 26.234 Release 9: "Transparent end-to-end packet switched streaming service (PSS); Protocols and codecs"). MPEG took 3GPP AHS Release 9 as the starting point for the MPEG DASH standard (ISO / IEC 23009-1: "Dynamic adaptive streaming over HTTP (DASH) Part 1: Media presentation description and segment formats", International Standard, 2nd Edition, 2014). 3GPP continued to work on adapting HTTP streaming for communication with MPEG and issued 3GP-DASH (Dynamic Adaptive Streaming over HTTP; 3GPP TS 26.247: "Transparent end-to-end packet switched streaming service (PSS); Progressive download and dynamic adaptive streaming over HTTP (3GP-DASH)"). MPEG DASH and 3GP-DASH are technically close to each other and can thus be collectively referred to as DASH. Streaming systems similar to MPEG DASH include, for example, HTTP Live Streaming as specified in IETF RFC 8216 (i.e., HLS). Some concepts, formats, and operations of DASH are described below as an example of a video streaming system in which embodiments can be implemented. Aspects of the present invention are not limited to DASH, but are described as a possible basis on which the present invention can be implemented in part or in whole.
[0258] In DASH, multimedia content can be stored on an HTTP server and delivered using HTTP. The content can be stored on the server in two parts: the Media Presentation Description (MPD), which describes the inventory of available content, its various alternatives, its URL addresses, and other characteristics; and the segments, which contain the actual multimedia bitstreams in the form of chunks in single or multiple files. The MDP provides the necessary information for the client via HTTP to establish dynamic adaptive streaming. The MPD contains information describing the media presentation, such as the HTTP Uniform Resource Locator (URL) of each segment to make GET segment requests. To play the content, a DASH client can obtain the MPD by, for example, using HTTP, email, thumb drive, broadcast, or other transport methods. By parsing the MPD, the DASH client can become aware of the program timing, media content availability, media type, resolution, minimum and maximum bandwidths, the presence of various encoded alternatives of the multimedia components, accessibility features, and the required Digital Rights Management (DRM), the location of the media components on the network, and other content characteristics. The DASH client can use this information to select a suitable encoded alternative and start streaming the content by obtaining segments using, for example, HTTP GET requests. After appropriate buffering to allow for network throughput variations, the client can continue to obtain subsequent segments and also monitor network bandwidth fluctuations. The client can decide how to adapt to the available bandwidth by obtaining segments of different alternatives (with lower or higher bitrates) to maintain sufficient buffering.
[0259] In DASH, a hierarchical data model is used to construct the media presentation as follows. The media presentation includes one or more periods of a sequence, each period containing one or more groups, each group containing one or more adaptation sets, each adaptation set containing one or more representations, and each representation including one or more segments. A representation is one of the alternative choices of the media content or a subset thereof, typically differing due to encoding choices (e.g., bitrate, resolution, language, codec, etc.). A segment contains media data of a specific duration, as well as metadata for decoding and presenting the contained media content. A segment is identified by a URI and can generally be requested by an HTTP GET request. A segment can be defined as a data unit associated with an HTTP URL and optionally can be defined as a byte range specified by the MPD.
[0260] The DASH MPD conforms to the Extensible Markup Language (XML) and is thus defined by elements and attributes as defined in XML.
[0261] In DASH, all descriptor elements are structured in the same way, i.e., they contain the @schemeIdUri attribute that provides a URI to identify the scheme, and the optional attributes @value and @id. The semantics of the element are specific to the scheme adopted. The URI identifying the scheme can be a URN or a URL.
[0262] In DASH, an independent representation can be defined as a representation that can be processed independently of any other representation. An independent representation can be understood to include an independent bitstream or an independent layer of a bitstream. A dependent representation can be defined as a representation for which segments from its complementary representation are necessary for the presentation and / or decoding of the media content components it contains. A dependent representation can be understood to include, for example, the predictive layers of a scalable bitstream. A complementary representation can be defined as a representation that complements at least one dependent representation. A complementary representation can be an independent representation or a dependent representation. A dependent representation can be described by a Representation element that contains the @dependencyId attribute. Dependent representations can be considered regular representations, except that they depend on a set of complementary representations for decoding and / or presentation. The @dependencyId contains the values of the @id attributes of all complementary representations (i.e., the representations necessary for presenting and / or decoding the media content components contained in that dependent representation).
[0263] The track references of ISOBMFF can be reflected in a list of four-character codes in the @associationType attribute of the DASH MPD that is mapped in a one-to-one manner to the list of @id values of the Representations given in @associationId. These attributes can be used to link media representations to metadata representations.
[0264] DASH services can be provided as on-demand services or live services. In the former, the MPD is static and all segments of the media presentation are available when the content provider publishes the MPD. However, in the latter, the MPD can be static or dynamic depending on the segment URL construction method adopted by the MPD, and segments are continuously created as the content is generated and published by the content provider to the DASH client. The segment URL construction method can be a template-based segment URL construction method or a segment list generation method. In the former, the DASH client is able to construct segment URLs without updating the MPD before requesting segments. In the latter, the DASH client must periodically download updated MPDs to obtain segment URLs. For live services, therefore, the template-based segment URL construction method is preferred over the segment list generation method.
[0265] An initialization segment can be defined as a segment that contains metadata necessary to present the media stream in the wrapped media segments. In the ISOBMFF-based segment format, the initialization segment can include a movie box (“moov”), which may not include metadata for any samples, i.e., any metadata for samples is provided in the “moof” box.
[0266] A media segment contains a specific duration of media data to be played back at normal speed, which is referred to as the media segment duration or segment duration. The content producer or service provider can select the segment duration according to the desired characteristics of the service. For example, a relatively short segment duration can be used in a live service to achieve a shorter end-to-end delay. The reason is that the segment duration is typically the lower bound of the end-to-end delay perceived by the DASH client, since segments are discrete units for generating media data in DASH. Content generation is usually done in such a way that the entire media data segment is available to the server. Additionally, many client implementations use segments as units for GET requests. Therefore, in a typical arrangement for a live service, the DASH client can only request segments when the entire duration of the media segment is available and has been encoded and encapsulated into the segment. For on-demand services, different strategies for selecting the segment duration can be used.
[0267] A segment can be further divided into multiple sub-segments, e.g., to allow the segment to be downloaded in multiple parts. The sub-segments may need to contain complete access units. The sub-segments can be indexed by a segment index box (i.e., SegmentIndexBox), which contains information mapping the presentation time range and byte range for each sub-segment. The segment index box can also describe the sub-segments and stream access points in the segment by signaling their durations and byte offsets. The DASH client can use the information obtained from the (multiple) segment index boxes to make an HTTP GET request for a specific sub-segment using a byte range HTTP request. If a relatively long segment duration is used, sub-segments can be used to keep the size of the HTTP response reasonable and for flexibility in bitrate adaptation. The index information for the segment can be placed in a single box at the beginning of the segment or spread across multiple index boxes in the segment. Different spreading methods are possible, such as hierarchical spreading, daisy-chain spreading, and hybrid spreading. This technique can avoid adding a large box at the beginning of the segment and thus can prevent possible initial download delays.
[0268] SegmentIndexBox can have the following syntax:
[0269]
[0270] The semantics of some syntax elements of SegmentIndexBox can be specified as follows.
[0271] reference_type: When set to 1, it indicates a reference to a SegmentIndexBox; otherwise, it references media content (e.g., in the case of a file based on this document, it references a MovieFragmentBox); if separate index segments are used, then entries with reference type 1 are in the index segment, and entries with reference type 0 are in the media file.
[0272] referenced_size: The byte distance from the first byte of the referenced item to the first byte of the next referenced item or to the end of the referenced material in the case of the last entry.
[0273] The term segment index can be defined as a compact index of time ranges to byte ranges mapped separately within media segments from the MPD. The segment index can include one or more SegmentIndexBoxes.
[0274] The symbol (sub) segment refers to a segment or a sub - segment. If the segment index box does not exist, then the symbol (sub) segment refers to a segment. If the segment index box exists, then the symbol (sub) segment can refer to a segment or a sub - segment, e.g., depending on whether the client requests on a segment basis or on a sub - segment basis.
[0275] MPEG - DASH defines a segment container format for both the ISO base media file format and the MPEG - 2 transport stream. Other specifications can specify segment formats based on other container formats. For example, a segment format based on the Matroska container file format has been proposed.
[0276] Sub - representations are implemented within a regular representation and are described by sub - representation elements. Sub - representation elements are contained within representation elements. Sub - representation elements describe the attributes of one or several media content components embedded within the representation. It can, for example, describe the exact attributes of an implemented audio component (such as codec, sampling rate, etc.), an implemented subtitle (such as codec), or it can describe some implemented lower - quality video layers (such as some lower frame rates or others). Sub - representations and representations share some common attributes and elements.
[0277] In the case where the @level attribute exists in the sub - representation element, the following applies:
[0278] Sub - representations provide the ability to access lower - quality versions of the representations in which they are contained. In this case, sub - representations, for example, allow extraction of audio tracks from a multiplexed representation or can allow efficient fast - forward or fast - reverse operations when provided with a lower frame rate;
[0279] The initialization segment and / or media segment and / or index segment shall provide sufficient information such that the data can be easily accessed via an HTTP partial GET request. Details regarding the provision of such information are defined by the media format in use.
[0280] When ISOBMFF segments are used for a representation that includes sub - representations, the following applies:
[0281] The initialization segment contains a level assignment box.
[0282] A sub - segment index box (“ssix”) exists for each sub - segment.
[0283] The attribute @level specifies the level to which the sub - representation described in the sub - segment index is associated. Information in the representation, sub - representation, and level assignment (“leva”) box contains information about the assignment of media data to levels.
[0284] The media data shall be ordered such that each level provides an enhancement over lower levels.
[0285] If the @level attribute is absent, then the sub - representation element is only used to provide a more detailed description of the media stream in the embedded representation.
[0286] ISOBMFF includes so-called levels to specify subsets of a file. The levels follow a dependency hierarchy such that samples mapped to level n can depend on any samples of level m, where m <= n, and do not depend on any samples of level p, where p > n. For example, levels can be specified according to temporal sublayers (e.g., TemporalId of HEVC). Levels can be advertised in a Level Assignment ("leva") box (i.e., LevelAssignmentBox) contained in the Movie Extensions ("mvex") box. Levels cannot be specified for the initial movie. When the Level Assignment box is present, it applies to all movie fragment segments after the initial movie. For the context of the Level Assignment box, a fragment is defined as including one or more movie fragment boxes and associated media data boxes, possibly including only the initial part of the last media data box. Within a fragment, data for each level appears continuously. Data for levels within a fragment appears in increasing order of level value. All data in a fragment is assigned to a level. The Level Assignment box provides a mapping from features such as scalability layers or temporal sublayers to levels. Features can be specified by tracks, sub-tracks within a track, or sample groupings of a track. For example, temporal level sample groupings can be used to indicate the mapping of pictures to temporal levels (which is equivalent to temporal sublayers in HEVC). That is, HEVC pictures with a specific TemporalId value can be mapped to a specific temporal level using temporal level sample groupings (and this can be repeated for all TemporalId values). The Level Assignment box can then refer to the temporal level sample groupings in the indicated mapping to levels.
[0287] The Subsegment Index Box ("ssix", i.e., SubsegmentIndexBox) provides a mapping from levels (such as specified by the Horizontal Assignment Box) to byte ranges of subsegments of the index. In other words, this box provides a compact index for how the data in the subsegment is sorted into partial subsegments according to the levels. It enables clients to easily access data for partial subsegments by downloading ranges of data in the subsegment. When a Subsegment Index Box exists, each byte in the subsegment is assigned to a level. If the range is not associated with any information in the horizontal assignment, any level not included in the horizontal assignment can be used. There are 0 or 1 Subsegment Index Boxes per Subsegment Index Box for only leaf subsegments (i.e., only index subsegments without segment indexes) according to the index. The Subsegment Index Box (if any) is the next box after the associated Subsegment Index Box. The Subsegment Index Box records the subsegment indicated in the immediately preceding Subsegment Index Box. Each level can be assigned to exactly one partial subsegment, i.e., the byte range for one level is continuous. The levels of the partial subsegments are assigned by increasing numbers within the subsegment, i.e., the samples of a partial subsegment can depend on any samples of the preceding partial subsegments within the same subsegment, but not vice versa. For example, each partial subsegment contains samples with the same temporal sublayer and the partial subsegments appear in increasing temporal sublayer order within the subsegment. When the partial subsegments are accessed in this way, the final Media Data Box may be incomplete, i.e., less data than indicated by the length of the Media Data Box is present. The length of the Media Data Box may need to be adjusted, or padding may be used. The padding_flag in the Horizontal Assignment Box indicates that the missing data can be replaced by zeros. If not, then the sample data for samples assigned to levels that are not accessed does not exist, and care should be taken.
[0288] DASH supports rate adaptation to match changing network bandwidths by dynamically requesting media segments of different representations within an adaptation set. When a DASH client switches up / down representations, the coding dependencies within the representation must be taken into account. Representation switches can occur at random access points (RAPs), which are commonly used in video coding technologies such as H.264 / AVC. In DASH, a more general concept named stream access point (SAP) is introduced to provide a codec-independent solution for accessing representations and switching between representations. In DASH, an SAP is specified as a position within a representation such that the playback of the media stream can be started using only the information contained in the representation data starting from that position (after initializing the data in the initialization segment, if any). Thus, representation switches can be performed at an SAP.
[0289] In DASH, the automated selection between representations in the same adaptation set has been performed based on, for example, but not limited to, one or more of the following: width and height (@width and @height); frame rate (@frameRate); bit rate (@bandwidth); the indicated quality ranking between representations (@qualityRanking). The semantics of @qualityRanking are specified as follows: Specify the quality ranking of representations in the same adaptation set relative to other presentations. A lower value indicates higher quality content. If not present, then no ranking is defined.
[0290] Several types of SAPs have been specified, including the following. SAP type 1 corresponds to what is referred to in some coding schemes as a "closed GOP random access point" (where all pictures in decoding order can be correctly decoded to yield a continuous time series of correctly decoded pictures without gaps), and furthermore, the first picture in decoding order is also the first picture in presentation order. SAP type 2 corresponds to what is referred to in some coding schemes as a "closed GOP random access point" (where all pictures in decoding order can be correctly decoded to yield a continuous time series of correctly decoded pictures without gaps), for which the first picture in decoding order may not be the first picture in presentation order. SAP type 3 corresponds to what is referred to in some coding schemes as an "open GOP random access point" (where all pictures in decoding order can be correctly decoded to yield a continuous time series of correctly decoded pictures without gaps), where there may be some pictures in decoding order that cannot be correctly decoded and have less presentation time than the intra-coded pictures associated with the SAP.
[0291] In some video coding standards such as MPEG-2, each intra picture has been a random access point within the coding sequence. In some video coding standards such as H.264 / AVC and H.265 / HEVC, the ability to flexibly use multiple reference pictures for inter prediction has the consequence that intra pictures may not be sufficient for random access. Thus, pictures can be marked with respect to their random access point functionality rather than inferring such functionality from the coding type; for example, IDR pictures as specified in the H.264 / AVC standard can be used as random access points. A group of pictures (GOP) is a set of pictures where all pictures can be correctly decoded. For example, in H.264 / AVC, a closed GOP can start from an IDR access unit.
[0292] A group of pictures (GOP) is a group of pictures in which pictures before the initial intra picture in output order may not be correctly decoded, but pictures after the initial intra picture in output order can be correctly decoded. Such an initial intra picture can be indicated in the bitstream and / or inferred from indications from the bitstream, e.g., by using the CRA NAL unit type in HEVC. Pictures before the initial intra picture that starts an open GOP in output order and after the initial intra picture in decoding order can be referred to as leading pictures. There are two types of leading pictures: decodable and non - decodable. Decodable leading pictures (such as RADL pictures in HEVC) enable correct decoding when decoding starts from the initial intra picture that starts an open GOP. In other words, decodable leading pictures use only the initial intra picture or subsequent pictures in decoding order as references in inter - prediction. Non - decodable leading pictures (such as RASL pictures in HEVC) do not enable correct decoding when decoding starts from the initial intra picture that starts an open GOP.
[0293] A stream access point (SAP) sample group as specified in ISOBMFF identifies samples as being of the indicated SAP type. The grouping_type_parameter for the SAP sample group includes the fields target_layers and layer_id_method_idc. target_layers specifies the target layer for the indicated SAP. The semantics of target_layers can depend on the value of layer_id_method_idc. layer_id_method_idc specifies the semantics of target_layers. A layer_id_method_idc equal to 0 specifies that the target layer includes all layers represented by the track. The sample group description entry for the SAP sample group includes the fields dependent_flag and SAP_type. For non - layered media, dependent_flag may be required to be 0. A dependent_flag equal to 1 specifies that the reference layer (if any) used to predict the target layer may have to be decoded to access the samples of the sample group. A dependent_flag equal to 0 specifies that the reference layer (if any) used to predict the target layer does not need to be decoded to access any SAP of the sample group. Sap_type values in the range from 1 to 6 (inclusive) specify the SAP type of the associated samples.
[0294] Synchronization samples can be defined as samples in a track that are SAPs of type 1 or 2. Synchronization samples can be signaled using a SyncSampleBox or by a sample_is_non_sync_sample equal to 0 in the signaling for a track fragment.
[0295] Implement a compact formation of a track that extracts NAL unit data by reference for the extractors specified in ISO / IEC 14496-15 for H.264 / AVC and HEVC. The extractor is a NAL unit-like structure. Like any NAL unit, the NAL unit-like structure can be specified to include a NAL unit header and a NAL unit payload, but the start code emulation prevention (required for NAL units) may not be followed in the NAL unit-like structure. For HEVC, the extractor contains one or more builders. The sample builder extracts NAL unit data by reference from samples of another track. The internal builder includes the NAL unit data. The term inline can be defined, for example, in association with a data unit to indicate that an inclusion syntax structure includes or carries the data unit (as opposed to including the data unit by reference or by a data pointer). When the extractor is processed by a file reader that requires it, the extractor is logically replaced by the bytes obtained when parsing the included builders in their order of appearance. Nested extractions may not be allowed; for example, the bytes referenced by the sample builder will not contain an extractor; an extractor will not reference another extractor directly or indirectly. The extractor can contain one or more builders for extracting data from the current track or from another track that is linked to the track in which the extractor resides by means of a track reference of type "scal". The bytes of the parsed extractor can represent one or more complete NAL units. The parsed extractor starts with a valid length field and a NAL unit header. The bytes of the sample builder are copied only from a single identified sample in the track referenced by the indicated "scal" track reference. Alignment is at decode time, i.e., only using the time-to-sample table, followed by a count offset in the number of samples. The extractor is a media-level concept and is thus applied to the destination track before any edit list is considered. (However, it will generally be expected that the edit lists in both tracks will be the same).
[0296] H.264 / AVC does not include the concept of tiles, but operations like MCTS can be implemented by vertically arranging regions as slices and restricting the encoding similar to that of MCTS. For simplicity, the terms tile and MCTS are used in this document but should be understood to apply to H.264 / AVC in a restricted manner. In general, the terms tile and MCTS should be understood to apply to similar concepts in any coding format or specification.
[0297] A tile base track (e.g., an HEVC tile base track) can be generated and stored in a file. The tile base track represents the bitstream by implicitly gathering a set of tiles with motion constraints from the tile tracks. The tile base track can include a track reference to the tile track, and / or the tile track can include a track reference to the tile base track. For example, in HEVC, the "sabt" track reference is used to refer to the tile track from the tile base track, and the tile ordering is indicated by the order of the tile tracks contained by the "sabt" track reference. Additionally, in HEVC, the tile track has a "tbas" track reference to the tile base track.
[0298] The extractor enables a compact formation of a track that extracts NAL unit data by reference. The extractor contains one or more builders, such as 1) a sample builder that extracts NAL unit data from samples of another track by reference, and 2) an internal builder that includes NAL unit data.
[0299] When the extractor is processed by a file reader that requires it, the extractor is logically replaced by the bytes obtained when parsing the contained builders in their order of appearance. In some embodiments, nested extractions may not be allowed, e.g., the bytes referenced by the sample builder may not contain an extractor; an extractor may not directly or indirectly reference another extractor. The extractor can contain one or more builders for extracting data from the current track or from another track that is linked to the track in which the extractor resides by means of a track reference of type "scal".
[0300] In an example, the bytes of the parsed extractor are one of the following:
[0301] - A complete NAL unit; note that when an integrator is referenced, both the included bytes and the referenced bytes are copied
[0302] - More than one complete NAL unit
[0303] In both cases, the bytes of the parsed extractor start with a valid length word segment and a NAL unit header.
[0304] The bytes of the sample builder are copied only from a single identified sample in the track referenced by the indicated "scal" track reference. Alignment is at decode time, i.e., only using the time-to-sample table followed by a count of samples in terms of sample number. The extractor is a media-level concept and thus is applied to the destination track before any edit lists are considered. Typically, the edit lists in both tracks will be the same. The following syntax can be used:
[0305]
[0306] NALUnitHeader() is the first two bytes of the HEVC NAL unit. A specific nal_unit_type value indicates the extractor, for example, a nal_unit_type equal to 49. The constructor_type specifies the constructor being used. EndOfNALUnit() is a function that returns 0 (false) when more data follows in the extractor and 1 (true) otherwise. The sample constructor (SampleConstructor) can have the following syntax:
[0307]
[0308] The track_ref_index identifies the source track from which the data is extracted. The track_ref_index is the index of a track reference of type "scal". This track reference has an index value of 1; the value 0 is reserved. The samples in the track from which the data is extracted are temporally aligned or the most recent in front in the media decoding timeline, i.e., using only the time-to-sample table, adjusted by the offset specified by sample_offset for the sample containing the extractor. The sample_offset gives the relative index of the sample in the track that can be used as a link to the information source. Sample 0 (zero) is the sample with the same or the most recent decoding time in front compared to the decoding time of the sample containing the extractor; sample 1 (one) is the next sample, sample -1 (negative one) is the previous sample, and so on. The data_offset is the offset of the first byte within the reference sample to be copied. If the extraction starts at the first byte of the data in the sample, then the offset takes the value 0. The data_length is the number of bytes to be copied.
[0309] The syntax of the inline constructor can be specified as follows:
[0310]
[0311] The length is the number of bytes belonging to the InlineConstructor after this field. The inline_data are the data bytes to be returned when parsing the inline constructor.
[0312] It needs to be understood that even though the extractors are described in the context of HEVC, they apply to other codecs and concepts similar to extractors.
[0313] An identified media data box may have the same semantics as that of the MediaDataBox but additionally contains an identifier that is used to establish a data reference to the contained media data. The identifier may be, for example, the first element contained by the identified media data box. The syntax of the identified media data box may be specified as follows, where imda_identifier is the identifier of the box. Note that although an imda_identifier of type 64-bit unsigned integer is used in the syntax, other field lengths and other basic data types (such as strings) are similarly possible. Example identified media data boxes are provided below:
[0314]
[0315] A box herein called DataEntryImdaBox may be used to reference data in an identified media data box. The DataEntryImdaBox identifies the IdentifiedMediaDataBox containing the data accessed via the data_reference_index corresponding to this DataEntryImdaBox. The DataEntryImdaBox contains the value of the imda_identifier of the referenced IdentifiedMediaDataBox. The media data offset is related to the first byte of the payload of the referenced IdentifiedMediaDataBox. In other words, media data offset 0 points to the first byte of the payload of the referenced IdentifiedMediaDataBox. A sample entry contains which data reference of the DataReferenceBox is in use for the data_reference_index of the sample containing the reference to this sample entry. When an IdentifiedMediaDataBox is used in a containing sample, the data_reference_index is set to the value pointing to the DataEntryImdaBox. The syntax of the DataEntryImdaBox may be specified as follows, where imda_ref_identifier provides the imda_identifier value and thus identifies a specific IdentifiedMediaDataBox:
[0316]
[0317] In an example, an identifier value for an identified media data box for a (sub)segment or movie fragment is determined and the identifier value is provided as a data reference basis for the identified media data for the (sub)segment or movie fragment. In an example, a template scheme for an identifier of an identified media data box is defined to be used as a data reference for sample data in, for example, a DataReferenceBox. The template scheme can be based on, but not limited to, a movie fragment sequence number (such as the sequence_number field of a MovieFragmentHeaderBox) or a track fragment decode time (such as the baseMediaDecodeTime field of a TrackFragmentBaseMediaDecodeTimeBox). It should be understood that any identifier provided for a movie fragment or track fragment, in addition to or in place of those described above, can be suitable for the template scheme. In an example, the following syntax can be used to reference an identified media data box using a template that yields an identifier.
[0318] aligned(8)class DataEntryTfdtBasedImdaBox(bit(24)flags)
[0319] extends FullBox('imdt', version = 0, flags){
[0320] }
[0321] The DataEntryTfdtBasedImdaBox identifies an IdentifiedMediaDataBox that contains media data accessed via the data_reference_index corresponding to this DataEntryTfdtBasedImdaBox. The media data offset 0 points to the first byte of the payload of the IdentifiedMediaDataBox with an imda_identifier having a baseMediaDecodeTime equal to that of the TrackFragmentBaseMediaDecodeTimeBox. The 64-bit imda_identifier value is used to carry the 64-bit value of the baseMediaDecodeTime. If the 32-bit baseMediaDecodeTime value is in use, the most significant bit of the 64-bit imda_identifier may be set to 0. For self-contained movie fragments, when the referenced data reference entry is of type DataEntryTfdtBasedImdaBox, the imda_identifier of the IdentifiedMediaDataBox is required to be equal to the baseMediaDecodeTime of the TrackFragmentBaseMediaDecodeTimeBox.
[0322] In another example, the following syntax can be used to reference an identified media data box using a template that yields an identifier.
[0323] aligned(8)class DataEntrySeqNumImdaBox(bit(24)flags)
[0324] extends FullBox('snim', version = 0, flags){
[0325] }
[0326] The DataEntrySeqNumImdaBox identifies the IdentifiedMediaDataBox that contains the media data accessed via the data_reference_index corresponding to this DataEntrySeqNumImdaBox. When the data_reference_index included in the sample entry refers to the DataEntrySeqNumImdaBox, each sample referring to the sample entry is included in the movie fragment, and the media data is offset 0 to point to the first byte of the payload of the IdentifiedMediaDataBox with an imda_identifier equal to the sequence_number of the MovieFragmentHeaderBox of the movie fragment containing the sample.
[0327] The size of the MovieFragmentBox does not need to be known at the time of determining the (multiple) base data offsets of the (multiple) tracks of the movie fragment, and thus sub-boxes of the MovieFragmentBox (such as TrackFragmentHeaderBox and TrackRunBoxes) can be "gradually" authorized before all the encoded media data for the movie fragment is available. In addition, the content encapsulator does not need to correctly estimate the size of the segment header and has some flexibility for dynamic variability of the segment duration.
[0328] The media segment header and the segment payload can be made separately available by a compile-time directive for a separate uniform resource locator (URL) for the segment header and a corresponding stream manifest for the segment payload. The stream manifest (such as the DASH media presentation description (MPD)) can provide a URL template, or the base URL given in the MPD can be indicated as applicable. In some embodiments, the stream manifest can also indicate that the data in the segment payload is tightly grouped and in decoding order. The segment payload can refer to, for example, the MediaDataBox. Tightly grouped refers to all the bytes belonging to the segment payload of the video bitstream, i.e., the segment payload includes a continuous byte range of the video bitstream. Such an indication can be provided as, for example, a supplementary property in the DASH MPD. The video bitstream in the segment payload can be an encapsulated video bitstream. For example, the segment payload can include a set of consecutive samples of the video track of an ISOBMFF file.
[0329] An index segment can be defined as a segment that mainly contains index information for media segments. The MPD can provide information indicating the URL that can be used to obtain the index segment. Examples of the information are given below:
[0330] The RepresentationIndex element within the SegmentBase element specifies a URL that includes the possible byte ranges for a segment of the representation index.
[0331] The SegmentList element includes a number of SegmentURL elements, which can include URLs for media segments (in the @media attribute), byte ranges within the resources identified by the URLs of the @media attribute, URLs for index segments (in the @index attribute), and / or byte ranges within the resources identified by the URLs of the @index attribute. The @media attribute, in combination with the URL in the @mediaRange attribute (if present), specifies the HTTP-URL for the media segment. The @index attribute, in combination with the URL in the @indexRange attribute (if present), specifies the HTTP-URL for the index segment.
[0332] The @index attribute of the SegmentTemplate element specifies that the template creates a list of index segments. The segment template includes a string from which a list of segments (identified by their URLs) can be obtained. The segment template can include specific identifiers that are replaced by dynamic values assigned to the segments to create a list of segments.
[0333] Each segment can have assigned segment index information that can be provided in an explicitly declared index segment. The presence of explicit index segment information can be indicated, for example, by any of the following:
[0334] - The presence of one RepresentationIndex element that provides the segment index for the entire representation, or
[0335] - The presence of at least one of the two attributes @index and @indexRange in the SegmentList.SegmentURL element, or
[0336] - The presence of the SegmentTemplate @index attribute.
[0337] The @indexRange attribute can also be used to provide the byte range for an index within a media segment, where this is permitted by the media segment format. In this case, the @index attribute is not present and the specified range lies entirely within any byte range specified for the media segment. The availability of index segments can be the same as the availability of the media segments to which they correspond.
[0338] Many types of content can be captured using multiple cameras, which can be temporally interleaved to produce video to be encoded. Such temporally interleaved shots from different cameras can be present in, for example, movies, sitcoms, late-night talk shows, conferences, and surveillance. Shots that are not temporally adjacent and originate from the same camera can have substantial similarity or correlation. However, since each shot conventionally starts with a random access picture, the similarity between shots has not been exploited in traditional video compression.
[0339] A method for exploiting the correlation of non-adjacent shots originating from the same camera has been presented in JVET-M0360 available at http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 13_Marr akech / wg11 / JVET-M0360-v4.zip. In JVET-M0360, many types of content can be captured using multiple cameras, which can be temporally interleaved to produce video to be encoded. Such temporally interleaved shots from different cameras can be present in, for example, movies, sitcoms, late-night talk shows, conferences, and surveillance. Shots that are not temporally adjacent and originate from the same camera can have substantial similarity or correlation. However, since each shot conventionally starts with a random access picture, the similarity between shots has not been exploited in traditional video compression.
[0340] A method for exploiting the correlation of non-adjacent shots originating from the same camera has been presented in JVET-M0360 available at http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 13_Marr akech / wg11 / JVET-M0360-v4.zip. In JVET-M0360, pictures that can be used by multiple random access segments or shots can be decoded into separate "library streams". The decoded pictures of the library streams can be input as additional reference pictures into the encoding and decoding of the "main stream".
[0341] Another method for exploiting the correlation of non-adjacent shots originating from the same camera has been presented in JVET-N0119 available at http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 14_Geneva / wg11 / JVET-N0119-v2.zip. Some aspects of JVET-N0119 are outlined in the following subparagraphs.
[0342] Pictures that can be used by multiple random access segments or shots are provided as coded pictures to the "main" encoding and decoding. The following aspects may exist:
[0343] An External Decoding Refresh (EDR) picture is a randomly accessible picture (in the "mainstream") for which one or more external pictures are required when randomly accessing the picture.
[0344] It is specified that an EDR picture can start a CVS and thus also start a bitstream.
[0345] External pictures are provided in the form of coded pictures rather than decoded pictures. External pictures are not output by the decoder; they are only used for inter-prediction reference.
[0346] The NAL unit type EDR_NUT is defined for EDR pictures.
[0347] For each EDR picture, the least significant bit (LSB) and the most significant bit of the picture order count (POC) value of the external pictures, as well as the most significant bit of the EDR picture itself, can be signaled in the tile group header.
[0348] Among the external pictures for an EDR picture, those that are not currently in the decoded picture buffer (DPB) can be provided by an external component. The access units containing these external pictures can be inserted into the bitstream to be decoded in decoding order immediately before the EDR access unit.
[0349] Some issues can be identified in some methods related to cross-shot. For example, the library pictures of JVET-M0360 and the external pictures of JVET-N0119 require separate encoding processes, specific signaling, and decoding process changes. Both methods are only applicable to codecs that implement the changes for encoding, signaling, and decoding.
[0350] The functionality is implemented as a codec change, which can be more generally applicable in cases where no change to the codec is required but the functionality can be implemented with system-level signaling applicable to any codec.
[0351] In JVET-M0360, library pictures may be provided to the "main" decoder as decoded pictures. Such an interface to the decoder may not generally be available.
[0352] In JVET-N0119, in the worst-case scenario, when randomly accessing at an External Decoding Refresh (EDR) picture, all external pictures are provided to the decoder as decoded pictures. The shortcut of using the already decoded external pictures available in the decoded picture buffer (DPB) may not be helpful.
[0353] An improved method is now introduced to at least mitigate the above problems. The improved method generally aims to exploit the correlation between non-adjacent shots from the same camera in video compression but may have other benefits and may be used for other purposes.
[0354] According to one embodiment, a processing entity obtains an intra random access point picture from a first position of an encoded video bitstream. The processing entity determines whether the intra random access point picture is reusable in at least a second position of the encoded video bitstream, the at least second position being different from the first position. In the case where it is determined that the intra random access point picture is reusable, the processing entity provides an identification assigned to the intra random access point picture that the intra random access point picture is a reusable intra random access point picture.
[0355] The processing entity can be, for example, an encoder, a file creator (encapsulating the encoded video bitstream into a container file), or an entity that converts a track (encapsulating the encoded video bitstream) into (sub)segments.
[0356] In an embodiment, obtaining the intra random access point picture includes encoding the intra random access point picture.
[0357] In an embodiment, the encoder reuses the intra random access point picture in a second position of the bitstream by including the same slice data in the first position and in the second position. As a result, the same intra random access point picture appears substantially multiple times in the bitstream, with potential differences in the picture segment header, such as different values for pic_output_flag and / or (a) syntax element(s) for deriving the delivery picture sequence number.
[0358] In an embodiment, the encoder determines that all intra random access point pictures are reusable when they are encoded. However, during the decoding process, the encoder can choose to reuse or not reuse a particular reusable intra random access point picture.
[0359] In an embodiment, the processing entity receives the encoded video bitstream. The processing entity parses or decodes at least a portion of the encoded video bitstream to obtain the intra random access point picture from a first position of the encoded video bitstream.
[0360] In an embodiment, the processing entity determines that all intra random access point pictures are reusable when they are processed (e.g., parsed).
[0361] In an embodiment, the processing entity analyzes the encoded video bitstream to find whether slice data (or picture segment payload or the like) is the same in another intra random access point picture of the encoded video bitstream. Thus, the processing entity determines that the intra random access point picture is a reusable intra random access point picture.
[0362] Identifiers for reusable IRAP pictures can be provided in or with the bitstream and can include, but are not limited to, one or more of the following:
[0363] - The identifier can be an ID value selected or given to the encoder (“IRAP ID” or “SAP ID”, which can be used interchangeably).
[0364] - The identifier can be a checksum of the coded picture content (e.g., all slice data). For example, an MD5 checksum can be used.
[0365] Identifiers for reusable IRAP pictures can be carried in or with the bitstream, for example, using but not limited to the following methods. When a reusable IRAP picture is included in the bitstream, the indication can be included in an SEI message. For example, a coded picture hash SEI message can be defined and the SEI message syntax can include the identifier of the coded picture as a checksum over all slice data.
[0366] A hash function can be defined as any function that can be used to map digital data of any size to digital data of a fixed size, where a slight difference in the input data can produce a large difference in the output data. A cryptographic hash function can be defined as a hash function that is designed to be practically irreversible (i.e., to create the input data solely based on the hash value). Cryptographic hash functions can include, for example, the MD5 function. The MD5 value can be a null-terminated string of UTF-8 characters containing the base64-encoded MD5 digest of the input data. One method of computing the string is specified in IETF RFC 1864. It should be understood that, instead of or in addition to MD5, other types of integrity check schemes can be used in various embodiments, such as different forms of cyclic redundancy check (CRC), such as the CRC scheme used in ITU-T Recommendation H.271.
[0367] A checksum or hash sum can be defined as a small piece of data from any block of digital data, which can be used for the purpose of detecting errors that may have been introduced during its transmission or storage. The actual process of generating a checksum for a given data input can be referred to as a checksum function or checksum algorithm. A checksum algorithm will typically output significantly different values even for very small changes to the input. This is especially true for cryptographic hash functions, which can be used to detect many data corruption errors and verify overall data integrity; if the checksum computed for the current data input matches the stored value of a previously computed checksum, then there is a high probability that the data has not been changed or corrupted.
[0368] The term checksum can be defined as being equivalent to a cryptographic hash value or the like.
[0369] Some example embodiments for transporting identifiers in ISOBMFF and / or DASH will be provided in this specification below.
[0370] The identifier of a reusable IRAP picture can enable a receiver and / or decoder to detect the reusable IRAP picture. Accordingly, the reusable IRAP picture does not need to be sent again if it has been received, and / or the reusable IRAP picture does not need to be decoded again if it has been decoded.
[0371] According to an embodiment, a processing entity receives an identifier associated with an intra random access point picture in a first position of an encoded video bitstream, the identifier determining whether the intra random access point picture is reusable in at least a second position of the encoded video bitstream, the at least second position being different from the first position. In the case where the identifier indicates that the intra random access point picture is reusable, the processing entity checks whether a corresponding previously received intra random access point picture is available. If the check reveals that the corresponding previously received intra random access point picture is available, then processing of the intra random access point picture is skipped.
[0372] In an embodiment, the corresponding previously received intra random access point picture has a second identifier associated therewith. The processing entity receives the second identifier. As part of the check, the processing entity compares the identifier with the second identifier, and in the case where they are equal and where the previously received intra random access point picture is available, the processing entity infers that the check reveals that the corresponding previously received intra random access point picture is available.
[0373] In an embodiment, if the check reveals that a first intra random access point picture is available, then the processing entity skips processing of the intra random access point picture, such as skipping a request, reception, storage, and / or decoding of the intra random access point picture. In an embodiment, if the check reveals that the intra random access point picture is available as a decoded picture, then the processing entity skips decoding of the intra random access point picture and instead uses the corresponding previously received intra random access point picture as a reference picture for decoding subsequent encoded pictures in decoding order.
[0374] According to an embodiment, a processing entity receives a first intra random access point (IRAP) picture from a first position of an encoded video bitstream and a first identifier associated with the first IRAP picture. The processing entity receives a second identifier associated with a second IRAP picture residing at a second position in the encoded video bitstream. In response to the first identifier being equal to the second identifier, the processing entity infers that the first IRAP picture is reused at the second position and checks whether the first IRAP picture is available (e.g., received and held in a buffer). If the check reveals that the first IRAP picture is available, the first IRAP picture is used instead of the second IRAP picture.
[0375] In an embodiment, if the check reveals that the first IRAP picture is available, the processing entity ignores the processing of the second IRAP picture, such as ignoring a request and / or receiving the second IRAP picture. In an embodiment, if the check reveals that the first IRAP picture is available as a decoded picture, the processing entity ignores the decoding of the first IRAP picture in place of the second IRAP picture and instead uses the decoded first IRAP picture as a reference picture for decoding subsequent encoded pictures in decoding order.
[0376] The processing entity can be, for example, a decoder, a receiver, a player, a streaming client, a file parser (which demultiplexes the encoded video bitstream from a container file), or an entity that processes received (sub)segments.
[0377] The identifier of a reusable IRAP picture can be considered as a piece of metadata characterizing the reusable IRAP picture. Additional metadata characterizing the reusable IRAP picture is described hereinafter. The additional metadata characterizing the reusable IRAP picture can include, but is not limited to:
[0378] - The stream access point type for the reusable IRAP picture
[0379] - The byte range for the reusable IRAP picture; if aligned with the start of a (sub)segment, the byte range or size of the initial SAP sample can be provided
[0380] - The byte range for non-IRAP pictures associated with the IRAP picture
[0381] ○ In an embodiment, the leading and trailing byte ranges for a picture can be provided separately
[0382] - The decoding timestamp and / or composition timestamp for the reusable IRAP picture
[0383] Additional metadata can be considered as part of the identification of reusable IRAP pictures, and can accompany the identification of reusable IRAP pictures, for example, within the same syntax structure and / or signaling mechanism, or can be carried separately from the identification of reusable IRAP pictures.
[0384] In the following, ISOBMFF and DASH embodiments will be described.
[0385] ISOBMFF and DASH embodiments can be classified into two categories of methods:
[0386] In the first category, reusable IRAP pictures (repeatedly) appear in traditional video tracks, DASH representations, or the like. Additional signaling exists, which enables avoiding unnecessary downloads of reusable IRAP pictures and identifies which reusable IRAP pictures can be used in place of non-downloaded IRAP pictures. This category is backward compatible, i.e., traditional players can handle the track or representation. Several embodiments falling into this category will be presented below.
[0387] In the second category, reusable IRAP pictures are located in their own representations. The "main" representation does not contain reusable IRAP pictures and is indicated as depending on the representation regarding reusable IRAP pictures, for example, using the @dependencyId attribute of DASH and / or track references. The SAP of the reusable IRAP picture representation can be defined as such a position: the reusable pictures before the SAP are not used as references for any pictures at or after the SAP. This method is backward compatible, but requires players to support the handling of dependent representations and / or track references.
[0388] In the following, some embodiments related to the above first category will now be explained in more detail.
[0389] In the first embodiment of the first category, the timing metadata track has metadata regarding reusable IRAP pictures, and the timing metadata track carries the identification and associated metadata of the reusable IRAP pictures.
[0390] The timing metadata track samples can indicate, but are not limited to, one or more of the following indications for reusable IRAP pictures:
[0391] - Output flag (containing the value of pic_output_flag in the slice header of the IRAP picture)
[0392] - Identification of the reusable IRAP picture, such as IRAP ID
[0393] - Byte range within a media segment
[0394] Each (sub)segment of the timing metadata track can be selected or required to cover an integral number of adjacent (sub)segments of the referenced media track.
[0395] The player can operate as follows:
[0396] The player stores received reusable IRAP pictures separately and can access them by their IDs. The player receives (sub)segments of the timing metadata track. When the player notices from the timing metadata track that the next media (sub)segment to be requested contains a reusable IRAP picture that the player has stored, the player requests a subset range of the next media (sub)segment that does not include the reusable IRAP picture. The player can concatenate the received byte range with the stored reusable IRAP picture (as originally stored in the media sample) and process the concatenated MediaDataBox in a conventional manner. If the value of the output flag of the stored reusable IRAP picture is different from the value of the output flag of the IRAP picture in that (sub)segment, then the value of pic_output_flag needs to be rewritten.
[0397] Alternatively, instead of requesting a subset range of the next media (sub)segment that does not include the reusable IRAP picture, the decoder can store a decoded version of the reusable IRAP picture and omit their decoding again if they are already available as decoded pictures.
[0398] In a second embodiment of the first category, an extension of the SegmentIndexBox is utilized.
[0399] In this embodiment, the SegmentIndexBox is extended at the end of the current syntax of the SegmentIndexBox. A conventional parser is expected to stop parsing before the extension.
[0400] The extension can provide, but is not limited to, one or more of the following:
[0401] - The movie fragment segment header size of an indexed (leaf) subsegment;
[0402] - If an indexed subsegment starts with an SAP, then the SAP can be considered a reusable IRAP picture, and the extension can provide, but is not limited to, one or more of the following for the reusable IRAP picture:
[0403] ○ The identification of the reusable IRAP picture (hereinafter SAP ID);
[0404] ○ For example, the size of the reusable IRAP picture in bytes;
[0405] ○ The SAP output flag (including the value of pic_output_flag).
[0406] The movie fragment header size can be used to infer the byte range to obtain the movie fragment header. The movie fragment header size can include, for example, all boxes from the start of the (sub)fragment up to and not including the start of the MediaDataBox that contains the media data for the movie fragment.
[0407] The SAP ID can be used to infer whether the corresponding reusable IRAP picture has been received.
[0408] (The SAP size in bytes) can be used to avoid its download, i.e., to request the byte range starting from the sample after the SAP sample.
[0409] If the SAP output flag has a value different from the value of the received SAP picture, then the pic_output_flag in the received SAP picture needs to be rewritten.
[0410] Several example embodiments including specific syntax are provided below. It is to be understood that other embodiments can be implemented similarly.
[0411] Examples for indicating the SAP ID and the SAP size are included below.
[0412]
[0413]
[0414] Using max_SAP_type_with_SAP_id, the signaling of the SAP ID and the SAP size can be limited to only specific SAP types, such as only for SAP type 1 and SAP type 2, for which max_SAP_type_with_SAP_id can be selected to be equal to 2.
[0415] In another example, the movie fragment header size, the SAP output flag, the SAP ID, and the SAP size are indicated. An optional number of bits is assigned to indicate the movie fragment header size, the SAP ID, and the SAP size. A function f() can be specified, for example, to obtain f(0)=0; f(1)=8; f(2)=16; and f(3)=32.
[0416]
[0417]
[0418] In a third embodiment of the first category, a sample group is utilized.
[0419] In this embodiment, a sample group is created to contain identifiers for reusable IRAP pictures and is carried in the "main" video track.
[0420] The sample group description (i.e., SampleGroupDescriptionBox) for the created sample group may include distinct sample group description entries for each IRAP ID value.
[0421] In an example, the following sample group description entry syntax is used (where sap_id contains the IRAP ID value):
[0422] class SAPIdentifierEntry()extends SampleGroupDescriptionEntry('said′){
[0423] unsigned int(32)sap_id;
[0424] }
[0425] In an embodiment, an output flag is also or may be provided in the sample group description entry. Thus, two sample group description entries may have the same IRAP ID value with different output flag values.
[0426] In an embodiment, the sample group description entry of the SAP sample group is appended with an IRAP ID. For example, the following sample group description entry syntax may be used:
[0427]
[0428] In another example, the following sample group description entry syntax may be used:
[0429]
[0430] The semantics of dependent_flag and sap_type may be specified as previously described. sap_id may contain an IRAP ID or a SAP ID value. A repeated_flag equal to 1 specifies that the encoded picture is the same as a previous encoded picture with the same sap_id value in decoding order. A repeated_flag equal to 0 specifies that the encoded picture may be the first occurrence of the encoded pictures in decoding order.
[0431] The player may operate as follows:
[0432] The player stores the received reusable IRAP pictures separately and can access them by their IDs. The player requests and receives the movie fragment headers of (sub-)segments of the media presentation. When the player notices from the SampleToGroupBox included in the movie fragment header that the next media (sub-)segment to be requested contains reusable IRAP pictures that the player has already stored, the player requests a subset range of the next media (sub-)segment that does not include the reusable IRAP pictures. The player can concatenate the received byte range with the stored reusable IRAP pictures (as originally stored in the media samples) and process the concatenated MediaDataBox in a conventional manner. If the value of the output flag of the stored reusable IRAP picture is different from the value of the output flag of the IRAP picture in the (sub-)segment, then the value of pic_output_flag needs to be rewritten.
[0433] Alternatively, instead of requesting a subset range of the next media (sub-)segment that does not include the reusable IRAP pictures, the decoder can store the decoded versions of the reusable IRAP pictures and omit their decoding again if they are already available as decoded pictures.
[0434] In a fourth embodiment of the first category, an extractor is utilized.
[0435] When a reusable IRAP picture already existed earlier in the bitstream, it can be included in the samples of the track by using, for example, a reference to an extractor NAL unit-like structure that specifies the samples from which the coded picture was obtained.
[0436] In the first method, the first occurrence of the reusable IRAP picture is included in the "main" video track. The extractor refers to the first occurrence, i.e., the same track that also contains the extractor itself. Using the current extractor syntax of ISO / IEC 14496-15, this method is limited by an 8-bit sample_offset, i.e., the first occurrence of the reusable IRAP picture cannot be far from later occurrences without a change in the extractor syntax. However, other extractor syntaxes can exist or can be created where the limitation may not exist or may be different from the 8-bit sample_offset.
[0437] In the second method, reusable IRAP pictures are included in their own (multiple) tracks. Thus, a reusable IRAP picture can be extracted to the current sample of the "main" video track as long as the extracted IRAP is within the -128 to 127 reusable IRAP pictures other than the current sample. The advantage of this method is that it uses existing mechanisms and some traditional players can easily support it. However, the (sub)segment containing the reusable IRAP picture to be extracted is not explicitly indicated but can be inferred, for example, from the sample_offset, frame rate, and (sub)segment duration.
[0438] In the fifth embodiment of the first category, the SubsegmentIndexBox and related structures and / or signaling are utilized. In this embodiment, the reusable IRAP picture starting the subsegment is assigned a different level from other pictures.
[0439] In an embodiment, a new level assignment type for the level assignment box is defined for the purpose of assigning reusable IRAP pictures to a different level from other pictures.
[0440] In an embodiment, a specific range of level values is used as the IRAP ID value. For example, a level equal to 255 can be specified to indicate pictures other than reusable IRAP pictures, and level values from 0 to 254 (inclusive) are used to indicate the IRAP ID value. Thus, when a level value within the range of 0 to 254 (inclusive) is indicated in the SubsegmentIndexBox, the corresponding range syntax element indicates the byte range of the reusable IRAP picture.
[0441] In an embodiment, the level assignment type for the sample group referring to the track can be used with a suitable sample group as discussed, for example, in the third embodiment of the first category above.
[0442] Thus, in the above embodiment, the SubsegmentIndexBox (for the subsegment) provides the byte range for the reusable IRAP picture starting the subsegment and the byte range for other pictures.
[0443] The third and fifth embodiments of the first category can be combined, where the sample group provides a mapping of samples as reusable IRAP pictures and the IRAP ID value or the like, and the SubsegmentIndexBox provides the byte range for the reusable IRAP pictures and other pictures.
[0444] In the sixth embodiment of the first category, a new box is defined as being accompanied by a SegmentIndexBox. The new box is hereinafter referred to as the SapIdentificationBox but could similarly have another name. The SapIdentificationBox provides the SAP ID values of the indexed subsegments and their byte ranges. The SapIdentificationBox can additionally provide additional metadata for reusable IRAP pictures. The reusable IRAP pictures described by the SapIdentificationBox can be limited to those starting with subsegments. The SapIdentificationBox (if any) can be required to be the next box after the associated SegmentIndexBox and SubsegmentIndexBox (if any). The SapIdentificationBox records the subsegments indicated in the immediately preceding SegmentIndexBox.
[0445] In the various embodiments above, it can be advantageous to infer the movie fragment header size to separately obtain the movie fragment header and the media data of the movie fragment (excluding reusable IRAP pictures). The movie fragment header size can, for example, include all boxes from the start of the (sub)segment up to and excluding the start of the MediaDataBox containing the media data for the movie fragment. In addition to or instead of the above mechanism for indicating the movie fragment header size, one or more of the following mechanisms can be used:
[0446] - Use the previously described DASH MPD mechanism (e.g., the RepresentationIndex element or the @indexRange attribute) to indicate the movie fragment header as an index segment.
[0447] - Use the previously described DASH MPD mechanism(s) to provide the byte range of the movie fragment. Specifically, the @indexRange attribute can be used to provide the byte range for the index within the media segment, and the index can be defined to include the movie fragment header.
[0448] Hereinafter, some alternative HRD parameters for omitting the buffering of reusable IRAP pictures will be described.
[0449] According to an embodiment that can be used independently of or in conjunction with other embodiments, an encoder encodes and / or a decoder decodes alternative Hypothetical Reference Decoder (HRD) parameters in or with a bitstream without buffering the coded picture buffer (CPB) of a reusable Instantaneous Random Access Picture (IRAP) picture. A decoder that already has buffered and / or available reusable IRAP pictures can use these alternative HRD parameters (instead of the “normal” HRD parameters that include reception of the reusable IRAP pictures into the CPB). In an embodiment, these alternative parameters apply when decoding starts from a reusable IRAP picture. The alternative parameters can include, but are not limited to, one or more initial delays (such as an initial CPB buffering delay) and / or one or more buffer sizes. The alternative parameters can be signaled, for example, in a buffering period SEI message.
[0450] Now, buffering of reusable IRAP pictures in the coded picture buffer (CPB) will be described.
[0451] According to an embodiment that can be used independently of or in conjunction with other embodiments, a reusable IRAP picture remains in the CPB after its decoding to be potentially re-decoded later.
[0452] An embodiment can include control signals to leave a coded picture in the CPB after decoding of the coded picture, re-decode the coded picture from the CPB, or remove the coded picture from the CPB.
[0453] An encoder according to an embodiment encodes one or more of the above control signals in or with a bitstream. The encoder operates its CPB according to the control signals.
[0454] A decoder according to an embodiment decodes one or more of the above control signals from or with a bitstream. The decoder operates its CPB according to the control signals.
[0455] Hereinafter, the use of multiple independent layers will be described.
[0456] According to an embodiment, an encoder selects to include temporally related pictures in the same independent layer.
[0457] Metadata indicating from which camera each picture originates can be provided with the video given as input to the encoder. In that case, the encoder can select to include pictures originating from the same camera in the same independent layer.
[0458] If camera source metadata is not provided along with the video given as input to the encoder, the encoder can operate as follows to infer which input pictures are temporally related to each other such that they are selected to be included in the same independent layer. The encoder can perform shot boundary detection between consecutive pictures. When a picture is determined to start a new shot, the encoder can perform shot boundary detection between that picture and a selected picture from an earlier shot (e.g., those encoded as IRAP pictures). When a correlation is found between the picture starting the shot and an earlier picture in the earlier shot, the encoder can determine that the shots are temporally related and select to include them in the same independent layer.
[0459] In an embodiment, pictures starting a temporally consecutive "segment" or shot within an independent layer are encoded as dependent random access point (DRAP) pictures, i.e., trailing pictures that reference only the previous IRAP picture within the independent layer during their decoding process. Assuming the necessary parameter sets are available when they need to be activated and the associated IRAP pictures are available, the DRAP pictures and all subsequent pictures in decoding order can be correctly decoded without performing the decoding process on any pictures in decoding order before the DRAP picture other than the associated IRAP picture. The encoder can indicate in the bitstream or along with the bitstream, e.g., using SEI messages, that a picture is a DRAP picture.
[0460] An embodiment can be illustrated as Figure 1b depicted in. In this example, there are three layers. In the bitstream of each layer, there are intra random access point pictures, followed by dependent random access point pictures at a later time instance. The dependent random access point pictures are predicted from the previous intra random access point pictures in the same layer.
[0461] This embodiment enables random access to the bitstream at a DRAP picture (in any independent layer) by obtaining the previous IRAP picture in that independent layer, decoding that previous IRAP picture, and then decoding the bitstream starting from the DRAP picture in decoding order. Random access can be performed, for example, in response to a end user seeking a DRAP picture or a time position close to a DRAP picture.
[0462] In an embodiment, there can be one or more reference pictures that are not necessarily previous IRAP pictures used in decoding a picture starting a temporally consecutive "segment" or shot within an independent layer. In other words, a picture starting a temporally consecutive "segment" or shot within an independent layer can be a trailing picture and does not need to be encoded as a dependent random access point (DRAP) picture. This embodiment can provide better compression performance in a similar simple way compared to embodiments where DRAP pictures are used but random access functionality is not provided.
[0463] An embodiment may include an indication or control signal (e.g., in a VPS) indicating one or more of the following:
[0464] - The layers are temporally interleaved. In other words, each access unit contains coded pictures of exactly one layer.
[0465] - The output of the layers is expected as a single picture sequence.
[0466] - Constraints on the use of the joint DPB for the layers (e.g., not exceeding the horizontal limit of any single layer).
[0467] - An indication or control signal that the layers share the same CPB.
[0468] - An indication or control signal that the layers share the same DPB.
[0469] - CPB parameters when a single CPB for all involved layers is used.
[0470] - DPB parameters when a single DPB for all involved layers is used.
[0471] - Profiles, levels, and / or tiers required for decoding using a single decoder instance.
[0472] In an embodiment, the above indication or control signal is indicated according to an output layer set. For example, the VPS syntax may include a loop over all output layer sets, and one or more of the above indications may be indicated for each loop entry, where the loop entry corresponds to an output layer set. In an embodiment, the above indication or control signal is indicated according to a layer set, possibly among multiple layer sets. In an embodiment, the above indication or control signal is indicated for a bitstream.
[0473] An encoder according to an embodiment encodes one or more of the above indication or control signals in or with the bitstream. The encoder operates in a manner that the indication or control signal is complied with.
[0474] A decoder according to an embodiment decodes one or more of the above indication or control signals from or with the bitstream. The decoder operates in a manner that the indication or control signal is complied with. For example, in response to an indication or control signal that the output of the layers is as a single picture sequence, the decoder outputs the decoded pictures as a single picture sequence.
[0475] In an embodiment, the decoder decodes an indication layer that is a temporally interleaved indication or control signal (e.g., from a VPS). As a result of the indication layer being a temporally interleaved indication or control signal (e.g., from a VPS), the decoder performs a single decoder instance for decoding all layers (within the scope of the indication or control signal). If the indication or control signal is specific to an output layer set or the like, the decoder may exclude layers that are not among the output layer set or the like from the bitstream and decode the remaining bitstream. Such a sub-bitstream extraction process may be performed for a portion of the bitstream at one time, or may be performed for the entire bitstream. As part of decoding with a single decoder instance, reference picture marking is performed layer by layer. For example, an IDR picture at a particular layer causes earlier pictures within that layer in decoding order to be marked as "not for reference" or requires RPL / RPS signaling that causes earlier pictures within that layer in decoding order to be marked as "not for reference", while the IDR picture has no effect on pictures on other layers.
[0476] In an embodiment, the encoder encodes in the bitstream or with the bitstream and / or the decoder decodes from the bitstream or with the bitstream (e.g., in an ERAS NAL unit) at the end of a random access segment (ERAS) signal (e.g., a NAL unit). The ERAS NAL unit or the like may be the last NAL unit of an encoded picture or access unit (which is the last encoded picture or access unit of a continuous sequence of RAS or shot or pictures not interleaved with pictures of other independent layers) or the last NAL unit associated with the encoded picture or access unit.
[0477] In an embodiment, during the encoding process and / or during the decoding process, the ERAS NAL unit or the like causes all pictures to be marked as "not for reference", except for previous IRAP pictures of the same independent layer in decoding order.
[0478] In an embodiment, the encoder encodes in the bitstream or with the bitstream (e.g., in an ERAS NAL unit) and / or the decoder decodes from the bitstream or with the bitstream (e.g., in an ERAS NAL unit) reference picture list (RPL) or reference picture set (RPS) signaling that is performed and / or applied after the last encoded picture associated with the ERAS NAL unit or the like has been decoded.
[0479] In an embodiment, the encoder encodes the RPL / RPS signaling applied after the last encoded picture of a slice in the following manner: the previous IRAP pictures of the same independent layer in decoding order are marked as "for reference" or "for short-term reference" or "for long-term reference", etc., and the other decoded reference pictures of the same independent layer are marked as "not for reference". This embodiment can be used in combination with encoding a DRAP picture as the first picture in a subsequent slice of the same independent layer, as described above in the embodiment.
[0480] In an embodiment, the encoder encodes the RPL / RPS signaling applied after the last encoded picture of a slice in the following manner: the selected picture(s) is / are marked as "for reference" or "for short-term reference" or "for long-term reference", etc., and the other decoded reference pictures of the same independent layer are marked as "not for reference". The number of the selected picture(s) can be chosen to be small to limit the memory consumption for decoded reference pictures. This embodiment can be used in combination with encoding a trailing picture as the first picture in a subsequent slice of the same independent layer, as described above in the embodiment.
[0481] In an embodiment, one or more specific indications are provided in or with the bitstream to indicate a cross-layer random access point (CL-RAP), which may be characterized in that pictures in any layer before the CL-RAP are not required during the decoding process for decoding any picture in any layer at or after the CL-RAP in decoding order. These indications can be generated by the encoder or other entities (such as a file creator or splicer that concatenates parts of the bitstream) and can be utilized and / or obeyed by the decoder. In some embodiments, these indications can be used for only one or more specific picture types or NAL unit types, such as only for IRAP pictures, while in other embodiments, these indications can be used for any picture type(s).
[0482] The indication(s) can reside, for example, in one or more of the following syntax structures in the bitstream:
[0483] - NAL unit header
[0484] - Slice header
[0485] - Picture parameter set
[0486] - Picture header
[0487] - Access unit delimiter
[0488] The indication(s) can be, for example, in one or more of the following syntax structures with the bitstream:
[0489] - Samples in a group within a file encapsulating a bitstream.
[0490] - Samples in a timing metadata track associated with a track of one or more layers of the encapsulated bitstream.
[0491] - A SegmentIndexBox, SubsegmentIndexBox, or another box that indexes or provides metadata for sub - segments of a portion of the encapsulated bitstream.
[0492] - It can be specified that the encapsulation indicates that the sync samples of the track(s) of the bitstream or output layer set involved correspond to the CL - RAP. Thus, the SyncSampleBox or sync sample indication within the track segment can be used to indicate the CL - RAP.
[0493] - It can be specified that the encapsulation indicates that the SAP of the track(s) of the bitstream or output layer set involved corresponds to the CL - RAP. Thus, the SAP sample group can be used to indicate the CL - RAP, where the grouping_type_parameter indicates the layer(s) involved.
[0494] In an embodiment, a decoder, player, or the like parses the above - described indication(s) to infer the CL - RAP position. The decoder, player, or the like uses the CL - RAP, for example, for random access, start - up decoding, seeking, fast - forward play, and / or rewind play.
[0495] Figure 2a FIG. is a flowchart illustrating an encoding method according to an embodiment. The method includes obtaining 161 an Intra Random Access Picture. Then it can check 162 whether the IRAP picture can be reused, and if so, then provide 163 an indication that it is a reusable IRAP picture associated with the IRAP picture.
[0496] Figure 2b FIG. is a flowchart illustrating a decoding method according to an embodiment. The method includes receiving 164 an Intra Random Access Picture. Then it can check 165 whether the IRAP picture can be reused, and if so, then store 166 the reusable IRAP picture into a memory.
[0497] Figure 2cFIG. is a flowchart illustrating a decoding method according to another embodiment. The method includes receiving 164 an Intra Random Access Picture (IRAP). Then it can check 165 whether the IRAP picture can be reused, and if so, further check 167 whether the IRAP picture has been stored in the memory. If the IRAP picture has not been stored in the memory or is not available for decoding for some other reason, then the IRAP picture is stored 168 in the memory. The method can also include requesting 169 the byte range of the next media (sub) segment that does not include a reusable IRAP picture.
[0498] Figure 4a FIG. shows some elements of a video encoding section 510 according to an embodiment. The checking block 511 can include obtaining an input 511a for encoding video. The checking block 511 checks whether the IRAP picture is reusable and if so, the indicating block 512 associates an indication that it is a reusable IRAP picture with the IRAP picture. The indication, the picture, and possibly other information are provided to the encoding block 513. The encoding block 513 can encode the signal and the video for storage and / or transmission. However, there can be separate encoding elements for signal encoding and visual information encoding.
[0499] Figure 4b FIG. shows a video decoding section 520 according to an embodiment. The video decoding section 520 can obtain signaling data via a first input 521 and encoded visual information (e.g., an IRAP picture) via a second input 522. The signaling data and the encoded visual information can be decoded by a decoding block 523. The decoded signaling data can be used by the checking block 524 to check, in particular, whether the IRAP picture should be requested, received, stored, and / or decoded; or whether the IRAP picture should be replaced by a corresponding previously received IRAP picture. The decoded video data can be used by a rendering block 525 to reconstruct a video based on the decoded video data.
[0500] An example of a device (e.g., a device for encoding and / or decoding) is illustrated in Figure 5 FIG. The general structure of the device will be explained in terms of the functional blocks of the system. Several functions can be performed by a single physical device, e.g., all computational processes can be executed in a single processor if needed. According to Figure 5 FIG., the data processing system of the example device includes a main processing unit 100, a memory 102, a storage device 104, an input device 106, an output device 108, and a graphics subsystem 110, all of which are connected to each other via a data bus 112.
[0501] The main processing unit 100 can be a conventional processing unit arranged to process data within a data processing system. The main processing unit 100 can include or be implemented as one or more processors or processor circuitry. The memory 102, storage device 104, input device 106, and output device 108 can include conventional components as recognized by those skilled in the art. The memory 102 and storage device 104 store data in the data processing system 100. Computer program code resides in the memory 102 for implementing, for example, a method according to an embodiment. The input device 106 inputs data into the system, while the output device 108 receives data from the data processing system and forwards the data to, for example, a display. The data bus 112 is a conventional data bus and, although shown as a single line, can be any combination of: a processor bus, a PCI bus, a graphics bus, an ISA bus. Thus, it is readily appreciated by those skilled in the art that the apparatus can be any data processing device, such as a computer device, a personal computer, a server computer, a mobile phone, a smartphone, or an Internet access device (e.g., an Internet tablet computer).
[0502] Various embodiments can be implemented by means of computer program code residing in a memory and causing a relevant apparatus to execute a method. For example, a device can include circuitry and electronics for processing, receiving, and sending data, computer program code in a memory, and a processor that, when the computer program code is run, causes the device to execute the features of an embodiment. Additionally, a network device, such as a server, can include circuitry for processing, receiving, and sending data, computer program code in a memory, and a processor that, when the computer program code is run, causes the network device to execute the features of an embodiment. The computer program code includes one or more operational characteristics. The operational characteristics are such that a system is connectable to the processor via a bus, and the programmable operational characteristics of the system include: receiving a data unit, the data unit being logically separated into a first bit stream and a second bit stream, combining the first bit stream and the second bit stream into a combined bit stream, where the combining includes writing a delimiter into the combined bit stream, the delimiter indicating to which of the first bit stream or the second bit stream one or more data units associated with the delimiter in the combined bit stream are assigned.
[0503] If desired, the different functions described herein can be performed in a different order and / or concurrently with each other. Additionally, if desired, one or more of the functions and embodiments described above can be optional or can be combined.
[0504] In the foregoing, although example embodiments have been described with reference to an encoder, it is to be understood that the resulting bit stream and decoder can have corresponding elements therein.
[0505] Similarly, although the example embodiments have been described with reference to a decoder, it is to be understood that an encoder may have a structure and / or a computer program for generating a bitstream to be decoded by the decoder.
[0506] Above, although the example embodiments have been described with reference to syntax and semantics or with reference to indications or signals in or accompanying the bitstream, it is to be understood that the embodiments similarly cover an encoder that outputs a portion of the bitstream or outputs indications or signals in or accompanying the bitstream according to syntax and semantics. Similarly, the embodiments similarly cover a decoder that decodes a portion of the bitstream or decodes or interprets indications or signals in or accompanying the bitstream according to syntax and semantics.
[0507] The embodiments of the invention described above describe a codec according to a single separate encoder and decoder device to assist in understanding the processes involved. However, it will be recognized that the device, structure, and operation may be implemented as a single encoder-decoder device / structure / operation. Additionally, it is possible that the encoder and decoder may share some or all common elements.
[0508] Although some embodiments of the invention describe codec operations within a device, it will be recognized that the invention as defined in the claims may be implemented as part of any video codec within any system or environment. Thus, for example, embodiments of the invention may be implemented in a video codec that may achieve video encoding via a fixed or wired communication path.
[0509] Above, although the example embodiments have been described with reference to ISOBMFF, it is to be understood that the embodiments similarly apply to other container file formats, such as Matroska.
[0510] Above, although the example embodiments have been described with reference to DASH MPD, it is to be understood that the embodiments similarly apply to other presentation formats, such as the playlist format specified in IETF RFC 8216, and / or other media description formats, such as the Session Description Protocol (SDP) format specified in IETF RFC 4566.
[0511] Although aspects of the embodiments are set out in the independent claims, other aspects include other combinations of features from the described embodiments and / or the dependent claims with the features of the independent claims, not only the combinations explicitly set out in the claims.
Claims
1. A method for encoding, comprising: Obtaining an intra random access point picture from a first position of a video bitstream; Determining whether the intra random access point picture is reusable in at least a second position of the video bitstream, wherein the at least second position is different from the first position; Providing an identification assigned to the intra random access point picture indicating that the intra random access point picture is a reusable intra random access point picture in case the determination indicates that the intra random access point picture is reusable; And Reusing the intra random access point picture in the second position of the video bitstream.
2. An apparatus for encoding, comprising: Means for obtaining an intra random access point picture from a first position of a video bitstream; Means for determining whether the intra random access point picture is reusable in at least a second position of the video bitstream, wherein the at least second position is different from the first position; Means for providing an identification assigned to the intra random access point picture indicating that the intra random access point picture is a reusable intra random access point picture in case the determination indicates that the intra random access point picture is reusable; And Means for reusing the intra random access point picture in the second position of the video bitstream.
3. The apparatus according to claim 2, further comprising means for: Encoding pictures into a plurality of independent layers, wherein each layer includes a representation of the picture; and Including temporally related pictures into the same independent layer, wherein pictures within the independent layer are not predicted from pictures of other layers.
4. The apparatus according to claim 3, further comprising means for: Encoding a picture that is the start of a temporally consecutive sequence of encoded pictures within an independent layer as a dependent random access point picture.
5. The apparatus according to claim 3, further comprising: Providing an indication or control signal indicating one or more of the following: The independent layers are time interleaved; The output of the layers is expected to be completed as a single picture sequence; Constraints on the use of the joint decoding picture buffer for the independent layers.
6. The apparatus according to any one of claims 2 to 5, comprising means for: Providing the identification in a container file encapsulating the video bitstream.
7. The apparatus according to claim 6, further comprising means for encoding the following in or linked to the identification: Byte ranges within a media segment.
8. The apparatus according to any one of claims 2 to 5, further comprising: Providing the identification as one of the following: An ID value; A checksum of the encoded picture content.
Citation Information
Patent Citations
A system and method for time optimized encoding
CN101682761A
An apparatus, a method and a computer program for video coding and decoding
WO2017140946A1