Methods, apparatuses, and computer readable media for processing media content
By adjusting the presentation time of samples in the media stream, the problem of unknown duration of sparse media content samples is solved, thus achieving accurate playback of media content.
Patent Information
- Application Number
- CN202211005687.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-08
- Filing Date
- 2019-07-09
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2039-07-09
AI Technical Summary
Existing technologies struggle to accurately render media content, especially when processing media streams where the sample duration of sparse media content is unknown.
By obtaining the current segment at the current time instance and determining the modified duration based on that segment, the presentation time of previous samples is adjusted to expand or shrink, ensuring accurate presentation of media content.
It achieves accurate presentation of sparse media content, improving the playback quality and efficiency of media streams.
Smart Images

Figure CN115550745B_ABST
Abstract
Description
[0001] This application is a divisional application of the application with the application date of July 9, 2019, the application number of 201980046379.0, and the invention name of “A method, apparatus and computer readable storage medium for processing media content”. TECHNICAL FIELD
[0002] The present application relates to systems and methods for media streaming. For example, aspects of the present disclosure are directed to time signaling for media streaming. BACKGROUND
[0003] Many devices and systems allow for processing and outputting media data for consumption. Media data can include video data and / or audio data. For example, digital video data can include a large amount of data to meet the needs of consumers and video providers. For example, consumers of video data desire video of optimal quality (with high fidelity, resolution, frame rate, etc.). As a result, the large amount of video data required to meet these demands places a burden on communication networks and devices that process and store the video data.
[0004] Various video coding techniques can be used to compress video data. Video coding is performed in accordance with one or more video coding standards. For example, video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 or ISO / IEC MPEG-4 AVC, including its Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions; and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 and ISO / IEC 23008-2, including its Scalable extension (i.e., Scalable High Efficiency Video Coding, SHVC) and Multiview extension (i.e., Multiview High Efficiency Video Coding, MV-HEVC). Video coding often uses prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit redundancy present in video images or sequences. An important goal of video coding techniques is to compress video data into a form that uses a lower bit rate, while avoiding or minimizing degradations to video quality.
[0005] After video data has been coded, the video data can be encapsulated for transmission or storage. Video data can be assembled into video files that conform to any of a variety of standards, such as the International Organization for Standardization (ISO) Base Media File Format and its extensions such as ITU-T H.264 / AVC. Such encapsulated video data can be transported in a variety of ways, such as transmission through computer networks using network streaming. SUMMARY
[0006] Techniques and systems are described herein for providing temporal signaling for media streams, such as low latency media streams or other media streams. For example, techniques and systems can present samples (e.g., samples of sparse media content or other media content) for which the sample duration can be unknown at the time of decoding the sample. According to some examples, the sample duration of a previous sample can be extended or reduced based on an indication or signaling provided in a current sample. The current sample can include a sample that is currently being processed, and the previous sample can include a sample that was received, decoded, and / or rendered prior to the current sample. In some examples, the previous sample can include sparse content of unknown duration. For example, the previous sample can be a media frame (e.g., a video frame) that includes subtitles or other sparse media content having an unknown duration. A previous segment that includes the previous sample can include a sample duration for the previous sample, where the sample duration is set to a reasonable estimated value.
[0007] Once the current sample is decoded, a modified duration can be obtained, the current sample can include signaling for extending or reducing the sample duration of the previous sample. For example, if a current segment that includes the current sample is decoded at a current time instance, the modified duration can be obtained from the current segment. The modified duration can indicate a duration by which a presentation of the previous sample is to be extended or reduced relative to the current time instance. The at least one media sample can be presented by a player device for a duration based on the modified duration. For example, presenting the at least one media sample can include presenting the previous media sample for an extended duration or presenting a new media sample that starts at the current time instance. In some examples, presenting the at least one media sample can include reducing the sample duration for presenting the previous media sample.
[0008] According to at least one example, a method of processing media content is provided. The method can include obtaining, at a current time instance, a current segment that includes at least a current time component. The method can further include determining, from the current time component, a modified duration for at least one media sample, the modified duration indicating a duration by which a presentation of a previous media sample of a previous segment is to be extended or reduced relative to the current time instance. The method can further include presenting the at least one media sample for a duration based on the modified duration.
[0009] In another example, an apparatus for processing media content is provided. The apparatus includes a memory and a processor implemented in circuitry. The apparatus is configured as and can obtain, at a current time instance, a current segment including at least a current time component. The apparatus is further configured as and can determine, from the current time component, a modified duration for at least one media sample, the modified duration indicating a duration for which presentation of a previous media sample of a previous segment is to be extended or reduced relative to the current time instance. The apparatus is further configured as and can present the at least one media sample for a duration based on the modified duration.
[0010] In another example, a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: obtain, at a current time instance, a current segment including at least a current time component; determine, from the current time component, a modified duration for at least one media sample, the modified duration indicating a duration for which presentation of a previous media sample of a previous segment is to be extended or reduced relative to the current time instance; and present the at least one media sample for a duration based on the modified duration.
[0011] In another example, an apparatus for processing media content is provided. The apparatus includes means for obtaining, at a current time instance, a current segment including at least a current time component; means for determining, from the current time component, a modified duration for at least one media sample, the modified duration indicating a duration for which presentation of a previous media sample of a previous segment is to be extended or reduced relative to the current time instance; and means for presenting the at least one media sample for a duration based on the modified duration.
[0012] In some aspects of the method, apparatuses, and computer-readable medium described above, the modified duration includes an extension duration indicating a duration for which the presentation of the previous media sample is to be extended relative to the current time instance.
[0013] In some aspects of the method, apparatuses, and computer-readable medium described above, the modified duration includes a reduction duration indicating a duration for which the presentation of the previous media sample is to be reduced relative to the current time instance.
[0014] In some aspects of the method, apparatuses, and computer-readable medium described above, presenting the at least one media sample includes extending a duration of the presentation of the previous media sample by at least the extension duration.
[0015] In some aspects of the method, apparatuses, and computer- readable medium described herein, presenting the at least one media sample includes presenting a new media sample for the extended duration at the current time instance.
[0016] In some aspects of the method, apparatuses, and computer- readable medium described herein, presenting the at least one media sample includes reducing a duration of presentation of a previous media sample by the reduced duration.
[0017] In some aspects of the method, apparatuses, and computer- readable medium described herein, the previous media sample is obtained at a previous time instance, the previous time instance being prior to the current time instance.
[0018] In some aspects of the method, apparatuses, and computer- readable medium described herein, the current segment is an empty segment having no media sample data. In some examples, the current segment includes a redundant media sample, where the redundant media sample matches a previous media sample.
[0019] In some aspects of the method, apparatuses, and computer- readable medium described herein, the current segment includes a redundant media sample field to provide an indication of the redundant media sample.
[0020] In some aspects of the method, apparatuses, and computer- readable medium described herein, presenting the at least one media sample includes displaying video content of the at least one media sample.
[0021] In some aspects of the method, apparatuses, and computer- readable medium described herein, presenting the at least one media sample includes presenting audio content of the at least one media sample.
[0022] In some aspects of the method, apparatuses, and computer- readable medium described herein, obtaining the current segment includes receiving and decoding the current segment.
[0023] In some aspects of the method, apparatuses, and computer- readable medium described herein, the current segment includes a track fragment decoding time (tfdt) box, the tfdt box including the current time component.
[0024] In some aspects of the method, apparatuses, and computer- readable medium described herein, the current time component includes a baseMediaDecodeTime value.
[0025] In some aspects of the method, apparatuses, and computer- readable medium described herein, the previous segment includes a sample duration for presenting the previous media sample, and the sample duration includes a predetermined reasonable duration.
[0026] In some aspects of the method, apparatuses, and computer- readable medium described above, the at least one media sample includes sparse content, where a duration for presentation of the sparse content is unknown at a previous time instance when the previous segment is decoded.
[0027] In some aspects of the method, apparatuses, and computer- readable medium described above, the apparatus includes a decoder.
[0028] In some aspects of the method, apparatuses, and computer- readable medium described above, the apparatus includes a player device for presenting media content.
[0029] According to at least one example, a method of providing media content is provided. The method can include providing, at a previous time instance, a previous segment including a previous media sample, where a time for presentation of the previous media sample is unknown at the previous time instance. The method can further include providing, at a current time instance, a current segment including at least a current time component, where the current time component includes a modified duration for the previous media sample, the modified duration indicating a duration by which presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0030] In another example, an apparatus for providing media content is provided. The apparatus includes a memory and a processor implemented in circuitry. The processor is configured to and can provide, at a previous time instance, a previous segment including a previous media sample, where a duration for presentation of the previous media sample is unknown at the previous time instance. The processor is further configured to and can provide, at a current time instance, a current segment including at least a current time component, where the current time component includes a modified duration for the previous media sample, the modified duration indicating a duration by which presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0031] In another example, a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: provide, at a previous time instance, a previous segment including a previous media sample, where a duration for presentation of the previous media sample is unknown at the previous time instance; and provide, at a current time instance, a current segment including at least a current time component, where the current time component includes a modified duration for the previous media sample, the modified duration indicating a duration by which presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0032] In another example, an apparatus for providing media content is provided. The apparatus comprises means for providing, at a previous time instance, a previous segment comprising a previous media sample, wherein a duration for presentation of the previous media sample is unknown at the previous time instance; and means for providing, at a current time instance, a current segment comprising at least a current time component, wherein the current time component comprises a modified duration for the previous media sample, the modified duration indicating a duration by which presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0033] According to at least one example, a method of processing media content is provided. The method comprises obtaining, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data; determining, based on the current time component, a modified duration for at least one media sample, the modified duration indicating a duration by which presentation of a previous media sample of a previous segment of the one or more previous segments is to be extended or reduced relative to the current time instance; and presenting the at least one media sample based on the modified duration.
[0034] According to at least one example, an apparatus for processing media content is provided. The apparatus comprises a memory; and a processor implemented in circuitry and configured to: obtain, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data; determine, based on the current time component, a modified duration for at least one media sample, the modified duration indicating a duration by which presentation of a previous media sample of a previous segment of the one or more previous segments is to be extended or reduced relative to the current time instance; and present the at least one media sample based on the modified duration.
[0035] According to at least one example, a non-transitory computer-readable medium having instructions stored therein that, when executed by one or more processors, cause the one or more processors to: obtain, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data; determine, based on the current time component, a modified duration for at least one media sample, the modified duration indicating a duration by which presentation of a previous media sample of a previous segment of the one or more previous segments is to be extended or reduced relative to the current time instance; and present the at least one media sample based on the modified duration.
[0036] According to at least one example, an apparatus for processing media content is provided. The apparatus comprises: means for obtaining, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data; means for determining, based on the current time component, a modified duration for at least one media sample, the modified duration indicating a duration by which a presentation of a previous media sample of a previous segment of the one or more previous segments is to be extended or reduced relative to the current time instance; and means for presenting the at least one media sample based on the modified duration.
[0037] According to at least one example, a method of providing media content is provided. The method comprises: providing, at a previous time instance, a previous segment comprising a previous media sample, wherein a duration for presenting the previous media sample is unknown at the previous time instance; and providing, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data, and wherein the current time component comprises a modified duration for a previous media sample of a previous segment of the one or more previous segments, the modified duration indicating a duration by which a presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0038] According to at least one example, an apparatus for providing media content is provided. The apparatus comprises: a memory; and a processor implemented in circuitry and configured to: provide, at a previous time instance, a previous segment comprising a previous media sample, wherein a duration for presenting the previous media sample is unknown at the previous time instance; and provide, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data, and wherein the current time component comprises a modified duration for a previous media sample of a previous segment of the one or more previous segments, the modified duration indicating a duration by which a presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0039] According to at least one example, a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform operations including providing, at a previous time instance, a previous segment comprising a previous media sample, wherein a duration for presentation of the previous media sample is unknown at the previous time instance; and providing, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data, and wherein the current time component comprises a modified duration for a previous media sample of a previous media segment of the one or more previous segments, the modified duration indicating a duration by which presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0040] According to at least one example, an apparatus for providing media content is provided. The apparatus comprises means for providing, at a previous time instance, a previous segment comprising a previous media sample, wherein a duration for presentation of the previous media sample is unknown at the previous time instance; and means for providing, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data, and wherein the current time component comprises a modified duration for a previous media sample of a previous media segment of the one or more previous segments, the modified duration indicating a duration by which presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0041] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in determining the scope of the claimed subject matter. The subject matter should be understood from reading this entire specification including the following sections:, any or all drawings and claims.
[0042] These and other features and embodiments will become more apparent from the following specification as it pertains to particular embodiments of the application. BRIEF DESCRIPTION OF DRAWINGS
[0043] Illustrative embodiments of the application are described below with reference to the following drawings:
[0044] Figure 1 to illustrate example block diagrams of an encoding device and a decoding device according to some examples;
[0045] Figure 2FIG. 1 is a diagram to illustrate an example file structure following the ISO base media file format (ISOBMFF) according to some examples;
[0046] Figure 3 FIG. 2 is a diagram to illustrate an example of an ISO base media file containing data and metadata for a video presentation (formatted according to the ISOBMFF) according to some examples;
[0047] Figure 4 FIG. 3 is a diagram to illustrate an example of segments for live streaming according to some examples;
[0048] Figure 5 FIG. 4 is a diagram to illustrate an example of segmenting for low latency live streaming according to some examples;
[0049] Figure 6 FIG. 5 is a diagram to illustrate another example of segmenting for low latency live streaming according to some examples;
[0050] Figure 7 FIG. 6 is a diagram to illustrate an example of a DASH encapsulator for normal operation of audio and video according to some examples;
[0051] Figure 8 FIG. 7 is a diagram to illustrate an example of a media presentation including sparse content according to some examples;
[0052] Figure 9 FIG. 8 is a diagram to illustrate an example of processing media content to reduce sample duration of sparse content according to some examples;
[0053] Figure 10 FIG. 9 is a diagram to illustrate an example of processing media content to extend sample duration of sparse content according to some examples;
[0054] Figure 11 FIG. 10 is a flowchart to illustrate an example of a process for processing media content according to some examples;
[0055] Figure 12 FIG. 11 is a flowchart to illustrate an example of a process for providing media content according to some examples;
[0056] Figure 13 FIG. 12 is a block diagram to illustrate an example video encoding device according to some examples; and
[0057] Figure 14 FIG. 13 is a block diagram to illustrate an example video decoding device according to some examples. DETAILED DESCRIPTION
[0058] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments can be independently applied and some of these aspects and embodiments can be applied in combination, as will be apparent to those skilled in the art. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the present application. It is apparent, however, that various embodiments can be practiced without
[0059] The following description of exemplary embodiments is provided as an enabling teaching only, and is not intended to be limiting of the scope of the application. Indeed, various modifications of the exemplary embodiments can be made that, remain within the scope of the application as defined by the appended claims. The exemplary embodiments can be practiced as solutions appropriate for the particular application or applications contemplated. The preferred embodiments will be described with reference to the figures.
[0060] Video encoding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques can include applying different prediction modes, including spatial prediction (e.g., intra-picture prediction or intra-frame prediction), temporal prediction (e.g., inter-picture prediction or inter-frame prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques, to reduce or remove inherent redundancies in video sequences. A video encoder can partition each picture of an original video sequence into rectangular regions called video blocks or coding units (described in greater detail below). These video blocks can be encoded using a particular prediction mode.
[0061] Video blocks can be divided into one or more groups of smaller blocks in one or more ways. Blocks can include coding tree blocks, prediction blocks, transform blocks, and / or other suitable blocks. Unless otherwise specified, a reference to a "block" can generally refer to these video blocks (e.g., coding tree blocks (CTBs), coding blocks, prediction blocks, transform blocks, or other suitable blocks or sub-blocks, as will be understood by one of ordinary skill in the art). Additionally, each of these blocks can also be interchangeably referred to herein as a "unit" (e.g., coding tree units (CTUs), coding units, prediction units (PUs), transform units (TUs), etc.). In some cases, a unit can indicate a coding logical unit encoded in a bitstream, while a block can indicate a portion of a video frame buffer to which a process is directed.
[0062] For inter prediction modes, the video encoder can search for a block similar to the block being encoded in a frame (or picture) positioned in another temporal location, referred to as a reference frame or reference picture. The search can be limited to a certain spatial displacement from the block to be encoded. A two-dimensional (2D) motion vector comprising a horizontal displacement component and a vertical displacement component can be used to locate the best match. For intra prediction modes, the video encoder can form a predicted block using spatial prediction techniques based on data from previously encoded neighboring blocks within the same picture.
[0063] The video encoder can determine a prediction error. For example, the prediction can be determined as a difference between image samples or pixel values of the block being encoded and the predicted block. The prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error (e.g., a discrete cosine transform (DCT) or other suitable transform) to produce transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements, and, along with control information, form an encoded representation of the video sequence. In some cases, the video encoder can entropy encode the syntax elements, further reducing the number of bits needed to represent them.
[0064] The video decoder can use the syntax elements and control information discussed above to construct predictive data (e.g., predicted blocks) for decoding the current frame. For example, the video decoder can add the predicted blocks to the compressed prediction error. The video decoder can determine the compressed prediction error by weighting a transform basis function using the quantized coefficients. The difference between the reconstructed frame and the original frame is referred to as the reconstruction error.
[0065] The techniques described herein can be applied to any of the existing video codecs (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs), and / or can be efficient coding tools for any video coding standard being developed and / or future video coding standards, such as Versatile Video Coding (VVC), Joint Exploration Model (JEM), and / or other video coding standards in development or to be developed. While video coding is used for purposes of illustration herein, in some cases, the techniques described herein can be performed using any coding device, such as an image encoder (e.g., a JPEG encoder and / or decoder, etc.), a video encoder (e.g., a video encoder and / or a video decoder), or other suitable coding device.
[0066] Figure 1A block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112 is shown. The encoding device 104 can be a part of a source device, and the decoding device 112 can be a part of a receiving device. The source device and / or the receiving device can include an electronic device, such as a mobile or stationary telephone handset (e.g., smart phone, cellular telephone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camcorder, a display device, a digital media player, a video gaming console, a video streaming device, an Internet Protocol (IP) camcorder, or any other suitable electronic device. In some examples, the source device and the receiving device can include one or more wireless transceivers for wireless communication. The encoding techniques described herein are applicable to video encoding in a variety of multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcast or transmission, encoding digital video for storage on a data storage medium, decoding digital video stored on a data storage medium, or other applications. In some examples, the system 100 can support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcast, gaming, and / or video telephony.
[0067] The encoding device 104 (or encoder) can be configured to encode video data using a video coding standard or protocol to generate an encoded video bitstream. Examples of video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including its Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions, and High-Efficiency Video Coding (HEVC) or ITU-T H.265. There are various extensions to HEVC that involve multi-layer video coding, including the range and screen content coding extensions, 3D video coding (3D-HEVC) and multi-view extension (MV-HEVC), and scalable extension (SHVC). HEVC and its extensions have been developed by the Joint Collaboration Team on Video Coding (JCT-VC) of ITU-T Video Coding Experts Group (VCEG) and ISO / IEC Motion Picture Experts Group (MPEG), and the Joint Collaboration Team on 3D Video Coding Extension Development (JCT-3V).
[0068] MPEG and ITU-T VCEG have also formed a Joint Exploration Video Team (JVET) to explore and develop new video coding tools for the next generation of video coding standard, referred to as Versatile Video Coding (VVC). The reference software is referred to as VVC Test Model (VTM). The goal of VVC is to provide a significant improvement in compression efficiency over the existing HEVC standard, to aid the deployment of higher quality video services and emerging applications such as 360° omnidirectional immersive multimedia, high dynamic range (HDR) video, etc.
[0069] Many of the embodiments described herein provide examples using VTM, VVC, HEVC, and / or extensions thereof. However, the techniques and systems described herein can also be applicable to other coding standards, such as AVC, MPEG, JPEG (or other coding standards for still images), extensions thereof, or other suitable coding standards that are already available or not yet available or developed. Thus, while the techniques and systems described herein can be described with reference to a particular video coding standard, one of ordinary skill in the art will appreciate that the description should not be interpreted as applying only to that particular standard.
[0070] Referring to Figure 1 , video source 102 can provide video data to encoding device 104. Video source 102 can be part of the source device, or can be part of a device other than the source device. Video source 102 can include a video capture device (e.g., a video camera, a camcorder phone, a video phone, etc.), a video archive containing stored video, a video server or content provider that provides video data, a video feed interface that receives video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.
[0071] Video data from video source 102 can include one or more input pictures. A picture can also be referred to as a “frame.” A picture or frame is a still image that is part of a video in some cases. In some examples, data from video source 102 can be a still image that is not part of a video. In HEVC, VVC, and other video coding specifications, a video sequence can include a series of pictures. A picture can include three two-dimensional arrays of samples, denoted as S L , S Cb , and S Cr . S L is a two-dimensional array of luma samples, S Cb is a two-dimensional array of Cb chroma samples and S CrA two-dimensional array of Cr chrominance image samples. Chrominance image samples can also be referred to herein as "chroma" image samples. An image sample can refer to an individual component of a pixel (e.g., a luma sample, a chroma blue sample, a chroma red sample, a blue sample, a green sample, a red sample, etc.). A pixel can refer to all components (e.g., including luma and chroma image samples) of a given location in a picture array (e.g., referred to as a pixel position). In other cases, a picture can be monochrome and can only include an array of luma image samples, in which case the terms pixel and image sample can be used interchangeably.
[0072] The encoder engine 106 (or encoder) of the encoding device 104 encodes video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more encoded video sequences. An encoded video sequence (CVS) includes a series of access units (AUs) starting with an AU that has a random access point picture in the base layer and has certain properties until and excluding the next AU that has a random access point picture in the base layer and has certain properties. For example, certain properties of a random access point picture that starts a CVS can include a RASL flag (e.g., NoRaslOutputFlag) equal to 1. Otherwise, a random access point picture (with a RASL flag equal to 0) does not start a CVS. An access unit (AU) includes one or more encoded pictures and control information corresponding to the encoded pictures that share the same output time. Encoded slices of a picture are encapsulated into data units called network abstraction layer (NAL) units at the bitstream level. For example, a HEVC video bitstream can include one or more CVSs that include NAL units. Each of the NAL units has a NAL unit header. In one example, the header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header take up specified bits and are thus visible to all kinds of systems and transport layers, such as transport streams, real-time transport (RTP) protocols, file formats, etc.
[0073] There are two categories of NAL units in the HEVC standard, including video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units include one slice or slice segment (as described below) of encoded picture data, and non-VCL NAL units include control information about one or more encoded pictures. In some cases, a NAL unit can be referred to as a packet. A HEVC AU includes VCL NAL units containing encoded picture data and non-VCL NAL units corresponding to the encoded picture data (if any).
[0074] A NAL unit can include a bit sequence (e.g., an encoded video bitstream, a CVS of a bitstream, etc.) that forms an encoded representation of video data, such as an encoded representation of a picture in a video. The encoder engine 106 produces an encoded representation of a picture by partitioning each picture into multiple slices. A slice is independent of other slices, such that information in the slice is encoded without dependency on data from other slices within the same picture. A slice includes one or more slice segments, including independent slice segments and, if present, one or more dependent slice segments that depend on a preceding slice segment.
[0075] In HEVC, a slice is then partitioned into coding tree blocks (CTBs) of luma and chroma image samples. A CTB of luma image samples and one or more CTBs of chroma image samples, along with syntax of the image samples, are referred to as a coding tree unit (CTU). A CTU can also be referred to as a "tree block" or a "largest coding unit" (LCU). A CTU is a basic processing unit for HEVC encoding. A CTU can be divided into multiple coding units (CUs) of different sizes. A CU contains arrays of luma and chroma image samples referred to as coding blocks (CBs).
[0076] Luma and chroma CBs can be further divided into prediction blocks (PBs). A PB is a block of image samples of a luma component or a chroma component that is inter-predicted or intra-predicted (when available or enabled for use) using the same motion parameters. A luma PB and one or more chroma PBs, along with associated syntax, form a prediction unit (PU). For inter-prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) are signaled in the bitstream for each PU and used for inter-prediction of the luma PB and the one or more chroma PBs. Motion parameters can also be referred to as motion information. A CB can also be partitioned into one or more transform blocks (TBs). A TB represents a square block of image samples of a color component on which the same two-dimensional transform is applied for encoding a prediction residual signal. A transform unit (TU) represents a TB of luma and chroma image samples and corresponding syntax elements.
[0077] The size of a CU corresponds to the size of the coding mode and can be square in shape. For example, the size of a CU can be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other appropriate size up to the size of the corresponding CTU. The phrase "NxN" is used herein to refer to the image sample or pixel dimensions of a video block in terms of vertical and horizontal dimensions (e.g., 8 pixels by 8 pixels or 8 samples by 8 samples). The image samples or pixels in a block can be arranged in columns and rows. In some embodiments, a block can not have the same number of image samples or pixels in the horizontal direction as in the vertical direction. Syntax data associated with a CU can describe, for example, partitioning of the CU into one or more PUs. The partitioning mode can differ between whether the CU is coded with an intra-prediction mode or an inter-prediction mode. The PUs can be partitioned into non-square shapes. Syntax data associated with a CU can also describe, for example, partitioning of the CU into one or more TUs according to a CTU. The TUs can be square or non-square in shape.
[0078] According to the HEVC standard, a transform unit (TU) can be used to perform a transform. The TU can vary for different CUs. The TU can be sized based on the size of the PUs within a given CU. The TU can be the same size or smaller than the PU. In some examples, a quad-tree structure known as a residual quad-tree (RQT) can be used to subdivide residual image samples corresponding to a CU into smaller units. Leaf nodes of the RQT can correspond to TUs. Pixel difference values (or image sample difference values) associated with a TU can be transformed to produce transform coefficients. The transform coefficients can then be quantized by the encoder engine 106.
[0079] Once a picture of video data is partitioned into CUs, the encoder engine 106 predicts each PU using a prediction mode. The predicted units or blocks are then subtracted from the original video data to arrive at residuals (as described below). For each CU, the prediction mode can be signaled inside the bitstream using syntax data. The prediction mode can include intra-prediction (or intra-picture prediction) or inter-prediction (or inter-picture prediction). Intra-prediction uses spatial neighboring image samples within a picture to exploit correlation between neighboring data. For example, when using intra-prediction, each PU is predicted from neighboring image data in the same picture using, for example, DC prediction to find a mean value for the PU, planar prediction to fit a planar surface to the PU, directional prediction to extrapolate from neighboring data, or any other suitable type of prediction. Inter-prediction uses temporal correlation between pictures to derive a motion compensated prediction of a block of image samples. For example, when using inter-prediction, each PU is predicted from image data in one or more reference pictures (that precede or follow the current picture in output order). Whether to use inter- or intra-picture prediction to encode a picture region can be decided, for example, at the CU level.
[0080] The encoder engine 106 and the decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, a video encoder, such as the encoder engine 106 and / or the decoder engine 116, partitions a picture into a plurality of coding tree units (CTUs). The video encoder can partition a CTU according to a tree structure, such as a quad tree binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concepts of multiple split types, such as the separation between CUs, PUs, and TUs of HEVC. The QTBT structure includes two levels: a first level partitioned according to quad tree partitioning, and a second level partitioned according to binary tree partitioning. A root node of the QTBT structure corresponds to a CTU. Leaf nodes of the binary trees correspond to coding units (CUs).
[0081] In the MTT partitioning structure, blocks can be partitioned using quad tree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. Ternary tree partitioning is a partitioning in which a block is divided into three sub-blocks. In some examples, ternary tree partitioning divides a block into three sub-blocks without dividing the original block via a center split. The partitioning types in the MTT (e.g., quad tree, binary tree, and ternary tree) can be symmetric or asymmetric.
[0082] In some examples, a video encoder can use a single QTBT or MTT structure to represent each of luma and chroma components, while in other examples, the video encoder can use two or greater QTBT or MTT structures, such as one QTBT or MTT structure for luma components and another QTBT or MTT structure for two chroma components (or two QTBT and / or MTT structures for respective chroma components).
[0083] In VVC, a picture can be partitioned into slices, tiles, and bricks. Generally, a brick can be a rectangular region of CTU columns within a particular tile in a picture. A tile can be a rectangular region of CTUs within a particular tile row and a particular tile column in a picture. A tile row is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in a picture parameter set. A tile column is a rectangular region of CTUs having a height specified by syntax elements in a picture parameter set and a width equal to the width of the picture. In some cases, a tile can be partitioned into multiple bricks, each of which can include one or more CTU columns within the tile. A tile that is not partitioned into multiple bricks is also referred to as a brick. However, a brick that is a proper subset of a tile is not referred to as a tile. A slice can be an integer number of bricks of a picture contained exclusively in a single NAL unit. In some cases, a slice can include multiple complete tiles or only a contiguous sequence of complete bricks of one tile.
[0084] A video encoder can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures according to HEVC. For illustrative purposes, the description herein can refer to QTBT partitioning. However, it should be understood that the techniques of this disclosure can also apply to video encoders configured to use quadtree partitioning or other types of partitioning.
[0085] In some examples, one or more slices of a picture are assigned a slice type. Slice types include I slices, P slices, and B slices. An I slice (intra, independently decodable) is a slice of a picture that is encoded with only intra prediction, and thus can be independently decodable because an I slice only requires intra data to predict any prediction units or prediction blocks of the slice. A P slice (uni-directional predictive frame) is a slice of a picture that can be encoded with intra prediction and uni-directional inter prediction. Each prediction unit or prediction block within a P slice is encoded with either intra prediction or inter prediction. When inter prediction applies, a prediction unit or prediction block is predicted from only one reference picture, and thus reference image samples come from only one reference region of one frame. A B slice (bi-directional predictive frame) is a slice of a picture that can be encoded with intra prediction and inter prediction (e.g., bi-directional prediction or uni-directional prediction). Prediction units or prediction blocks of a B slice can be bi-directionally predicted from two reference pictures, where each picture contributes one reference region and the image sample sets of the two reference regions are weighted (e.g., with equal weights or with different weights) to produce a prediction signal for the bi-directional prediction block. As explained above, a slice of a picture is independently encoded. In some cases, a picture can be encoded as only one slice.
[0086] As mentioned above, intra prediction uses the correlation between spatially neighboring samples within a picture. Inter prediction uses temporal correlation between pictures in order to derive a motion-compensated prediction of a block of samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector (Ax, Ay), where Ax specifies the horizontal displacement of the reference block relative to the position of the current block, and Ay specifies the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector (Ax, Ay) can be in integer sample accuracy (also referred to as integer accuracy), in which case the motion vector points to an integer pixel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector (Ax, Ay) can have fractional sample accuracy (also referred to as fractional pixel accuracy or non-integer accuracy) to more accurately capture the movement of the underlying object without being limited to the integer pixel grid of the reference frame. The accuracy of the motion vector can be expressed by the quantization level of the motion vector. For example, the quantization level can be integer accuracy (e.g., 1 pixel) or fractional pixel accuracy (e.g., ¼ pixel, ½ pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample accuracy, interpolation is applied to the reference picture to derive the prediction signal. For example, image samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate the values at fractional positions. The previously decoded reference picture is indicated by a reference index (refldx) of a reference picture list. The motion vector and the reference index can be referred to as motion parameters. Two types of inter prediction can be performed, including uni-prediction and bi-prediction.
[0087] In the case that inter prediction uses bi-prediction, two sets of motion parameters (Ax0, y0, refldx0 and Ax1, y1, refldx1) are used to produce two motion-compensated predictions (from the same reference picture or possibly from different reference pictures). For example, with bi-prediction, each prediction block uses two motion-compensated prediction signals, and a B prediction unit is produced. The two motion-compensated predictions are then combined to get the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by taking an average. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference pictures that can be used in bi-prediction are stored in two separate lists, denoted as List 0 and List 1. The motion parameters can be derived at the encoder using a motion estimation process.
[0088] In the case that inter prediction uses uni-prediction, one set of motion parameters (Ax0, y0, refldx0) is used to produce a motion-compensated prediction from a reference picture. For example, with uni-prediction, each prediction block uses at most one motion-compensated prediction signal, and a P prediction unit is produced.
[0089] The PU can include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when the PU is encoded using intra-prediction, the PU can include data describing an intra-prediction mode used for the PU. As another example, when the PU is encoded using inter-prediction, the PU can include data defining a motion vector used for the PU. The data defining the motion vector used for the PU can describe, for example, a horizontal component of the motion vector (Ax), a vertical component of the motion vector (Ay), a resolution for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), a reference picture to which the motion vector points, a reference index, a reference picture list (e.g., List 0, List 1, or List C) for the motion vector, or any combination thereof.
[0090] The encoding device 104 can then perform a transform and quantization. For example, after prediction, the encoder engine 106 can calculate residual values corresponding to the PU. The residual values can include pixel difference values (or image sample difference values) between the current block (PU) of pixels (or image samples) being encoded and a prediction block (e.g., a predicted version of the current block) used to predict the current block. For example, after generating the prediction block (e.g., using inter-prediction or intra-prediction), the encoder engine 106 can generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel difference values (or image sample difference values) quantizing the differences between the pixel values (or image sample values) of the current block and the pixel values (or image sample values) of the prediction block. In some examples, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values or sample values). In such examples, the residual block is a two-dimensional representation of pixel values (or image sample values).
[0091] The residual data remaining after performing prediction can be transformed using a block transform, which can be based on a discrete cosine transform, a discrete sines transform, an integer transform, a wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., size 32x32, 16x16, 8x8, 4x4, or other suitable size) can be applied to the residual data in each CU. In some embodiments, TUs can be used for the transform and quantization processes implemented by the encoder engine 106. A given CU having one or more PUs can also include one or more TUs. As described in further detail below, the residual values can be transformed into transform coefficients using a block transform, and then the TUs can be used to quantize and scan the residual values to produce serialized transform coefficients for entropy encoding.
[0092] In some embodiments, after intra-predictive or inter-predictive encoding using the PUs of a CU, the encoder engine 106 can calculate residual data for the TUs of the CU. The PUs can comprise pixel data (or image samples) in the spatial domain (or pixel domain). After application of the block transform, the TUs can comprise coefficients in the transform domain. As previously mentioned, the residual data can correspond to pixel differences (or image sample differences between image samples) between pixels of the unencoded picture and prediction values corresponding to the PUs. The encoder engine 106 can form the TUs comprising the residual data for the CU, and can then transform the TUs to produce transform coefficients for the CU.
[0093] The encoder engine 106 can perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient having an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.
[0094] After quantization is performed, the encoded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), partitioning information, and any other suitable data, such as other syntax data. The different elements of the encoded video bitstream can then be entropy encoded by the encoder engine 106. In some examples, the encoder engine 106 can scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy encoded. In some examples, the encoder engine 106 can perform an adaptive scan. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 can entropy encode the vector. For example, the encoder engine 106 can use context adaptive variable length coding, context adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval partitioning entropy coding, or another suitable entropy encoding technique.
[0095] The output 110 of the encoding device 104 can send the NAL units that make up the encoded video bitstream data to a decoding device 112 of a receiving device via a communication link 120. The input 114 of the decoding device 112 can receive the NAL units. The communication link 120 can comprise a channel provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The wireless network can comprise any wireless interface or combination of interfaces, and can include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi networks, Bluetooth networks, ZigBee networks, wireless personal area networks, wireless local area networks, or wireless metropolitan area networks). The wired network can comprise any wired interface or combination of interfaces, and can include any suitable wired network (e.g., the Internet or other wide area networks, packet-based networks, or local area networks). TM , radio frequency (RF), UWB, WiFi-Direct, cellular, long term evolution (LTE), WiMax TMA wired network can include any wired interfaces (e.g., fiber, Ethernet, powerline Ethernet, Ethernet over coaxial cable, digital signal line (DSL), etc.). Various equipment such as base stations, routers, access points, bridges, gateways, switches, etc. can be used to implement wired and / or wireless networks. The encoded video bitstream data can be modulated according to a communications standard, such as a wireless communication protocol, and transmitted to a receiving device.
[0096] In some examples, the encoding device 104 can store the encoded video bitstream data in a storage device 108. The output 110 can retrieve the encoded video bitstream data from the encoder engine 106 or from the storage device 108. The storage device 108 can include any of a variety of distributed or locally accessed data storage media. For example, the storage device 108 can include a hard drive, storage optical disc, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.
[0097] The input 114 of the decoding device 112 receives the encoded video bitstream data and can provide the video bitstream data to the decoder engine 116, or to the storage device 118 for later use by the decoder engine 116. The decoder engine 116 can decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more encoded video sequences that make up the encoded video data. The decoder engine 116 can then rescale the encoded video bitstream data and perform inverse transforms on the encoded video bitstream data. The residual data is then passed to a prediction stage of the decoder engine 116. The decoder engine 116 then predicts blocks (e.g., PUs) of pixels or image samples. In some examples, the prediction is added to the inverse transformed output (residual data).
[0098] The decoding device 112 can output the decoded video to a video destination device 122, which can include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, the video destination device 122 can be part of the receiving device that includes the decoding device 112. In some aspects, the video destination device 122 can be part of a separate device than the receiving device.
[0099] In some embodiments, video encoding device 104 and / or video decoding device 112 can be integrated with an audio encoding device and an audio decoding device, respectively. Video encoding device 104 and / or video decoding device 112 can also include other hardware or software necessary to implement the encoding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combination thereof. Video encoding device 104 and video decoding device 112 can be integrated as part of a combined encoder / decoder (CODEC) in the respective device. See below for further details. Figure 13 An example of specific details describing encoding device 104 is described. See below for further details. Figure 14 An example of specific details describing decoding device 112 is described.
[0100] Extensions to the HEVC standard include a multiview video coding extension (referred to as MV-HEVC) and a scalable video coding extension (referred to as SHVC). The MV-HEVC and SHVC extensions share the concept of layered coding, where different layers are included in the encoded video bitstream. Each layer in the encoded video sequence is addressed by a unique layer identifier (ID). The layer ID can be present in the header of a NAL unit to identify the layer with which the NAL unit is associated. In MV-HEVC, different layers can represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided that represent the video bitstream at different spatial resolutions (or picture resolutions) or at different reconstruction fidelity. The scalable layers can include a base layer (where the layer ID = 0) and one or more enhancement layers (where the layer ID = 1, 2,... n). The base layer can conform to the profile of the first version of HEVC and represents the lowest available layer in the bitstream. The enhancement layers have increased spatial resolution, temporal resolution or frame rate, and / or reconstruction fidelity (or quality) compared to the base layer. The enhancement layers are hierarchically organized and can (or can not) depend on lower layers. In some examples, a single standard CODEC can be used to encode the different layers (e.g., all layers are encoded using HEVC, SHVC, or other encoding standard). In some examples, multi-standard CODECs can be used to encode the different layers. For example, the base layer can be encoded using AVC, while one or more enhancement layers can be encoded using SHVC and / or MV-HEVC extensions to the HEVC standard.
[0101] In general, a layer includes a set of VCL NAL units and a corresponding set of non-VCL NAL units. The NAL units are assigned a particular layer ID value. Layers can be hierarchical in the sense that a layer can depend on a lower layer. A layer set refers to a set of independent layers represented within a bitstream, meaning that layers within a layer set can depend on other layers in the layer set but not on any other layers for decoding. Thus, layers in a layer set can form an independent bitstream that can represent video content. A set of layers in a layer set can be obtained from another bitstream through the operation of a sub-bitstream extraction process. A layer set can correspond to a set of layers to be decoded when a decoder wishes to operate according to certain parameters.
[0102] As previously described, an HEVC bitstream includes groups of NAL units, including VCL NAL units and non-VCL NAL units. VCL NAL units include encoded picture data that forms an encoded video bitstream. For example, a bit sequence that forms an encoded video bitstream is present in VCL NAL units. Non-VCL NAL units can also contain parameter sets with high-level information related to the encoded video bitstream, among other information. For example, parameter sets can include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). Examples of the goals of parameter sets include bitrate efficiency, error resilience, and providing a system layer interface. Each slice references a single active PPS, SPS, and VPS to access information that the decoding device 112 can use to decode the slice. An identifier (ID) can be encoded for each parameter set, including a VPS ID, an SPS ID, and a PPS ID. An SPS includes an SPS ID and a VPS ID. A PPS includes a PPS ID and an SPS ID. Each slice header includes a PPS ID. Using the IDs, the active parameter sets for a given slice can be identified.
[0103] A PPS includes information applicable to all slices in a given picture. Thus, all slices in a picture refer to the same PPS. Slices in different pictures can also refer to the same PPS. An SPS includes information applicable to all pictures in the same coded video sequence (CVS) or bitstream. As described previously, a coded video sequence is a series of access units (AUs) that starts with a random access point picture (e.g., an instantaneous decoding reference (IDR) picture or a broken link access (BLA) picture, or other appropriate random access point picture) in the base layer and having certain properties (described above) up to and not including the next AU with a random access point picture in the base layer and having certain properties (or the end of the bitstream). The information in an SPS can not change from picture to picture within a coded video sequence. Pictures in a coded video sequence can use the same SPS. A VPS includes information applicable to all layers within a coded video sequence or bitstream. A VPS includes syntax structures with syntax elements that apply to all coded video sequences. In some embodiments, a VPS, SPS, or PPS can be transmitted in-band with a coded bitstream. In some embodiments, a VPS, SPS, or PPS can be transmitted out-of-band in a separate transmission from NAL units containing coded video data.
[0104] A video bitstream can also include supplemental enhancement information (SEI) messages. For example, SEI NAL units can be part of a video bitstream. In some cases, SEI messages can contain information that is not necessarily required by the decoding process. For example, information in SEI messages can not be important for a decoder to decode video pictures of the bitstream, but the decoder can use the information to improve the display or processing (e.g., decoded output) of the pictures. Information in SEI messages can be embedded metadata. In one illustrative example, information in SEI messages can be used by a decoder-side entity to improve the visibility of content. In some cases, certain application standards can mandate the presence of such SEI messages in bitstreams so that all devices conforming to the application standard can achieve quality improvements (e.g., carrying of frame packing SEI messages for frame-compatible planar stereoscopic 3D TV video formats, where an SEI message is carried for each frame of video; handling of recovery point SEI messages; use of pull-down scan rectangular SEI messages in DVB; among many other examples).
[0105] As previously described, media formats can be used to encapsulate encoded video. One example of a media format includes the ISO Base Media File Format (ISOBMFF, specified in IS0 / IEC 14496-12, which is hereby incorporated by reference in its entirety and for all purposes). There are also other media file formats derived from ISOBMFF (ISO / IEC 14496-12), including the MPEG-4 File Format (ISO / IEC 14496-14), the 3GPP File Format (3GPP TS 26.244), and the AVC File Format (ISO / IEC 14496-15). For example, an encoded video bitstream as discussed above can be written or encapsulated into one or more files using ISOBMFF, a file format derived from ISOBMFF, some other file format, and / or a combination of file formats including ISOBMFF. ISOBMFF files can be played using a video player device, can be transmitted by an encoding device (or file generating device) and then displayed by a player device, can be stored, and / or can be used in any other suitable manner.
[0106] ISOBMFF is used as a basis for many codec encapsulation formats, such as the AVC File Format, and for many multimedia container formats, such as the MPEG-4 File Format, the 3GPP File Format (3GPP), the DVB File Format, and others. Continuous media (e.g., audio and video), still media (e.g., images), and metadata can be stored in files conforming to ISOBMFF. Files structured according to ISOBMFF can be used for many purposes, including local media file playback, progressive download of remote files, segments for Dynamic Adaptive Streaming over HTTP (DASH), containers for content to be streamed and encapsulation instructions for that content, and recording of received real-time media streams, among other suitable purposes. For example, although originally designed for storage, ISOBMFF has proven to be extremely valuable for media streaming (e.g., for progressive download or DASH). For streaming purposes, movie fragments defined in ISOBMFF can be used.
[0107] ISOBMFF is designed to contain timed media information in a flexible and extensible format that facilitates interchange, management, editing, and presentation of media. Presentations of media can be "local" to the system that contains the presentation, or the presentation can be via a network or other streaming delivery mechanism (e.g., DASH or other suitable streaming mechanism). A "presentation" as defined by the ISOBMFF specification can include media files related to a sequence of pictures, often related by having been sequentially captured by a video capture device, or related for some other reason. In some examples, a presentation can also be referred to as a movie, video presentation, or rendering. In some examples, a presentation can include audio. A single presentation can be contained in one or more files, with one file containing metadata for the entire presentation. Metadata includes information such as timing and frame data, descriptors, pointers, parameters, and other information that describes the presentation. Metadata does not itself include video and / or audio data. Files other than the file containing metadata need not be formatted according to ISOBMFF, and only need to be formatted so that they can be referenced by the metadata.
[0108] The file structure of an ISO Base Media File is object-oriented, and the structure of individual objects in the file can be inferred directly from the type of the object. The ISOBMFF specification refers to objects in an ISO Base Media File as "boxes." An ISO Base Media File is constructed as a series of boxes that can contain other boxes. A box is the basic syntax structure in ISOBMFF, including a four-character coded box type, the number of bytes of the box, and the payload. A box can include a header that provides the size and type of the box. The size describes the entire size of the box, including the header, fields, and all boxes contained within the box. Boxes of types that are not recognized by a player device are typically ignored and skipped.
[0109] An ISOBMFF file can contain different kinds of boxes. For example, a movie box ("moov") contains metadata for continuous media streams present in the file, where each media stream is represented as a track in the file. Metadata for a track is enclosed in a track box ("trak"), while the media content of a track is enclosed in a media data box ("mdat") or directly in a separate file. There can be different kinds of tracks. For example, the ISOBMFF specifies the following types of tracks: a media track, which contains a basic media stream; a hint track, which includes media transport instructions or represents a received packet stream; and a timed metadata track, which includes time-synchronized metadata.
[0110] The media content for a track includes a sequence of samples, such as audio or video access units or frames, referred to as media samples. Such media samples are distinguished from the image samples described above, where an image sample is a single color component of a pixel. As used herein, the term "media sample" refers to media data (audio or video) associated with a single time (e.g., a single time point, a range of times, or other time). The metadata for each track includes a list of sample description entries, each providing the encoding or encapsulation format used in the track and initialization data needed to process the format. Each sample is associated with one of the sample description entries for the track.
[0111] The ISOBMFF uses various mechanisms to enable the use of sample-specific metadata. Specific boxes within the sample table box ("stbl") have been standardized to respond to common needs. For example, the sync sample box ("stss") can be used to enumerate the random access samples of a track. The sample grouping mechanism enables mapping of samples into groups that share the same properties, designated as sample group description entries in the file, according to a four-character grouping type. Several grouping types have been specified in the ISOBMFF.
[0112] Figure 2 A diagram showing an example of a file 200 having an example file structure that conforms to the ISO base media file format. The ISO base media file 200 can also be referred to as a media format file. A media presentation can (but does not always) be contained in one file, in which case the media presentation is self-contained in the file. The file 200 includes a movie container 202 (or "movie box"). The movie container 202 can contain metadata for the media, which can include, for example, one or more video tracks and one or more audio tracks. For example, a video track 210 can contain information about the layers of a video, which can be stored in one or more media information containers 214. For example, a media information container 214 can include a sample table that provides information about the video samples of a video. In various implementations, video data chunks 222 and audio data chunks 224 are contained in a media data container 204. In some implementations, the video data chunks 222 and audio data chunks 224 can be contained in one or more other files (other than the file 200).
[0113] In various implementations, a presentation (e.g., a motion sequence) can be contained in several files. All timing and frame (e.g., position and size) information can be in the ISO base media file, and the auxiliary files can use substantially any format.
[0114] An ISO file has a logical structure, a temporal structure, and a physical structure. The different structures do not need to be coupled. The logical structure of the file has a movie that in turn contains a set of temporally parallel tracks (e.g., video track 210). The temporal structure of the file is such that a track contains a sequence of samples in time, and those sequences are mapped to the overall movie schedule by an optional edit list.
[0115] The physical structure of the file separates the data required for logical, temporal, and structural decomposition from the media data samples themselves. This structural information is concentrated in the movie box (e.g., movie container 202), which may be extended temporally by the movie fragment box. The movie box records the logical and temporal relationships of the samples and also contains a pointer to the location where the sample is located. The pointer can point to the same file or to another file, which can be referenced by, for example, a uniform resource locator (URL).
[0116] Each media stream is contained in a playback track specific to that media type. For example, Figure 2 In the example shown in FIG, movie container 202 includes video track 210 and audio track 216. Movie container 202 may also include hint track 218, which may include transport instructions from video track 210 and / or audio track 216, or may represent other information about other tracks in movie container 202 or other movie containers (not shown) in file 200. Each track may be further parameterized by sample entries. For example, in the example shown, video track 210 includes media information container 214, which includes a table of samples (referred to as a "sample table"). A sample entry contains the exact media type (e.g., the type of decoder required to decode the stream) and any parameterized "name" required by the decoder for that media type. The name may be in the form of a four-character code (e.g., moov, trak, or other suitable name code). Defined sample entry formats exist for various media types. A sample entry may further include a pointer to a video data block (e.g., video data block 222) in a logical box 220 in media data container 204. Logic block 220 includes interleaved time-sequenced video samples (organized into video data blocks, such as video data block 222 ), audio frames (eg, in audio data block 224 ), and hint instructions (eg, in hint instruction block 226 ).
[0117] Support for metadata can take different forms. In one example, timing metadata can be stored in the appropriate track and synchronized with the media data described by the metadata as needed. In a second example, there is general support for non-timing metadata attached to movies or individual tracks. Structural support is general and, as in the media data, allows metadata resources to be stored elsewhere in the file or in another file.
[0118] In some cases, one track in a video file can include multiple layers. A video track can also include a track header (e.g., track header 212) that can include some information about the contents of the video track (e.g., video track 210). For example, the track header can include a track content information (also referred to as "tcon") box. The tcon box can list all of the layers and sub-layers in the video track. The video file can also include an operation point information box (also referred to as "oinf" box). The oinf box records information about the operation points, such as the layers and sub-layers that make up the operation points, dependencies between operation points if there are any, configuration file, level, layer parameters of the operation points, and other such operation point related information. In some cases, an operating point can also be referred to as an operation point.
[0119] Figure 3 A diagram is shown to illustrate another example of an ISO base media file 300 formatted according to the ISOBMFF format. The ISO base media file 300 can also be referred to as a media format file. The ISO base media file 300 contains data and metadata for a video presentation. At the top level of the file 300, there is included a file type box 310, a movie box 320, and one or more fragments 330a, 330b, 330c through 330n. Other boxes that can be included at this level but are not represented in this example include a free space box, a metadata box, and a media data box, among others.
[0120] The file type box 310 is identified by the box type "ftyp." The file type box 310 is typically placed as early as possible in the ISO base media file 300. The file type box 310 identifies the ISOBMFF specification that is best suited for parsing the file. "Best" in this case means that the ISO base media file 300 can have been formatted according to a particular ISOBMFF specification, but is likely to be compatible with other iterations of the specification. This best suited specification is referred to as the major brand. A player device can use the major brand to determine whether the device is capable of decoding and displaying the contents of the file. The file type box 310 can also include a version number that can be used to indicate the version of the ISOBMFF specification. The file type box 310 can also include a list of compatible brands that includes a list of other brands that the file is compatible with. An ISO base media file can be compatible with more than one major brand.
[0121] When an ISO Base Media File includes a file type box (as in ISO Base Media File 300), only one file type box exists. In some cases, an ISO Base Media File can omit the file type box in order to be compatible with other early player devices. When an ISO Base Media File does not include a file type box, a player device can assume a pre-set major brand (e.g., mp41), minor version (e.g., "0"), and compatible brands (e.g., mp41, isom, iso2, avcl, etc.).
[0122] ISO Base Media File 300 further includes movie box 320, which contains metadata for a presentation. Movie box 320 is identified by the box type "moov." ISO / IEC 14496-12 specifies that a presentation contained in one file or multiple files can include only one movie box 320. Movie boxes are typically located near the beginning of an ISO Base Media File (e.g., as indicated by the placement of movie box 320 in ISO Base Media File 300). Movie box 320 includes movie header box 322, and can include one or more track box 324, as well as other boxes.
[0123] Movie header box 322, identified by the box type "mvhd," can include information that is media-independent and related to the presentation as a whole. For example, movie header box 322 can include information such as creation time, modification time, time scale, and / or duration of the presentation, etc. Movie header box 322 can also include an identifier that identifies the next track in the presentation. For example, in the illustrated example, the identifier can point to track box 324, which is contained by movie box 320.
[0124] Track box 324, identified by the box type "trak," can contain information for a track of a presentation. A presentation can include one or more tracks, where each track is independent of the other tracks in the presentation. Each track can include temporal and spatial information specific to the content in the track, and each track can be associated with a media box. The data in a track can be media data, in which case the track is a media track; or the data can be encapsulation information for a streaming protocol, in which case the track is a hint track. For example, media data includes video and audio data. In the example shown in FIG. 3, example track box 324 includes track header box 324a and media box 324b. A track box can include other boxes, such as a track reference box, a track group box, an edit box, a user data box, a meta box, etc. Figure 3 In the example shown in FIG. 3, example track box 324 includes track header box 324a and media box 324b. A track box can include other boxes, such as a track reference box, a track group box, an edit box, a user data box, a meta box, etc.
[0125] A track header box 324a identified by the logical box type "tkhd" can specify characteristics of the track contained in the track box 324. For example, the track header box 324a can include the creation time, modification time, duration, track identifier, layer identifier, group identifier, volume, width, and / or height of the track, among others. For a media track, the track header box 324a can further identify whether the track is enabled, whether the track should be played as part of a presentation, whether the track can be used for previewing a presentation, and other uses of the track. A presentation of the track is typically assumed to be at the start of the presentation. The track box 324 can include an edit list box (not shown), which can include an explicit timeline map. The timeline map can specify the offset time of the track, among others, where the offset indicates the start time of the track after the start of the presentation.
[0126] In the illustrated example, the track box 324 also includes a media box 324b identified by the logical box type "mdia." The media box 324b can contain objects and information about the media data in the track. For example, the media box 324b can contain a handler reference box, which can identify the media type of the track and the process by which the media in the track is presented. As another example, the media box 324b can contain a media information box, which can specify characteristics of the media in the track. The media information box can further include a table of samples, as described above with respect to the movie box 322, where each sample description includes a chunk of media data (e.g., video or audio data), such as the location of the data for the sample. The data for the samples is stored in a media data box, which is discussed further below. Like most other boxes, the media box 324b can also include a media header box. Figure 2
[0127] In the illustrated example, the example ISO base media file 300 also includes a presentation of a plurality of segments 330a, 330b, 330c through 330n. Segments can also be referred to as movie fragments. Segments (e.g., which can include Common Media Application Format (CMAF) boxes in some cases) can extend a presentation in time. In some examples, segments can provide information that can have been included in a movie box ("moov"). A movie fragment (or CMAF box) can include at least a movie fragment box (identified by the box type "moof"), followed by a media data box (identified by the box type "mdat"). For example, segment 330a can include a movie fragment ("moof") box 332 and a media data ("mdat") box 338, and can extend a presentation by including additional information that would otherwise be stored in the movie box 320. Segments 330a, 330b, 330c through 330n are not ISOBMFF boxes, but rather describe a movie fragment box and a media data box (e.g., movie fragment box 332 and media data box 338 referenced by movie fragment box 332) referenced by the movie fragment box. Movie fragment box 332 and media data box 338 are top-level boxes, but are grouped here to indicate the relationship between movie fragment box 332 and media data box 338. In cases where movie fragment boxes (e.g., movie fragment box 332) are used, a presentation can be built incrementally.
[0128] In some examples, movie fragment box 332 can include a movie fragment header box 334 and a track fragment box 336, as well as other boxes not shown here. Movie fragment header box 334, identified by the box type "mfhd," can include a sequence number. A sequence number can be used by a player device to verify that segment 330a includes the next segment of data for presentation. In some cases, the contents of a file or a file for presentation can be provided out of order to a player device. For example, network packets can frequently arrive in a different order than the order in which the packets were originally transmitted. In these cases, a sequence number can assist a player device in determining the correct order of segments.
[0129] Movie fragment box 332 can also include one or more track fragment boxes 336, identified by the box type "traf." Movie fragment box 332 can include a set of track fragments (zero or more per track). A track fragment can contain zero or more track runs, each of which describes a contiguous run of samples of a track. In addition to adding samples to a track, a track fragment can be used to add empty time to a track.
[0130] The media data box 338, identified by the logical box type "mdat," contains media data. In a video track, the media data box 338 can contain video frames, access units, NAL units, or other forms of video data. The media data box can alternatively or additionally include audio data. A presentation can include zero or more than zero media data boxes contained in one or more separate files. The media data is described by metadata. In the example shown, the media data in the media data box 338 can be described by metadata included in the track fragment box 336 of the movie fragment box 332. In other examples, the media data in the media data box can be described by metadata in the movie box 320. Metadata can refer to specific media data in terms of absolute offsets within the file 300, such that media data headers and / or free space within the media data box 338 can be skipped.
[0131] Other fragments 330b, 330c through 330n in the ISO base media file 300 can contain similar logical boxes as shown for the first fragment 330a, and / or can contain other logical boxes.
[0132] As mentioned above, the ISOBMFF includes support for streaming media data via a network, in addition to supporting local playback of media. One or more files that include one movie presentation can include additional tracks called hint tracks that contain instructions that can assist a streaming server in forming and transmitting one or more files as packets. For example, the instructions can include data for the server to send (e.g., header information) or references to fragments of media data. A fragment can include a portion of an ISO base media file format file, including a movie box and associated media data and other logical boxes, if present. A fragment can also include a portion of an ISO base media file format file, including one or more movie fragment boxes and associated media data and other logical boxes, if present. A file can include separate hint tracks for different streaming protocols. Hint tracks can also be added to a file without requiring reformatting of the file.
[0133] One approach for streaming data is dynamic adaptive streaming over Hypertext Transfer Protocol (HTTP) or DASH (defined in ISO / IEC 23009-1:2014). DASH, known as MPEG-DASH, is an adaptive bitrate streaming technology that enables high quality streaming of media content using conventional HTTP web servers. DASH operates by breaking media content into a series of HTTP-based small file segments, where each segment contains content for a short time interval. Using DASH, a server can provide media content at different bitrates. A client device playing the media can choose among alternative bitrates as it downloads the next segment, and thus adapt to changing network conditions. DASH uses the HTTP web server infrastructure of the Internet to deliver content over the World Wide Web. DASH is independent of codecs used to encode and decode media content, and thus operates with codecs such as H.264 and HEVC, among others.
[0134] As mentioned above, ISO / IEC 14496-12 ISO Base Media File Format (ISOBMFF) specifies a carriage format for media and is used for many streaming applications, including MPEG DASH, in most cases. These applications of MPEG DASH and Common Media Application Format (CMAF) are also suitable for low latency streaming, which has the goal of reducing file format related delays down to the typical sample duration of audio and video (e.g., in the tens of milliseconds range, as compared to orders of seconds in conventional streaming).
[0135] Conventionally, for real-time streaming applications, "low latency" is used to refer to encapsulation delays on the order of seconds. To achieve this, media files can be segmented into individually addressable segments with a duration of approximately 1 to 2 seconds. Each segment can be addressed, for example, by a uniform resource locator (URL).
[0136] Figure 4A diagram to illustrate another example of an ISO Base Media File 400 formatted according to the ISOBMFF format. The ISO Base Media File 400 can be segmented for live streaming. The ISO Base Media File 400 is shown to contain two example segments 402x and 402y. The segments 402x and 402y can be segments that are sequentially streamed, with segment 402y following segment 402x. In an example, each of the segments 402x and 402y contains a single movie fragment, shown as fragments 430x and 430y, respectively. The fragments 430x and 430y can include respective movie fragment (moof) boxes 432x and 432y, and media data (mdat) boxes 438x and 438y. The separate mdat boxes 438x and 438y can each contain more than one media data sample, shown as samples 1 through M, labeled as samples 440xa through 440xm in mdat box 438x and samples 440ya through 440ym in mdat box 438y, respectively. The samples 1 through M in the mdat boxes 438x through 438y can be time-ordered video samples (e.g., organized into video data chunks) or audio frames.
[0137] File format related latency can be associated with the format of the ISO Base Media File 400, where the data in the fragments 430x and 430y can only be decoded by a decoder or player device after the respective data of the fragments 430x and 430y is completely encoded. Since each of the segments 402x and 402y contains a single fragment 430x and 430y, respectively, the latency for completely encoding the fragment 430x or 430y corresponds to the duration of the respective segment 402x or 402y. The respective moof boxes 432x and 432y of the fragments 430x and 430y contain signaling of the duration, size, etc. for all samples in the respective fragments 430x and 430y. Thus, data from all M samples, including the last sample M (440xm and 440ym) in the respective fragments 430x and 430y, will be needed at the encoder or packager before the respective moof boxes 432x and 432y can be completely written. The moof boxes 432x and 432y will be needed for processing or decoding the respective fragments 430x and 430y by a decoder or player device. Thus, the time for encoding the entire segment data of the segments 402x and 402y can include the duration for encoding all of the samples 1 through M in the mdat boxes 438x through 438y. This duration for encoding all of the samples 1 through M can constitute a significant delay or file format related latency associated with each of the segments 402x through 402y. This type of delay can be present in typical live streaming examples for on-demand playback of video and / or other media systems.
[0138] Figure 5 FIG. 2 is a diagram illustrating another example of an ISO base media file 500 formatted according to the ISOBMFF. The ISO base media file 500 can include optimizations relative to the ISO base media file 400 of FIG. 1. Figure 4 For example, the format of the ISO base media file 500 can result in lower latency compared to the latency associated with encoding all of samples 1 through M in the mdat box 438x through 438y discussed above.
[0139] As shown, the format of the ISO base media file 500 divides each segment into a larger number of fragments (also referred to as "fragmentation" of the segments), such that a smaller number of samples are in each fragment of a segment, and the total number of samples in each segment can be the same or similar to the number of the ISO base media file 400. Since the number of samples in the segments of the ISO base media file 500 can remain the same as in the ISO base media file 400, the fragmentation of the segments in the ISO base media file 500 does not adversely affect the addressing scheme of the samples. In the example shown, segments 502x and 502y are shown in the ISO base media file 500. In some examples, the segments 502x and 502y can be sequentially streamed, with the segment 502y following the segment 502x. The segments 502x and 502y can each include a plurality of fragments, such as fragments 530xa through 530xm included in the segment 502x and fragments 530ya through 530ym included in the segment 502y. Each of these fragments can include a movie fragment (moof) box and a media data (mdat) box, where each mdat box can contain a single sample. For example, the fragments 530xa through 530xm each contain a respective moof box 532xa through 532xm and an mdat box 538xa through 538xm, where each of the mdat boxes 538xa through 538xm contains a respective sample 540xa through 540xm. Similarly, the fragments 530ya through 530ym each contain a respective moof box 532ya through 532ym and an mdat box 538ya through 538ym, where each of the mdat boxes 538ya through 538ym contains a respective sample 540ya through 540ym. While a single sample is shown in each of the mdat boxes 538xa through 538xm and 538ya through 538ym, in some examples it is possible to have a higher, but still low, number of samples in each of the mdat boxes (e.g., 1 to 2 samples per mdat box).
[0140] With a low number of samples contained in each of the mdat boxes 538xa through 538xm and 538ya through 538ym, the format of the ISO base media file 500 can result in lower latency compared to the latency associated with encoding all of samples 1 through M in the mdat box 438x through 438y discussed above. Figure 4compared to fragments 430x to 430y, corresponding fragments 530xa to 530xm and 530ya to 530ym can be decoded at a lower latency or higher speed. Thus, file format related latency can be reduced Figure 4 of each movie fragment can be decoded by the client or player device. For example, the file format related latency of the duration of the complete segment is reduced to the latency of encoding a low number of samples in each fragment. For example, in the illustrated case of a single sample 540xa in fragment 530xa as shown in segment 502x, the latency for decoding fragment 530xa can be based on the duration of single sample 540xa compared to the set duration of multiple samples. Although there can be a small increase in latency for segments 502x to 502y in the case of a larger number of fragments included in segments 502x to 502y, the increase can not be significant for typical high quality media bit rates.
[0141] Figure 6 FIG. is a diagram illustrating another example of an ISO base media file 600 formatted according to the ISOBMFF. ISO base media file 600 can include Figure 5 of ISO base media file 500, where ISO base media file 600 can include a single segment 602 instead of the multiple segments 502x to 502y shown in ISO base media file 500. Figure 5 Segment 602 can be fragmented to include multiple fragments 630a to 630m, where each fragment has a corresponding movie fragment (moof) box 632 to 632m and media data (mdat) box 638a to 638m. Mdat boxes 638a to 638m can each have a low number of samples, such as a single sample. For example, samples 640a to 640m can each be included in a corresponding mdat box 638a to 638m as shown. Similar to the optimizations discussed with respect to Figure 5 ISO base media file 600, Figure 6 Fragmentation in ISO base media file 600 can also enable low latency for decoding each fragment, as the latency is based on the sample duration of a single sample (or a low number of samples).
[0142] While ISO base media file 600 can be used by a player device for presentation of traditional media, such as audio or video, there are challenges involved with sparse media. As will be discussed further below, sparse media can include subtitles, interactive graphics, or other displayable content that can remain constant over multiple segments. In presenting sparse media that remains constant across multiple samples, removing a sample and then possibly providing another sample with the same content can be possible, but can also result in undesirable behavior such as flickering because of the sample being removed and re-presented. For such sparse media or sparse tracks included in samples, such as samples 640a-640m, the related sparse metadata (e.g., moof boxes 632a-632m, mdat boxes 638a-638m, etc.) can need to be customized to address these challenges. For example, it can be desirable to have an indication at the beginning of a sample, segment, segment, or file (e.g., at a random access point) to render the sample, segment, segment, or file until further indicated. For example, indication 604a can identify the beginning of segment 602, indication 604b can identify the beginning of mdat box 638a, and indications 604c, 604d, and 604e can identify the beginning of segments 630b, 630c, and 630m, respectively. However, there is currently no existing mechanism for providing indications, such as indications 604a-604e or others, for ISO base media files formatted according to the ISOBMFF
[0143] As mentioned previously, media data can also be streamed and delivered using dynamic adaptive streaming over hypertext transfer protocol (HTTP) (DASH) using a traditional HTTP web server. For example, it can be possible that each movie segment or concatenation of multiple movie segments is to be delivered using HTTP chunk transfer. HTTP chunk transfer can allow media to be requested by a DASH client device (e.g., a player device). Segments of the requested media can be delivered to the DASH client by a host or origin server before the segment is completely encoded. Allowing such HTTP chunk transfer can reduce the latency or end-to-end delay involved in transferring media data.
[0144] Figure 7 A diagram showing an example of a DASH encapsulator 700. DASH encapsulator 700 can be configured for using HTTP chunk transfer of media data, such as video and / or audio. A server or host device (e.g., an encoder) can transfer media data to be encapsulated by DASH encapsulator 700, where DASH encapsulator can generate HTTP chunks to be transferred to a client or player device (e.g., a decoder). In Figure 7The various media data chunks that can be obtained by the DASH encapsulator 700 from the encoder are shown. These chunks can be provided as Common Media Application Format (CMAF) chunks. These CMAF chunks can include a CMAF header (CH) 706, one or more CMAF initial chunks (CICs) 704a, 704b with random access, and one or more CMAF non-initial chunks (CNCs) 702a, 702b, 702c, and 702d. The CICs 704a, 704b can include media data contained at the beginning of a segment and can be delivered as HTTP chunks to a client. The CNCs 702a-702d can be delivered as HTTP chunks of the same segment.
[0145] For example, the DASH encapsulator 700 can generate DASH segments 720a and 720b, each containing HTTP chunks generated from the media data received from the encoder. For example, in the DASH segment 720a, the HTTP chunks 722a-722c include CNCs 712a-712b and a CIC 714a corresponding to the CNCs 702a-702b and the CIC 704a received from the encoder. Similarly, in the DASH segment 720b, the HTTP chunks 722d-722f include CNCs 712c-712d and a CIC 714b corresponding to the CNCs 702c-702d and the CIC 704b received from the encoder. A media presentation description (MPD) 724 can include a list of instructions (e.g., an extensible markup language (XML) file) that includes information about the media segments, information about the relationships of the media segments (e.g., the order of the media segments) that a client device can use to select between the media segments, and information about other metadata that can be needed by the client device. For example, the MPD 724 can include an address (e.g., a uniform resource location (URL) or other type of address) for each media segment and can also provide an address for an initialization segment (IS) 726. The IS 726 can include information needed to initialize a video decoder on the client device. In some cases, the IS 726 can not be presented.
[0146] However, conventional implementations of the DASH encapsulator 700 are less suitable for low latency optimization that can be involved in delivery and presentation of sparse playitems. For example, each sample in a playitem of transmitted HTTP chunks can have an associated decoding time. The ISOBMFF specifies that the decoding time is encoded as a decoding time delta. For example, the decoding time delta for a sample can include a change (e.g., an increase) relative to the decoding time of a previous sample. These decoding time deltas can be included in metadata associated with the sample. For example, the decoding time delta can be included in a decoding time to sample (Stts) box of a segment that includes the sample. The decoding time delta is specified in subclause 8.6.1.2.1 of the ISOBMFF specification (e.g., ISO / IEC 14496-12) as:
[0147] DT(n + 1) = DT(n) + STTS(n) Equation (1)
[0148] where DT(n) is the decoding time of the current sample "n" and STTS(n) includes the decoding time delta to be added to DT(n) to obtain the decoding time of the next sample "n + 1", DT(n + 1).
[0149] Thus, the decoding time delta can be used to convey the decoding time difference between the current sample and the next sample to a player device. However, the ability to encode the decoding time delta STTS(n) in the current sample n requires knowledge of the decoding time of the subsequent sample DT(n + 1) at the time of encoding the current sample in order to determine the decoding time delta STTS(n) relative to the decoding time of the current sample DT(n). While it can be possible to know or determine the decoding time of subsequent samples for typical media content, sparse media playitems present a challenge in this regard.
[0150] For example, referring back to Figure 6 , the current sample can be, for example, sample 640a of segment 630a, where segment 630a can be the current segment. The next sample can be sample 640b in subsequent segment 630b. For typical media content (such as typical video and / or audio files), the duration of a sample is constant and / or known and / or determinable. If the duration of sample 640a is known and the decoding time of segment 630a is known, then the decoding time of segment 630b and sample 640b included therein can be estimated or determinable. Thus, it can be possible to obtain the decoding time delta between sample 640a and sample 640b and encode the decoding time delta of sample 640a (or metadata of segment 630a). However, for sparse content, it is not determinable to know the duration and / or decoding time delta of sample 640a, as will be further discussed below with respect to Figure 8
[0151] Figure 8 A diagram showing an example of a media presentation 800. The media presentation 800 can be presented in a player device, where the video component of the natural scene is rendered on the player device. The media presentation 800 can also include an audio component, even though it is not shown. The media presentation 800 can also include sparse content, such as a company logo 802 and a caption 804. Other types of sparse content can include interactive graphics, audio content (such as a siren or alarm that sounds for a different duration), etc. A reference to typical media data in the media presentation 800 excludes the sparse content. For example, the typical media data can include video and / or audio related to the scenery being presented, rather than the company logo 802 or the caption 804. In some examples, the sparse content can be overlaid on the media data, but this is not necessary.
[0152] The media data and sparse content in the media presentation 800 can be provided to the player device in the ISO Base Media File Format in some examples. In some examples, the media data can be encapsulated by a DASH encapsulator and delivered as HTTP chunks to the player device. In some examples, the media data can be encoded in samples, such as in the ISO Base Media File 400, 500, 600 of Figure 4 to Figure 6 For example, the media data can be encoded in segments and further fragmented, where each fragment includes one or more samples. For typical media data that contains audio and / or video content, as mentioned above, the duration of each sample can be constant, even though the content of each sample can change.
[0153] However, for sparse content, the data can remain the same for different durations. For example, there can be an indeterminate gap or period of silence in an audio play-out track. In another example, it can not be necessary to update certain sparse content, such as interactive content, where an interactive screen can remain static for a variable period of time. In yet another example, it can not be necessary to update sparse content, such as the caption 804, from one fragment to the next. Thus, for this sparse content, it is a challenge to know the decoding time of the first sample of the subsequent fragment when encoding the current fragment.
[0154] To address the problems associated with delivering low-latency presentations of sparse content, it can be desirable to provide an indication to the player device that it is to hold the sparse content in the presentation of the sample until the player device receives a new sample. However, there is no known mechanism in the current ISOBMFF standard to communicate to the player device that it is to present the sample indefinitely. In the case where the sample duration is not included in the metadata of the sample, the player device will not know how long to present the sample when it decodes the sample. If the estimated sample duration is assumed when the sample is decoded, it is possible that the estimated sample duration can elapse before a new sample is received. For example, in the case of presenting a subtitle 804, if the current sample and the new sample include the same subtitle 804, the player device can stop presenting the subtitle 804 when the estimated sample duration elapses, and then can resume presenting the same subtitle 804 when the new sample is received. This can result in flickering and unnecessary processing load. On the other hand, the estimated sample duration can also be too long such that an error occurs when the sample is presented longer than needed. There is also no known mechanism to indicate to the player device to reduce the sample duration after the sample has been decoded.
[0155] One approach to address this dilemma involves the use of a polling mechanism, where the resolution of time can be defined according to the type of media being presented. For example, the time resolution can be set to 100 milliseconds (ms) for sparse content such as interactive content. When there is no data for the interactive content to be presented, empty samples can be sent at a frequency of one sample per 100 ms during a period of silence. However, there is a significant additional burden introduced by the file format of the ISO base media format file encoded in this manner. For example, over a large time period, encoding samples with no data and transmitting these empty samples for presentation at the player device can result in associated costs. In addition, high quality content can require a higher accuracy or a higher refresh (or update) rate, which means that for sparse content such as subtitles, the time resolution can need to be set to a much lower value to achieve the desired user experience. In addition, the lower update rate based on the larger time resolution can result in a low accuracy presentation, resulting in a poor user experience.
[0156] Another method can include setting a time in a next segment, which can correspond to a cumulative decoding time. The cumulative decoding time can be set in a baseMediaDecodeTime of the next segment, which can be included in metadata of an ISO base media format file. A sum of durations of samples in a current segment is calculated. If the cumulative decoding time set in the baseMediaDecodeTime of the next segment exceeds the sum of the durations of the samples in the current segment, a duration of a last sample of the current segment is extended. The extension of the duration of the last sample of the current segment is designed to make the sum of the durations of the samples in the current segment equal to the cumulative decoding time set in the baseMediaDecodeTime of the next segment. Thus, it is possible to extend the time of the current segment without knowing the decoding time of the next segment.
[0157] In some examples, setting the cumulative decoding time in the baseMediaDecodeTime of the next segment can include setting a duration of a sample of the next segment to a nominal value (e.g., a typical sample duration of a playitem included in the sample). For example, for a video playitem, the sample duration can be determined based on a frame rate of the video playitem. Subsequently, when an actual sample duration of the sample becomes known (e.g., when a next sample arrives), the duration is updated by using signaling included in the baseMediaDecodeTime box.
[0158] However, according to this method, the current segment can be encoded without knowing the decoding time of the next segment. Since a player device or a decoder does not have information about the decoding time of the next segment when decoding the samples in the current segment, the player device can stop presenting the samples in the current segment after the duration of the elapsed samples. The ISOBMFF specification does not currently address this issue. In addition, this method can also encounter a situation of lack of signaling mechanism for reducing the duration of the samples that are currently being presented.
[0159] Systems and methods described herein provide solutions to the problems described above. For example, a sample duration of a previous sample can be extended or reduced based on an indication or signaling provided in a current sample. The current sample can include a sample that is currently being processed, and the previous sample can include a sample that was received, decoded, and / or rendered by a client device before the current sample. The previous sample can include sparse content of unknown duration. A previous segment that includes the previous sample can include a sample duration of the previous sample, where the sample duration is set to a reasonable estimate value. The previous sample can include a sparse playitem of unknown duration. In some examples, the reasonable estimate value can be derived based on a type of the sparse content.
[0160] In some cases, a reasonable estimate value can be based on experience or statistical information about the duration of the type of sparse content. In some examples, a reasonable estimate value can include a nominal estimate value (e.g., a typical sample duration of a playitem included in a sample). For example, for a video playitem, the sample duration can be determined based on the frame rate of the video playitem. Subsequently, when the actual sample duration of the sample becomes known (e.g., when the next sample arrives), the estimate value can be updated by using signaling such as that included in the baseMediaDecodeTime box. For example, after encapsulating a sample, when the accurate sample duration becomes known (e.g., when the next sample arrives), the encapsulator can include signaling to either reduce the signaled sample duration or extend the signaled sample duration.
[0161] For example, once a current sample is decoded, a modified duration can be obtained, which can include signaling to extend or reduce the sample duration of a previous sample. For example, if a current fragment including a current sample is decoded at a current time instance, the modified duration can be obtained from the current fragment. The modified duration can indicate a duration by which a presentation of a previous sample is to be extended or reduced relative to the current time instance. The at least one media sample can be presented by the player device for a duration based on the modified duration. For example, presenting the at least one media sample can include presenting the previous media sample for an extended duration or presenting a new media sample that starts at the current time instance. In some examples, presenting the at least one media sample can include reducing the sample duration for presenting the previous media sample.
[0162] Figure 9 A diagram showing an example of an ISO base media file 900 formatted according to the ISOBMFF. The example ISO base media file 900 can include a plurality of fragments such as fragments 910a-910d. Fragments 910a-910d can be encoded by a host device and decoded and presented by a player device. In some examples, fragments 910a-910d can not be ISOBMFF boxes, but rather describe movie fragment (moof) boxes 902a-902d and media data (mdat) boxes 906a-906d referenced by moof boxes 902a-902d, respectively.
[0163] The moof boxes 902a-d can be extensible presentations such that the presentation can be built incrementally. The moof boxes 902a-d can each include additional fields such as movie fragment header boxes, track fragment boxes, and other boxes not described herein. The moof boxes 902a-d are shown to include respective time boxes 904a-d. The time boxes 904a-d can each contain one or more time structures or values regarding absolute time, relative time, duration, etc.
[0164] In one example, one or more of the time boxes 904a-d can contain a TimeToSampleBox. In some examples, the TimeToSampleBox can include a sample duration. The sample duration is a duration (also referred to as an "increment") in the TimeToSampleBox. The sample duration of a track can refer to the duration of a sample in the track. A track can include a sequence of samples in decoding order. Each sample can have a decoding time calculated by adding the duration of the previous sample (as given by the value in the TimeToSampleBox or equivalent field in the movie fragment) to the decoding time of the previous sample. The decoding time of the first sample in a track or fragment can be defined as at time zero. This forms the decoding timeline of the track. In some examples, the sample duration of a sample in a fragment can be modified based on modification information contained in a subsequent fragment. For example, the modification information can be obtained from a track fragment decoding time (tfdt) contained in the subsequent fragment.
[0165] In some examples, one or more of the time logic boxes 904a-904d can also include a tfdt logic box having one or more tfdt values. Example tfdt values that can be used to signal modification information can include an absolute decode time or a baseMediaDecodeTime. The baseMediaDecodeTime is an integer equal to the sum of the decode durations of all earlier samples in the media, expressed in the time scale of the media. In some examples, the tfdt logic box can include an absolute decode time of a first sample in a playitem fragment in decode order, measured on a decode timeline. The absolute decode time can be useful, for example, when performing random access in a file. For example, in the case of random access, if the absolute decode time is known, it is not necessary to sum the sample durations of all preceding samples in a previous fragment to find the decode time of the first sample in the fragment. In examples where a fragment includes a single sample, the baseMediaDecodeTime or the absolute decode time provided by the tfdt can provide the decode time of the sample. In some examples, one or more of the time logic boxes 904a-904d can include a TrackFragmentBaseMediaDecodeTimeBox, where the tfdt logic box can be present within the TrackFragmentBox container of the TrackFragmentBaseMediaDecodeTimeBox.
[0166] The mdat logic boxes 906a-906d can include, for example, media data included in the respective samples 908a-908d. In a video track, the media data can include video frames, access units, NAL units, and / or other forms of video data. The media can alternatively or additionally include audio data. A presentation can include zero or more media data logic boxes included in one or more separate files. In some examples, the media data in one or more of the samples 908a-908d can include sparse data, such as the company logo 802, the subtitles 804, and / or any other type of sparse content.
[0167] Although not shown, the ISO base media file 900 can include one or more sections, each section having one or more fragments. Each of the fragments can include a single media sample, or in some cases, more than one sample with a known duration can precede each fragment. In the example shown, fragments 910a-910d are shown as each having a single sample. Each fragment has a corresponding one of moof boxes 902a-902d and mdat boxes 906a-906d (e.g., fragment 910a has mdat box 906a and moof box 902a, fragment 910b has mdat box 906b and moof box 902b, etc.). Fragments 910a-910d can be decoded by a player device at times tl-t4, respectively (e.g., fragment 910a can be decoded at time tl, fragment 910b can be decoded at time t2, etc.).
[0168] In an example, fragment 910a decoded at time tl with sample 908a can include media data such as typical media data (video and / or audio related to the scene being presented) or sparse data. In an example, the sample duration of sample 908a can be modified or can remain unmodified. Similarly, fragment 910d decoded at time t4 with sample 908d can include media data such as typical media data or sparse data. In an example, the sample duration of sample 908a can be modified or can remain unmodified. Presentation of samples 908a and 908d can be based on the sample duration or other time information obtained from corresponding time boxes 904a and 908d.
[0169] In an example, the duration of sample 908b decoded at time t2 can be modified. For purposes of illustrating example aspects, fragment 910b can be referred to as a previous fragment and sample 908b therein can be referred to as a previous sample. Sample 908b can include sparse data such as data associated with subtitles, interactive graphics (e.g., logos), or other sparse data.
[0170] According to example aspects, the segment 910b can have a sample duration associated with the sample 908b. The sample duration can be set to a reasonable duration based on an estimate. The estimate of the reasonable duration can be based on the type of sparse data and / or other factors. It can be desirable to modify the duration of the presentation of the sample 908b as needed. In one example, a dynamic need can arise after the decoding time t2 of the previous segment (segment 910b) to reduce the segment duration for content insertion. For example, a need can arise at time t3 after time t2 to insert third party content (e.g., an advertisement, product information, or other data). If the sample duration (set to a reasonable duration) is greater than t3-t2 (in which case the presentation of the previous sample can extend beyond time t3), the sample duration of the previous sample can be reduced. For example, reducing the sample duration of the previous sample can prevent the previous sample from being presented beyond time t3.
[0171] As mentioned above, the sample duration of the previous sample (sample 908b) can be modified to reduce the sample duration. In one illustrative example, to reduce the signaled sample duration (e.g., in the time logic box 904b of the previous sample), a new segment can be provided with modification information. For example, the new segment can include a segment 910c with a decoding time of t3. The segment 910c is referred to as a current segment to illustrate example aspects. The current segment (segment 910c) can include a current time component in a time logic box 910c.
[0172] In one illustrative example, the current time component can be a tfdt that includes an updated decoding time or modified decoding time signaled by the baseMediaDecodeTime field. In examples, the tfdt for the current segment (segment 910c) can be encoded or set to t3. In some examples, the segment 910c does not necessarily include a sample. The mdat logic box 906c and the sample 908c contained therein are shown in dashed logic boxes to indicate that the contents in the segment 910c are optional. If the segment 910c does not contain any sample data, the segment 910c can be referred to as an empty segment. Whether or not sample data is presented in the current segment (segment 910c), a decoder or player device can modify the sample duration of the previous sample based on the tfdt value t3. For example, the player device can update the sample duration of the previous sample (sample 910b) from a reasonable duration (which is set to time t2) to a modified duration. In some cases, the modified duration can correspond to t3-t2. Thus, the modified duration can reduce the sample duration of the previous sample from extending beyond t3.
[0173] In some examples, there can also be a need to extend the sample duration of the previous sample. As will be discussed in more detail below, the need to extend the sample duration of the previous sample can arise after the decoding time t3 of the current segment (segment 910c). For example, a need can arise at time t4 after time t3 to insert third party content (e.g., an advertisement, product information, or other data). If the sample duration of the previous sample (set to a modified duration) is less than t4-t3 (in which case the presentation of the previous sample can not extend beyond time t3), the sample duration of the previous sample can be extended. For example, extending the sample duration of the previous sample can prevent the previous sample from being presented before time t3. Figure 10To further explain, a modified duration can be signaled in a subsequent segment (or the current segment) to extend the duration of a previous sample. In some examples, the current segment can include the same data included in the previous sample to affect the extension. Thus, the current segment can include a current sample referred to as a redundant sample that carries the same data as the previous sample. The current segment can include a field or box with sample has redundancy set to 1, or can include any other suitable value or mechanism that indicates that the current sample is a redundant sample. The current segment can be sent at the time instance at which the previous sample is to be extended. For example, an extension of this nature can be needed when the encapsulator (e.g., DASH encapsulator 700) needs to start and send a new segment but there is no new sample to send.
[0174] Figure 10 A diagram showing an example of an ISO base media file 1000 formatted according to the ISOBMFF. The ISO base media file 1000 is similar in some respects to the ISO base media file 900 of Figure 9 , and thus the following discussion will focus more on the differences between the ISO base media file 1000 and the ISO base media file 900. The ISO base media file 1000 can also include a plurality of segments such as segments 1010a-1010d described by movie fragment (moof) boxes 1002a-1002d and media data (mdat) boxes 1006a-1006d referenced by moof boxes 1002a-1002d, respectively.
[0175] Moof boxes 1002a-1002d are shown to include respective time boxes 1004a-1004d. Time boxes 1004a-1004d can each contain one or more time structures or values regarding absolute time, relative time, duration, etc. In some examples, one or more of time boxes 1004a-1004d can contain a sample duration. In some examples, one or more of time boxes 1004a-1004d can also contain a tfdt box with one or more tfdt values. Example tfdt values that can be used to signal modification information can include an absolute decoding time or a baseMediaDecodeTime.
[0176] Mdat boxes 1006a-1006d can contain media data included in respective samples 1008a-1008d, for example. In some examples, the media data in one or more of samples 1008a-1008d can include sparse data such as a company logo 802 and / or a subtitle 804 from Figure 8 , and / or any other type of sparse content.
[0177] Although not shown, the ISO base media file 1000 can include one or more sections, each having one or more fragments. In the example shown, fragments 1010a-1010d are shown, each having a respective one of moof boxes 1002a-1002d and mdat boxes 1006a-1006d. Fragments 1010a-1010d can be decoded by a player device at times tl-t4, respectively.
[0178] In an example, fragment 1010a, decoded at time tl, having sample 1008a, can include media data such as typical media data (video and / or audio related to a scene being presented) or sparse data. In an example, the sample duration of sample 1008a can be modified or can remain unmodified. Presentation of sample 1008a can be based on the sample duration or other time information obtained from time box 1004a.
[0179] In an example, the duration of sample 1008b, decoded at time t2, can be modified. For purposes of illustrating example aspects, fragment 1010b can be referred to as a previous fragment and sample 1008b therein can be referred to as a previous sample. Sample 1008b can include sparse data.
[0180] According to an example aspect, fragment 1010b can have a sample duration associated with sample 1008b. The sample duration can be set to a reasonable duration based on an estimate. The estimate of a reasonable duration can be based on the type of sparse data and / or other factors. In an example, the reasonable estimate can include a duration that extends from time t2 until a subsequent time t3. It can be desirable to modify the duration of presenting sample 1008b as needed. In an example, the sample duration of sample 1008b can be modified to extend from time t3 to a time t4 after time t3. For example, as previously mentioned, an encapsulator (e.g., DASH encapsulator 700) can need to include an extension of the duration of a previous sample in a previous fragment when a subsequent new fragment is to be sent that does not have sample data. Extending the duration of a previous fragment in this manner can allow a continuous stream to be maintained until data for a new section becomes available. In an example, the sample duration of the previous sample (sample 1008b) can be modified as described below.
[0181] In one illustrative example, the sample duration of the previous sample (sample 1008b) can be modified to increase the sample duration. For example, to increase the sample duration of the previous sample to extend the elapsed time t3, the new segment can be provided with modification information. In an example, the new segment can include segment 1010c with a decoding time of t3. Segment 1010c is referred to as the current segment to illustrate example aspects in which segment 1010c is currently being processed. As previously mentioned, the current segment (segment 1010c) can also include a redundant sample whose sample data matches or repeats data in the previous sample (sample 1010b). The current segment (segment 1010c) can also include a field or logical box (not shown in the figure) with sample has redundancy set to 1, or can include any other suitable value or mechanism that indicates that the current sample is a redundant sample.
[0182] The current segment (segment 1010c) can also include a current time component in the time logical box 1010c. In an example, the current time component can be a tfdt that includes an updated decoding time or modified decoding time signaled by the baseMediaDecodeTime field. In an example, the tfdt for the current segment (segment 1010c) can be set to t4.
[0183] At time t3, the player device or decoder can ignore the redundant sample (sample 1010c) for purposes of decoding and presentation based on, for example, the sample has redundancy field being set to 1. The player device or decoder can also extend the sample duration of the previous sample from the initial value or a reasonable estimate set to t3-t2 to an extended duration. The extended duration can include the duration contained in the tfdt field of the current segment 1010c from time t3 to time t4. Based on the extension, the duration of the previous sample can thus be modified to t4-t2. The use of the redundant sample and the player device's ignoring of the redundant sample for purposes of decoding and presentation can allow for continuous uninterrupted presentation of the previous sample.
[0184] In some examples, another player device can start receiving samples at time t3, after time t2. Based on the fragment 1010c received at time t3, this other player device can also decode and render the samples contained in fragment 1010c from time t3 to time t4. For example, this other player device can ignore the sample has redundancy field set to 1, because this other player device did not receive samples prior to time t3, and thus did not receive the previous sample (sample 1008b). However, because the data in sample 1008c is the same as the data in sample 1008b, the other player device can decode sample 1008c at time t3, and render sample 1008c (which is the same as the previous sample) for the extended duration from t3 to t4.
[0185] In some cases, an exceptional value can be used for low latency rendering of sparse content. For example, an exceptional value can be defined for the sample duration of media samples that contain sparse content or whose sample duration is unknown. The exceptional value can be included in fragments that include these media samples. Once the exceptional value is exactly the sample duration of any media sample, a decoder or player device can render the media sample until the presentation time of the next media sample. One example of an exceptional value can be 0. Another example of an exceptional value can be all 1s.
[0186] The absolute decoding time of the next media sample can then be signaled using a media decoding time logic box (e.g., using a tftd logic box or using its baseMediaDecodeTime value). The presentation time of the next media sample can be based on a composition offset, where the composition offset can include a sample-by-sample mapping of decoding to presentation time. In some cases, the composition offset can be provided in a composition time to sample logic box ("ctts").
[0187] Additionally, it is desirable to provide the ability to signal the termination of rendering at the end of a file or representation, as well as the end of a media item (e.g., a movie, program, or other media item). Signaling the termination of rendering and the end of a media item can be implemented using yet another exception. Such an exception can be referred to as a file end exception. In one illustrative example, such an exception can be implemented by sending a moof with a sample duration (e.g., with just a sample duration) and setting the sample duration to 0, in order to signal that this is the end of the media file. For example, the file end exception can indicate that the sample is the end of the media file.
[0188] At the end of a section, it can be even more difficult to present sparse content, since the section boundary is needed to avoid causing the end of the sample duration. Another exception signal can be provided, signaling the end of the section, but the presentation of the samples of the section will only stop when the instruction to terminate the presentation is not followed by an instruction to resume it again. This exception can be referred to as a section end exception.
[0189] Figure 11 A flowchart illustrating an example of a process 1100 for processing media content is described herein. In some examples, the process 1100 can be performed at a decoding device 112 (e.g., of a decoder, such as the decoder 112 of FIG. 1) or a player device (e.g., of a player device, such as the video destination device 122 of FIG. 1). Figure 1 or Figure 14 of FIG. 1).
[0190] At 1102, the process 1100 includes obtaining, at a current time instance, a current segment including at least a current time component. For example, at time t3, the segment 910c of FIG. 9 Figure 9 may be obtained by the player device, where the segment 910c can include the time 904c. In another example, the segment 1010c of FIG. 10 Figure 10 may be obtained by the player device, where the segment 1010c can include the time 1004c.
[0191] At 1104, the process 1100 includes determining, from the current time component, a modified duration for at least one media sample, the modified duration indicating a duration by which a presentation of a previous media sample of a previous segment is to be extended or reduced relative to the current time instance. In some examples, the previous segment can include a sample duration for presenting the previous media sample, where the sample duration can have been set to a predetermined reasonable duration.
[0192] For example, the player device can determine or decode, from the baseMediaDecodeTime, a time (time t3) contained in the tftd field of the time 904c. The time t3 can correspond to the current time instance. The player device can determine that a sample duration of a previous media sample (sample 908b) contained in a previous segment (segment 910b) is to be reduced relative to the current time instance. For example, as indicated in the sample duration field in the time logic box 904b of the previous segment 910b, the sample duration of the sample 908b can extend past the time t3. Based on the time t3 contained in the current time component, the player device can determine a modified duration indicating a duration by which the sample duration of the previous sample is to be reduced. The modified duration can include reducing the sample duration from extending past the time t3, such that the sample duration of the previous sample is t3-t2.
[0193] In some examples, the current segment can be an empty segment that does not have media sample data. For example, segment 910c can not include Figure 9 The mdat box 906c or the sample 908c shown as optional fields in
[0194] In another example, the player device can determine or decode a time (time t4) contained in the tftd field in the time 1004c from the baseMediaDecodeTime. The player device can determine that the sample duration of a previous media sample (sample 1008b) contained in a previous segment (segment 1010b) is to be extended relative to the current time instance t3. For example, as indicated in the sample duration field in the time box 1004b of the previous segment 1010b, the sample duration of the sample 1008b can be extended through time t3 to time t4. Based on the time t4 contained in the current time component, the player device can determine a modified duration that indicates the duration that the sample duration of the previous sample is to be extended. The modified duration can include extending the sample duration from time t3 to time t4, such that the sample duration of the previous sample is t4-t2.
[0195] In some examples, the current segment can include a redundant media sample, where the redundant media sample matches a previous media sample. For example, sample 1008c can match or contain the same sample data as sample 1008b in Figure 10 In addition, in some examples, the current segment can include a redundant media sample field to provide an indication of the redundant media sample. For example, segment 1010c can include a field such as sample has redundancy, which has a value set to 1 to indicate that sample 1008c is a redundant sample.
[0196] At 1106, the process 1100 includes presenting the at least one media sample for a duration based on the modified duration. For example, the player device can present sample 908b for a duration that is reduced from the sample duration by a reduction duration. For example, the player device can present sample 908b for a duration of t3-t2. In another example, the player device can present sample 1008b for a duration that is extended from the sample duration by an extension duration. For example, the player device can present sample 1008b for a duration of t4-t2. In some examples, the player device can present a new media sample that starts at the current time instance t3 for an extension duration of t4-t3.
[0197] In some examples, the at least one media sample presented by the player device can include sparse content, where a duration for presenting the sparse content is unknown at a previous time instance when a previous segment is decoded. For example, samples 908b and / or 1008b can include sparse content, such as Figure 8 the company logo 802 or the subtitle 804 shown in FIG. 8B. The duration for presenting the sparse content can not be known by the player device at the previous time instance t2 when the segment 910b or 1010b is decoded.
[0198] Figure 12 A flowchart illustrating an example of a process 1200 to provide media content as described herein is shown. In some examples, the process 1200 can be performed at an encoder (e.g., the encoding device 102 of FIG. 1A) or a decoding device (e.g., the decoding device 104 of FIG. 1B). Figure 1 or Figure 13 FIG. 1B).
[0199] At 1202, the process 1200 includes providing, at a previous time instance, a previous segment including a previous media sample, where a duration for presenting the previous media sample is unknown at the previous time instance.
[0200] For example, at the time instance t2 shown in Figure 9 , the segment 910b including the sample 908b can be provided to a decoder or player device. At the time instance t2, the duration for presenting the sample 908b can be unknown. The sample 908b can include sparse content, and the duration can be set at the time instance t2 to a reasonable duration for the sparse content.
[0201] Similarly, in another example, at the time instance t2 shown in Figure 10 , the segment 1010b including the sample 908b can be provided to a decoder or player device. At the time instance t2, the duration for presenting the sample 1008b can be unknown. The sample 1008b can include sparse content, and the duration can be set at the time instance t2 to a reasonable duration for the sparse content.
[0202] At 1204, the process 1200 includes providing, at a current time instance, a current segment including at least a current time component, where the current time component includes a modified duration for the previous media sample, the modified duration indicating a duration by which presentation of the previous media sample is to be extended or reduced relative to the current time instance.
[0203] For example, at the time instance t2 shown in Figure 9At time instance t3 shown in FIG. 10, segment 1010c including temporal logic box 1004c can be provided to a decoder or player device. Temporal logic box 1004c can include time t4, which can indicate that time for presentation of sample 1008b is to be extended to a duration that extends beyond time instance t3 to time instance t4.
[0204] In another example, in Figure 10 At time instance t3 shown in FIG. 10, segment 1010c including temporal logic box 1004c can be provided to a decoder or player device. Temporal logic box 1004c can include time t4, which can indicate that time for presentation of sample 1008b is to be extended to a duration that extends beyond time instance t3 to time instance t4.
[0205] In some implementations, the processes (or methods) described herein can be performed by a computing device or apparatus, such as Figure 1 system 100 shown in FIG. 1. For example, the processes can be performed by Figure 1 and Figure 13 encoding device 104 shown in FIG. 1, by another video source-side device or video transmission device, by Figure 1 and Figure 14 decoding device 112 shown in FIG. 1, and / or by another client-side device, such as a player device, a display, or any other client-side device. In some cases, the computing device or apparatus can include a processor, microprocessor, microcomputer, or other component of a device configured to implement the steps of the processes described herein. In some examples, the computing device or apparatus can include a camera configured to capture video data (e.g., a video sequence) including video frames. In some examples, the camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives or obtains the captured video data. The computing device can further include a network interface configured to communicate the video data. The network interface can be configured to communicate Internet Protocol (IP)-based data or other types of data. In some examples, the computing device or apparatus can include a display for displaying output video content, such as samples of pictures of a video bitstream.
[0206] Components of computing devices (e.g., one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) can be implemented in circuitry. For example, a component can include or be implemented using electronic circuitry or other electronic hardware and / or can include or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.
[0207] Processes can be described with respect to logical flow diagrams, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel with one another to implement the processes.
[0208] Additionally, processes can be performed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As mentioned above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0209] Encoding techniques discussed herein can be implemented in an example video encoding and decoding system (e.g., system 100). In some examples, a system includes a source device that provides encoded video data for later decoding by a destination device. In particular, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device can comprise any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called smartphones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, and the like. In some cases, the source device and the destination device can be equipped for wireless communication.
[0210] The destination device can receive, via the computer-readable medium, the encoded video data to be decoded. The computer-readable medium can comprise any type of medium or device capable of moving the encoded video data from source device to destination device. In one example, computer-readable medium can comprise a communication medium to enable source device to transmit encoded video data directly to destination device in real-time. The encoded video data can be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be used to facilitate communication from source device to destination device.
[0211] In some examples, encoded data can be output from output interface to a storage device. Similarly, encoded data can be accessed from the storage device by input interface. The storage device can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In a further example, the storage device can correspond to a file server or another intermediate storage device that stores the encoded video generated by source device. The destination device can access stored video data from the storage device via streaming or download. The file server can be any type of server capable of storing encoded video data and transmitting that encoded video data to the destination device. Example file servers include web servers (e.g., for a website), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device can access the encoded video data through any standard data connection, including an Internet connection. This connection can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on a file server. The transmission of encoded video data from the storage device can be a streaming transmission, a download transmission, or a combination thereof.
[0212] The techniques of this disclosure are not necessarily limited to wireless applications or contexts. The techniques can be applied to video coding in support of any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions, such as dynamic adaptive streaming over HTTP (DASH), digital video that is encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0213] In one example, a source device includes a video source, a video encoder, and an output interface. A destination device can include an input interface, a video decoder, and a display device. The video encoder of the source device can be configured to apply the techniques disclosed herein. In other examples, a source device and a destination device can include other components or arrangements. For example, the source device can receive video data from an external video source, such as an external camera. Likewise, the destination device can interface with an external display device, rather than include an integrated display device.
[0214] The example system above is merely one example. Techniques for processing video data in parallel can be performed by any digital video encoding and / or decoding device. Although generally the techniques of this disclosure are performed by a video encoding device, the techniques can also be performed by a video encoder / decoder, typically referred to as a "CODEC." Moreover, the techniques of this disclosure can also be performed by a video preprocessor. The source device and destination device are merely examples of such coding devices in which the source device generates encoded video data for transmission to the destination device. In some examples, the source and destination devices can operate in a generally symmetrical manner, such that each of the devices includes video encoding and decoding components. Therefore, the example system can support one-way or two-way transmission between video devices, e.g., for video streaming, video playback, video broadcasting, or video telephony.
[0215] The video source can include a video capture device, such as a video camera, a video archive containing previously captured video, and / or a video feed interface to receive video from a video content provider. As a further alternative, the video source can generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a video camera, the source device and the destination device can form so-called camera phones or video phones. However, as mentioned, the techniques described in this disclosure can be applicable to video coding in general, and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by the video encoder. The encoded video information can then be output by the output interface onto a computer-readable medium.
[0216] As noted, the computer-readable media can include transitory media, such as a wireless broadcast or wired network transmission, or storage media (i.e., non-transitory storage media), such as a hard disk, flash drive, compact disk, digital video disk, Blu-ray disk, or other computer-readable media. In some examples, a network server (not shown) can receive encoded video data from the source device and provide the encoded video data to the destination device, e.g., via network transmission, or a combination of network and device storage. Similarly, a computing device of a medium production facility, such as a disc stamping facility, can receive encoded video data from the source device and produce a disc containing the encoded video data. Therefore, computer-readable media, in various examples, includes one or more computer-readable media of any type.
[0217] The input interface of the destination device receives information from the computer- readable media. The information of the computer-readable media can include syntax information defined by the video encoder, which is also used by the video decoder, that includes syntax elements that describe characteristics and / or processing of blocks and other coded units, e.g., groups of pictures (GOPs). The display device displays the decoded video data to a user, and can comprise any of a variety of display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device. Various embodiments of the application have been described.
[0218] Particular details of the encoding device 104 and the decoding device 112 are shown in Figure 13 and Figure 14 respectively. Figure 13 A block diagram of an example encoding device 104 that can implement one or more of the techniques described in this disclosure is shown. The encoding device 104 may, for example, generate the syntax structures described herein (e.g., syntax structures for VPS, SPS, PPS, or other syntax elements). The encoding device 104 can perform intra prediction and inter prediction encoding of video blocks within a video slice. As described previously, intra coding relies at least in part upon spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter coding relies at least in part upon temporal prediction to reduce or remove temporal redundancy among neighboring or surrounding frames of a video sequence. Intra modes (I modes) can refer to any of a number of spatial-based compression modes. Inter modes, such as uni-prediction (P modes) or bi-prediction (B modes) can refer to any of a number of temporal-based compression modes.
[0219] The encoding device 104 includes a partitioning unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-prediction processing unit 46. For video block reconstruction, the encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. The filter unit 63 is intended to represent one or more loop filters such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 63 is shown in Figure 13 FIG. 1 as a loop filter, in other configurations, the filter unit 63 can be implemented as a post loop filter. A post processing device 57 can perform additional processing on the encoded video data produced by the encoding device 104. In some cases, the techniques of this disclosure can be implemented by the encoding device 104. However, in other cases, one or more of the techniques of this disclosure can be implemented by the post processing device 57.
[0220] As shown in Figure 13 FIG. 1, the encoding device 104 receives video data and the partitioning unit 35 partitions the data into video blocks. The partitioning can also include partitioning into slices, slice segments, pictures, or other larger units according to a quad tree structure of LCUs and CUs, for example, as well as video block partitioning. The encoding device 104 generally shows the components that encode a video block within a slice of video to be encoded. The slice can be divided into a plurality of video blocks (and possibly into sets of video blocks referred to as picture blocks). The prediction processing unit 41 can select one of a plurality of possible encoding modes (such as one of a plurality of intra-prediction encoding modes or one of a plurality of inter-prediction encoding modes) for the current video block based on error results (e.g., rate-distortion levels, etc.). The prediction processing unit 41 can provide the resulting intra- or inter-encoded block to the summer 50 to generate residual block data and to the summer 62 to reconstruct the encoded block for use as a reference picture.
[0221] The intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-predictive encoding of a current video block relative to one or more neighboring blocks in the same frame or slice as the current block being encoded to provide spatial compression. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive encoding of the current video block relative to one or more predictive blocks in one or more reference pictures to provide temporal compression.
[0222] Motion estimation unit 42 can be configured to determine an inter- frame prediction mode for a video slice according to a predetermined pattern for a video sequence. The predetermined pattern can designate video slices in the sequence as P slices, B slices, or GPB slices. Motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are shown separately for conceptual purposes. Motion estimation, performed by motion estimation unit 42, is a process that produces a motion vector that estimates the motion of a video block. A motion vector, for example, can indicate a displacement of a prediction unit (PU) of a video block within a current video frame or intra picture relative to a predictive block within a reference picture.
[0223] A predictive block is a block that is found to closely match a PU of a video block to be encoded in terms of pixel differences (or image sample differences), which can be determined by a sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metric. In some examples, encoding device 104 can calculate values for sub-integer pixel positions of a reference picture stored in picture memory 64. For example, encoding device 104 can interpolate values for quarter-pel positions, eighth-pel positions, or other fractional-pel positions of a reference picture. Thus, motion estimation unit 42 can perform a motion search with respect to full-pel positions and fractional-pel positions and output motion vectors with fractional-pel precision.
[0224] Motion estimation unit 42 calculates motion vectors for PUs of video blocks in a slice that is inter-coded by comparing the locations of the PUs to the locations of predictive blocks of reference pictures. The reference pictures can be selected from a first reference picture list (List 0) or a second reference picture list (List 1), each of which identifies one or more reference pictures stored in picture memory 64. Motion estimation unit 42 sends the calculated motion vectors to entropy encoding unit 56 and motion compensation unit 44.
[0225] Motion compensation, performed by motion compensation unit 44, can involve fetching or generating a predictive block based on a motion vector determined by motion estimation, possibly performing interpolation to sub-pixel precision. Upon receiving a motion vector for a PU of a current video block, motion compensation unit 44 can find the location of the predictive block to which the motion vector points in the reference picture list. Encoding device 104 forms a residual video block by subtracting pixel values (or image sample values) of the predictive block from pixel values of the current video block being encoded (thus forming pixel difference values (or image sample difference values)). The pixel difference values (or image sample difference values) form residual data for the block, and can include both luma and chroma difference components. Summer 50 represents the component or components that perform this subtraction operation. Motion compensation unit 44 can also generate syntax elements associated with video blocks and video slices for use by decoding device 112 in decoding the video blocks of the video slice.
[0226] As described above, intra-prediction processing unit 46 can intra-predict the current block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44. In particular, intra-prediction processing unit 46 can determine an intra-prediction mode to use for encoding the current block. In some examples, intra-prediction processing unit 46 can encode the current block using various intra-prediction modes during separate encoding passes, for example, and intra-prediction processing unit 46 can select an appropriate intra-prediction mode to use from among the tested modes. For example, intra-prediction processing unit 46 can use rate-distortion analysis of the various tested intra-prediction modes to calculate rate-distortion values, and can select the intra-prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original, unencoded block that was encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. Intra-prediction processing unit 46 can calculate a ratio from the distortion and rate of various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion values for that block.
[0227] In any case, after selecting an intra-prediction mode for a block, intra-prediction processing unit 46 can provide information indicative of the selected intra-prediction mode for the block to entropy encoding unit 56. Entropy encoding unit 56 can encode the information indicative of the selected intra-prediction mode. Encoding device 104 can include a definition of the encoding contexts for various blocks, as well as an indication of the most probable intra-prediction mode to be used in each of the contexts, an intra-prediction mode index table, and a modified intra-prediction mode index table in the transmitted bitstream configuration data. The bitstream configuration data can include a plurality of intra-prediction mode index tables and a plurality of modified intra-prediction mode index tables (also referred to as codeword mapping tables).
[0228] After prediction processing unit 41 generates a predictive block for a current video block via inter-prediction or intra-prediction, encoding device 104 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block can be included in one or more TUs and applied to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT), or a conceptually similar transform. Transform processing unit 52 can convert the residual video data from a pixel domain to a transform domain, such as a frequency domain.
[0229] The transform processing unit 52 can send the resulting transform coefficients to quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting a quantization parameter. In some examples, the quantization unit 54 can then perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy encoding unit 56 can perform the scan.
[0230] Following quantization, the entropy encoding unit 56 entropy encodes the quantized transform coefficients. For example, the entropy encoding unit 56 can perform Context- Adaptive Variable Length Coding (CAVLC), Context- Adaptive Binary Arithmetic Coding (CABAC), Syntax-Based Context- Adaptive Binary Arithmetic Coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding, or another entropy encoding technique. Following the entropy encoding by the entropy encoding unit 56, the encoded bitstream can be transmitted to the decoding device 112, or archived for later transmission by or retrieval from the decoding device 112. The entropy encoding unit 56 can also entropy encode motion vectors and other syntax elements for the current video slice being encoded.
[0231] The inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual blocks in the pixel domain for use as reference blocks of reference pictures at a later time. The motion compensation unit 44 can calculate a reference block by adding the residual block to a predictive block of one of the reference pictures within the reference picture list. The motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values (or image sample values) for use in motion estimation. The summer 62 adds the reconstructed residual block to the motion compensated predictive block produced by the motion compensation unit 44 to produce a reference block that is stored in the picture memory 64. The reference block can be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block to inter predict blocks in subsequent video frames or pictures.
[0232] In this way, Figure 13 The encoding device 104 represents an example of a video encoder configured to derive LIC parameters, adaptively determine a size of a template, and / or adaptively select weights. As described above, the encoding device 104 may, for example, derive LIC parameters, adaptively determine a size of a template, and / or adaptively select a set of weights. For example, the encoding device 104 can perform any of the techniques described herein, including the processes described above with reference to Figure 9-12 In some cases, some of the techniques of this disclosure can also be implemented by the post-processing device 57.
[0233] Figure 14A block diagram of an example decoding device 112 is shown. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-prediction processing unit 84. In some examples, the decoding device 112 can perform a decoding pass that is generally reciprocal to the encoding passes described with respect to the encoding device 104 from Figure 13 above.
[0234] During the decoding process, the decoding device 112 receives an encoded video bitstream that represents video blocks of encoded video slices sent by the encoding device 104 and associated syntax elements sent by the encoding device 104. In some embodiments, the decoding device 112 can receive the encoded video bitstream from the encoding device 104. In some embodiments, the decoding device 112 can receive the encoded video bitstream from a network entity 79, such as a server, a media aware network element (MANE), a video editor / splicer, or other such device configured to implement one or more of the techniques described above. The network entity 79 can or can not include the encoding device 104. Some of the techniques described in this disclosure can be implemented by the network entity 79 prior to the network entity 79 transmitting the encoded video bitstream to the decoding device 112. In some video decoding systems, the network entity 79 and the decoding device 112 can be parts of separate devices, while in other cases the functionality described with respect to the network entity 79 can be performed by the same device that includes the decoding device 112.
[0235] The entropy decoding unit 80 of the decoding device 112 entropy decodes the bitstream to produce quantized coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or the video block level. The entropy decoding unit 80 can process and parse both fixed-length syntax elements and variable-length syntax elements in one or more parameter sets, such as VPS, SPS, and PPS.
[0236] When the video slice is coded as an intra-coded (I) slice, intra-prediction processing unit 84 of prediction processing unit 81 can generate prediction data for a video block of the current video slice based on a signaled intra-prediction mode and data from previously decoded blocks of the current frame or picture. When the video frame is coded as an inter-coded (i.e., B, P, or GPB) slice, motion compensation unit 82 of prediction processing unit 81 produces predictive blocks for a video block of the current video slice based on the motion vectors and other syntax elements received from entropy decoding unit 80. The predictive blocks can be produced from one of the reference pictures within the reference picture list. Decoding device 112 can construct the reference frame lists, i.e., List 0 and List 1, based on the reference pictures stored in picture memory 92 using default construction techniques.
[0237] Motion compensation unit 82 determines the prediction information for the video blocks of the current video slice by parsing the motion vectors and other syntax elements, and uses this prediction information to produce predictive blocks for the decoded current video blocks. For example, motion compensation unit 82 can use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra-prediction or inter-prediction) for the video blocks of the coded video slice, the inter-prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information for one or more reference picture lists of the slice, the motion vectors for each inter-coded video block of the slice, the inter-prediction status for each inter-coded video block of the slice, and other information used in decoding the video blocks in the current video slice.
[0238] Motion compensation unit 82 can also perform interpolation based on an interpolation filter. Motion compensation unit 82 can use the interpolation filter used by encoding device 104 during encoding of the video blocks to calculate interpolated values for sub-integer pixels of the reference blocks. In this case, motion compensation unit 82 can determine the interpolation filter used by encoding device 104 from the received syntax elements, and can use the interpolation filter to produce the predictive blocks.
[0239] Inverse quantization unit 86 inverse quantizes, or de-quantizes, the quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 80. The inverse quantization process can include use of a quantization parameter calculated by encoding device 104 for each video block in the video slice to determine a degree of quantization and, likewise, a degree of inverse quantization that should be applied. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT or other suitable inverse transform, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to produce residual blocks in the pixel domain.
[0240] After motion compensation unit 82 generates the predictive block for the current video block based on the motion vectors and other syntax elements, the decoded video block is formed by decoding device 112 by summing the residual block from inverse transform processing unit 88 with the corresponding predictive block generated by motion compensation unit 82. Summer 90 represents the component or components that perform this summation operation. If desired, loop filtering (either in the coding loop or after the coding loop) can also be used to smooth out pixel transitions or otherwise improve the video quality. Filter unit 91 is intended to represent one or more loop filters such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although filter unit 91 is shown in Figure 14 FIG. 5 as a loop filter, in other configurations, filter unit 91 can be implemented as a post loop filter. The decoded video blocks in a given frame or picture are then stored in picture memory 92, which stores reference pictures that are used for subsequent motion compensation. Picture memory 92 also stores the decoded video for later presentation on a display device (such as display 120 shown in Figure 1 FIG. 5).
[0241] In this way, Figure 14 Decoding device 112 represents an example of a video decoder that is configured to derive LIC parameters, adaptively determine a size of a template, and / or adaptively select weights. As described above, decoding device 112 may, for example, derive LIC parameters, adaptively determine a size of a template, and / or adaptively select a set of weights. For example, decoding device 112 can perform any of the techniques described herein, including the processes described above with reference to Figure 9-12 FIG. 4.
[0242] As used herein, the term "computer readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, various types of memory, and the like. A computer readable medium can include a non-transitory medium in which data can be stored and which does not include carrier waves and / or transitory electronic signals propagating through a processing machine or a carrier wave derived from a processing machine. Examples of a non-transitory medium can include, but are not limited to, a magnetic disk or tape, optical storage medium, flash memory, memory or memory devices, and the like. A computer readable medium can have stored thereon code and / or machine-executable instructions that can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0243] In some embodiments, computer readable storage devices, media and memory can include cables or wireless signals containing bitstreams, etc. However, when referred to, non-transitory computer readable storage media expressly excludes media such as energy, carrier waves, electromagnetic waves, and signals per se.
[0244] In the above description, specific details are provided to provide a thorough understanding of the embodiments and examples provided herein. It will be apparent, however, to one skilled in the art that the embodiments can be practiced without these specific details. In some instances, well-known structures and functions have not been described in detail in order to avoid obscuring the concepts of the subject technology. For the sake of clarity, in the description above and in the description of examples, functional blocks are presented as having a generic function that can or can not be implemented as described. In some instances, detailed examples can provide one or more specific implementations in which a function can be implemented. Other examples can provide further implementations in which the function can be implemented differently. This description and the examples provided should not be interpreted as a complete enumeration of all manner of implementations or all manner of uses of the subject technology.
[0245] Individual embodiments can be described as a process or method although the process or method can be embodied in several configurations or modes. (a) The various steps or acts in a process or method can be embodied in computer program code, software, firmware, microcode, and / or hardware, and any combination of the preceding. In some embodiments, a process or method can be embodied in a system, apparatus, or device. (b) The various components of the systems described herein can be implemented, in part, as a set of computer executable instructions executed by a computer, processor, or controller. (c) The computer readable medium can include a floppy disk, CD-ROM, hard disk, optical, computer memory, firmware, physical valve, non-transitory storage, non-transitory processor, etc. The computer readable medium can have stored thereon computer executable instructions or code that, when executed by the computer, processor, or controller, cause the computer, processor, or controller to perform steps described as a process or method. In other embodiments, hard-wired circuitry can be used in place of or in combination with computer executable instructions.
[0246] Processes and methods according to the examples described above can be implemented using computer-executable instructions, which can be stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used by a program can be accessed via a network. Computer-executable instructions can be, for example, binary, intermediate format instructions such as produced by a compiler, assembly language, firmware, source code, or other. Examples of computer-readable media include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-attached storage devices, and the like.
[0247] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks can be stored in a computer-readable or machine-readable medium. A processor(s) can execute the necessary tasks. Typical examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-on
[0248] The instructions, media for transporting those instructions, computing resources for executing those media, and other structures for supporting those computing resources are example components for providing the functionality described in this disclosure.
[0249] In the foregoing description, aspects of the application are described with reference to particular embodiments thereof, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative embodiments of the application have been described in detail herein, it is to be understood that the inventive concepts can be otherwise variously embodied and employed, and that the appended claims are intended to encompass such variations as falling within the spirit and scope of the prior art. Various features and aspects of the above-described application can be used individually or jointly. Further, embodiments can be utilized in any number of environments and applications beyond the type described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative in nature and not as restrictive. The methods described herein can be described in a particular, sequential order, which should not be understood as a restriction unless otherwise specifically indicated. One of ordinary skill in the art will recognize that steps outside the described order are possible and could also be employed. Further, some steps that are, but need not be, executed continuously such as, for example, operations carried out in real time versus those techniques that can be carried out in an iterative manner until a prior operation is completed. Like reference numerals can be used to denote like elements throughout the specification and figures.
[0250] Those of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terms used herein can be replaced by less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of this description.
[0251] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing the electronic circuitry or other hardware of the components to perform the operation, by programming the components (e.g., microprocessors or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0252] The phrase “coupled to” means any direct or indirect electrical, optical, or other physical connection made between entities, and / or any two components that are coupled together to exchange electrical signals (e.g., via wired or wireless connection and / or other suitable communication interface).
[0253] Claim language or other language reciting “at least one of’ a listed set of items indicates that a single item from the set can be used or claimed, or multiple items from the set can be used or claimed, collectively. For example, the claim language “at least one of A and B” means A, B, or A and B.
[0254] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0255] The techniques described herein can be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of various devices such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having many uses including application in wireless communication device handsets and other devices. Any features described as modules or components can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which can include packaging material. The computer-readable medium can comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, can be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0256] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor can be configured to perform any of the techniques described in this disclosure. A general purpose processor can be a microprocessor; but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term "processor," as used herein can refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein can be provided within dedicated software modules or hardware modules configured for encoding and decoding, or incorporated in a combined video encoder-decoder (CODEC).
Claims
1. A method for processing media content, the method comprising: obtaining, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data; determining, based on the current time component, a modified duration for at least one media sample forming part of the one or more previous media samples, the modified duration indicating a duration by which the at least one media sample is to be reduced relative to the current time instance; as well as The at least one media sample is presented based on the modified duration.
2. The method according to claim 1, wherein Presenting the at least one media sample includes reducing a duration of presentation of the at least one media sample by a reduced duration.
3. The method according to claim 1, wherein The previous media sample is obtained at a previous time instance, which is before the current time instance.
4. The method of claim 1 , further comprising: An additional segment is obtained that includes at least an additional temporal component associated with a media sample of an additional previous segment, wherein the additional segment includes a redundant media sample, wherein the redundant media sample matches the media sample of the additional previous segment.
5. The method according to claim 4, wherein: The additional segment includes a redundant media sample field, and the redundant media sample field is used to provide an indication of the redundant media sample.
6. The method of claim 1, wherein: Presenting the at least one media sample includes displaying video content of the at least one media sample.
7. The method of claim 1, wherein: Rendering the at least one media sample includes rendering audio content of the at least one media sample.
8. The method of claim 1, wherein: Obtaining the current segment includes receiving and decoding the current segment.
9. The method of claim 1, wherein: The current segment includes a track segment decode time tfdt logic box, and the tfdt logic box includes the current time component.
10. The method of claim 1, wherein: The current time component includes a baseMediaDecodeTime value.
11. The method of claim 1, wherein: The previous segment comprises a sample duration for presenting the previous media sample, and wherein the sample duration comprises a predetermined reasonable duration.
12. The method of claim 1, wherein: The at least one media sample comprises sparse content, wherein a duration for presenting the sparse content is unknown at a previous time instance when the previous segment was decoded.
13. An apparatus for processing media content, the apparatus comprising: Memory; as well as A processor, implemented in circuitry and configured to: obtaining, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data; determining, based on the current time component, a modified duration for at least one media sample forming part of the one or more previous media samples, the modified duration indicating a duration by which the at least one media sample is to be reduced relative to the current time instance; as well as The at least one media sample is presented based on the modified duration.
14. The apparatus of claim 13, wherein: For presenting the at least one media sample, the processor is configured to reduce a duration of presentation of the at least one media sample by a reduced duration.
15. The apparatus of claim 13, wherein: The previous media sample is obtained at a previous time instance, which is before the current time instance.
16. The apparatus of claim 13, wherein: The processor is configured to: An additional segment is obtained that includes at least an additional temporal component associated with a media sample of an additional previous segment, wherein the additional segment includes a redundant media sample, wherein the redundant media sample matches the media sample of the additional previous segment.
17. The apparatus of claim 16, wherein: The current segment includes a redundant media sample field, and the redundant media sample field is used to provide an indication of the redundant media sample.
18. The apparatus of claim 13, wherein: To present the at least one media sample, the processor is configured to display video content of the at least one media sample.
19. The apparatus of claim 13, wherein: To present the at least one media sample, the processor is configured to present audio content of the at least one media sample.
20. The apparatus of claim 13, wherein: To obtain the current segment, the processor is configured to receive and decode the current segment.
21. The apparatus of claim 13, wherein: The current segment includes a track segment decode time tfdt logic box, and the tfdt logic box includes the current time component.
22. The apparatus of claim 13, wherein: The current time component includes a baseMediaDecodeTime value.
23. The apparatus of claim 13, wherein: The previous segment comprises a sample duration for presenting the previous media sample, and wherein the sample duration comprises a predetermined reasonable duration.
24. The apparatus of claim 13, wherein: The at least one media sample comprises sparse content, wherein a duration for presenting the sparse content is unknown at a previous time instance when the previous segment was decoded.
25. The apparatus of claim 13, wherein: The apparatus comprises a decoder.
26. The apparatus of claim 13, wherein: The apparatus comprises a player device for presenting the media content.
27. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: obtaining, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data; determining, based on the current time component, a modified duration for at least one media sample forming part of the one or more previous media samples, the modified duration indicating a duration by which the at least one media sample is to be reduced relative to the current time instance; as well as The at least one media sample is presented based on the modified duration.
28. The non-transitory computer readable medium of claim 27, wherein: Presenting the at least one media sample includes reducing a duration of presentation of the at least one media sample by a reduced duration.
29. The non-transitory computer readable medium of claim 27, wherein: The previous media sample is obtained at a previous time instance, which is before the current time instance.
30. The non-transitory computer-readable medium of claim 27, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to: An additional segment is obtained that includes at least an additional temporal component associated with a media sample of an additional previous segment, wherein the additional segment includes a redundant media sample, wherein the redundant media sample matches the media sample of the additional previous segment.
31. The non-transitory computer readable medium of claim 30, wherein: The additional segment includes a redundant media sample field, and the redundant media sample field is used to provide an indication of the redundant media sample.
32. The non-transitory computer readable medium of claim 27, wherein: Presenting the at least one media sample includes displaying video content of the at least one media sample.
33. The non-transitory computer readable medium of claim 27, wherein: Rendering the at least one media sample includes rendering audio content of the at least one media sample.
34. The non-transitory computer readable medium of claim 27, wherein: Obtaining the current segment includes receiving and decoding the current segment.
35. The non-transitory computer readable medium of claim 27, wherein: The current segment includes a track segment decode time tfdt logic box, and the tfdt logic box includes the current time component.
36. The non-transitory computer readable medium of claim 27, wherein: The current time component includes a baseMediaDecodeTime value.
37. The non-transitory computer readable medium of claim 27, wherein: The previous segment comprises a sample duration for presenting the previous media sample, and wherein the sample duration comprises a predetermined reasonable duration.
38. The non-transitory computer readable medium of claim 27, wherein: The at least one media sample comprises sparse content, wherein a duration for presenting the sparse content is unknown at a previous time instance when the previous segment was decoded.
39. An apparatus for processing media content, the apparatus comprising: means for obtaining, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data; means for determining, based on the current time component, a modified duration for at least one media sample forming part of the one or more previous media samples, the modified duration indicating a duration by which the at least one media sample is to be reduced relative to the current time instance; as well as Means for presenting the at least one media sample based on the modified duration.
40. The apparatus of claim 39, wherein The means for presenting the at least one media sample comprises means for reducing a duration of presentation of the at least one media sample by a reduced duration.
41. The apparatus of claim 39, wherein: The previous media sample is obtained at a previous time instance, which is before the current time instance.
42. The apparatus of claim 39, further comprising: Means for obtaining an additional segment comprising at least an additional temporal component associated with a media sample of an additional previous segment, wherein the additional segment comprises redundant media samples, wherein the redundant media samples match the media samples of the additional previous segment.
43. The apparatus of claim 42, wherein: The additional segment includes a redundant media sample field, and the redundant media sample field is used to provide an indication of the redundant media sample.
44. The apparatus of claim 39, wherein: The means for presenting the at least one media sample includes means for displaying video content of the at least one media sample.
45. The apparatus of claim 39, wherein: The means for presenting the at least one media sample includes means for presenting audio content of the at least one media sample.
46. The apparatus of claim 39, wherein: The means for obtaining the current segment includes means for receiving the current segment and means for decoding the current segment.
47. The apparatus of claim 39, wherein: The current segment includes a track segment decode time tfdt logic box, and the tfdt logic box includes the current time component.
48. The apparatus of claim 39, wherein The current time component includes a baseMediaDecodeTime value.
49. The apparatus of claim 39, wherein: The previous segment comprises a sample duration for presenting the previous media sample, and wherein the sample duration comprises a predetermined reasonable duration.
50. The apparatus of claim 39, wherein The at least one media sample comprises sparse content, wherein a duration for presenting the sparse content is unknown at a previous time instance when the previous segment was decoded.
51. A method for providing media content, the method comprising: providing a previous segment comprising a previous media sample at a previous time instance, wherein a duration for presenting the previous media sample was unknown at the previous time instance; as well as A current segment is provided at a current time instance that includes at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data, and wherein the current time component includes a modified duration of a previous media sample of a previous media segment of the one or more previous segments, the modified duration indicating a duration by which presentation of the previous media sample is to be reduced relative to the current time instance.
52. An apparatus for providing media content, the apparatus comprising: Memory; as well as A processor, implemented in circuitry and configured to: providing a previous segment comprising a previous media sample at a previous time instance, wherein a duration for presenting the previous media sample was unknown at the previous time instance; as well as A current segment is provided at a current time instance that includes at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data, and wherein the current time component includes a modified duration of a previous media sample of a previous media segment of the one or more previous segments, the modified duration indicating a duration by which presentation of the previous media sample is to be reduced relative to the current time instance.
53. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: providing a previous segment including a previous media sample at a previous time instance, wherein a duration for presenting the previous media sample was unknown at the previous time instance; and A current segment is provided at a current time instance that includes at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data, and wherein the current time component includes a modified duration of a previous media sample of a previous media segment of the one or more previous segments, the modified duration indicating a duration by which presentation of the previous media sample is to be reduced relative to the current time instance.
54. An apparatus for providing media content, the apparatus comprising: means for providing a previous segment comprising a previous media sample at a previous time instance, wherein a duration for presenting the previous media sample is unknown at the previous time instance; as well as Means for providing, at a current time instance, a current segment comprising at least a current time component associated with one or more previous media samples of one or more previous segments, wherein at least one previous segment is an empty segment having no media sample data, and wherein the current time component comprises, for a previous media sample of a previous media segment of the one or more previous segments, a modified duration indicating a duration by which presentation of the previous media sample is to be reduced relative to the current time instance.