Random Access at Resynchronization Points in DASH Segments
By introducing resynchronization points in DASH and CMAF, the problem of random access in segmented non-start positions is solved, efficient retrieval and presentation of media data is realized, and low latency and real-time streaming is supported.
Patent Information
- Application Number
- CN202080066431.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-01
- Filing Date
- 2020-10-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-10-02
AI Technical Summary
The prior art is difficult to implement random access in DASH and CMAF, especially in non-start positions of segments, affecting the efficient retrieval and presentation of media data.
By introducing a resynchronization point in HTTP streaming, random access is allowed anywhere in the segment. The resync point is the beginning of the chunk boundary, contains the parsing points of the file-level container, and notifies the client via signaling in the manifest file.
The ability to access randomly in DASH and CMAF is realized, the retrieval efficiency and presentation quality of media data is improved, and the support of low latency and real-time media streaming is supported.
Smart Images

Figure CN114430911B_ABST
Abstract
Description
[0001] This application claims the benefit of U.S. Application No. 17 / 061,152, filed Oct. 1, 2020, and U.S. Provisional Application No. 62 / 909,642, filed Oct. 2, 2019, the entire contents of which are hereby incorporated by reference. Technical Field
[0002] This disclosure relates to the storage and transmission of encoded video data. Background Art
[0003] Digital video capabilities can be incorporated into a wide variety of devices, including digital televisions, digital live systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, video teleconferencing devices, and the like. Digital video devices implement video compression techniques (such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, or ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Coding (AVC)), ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)), and extensions to such standards) to more efficiently send and receive digital video information.
[0004] Video compression techniques perform spatial prediction and / or temporal prediction to reduce or remove redundancy inherent in a video sequence. For block-based video coding, a video frame or slice can be partitioned into macroblocks. Each macroblock can be further divided. Macroblocks in an intra-coded (I) frame or slice are coded using spatial prediction relative to adjacent macroblocks. Macroblocks in an inter-coded (P or B) frame or slice can use spatial prediction relative to adjacent macroblocks in the same frame or slice, or temporal prediction relative to other reference frames.
[0005] After the video data has been encoded, the video data can be packetized for transmission or storage. The video data can be assembled into a video file that conforms to any of various standards, such as the ISO Base Media File Format and its extensions (such as AVC). Summary of the Invention
[0006] In general, the present disclosure describes techniques for accessing data (e.g., for random access) in segments (not only at the beginning of a segment but also elsewhere within a segment) in HTTP-based Dynamic Adaptive Streaming over HTTP (DASH) and / or Common Media Application Format (CMAF). The present disclosure also describes techniques related to signaling the ability to perform random access within a segment. The present disclosure describes various use cases related to these techniques. For example, the present disclosure defines resync points in DASH and ISO Base Media File Format (BMFF) and the signaling of resync points.
[0007] In one example, a method of retrieving media data includes: retrieving a manifest file for a media presentation, the manifest file indicating that container parsing of media data of a bitstream can start at a resync point of a segment of a representation of the media presentation, the resync point being located at a position other than the beginning of the segment and representing a point at which container parsing of media data of the bitstream can start; using the manifest file to form a request to start retrieving the media data of the representation at the resync point; sending the request to initiate retrieval of the media data of the media presentation at the resync point; and presenting the retrieved media data.
[0008] In another example, a device for retrieving media data includes: a memory configured to store media data of a media presentation; and one or more processors implemented in circuitry and configured to: retrieve a manifest file for a media presentation, the manifest file indicating that container parsing of media data of a bitstream can start at a resync point of a segment of a representation of the media presentation, the resync point being located at a position other than the beginning of the segment and representing a point at which container parsing of media data of the bitstream can start; use the manifest file to form a request to start retrieving the media data of the representation at the resync point; send the request to initiate retrieval of the media data of the media presentation at the resync point; and present the retrieved media data.
[0009] In another example, a computer-readable storage medium stores instructions that, when executed, cause a processor to: retrieve a manifest file for a media presentation, the manifest file indicating that container parsing of media data of a bitstream can start at a resync point of a segment of a representation of the media presentation, the resync point being located at a position other than the beginning of the segment and representing a point at which container parsing of media data of the bitstream can start; use the manifest file to form a request to start retrieving the media data of the representation at the resync point; send the request to initiate retrieval of the media data of the media presentation at the resync point; and present the retrieved media data.
[0010] In another example, a device for retrieving media data includes: a unit for retrieving a manifest file for media presentation, the manifest file indicating that container parsing of the media data of the bitstream can start at a resynchronization point of a segment of the representation of the media presentation, the resynchronization point being located at a position other than the start of the segment and representing a point at which container parsing of the media data of the bitstream can start; a unit for using the manifest file to form a request for retrieving the media data of the representation starting at the resynchronization point; a unit for sending the request to initiate retrieval of the media data of the media presentation starting at the resynchronization point; and a unit for presenting the retrieved media data.
[0011] Details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a block diagram illustrating an example system implementing techniques for streaming media data over a network.
[0013] Figure 2 is a block diagram of an example set of components of a retrieval unit.
[0014] Figure 3 is a conceptual diagram showing elements of an example multimedia content.
[0015] Figure 4 is a block diagram showing elements of an example video file that may correspond to a segment of a representation.
[0016] Figure 5 is a conceptual diagram showing an example low-latency architecture that can be used in a first use case according to the present disclosure.
[0017] Figure 6 is a further detailed conceptual diagram showing an example of the use case described with respect to Figure 5 is a conceptual diagram showing an example of the use case described.
[0018] Figure 7 is a conceptual diagram showing an example second use case of using DASH and CMAF random access in the context of a broadcast protocol.
[0019] Figure 8 is a conceptual diagram showing an example signaling of a stream access point (SAP) in a manifest file.
[0020] Figure 9 is a flowchart showing an example method of retrieving media data according to the techniques of the present disclosure. DETAILED DESCRIPTION
[0021] The technology of the present disclosure can be applied to video files that conform to video data encapsulated according to any one of the following: ISO Base Media File Format, Scalable Video Coding (SVC) File Format, Advanced Video Coding (AVC) File Format, 3rd Generation Partnership Project (3GPP) File Format, and / or Multi-View Video Coding (MVC) File Format, or other similar video file formats.
[0022] In HTTP streaming, frequently used operations include HEAD, GET, and partial GET. The HEAD operation retrieves the header of a file associated with a given Uniform Resource Locator (URL) or Uniform Resource Name (URN) without retrieving the payload associated with the URL or URN. The GET operation retrieves the entire file associated with a given URL or URN. The partial GET operation receives a byte range as an input parameter and retrieves a consecutive number of bytes of the file, where the number of bytes corresponds to the received byte range. Thus, movie fragments can be provided for HTTP streaming because the partial GET operation can obtain one or more individual movie fragments. In a movie fragment, there can be several track fragments of different tracks. In HTTP streaming, a media presentation can be a structured collection of data accessible to a client. The client can request and download media data information to present a streaming service to a user.
[0023] In an example of streaming 3GPP data using HTTP streaming, there can be multiple representations for the video and / or audio data of the multimedia content. As explained below, different representations can correspond to different coding characteristics (e.g., different profiles or levels of a video coding standard), different coding standards or extensions of a coding standard (such as multi-view and / or scalable extensions), or different bitrates. A listing of such representations can be defined in a Media Presentation Description (MPD) data structure. A media presentation can correspond to a structured collection of data accessible to an HTTP streaming client device. The HTTP streaming client device can request and download media data information to present a streaming service to a user of the client device. The media presentation can be described in the MPD data structure, and the MPD data structure can include an update of the MPD.
[0024] A media presentation can include a sequence of one or more periods. Each period can extend until the start of the next period, or until the end of the media presentation (in the case of the last period). Each period can include one or more representations of the same media content. A representation can be one version among several alternative encoded versions of audio, video, timed text, or other such data. Representations can differ in encoding type (e.g., for video data, bitrate, resolution, and / or codec, and for audio data, bitrate, language, and / or codec). The term representation can be used to refer to the portion of the encoded audio or video data that corresponds to a particular period of the multimedia content and is encoded in a particular way.
[0025] A representation of a particular period can be assigned to a group indicated by an attribute in the MPD that indicates the adaptation set to which the representation belongs. Representations within the same adaptation set are generally considered alternatives to each other because a client device can dynamically and seamlessly switch between these representations, e.g., to perform bandwidth adaptation. For example, each representation of video data for a particular period can be assigned to the same adaptation set such that any of these representations can be selected for decoding to present the media data, such as video data or audio data, of the multimedia content for the corresponding period. In some examples, the media content within a period can be represented by one representation (if any) from group 0 or a combination of at most one representation from each non-zero group. The timing data for each representation of a period can be expressed relative to the start time of that period.
[0026] A representation can include one or more segments. Each representation can include an initialization segment, or each segment of a representation can be self-initializing. When present, the initialization segment can contain initialization information for accessing the representation. Generally, the initialization segment does not contain media data. Segments can be uniquely referenced by an identifier, such as a Uniform Resource Locator (URL), Uniform Resource Name (URN), or Uniform Resource Identifier (URI). The MPD can provide an identifier for each segment. In some examples, the MPD can also provide byte ranges in the form of a range attribute, which can correspond to the data of the segment accessible within a file by a URL, URN, or URI.
[0027] For different types of media data, different representations can be selected for substantially simultaneous retrieval. For example, a client device can select an audio representation, a video representation, and a timed text representation from which to retrieve segments. In some examples, the client device can select a particular adaptation set for performing bandwidth adaptation. That is, the client device can select an adaptation set that includes a video representation, an adaptation set that includes an audio representation, and / or an adaptation set that includes timed text. Alternatively, the client device can select an adaptation set for some types of media (e.g., video) and directly select representations for other types of media (e.g., audio and / or timed text).
[0028] HTTP-based Low-Latency Dynamic Adaptive Streaming over HTTP (LL-DASH) is a profile for DASH that attempts to provide media data to DASH clients with low latency. Certain techniques for LL-DASH are briefly outlined below:
[0029] · Encoding is based on fragmented ISO BMFF files and generally assumes CMAF segments and CMAF chunks.
[0030] · Each chunk is individually accessible by a DASH packager and is mapped to an HTTP chunk that is uploaded to the origin server. This 1-to-1 mapping is a recommendation for low-latency operation but not a requirement. In any case, the client should not assume that this 1-to-1 mapping is reserved for the client.
[0031] · A low-latency protocol for partially available segments (e.g., HTTP chunked transfer encoding) is used so that the client can access the segments before they are complete. The availability start time is adjusted for clients that can take advantage of this feature.
[0032] · Allows two operating modes:
[0033] ο Use a simple live product by applying @duration signaling and $Number$-based templating
[0034] ο Support a major live product with a SegmentTimeline that is either $Number$ or $Time$ through an update proposed in DASH version 4.
[0035] · MPD validity expiration events can be used, but the client does not have to understand these events.
[0036] · Generally, in-band event messages may exist, but only the client is expected to resume messages at the start of a segment (not at any chunk). The DASH packer can receive notifications from the encoder completely asynchronously at chunk boundaries or using the timing metadata track.
[0037] · It is allowed to have an adaptation set using the chunk low-latency mode and an adaptation set using short segments for different media types within a single media presentation and within a period of a media presentation.
[0038] · A specific amount of play control on the media pipeline can be available for the DASH client and should be used for the robustness of the DASH client. For example, playback can be accelerated or decelerated within a certain time period, or the DASH client can perform a search for segments.
[0039] · The system is designed to work with standard HTTP / 1.1, but should also be applicable to HTTP extensions and other protocols for improved low-latency operation.
[0040] · The MPD includes explicit signaling about the service configuration and service characteristics (e.g., including the target latency of the service).
[0041] · The MPD and possibly also the segments include an anchor time, which allows the DASH client to measure the current latency compared to real time and make adjustments to meet the service expectations.
[0042] · For example, in the case of an encoder failure, the operational robustness is addressed.
[0043] · Existing DRM and encryption modes are compatible with the proposed low-latency operation.
[0044] Based on the above high-level overview, the following is defined: Segments can be used for random access to the representation not only at segment boundaries but also within segments. If such random access is provided, it should be signaled in the MPD.
[0045] Figure 1 FIG. 10 is a block diagram illustrating an example system 10 that implements techniques for streaming media data over a network. In this example, system 10 includes a content preparation device 20, a server device 60, and a client device 40. The client device 40 and the server device 60 are communicatively coupled via a network 74 that can include the Internet. In some examples, the content preparation device 20 and the server device 60 may also be coupled via network 74 or another network, or may be directly communicatively coupled. In some examples, the content preparation device 20 and the server device 60 may include the same device.
[0046] In Figure 1In the example, the content preparation device 20 includes an audio source 22 and a video source 24. The audio source 22 can include, for example, a microphone that generates an electrical signal representing the captured audio data to be encoded by the audio encoder 26. Alternatively, the audio source 22 can include a storage medium storing previously recorded audio data, an audio data generator (such as a computerized synthesizer), or any other audio data source. The video source 24 can include a camera that generates video data to be encoded by the video encoder 28, a storage medium encoding using previously recorded video data, a video data generation unit (such as a computer graphics source), or any other video data source. In all examples, the content preparation device 20 is not necessarily communicatively coupled to the server device 60, but can store the multimedia content to a separate medium readable by the server device 60.
[0047] The original audio and video data can include analog data or digital data. The analog data can be digitized before being encoded by the audio encoder 26 and / or the video encoder 28. The audio source 22 can obtain audio data from a speaking participant while the speaking participant is speaking, and the video source 24 can simultaneously obtain video data of the speaking participant. In other examples, the audio source 22 can include a computer-readable storage medium that includes stored audio data, and the video source 24 can include a computer-readable storage medium that includes stored video data. In this way, the techniques described in the present disclosure can be applied to live, streaming, real-time audio and video data or to archived, pre-recorded audio and video data.
[0048] An audio frame corresponding to a video frame is typically an audio frame containing audio data that is captured (or generated) by the audio source 22 simultaneously with the video data captured (or generated) by the video source 24 and included in the video frame. For example, when a speaking participant typically generates audio data by speaking, the audio source 22 captures the audio data, and the video source 24 simultaneously (i.e., when the audio source 22 is capturing the audio data) captures video data of the speaking participant. Thus, the audio frame can temporally correspond to one or more specific video frames. Accordingly, an audio frame corresponding to a video frame typically corresponds to a situation where the audio data and the video data are captured simultaneously, and for which the audio frame and the video frame respectively include the simultaneously captured audio data and video data.
[0049] In some examples, the audio encoder 26 may encode a timestamp representing the time at which the audio data for each encoded audio frame is recorded in the encoded audio frame, and similarly, the video encoder 28 may encode a timestamp representing the time at which the video data for each encoded video frame is recorded in the encoded video frame. In such examples, the audio frame corresponding to the video frame may include an audio frame containing the timestamp and a video frame containing the same timestamp. The content preparation device 20 may include an internal clock, and the audio encoder 26 and / or the video encoder 28 may generate timestamps based on the internal clock, or the audio source 22 and the video source 24 may associate the audio data and the video data with timestamps respectively using the internal clock.
[0050] In some examples, the audio source 22 may send data corresponding to the time at which the audio data is recorded to the audio encoder 26, and the video source 24 may send data corresponding to the time at which the video data is recorded to the video encoder 28. In some examples, the audio encoder 26 may encode a sequence identifier in the encoded audio data to indicate the relative temporal ordering of the encoded audio data, but not necessarily the absolute time at which the audio data is recorded, and similarly, the video encoder 28 may also use the sequence identifier to indicate the relative temporal ordering of the encoded video data. Similarly, in some examples, the sequence identifier may be mapped or otherwise related to the timestamp.
[0051] The audio encoder 26 generally produces a stream of encoded audio data, while the video encoder 28 produces a stream of encoded video data. Each individual data stream (whether audio or video) may be referred to as a elementary stream. An elementary stream is a single, decoded (possibly compressed) component of a representation. For example, the decoded video or audio portion of a representation may be an elementary stream. Before being encapsulated within a video file, the elementary stream may be converted to a Packetized Elementary Stream (PES). Within the same representation, a stream ID may be used to distinguish PES packets belonging to one elementary stream from those belonging to another elementary stream. The basic data unit of an elementary stream is a Packetized Elementary Stream (PES) packet. Thus, the decoded video data generally corresponds to an elementary video stream. Similarly, the audio data corresponds to one or more corresponding elementary streams.
[0052] Many video coding standards, such as ITU-T H.264 / AVC and the upcoming High Efficiency Video Coding (HEVC) standard, define the syntax, semantics, and decoding processes for an error-free bitstream, any of which conforms to a certain profile or level. Video coding standards generally do not specify an encoder, but it is the task of the encoder to ensure that the generated bitstream is standard-compatible for the decoder. In the context of video coding standards, a "profile" corresponds to a subset of algorithms, features, or tools and constraints applied to them. As defined by the H.264 standard, for example, a "profile" is a subset of the entire bitstream syntax specified by the H.264 standard. A "level" corresponds to a limit on decoder resource consumption related to the resolution, bit rate, and block processing rate of pictures, such as, for example, decoder memory and computation. A profile can be signaled using a profile_idc (profile indicator) value, while a level can be signaled using a level_idc (level indicator) value.
[0053] For example, the H.264 standard recognizes that within the bounds imposed by the syntax of a given profile, there may still be a need for large variations in the performance of the encoder and decoder, depending on the values taken by the syntax elements in the bitstream, such as the specified size of the decoded pictures. The H.264 standard further recognizes that in many applications, it is neither practical nor economical to implement a decoder that can handle all the assumed uses of the syntax within a particular profile. Therefore, the H.264 standard defines a "level" as a specified set of constraints imposed on the values of the syntax elements in the bitstream. These constraints can be simple limits on values. Alternatively, these constraints can take the form of constraints on arithmetic combinations of values (e.g., picture width times picture height times the number of decoded pictures per second). The H.264 standard also specifies that individual implementations can support different levels for each supported profile.
[0054] A decoder compliant with a profile typically supports all features defined in the profile. For example, as a decoding feature, B picture decoding is not supported in the baseline profile of H.264 / AVC, but is supported in other profiles of H.264 / AVC. A decoder compliant with a level should be able to decode any bitstream that does not require resources beyond the limits defined in that level. The definition of profiles and levels can contribute to interpretability. For example, during a video transmission, a pair of profile and level definitions can be negotiated and agreed upon for the entire transmission session. More specifically, in H.264 / AVC, a level can define limits on: the number of macroblocks that need to be processed, the size of the decoded picture buffer (DPB), the size of the coded picture buffer (CPB), the vertical motion vector range, the maximum number of motion vectors per two consecutive MBs, and whether a B block can have a sub-macroblock partition smaller than 8x8 pixels. In this way, a decoder can determine whether the decoder can correctly decode a bitstream.
[0055] In Figure 1 the example of, the encapsulation unit 30 of the content preparation device 20 receives a elementary stream including coded video data from the video encoder 28, and receives a elementary stream including coded audio data from the audio encoder 26. In some examples, the video encoder 28 and the audio encoder 26 may each include a packetizer for forming PES packets from the encoded data. In other examples, the video encoder 28 and the audio encoder 26 may each interface with a respective packetizer for forming PES packets from the encoded data. In other examples, the encapsulation unit 30 may include a packetizer for forming PES packets from the encoded audio data and video data.
[0056] The video encoder 28 can encode the video data of the multimedia content in various ways to produce different representations of the multimedia content at various bitrates and having various characteristics (such as pixel resolution, frame rate, compliant with various decoding standards, compliant with respective profiles and / or levels of various decoding standards, having a representation with one or more views (e.g., for 2D or 3D playback) or other such characteristics). A representation as used in this disclosure can include one of audio data, video data, text data (e.g., for closed captioning) or other such data. A representation can include an elementary stream, such as an audio elementary stream or a video elementary stream. Each PES packet can include a stream_id identifying the elementary stream to which the PES packet belongs. The encapsulation unit 30 is responsible for assembling the elementary streams into video files (e.g., segments) of respective representations.
[0057] The encapsulation unit 30 receives PES packets for the elementary streams to be represented from the audio encoder 26 and the video encoder 28, and forms corresponding Network Abstraction Layer (NAL) units from the PES packets. The decoded video segments can be organized into NAL units, which provide a "network-friendly" video representation addressed to applications such as videotelephony, storage, broadcast, or streaming. NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units can contain the core compression engine and can include block, macroblock, and / or slice-level data. Other NAL units can be non-VCL NAL units. In some examples, the pictures decoded at a time instance (commonly presented as the base decoded pictures) can be contained in access units, which can include one or more NAL units.
[0058] In addition, non-VCL NAL units can include parameter set NAL units and SEI NAL units. Parameter sets can contain sequence-level header information (in the Sequence Parameter Set (SPS)) and picture-level header information that does not change frequently (in the Picture Parameter Set (PPS)). With parameter sets (e.g., PPS and SPS), it is not necessary to repeat the information that does not change frequently for each sequence or picture; thus, the decoding efficiency can be improved. In addition, the use of parameter sets can enable out-of-band transmission of important header information, thereby avoiding the need for redundant transmission for error recovery. In an out-of-band transmission example, the parameter set NAL units can be sent on a different channel from other NAL units (such as SEI NAL units).
[0059] Supplemental Enhancement Information (SEI) can contain information that is not necessary for decoding the decoded picture samples from VCL NAL units, but may be helpful for processes related to decoding, display, error recovery, and other purposes. SEI messages can be contained in non-VCL NAL units. SEI messages are a normative part of some standard specifications and are not always mandatory for standard-compliant decoder implementations. SEI messages can be sequence-level SEI messages or picture-level SEI messages. Some sequence-level information can be contained in SEI messages (such as the scalability information SEI message in the example of SVC and the view scalability information SEI message in MVC). These example SEI messages can convey information about, for example, the extraction of operation points and the characteristics of operation points. Additionally, the encapsulation unit 30 can form a manifest file, such as a Media Presentation Descriptor (MPD) that describes the characteristics of the representation. The encapsulation unit 30 can format the MPD according to the Extensible Markup Language (XML).
[0060] The encapsulation unit 30 may provide data for one or more representations of multimedia content, as well as a manifest file (e.g., MPD), to the output interface 32. The output interface 32 may include a network interface, or an interface for writing to a storage medium (such as a Universal Serial Bus (USB) interface, a CD or DVD burner or writer, an interface to a magnetic storage medium or a flash storage medium, or other interfaces for storing or transmitting media data). The encapsulation unit 30 may provide the data for each representation in the representation of the multimedia content to the output interface 32, and the output interface 32 may send the data to the server device 60 via a network transmission or a storage medium. In Figure 1 the example of, the server device 60 includes a storage medium 62 storing various multimedia contents 64, each multimedia content 64 including a corresponding manifest file 66 and one or more representations 68A - 68N (representations 68). In some examples, the output interface 32 may also directly send data to the network 74.
[0061] In some examples, the representation 68 may be divided into adaptation sets. That is, the respective subsets of the representation 68 may include corresponding common characteristic sets, such as codec, profile and level, resolution, number of views, file format for segmentation, text type information that may identify a language or other characteristics of the text to be displayed together with the representation and / or the audio data to be decoded and presented, for example, by speakers, camera angle information that may describe the camera angle or real - world perspective of the scene for the representation in the adaptation set, rating information that describes the suitability of the content for a particular audience, etc.
[0062] The manifest file 66 may include data indicating the subset of the representation 68 corresponding to a particular adaptation set and the common characteristics for the adaptation set. The manifest file 66 may also include data representing the individual characteristics (such as bitrate) of the individual representations for the adaptation set. In this way, the adaptation set may provide simplified network bandwidth adaptation. The sub - elements of the adaptation set of the manifest file 66 may be used to indicate the representations in the adaptation set.
[0063] The server device 60 includes a request processing unit 70 and a network interface 72. In some examples, the server device 60 may include multiple network interfaces. Additionally, any or all features of the server device 60 may be implemented on other devices of the content delivery network (such as routers, bridges, proxy devices, switches, or other devices). In some examples, intermediate devices of the content delivery network may cache the data of the multimedia content 64 and include components substantially consistent with the components of the server device 60. Generally, the network interface 72 is configured to send and receive data via the network 74.
[0064] The request processing unit 70 is configured to receive a network request for data of the storage medium 62 from a client device such as the client device 40. For example, the request processing unit 70 may implement the Hypertext Transfer Protocol (HTTP) version 1.1 as described in RFC 2616 (June 1999, IETF, Network Working Group, “Hypertext Transfer Protocol – HTTP / 1.1” by R. Fielding et al.). That is, the request processing unit 70 may be configured to receive an HTTP GET or partial GET request and, in response to the request, provide data of the multimedia content 64. The request may specify the segment, for example, using a segmented URL of one of the representations 68 in the representation 68. In some examples, the request may also specify one or more byte ranges of the segment, thereby including a partial GET request. The request processing unit 70 may also be configured to service an HTTP HEAD request to provide header data of a segment of one of the representations 68. In any case, the request processing unit 70 may be configured to process the request to provide the requested data to the requesting device, such as the client device 40.
[0065] Additionally or alternatively, the request processing unit 70 may be configured to deliver media data via a broadcast or multicast protocol such as eMBMS. The content preparation device 20 may create DASH segments and / or sub-segments in substantially the same manner as described, but the server device 60 may use eMBMS or another broadcast or multicast network transport protocol to deliver these segments or sub-segments. For example, the request processing unit 70 may be configured to receive a multicast group join request from the client device 40. That is, the server device 60 may announce an Internet Protocol (IP) address associated with a multicast group to client devices (including the client device 40), the multicast group being associated with a particular media content (e.g., a broadcast of a live event). The client device 40 may then submit a request to join the multicast group. The request may propagate throughout the network 74 (e.g., routers that make up the network 74), such that the routers direct traffic destined for the IP address associated with the multicast group to the subscribing client devices (such as the client device 40).
[0066] As shown in the example of Figure 1 the multimedia content 64 includes a manifest file 66, which may correspond to a Media Presentation Description (MPD). The manifest file 66 may contain descriptions of different alternative representations 68 (e.g., video services with different qualities), and the descriptions may include, for example, codec information, profile values, level values, bitrates, and other descriptive characteristics of the representations 68. The client device 40 may retrieve the MPD of the media presentation to determine how to access segments of the representations 68.
[0067] Specifically, the retrieval unit 52 may retrieve configuration data (not shown) of the client device 40 to determine the decoding capabilities of the video decoder 48 and the rendering capabilities of the video output 44. The configuration data may also include any one or all of the following: language preferences selected by a user of the client device 40, one or more camera perspectives corresponding to depth preferences set by a user of the client device 40, and / or rating preferences selected by a user of the client device 40. The retrieval unit 52 may include, for example, a web browser or media client configured to submit HTTP GET and partial GET requests. The retrieval unit 52 may correspond to software instructions executed by one or more processors or processing units (not shown) of the client device 40. In some examples, all or part of the functions described with respect to the retrieval unit 52 may be implemented in hardware, or in a combination of hardware, software, and / or firmware, where the necessary hardware may be provided to execute the instructions for the software or firmware.
[0068] The retrieval unit 52 may compare the decoding and rendering capabilities of the client device 40 with the characteristics of the representation 68 indicated by the information in the manifest file 66. The retrieval unit 52 may initially retrieve at least a portion of the manifest file 66 to determine the characteristics of the representation 68. For example, the retrieval unit 52 may request a portion of the manifest file 66 that describes the characteristics of one or more adaptation sets. The retrieval unit 52 may select a subset (e.g., an adaptation set) of the representations 68 that have characteristics that can be satisfied by the decoding and rendering capabilities of the client device 40. The retrieval unit 52 may then determine the bitrate for the representations in the adaptation set, determine the amount of network bandwidth currently available, and retrieve segments from one of the representations that has a bitrate that can be satisfied by the network bandwidth.
[0069] Generally, representations with higher bitrates may produce higher quality video playback, while representations with lower bitrates may provide video playback of sufficient quality when the available network bandwidth decreases. Accordingly, when the available network bandwidth is relatively high, the retrieval unit 52 may retrieve data from a representation with a relatively high bitrate, and when the available network bandwidth is low, the retrieval unit 52 may retrieve data from a representation with a relatively low bitrate. In this way, the client device 40 may stream multimedia data over the network 74 while also adapting to the changing network bandwidth availability of the network 74.
[0070] Additionally or alternatively, the retrieval unit 52 may be configured to receive data according to a broadcast or multicast network protocol such as eMBMS or IP multicast. In such an example, the retrieval unit 52 may submit a request to join a multicast network group associated with a particular media content. After joining the multicast group, the retrieval unit 52 may receive the data of the multicast group without issuing an additional request to the server device 60 or the content preparation device 20. When the data of the multicast group is no longer needed (e.g., stop playback or change the channel to a different multicast group), the retrieval unit 52 may submit a request to leave the multicast group.
[0071] The network interface 54 may receive the segmented data of the selected representation and provide it to the retrieval unit 52, which in turn may provide the segments to the de-encapsulation unit 50. The de-encapsulation unit 50 may de-encapsulate the elements of the video file into the constituent PES streams, de-packetize the PES streams to retrieve the encoded data, and send the encoded data to the audio decoder 46 or the video decoder 48 (depending on whether the encoded data is part of an audio stream or a video stream, e.g., as indicated by the PES packet header of the stream). The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, while the video decoder 48 decodes the encoded video data and sends the decoded video data (which may include multiple views of the stream) to the video output 44.
[0072] The video encoder 28, the video decoder 48, the audio encoder 26, the audio decoder 46, the encapsulation unit 30, the retrieval unit 52, and the de-encapsulation unit 50 may each be implemented as any one of a variety of suitable processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. Each of the video encoder 28 and the video decoder 48 may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined video encoder / decoder (CODEC). Similarly, each of the audio encoder 26 and the audio decoder 46 may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined CODEC. A device including the video encoder 28, the video decoder 48, the audio encoder 26, the audio decoder 46, the encapsulation unit 30, the retrieval unit 52, and / or the de-encapsulation unit 50 may include an integrated circuit, a microprocessor, and / or a wireless communication device (such as a cellular phone).
[0073] The client device 40, the server device 60, and / or the content preparation device 20 may be configured to operate according to the techniques of the present disclosure. For purposes of example, the present disclosure describes these techniques with respect to the client device 40 and the server device 60. However, it should be understood that the content preparation device 20 may be configured to perform these techniques in place of (or in addition to) the server device 60.
[0074] The encapsulation unit 30 may form NAL units, where a NAL unit includes a header identifying the program to which the NAL unit belongs and a payload (e.g., audio data, video data, or data describing the transport or program stream corresponding to the NAL unit). For example, in H.264 / AVC, a NAL unit includes a 1-byte header and a variable-size payload. A NAL unit including video data in its payload may include video data at various granularity levels. For example, a NAL unit may include blocks of video data, multiple blocks, slices of video data, or an entire picture of video data. The encapsulation unit 30 may receive encoded video data from the video encoder 28 in the form of PES packets of an elementary stream. The encapsulation unit 30 may associate each elementary stream with a corresponding program.
[0075] The encapsulation unit 30 may also assemble access units from multiple NAL units. Generally, an access unit may include one or more NAL units, which are used to represent a frame of video data and the corresponding audio data (when such audio data is available). An access unit generally includes all NAL units for one output time instance, e.g., all audio and video data for one time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, a particular frame of all views for the same access unit (same time instance) may be rendered simultaneously. In one example, an access unit may include a decoded picture at one time instance, which may be presented as a basic decoded picture.
[0076] Accordingly, an access unit may include all audio frames and video frames of a common time instance, e.g., all views corresponding to time X. The present disclosure also refers to an encoded picture of a particular view as a "view component". That is, a view component may include an encoded picture (or frame) for a particular view at a particular time. Accordingly, an access unit may be defined as including all view components of a common time instance. The decoding order of an access unit does not necessarily need to be the same as the output order or the display order.
[0077] A media presentation may include a Media Presentation Description (MPD), which may contain descriptions of different alternative representations (e.g., video services with different qualities), and the description may include, for example, codec information, profile values, and level values. The MPD is an example of a manifest file, such as manifest file 66. The client device 40 may retrieve the MPD of the media presentation to determine how to access the movie segments of each presentation. The movie segments may be located in the movie fragment box (moof box) of the video file.
[0078] The manifest file 66 (which may include, for example, the MPD) may announce the availability of the segments of the representation 68. That is, the MPD may include information indicating the wall clock time at which the first segment of one of the representations in the representation 68 becomes available, and information indicating the duration of the segments within the representation 68. In this way, the retrieval unit 52 of the client device 40 may determine when each segment is available based on the start time and duration of the segments preceding a particular segment.
[0079] After the encapsulation unit 30 has assembled the NAL units and / or access units into a video file based on the received data, the encapsulation unit 30 passes the video file to the output interface 32 for output. In some examples, instead of sending the video file directly to the client device 40, the encapsulation unit 30 may locally store the video file or send the video file to a remote server via the output interface 32. The output interface 32 may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium (such as, for example, an optical drive, a magnetic medium drive (e.g., a floppy disk drive)), a Universal Serial Bus (USB) port, a network interface, or other output interfaces. The output interface 32 outputs the video file to a computer-readable medium, such as, for example, a transmission signal, a magnetic medium, an optical medium, a memory, a flash drive, or other computer-readable media.
[0080] The network interface 54 may receive NAL units or access units via the network 74 and provide the NAL units or access units to the de-encapsulation unit 50 via the retrieval unit 52. The de-encapsulation unit 50 may de-encapsulate the elements of the video file into constituent PES streams, de-packetize the PES streams to retrieve the encoded data, and send the encoded data to the audio decoder 46 or the video decoder 48 (depending on whether the encoded data is part of an audio stream or a video stream, e.g., as indicated by the PES packet header of the stream). The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, while the video decoder 48 decodes the encoded video data and sends the decoded video data (which may include multiple views of the stream) to the video output 44.
[0081] According to the technology of the present disclosure, the content preparation device 20 and / or the server device 60 may add additional random access points in DASH / CMAF segments. Random access includes clean random access and open or progressive decoder refresh, up to and including only providing resynchronization for file format parsing. This can be addressed by providing: chunk boundaries that provide information enabling resynchronization and decryption to start at that point; and signaling regarding the type of subsequent random access points. The availability of tfdt, along with moof header information and potentially the use of initialization segments, allows for temporal resynchronization at the presentation time level. The present disclosure refers to this new point as a "resynchronization point". That is, a resynchronization point represents a point at which the file-level container (e.g., boxes in ISO BMFF) can be correctly parsed and a random access point (e.g., an I-frame) in the media data will appear following it. Thus, the client device 40 can randomly access the multimedia content 64, for example, at one of these random access points.
[0082] The content preparation device 20 and / or the server device 60 may also add appropriate signaling in the manifest file 66 (e.g., MPD) that indicates the availability of random access points and resynchronization in each DASH segment and provides information regarding the location, type, and timing of the random access points. The content preparation device 20 and / or the server device 60 may provide signaling in the manifest file 66 (MPD) indicating the availability of additional resynchronization points in the segments, potentially adding characteristics regarding the location, timing, and type of random access, as well as whether the information is accurate or estimated. Thus, the client device 40 can use this signaled data to determine whether such random access points are available and resynchronize retrieval and playback accordingly.
[0083] The client device 40 may be configured to have the ability to resynchronize to unpack, decrypt, and decode by finding a resynchronization point in any starting situation. The content preparation device 20 and / or the server device 60 may provide appropriate chunks that meet the above requirements as proposed in CMAF TUC. Different types may be defined later.
[0084] The client device 40 may be configured to start processing in a restricted receiver environment such as is available, for example, in HTML-5 / MSE-based playback. This issue can be addressed by receiver implementation. However, it would be appropriate to provide a resynchronization trigger and information to a receiver pipeline that has the ability to obtain the mapping, timing, and type of resynchronization points in the data structure, and providing the resynchronization trigger and information to the receiver pipeline allows the user of the decoding pipeline to initialize playback at a random access point.
[0085] In addition, backward-compatible signaling can be provided in the manifest file 66. Client devices not configured to have the ability to parse the signaling can ignore the signaling and perform the methods discussed above. In addition, the manifest file 66 can include signaling that associates a location with the value of @bandwidth to enable signaling at the adaptation set level.
[0086] Figure 2 is a more detailed illustration Figure 1 of an example set of components of the retrieval unit 52. In this example, the retrieval unit 52 includes an eMBMS middleware unit 100, a DASH client 110, and a media application 112.
[0087] In this example, the eMBMS middleware unit 100 further includes an eMBMS receiving unit 106, a cache 104, and a proxy server unit 102. In this example, the eMBMS receiving unit 106 is configured to receive data via eMBMS, for example, according to File Delivery over Unidirectional Transport (FLUTE), which is described in "FLUTE—File Delivery over Unidirectional Transport" by T. Paila et al., Network Working Group, November 2012, RFC 6726, and is available at tools.ietf.org / html / rfc6726. That is, the eMBMS receiving unit 106 can receive files via broadcast from, for example, a server device 60, which can act as a Broadcast / Multicast Service Center (BM-SC).
[0088] As the eMBMS middleware unit 100 receives data for a file, the eMBMS middleware unit can store the received data in the cache 104. The cache 104 can include a computer-readable storage medium, such as flash memory, a hard disk, RAM, or any other suitable storage medium.
[0089] The proxy server unit 102 can act as a server for the DASH client 110. For example, the proxy server unit 102 can provide the MPD file or other manifest files to the DASH client 110. The proxy server unit 102 can announce in the MPD file the availability time for the segments and the hyperlinks from which the segments can be retrieved. These hyperlinks can include a local host address prefix corresponding to the client device 40 (e.g., for IPv4, 127.0.0.1). In this way, the DASH client 110 can request segments from the proxy server unit 102 using HTTP GET or partial GET requests. For example, for a segment available from the link http: / / 127.0.0.1 / rep1 / seg3, the DASH client 110 can construct an HTTP GET request including a request for http: / / 127.0.0.1 / rep1 / seg3 and submit this request to the proxy server unit 102. The proxy server unit 102 can retrieve the requested data from the cache 104 in response to such a request and provide this data to the DASH client 110.
[0090] Figure 3 is a conceptual diagram showing elements of an example multimedia content 120. The multimedia content 120 can correspond to the multimedia content 64 ( Figure 1 ) or another multimedia content stored in the storage medium 62. In Figure 3 's example, the multimedia content 120 includes a media presentation description (MPD) 122 and a plurality of representations 124A - 124N (representations 124). The representation 124A includes optional header data 126 and segments 128A - 128N (segments 128), while the representation 124N includes optional header data 130 and segments 132A - 132N (segments 132). For convenience, the letter N is used to designate the last movie segment in each of the representations 124 in the representations 124. In some examples, there can be a different number of movie segments between the representations 124.
[0091] The MPD 122 can include a data structure separate from the representations 124. The MPD 122 can correspond to Figure 1 's manifest file 66. Similarly, the representations 124 can correspond to Figure 1 's representations 68. Generally, the MPD 122 can include data that generally describes the characteristics of the representations 124 (such as decoding and rendering characteristics, adaptation sets, the profile to which the MPD 122 corresponds, text type information, camera angle information, rating information, stunt mode information (e.g., information indicating a representation including a time subsequence), and / or information for retrieving remote periods (e.g., for inserting targeted advertisements into the media content during playback)).
[0092] The header data 126 (when present) may describe characteristics of the segment 128, such as, the time position of a random access point (RAP, also referred to as a stream access point (SAP)), which segment 128 within the segment 128 includes the random access point, the byte offset of the random access point within the segment 128, the uniform resource locator (URL) of the segment 128, or other aspects of the segment 128. The header data 130 (when present) may describe similar characteristics of the segment 132. Additionally or alternatively, such characteristics may be fully included within the MPD 122.
[0093] The segments 128, 132 include one or more decoded video samples, each of which may include a frame or a slice of video data. Each of the decoded video samples in the segment 128 of video samples may have similar characteristics, such as, height, width, and bandwidth requirements. Such characteristics may be described by data of the MPD 122, although such data is not shown in the Figure 3 example. The MPD 122 may include characteristics as described by the 3GPP specifications, with any or all of the information signaled in the present disclosure added.
[0094] Each of the segments 128, 132 may be associated with a unique uniform resource locator (URL). Thus, each of the segments 128, 132 may be independently retrievable using a streaming network protocol such as DASH. In this manner, a destination device such as the client device 40 may use an HTTP GET request to retrieve the segment 128 or the segment 132. In some examples, the client device 40 may use an HTTP partial GET request to retrieve a specific byte range of the segment 128 or the segment 132.
[0095] According to the techniques of the present disclosure, the MPD 122 (which may again correspond to the Figure 1 manifest file 66) may include signaling for solving the problems discussed above. For example, for each of the representations 124 in the representation 124 (and possibly defaulted at the adaptation set level), the MPD 122 may include one or more resynchronization elements (which may enable backward compatibility). Each resynchronization element may indicate that for each of the segments 128, 132 in the corresponding representation 124 in the representation 124, the following holds:
[0096] · A resynchronization point of stream access point (SAP) type @type or less (but greater than 0) exists in each segment, where the maximum deltaT is signaled by @dT, and the maximum byte offset difference is signaled by @dImax, and the minimum byte offset difference is signaled by @dImin, where two of these values need to be multiplied by the value of the @bandwidth attribute assigned to the representation to obtain the value. "deltaT" refers to the difference in the earliest presentation time of any data following the resynchronization point, in units of the @timescale of the representation. If @type is set to 0, only resynchronization at the container and decryption levels is guaranteed.
[0097] · The resynchronization marker flag @marker can be set using a resynchronization pattern as defined by the segment format in use to indicate the inclusion of a resynchronization point at each resynchronization point.
[0098] · For different SAP types, there can be multiple resynchronization elements.
[0099] · Resynchronization points require that the processing of the media can occur on ISO BMFF and decryption information combined with CMAF headers / initialization segments.
[0100] The following describes Figure 8 an example use of resynchronization points.
[0101] For any segment in segments 128, 132 (for which resynchronization elements are provided in the MPD 122 and the segment is based on ISO BMFF or CMAF), the following can hold:
[0102] · The segment can be a sequence of one or more chunks defined as follows.
[0103] · Additionally, for any two consecutive resynchronization points in a segment of @type specified in the element, the following can hold:
[0104] ο The difference in the earliest presentation times of the two can be at most the value of @dT.
[0105] ο The difference in the byte offset from the start can be at most @dImax normalized with the @bandwidth value.
[0106] ο The difference in the byte offset from the start can be at least @dImin normalized with the @bandwidth value.
[0107] · If the resynchronization marker flag is set, each resynchronization point can include a resynchronization box / styp.
[0108] Figure 4 is a block diagram showing elements of an example video file 150, which may correspond to segments of a representation, such as Figure 3 one of segment 128 and segment 132. Each of segment 128 and segment 132 may include data that substantially conforms to the arrangement of data shown in the example of Figure 4 . The video file 150 may be considered an encapsulated segment. As described above, video files according to the ISO base media file format and its extensions store data in a series of objects called "boxes". In Figure 4 the example, the video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a segment index (sidx) box 162, a movie fragment (MOOF) box 164, and a movie fragment random access (MFRA) box 166. Although Figure 4 shows an example of a video file, it should be understood that other media files may include other types of media data (e.g., audio data, timed text data, etc.) that are similarly structured to the data of the video file 150 according to the ISO base media file format and its extensions.
[0109] The file type (FTYP) box 152 generally describes the file type for the video file 150. The file type box 152 may include data identifying a specification for the best use of the video file 150. The file type box 152 may alternatively be placed before the MOOV box 154, the movie fragment box 164, and / or the MFRA box 166.
[0110] In some examples, segments such as the video file 150 may include an MPD update box (not shown) before the FTYP box 152. The MPD update box may include information indicating that the MPD corresponding to the representation including the video file 150 is to be updated and information for updating the MPD. For example, the MPD update box may provide a URI or URL for a resource to be used to update the MPD. As another example, the MPD update box may include data for updating the MPD. In some examples, the MPD update box may follow immediately after a segment type (STYP) box (not shown) of the video file 150, where the STYP box may define the segment type for the video file 150.
[0111] In Figure 4In the example, the MOOV box 154 includes a Movie Header (MVHD) box 156, a Track (TRAK) box 158, and one or more Movie Extensions (MVEX) boxes 160. Generally, the MVHD box 156 can describe the general characteristics of the video file 150. For example, the MVHD box 156 can include data describing when the video file 150 was initially created, when the video file 150 was most recently modified, the time scale used for the video file 150, the duration for playback of the video file 150, or other data that generally describes the video file 150.
[0112] The TRAK box 158 can include data for the tracks of the video file 150. The TRAK box 158 can include a Track Header (TKHD) box that describes the characteristics of the track corresponding to the TRAK box 158. In some examples, the TRAK box 158 can include decoded video pictures, while in other examples, the decoded video pictures of the track can be included in a movie fragment 164, and the movie fragment 164 can be referenced by the data of the TRAK box 158 and / or the sidx box 162.
[0113] In some examples, the video file 150 can include more than one track. Thus, the MOOV box 154 can include a number of TRAK boxes, and the number of TRAK boxes is equal to the number of tracks in the video file 150. The TRAK box 158 can describe the characteristics of the corresponding track of the video file 150. For example, the TRAK box 158 can describe the time information and / or spatial information for the corresponding track. When the encapsulation unit 30 ( Figure 3 ) includes a parameter set track in a video file such as the video file 150, a TRAK box similar to the TRAK box 158 of the MOOV box 154 can describe the characteristics of the parameter set track. The encapsulation unit 30 can signal the presence of a sequence level SEI message in the parameter set track within the TRAK box that describes the parameter set track.
[0114] The MVEX box 160 can describe the characteristics of the corresponding movie fragment 164, for example, to signal that in addition to the video data (if any) included within the MOOV box 154, the video file 150 includes the movie fragment 164. In the context of streaming video data, the decoded video pictures can be included in the movie fragment 164 rather than in the MOOV box 154. Thus, all decoded video samples can be included in the movie fragment 164 rather than in the MOOV box 154.
[0115] The MOOV box 154 may include a number of MVEX boxes 160, the number of MVEX boxes 160 being equal to the number of movie segments 164 in the video file 150. Each of the MVEX boxes 160 may describe the characteristics of the corresponding movie segment in the movie segments 164. For example, each MVEX box may include a Movie Extension Header Box (MEHD) box that describes the temporal duration for the corresponding movie segment in the movie segments 164.
[0116] As described above, the encapsulation unit 30 may store the sequence data set in a video sample that does not include actual decoded video data. The video sample may generally correspond to an access unit, which is a representation of a decoded picture at a particular time instance. In the context of AVC, the decoded picture includes one or more VCL NAL units and other associated non-VCL NAL units (such as SEI messages), and the one or more VCL NAL units include information for all the pixels that make up the access unit. Accordingly, the encapsulation unit 30 may include the sequence data set in one of the movie segments 164, and the sequence data set may include a sequence-level SEI message. The encapsulation unit 30 may also signal the presence of the sequence data set and / or the sequence-level SEI message within the MVEX box 160 corresponding to one of the movie segments 164 in the movie segments 164 as being present in that movie segment 164 in the movie segments 164.
[0117] The SIDX box 162 is an optional element of the video file 150. That is, a video file conforming to the 3GPP file format or other such file formats does not necessarily include the SIDX box 162. According to an example of the 3GPP file format, the SIDX box may be used to identify sub-segments of a segment (e.g., a segment contained within the video file 150). The 3GPP file format defines a sub-segment as "a self-contained set of one or more consecutive movie segment boxes with corresponding media data boxes, and the media data boxes containing the data referenced by the movie segment boxes must follow the movie segment boxes and precede the next movie segment box containing information about the same track." The 3GPP file format also indicates that the SIDX box "contains a sequence of references to the sub-segments of the (sub) segment documented by the box. The referenced sub-segments are consecutive in presentation time. Similarly, the bytes referenced by the segment index box are always consecutive within the segment. The referenced size gives a count of the number of bytes in the referenced material."
[0118] The SIDX box 162 typically provides information representing one or more sub - segments of segments included in the video file 150. For example, such information may include the playback time at which the sub - segment starts and / or ends, the byte offset for the sub - segment, whether the sub - segment includes a stream access point (SAP) (e.g., starting from an SAP), the type used for the SAP (e.g., the SAP is an Instant Decoder Refresh (IDR) picture, a Clean Random Access (CRA) picture, a Broken Link Access (BLA) picture, etc.), the position of the SAP within the sub - segment (based on playback time and / or byte offset), etc.
[0119] The movie fragment 164 may include one or more decoded video pictures. In some examples, the movie fragment 164 may include one or more Groups of Pictures (GOPs), each of which may include several decoded video pictures, e.g., frames or pictures. Additionally, as described above, in some examples, the movie fragment 164 may include sequence data sets. Each movie fragment within the movie fragment 164 may include a Movie Fragment Header box (MFHD, not shown in Figure 4 ). The MFHD box may describe characteristics of the corresponding movie fragment, such as the sequence number for the movie fragment. The movie fragments 164 may be included in the video file 150 in the order of the sequence numbers.
[0120] The MFRA box 166 may describe random access points within the movie fragments 164 of the video file 150. This can assist in performing trick modes, such as searching for a specific time position (i.e., playback time) within segments encapsulated by the video file 150. The MFRA box 166 is typically optional and in some examples does not need to be included in the video file. Similarly, a client device (such as client device 40) does not necessarily need to reference the MFRA box 166 to correctly decode and display the video data of the video file 150. The MFRA box 166 may include a number of Track Fragment Random Access (TFRA) boxes (not shown), the number of TFRA boxes being equal to the number of tracks of the video file 150, or in some examples, equal to the number of media tracks (e.g., non - hint tracks) of the video file 150.
[0121] In some examples, the movie fragment 164 may include one or more stream access points (SAPs), such as IDR pictures. Similarly, the MFRA box 166 may provide an indication of the location of the SAP within the video file 150. Accordingly, a temporal subsequence of the video file 150 may be formed from the SAPs of the video file 150. The temporal subsequence may also include other pictures, such as P-frames and / or B-frames that depend on the SAPs. The frames and / or slices of the temporal subsequence may be arranged within segments such that frames / slices of the temporal subsequence that depend on other frames / slices of the subsequence can be decoded correctly. For example, in a hierarchical arrangement of data, data used for prediction of other data may also be included in the temporal subsequence.
[0122] In one example, the present disclosure defines "chunking" as follows in the table below. The table provides example definitions for both the cardinality and ordinality of chunking:
[0123]
[0124] In one example, the present disclosure defines a resynchronization point as the start of a chunk. Additionally, the resynchronization point may be assigned the following characteristics:
[0125] · It has a byte offset from the start of the segment that points to the first byte of the chunk.
[0126] · It has an earliest presentation time that is derived from the information in the moof and possibly a movie header assigned to it.
[0127] · It has an SAP type assigned to it, as defined in ISO / IEC 14496-12.
[0128] · There is an indication of whether the chunk contains a resynchronization box (styp as above).
[0129] · Starting from the resynchronization point, together with the information in the movie header, file format parsing and decryption can be performed.
[0130] In some examples, a resynchronization marker box may be defined that synchronizes with the start of a segment (e.g., video file 150) by scanning the byte stream for a marker. The resynchronization box may have the following characteristics:
[0131] · It defines a unique pattern with a very high probability for resynchronization.
[0132] · It defines an SAP type.
[0133] The resynchronization marker box can be a new box, or it can reuse an existing box, such as a styp box. In the present disclosure, it is assumed that a styp with specific restrictions can be used as the resynchronization marker box. Research is currently being conducted on the robustness of this method.
[0134] The following Figures 5 - 7 is used to describe several use cases for which random access to DASH segments at points other than the start of a segment can be useful.
[0135] Figure 5 is a conceptual diagram showing an example low-latency architecture 200 that can be used in a first use case according to the present disclosure. That is, Figure 5 shows the basic information flow for operating a low-latency DASH service according to the DASH-IF IOP. The low-latency architecture 200 includes a DASH packager 202, an encoder 216, a content delivery network (CDN) 220, a conventional DASH client 230, and a low-latency DASH client 232. The encoder 216 can generally correspond to Figure 1 either or both of the audio encoder 26 and the video encoder 28, and the DASH packager 202 can correspond to Figure 1 the encapsulation unit 30.
[0136] In this example, the encoder 216 encodes the received media data to form a CMAF header (CH) (such as CH 208), CMAF initial chunks 206A, 206B (CIC 206), and CMAF non-initial chunks 204A - 204D (CNC 204). The encoder 216 provides the CH 208, CIC 206, and CNC 204 to the DASH packager 202. The DASH packager 202 also receives a service description that includes information about the general description of the service and the encoder configuration of the encoder 216.
[0137] The DASH packager 202 uses the service description, CH 208, CIC 206, and CNC 204 to form a Media Presentation Description (MPD) 210 and an initialization segment 212. The DASH packager 202 also generates mappings of CH 208, CIC 206, and CNC 204 into segments 214A, 214B (segments 214), and provides the segments 214 to the CDN 220 incrementally. As segments 214 are generated, the DASH packager 202 can deliver segments 214 in chunks. The CDN 220 includes a segment store 222 for storing the MPD 210, IS 212, and segments 214. For example, in response to HTTP Get or partial Get requests from a regular DASH client 230 and a low-latency DASH client 232, the CDN 220 delivers complete segments to the regular DASH client 230, but delivers individual chunks (e.g., CH 208, CIC 206, and CNC 204) to the low-latency DASH client 232.
[0138] Figure 6 is a conceptual diagram that further shows in detail an example of the use case described with respect to Figure 5 the example described. Figure 6 The example of shows segments 250A - 250E (segments 250), which include corresponding sets of chunks 252A - 252E (chunks 252). A client device (such as Figure 1 client device 40) can retrieve a complete segment 250 or an individual chunk 252. For example, as Figure 5 shown, a regular DASH client 230 can retrieve segment 250, while a low-latency DASH client 232 can retrieve an individual chunk 252 (at least initially).
[0139] Figure 6 further depicts how latency can be reduced by retrieving individual chunks 252 instead of complete segments 250. For example, retrieving a complete segment at the current time may result in a high latency. Simply retrieving the latest fully available segment can reduce latency, but still results in a relatively high latency.
[0140] Alternatively, by retrieving chunks, these latencies can be greatly reduced. For example, at the current time indicated by "now" in Figure 6 segment 250E is not fully formed. However, even before segment 250E is fully formed, the client device can retrieve the formed chunks, such as chunks 252E - 1 and 252E - 2 of segment 250E, assuming that chunks 252E - 1 and 252E - 2 are formed and available for retrieval.
[0141] When adding a live stream, both low latency and fast startup should generally be achieved. However, this is no small feat, and the following is based on Figure 6 discusses some strategies:
[0142] · In the first case, at the live edge, segments that are 3 segments behind in the time history (i.e., segments 250B, 250C, and 250D) are loaded into the buffer. Once a segment becomes available, playback begins. This results in significant latency, but playback can start relatively quickly because of the random access at the start of the segments that are loaded.
[0143] · In the second case, the most recently available segment (segment 250D) is selected instead of the three-segment-old segment. In this case, the playback latency is at least the segment duration, but can be higher. The startup may be similar to the above case.
[0144] · In the other three cases, segments that include multiple chunks (e.g., segment 250E) are played back while still being generated. This reduces latency, but there is a problem that the start of playback may be affected, especially if the difference between the segment availability start time of the most recently published segment and the wall clock time is greater than the target latency. In this case, the client device may have to wait until the next second segment is published. In the case of 6-second segments, this can result in a startup latency of 4 - 5 seconds.
[0145] · There are other techniques and use cases. For example, the client can access old segments at the start, download all the content, and accelerate playback and perform fast-forward decoding. However, such an approach has the following disadvantages: a large amount of data needs to be downloaded before accelerated decoding can occur. In addition, it is not widely supported in the decoder interface.
[0146] A suitable solution could be:
[0147] · At least one representation of the adaptation set includes more frequent random access points and non-initial chunks in the segments / fragments.
[0148] · The DASH client can use the information from the MPD to determine that such a random access method exists, but may inaccurately signal the location / byte offset of the random access points.
[0149] · The DASH client can access such a representation at startup, but can only start downloading from the byte range of the most recently available non-initial chunk, or at least close to that range.
[0150] · Once downloaded, the DASH client can determine the random access points and start processing the data and the initialization segments / CMAF headers of the same representation that have been downloaded. The location of the random access points is discussed below.
[0151] However, the latter approach may encounter various problems, as outlined below.
[0152] Thus, as shown in the example of Figure 6 and discussed below, using chunks as discussed in the present disclosure can significantly reduce latency. Signaling the beginning of a chunk in advance can allow for the early generation of a manifest file that does not need to be updated frequently, but can still indicate the approximate locations of stream access points (SAPs) within the chunk. In this way, a client device can use the manifest file to determine the locations of chunk boundaries without continuous manifest file updates, and at the same time still allow the client device to initiate media streaming at the beginning of a chunk boundary (e.g., a resynchronization point). That is, the client device can determine the byte range of a segment including a resynchronization point from the manifest file even before the segment has been fully formed, because the manifest file can signal the byte range or other data representing the approximate location of the resynchronization point within the segment.
[0153] Figure 7 is a conceptual diagram showing an example second use case of using DASH and CMAF random access in the context of a broadcast protocol. Figure 7 An example is shown that includes a media encoder 280, a CMAF / File Format (FF) packer 282, a DASH packer 284, a ROUTE transmitter 286, a CDN source server 288, a ROUTE receiver 290, a DASH client 292, a CMAF / FF parser 294, and a media decoder 296. The media encoder 280 encodes media data (such as audio data or video data). The media encoder 280 may correspond to Figure 1 the audio encoder 26 or video encoder 28 of Figure 5 or the encoder 216 of
[0154] The CMAF / FF packager 282 provides these files (e.g., chunks) to the DASH packager 284, which aggregates the files / chunks into DASH segments. The DASH packager 284 can also form a manifest file (such as an MPD) that includes data describing the files / chunks / segments. Additionally, according to the techniques of the present disclosure, the DASH packager 284 can determine an approximate location of a future stream access point (SAP) or random access point (RAP) and signal that approximate location in the MPD. The CMAF / FF packager 282 and the DASH packager 284 can correspond to Figure 1 the encapsulation unit 30 of Figure 5 or the DASH packager 202 of
[0155] The DASH packager 284 provides the segments and the MPD to the ROUTE sender 286 and the CDN source server 288. The ROUTE sender 286 and the CDN source server 288 can correspond to Figure 1 the server device 60 of Figure 5 or the CDN 220 of. Generally, in this example, the ROUTE sender 286 can send media data to the ROUTE receiver 290 according to ROUTE. In other examples, other file-based delivery protocols can be used for broadcasting or multicasting, such as FLUTE. Additionally or alternatively, the CDN source server 288 can send media data to the ROUTE receiver 290 according to HTTP, for example, and / or directly to the DASH client 292.
[0156] The ROUTE receiver 290 can be implemented in middleware (such as Figure 2 the eMBMS middleware unit 100 of). The ROUTE receiver 290 can buffer the received media data in a cache 104 as shown in Figure 2 . The DASH client 292 (which can correspond to Figure 2 the DASH client 110 of) can retrieve the cached media data from the ROUTE receiver 290 using HTTP. Alternatively, the DASH client 292 can retrieve media data directly from the CDN source server 288 according to HTTP, as discussed above.
[0157] In addition, according to the technology of the present disclosure, the DASH client 292 can use a manifest file (such as an MPD) to determine, for example, the position of an SAP or RAP following a resynchronization point signaled in the manifest file. The DASH client 292 can initiate the retrieval of a media presentation starting from the next fastest resynchronization point. A resynchronization point can generally indicate a position in the bitstream where the file container level data can be correctly parsed. Thus, the DASH client 292 can initiate streaming at the resynchronization point and deliver the media data received starting from the resynchronization point to the CMAF / FF parser 294.
[0158] The CMAF / FF parser 294 can start parsing the media data starting from the resynchronization point. The CMAF / FF parser 294 can correspond to Figure 1 the demultiplexing unit 50. In addition, the CMAF / FF parser 294 can extract decodable media data from the parsed data and deliver the decodable media data to the media decoder 296, and the media decoder 296 can correspond to Figure 1 the audio decoder 46 or the video decoder 48. The media decoder 296 can decode the media data and deliver the decoded media data to a corresponding output device, such as Figure 1 the audio output 42 or the video output 44.
[0159] In the case of broadcasting, an example for the combination of DASH / CMAF and ROUTE is shown in Figure 7 In the combination of the low-latency DASH mode and ROUTE (as considered, for example, for the DVB TM-IPI task group over ABR multicast and for the ATSC profile), the following problem may occur. If the ROUTE receiver 290 joins in the middle of a DASH / CMAF low-latency segment, it cannot start processing the data because there is no synchronization available and there is no random access for any other purpose. Thus, even if more frequent random access is provided in the middle of the segment, the startup will be delayed.
[0160] A suitable solution can be:
[0161] · The broadcast / multicast representation includes more frequent random access points and non-initial chunks in the segments / fragments.
[0162] · The DASH client 292 uses the information of the MPD and / or possibly the information from the ROUTE receiver 290 to determine that such a random access method exists. The DASH client 292 can use such information to precisely locate the random access point, but this may not always be the case.
[0163] · The DASH client 292 can access the representation at startup, but may not be able to access all the information from the beginning.
[0164] · Once the DASH client 292 starts accessing the received segmented parts, it can find the random access points and start processing the data and the initialization non-segmented / CMAF headers of the same download of the representation. The location of the random access points is discussed below.
[0165] However, the latter method may encounter various problems as outlined below.
[0166] In a similar situation as discussed in the second use case above, packet loss can also be problematic, not only during random access resynchronization. In the third use case of this example, a similar process as discussed above can apply. In addition, it can even be the case that not only a clean random access is attempted, but also after sufficient box parsing is possible, non-random access chunks (e.g., without IDR frames), decoding, and presentation events can be attempted. Therefore, not only is it important to resynchronize to a clean random access, but also random access for file format parsing is important.
[0167] Another fourth use case may occur if real-time media content is typically distributed with low latency, but then the same media content is used for delayed playback in time-shifting. The client may wish to access the media presentation at a specific time, but that time may not (and typically will not) coincide with the start of a segment / CMAF fragment.
[0168] A suitable solution can be:
[0169] · At least one representation of the adaptation set can include more frequent random access points and non-initial chunks in the segments / fragments.
[0170] · The DASH client 292 can use the information from the MPD to determine the existence of such a random access method, but the location / byte offset of the random access points may not be precisely known.
[0171] · The DASH client 292 can access the representation during seeking, but can only allow the download to start from the byte range of the most recently available non-initial chunk or at least close to that byte range.
[0172] · Once downloaded, the DASH client 292 can find the random access points and start processing the data and the initialization segment / CMAF headers of the same download of the representation. The location of the random access points is discussed below.
[0173] However, the latter method may encounter various problems as outlined below.
[0174] Resynchronization in the case of ISO BMFF / DASH / CMAF segmentation generally involves multiple processes outlined below:
[0175] 1) Locate the box structure.
[0176] 2) Locate the CMAF chunk / fragment with all relevant information.
[0177] 3) Find the timing via mdat and tfdt.
[0178] 4) Obtain all decryption-related information (if applicable).
[0179] 5) Possibly process event messages.
[0180] 6) Start decoding at the elementary stream level.
[0181] The following outlines an example way to find resynchronization points in the box structure at a specific time:
[0182] · If a segment index (SIDX box) exists, such a resynchronization point is provided as the presentation time and byte offset. However, since the segments are not fully formed in advance, the segment index is typically not available for low-latency live.
[0183] · If the start of the segment is available, the client can download the minimum set of byte ranges such that the box structure can be processed.
[0184] · Resynchronization is provided by the underlying protocol, which for example specifies signaling the boundaries of chunks and the client can start parsing.
[0185] · If the start of the segment cannot be easily determined by signaled data, the client can find a synchronization pattern that allows the client to randomly access the data. Then, the client can start parsing and find an appropriate box structure that allows processing, such as for example emsg, prft, mdat, moof, and / or mdat.
[0186] This disclosure describes techniques that can be applied to the fourth example above. The first three represent example simplifications when the corresponding information is available.
[0187] Based on the above discussion, this disclosure recognizes the following problems, and these problems require solutions:
[0188] 1) Add additional random access points in DASH / CMAF segmentation. Random access can include clean random access as well as open or progressive decoder refreshes, up to just providing resynchronization regarding file format parsing.
[0189] 2) Add appropriate signaling in the MPD (or other manifest file) that indicates the availability of random access points and resynchronization in each DASH segment and provides information about the location, type, and timing of the random access points. The information can be exact or may be within a range.
[0190] 3) The ability to resynchronize at any starting point by finding a resynchronization point for the ability to unpack, decrypt, and decode.
[0191] 4) The ability to start processing in a restricted receiver environment such as is available in, for example, HTML-5 / MSE-based playback.
[0192] Figure 8 is a conceptual diagram showing an example signaling of stream access points (SAPs) in a manifest file. Specifically, Figure 8 shows a bitstream 300 including SAPs 302A - 302D (SAP 302) and segments 304A - 304D (segments 304), and a bitstream 310 including SAPs 312A - 312D (SAP 312), SAPs 316A - 316D (SAP 316), and segments 314A - 314D (segments 314). That is, in this example, segments 314 of bitstream 310 include more frequent SAPs 312, SAPs 316 compared to segments 304 of bitstream 300. Each of SAP 302, SAP 312 can correspond to the start of the corresponding segment in segments 304, segments 314 and the first set of chunks of these segments. SAP 316 can correspond to the start of a chunk within the corresponding segment 316 (but not at the start of the corresponding segment 316).
[0193] To provide a simple technique for implementing a constant bitrate representation (with chunks having an equal distance of 1000 samples (and @timescale = 1000, in sample duration units) and SAP type 1 (which can be, for example, an audio representation)), resynchronization elements can be added, where:
[0194] · @type = 1
[0195] · @dT = 1000
[0196] · @dImin = 100
[0197] · @dImax = 100
[0198] A client that receives such information (e.g., Figure 1The client device 40) may not be able to identify a random access point that can be accessed per second within an exact byte range for a segment with @duration = 10000. If the bitrate is variable, the receiver (e.g., client device 40) can use @dIMin and @dIMax to identify the range in which to find the random access point. As an alternative to signaling the maximum @dT, it can also signal the nominal chunk duration.
[0199] The resynchronization element of the manifest file can also include a URL@index that points to the binary resynchronization index of the resynchronization points in each segment using the same template function as the regular segments. If present, this resynchronization can provide the exact location of all resynchronization points in the segment in a similar manner to the segment index. If this index exists, the resynchronization index can be available for all segments of the period available at the time of publication of the manifest file / MPD.
[0200] In one method, the resynchronization index can be the same as the segment index, but can be changed.
[0201] Figure 1 The client device 40 can use the ISO BMFF 4-character box type as a basis for resynchronization into a media file (e.g., Figure 4 the video file 150, which can be a segment). In one example, the selected box type is the "styp" box, but it can also be the "moof" box itself. Random emulation of box string types is very rare. A test report of styp emulation is described below. This emulation is then avoided by checking against known expected box types. The client device 40 can perform the resynchronization mechanism outlined as follows:
[0202] 1) Find the occurrence of the "styp" byte string in the segment, e.g., at byte offset B1.
[0203] 2) Verify against random emulation as follows: Compare the next box type with a list of expected box types: "styp", "sidx", "ssix", "prft", "moof", "mdat", "free", "mfra", "skip", "meta", "meco".
[0204] a. If one of the known box types is found, the byte offset B1 - 4 bytes is the byte offset of the resynchronization point.
[0205] b. If this is not one of the previously mentioned known box types, the occurrence of the styp box is considered an invalid synchronization point and is ignored. Restart from step 1 above.
[0206] The techniques of the present disclosure were tested on 30,282 segmented tests from the scanned DASH-IF test assets. The scan revealed 28,408 occurrences of the "styp" string in the file, and only 10 of these 28,408 occurrences (about 1 in 2,840 occurrences) were emulations, which were discarded once it was determined that the following boxes were not one of the expected box types: "styp", "sidx", "ssix", "prft", "moof", "mdat", "free", "mfra", "skip", "meta", "meco".
[0207] Based on these results, it is considered that using styp resynchronization point detection and chunk structure is sufficient. Limiting to the set of boxes that can follow styp (such as prft, emsg, free, skip, and moof) would be appropriate.
[0208] The remaining problem is to determine the SAP type and the earliest presentation time. The latter is easily achieved by using tfdt and other information in the movie fragment header. It would be appropriate to document this algorithm.
[0209] There are several options for determining the SAP type as described below:
[0210] · Detect based on the information in moof. Simple techniques can be documented and implemented.
[0211] · Use the compatibility brand for the SAP type. By using CMAF, the following can be obtained:
[0212] οcmff: Indicates that the SAP is 1 or 2
[0213] οcmfl: Indicates that the SAP is 0 (correct for encryption?)
[0214] οcmfr:: Indicates that the SAP is 1, 2, or 3
[0215] · If used consistently, this signaling can be sufficient. Compatibility brands for other SAP types can be defined.
[0216] · Other techniques can be used to indicate the SAP type.
[0217] Existing options can be used to determine the SAP type.
[0218] In this way, the techniques of the present disclosure can be summarized as follows and can be performed by devices such as Figure 1 content preparation device 20, server device 60, and / or client device 40 as discussed above:
[0219] In the DASH context, in some cases, segments are considered as single units for downloading, accessing media presentations, and also via the addressed URLs. However, segments can be structured to enable resynchronization at the container level and even random access to the corresponding representations within the segments. The resynchronization mechanism is supported and signaled by resynchronization elements.
[0220] Resynchronization elements signal resynchronization points within the segments. A resynchronization point is the start of a chunk (in terms of byte position), where a chunk is defined as a structured consecutive byte range of media data within the segment that contains a specific presentation duration and can be accessed (including potential decryption) independently in the container format. The resynchronization points in a segment can be defined as follows:
[0221] · A resynchronization point is the start of a chunk.
[0222] · Additionally, resynchronization points have been assigned the following characteristics:
[0223] ο It has a byte offset or index value from the start of the segment that points to the first byte of the chunk.
[0224] ο It has an assigned earliest presentation time in the representation.
[0225] ο It has an assigned SAP type, e.g., defined by the SAP type in ISO / IEC 14496-12.
[0226] ο It has been assigned a flagging characteristic that indicates whether the resynchronization point can be detected when parsing the segment via a specific flag, or whether the resynchronization point needs to be signaled by an external unit.
[0227] · Processing and initializing the information in the segment starting from the resynchronization point (if any) allows the container to be parsed and decrypted. The ability to access the contained elementary video stream, if and how, is defined by the SAP type.
[0228] Due to causal reasons, it may be difficult to signal every resynchronization point in the MPD because resynchronization points can be added by the segment packager independently of MPD updates. For example, resynchronization points can be generated by the encoder and packager independently of the MPD. Additionally, in low-latency scenarios, MPD signaling may not be available to DASH clients (e.g., Figure 2 DASH client 110 or Figure 7 DASH client 292). Therefore, there are two ways to signal the resynchronization points provided in the segments in the MPD:
[0229] ·By providing a binary map for resynchronization points in the resynchronization index segment for each segment. This is most easily applied to segments that are fully available over the network.
[0230] ·By signaling the presence of a resynchronization point in the segment and also some additional information that allows the resynchronization point to be easily located based on byte position and presentation time.
[0231] To signal the above features, the resynchronization element has different attributes that are explained in more detail in Clause 5.3.12.2 of the DASH specification.
[0232] Random access refers to processing, decoding, and presenting the representation forward starting from a random access point at time t by initializing the representation using the initialization segment (if present) and decoding and presenting the representation forward from the signaled segment. The random access point can be signaled using the RandomAccess element, as defined in Table 10 below.
[0233] Table 10—Random Access Signaling
[0234]
[0235]
[0236] Table 11 provides different types of random access points.
[0237] Table 11
[0238]
[0239] The resynchronization index segment contains information related to the media segments. The resynchronization index segment provides the exact positions of all resynchronization points in the segment in a manner similar to the segment index. Resynchronization points are defined in Clause 5.3.12.1 of the DASH specification.
[0240] Resynchronization points for ISO BMFF can be defined as the start of an ISO BMFF segment, with the following restrictions in both cardinality and ordinality:
[0241]
[0242]
[0243] For resynchronization points based on ISO BMFF, the features can be defined as follows:
[0244] ·The Index is defined as the offset of the first byte of the above - constrained ISO BMFF segment.
[0245] · The earliest presentation time, Time, is defined as the minimum of the decoding time, composition offset, and edit list for any sample in the chunk.
[0246] · The SAP type is defined according to Section 4.5.2 of the DASH specification.
[0247] · If styp is presented with "cmfl" as the primary compatibility brand, then there is a marker.
[0248] A resynchronization index segment can index a media segment of a representation and can be defined as follows:
[0249] · Each representation index segment shall start with an "styp" box, and the brand "risg" shall be present in the "styp" box. The compliance requirements for the brand "risg" are defined by this subclause.
[0250] · Each media segment is indexed by one or more segment index boxes; the boxes for a given media segment are consecutive.
[0251] Figure 9 is a flowchart showing an example method of retrieving media data according to the technology of the present disclosure. It is explained with respect to Figure 1 client device 40 of Figure 9 the method. However, Figure 5 the low-latency DASH client 232 or a client device including a media decoder 296, a CMAF / FF parser 294, a DASH client 292, and Figure 7 a ROUTE receiver 290 of
[0252] Initially, client device 40 may retrieve a manifest file (350) of a media presentation, such as an MPD. Client device 40 may retrieve the manifest file from, for example, server device 60. The manifest file may include data indicating resynchronization points at chunk boundaries within segments of the representations included in the media presentation. Thus, client device 40 may determine the resynchronization points (352) of the media presentation, e.g., the most recently available resynchronization points. Generally, the resynchronization points may indicate the start of chunk boundaries, which are randomly accessible points in the representation at which the file-level container (e.g., the data structures such as boxes discussed above) can be correctly parsed.
[0253] Specifically, the manifest file can indicate the location of resynchronization points, such as byte offsets from the beginning of a segment. This information may not precisely identify the location of the resynchronization point within the segment, but it can ensure that the resynchronization point will be available within a range of bytes from the byte offset. Thus, the client device 40 can form a request (such as an HTTP partial Get request) (354) specifying the indicated byte offset to start retrieving at the resynchronization point. Then, the client device 40 can send this request to the server device 60 (356).
[0254] In response to this request, the client device 40 can receive the requested media data (358), which includes a resynchronization point. As described above, the byte offset may not precisely identify the location of the resynchronization point, and thus, the client device 40 can parse the data until the actual location of the resynchronization point is detected. The client device 40 can start from the resynchronization point and parse the file-level data structure (such as a file format box) to determine the location of the media data for the corresponding chunks of the retrieved media data. Specifically, the client device 40 can identify the resynchronization point as the start of a chunk by detecting, for example, segment type values, maker reference time values, event messages, movie fragments, and media data container boxes. The movie fragment can include encoded media data.
[0255] The de-encapsulation unit 50 can, for example, extract the encoded media data for the corresponding chunk from the movie fragment (360) and provide the encoded media data to, for example, the video decoder 48. The chunk can start from a random access point (RAP) (such as an intra-frame predicted frame (I-frame) of video data). The manifest file can also indicate whether the RAP is the start of a closed group of pictures (GOP) or an open GOP, and thereby indicate the type of random access that can be performed starting from the RAP (e.g., whether the leading picture of the I-frame is decodable or not). The video decoder 48 can then decode the encoded media data (362) and send the decoded media data to, for example, the video output 44 to present the media data (364).
[0256] In this way, Figure 9 the method represents an example of a method for retrieving media data, the method comprising: retrieving a manifest file for media presentation, the manifest file indicating a resynchronization point at which container parsing of media data of a bitstream can start for a segment of a representation of the media presentation, the resynchronization point being located at a position other than the start of the segment and representing a point at which container parsing of media data of the bitstream can start; using the manifest file to form a request for retrieving the media data of the representation starting at the resynchronization point; sending the request to initiate retrieval of the media data of the media presentation starting at the resynchronization point; and presenting the retrieved media data.
[0257] Certain techniques of the present disclosure are outlined in the following examples:
[0258] Example 1: A method for retrieving media data, the method comprising: retrieving a manifest file for media presentation, the manifest file indicating that resynchronization and decryption can start at a resynchronization point of a representation of the media presentation; starting to retrieve the media data of the representation at the resynchronization point; and presenting the retrieved media data.
[0259] Example 2: The method according to Example 1, wherein the resynchronization point includes the start of a chunk boundary.
[0260] Example 3: The method according to Example 2, wherein the chunk boundary includes the start of a chunk, the chunk including a zero or one segment type value, a zero or one producer reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
[0261] Example 4: The method according to any one of Examples 1-3, wherein the manifest file indicates the availability of the resynchronization point in a segment of the representation.
[0262] Example 5: The method according to Example 4, wherein the resynchronization point is located at a position other than the start of the segment.
[0263] Example 6: The method according to any one of Examples 4 and 5, wherein the manifest file indicates the type of random access that can be performed at the resynchronization point.
[0264] Example 7: The method according to any one of Examples 4-6, wherein the manifest file indicates the location and timing of the resynchronization point and whether the location and timing information is accurate or estimated.
[0265] Example 8: The method according to any one of Examples 1-7, wherein the manifest file includes a Media Presentation Description (MPD).
[0266] Example 9: A device for retrieving media data, the device including one or more units for performing the method according to any one of Examples 1-8.
[0267] Example 10: The device according to Example 9, wherein the one or more units include one or more processors implemented in a circuit and a memory configured to store media data.
[0268] Example 11: The device according to Example 9, wherein the device includes at least one of the following: an integrated circuit; a microprocessor; or a wireless communication device.
[0269] Example 12: A computer-readable storage medium storing instructions thereon, the instructions, when executed, causing a processor to execute the method according to any one of Examples 1-8.
[0270] Example 13: A device for retrieving media data, the device comprising: a unit for retrieving a manifest file for media presentation, the manifest file indicating that resynchronization and decryption can start at a resynchronization point of a representation of the media presentation; a unit for starting to retrieve media data of the representation at the resynchronization point; and a unit for presenting the retrieved media data.
[0271] Example 14: A method of sending media data, the method comprising: sending a manifest file for media presentation to a client device, the manifest file indicating that resynchronization and decryption can start at a resynchronization point of a representation of the media presentation; receiving, from the client device, a request for media data starting at the resynchronization point; and in response to the request, sending to the client device the media data of the representation starting at the resynchronization point.
[0272] Example 15: The method according to Example 14, further comprising: generating the manifest file.
[0273] Example 16: The method according to any one of Examples 14 and 15, wherein the resynchronization point includes the start of a chunk boundary.
[0274] Example 17: The method according to Example 16, wherein the chunk boundary includes the start of a chunk, the chunk including zero or one segment type value, zero or one producer reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
[0275] Example 18: The method according to any one of Examples 14-17, wherein the manifest file indicates the availability of the resynchronization point in a segment of the representation.
[0276] Example 19: The method according to Example 18, wherein the resynchronization point is located at a position other than the start of the segment.
[0277] Example 20: The method according to any one of Examples 18 and 19, wherein the manifest file indicates the type of random access that can be performed at the resynchronization point.
[0278] Example 21: The method according to any one of Examples 18-20, wherein the manifest file indicates the location and timing of the resynchronization point and whether the location and timing information is accurate or estimated.
[0279] Example 22: The method according to any one of Examples 14-21, wherein the manifest file includes a Media Presentation Description (MPD).
[0280] Example 23: A device for transmitting media data, the device including one or more units for performing the method according to any one of Examples 14-22.
[0281] Example 24: The device according to Example 23, wherein the one or more units include one or more processors implemented in a circuit and a memory configured to store media data.
[0282] Example 25: The device according to Example 23, wherein the device includes at least one of the following: an integrated circuit; a microprocessor; or a wireless communication device.
[0283] Example 26: A computer-readable storage medium having instructions stored thereon, which when executed cause a processor to perform the method according to any one of Examples 1-8.
[0284] Example 27: A device for transmitting media data, the device including: a unit for sending to a client device a manifest file for media presentation, the manifest file indicating that resynchronization and decryption can start at a resynchronization point of a representation of the media presentation; a unit for receiving from the client device a request for media data starting at the resynchronization point; and a unit for, in response to the request, sending to the client device the requested media data of the representation starting at the resynchronization point.
[0285] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted through a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium, or a communication medium including, for example, any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. In this way, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0286] By way of example, and not limitation, such a computer-readable storage medium can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that is accessible by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transitory, tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0287] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, as used herein, the term "processor" can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Further, the techniques can be fully implemented in one or more circuits or logic elements.
[0288] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a group of ICs (e.g., a chipset). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but need not be implemented by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperating hardware units, including one or more processors as described above, in conjunction with appropriate software and / or firmware.
[0289] Various examples have been described. These examples and other examples are within the scope of the appended claims.
Claims
1. A method for retrieving media data, the method comprising: Retrieving a manifest file for media presentation, the manifest file indicating positions of resynchronization points in segments of a representation of the media presentation, the manifest file further indicating that container parsing of media data of a bitstream can be started at the resynchronization points, the resynchronization points being located at positions other than the start of the segments and representing points at which container parsing of the media data of the bitstream can be started and representing that a random access point (RAP) follows after the resynchronization points; Using the manifest file to form a request for retrieving the media data of the representation starting at the resynchronization point; Sending the request to initiate retrieval of the media data of the media presentation starting at the resynchronization point; Parsing the retrieved media data starting at the resynchronization point and determining that the RAP of the media data follows after the resynchronization point; And Presenting the retrieved media data including the RAP of the media data and following the RAP of the media data and not presenting the retrieved media data from the resynchronization point to the RAP of the media data.
2. The method according to claim 1, wherein Presenting the retrieved media data includes: parsing a file-level media data container of the retrieved media data at the resynchronization point.
3. The method according to claim 2, wherein Parsing includes: Parsing the file-level media data container until the RAP of the media presentation is detected; And Sending the RAP to a media decoder.
4. The method according to claim 1, wherein The resynchronization point includes the start of a chunk boundary.
5. The method according to claim 4, wherein The chunk boundary includes the start of a chunk, the chunk including zero or one segment type value, zero or one producer reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
6. The method according to claim 1, wherein, The manifest file indicates the availability of the resynchronization point in the segment of the representation.
7. The method according to claim 1, wherein, Retrieving the manifest file includes retrieving the manifest file before the segment is fully formed, wherein sending the request for initiating retrieval of the media data includes: sending the request for initiating retrieval of the media data as a live stream, wherein the segments of the media presentation are not fully formed before at least some segments in the segment are retrieved.
8. The method according to claim 6, wherein, The manifest file indicates the type of random access that can be performed at the resynchronization point.
9. The method according to claim 6, wherein The manifest file indicates the position and timing of the resynchronization point and whether the position and timing information is accurate or estimated.
10. The method according to claim 1, wherein, The manifest file includes a media presentation description (MPD).
11. A device for retrieving media data, the device comprising: A memory configured to store media data of a media presentation; And One or more processors implemented in circuitry and configured to: Retrieve a manifest file for media presentation, the manifest file indicating the positions of resynchronization points in segments of a representation of the media presentation, the manifest file further indicating that container parsing of media data of a bitstream can be started at the resynchronization points, the resynchronization points being located at positions other than the start of the segments and representing points at which container parsing of the media data of the bitstream can be started and representing that a random access point (RAP) follows after the resynchronization points; Use the manifest file to form a request for retrieving the media data of the representation starting at the resynchronization point; Send the request to initiate retrieval of the media data of the media presentation starting at the resynchronization point; Parse the retrieved media data starting at the resynchronization point and determine that the RAP of the media data follows after the resynchronization point; And Present the retrieved media data including the RAP of the media data and following after the RAP of the media data and do not present the retrieved media data from the resynchronization point to the RAP of the media data.
12. The device according to claim 11, wherein, To present the retrieved media data, the one or more processors are configured to: Parse a file-level media data container of the retrieved media data at the resynchronization point.
13. The device according to claim 12, wherein, To parse the file-level media data container, the one or more processors are configured to: Parse the file-level media data container until the RAP of the media presentation is detected; And Send the RAP to a media decoder.
14. The device according to claim 11, wherein, The resynchronization point includes the start of a chunk boundary.
15. The device according to claim 14, wherein, The chunk boundary includes the start of a chunk, the chunk including zero or one segment type value, zero or one producer reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
16. The apparatus according to claim 11, wherein, The manifest file indicates the availability of the resynchronization points in the segments of the representation.
17. The apparatus according to claim 11, wherein, The one or more processors are configured to retrieve the manifest file before the segment is fully formed, and wherein the one or more processors are configured to send the request for initiating retrieval of the media data as a real-time stream, where the segments of the media presentation are not fully formed until at least some segments in the segment are retrieved.
18. The device according to claim 16, wherein, The manifest file indicates the type of random access that can be performed at the resynchronization point.
19. The apparatus according to claim 16, wherein, The manifest file indicates the position and timing of the resynchronization point and whether the position and timing information is accurate or estimated.
20. The apparatus according to claim 11, wherein The manifest file includes a media presentation description (MPD).
21. A computer-readable storage medium storing instructions that, when executed, cause a processor to perform the following operations: Retrieve a manifest file for media presentation, the manifest file indicating the positions of resynchronization points in segments of a representation of the media presentation, the manifest file further indicating that container parsing of media data of a bitstream can be started at the resynchronization points, the resynchronization points being located at positions other than the start of the segments and representing points at which container parsing of the media data of the bitstream can be started and representing that a random access point (RAP) follows after the resynchronization points; Use the manifest file to form a request for starting to retrieve the media data of the representation at the resynchronization points; Send the request to initiate retrieval of the media data of the media presentation at the resynchronization points; Parse the retrieved media data starting at the resynchronization points and determine that the RAP of the media data follows after the resynchronization points; And Present the retrieved media data including the RAP of the media data and following after the RAP of the media data and do not present the retrieved media data from the resynchronization points to the RAP of the media data.
22. The computer-readable storage medium according to claim 21, wherein, The resynchronization points include the start of chunk boundaries.
23. The computer-readable storage medium according to claim 22, wherein, The chunk boundaries include the start of a chunk, the chunk including zero or one segment type value, zero or one producer reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
24. The computer-readable storage medium according to claim 21, wherein, The manifest file indicates the availability of the resynchronization points in the segments of the representation.
25. The computer-readable storage medium according to claim 21, wherein, The instructions for causing the processor to retrieve the manifest file include instructions for causing the processor to retrieve the manifest file before the segments are fully formed, wherein the instructions for causing the processor to send the request for initiating retrieval of the media data include: instructions for causing the processor to send the request for initiating retrieval of the media data as a real-time stream, wherein the segments of the media presentation are not fully formed before at least some of the segments are retrieved.
26. The computer-readable storage medium according to claim 24, wherein, The manifest file indicates the type of random access that can be performed at the resynchronization points.
27. The computer-readable storage medium according to claim 24, wherein, The manifest file indicates the position and timing of the resynchronization points and whether the position and timing information is accurate or estimated.
28. The computer-readable storage medium according to claim 21, wherein, The manifest file includes a media presentation description (MPD).
29. A device for retrieving media data, the device comprising: A unit for retrieving a manifest file for media presentation, the manifest file indicating the positions of resynchronization points in segments of a representation of the media presentation, the manifest file further indicating that container parsing of media data of a bitstream can be started at the resynchronization points, the resynchronization points being located at positions other than the start of the segments and representing points at which container parsing of the media data of the bitstream can be started and representing that a random access point (RAP) follows after the resynchronization points; A unit for using the manifest file to form a request for starting to retrieve the media data of the representation at the resynchronization points; A unit for sending the request to initiate retrieval of the media data of the media presentation at the resynchronization point; A unit for parsing the retrieved media data starting at the resynchronization point and determining that the RAP of the media data follows after the resynchronization point; and A unit for presenting the RAP of the media data including the media data and the retrieved media data following after the RAP of the media data and not presenting the retrieved media data from the resynchronization point to the RAP of the media data.
30. The apparatus according to claim 29, wherein, The resynchronization point includes the start of a chunk boundary.
31. The apparatus according to claim 30, wherein, The chunk boundary includes the start of a chunk, and the chunk includes zero or one segment type value, zero or one producer reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
32. The apparatus according to claim 29, wherein, The manifest file indicates the availability of the resynchronization point in the segments of the representation.
33. The apparatus according to claim 29, wherein, The unit for retrieving the manifest file includes a unit for retrieving the manifest file before the segment is fully formed, wherein the unit for sending the request for initiating retrieval of the media data includes: a unit for sending the request for initiating retrieval of the media data as a live stream, wherein the segments of the media presentation are not fully formed until at least some of the segments in the segment are retrieved.
34. The apparatus according to claim 32, wherein, The manifest file indicates the type of random access that can be performed at the resynchronization point.
35. The device according to claim 32, wherein, The manifest file indicates the location and timing of the resynchronization point and whether the location and timing information is accurate or estimated.
36. The apparatus according to claim 29, wherein, The manifest file includes a Media Presentation Description (MPD).