Random access at the resynchronization point of the DASH segment
By introducing resynchronization points and signaling within DASH segments, the solution addresses the limitations of random access in existing video streaming technologies, enhancing flexibility and reducing latency in DASH and CMAF systems.
Patent Information
- Authority / Receiving Office
- JP ยท JP
- Patent Type
- Patents
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2025-03-26
- Publication Date
- 2026-07-24
AI Technical Summary
Existing video streaming technologies, such as DASH and CMAF, lack efficient mechanisms for random access within segments, limiting flexibility and adaptability in video playback, especially in low-latency scenarios.
Incorporating resynchronization points and signaling within DASH segments to enable random access at points other than the segment start, allowing flexible initiation of container analysis and data retrieval, and providing additional random access points through chunk boundaries and manifest file signaling.
Enhances video streaming by enabling flexible random access within segments, improving adaptability to network conditions and reducing latency in DASH and CMAF systems.
Smart Images

Figure 0007894968000006 
Figure 0007894968000007 
Figure 0007894968000008
Abstract
Description
Technical Field
[0001]
[0001] This application claims the benefit of U.S. Application No. 17 / 061,152, filed on Oct. 1, 2020, and U.S. Provisional Application No. 62 / 909,642, filed on Oct. 2, 2019, the entire contents of which are incorporated herein by reference.
[0002]
[0002] This disclosure relates to the storage and transport of encoded video data.
Background Art
[0003]
[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular phones or satellite radiotelephones, video teleconferencing devices, and the like. Digital video devices implement video compression techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)), and extensions to such standards, to transmit and receive digital video information more efficiently.
[0004]
[0004] Video compression techniques perform spatial and / or temporal predictions to reduce or eliminate redundancy inherent in video sequences. In block-based video coding, a video frame or slice may be divided into macroblocks. Each macroblock may be further divided. Macroblocks in an intracoded (I) frame or slice are encoded using spatial predictions with respect to adjacent macroblocks. Macroblocks in an intercoded (P or B) frame or slice may use spatial predictions with respect to adjacent macroblocks in the same frame or slice, or temporal predictions with respect to other reference frames.
[0005]
[0005] After the video data has been encoded, the video data may be packetized for transmission or storage. The video data may be assembled into a video file that conforms to any of the various standards, such as AVC, an International Organization for Standardization (ISO) based media file format, and its extensions. [Overview of the Initiative]
[0006]
[0006] Generally speaking, this disclosure describes techniques for accessing data in a segment not only at the beginning of the segment but also at other points within the segment (for example, for random access) in Dynamic Adaptive Streaming over HTTP (DASH) and / or Common Media Access Format (CMAF). This disclosure also describes techniques for signaling the ability to perform random access within a segment. This disclosure describes various use cases for these techniques. For example, this disclosure defines resync points and resync point signaling in DASH and ISO-based Media File Format (BMFF).
[0007]
[0007] In one example, a method for retrieving media data includes retrieving a media presentation manifest file indicating that container analysis of the bitstream media data may be initiated at a resynchronization point of a segment of the media presentation's representation; forming a request to retrieve media data for a representation that starts at a resynchronization point using the manifest file, which indicates a point other than the start of a segment where container analysis of the bitstream media data may be initiated; sending a request to initiate the retrieval of media data for a media presentation that starts at a resynchronization point; and presenting the retrieved media data.
[0008]
[0008] In another example, a device for retrieving media data includes a memory configured to store media data of a media presentation, a manifest file for a media presentation indicating that container parsing of the media data of a bitstream may be initiated at a resynchronization point of a segment of the representation of the media presentation, using the manifest file to form a request to retrieve media data of a representation that starts at a resynchronization point, where the resynchronization point is located elsewhere than the start of a segment and represents a point where container parsing of the media data of a bitstream may be initiated, sending a request to initiate the retrieval of media data of a media presentation that starts at a resynchronization point, and presenting the retrieved media data.
[0009]
[0009] In another example, a computer-readable storage medium stores instructions thereon causing a processor to retrieve a media presentation manifest file indicating that, when executed, container analysis of the media data of a bitstream may be initiated at a resynchronization point of a segment of the media presentation's representation; to use the manifest file to form a request to retrieve media data of a representation that starts at a resynchronization point, where the resynchronization point is located elsewhere than the start of a segment and represents a point where container analysis of the media data of a bitstream may be initiated; to send a request to initiate retrieval of media data of a media presentation that starts at a resynchronization point; and to present the retrieved media data.
[0010]
[0010] In another example, a device for retrieving media data includes means for retrieving a media presentation manifest file indicating that container analysis of the media data of a bitstream may be initiated at a resynchronization point of a segment of the representation of a media presentation; means for forming a request to retrieve media data of a representation that starts at a resynchronization point, using the manifest file, where the resynchronization point is located elsewhere than the start of a segment and represents a point where container analysis of the media data of a bitstream may be initiated; means for sending a request to initiate the retrieval of media data of a media presentation that starts at a resynchronization point; and means for presenting the retrieved media data.
[0011]
[0011] Details of one or more examples are described in the accompanying drawings and the following description. Other features, purposes, and advantages will become apparent from the description and drawings, as well as from the claims. [Brief explanation of the drawing]
[0012] [Figure 1]
[0012] A block diagram illustrating an exemplary system that implements techniques for streaming media data over a network. [Figure 2]
[0013] A block diagram showing an exemplary set of components for the extraction unit. [Figure 3]
[0014] A conceptual diagram illustrating elements of exemplary multimedia content. [Figure 4]
[0015] A block diagram showing elements of an exemplary video file that could correspond to segments of representation. [Figure 5]
[0016] A conceptual diagram illustrating an exemplary low-latency architecture that may be used in the first use case described herein. [Figure 6]
[0017] A conceptual diagram that provides a more detailed example of the use case described in Figure 5. [Figure 7]
[0018] A conceptual diagram illustrating a second exemplary specification example of using DASH and CMAF random access in the context of a broadcast protocol. [Figure 8]
[0019] A conceptual diagram illustrating exemplary signaling for a Stream Access Point (SAP) in a manifest file. [Figure 9]
[0020] A flowchart illustrating an exemplary method for extracting media data using the techniques of this disclosure. [Modes for carrying out the invention]
[0013]
[0021] The techniques of this disclosure may be applied to video files conforming to video data encapsulated according to any of the following: ISO-based media file formats, Scalable Video Coding (SVC) file formats, Advanced Video Coding (AVC) file formats, Third Generation Partnership Project (3GPPยฎ) file formats, and / or Multiview Video Coding (MVC) file formats, or other similar video file formats.
[0014]
[0022] In HTTP streaming, frequently used operations include HEAD, GET, and Partial GET. The HEAD operation retrieves the header of a file associated with a given Uniform Resource Locator (URL) or Uniform Resource Name (URN) without retrieving the payload associated with the URL or URN. The GET operation retrieves the entire file associated with a given URL or URN. The Partial GET operation receives a range of bytes as an input parameter and retrieves several consecutive bytes from the file, the number of bytes corresponding to the received byte range. Thus, a Partial GET operation can obtain one or more individual movie fragments, so a movie fragment for HTTP streaming can be given. Within a movie fragment, there may be several track fragments from different tracks. In HTTP streaming, a media presentation can be a structured collection of data accessible to the client. The client may request and download media data information in order to present the streaming service to the user.
[0015]
[0023] In an example of streaming 3GPP data using HTTP streaming, multiple representations of video and / or audio data of multimedia content may exist. As described below, different representations may correspond to different coding characteristics (e.g., different profiles or levels of the video coding standard), different coding standards or extensions of coding standards (such as multi-view and / or scalable extensions), or different bitrates. Manifests for such representations may be defined in a Media Presentation Description (MPD) data structure. A media presentation may correspond to a structured set of data accessible to an HTTP streaming client device. An HTTP streaming client device may request and download media data information in order to present the streaming service to the user of the client device. A media presentation may be described in an MPD data structure, which may include updates to the MPD.
[0016]
[0024] A media presentation may consist of a sequence of one or more periods. Each period may continue until the start of the next period, or, in the case of the last period, until the end of the media presentation. Each period may contain one or more representations in the same media content. A representation may be audio, video, timed text, or one of several alternative encoded versions of such data. A representation may differ by encoding type, for example, by bitrate, resolution, and / or codec for video data, or by bitrate, language, and / or codec for audio data. The term "representation" may be used to refer to a section of encoded audio or video data that corresponds to a particular period of multimedia content and is encoded in a particular way.
[0017]
[0025] The representation of a particular period can be assigned to a group indicated by an attribute in the MPD that indicates the adaptation set to which the representation belongs. Representations within the same adaptation set are generally considered to be alternatives to each other in that a client device can dynamically and seamlessly switch between these representations, for example, to perform bandwidth adaptation. For example, each representation of video data for a particular period can be assigned to the same adaptation set such that any one of the representations can be selected to decode to present media data such as video data or audio data of the multimedia content for the corresponding period. In some examples, the media content within a period can be represented by either one representation from group 0, if it exists, or a combination of at most one representation from each non-zero group. The timing data for each representation of a period can be expressed relative to the start time of the period.
[0018]
[0026] A representation can include one or more segments. Each representation may include an initialization segment, or each segment of a representation may be self-initializing. When present, the initialization segment may contain initialization information for accessing the representation. Generally, the initialization segment does not contain media data. A segment can be uniquely referenced by an identifier such as a Uniform Resource Locator (URL), Uniform Resource Name (URN), or Uniform Resource Identifier (URI). The MPD can assign an identifier to each segment. In some examples, the MPD can also provide a byte range, in the form of a range attribute, corresponding to the data for a segment within a file accessible by a URL, URN, or URI.
[0019]
[0027] Different representations may be selected for substantially simultaneous retrieval of different types of media data. For example, a client device may select an audio representation, a video representation, and a timed text representation that retrieve segments. In some examples, the client device may select a particular set of adaptations for performing bandwidth adaptation. That is, the client device may select a set of adaptations that includes a video representation, a set of adaptations that includes an audio representation, and / or a set of adaptations that includes timed text. Alternatively, the client device may select a set of adaptations for one type of media (e.g., video) and directly select representations for other types of media (e.g., audio and / or timed text).
[0020]
[0028] Low Latency Dynamic Adaptive Streaming over HTTP (LL-DASH) is a profile for DASH that attempts to provide media data to DASH clients with low latency. Some of the techniques for LL-DASH are briefly summarized below. ยท Encoding is based on fragmented ISO BMFF files, typically assuming CMAF fragments and CMAF chunks. ยท Each chunk is individually accessible by a DASH packager and is mapped to an HTTP chunk that is uploaded to an origin server. This one-to-one mapping is a recommendation for low latency operation but not a requirement. The client should never assume that this one-to-one mapping is stored at the client. ยท A low latency protocol for partially available segments, such as HTTP chunk transfer encoding, is used so that a client can access a segment before it is complete. The available start time is adjusted for clients that can utilize this feature. ยท The following two operating modes are permitted.
[0021] โSimple live offerings are used by applying @duration signaling and $Number$-based templating.
[0022] Main live offerings that have SegmentTimeline as either $Number$ or $Time$ are supported by the updates proposed in DASH version 4. โข The MPD validity expiration event can be used, but it is not essential for the client to understand it. Generally, in-band event messages can exist, but clients are only expected to retrieve them at the beginning of a segment, not at any given chunk. DASH packagers can receive notifications from encoders at chunk boundaries, or entirely asynchronously using timed metadata tracks. โข In a single media presentation, it is permitted to have both an adaptive set using chunked low-latency modes and an adaptive set using shorter segments for different media types within a single duration of the media presentation. A certain amount of playback control for DASH clients on the media pipeline may be available and should be used for the robustness of DASH clients. For example, playback may be accelerated or decelerated for a certain period of time, or DASH clients may perform seeking to segments. The system is designed to run on standard HTTP / 1.1, but should also be applicable to HTTP extensions and other protocols to improve low-latency operation. MPD includes explicit signaling in service configuration and service properties (e.g., target latency of the service). โข MPD and, in some cases, segments also include anchor time, which allows DASH clients to measure current latency compared to live and adjust to meet service expectations. โข For example, operational robustness is addressed in the event of encoder failure. โข Existing DRM and encryption modes are compatible with the proposed low-latency operation.
[0023]
[0029] Based on the high-level overview above, the following is defined: Segments can be used to randomly access representations not only at segment boundaries but also within segments. When such random access is given, this should be signaled in MPD.
[0024]
[0030] Figure 1 is a block diagram illustrating an exemplary system 10 that implements a technique for streaming media data over a network. In this example, system 10 includes a content creation device 20, a server device 60, and a client device 40. The client device 40 and the server device 60 are connected communicatively by a network 74, which may include the Internet. In some examples, the content creation device 20 and the server device 60 may also be connected by network 74 or another network, or they may be connected communicatively directly. In some examples, the content creation device 20 and the server device 60 may be the same device.
[0025]
[0031] In the example shown in Figure 1, the content creation device 20 comprises an audio source 22 and a video source 24. The audio source 22 may comprise, for example, a microphone that generates electrical signals representing captured audio data to be encoded by an audio encoder 26. Alternatively, the audio source 22 may comprise a storage medium for storing previously recorded audio data, an audio data generator such as a computerized synthesizer, or any other source of audio data. The video source 24 may comprise a video camera that generates video data to be encoded by a video encoder 28, a storage medium encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other source of video data. The content creation device 20 may not necessarily be communicatively coupled to the server device 60 in all examples, but it may store multimedia content on separate media that can be read by the server device 60.
[0026]
[0032] Raw audio and video data may comprise analog or digital data. Analog data may be digitized before being encoded by the audio encoder 26 and / or video encoder 28. Audio source 22 may acquire audio data from a call participant while the participant is speaking, and simultaneously, video source 24 may acquire video data of the call participant. In other examples, audio source 22 may comprise a computer-readable storage medium containing stored audio data, and video source 24 may comprise a computer-readable storage medium containing stored video data. Thus, the techniques described herein may be applied to live, streaming, real-time audio and video data, or archived, pre-recorded audio and video data.
[0027]
[0033] An audio frame corresponding to a video frame is generally an audio frame that contains audio data captured (or generated) by an audio source 22 simultaneously with video data captured (or generated) by a video source 24 contained within the video frame. For example, while a call participant generates audio data by speaking, audio source 22 captures the audio data, and simultaneously, i.e., while audio source 22 is capturing the audio data, video source 24 captures the video data of the call participant. Thus, an audio frame may temporally correspond to one or more specific video frames. Therefore, an audio frame corresponding to a video frame generally corresponds to a situation in which audio data and video data are captured simultaneously, and a situation in which an audio frame and a video frame each contain simultaneously captured audio data and video data, respectively.
[0028]
[0034] In some examples, the audio encoder 26 can encode a timestamp in each encoded audio frame, representing the time the audio data of the encoded audio frame was recorded, and similarly, the video encoder 28 can encode a timestamp in each encoded video frame, representing the time the video data of the encoded video frame was recorded. In such examples, the audio frames corresponding to the video frames may include an audio frame with a certain timestamp and a video frame with the same timestamp. The content creation device 20 may include an internal clock from which the audio encoder 26 and / or the video encoder 28 can generate timestamps, or which the audio source 22 and video source 24 can use to associate audio data and video data with timestamps, respectively.
[0029]
[0035] In some examples, an audio source 22 can send data to an audio encoder 26 corresponding to the time the audio data was recorded, and a video source 24 can send data to a video encoder 28 corresponding to the time the video data was recorded. In some examples, the audio encoder 26 can encode a sequence identifier in the encoded audio data to indicate the relative time order of the encoded audio data without necessarily indicating the absolute time the audio data was recorded, and similarly, the video encoder 28 can use a sequence identifier to indicate the relative time order of the encoded video data. Similarly, in some examples, the sequence identifier can be mapped to a timestamp or, in some cases, correlated with a timestamp.
[0030]
[0036] The audio encoder 26 generally generates a stream of encoded audio data, while the video encoder 28 generates a stream of encoded video data. Each individual stream of data (whether audio or video) is sometimes called an elementary stream. An elementary stream is a single, digitally encoded (and possibly compressed) component of a representation. For example, the encoded video or audio portion of a representation can be an elementary stream. An elementary stream can be converted to a packetized elementary stream (PES) before being encapsulated within a video file. Within the same representation, a stream ID may be used to distinguish PES packets belonging to one elementary stream from others. The basic data unit of an elementary stream is a packetized elementary stream (PES) packet. Therefore, encoded video data generally corresponds to an elementary video stream. Similarly, audio data corresponds to one or more of each elementary stream.
[0031]
[0037] Many video coding standards, such as the ITU-T H.264 / AVC and the upcoming High Efficiency Video Coding (HEVC) standard, define syntax, semantics, and a decoding process for error-free bitstreams, all of which conform to a specific profile or level. While video coding standards typically do not specify an encoder, the encoder is obligated to ensure that the resulting bitstream conforms to the decoder's standard. In the context of video coding standards, a โprofileโ corresponds to a subset of algorithms, features, or tools, and the constraints applied to them. For example, a โprofileโ defined by the H.264 standard is a subset of the entire bitstream syntax specified by the H.264 standard. A โlevelโ corresponds to, for example, the picture resolution and the limits on decoder resource consumption, such as decoder memory and computation related to the block processing rate. Profiles can be signaled by profile_idc (profile indicator) values, while levels can be signaled by level_idc (level indicator) values.
[0032]
[0038] The H.264 standard acknowledges that significant variations in encoder and decoder performance may still be required depending on the values โโtaken by syntax elements in the bitstream, such as a specified size of the decoded picture, within the boundaries imposed by the syntax of a given profile. The H.264 standard further acknowledges that in many applications, it is neither practical nor economical to implement a decoder capable of handling all hypothetical uses of the syntax within a particular profile. Therefore, the H.264 standard defines a "level" as a specified set of constraints imposed on the values โโof syntax elements in the bitstream. These constraints may be simple restrictions on values. Alternatively, these constraints may take the form of constraints on combinations of value operations (e.g., picture width ร picture height ร number of pictures decoded per second). The H.264 standard further specifies that individual implementations may support different levels for each supported profile.
[0033]
[0039] A profile-compliant decoder typically supports all features defined in the profile. For example, as a coding feature, B-picture coding is not supported in the H.264 / AVC baseline profile, but is supported in other H.264 / AVC profiles. A level-compliant decoder should be able to decode any bitstream that does not require resources beyond the limits defined in the level. Profile and level definitions can serve an explainability purpose. For example, during a video transmission, pairs of profile and level definitions may be negotiated and agreed upon for the entire transmission session. More specifically, in H.264 / AVC, a level may define limits on the number of macroblocks that need to be processed, the decoded picture buffer (DPB) size, the coded picture buffer (CPB) size, the vertical motion vector range, the maximum number of motion vectors per two consecutive MBs, and whether a B-block can have sub-macroblock divisions smaller than 8x8 pixels. In this way, a decoder can determine whether it is capable of properly decoding a bitstream.
[0034]
[0040] In the example shown in Figure 1, the encapsulation unit 30 of the content creation device 20 receives an elementary stream containing encoded video data from the video encoder 28 and an elementary stream containing encoded audio data from the audio encoder 26. In some examples, the video encoder 28 and the audio encoder 26 may each include a packetizer for forming PES packets from the encoded data. In other examples, the video encoder 28 and the audio encoder 26 may each interface with their respective packetizers for forming PES packets from the encoded data. In yet another example, the encapsulation unit 30 may include a packetizer for forming PES packets from the encoded audio data and the encoded video data.
[0035]
[0041] The video encoder 28 can encode video data of multimedia content in various ways to produce different representations of multimedia content using various characteristics such as pixel resolution, frame rate, compliance with various coding standards, compliance with various profiles and / or levels of profiles for various coding standards, representations having one or more views (for example, for two-dimensional or three-dimensional playback), or other such characteristics at various bitrates. The representations used in this disclosure may comprise one of audio data, video data, text data (for example, for closed captions), or other such data. The representations may include elementary streams, such as audio elementary streams or video elementary streams. Each PES packet may include a stream_id that identifies the elementary stream to which the PES packet belongs. The encapsulation unit 30 is responsible for assembling the elementary streams into video files (for example, segments) of the various representations.
[0036]
[0042] The encapsulation unit 30 receives PES packets of the elementary stream of the representation from the audio encoder 26 and the video encoder 28, and forms corresponding network abstraction layer (NAL) units from the PES packets. The coded video segment can be organized into NAL units that provide a โnetwork-friendlyโ video representation to address applications such as videophone, storage, broadcast, or streaming. NAL units can be classified into video coding layer (VCL) NAL units and non-VCL NAL units. VCL units may include a core compression engine and may include block, macroblock, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some examples, a coded picture in one time instance, typically presented as a primary coded picture, may be contained within an access unit that may contain one or more NAL units.
[0037]
[0043] Non-VCL NAL units may include, in particular, parameter set NAL units and SEI NAL units. Parameter sets may include sequence-level header information (in a sequence parameter set (SPS)) and rarely changing picture-level header information (in a picture parameter set (PPS)). With parameter sets (e.g., PPS and SPS), rarely changing information does not need to be repeated per sequence or per picture, thus improving coding efficiency. Furthermore, the use of parameter sets allows for out-of-band transmission of critical header information, avoiding the need for redundant transmission for error tolerance. In an example of out-of-band transmission, a parameter set NAL unit, such as an SEI NAL unit, may be transmitted on a different channel than other NAL units.
[0038]
[0044] Supplemental Enhancement Information (SEI) is not required to decode coded picture samples from VCL NAL units, but may contain information that can assist processes related to decoding, display, error tolerance, and other purposes. SEI messages may be included in non-VCL NAL units. SEI messages are a normative part of some standards specifications and are therefore not always required for standard-compliant decoder implementations. SEI messages can be sequence-level SEI messages or picture-level SEI messages. SEI messages may contain some sequence-level information, such as the scalability information SEI message in the SVC example and the view scalability information SEI message in MVC. These exemplary SEI messages may carry, for example, information about the extraction of operating points and the characteristics of those operating points. In addition, the encapsulation unit 30 may form a manifest file, such as a Media Presentation Descriptor (MPD) that describes the characteristics of the representation. The encapsulation unit 30 may format the MPD according to Extensible Markup Language (XML).
[0039]
[0045] The encapsulation unit 30 may provide data for one or more representations of multimedia content to the output interface 32, along with a manifest file (e.g., MPD). The output interface 32 may include a Universal Serial Bus (USB) interface, a network interface or interface for writing to a storage medium such as a CD or DVD writer or burner, an interface to a magnetic or flash storage medium, or other interfaces for storing or transmitting media data. The encapsulation unit 30 can provide data for each representation of the multimedia content to the output interface 32, which can then send that data to the server device 60 via network transmission or storage medium. In the example in Figure 1, the server device 60 includes a storage medium 62 for storing various multimedia content 64, each multimedia content 64 including its own manifest file 66 and one or more representations 68A-68N (representation 68). In some examples, the output interface 32 may also send data directly to the network 74.
[0040]
[0046] In some examples, representation 68 can be separated into adaptive sets. That is, various subsets of representation 68 may include each common set of characteristics, such as codecs, profiles and levels, resolution, number of views, segment file formats, text type information that can identify the language or other characteristics of the text and / or audio data to be displayed using the representation and / or decoded and presented, for example, by a speaker, camera angle information that can describe the camera angle of the scene or the real-world camera perspective for the representation in the adaptive set, and rating information that describes the content suitability for a particular audience.
[0041]
[0047] The manifest file 66 may contain data indicating a subset of representations 68 corresponding to a particular adaptation set, and common characteristics of the adaptation set. The manifest file 66 may also contain data representing individual characteristics of individual representations of the adaptation set, such as bitrate. In this way, the adaptation set can enable simplified network bandwidth adaptation. Representations within the adaptation set may be indicated using child elements of the adaptation set element in the manifest file 66.
[0042]
[0048] The server device 60 includes a request processing unit 70 and a network interface 72. In some examples, the server device 60 may include multiple network interfaces. Furthermore, some or all of the functions of the server device 60 may be implemented on other devices in the content distribution network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, intermediate devices in the content distribution network may include components that cache data for multimedia content 64 and substantially conform to the components of the server device 60. Generally, the network interface 72 is configured to send and receive data over the network 74.
[0043]
[0049] The request processing unit 70 is configured to receive network requests for data on the storage medium 62 from client devices, such as the client device 40. For example, the request processing unit 70 may implement the Hypertext Transfer Protocol (HTTP) version 1.1, as described in RFC 2616, "Hypertext Transfer Protocol - HTTP / 1.1", Network Working Group, IETF, June 1999, by R. Fielding et al. That is, the request processing unit 70 may be configured to receive HTTP GET requests or partial GET requests and to provide data for the multimedia content 64 in response to these requests. A request may specify one segment of the representation 68, for example, using the segment's URL. In some examples, a request may also specify one or more byte ranges of a segment, thus comprising a partial GET request. The request processing unit 70 may be further configured to service HTTP HEAD requests for header data for one segment of the representation 68. In any case, the request processing unit 70 may be configured to process requests for providing the requested data to the requesting device, such as the client device 40.
[0044]
[0050] As an addition or alternative, the request processing unit 70 may be configured to deliver media data via a broadcast or multicast protocol such as eMBMS. The content creation device 20 may create DASH segments and / or subsegments in substantially the same manner as described, but the server device 60 may deliver these segments or subsegments using eMBMS or another broadcast or multicast network transport protocol. For example, the request processing unit 70 may be configured to receive multicast group join requests from client devices 40. That is, the server device 60 can advertise the Internet Protocol (IP) address associated with the multicast group to client devices associated with specific media content (e.g., a broadcast of a live event), including client device 40. Client device 40 can then submit a request to join the multicast group. This request may be propagated across the network 74, for example, to routers that make up the network 74, so that the routers direct traffic destined for the IP address associated with the multicast group to joining client devices such as client device 40.
[0045]
[0051] As shown in the example in Figure 1, the multimedia content 64 includes a manifest file 66 which may correspond to a Media Presentation Description (MPD). The manifest file 66 may include descriptions of different alternative representations 68 (for example, video services with different qualities), which may include, for example, codec information, profile values, level values, bitrate, and other description characteristics of the representation 68. The client device 40 may retrieve the MPD of the media presentation to determine how to access segments of the representation 68.
[0046]
[0052] More specifically, the retrieval unit 52 may retrieve configuration data (not shown) of the client device 40 to determine the decoding capability of the video decoder 48 and the rendering capability of the video output 44. The configuration data may also include any or all of the following: language preferences selected by the user of the client device 40, one or more camera perspectives corresponding to depth preferences set by the user of the client device 40, and / or rating preferences selected by the user of the client device 40. The retrieval unit 52 may comprise, for example, a web browser or media client configured to submit HTTP GET requests and partial GET requests. The retrieval unit 52 may correspond to software instructions executed by one or more processors or processing units (not shown) of the client device 40. In some examples, all or part of the functions described with respect to the retrieval unit 52 may be implemented in hardware, or in a combination of hardware, software, and / or firmware, and hardware may be provided that is essential for executing software or firmware instructions.
[0047]
[0053] The extraction unit 52 may compare the decoding and rendering capabilities of the client device 40 with the characteristics of the representation 68 indicated by the information in the manifest file 66. The extraction unit 52 may first extract at least a portion of the manifest file 66 to determine the characteristics of the representation 68. For example, the extraction unit 52 may request a portion of the manifest file 66 that describes the characteristics of one or more adaptive sets. The extraction unit 52 may select a subset of representations 68 (e.g., adaptive sets) that have characteristics that can be satisfied by the coding and rendering capabilities of the client device 40. The extraction unit 52 may then determine the bitrate of the representations in the adaptive sets, determine the amount of network bandwidth currently available, and extract a segment from one of the representations that has a bitrate that can be satisfied by the network bandwidth.
[0048]
[0054] Generally, higher bitrate representations can result in higher quality video playback, but when available network bandwidth decreases, lower bitrate representations can provide sufficient quality video playback. Therefore, when available network bandwidth is relatively high, the extraction unit 52 can extract data from a relatively high bitrate representation, but when available network bandwidth is low, the extraction unit 52 can extract data from a relatively low bitrate representation. In this way, the client device 40 can stream multimedia data over the network 74 while also adapting to the changing network bandwidth availability of the network 74.
[0049]
[0055] As an addition or alternative, the retrieval unit 52 may be configured to receive data according to a broadcast or multicast network protocol such as eMBMS or IP multicast. In such an example, the retrieval unit 52 may submit a request to join a multicast network group associated with a particular media content. After joining the multicast group, the retrieval unit 52 may receive data from the multicast group without further requests issued to the server device 60 or content creation device 20. When the data from the multicast group is no longer needed, the retrieval unit 52 may submit a request to leave the multicast group, for example, to stop playback or to change the channel to a different multicast group.
[0050]
[0056] The network interface 54 can receive data from a selected representation segment and provide it to the extraction unit 52, which in turn can provide the segment to the decapsulation unit 50. The decapsulation unit 50 decapsulates the video file elements into a constituent PES stream, depackets the PES stream to extract the encoded data, and can send the encoded data to either the audio decoder 46 or the video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream, as indicated by the PES packet header of the stream. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, and the video decoder 48 decodes the encoded video data and sends the decoded video data, which may contain multiple views of the stream, to the video output 44.
[0051]
[0057] The video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, extraction unit 52, and decapsulation unit 50 may each be implemented as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof, where applicable. Each of the video encoder 28 and video decoder 48 may be included in one or more encoders or decoders, and any of them may be integrated as part of a composite video encoder / decoder (codec). Similarly, each of the audio encoder 26 and audio decoder 46 may be included in one or more encoders or decoders, and any of them may be integrated as part of a composite codec. The device including the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, extraction unit 52, and / or decapsulation unit 50 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular phone.
[0052]
[0058] The client device 40, the server device 60, and / or the content creation device 20 may be configured to operate in accordance with the techniques of this disclosure. As an example, this disclosure describes these techniques with respect to the client device 40 and the server device 60. However, it should be understood that the content creation device 20 may be configured to perform these techniques instead of (or in addition to) the server device 60.
[0053]
[0059] The encapsulation unit 30 may form a NAL unit comprising a header that identifies the program to which the NAL unit belongs, and a payload, for example, audio data, video data, or data describing the transport or program stream to which the NAL unit corresponds. For example, in H.264 / AVC, the NAL unit includes a 1-byte header and a payload of varying size. A NAL unit that includes video data in its payload may have video data at various granularity levels. For example, the NAL unit may have blocks of video data, multiple blocks, slices of video data, or entire pictures of video data. The encapsulation unit 30 may receive encoded video data from the video encoder 28 in the form of PES packets of elementary streams. The encapsulation unit 30 may associate each elementary stream with a corresponding program.
[0054]
[0060] The encapsulation unit 30 can also assemble access units from multiple NAL units. Generally, an access unit may comprise one or more NAL units for representing frames of video data, and, if audio data corresponding to those frames is available, such audio data. Generally, an access unit includes all NAL units across one output time instance, for example, all audio and video data across one time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, the unique frames of all views of the same access unit (same time instance) can be rendered simultaneously. In one example, an access unit may comprise a coded picture in one time instance, which may be presented as a primary coded picture.
[0055]
[0061] Therefore, an access unit may contain all audio and video frames of a common time instance, for example, all views corresponding to time X. This disclosure also refers to the encoded pictures of a particular view as โview components.โ That is, a view component may contain the encoded pictures (or frames) of a particular view at a particular time. Therefore, an access unit can be defined as one that contains all view components of a common time instance. The decoding order of an access unit does not necessarily have to be the same as the output order or display order.
[0056]
[0062] A media presentation may include a media presentation description (MPD) which may contain descriptions of different alternative representations (e.g., video services with different qualities), and this description may include, for example, codec information, profile values, and level values. An MPD is an example of a manifest file, such as manifest file 66. To determine how to access movie fragments of various presentations, a client device 40 may retrieve the MPD of the media presentation. Movie fragments may be placed in movie fragment boxes (moof boxes) of a video file.
[0057]
[0063] A manifest file 66 (which may, for example, include an MPD) may advertise the availability of segments in representation 68. That is, the MPD may include information indicating the wall clock time when one of the first segments in representation 68 becomes available, and information indicating the duration of the segments in representation 68. In this way, the retrieval unit 52 of the client device 40 may determine when each segment is available based on the start time and the duration of the segments preceding a particular segment.
[0058]
[0064] After the encapsulation unit 30 assembles the NAL unit and / or access unit into a video file based on the received data, the encapsulation unit 30 passes the video file to the output interface 32 for output. In some examples, the encapsulation unit 30 may store the video file locally or send the video file to a remote server via the output interface 32, rather than sending the video file directly to the client device 40. The output interface 32 may comprise, for example, a transmitter, transceiver, a device for writing data to a computer-readable medium such as an optical drive or magnetic media drive (e.g., a floppy disk drive), a Universal Serial Bus (USB) port, a network interface, or other output interface. The output interface 32 outputs the video file to a computer-readable medium such as a transmission signal, magnetic media, optical media, memory, a flash drive, or other computer-readable medium.
[0059]
[0065] The network interface 54 can receive NAL units or access units via the network 74 and provide the NAL units or access units to the decapsulation unit 50 via the extraction unit 52. The decapsulation unit 50 decapsulates the elements of the video file into a constituent PES stream, depackets the PES stream to extract the encoded data, and can send the encoded data to either the audio decoder 46 or the video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream, as indicated by the PES packet header of the stream, for example. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, and the video decoder 48 decodes the encoded video data and sends the decoded video data, which may contain multiple views of the stream, to the video output 44.
[0060]
[0066] According to the techniques of this disclosure, the content creation device 20 and / or server device 60 may add additional random access points during the DASH / CMAF segment. Random access includes clean random access and open or incremental decoder refreshes, until only resynchronization is given during file format parsing. This can be addressed by providing chunk boundaries that provide information that resynchronization and decoding may be initiated at this point, and signaling regarding the type of random access point. The availability of tfdt, along with the use of moof header information and possibly initialization segments, enables time resynchronization at the presentation time level. In this disclosure, this new point is referred to as a โresynchronization point.โ That is, a resynchronization point represents a point where a file-level container (e.g., a box in ISO BMFF) can be properly parsed, followed by random access points in the media data (e.g., I-frames). Thus, the client device 40 may randomly access the multimedia content 64 at, for example, one of these random access points.
[0061]
[0067] The content creation device 20 and / or the server device 60 may also add appropriate signaling in the manifest file 66 (e.g., MPD) that indicates the availability of random access points and resynchronization in each DASH segment, and provides information about the location, type, and timing of random access points. The content creation device 20 and / or the server device 60 may optionally provide signaling in the manifest file 66 (MPD) that indicates additional resynchronization points are available in the segment, adding characteristics about the location, timing, and type of random access, as well as whether the information is accurate or estimated. Thus, the client device 40 may use this signaled data to determine whether such random access points are available and to resynchronize retrieval and playback accordingly.
[0062]
[0068] The client device 40 may be configured with the ability to resynchronize to decapsulation, decryption, and decoding by finding a resynchronization point, given any starting point. The content creation device 20 and / or server device 60 may provide appropriate chunks that satisfy the above requirements addressed in CMAF TUC. Different types may be defined later.
[0063]
[0069] The client device 40 may be configured to initiate processing in a restricted receiver environment, for example, as is available in HTML-5 / MSE-based playback. This issue can be addressed through the receiver implementation. However, it is appropriate to provide resynchronization triggers and information to the receiver pipeline, which has the ability to obtain a map of resynchronization points in data structures, timing, and types, thereby enabling the user of the decoding pipeline to initialize playback at random access points.
[0064]
[0070] Furthermore, signaling may be provided in a backward-compatible manifest file 66. Client devices not configured with the ability to parse the signaling may ignore the signaling and perform the actions described above. Additionally, manifest file 66 may include signaling that associates locations with @bandwidth values โโto enable signaling at the adaptive set level.
[0065]
[0071] Figure 2 is a block diagram showing in more detail an exemplary set of components of the retrieval unit 52 of Figure 1. In this example, the retrieval unit 52 includes an eMBMS middleware unit 100, a DASH client 110, and a media application 112.
[0066]
[0072] In this example, the eMBMS middleware unit 100 further includes an eMBMS receiving unit 106, a cache 104, and a proxy server unit 102. In this example, the eMBMS receiving unit 106 is configured to receive data via eMBMS in accordance with File Delivery over Unidirectional Transport (FLUTE), as described by T. Paira et al., "FLUTE - File Delivery over Unidirectional Transport," Network Working Group, RFC6726, November 2012, available, for example, at tools.ietf.org / html / rfc6726. That is, the eMBMS receiving unit 106 can receive files via broadcast from, for example, a server device 60 that can act as a Broadcast / Multicast Service Center (BM-SC).
[0067]
[0073] When the eMBMS middleware unit 100 receives data relating to a file, the eMBMS middleware unit may store the received data in the cache 104. The cache 104 may comprise a computer-readable storage medium such as flash memory, a hard disk, RAM, or any other suitable storage medium.
[0068]
[0074] The proxy server unit 102 can act as a server for the DASH client 110. For example, the proxy server unit 102 can provide the DASH client 110 with an MPD file or other manifest file. The proxy server unit 102 can advertise the availability time for segments in the MPD file and hyperlinks from which the segments can be retrieved. These hyperlinks may include a local host address prefix corresponding to the client device 40 (for example, 127.0.0.1 for IPv4). In this way, the DASH client 110 can request a segment from the proxy server unit 102 using an HTTP GET or partial GET request. For example, for a segment available from the link http: / / 127.0.0.1 / rep1 / seg3, the DASH client 110 can construct an HTTP GET request that includes a request for http: / / 127.0.0.1 / rep1 / seg3 and submit that request to the proxy server unit 102. The proxy server unit 102 may retrieve the requested data from the cache 104 and, in response to such a request, provide that data to the DASH client 110.
[0069]
[0075] Figure 3 is a conceptual diagram showing the elements of an exemplary multimedia content 120. The multimedia content 120 may correspond to multimedia content 64 (Figure 1) or another multimedia content stored in the storage medium 62. In the example in Figure 3, the multimedia content 120 includes a media presentation description (MPD) 122 and multiple representations 124A to 124N (representation 124). Representation 124A includes arbitrary header data 126 and segments 128A to 128N (segment 128), and representation 124N includes arbitrary header data 130 and segments 132A to 132N (segment 132). The character N is used for convenience to specify the last movie fragment in each of the representations 124. In some examples, there may be different numbers of movie fragments between representations 124.
[0070]
[0076] MPD122 may have a different data structure from representation 124. MPD122 may correspond to manifest file 66 in Figure 1. Similarly, representation 124 may correspond to representation 68 in Figure 1. Generally, MPD122 may contain data that collectively describes the characteristics of representation 124, such as coding and rendering characteristics, adaptive sets, the profile to which MPD122 corresponds, text type information, camera angle information, rating information, trick mode information (e.g., information indicating representations that include time subsequences), and / or information for extracting remote periods (e.g., for inserting targeted advertisements into media content during playback).
[0071]
[0077] Header data 126, if present, may describe characteristics of segment 128, such as the time location of a random access point (RAP, also known as a stream access point (SAP)), and the random access points of segment 128 may include the random access point, the byte offset to the random access point within segment 128, the uniform resource locator (URL) of segment 128, or other aspects of segment 128. Header data 130, if present, may describe similar characteristics of segment 132. Additional or alternative, such characteristics may be entirely contained within MPD 122.
[0072]
[0078] Segments 128 and 132 include one or more coded video samples, each of which may include a frame or slice of video data. Each coded video sample in segment 128 may have similar characteristics, such as height, width, and bandwidth requirements. Such characteristics may be described by data in MPD122, but such data is not shown in the example in Figure 3. MPD122 may include characteristics described by the 3GPP specification in addition to any or all of the signaled information described herein.
[0073]
[0079] Each of segments 128 and 132 can be associated with a unique uniform resource locator (URL). Thus, each of segments 128 and 132 can be retrieved independently using a streaming network protocol such as DASH. In this way, a destination device, such as client device 40, can use an HTTP GET request to retrieve segment 128 or 132. In some examples, client device 40 can use an HTTP partial GET request to retrieve a specific byte range of segment 128 or 132.
[0074]
[0080] According to the techniques of this disclosure, MPD122 (which may also correspond to manifest file 66 in Figure 1) may include signaling to address the problems described above. For example, for each of the representations 124 (and possibly defaulted at the adaptive set level), MPD122 may include one or more resynchronization elements (which may enable backward compatibility). Each resynchronization element may indicate that for each of the segments 128, 132 in one of the corresponding representations 124, the following holds: Resynchronization points for Stream Access Point (SAP) types of @type or smaller (but greater than 0) exist within each segment having a maximum delta T signaled by @dT, a maximum byte offset difference signaled by @dImax, and a minimum byte offset difference signaled by @dImin, where both @dImax and @dImin values โโmust be multiplied by the value of the @bandwidth attribute assigned to this representation in order to obtain a value. "Delta T" refers to the difference in the earliest presentation time of any data following the resynchronization point in units of the @timescale of the representation. If @type is set to 0, only resynchronization with respect to the container and decryption level is guaranteed. The resynchronization marker flag @marker can be set to indicate that each resynchronization point contains another resynchronization point, using a resynchronization pattern defined by the segment format being used. โข Multiple resynchronization elements may exist for different SAP types. The resynchronization point requires that media processing, in combination with the CMAF header / initialization segment, can occur on the ISO BMFF and decoded information.
[0075]
[0081] An illustrative use of the resynchronization point is illustrated with reference to Figure 8 below.
[0076]
[0082] A resynchronization element is provided in MPD122, and with respect to either segment 128 or 132, which is based on ISO BMFF or CMAF, the following may hold: โข A segment can be a sequence of one or more chunks as defined below. โข Furthermore, with respect to any two consecutive resynchronization points in a segment of @type specified in an element, the following may hold: โThe difference in the shortest presentation time between the two can be at most the value of @dT.
[0077] โThe difference in byte offset from the start can be at most @dImax normalized by the @bandwidth value.
[0078] โThe difference in byte offsets from the start can be at least @dImin normalized by the @bandwidth value. โข If the resynchronization marker flag is set, each resynchronization point may contain a resynchronization box / styp.
[0079]
[0083] Figure 4 is a block diagram showing elements of an exemplary video file 150 that may correspond to a segment of representation, such as one of segments 128 and 132 in Figure 3. Each of segments 128 and 132 may contain data that substantially matches the sequence of data shown in the example in Figure 4. The video file 150 is sometimes said to encapsulate segments. As explained above, video files by the ISO-based media file format and its extensions store data in a set of objects called "boxes". In the example in Figure 4, the video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a segment index (sidx) box 162, a movie fragment (MOOF) box 164, and a movie fragment random access (MFRA) box 166. While Figure 4 represents an example of a video file, it should be understood that other media files may contain other types of media data (e.g., audio data, timed text data, etc.) structured similarly to the data in the video file 150, according to the ISO-based media file format and its extensions.
[0080]
[0084] The File Type (FTYP) box 152 generally describes the file type of the video file 150. The File Type box 152 may contain data identifying a specification that describes the best use of the video file 150. Alternatively, the File Type box 152 may be placed before the MOOV box 154, the Movie Fragment box 164, and / or the MFRA box 166.
[0081]
[0085] In some examples, a segment such as video file 150 may include an MPD update box (not shown) before the FTYP box 152. The MPD update box may contain information indicating that the MPD corresponding to the representation containing video file 150 should be updated, along with information for updating the MPD. For example, the MPD update box may provide a URI or URL for the resource used to update the MPD. In another example, the MPD update box may contain data for updating the MPD. In some examples, the MPD update box may come immediately after the segment type (STYP) box (not shown) for video file 150, where the STYP box may define the segment type for video file 150.
[0082]
[0086] In the example in Figure 4, the MOOV box 154 includes a movie header (MVHD) box 156, a track (TRAK) box 158, and one or more movie extension (MVEX) boxes 160. Generally, the MVHD box 156 may describe the general characteristics of the video file 150. For example, the MVHD box 156 may include data describing when the video file 150 was first created, when the video file 150 was last modified, the timeline of the video file 150, the duration of playback of the video file 150, or other data that generally describes the video file 150.
[0083]
[0087] The TRAK box 158 may contain data about the track of the video file 150. The TRAK box 158 may contain a Track Header (TKHD) box that describes the characteristics of the track corresponding to the TRAK box 158. In some examples, the TRAK box 158 may contain an encoded video picture, while in other examples, the encoded video picture of the track may be contained in a movie fragment 164 that can be referenced by the data in the TRAK box 158 and / or the sidx box 162.
[0084]
[0088] In some examples, video file 150 may contain two or more tracks. Therefore, MOOV box 154 may contain several TRAK boxes equal to the number of tracks in video file 150. TRAK box 158 may describe the characteristics of the corresponding track in video file 150. For example, TRAK box 158 may describe temporal and / or spatial information about the corresponding track. When the encapsulation unit 30 (Figure 3) includes a parameter set track in a video file such as video file 150, TRAK boxes similar to TRAK box 158 in MOOV box 154 may describe the characteristics of the parameter set track. The encapsulation unit 30 may signal the presence of a sequence-level SEI message in the parameter set track within the TRAK box describing the parameter set track.
[0085]
[0089] The MVEX box 160 may, for example, describe the properties of the corresponding movie fragment 164 so as to signal that the video file 150 contains the movie fragment 164 in addition to the video data contained in the MOOV box 154, if any. In the context of streaming video data, coded video pictures may be contained in the movie fragment 164 rather than in the MOOV box 154. Therefore, all coded video samples may be contained in the movie fragment 164 rather than in the MOOV box 154.
[0086]
[0090] A MOOV box 154 may contain several MVEX boxes 160 equal to the number of movie fragments 164 in the video file 150. Each MVEX box 160 may describe one corresponding characteristic of one of the movie fragments 164. For example, each MVEX box may contain a movie extension header box (MEHD) box that describes the duration of one corresponding movie fragment 164.
[0087]
[0091] As described above, the encapsulation unit 30 may store a sequence dataset in a video sample that does not contain the actual coded video data. A video sample may generally correspond to an access unit, which is a representation of a coded picture in a particular time instance. In the context of AVC, a coded picture includes one or more VCL NAL units containing information to constitute all the pixels of the access unit, and other related non-VCL NAL units such as SEI messages. Thus, the encapsulation unit 30 may include a sequence dataset in one of the movie fragments 164, which may contain a sequence-level SEI message. The encapsulation unit 30 may further signal the presence of the sequence dataset and / or the sequence-level SEI message as being present in one of the movie fragments 164 within one of the MVEX boxes 160 corresponding to one of the movie fragments 164.
[0088]
[0092] The SIDX box 162 is an optional element of the video file 150. That is, a video file conforming to the 3GPP file format, or any other such file format, does not necessarily contain a SIDX box 162. According to the example of the 3GPP file format, a SIDX box may be used to identify a subsegment of a segment (for example, a segment contained within the video file 150). The 3GPP file format defines a subsegment as "a self-contained set of one or more consecutive movie fragment boxes, each having a corresponding media data box and a media data box containing data referenced by a movie fragment box, which must follow that movie fragment box and precede the next movie fragment box containing information about the same track." The 3GPP file format also indicates that a SIDX box "contains a sequence of references to subsegments of the (sub)segment documented by the box. The referenced subsegments are consecutive in presentation time. Similarly, bytes referenced by a segment index box are always consecutive within a segment. The referenced size gives a count of the number of bytes in the referenced material."
[0089]
[0093] The SIDX box 162 generally provides information representing one or more subsegments of a segment contained within the video file 150. For example, such information may include the playback time at which the subsegment begins and / or ends, the byte offset of the subsegment, whether the subsegment contains a stream access point (SAP) (e.g., starts with one), the type of SAP (e.g., whether the SAP is an instantaneous decoder refresh (IDR) picture, a clean random access (CRA) picture, or a broken link access (BLA) picture), the position of the SAP within the subsegment (relative to playback time and / or byte offset), and so on.
[0090]
[0094] A movie fragment 164 may contain one or more coded video pictures. In some examples, a movie fragment 164 may contain one or more picture groups (GOPs), each of which may contain several coded video pictures, such as frames or pictures. Furthermore, as described above, a movie fragment 164 may, in some examples, contain a sequence dataset. Each movie fragment 164 may contain a movie fragment header box (MFHD, not shown in Figure 4). The MFHD box may describe characteristics of the corresponding movie fragment, such as the sequence number of the movie fragment. The movie fragments 164 may be included in the order of their sequence numbers in the video file 150.
[0091]
[0095] The MFRA box 166 may describe random access points within movie fragments 164 of the video file 150. This can help perform trick modes, such as performing seeks to specific time locations (i.e., playback time) within segments encapsulated by the video file 150. The MFRA box 166 is generally optional and, in some examples, does not need to be included in the video file. Similarly, a client device, such as the client device 40, does not necessarily need to refer to the MFRA box 166 in order to correctly decode and display the video data of the video file 150. The MFRA box 166 may include several Track Fragment Random Access (TFRA) boxes (not shown), equal to the number of tracks in the video file 150, or, in some examples, equal to the number of media tracks (e.g., non-hint tracks) in the video file 150.
[0092]
[0096] In some examples, a movie fragment 164 may contain one or more stream access points (SAPs), such as IDR pictures. Similarly, an MFRA box 166 may provide location instructions for the SAPs within the video file 150. Thus, a time subsequence of the video file 150 may be formed from the SAPs of the video file 150. The time subsequence may also contain other pictures, such as P frames and / or B frames, dependent on the SAPs. Frames and / or slices of the time subsequence may be structured within a segment so that frames / slice of the time subsequence that depend on other frames / slice of the subsequence can be properly decoded. For example, in a data hierarchy, data used for predicting other data may also be included in the time subsequence.
[0093]
[0097] For example, this disclosure defines โchunkโ in the table below as follows. This table provides exemplary definitions of both the cardinality and ordering of chunks.
[0094] [Table 1]
[0095]
[0098] In one example, this disclosure defines a resynchronization point as the start of a chunk. Furthermore, a resynchronization point may be assigned the following properties: It has a byte offset from the start of the segment that points to the first byte of the chunk. It has the earliest presentation time, which is derived from the information in the movie and, if applicable, the movie header assigned to it. It has an SAP type assigned to it, as defined in ISO / IEC 14496-12. โข There is an indication of whether the chunk contains a resynchronization box (the styp shown above). Starting from the resynchronization point, file format analysis and decoding can be performed along with the information in the movie header.
[0096]
[0099] In some examples, a resynchronization marker box may be defined that allows synchronization to the start of a segment (e.g., video file 150) by scanning a byte stream of markers. This resynchronization box may have the following properties: It defines a unique pattern for resynchronization with extremely high likelihood. โข It defines the SAP type.
[0097]
[0100] The resynchronization marker box may be a new box or an existing box, such as a styp box, may be reused. This disclosure assumes that a styp with certain restrictions may be used as the resynchronization marker box. Research on the robustness of this method is ongoing.
[0098]
[0101] Figures 5-7 below are used to illustrate some use cases where random access to a DASH segment, other than at the start of the segment, may be useful.
[0099]
[0102] Figure 5 is a conceptual diagram showing an exemplary low-latency architecture 200 that may be used in a first use case of the present disclosure. Specifically, Figure 5 shows the basic flow of information for operating a low-latency DASH service with DASH-IF IOP. The low-latency architecture 200 includes a DASH packager 202, an encoder 216, a content delivery network (CDN) 220, a regular DASH client 230, and a low-latency DASH client 232. The encoder 216 may generally correspond to either or both of the audio encoder 26 and the video encoder 28 in Figure 1, while the DASH packager 202 may correspond to the encapsulation unit 30 in Figure 1.
[0100]
[0103] In this example, encoder 216 encodes the received media data to form a CMAF header (CH) including CH208, CMAF initial chunks 206A and 206B (CIC206), and CMAF non-initial chunks 204A to 204D (CNC204). Encoder 216 provides CH208, CIC206, and CNC204 to DASH packager 202. DASH packager 202 also receives a service description containing a general description of the service and information about the encoder configuration of encoder 216.
[0101]
[0104] The DASH packager 202 uses the service description, CH208, CIC206, and CNC204 to form a media presentation description (MPD) 210 and an initialization segment 212. The DASH packager 202 also generates maps of CH208, CIC206, and CNC204 within segments 214A, 214B (segment 214) and provides segment 214 to the CDN220 in an incremental manner. The DASH packager 202 may deliver segment 214 in the form of chunks as they are generated. The CDN220 includes segment storage 222 for storing the MPD 210, IS212, and segment 214. CDN220 responds to HTTP Get or partial Get requests from, for example, a regular DASH client 230 and a low-latency DASH client 232 by delivering the complete segment to the regular DASH client 230, but individual chunks (e.g., CH208, CIC206, and CNC204) to the low-latency DASH client 232.
[0102]
[0105] Figure 6 is a conceptual diagram that provides further detail to an example of the use case described in relation to Figure 5. The example in Figure 6 shows segments 250A to 250E (segment 250), each containing a set of chunks 252A to 252E (chunk 252). A client device, such as client device 40 in Figure 1, can retrieve either the complete segment 250 or individual chunks 252. For example, as shown in Figure 5, a normal DASH client 230 can retrieve segment 250, while a low-latency DASH client 232 can retrieve individual chunks 252 (at least initially).
[0103]
[0106] Figure 6 further illustrates how latency can be reduced by retrieving individual chunks 252 rather than the complete segment 250. For example, retrieving the complete segment at the present time may result in higher latency. Simply retrieving the most recent fully available segment reduces latency, but may still result in relatively high latency.
[0104]
[0107] These latencies can be significantly reduced by extracting chunks instead. For example, at the present time indicated by "now" in Figure 6, segment 250E is not yet fully formed. Nevertheless, a client device can extract formed chunks such as chunks 252E-1 and 252E-2 of segment 250E, even before segment 250E is fully formed, assuming that chunks 252E-1 and 252E-2 are formed and available for extraction.
[0105]
[0108] When joining a live stream, typically both low latency and fast startup should be achieved. However, this is not self-evident, and several strategies are described below based on Figure 6. In the first example, at the live edge, segments that are three segments behind in the time history (i.e., segments 250B, 250C, and 250D) are loaded into the buffer. When one segment becomes available, playback begins. This introduces considerable latency, but playback can start relatively quickly because random access is loaded at the beginning of the segment. In the second example, segment 250D, the most recent available segment, is selected instead of the three older segments. The regeneration latency in this example is at least the segment duration, but may be longer. The startup may be similar to the example above. โข In the other three cases, segments containing multiple chunks (e.g., segment 250E) are replayed while they are still being generated. This reduces latency, but there is a problem in that the start of replay can be affected, especially if the difference between the segment availability start time and the wall clock time of the most recently published segment is greater than the target latency. In this case, the client device may have to wait until the next second is issued. In the case of the 6-second segment, this can result in a start latency of 4-5 seconds. Other techniques and use cases exist. For example, a client can access older segments at the start, download everything, accelerate playback, and perform fast-forward decoding. However, such methods have the drawback that significant data must be downloaded before accelerated decoding can occur. Furthermore, it is not widely supported in decoder interfaces.
[0106]
[0109] The appropriate solutions may be the following: โข At least one representation of the adaptation set includes more frequent random access points and non-initial chunks within the segment / fragment. โข A DASH client may use information from MPD to determine that such a random access method exists, but the location / byte offset of the random access point may not be accurately signaled. The DASH client can access this representation at startup, but will only download it starting from the most recent available non-initial chunk's byte range, or at least close to it. โข Once downloaded, the DASH client can determine a random access point and begin processing data along with the same downloaded initialization segment / CMAF header. The location of the random access point is described below.
[0107]
[0110] However, the latter approach can encounter various problems, as summarized below.
[0108]
[0111] Therefore, as illustrated in the example in Figure 6 and described below, the use of chunks described in this disclosure can substantially reduce latency. Pre-signaling the start of chunks may allow for the pre-generation of a manifest file that does not require frequent updates but can still indicate the approximate location of stream access points (SAPs) within the chunk. In this way, a client device can determine the location of chunk boundaries using the manifest file without requiring continuous manifest file updates, while still allowing the client device to start media streaming at the beginning of a chunk boundary, for example, at a resynchronization point. That is, a client device can determine the byte range of a segment containing a resynchronization point from the manifest file even before the segment is fully formed, because the manifest file can signal the byte range or other data representing the approximate location of the resynchronization point within the segment.
[0109]
[0112] Figure 7 is a conceptual diagram illustrating a second exemplary specification example using DASH and CMAF random access in the context of a broadcast protocol. Figure 7 shows an example including a media encoder 280, a CMAF / File Format (FF) packager 282, a DASH packager 284, a ROUTE sender 286, a CDN origin server 288, a ROUTE receiver 290, a DASH client 292, a CMAF / FF parser 294, and a media decoder 296. The media encoder 280 encodes media data, such as audio or video data. The media encoder 280 may correspond to the audio encoder 26 or video encoder 28 in Figure 1, or the encoder 216 in Figure 5. The media encoder 280 provides the encoded media data to the CMAF / FF packager 282, which formats the encoded media data in a file according to CMAF and a specific file format such as ISO BMFF or its extensions.
[0110]
[0113] The CMAF / FF packager 282 provides these files (e.g., chunks) to the DASH packager 284, which aggregates the files / chunks into a DASH segment. The DASH packager 284 may also form a manifest file, such as an MPD, containing data describing the files / chunks / segments. Furthermore, according to the techniques of this disclosure, the DASH packager 284 may determine the approximate location of a future Stream Access Point (SAP) or Random Access Point (RAP) and signal the approximate location in the MPD. The CMAF / FF packager 282 and the DASH packager 284 may correspond to the encapsulation unit 30 in Figure 1 or the DASH packager 202 in Figure 5.
[0111]
[0114] The DASH packager 284, along with MPD, provides the segment to the ROUTE sender 286 and the CDN origin server 288. The ROUTE sender 286 and the CDN origin server 288 may correspond to the server device 60 in Figure 1 or the CDN 220 in Figure 5. Generally, the ROUTE sender 286 can send media data to the ROUTE receiver 290 according to the ROUTE in this example. In other examples, other file-based delivery protocols, such as FLUTE, may be used for broadcast or multicast. As an addition or alternative, the CDN origin server 288 can send media data to the ROUTE receiver 290 and / or directly to the DASH client 292, for example, according to HTTP.
[0112]
[0115] The ROUTE receiver 290 may be implemented in middleware such as the eMBMS middleware unit 100 in Figure 2. The ROUTE receiver 290 can buffer the received media data, for example, in the cache 104 shown in Figure 2. The DASH client 292 (which may correspond to the DASH client 110 in Figure 2) can retrieve the cached media data from the ROUTE receiver 290 using HTTP. Alternatively, the DASH client 292 can retrieve the media data directly from the CDN origin server 288 via HTTP, as described above.
[0113]
[0116] Furthermore, according to the techniques of this disclosure, the DASH client 292 may use a manifest file, such as an MPD, to determine the location of SAP or RAP following a resynchronization point signaled in the manifest file. The DASH client 292 may begin retrieving a media presentation starting from the next earliest resynchronization point. A resynchronization point generally indicates a location in the bitstream where file container-level data can be correctly parsed. Thus, the DASH client 292 may begin streaming at the resynchronization point and deliver the received media data starting from the resynchronization point to the CMAF / FF parser 294.
[0114]
[0117] The CMAF / FF parser 294 can begin analyzing media data starting from a resynchronization point. The CMAF / FF parser 294 may correspond to the deencapsulation unit 50 in Figure 1. Furthermore, the CMAF / FF parser 294 can extract decryptable media data from the analyzed data and distribute the decryptable media data to a media decoder 296, which may correspond to the audio decoder 46 or video decoder 48 in Figure 1. The media decoder 296 can decrypt the media data and distribute the decrypted media data to a corresponding output device, such as the audio output 42 or video output 44 in Figure 1.
[0115]
[0118] In the case of broadcasting, an example of the combination of DASH / CMAF and ROUTE is shown in Figure 7. The combination of low-latency DASH mode and ROUTE (for example, as considered for the DVB TM-IPI task force and ATSC profile in ABR multicast) can lead to the following problems: If ROUTE receiver 290 joins a DASH / CMAF low-latency segment midway through, it cannot begin processing data because synchronization is not available and there is no random access for other purposes. Therefore, even if more frequent random access is provided midway through the segment, the start will be delayed.
[0116]
[0119] The appropriate solutions may be the following: โข Broadcast / multicast representations include more frequent random access points and non-initial chunks within a segment / fragment. The DASH client 292 determines, using MPD information and / or possibly information from the ROUTE receiver 290, that such a random access method exists. The DASH client 292 may, but may not, use such information precisely to pinpoint the location of the random access point. โข DASH client 292 can access this representation at startup, but may not have access to all information from the beginning. When accessing the received portion of the segment begins, the DASH client 292 may find a random access point and begin processing the data along with the same downloaded initialization segment / CMAF header. The location of the random access point is described below.
[0117]
[0120] However, the latter approach can encounter various problems, as summarized below.
[0118]
[0121] In a scenario similar to the second use case described above, packet loss can be a problem not only during random access resynchronization. In this exemplary third use case, the same procedure as described above may be applied. In addition, after clean random access is attempted and sufficient box analysis is possible, events, decoding, and presentation in non-random access chunks (e.g., without IDR frames) may be attempted. Therefore, resynchronization to random access for file format analysis is important, as is resynchronization to clean random access.
[0119]
[0122] A fourth use case can typically occur when live media content is delivered with low latency, but then the same media content is used time-shifted for delayed playback. A client may want access to a media presentation at a specific time, which may not coincide with (and generally does not coincide with) the start of a segment / CMAF fragment.
[0120]
[0123] The appropriate solutions may be the following: โข At least one representation of the adaptation set may contain more frequent random access points and non-initial chunks within a segment / fragment. โข DASH client 292 may use information from MPD to determine that such a random access method exists, but the exact location / byte offset of the random access point may not be known. โข DASH client 292 can access this representation during seeking, but may only be allowed to download from the byte range of the most recent available non-initial chunk, or at least a byte range close to it. โข Once downloaded, the DASH client 292 may find a random access point and begin processing data along with the same downloaded initialization segment / CMAF header. The location of the random access point is described below.
[0121]
[0124] However, the latter approach can encounter various problems, as summarized below.
[0122]
[0125] Resynchronization of ISO BMFF / DASH / CMAF segments generally involves several processes, summarized below. 1) Find the box structure. 2) Find the CMAF chunk / fragment that contains all the relevant information. 3) Find the timing via mdat and tfdt. 4) Obtain all decryption-related information, where applicable. 5) Process event messages as needed. 6) Initiate decryption at the elementary stream level.
[0123]
[0126] An exemplary method for finding resynchronization points in a box structure at a specific time is summarized below. โข If a segment index (SIDX box) exists, such resynchronization points are provided as presentation time and byte offsets. However, since segments are not fully formed beforehand, segment indexes are typically not available for low-latency live streaming. If the start of a segment is available, the client may download a minimum set of bytes so that the box structure can be processed. Resynchronization is provided by the underlying protocol, for example, by signaling chunk boundaries so that the client can begin analysis. If the start of a segment cannot be easily determined through signaled data, the client can find a synchronization pattern that allows the client to randomly access the data. The client can then initiate analysis and find a suitable box structure that enables processing such as emsg, prft, mdat, moof, and / or mdat.
[0124]
[0127] This disclosure describes techniques that may be applied to the fourth example described above. The first three represent exemplary simplifications where the corresponding information is available.
[0125]
[0128] This disclosure acknowledges the following issues based on the above description and recognizes that these issues require solutions. 1) Add additional random access points within the DASH / CMAF segment. Random access may include clean random access and open or incremental decoder refreshes, continuing until only resynchronization is given during file format analysis. 2) Add appropriate signaling to the MPD (or other manifest file) that indicates the availability of random access points and resynchronization in each DASH segment, and provides information about the location, type, and timing of random access points. The information may be accurate or within a certain range. 3) The ability to resynchronize to decapsulation, decoding, and decryption by finding a resynchronization point, given any starting point. 4) The ability to initiate processing in a restricted receiver environment, for example, as available in HTML-5 / MSE-based playback.
[0126]
[0129] Figure 8 is a conceptual diagram illustrating exemplary signaling of Stream Access Points (SAPs) in a manifest file. Specifically, Figure 8 shows bitstream 300 containing SAP302A-302D (SAP302) and segments 304A-304D (segment 304), and bitstream 310 containing SAP312A-312D (SAP312), SAP316A-316D (SAP316), and segments 314A-314D (segment 314). That is, in this example, segment 314 of bitstream 310 contains SAP312 and 316 more frequently than segment 304 of bitstream 300. Each of SAP302 and 312 may correspond to the start of both the corresponding one of segments 304 and 314, and to the first chunk of these segments. SAP316 may correspond to the start of a chunk within the corresponding segment 316, but not to the beginning of the corresponding segment 316.
[0127]
[0130] To provide a simple technique for achieving a constant bitrate representation using equidistant chunks of 1000 samples (and @timescale=1000 in sample duration) and SAP type 1 (which can be, for example, an audio representation), the following resynchronization elements may be added:
[0128]
number
[0129]
[0131] A client receiving such information, for example, client device 40 in Figure 1, may not be able to identify that random access points can be accessed per second within a precise byte range for a segment with @duration=10000. If the bitrate is variable, the receiver (e.g., client device 40) may use @dIMin and @dIMax to identify the range over which random access points should be found. As an alternative to @dT, which signals the maximum value, it may also signal the nominal chunk duration.
[0130]
[0132] The resynchronization element in the manifest file may also include a URL@index pointing to a binary resynchronization index for each resynchronization point within the segment, using the same template functions as regular segments. This resynchronization, if present, can provide the precise location of all resynchronization points within the segment, similar to the segment index. If this index exists, the resynchronization index may be available for all segments during the period available at the time of manifest file / MPD issuance.
[0131]
[0133] In one approach, the resynchronization index may be identical to the segment index, but it may also be modified.
[0132]
[0134] The client device 40 in Figure 1 may use an ISO BMFF 4-letter box type as the basis for resynchronizing to a media file (for example, a video file 150 in Figure 4, which may be a segment). In one example, the selected box type is the "styp" box, but it may also be the "moof" box itself. Random emulation of box column types is extremely rare. A test report of styp emulation is described below. This emulation is then avoided by checking against known expected box types. The client device 40 may perform a resynchronization mechanism outlined below. 1) For example, find the occurrence of the "styp" byte sequence in the segment at byte offset B1. 2) Verify against random emulation as follows: The following box types are compared to the expected list of box types, namely, "styp", "sidx", "ssix", "prft", "moof", "mdat", "free", "mfra", "skip", "meta", and "meco".
[0133] a. If one of the known box types is found, the byte offset B1-4 bytes is the byte offset of the resynchronization point.
[0134] b. If this is not one of the known box types mentioned above, this occurrence of the styp box is considered an invalid synchronization point and is ignored. Restart from step 1 above.
[0135]
[0135] The technique of the present disclosure was tested on 30,282 segments from scanned DASH-IF test assets. This scan revealed 28,408 occurrences of the column "styp" in the file, and only 10 of these 28,408 occurrences (about 1 out of 2,840 occurrences) were emulations that were discarded when it was determined that the next box was not one of the expected box types, i.e., "styp", "sidx", "ssix", "prft", "moof", "mdat", "free", "mfra", "skip", "meta", or "meco".
[0136]
[0136] Based on these results, it is considered sufficient to use styp resynchronization point detection along with chunk structure. It is appropriate to restrict it to only a subset of boxes that may follow styp, such as prft, emsg, free, skip, and moof.
[0137]
[0137] The remaining issues are determining the SAP type and the earliest presentation time. The latter can be easily achieved by using the tfdt and other information in the movie fragment header. It would be appropriate to document the algorithm.
[0138]
[0138] There are several options for determining the SAP type, as shown below: โข Detection based on information in moof. A simple technique can be documented and implemented. โข Use of compatible brands in SAP types. By already using CMAF, the following can be inferred:
[0139] โcmff: Indicates that SAP is 1 or 2 โcmfl: Indicates that SAP is 0 (Is this correct for decoding?) โcmfr: Indicates that SAP is 1, 2, or 3. This signaling may be sufficient if used consistently. Compatibility brands for other SAP types may be defined. Other techniques may be used to indicate the SAP type.
[0140]
[0139] Existing options may be used to determine the SAP type.
[0141]
[0140] Thus, the techniques of the present disclosure can be summarized as follows and can be performed by devices such as the content creation device 20, server device 60, and / or client device 40 in Figure 1, as described above.
[0142]
[0141] In the DASH context, in some cases, a segment is treated as a single unit for downloading and accessing media presentations, and is also treated by an addressed URL. However, segments can be structured to allow resynchronization at the container level and random access to each representation within the segment. The resynchronization mechanism is supported and signaled by the resynchronization element.
[0143]
[0142] The resynchronization element signals resynchronization points within a segment. A resynchronization point is the start of a chunk (at a byte position), and a chunk is defined as a structured, contiguous range of bytes within a segment containing media data for a particular presentation duration, which can be accessed independently on a container format, including the possibility of decryption. A resynchronization point within a segment may be defined as follows: The resynchronization point is the start of a chunk. Furthermore, the resynchronization point is assigned the following properties:
[0144] โIt has a byte offset or index value from the start of the segment that points to the first byte of the chunk.
[0145] โIt has the earliest presentation time allocated in the presentation.
[0146] โIt has an assigned SAP type, as defined, for example, by the SAP type in ISO / IEC 14496-12.
[0147] โIt assigned a marker property that indicates whether a resynchronization point can be detected while analyzing a segment through a specific marker, or whether the resynchronization point needs to be signaled by an external means. Starting processing from the resynchronization point allows for container parsing and decryption, if present, along with information in the initialization segment. The ability to access the included elementary video stream, and how to access it, is defined by the SAP type.
[0148]
[0143] Signaling each resynchronization point in MPD can be difficult for causal reasons, as resynchronization points may be added by the segment packager independently of MPD updates. For example, resynchronization points may be generated by the encoder and packager independently of MPD. Also, at low latency, MPD signaling may not be available to DASH clients, for example, DASH client 110 in Figure 2 or DASH client 292 in Figure 7. Therefore, there are two methods for signaling resynchronization points provided in a segment in MPD. โข By providing a binary map of resynchronization points within the resynchronization index segment for each segment. This is most easily used for segments that are fully available on the network. This involves signaling the presence of resynchronization points within a segment, as well as providing additional information that facilitates the identification of these points in terms of byte position and presentation time.
[0149]
[0144] In order to signal the above characteristics, the resynchronization element has different attributes which are described in more detail in Section 5.3.12.2 of the DASH specification.
[0150]
[0145] Random access, if present, refers to initiating the processing, decoding, and presentation of a representation from a random access point after time t by initializing the representation using an initialization segment and decoding and presenting the representation from the signaled segment onward. Random access points can be signaled using the RandomAccess element, as defined in Table 10 below.
[0151] [Table 2]
[0152]
[0146] Table 11 provides various types of random access points.
[0153] [Table 3]
[0154]
[0147] The resynchronization index segment contains information related to the media segment. Like the segment index, the resynchronization index segment provides the exact location of all resynchronization points within the segment. Resynchronization points are defined in section 5.3.12.1 of the DASH specification.
[0155]
[0148] The resynchronization point of an ISO BMFF can be defined as the start of an ISO BMFF segment, subject to the following restrictions with respect to both cardinality and order:
[0156] [Table 4]
[0157]
[0149] For ISO BMFF-based resynchronization points, the properties can be defined as follows: The index is defined as the offset of the first byte of the constrained ISO BMFF segment described above. The earliest presentation time, Time, is defined as the minimum time for the combination of the decoding time of any sample in the chunk, the configuration offset, and the edit list. โข SAP types are defined according to Section 4.5.2 of the DASH specification. If styp exists alongside "cmfl" as a primary compatible brand, then the marker exists.
[0158]
[0150] A resynchronization index segment can index one media segment of one representation and can be defined as follows: Each expression index segment should begin with the "styp" box, and the brand "risg" should reside within the "styp" box. The conformance requirements for the brand "risg" are defined by this sub-clause. Each media segment is indexed by one or more segment index boxes, and the boxes for a given media segment are contiguous.
[0159]
[0151] Figure 9 is a flowchart illustrating an exemplary method for retrieving media data using the technique of the present disclosure. The method of Figure 9 is described with respect to the client device 40 of Figure 1. However, a low-latency DASH client 232 of Figure 5, or a client device including a media decoder 296, a CMAF / FF parser 294, a DASH client 292, and a ROUTE receiver 290 of Figure 7, may also be configured to perform this method or a similar method.
[0160]
[0152] First, the client device 40 can retrieve a manifest file, such as an MPD, of the media presentation (350). The client device 40 can retrieve a manifest file from, for example, the server device 60. The manifest file may contain data indicating that the media presentation includes resynchronization points at chunk boundaries within segments of the representation of the media presentation. Thus, the client device 40 can determine a resynchronization point of the media presentation, for example, the most recently available resynchronization point (352). Generally, a resynchronization point can indicate the start of a chunk boundary and is a randomly accessible point in the representation that can be properly parsed by a file-level container (for example, a data structure such as a box, as described above).
[0161]
[0153] In particular, the manifest file can indicate the location of a resynchronization point, such as a byte offset from the start of the segment. This information cannot pinpoint the exact location of the resynchronization point within the segment, but it can ensure that the resynchronization point is available within a byte range from the byte offset. Thus, the client device 40 can form a request, such as an HTTP partial GET request, specifying the indicated byte offset to begin retrieving at the resynchronization point (354). The client device 40 can then send the request to the server device 60 (356).
[0162]
[0154] In response to the request, the client device 40 may receive the requested media data including the resynchronization point (358). As described above, the byte offset may not accurately identify the location of the resynchronization point, and therefore the client device 40 may parse the data until it detects the actual location of the resynchronization point. The client device 40 may parse file-level data structures, such as file format boxes, starting at the resynchronization point, in order to determine the location of the media data in the corresponding chunk of the retrieved media data. In detail, the client device 40 may identify the resynchronization point as the start of a chunk by, for example, detecting a segment type value, a generator reference time value, an event message, a movie fragment, and a media data container box. The movie fragment may contain encoded media data.
[0163]
[0155] The deencapsulation unit 50 can, for example, extract encoded media data for corresponding chunks from a movie fragment (360) and provide the encoded media data to, for example, the video decoder 48. A chunk may begin with a random access point (RAP), such as an intra-prediction frame (I-frame) of video data. The manifest file may further indicate whether the RAP is the beginning of a closed picture group (GOP) or an open GOP, thereby indicating the type of random access that may be initiated and performed at the RAP (for example, whether the leading picture of the I-frame is decryptable or not). The video decoder 48 can then decode the encoded media data (362) and send the media data to, for example, the video output 44 to present the decoded media data (364).
[0164]
[0156] Thus, the method in Figure 9 represents an example of a method for retrieving media data, which includes: retrieving a media presentation manifest file indicating that container analysis of the media data of the bitstream may be initiated at a resynchronization point of a segment of the media presentation's representation; using the manifest file, which indicates that the resynchronization point is located somewhere other than the start of a segment and represents a point where container analysis of the media data of the bitstream may be initiated, to form a request to retrieve media data of a representation that starts at a resynchronization point; sending a request to initiate the retrieval of media data of a media presentation that starts at a resynchronization point; and presenting the retrieved media data.
[0165]
[0157] Some of the techniques of this disclosure are summarized in the following examples.
[0166]
[0158] Example 1: A method for extracting media data, comprising: extracting a manifest file of a media presentation indicating that resynchronization and decoding can be initiated at a resynchronization point of the representation of the media presentation; extracting media data of the representation to be initiated at the resynchronization point; and presenting the extracted media data.
[0167]
[0159] Example 2: The method of Example 1, wherein the resynchronization point includes the start of a chunk boundary.
[0168]
[0160] Example 3: The method of Example 2, wherein the chunk boundary comprises the start of a chunk having zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
[0169]
[0161] Example 4: The manifest file indicates the availability of resynchronization points in the representation segment, using one of the methods from Examples 1 to 3.
[0170]
[0162] Example 5: The method of Example 4, where the resynchronization point is located at a position other than the start of the segment.
[0171]
[0163] Example 6: The manifest file indicates the type of random access that may be performed at the resynchronization point, using one of the methods in Examples 4 and 5.
[0172]
[0164] Example 7: The manifest file indicates the location and timing of the resynchronization point, and whether the location and timing information is accurate or estimated, using one of the methods in Examples 4 to 6.
[0173]
[0165] Example 8: The manifest file comprises a Media Presentation Description (MPD) using any of the methods in Examples 1 to 7.
[0174]
[0166] Example 9: A device for extracting media data, comprising one or more means for performing any of the methods in Examples 1 to 8.
[0175]
[0167] Example 10: The device of Example 9, wherein one or more means comprises one or more processors implemented in the circuit and a memory configured to store media data.
[0176]
[0168] Example 11: The device of Example 9, comprising at least one of an integrated circuit, a microprocessor, or a wireless communication device.
[0177]
[0169] Example 12: A computer-readable storage medium that stores instructions that, when executed, cause the processor to perform one of the methods in Examples 1 to 8.
[0178]
[0170] Example 13: A device for retrieving media data, comprising: means for retrieving a manifest file of a media presentation indicating that resynchronization and decoding can be initiated at a resynchronization point of the representation of the media presentation; means for retrieving media data of the representation to be initiated at the resynchronization point; and means for presenting the retrieved media data.
[0179]
[0171] Example 14: A method for sending media data, comprising: sending a manifest file of a media presentation to a client device indicating that resynchronization and decoding can be initiated at a resynchronization point of the representation of the media presentation; receiving a request from the client device for media data to be started at a resynchronization point; and, in response to the request, sending the requested media data of the representation to be started at a resynchronization point to the client device.
[0180]
[0172] Example 15: The method of Example 14, further comprising generating a manifest file.
[0181]
[0173] Example 16: The resynchronization point includes the start of a chunk boundary, in any way of Examples 14 and 15.
[0182]
[0174] Example 17: The method of Example 16, wherein the chunk boundary comprises the start of a chunk having zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
[0183]
[0175] Example 18: The manifest file indicates the availability of resynchronization points in the segment of the representation, in any of the methods from Examples 14 to 17.
[0184]
[0176] Example 19: The resynchronization point is located at a position other than the start of the segment, as in Example 18.
[0185]
[0177] Example 20: The manifest file indicates the type of random access that may be performed at the resynchronization point, in one of the methods of Examples 18 and 19.
[0186]
[0178] Example 21: The manifest file indicates the location and timing of the resynchronization point, and whether the location and timing information is accurate or estimated, in any of the methods shown in Examples 18 to 20.
[0187]
[0179] Example 22: The manifest file comprises a Media Presentation Description (MPD) in any of the manner described in Examples 14 through 21.
[0188]
[0180] Example 23: A device for transmitting media data, comprising one or more means for performing any of the methods in Examples 14 to 22.
[0189]
[0181] Example 24: The device of Example 23, wherein one or more means comprises one or more processors implemented in the circuit and a memory configured to store media data.
[0190]
[0182] Example 25: The device of Example 23, comprising at least one of an integrated circuit, a microprocessor, or a wireless communication device.
[0191]
[0183] Example 26: A computer-readable storage medium that stores instructions that, when executed, cause the processor to execute one of the methods from Example 1 to 8.
[0192]
[0184] Example 27: A device for sending media data, comprising: means for sending a media presentation manifest file to a client device indicating that resynchronization and decoding can be initiated at a resynchronization point of the representation of the media presentation; means for receiving a request from the client device for media data to be initiated at a resynchronization point; and means for sending the requested media data of the representation to be initiated at a resynchronization point to the client device in response to the request.
[0193]
[0185] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via computer-readable media as one or more instructions or codes and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media corresponding to tangible media such as data storage media, or communication media including any media that facilitates the transfer of computer programs from one place to another according to a communication protocol, for example. Thus, computer-readable media may generally correspond to (1) non-transient tangible computer-readable storage media, or (2) communication media such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes and / or data structures for implementing the techniques described herein. Computer program products may include computer-readable media.
[0194]
[0186] As an example, and not an limitation, such computer-readable storage media may include RAM, ROM, EEPROMยฎ, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. Furthermore, any connection is appropriately called a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. However, it should be understood that computer-readable storage media and data storage media refer to non-temporary tangible storage media instead of connections, carriers, signals, or other temporary media. As used herein, the terms "disk" and "disc" include Compact Disc (CD), LaserDiscยฎ (disc), Optical Disc (disc), Digital Multipurpose Disc (disc) (DVD), Floppy Disk (disk), and Blu-rayยฎ Disc (disc), where a Disk typically reproduces data magnetically, and a Disc (disc) reproduces data optically using a laser. Any combination of the above should also be included within the scope of computer-readable media.
[0195]
[0187] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated circuits or discrete logic circuits. Thus, the term โprocessorโ as used herein may refer to any of the above-described structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some embodiments, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a composite codec. The techniques may also be fully implemented by one or more circuits or logic elements.
[0196]
[0188] The techniques of the Disclosure may be implemented in a wide variety of devices or apparatus, including wireless handsets, integrated circuits (ICs) or sets of ICs (e.g., chipsets). While the Disclosure has described various components, modules, or units to highlight the functional aspects of devices configured to perform the techniques disclosed, these components, modules, or units do not necessarily have to be implemented by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit, including one or more processors as described above, along with suitable software and / or firmware, or may be provided by a set of interoperable hardware units.
[0197]
[0189] Various examples have been described. These and other examples fall within the scope of the following claims. The invention described in the original claims of this application is listed below. [C1] A method for extracting media data, Retrieving the media presentation manifest file which indicates that container analysis of the media data of the bitstream may be initiated at a resynchronization point of a segment of the media presentation's representation, and the resynchronization point is located at a position other than the start of the segment and represents a point where the container analysis of the media data of the bitstream may be initiated. Using the manifest file, a request is formed to retrieve the media data of the representation that starts at the resynchronization point. Sending the request to start retrieving the media data of the media presentation that starts at the aforementioned resynchronization point, To present the extracted media data A method that includes [a certain feature]. [C2] The method according to C1, wherein presenting the extracted media data comprises analyzing the file-level media data container of the extracted media data at the resynchronization point. [C3] To analyze, The process involves analyzing the file-level media data container until a random access point (RAP) of the media presentation is detected. Sending the aforementioned RAP to the media decoder A method of C2 comprising: [C4] The method according to C1, wherein the resynchronization point includes the start of a chunk boundary. [C5] The method of C2, wherein the chunk boundary comprises the start of a chunk having zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box. [C6] The method of C1, wherein the manifest file indicates the availability of the resynchronization point in the segment of the representation. [C7] The method according to C6, wherein the resynchronization point is located at a position other than the start of the segment. [C8] The method of C6, wherein the manifest file indicates the type of random access that may be performed at the resynchronization point. [C9] The method of C6, wherein the manifest file indicates the location and timing of the resynchronization point and whether the location and timing information is accurate or estimated. [C10] The method according to C1, wherein the manifest file comprises a media presentation description (MPD). [C11] A device for extracting media data, A memory configured to store media data of a media presentation; retrieving a manifest file of the media presentation indicating that container analysis of the media data of the bitstream can be initiated at a resynchronization point of a segment of the representation of the media presentation; and the resynchronization point being located at a position other than the beginning of the segment, and representing a point where the container analysis of the media data of the bitstream can be initiated. Using the manifest file to form a request to retrieve the media data of the representation that starts at the aforementioned resynchronization point, Sending the request to start retrieving the media data of the media presentation that starts at the aforementioned resynchronization point, To present the extracted media data One or more processors implemented within the circuit are configured to perform the following: A device equipped with the following features. [C12] The device according to C11, wherein one or more processors are configured to analyze the file-level media data container of the retrieved media data at the resynchronization point in order to present the retrieved media data. [C13] In order to analyze the file-level media data container, one or more processors The process involves analyzing the file-level media data container until a random access point (RAP) of the media presentation is detected. Sending the aforementioned RAP to the media decoder A device as described in C12, configured to perform the following actions. [C14] The device according to C11, wherein the resynchronization point includes the start of a chunk boundary. [C15] The device according to C14, wherein the chunk boundary comprises the start of a chunk having zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box. [C16] The manifest file is the device described in C11, which indicates the availability of the resynchronization point in the segment of the representation. [C17] The device according to C16, wherein the resynchronization point is located at a position other than the start of the segment. [C18] The manifest file is the device described in C16, which indicates the type of random access that may be performed at the resynchronization point. [C19] The manifest file is the device described in C16, which indicates the location and timing of the resynchronization point and whether the location and timing information is accurate or estimated. [C20] The manifest file comprises a media presentation description (MPD) and is the device described in C11. [C21] A computer-readable storage medium storing instructions, wherein, when an instruction is executed, the processor, Retrieving the media presentation manifest file which indicates that container analysis of the media data of the bitstream may be initiated at a resynchronization point of a segment of the media presentation's representation, and the resynchronization point is located at a position other than the start of the segment and represents a point where the container analysis of the media data of the bitstream may be initiated. Using the manifest file to form a request to retrieve the media data of the representation that starts at the aforementioned resynchronization point, Sending the request to start retrieving the media data of the media presentation that starts at the aforementioned resynchronization point, To present the extracted media data A computer-readable storage medium that enables the following process. [C22] The computer-readable storage medium according to C21, wherein the resynchronization point includes the start of a chunk boundary. [C23] The computer-readable storage medium according to C22, wherein the chunk boundary comprises the start of a chunk comprising zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box. [C24] The manifest file is a computer-readable storage medium as described in C21, indicating the availability of the resynchronization point in the segment of the representation. [C25] The computer-readable storage medium according to C24, wherein the resynchronization point is located at a position other than the start of the segment. [C26] The manifest file is a computer-readable storage medium as described in C24, indicating the type of random access that may be performed at the resynchronization point. [C27] The manifest file is a computer-readable storage medium as described in C24, indicating the location and timing of the resynchronization point and whether the location and timing information is accurate or estimated. [C28] The manifest file comprises a media presentation description (MPD) on a computer-readable storage medium as described in C21. [C29] A device for extracting media data, Means for retrieving a media presentation manifest file indicating that container analysis of the media data of a bitstream may be initiated at a resynchronization point of a segment of the media presentation's representation, and that the resynchronization point is located at a position other than the start of the segment and represents a point where the container analysis of the media data of the bitstream may be initiated. Means for using the manifest file to form a request to retrieve the media data of the representation to be started at the resynchronization point, means for sending the request to start retrieving the media data of the media presentation which starts at the resynchronization point, Means for presenting the extracted media data A device equipped with the following features.
Claims
1. A method for extracting media data, Retrieving the media presentation manifest file which indicates the location of a resynchronization point within a segment of the media presentation's representation, and further indicating that the manifest file indicates that container analysis of the bitstream's media data may be initiated at the resynchronization point, that the resynchronization point is located elsewhere than the beginning of the segment and represents a point where container analysis of the bitstream's media data may be initiated, and that a random access point (RAP) follows the resynchronization point. Using the manifest file, a request is formed to retrieve the media data of the representation that starts at the resynchronization point. Sending the request to start retrieving the media data of the media presentation that starts at the aforementioned resynchronization point, Analyzing the extracted media data that starts at the aforementioned resynchronization point. A method that includes [a certain feature].
2. The method according to claim 1, wherein analyzing the extracted media data comprises analyzing the file-level media data container of the extracted media data at the resynchronization point.
3. To analyze, The process involves analyzing the file-level media data container until the RAP of the media presentation is detected, Sending the aforementioned RAP to the media decoder The method according to claim 2, comprising:
4. The method according to claim 1, wherein the resynchronization point includes the start of a chunk boundary.
5. The method according to claim 4, wherein the chunk boundary comprises the start of a chunk having zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
6. A device for extracting media data, A memory configured to store media data for a media presentation, Retrieving the media presentation manifest file which indicates the location of a resynchronization point within a segment of the media presentation's representation, and further indicating that the manifest file indicates that container analysis of the bitstream's media data may be initiated at the resynchronization point, that the resynchronization point is located elsewhere than the beginning of the segment and represents a point where container analysis of the bitstream's media data may be initiated, and that a random access point (RAP) follows the resynchronization point. Using the manifest file to form a request to retrieve the media data of the representation that starts at the aforementioned resynchronization point, Sending the request to start retrieving the media data of the media presentation that starts at the aforementioned resynchronization point, Analyzing the extracted media data that starts at the aforementioned resynchronization point. One or more processors implemented within the circuit are configured to perform the following: A device equipped with the following features.
7. In order to analyze the extracted media data, one or more processors are configured to analyze the file-level media data container of the extracted media data at the resynchronization point. In order to analyze the file-level media data container, one or more processors The process involves analyzing the file-level media data container until the RAP of the media presentation is detected, Sending the aforementioned RAP to the media decoder The device according to claim 6, configured to perform the following:
8. The device according to claim 6, wherein the resynchronization point includes the start of a chunk boundary.
9. The device according to claim 8, wherein the chunk boundary comprises the start of a chunk comprising zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.
10. A computer-readable storage medium storing instructions, wherein, when the instructions are executed, the computer-readable storage medium causes a processor to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
JP2015111898A
JP2015208018A
JP2019523600A
JP2019525677A