Random access at resync point of dash segment

The implementation of resynchronization points in video streaming technologies enables random access within segments, enhancing flexibility and efficiency in video data retrieval and adaptation to network conditions.

JP2025111451AActive Publication Date: 2025-07-30QUALCOMM INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025051693
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-10-01
Filing Date
2025-03-26
Publication Date
2025-07-30
Estimated Expiration
2040-10-02

AI Technical Summary

Technical Problem

Existing video streaming technologies, such as DASH and CMAF, lack the ability to perform random access within a segment other than at the segment start, limiting flexibility and efficiency in video data retrieval.

Method used

Implementing techniques to allow random access at resynchronization points within a segment by signaling resynchronization points in the manifest file, enabling container parsing to start at positions other than the segment start, and facilitating media data retrieval and presentation from these points.

Benefits of technology

Enhances video streaming flexibility and efficiency by allowing random access within segments, improving adaptability and responsiveness to network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111451000001_ABST
    Figure 2025111451000001_ABST
Patent Text Reader

Abstract

To provide a method of retrieving media data, a device, and a storage medium.SOLUTION: A method of retrieving media includes: retrieving a manifest file for a media presentation indicating that container parsing of media data of a bitstream can be started at a resync point of a segment of a representation of the media presentation, the resync point being at a position other than a start of the segment and representing a point at which the container parsing of the media data of the bitstream can be started; using the manifest file to form a request to retrieve the media data of the representation starting at the resync point; sending the request to initiate retrieval of the media data of the media presentation starting at the resync point; and presenting the retrieved media data.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] This application claims the benefit of U.S. Application No. 17 / 061,152, filed Oct. 1, 2020, and U.S. Provisional Application No. 62 / 909,642, filed Oct. 2, 2019, the entire contents of which are incorporated herein by reference.

[0002]

[0002] This disclosure relates to the storage and transport of encoded video data.

Background Art

[0003]

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite telephones, video teleconferencing devices, and the like. Digital video devices implement video compression techniques, such as those defined by standards such as MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)), and extensions to such standards, to transmit and receive digital video information more efficiently.

[0004]

[0004] Video compression techniques perform spatial prediction and / or temporal prediction to reduce or remove redundancy inherent in a video sequence. In the case of block-based video coding, a video frame or slice can be partitioned into macroblocks. Each macroblock can be further partitioned. Macroblocks in an intra-coded (I) frame or slice are coded using spatial prediction with respect to adjacent macroblocks. Macroblocks in an inter-coded (P or B) frame or slice can use spatial prediction with respect to adjacent macroblocks in the same frame or slice, or temporal prediction with respect to other reference frames.

[0005]

[0005] After the video data is encoded, the video data can be packetized for transmission or storage. The video data can be assembled into a video file conforming to any of various standards, such as an International Organization for Standardization (ISO)-based media file format like AVC and its extensions.

Summary of the Invention

[0006]

[0006] Generally, the present disclosure describes techniques for accessing segment data not only at the start of a segment but also at other locations within the segment (e.g., for random access) in Dynamic Adaptive Streaming over HTTP (DASH) and / or Common Media Access Format (CMAF). The present disclosure also describes techniques related to signaling the ability to perform random access within a segment. The present disclosure describes various use cases related to these techniques. For example, the present disclosure defines resync points and the signaling of resync points in DASH and ISO-based media file format (BMFF).

[0007]

[0007] In one example, a method for retrieving media data includes retrieving a manifest file of a media presentation indicating that container parsing of the media data in a bitstream can be started at a resynchronization point of a segment of the media presentation's representation, where the resynchronization point is at a position other than the start of the segment and represents a point at which container parsing of the media data in the bitstream can be started, forming a request to retrieve the media data of the representation starting at the resynchronization point using the manifest file, sending a request to start retrieving the media data of the media presentation starting at the resynchronization point, and presenting the retrieved media data.

[0008]

[0008] In another example, a device for retrieving media data includes a memory configured to store the media data of a media presentation, and one or more processors implemented in a circuit configured to retrieve a manifest file of a media presentation indicating that container parsing of the media data in a bitstream can be started at a resynchronization point of a segment of the media presentation's representation, where the resynchronization point is at a position other than the start of the segment and represents a point at which container parsing of the media data in the bitstream can be started, use the manifest file to form a request to retrieve the media data of the representation starting at the resynchronization point, send a request to start retrieving the media data of the media presentation starting at the resynchronization point, and present the retrieved media data.

[0009]

[0009] In another example, a computer-readable storage medium stores instructions thereon that, when executed, cause a processor to retrieve a media presentation manifest file indicating that container parsing of media data of a bitstream can be started at a resynchronization point of a segment of the media presentation representation, where the resynchronization point is at a position other than the start of the segment and represents a point at which container parsing of the media data of the bitstream can be started, use the manifest file to form a request to retrieve media data of the representation starting at the resynchronization point, send a request to start retrieving the media data of the media presentation starting at the resynchronization point, and present the retrieved media data.

[0010]

[0010] In another example, a device for retrieving media data includes means for retrieving a media presentation manifest file indicating that container parsing of media data of a bitstream can be started at a resynchronization point of a segment of the media presentation representation, where the resynchronization point is at a position other than the start of the segment and represents a point at which container parsing of the media data of the bitstream can be started, means for using the manifest file to form a request to retrieve media data of the representation starting at the resynchronization point, means for sending a request to start retrieving the media data of the media presentation starting at the resynchronization point, and means for presenting the retrieved media data.

[0011]

[0011] Details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.

Brief Description of the Drawings

[0012]

Figure 1

[0012] A block diagram illustrating an exemplary system that implements techniques for streaming media data over a network.

Figure 2

[0013] Block diagram showing an exemplary set of components of an extraction unit.

Figure 3

[0014] Conceptual diagram showing elements of exemplary multimedia content.

Figure 4

[0015] Block diagram showing elements of an exemplary video file that can correspond to segments of a representation.

Figure 5

[0016] Conceptual diagram showing an exemplary low-latency architecture that can be used in a first use case according to the present disclosure.

Figure 6

[0017] Conceptual diagram showing in more detail an example of the use case described with respect to FIG. 5.

Figure 7

[0018] Conceptual diagram showing an exemplary second specification case using DASH and CMAF random access in the context of a broadcast protocol.

Figure 8

[0019] Conceptual diagram showing exemplary signaling of stream access points (SAPs) in a manifest file.

Figure 9

[0020] Flowchart showing an exemplary method of extracting media data according to the techniques of the present disclosure.

DETAILED DESCRIPTION OF THE INVENTION

[0013]

[0021] The techniques of the present disclosure can be applied to video files compliant with video data encapsulated according to any of the ISO base media file format, scalable video coding (SVC) file format, advanced video coding (AVC) file format, 3rd Generation Partnership Project (3GPP (registered trademark)) file format, and / or multi-view video coding (MVC) file format, or other similar video file formats.

[0014]

[0022] In HTTP streaming, frequently used operations include HEAD, GET, and partial GET. The HEAD operation retrieves the header of a file associated with a given Uniform Resource Locator (URL) or Uniform Resource Name (URN) without retrieving the payload associated with the URL or URN. The GET operation retrieves the entire file associated with a given URL or URN. The partial GET operation receives a byte range as an input parameter and retrieves several consecutive bytes of the file, where the number of bytes corresponds to the received byte range. Thus, since the partial GET operation can obtain one or more individual movie fragments, movie fragments for HTTP streaming can be provided. In a movie fragment, there may be several track fragments of different tracks. In HTTP streaming, a media presentation can be a structured set of data accessible to a client. A client can request and download media data information to present a streaming service to a user.

[0015]

[0023] In an example of streaming 3GPP data using HTTP streaming, there may be multiple representations for the video and / or audio data of multimedia content. As described below, different representations may correspond to different coding characteristics (e.g., different profiles or levels of a video coding standard), different coding standards or extensions of coding standards (such as multi-view and / or scalable extensions), or different bitrates. The manifest of such representations may be defined in a Media Presentation Description (MPD) data structure. A media presentation may correspond to a structured set of data accessible by an HTTP streaming client device. The HTTP streaming client device may request and download media data information to present the streaming service to the user of the client device. The media presentation may be described in an MPD data structure that may include updates to the MPD.

[0016]

[0024] A media presentation may include a sequence of one or more periods. Each period may continue until the start of the next period or, in the case of the last period, until the end of the media presentation. Each period may include one or more representations of the same media content. A representation may be one of several alternative encoded versions of audio, video, timed text, or other such data. Representations may differ by coding type, e.g., for video data, by bitrate, resolution, and / or codec, and for audio data, by bitrate, language, and / or codec. The term representation may be used to refer to a section of encoded audio or video data corresponding to a particular period of multimedia content and encoded in a particular way.

[0017]

[0025] The representation of a particular period can be assigned to a group indicated by an attribute in the MPD that indicates the adaptation set to which the representation belongs. Representations within the same adaptation set are generally considered to be alternatives to each other in that a client device can dynamically and seamlessly switch between these representations, for example, to perform bandwidth adaptation. For example, each representation of video data for a particular period can be assigned to the same adaptation set such that any of the representations can be selected to decode to present media data such as video data or audio data of the multimedia content for the corresponding period. The media content within one period can, in some examples, if present, be represented by either one representation from group 0 or a combination of at most one representation from each non-zero group. The timing data for each representation of the period can be expressed relative to the start time of the period.

[0018]

[0026] A representation can include one or more segments. Each representation may include an initialization segment, or each segment of a representation may be self-initializing. When present, the initialization segment can include initialization information for accessing the representation. Generally, the initialization segment does not contain media data. A segment can be uniquely referenced by an identifier such as a Uniform Resource Locator (URL), Uniform Resource Name (URN), or Uniform Resource Identifier (URI). The MPD can give an identifier to each segment. In some examples, the MPD can also give a byte range in the form of a range attribute corresponding to data for a segment within a file accessible by a URL, URN, or URI.

[0019]

[0027] Different representations can be selected for substantially simultaneous retrieval of different types of media data. For example, a client device may select an audio representation, a video representation, and a timed text representation that retrieve segments. In some examples, the client device may select a particular set of adaptations for performing bandwidth adaptation. That is, the client device may select a set of adaptations that includes a video representation, a set of adaptations that includes an audio representation, and / or a set of adaptations that includes timed text. Alternatively, the client device may select a set of adaptations for one type of media (e.g., video) and directly select representations for other types of media (e.g., audio and / or timed text).

[0020]

[0028] Low-Latency Dynamic Adaptive Streaming over HTTP (LL-DASH) is a profile for DASH that attempts to provide media data to DASH clients with low latency. Some of the techniques for LL-DASH are briefly summarized below. · Encoding is based on fragmented ISO BMFF files, typically assuming CMAF fragments and CMAF chunks. · Each chunk is individually accessible by a DASH packager and is mapped to an HTTP chunk that is uploaded to the origin server. This one-to-one mapping is a recommendation for low-latency operation but not a requirement. The client should never assume that this one-to-one mapping is stored at the client. · A low-latency protocol for partially available segments, such as HTTP chunk transfer encoding, is used so that the client can access the segment before it is complete. The available start time is adjusted for clients that can utilize this feature. · The following two operating modes are permitted.

[0021] ○ Simple live offerings are used by applying @duration signaling and $Number$-based templating.

[0022] ○ Main live offerings having a SegmentTimeline as either $Number$ or $Time$ are supported by the updates proposed in DASH version 4. · The MPD validity expiration event may be used but is not essential for client understanding. · In general, in-band event messages may exist, but clients are only expected to recover them at the start of segments, not at any chunk. The DASH packager may receive notifications from the encoder at chunk boundaries or fully asynchronously using a timed metadata track. · In a single media presentation, it is permitted to have an adaptation set using chunked low-latency mode within one period of the media presentation and an adaptation set using short segments for a different media type. · A certain amount of playback control of the DASH client on the media pipeline may be available and should be used for the robustness of the DASH client. For example, playback may be accelerated or decelerated for a period of time, or the DASH client may perform a seek to a segment. · The system is designed to be executable with standard HTTP / 1.1 but should also be applicable to HTTP extensions and other protocols for improved low-latency operation. · The MPD includes explicit signaling in service configuration and service properties (including, for example, the target latency of the service). · The MPD and possibly also segments include an anchor time that enables the DASH client to measure the current latency compared to live and adjust to meet service expectations. ·For example, operational robustness is targeted, for example, in the case of encoder failures. ·Existing DRM and encryption modes are compatible with the proposed low-latency operation.

[0023]

[0029] Based on the above high-level overview, the following is defined. A segment can be used to randomly access the representation not only at the segment boundary but also within the segment. If such random access is provided, this should be signaled in the MPD.

[0024]

[0030] FIG. 1 is a block diagram showing an exemplary system 10 that implements a technique for streaming media data over a network. In this example, system 10 includes a content creation device 20, a server device 60, and a client device 40. The client device 40 and the server device 60 are communicatively coupled by a network 74 that may comprise the Internet. In some examples, the content creation device 20 and the server device 60 may also be coupled by the network 74 or another network, or may be communicatively coupled directly. In some examples, the content creation device 20 and the server device 60 may comprise the same device.

[0025]

[0031] In the example of FIG. 1, the content creation device 20 includes an audio source 22 and a video source 24. The audio source 22 may include, for example, a microphone that generates an electrical signal representing captured audio data to be encoded by an audio encoder 26. Alternatively, the audio source 22 may include a storage medium storing previously recorded audio data, an audio data generator such as a computerized synthesizer, or any other source of audio data. The video source 24 may include a video camera that generates video data to be encoded by a video encoder 28, a storage medium encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other source of video data. The content creation device 20 is not necessarily communicatively coupled to the server device 60 in all examples, but may store multimedia content in a separate medium readable by the server device 60.

[0026]

[0032] The raw audio and video data may comprise analog or digital data. The analog data may be digitized before being encoded by the audio encoder 26 and / or the video encoder 28. The audio source 22 may obtain audio data from a call participant while the call participant is speaking, and at the same time, the video source 24 may obtain video data of the call participant. In other examples, the audio source 22 may include a computer-readable storage medium storing stored audio data, and the video source 24 may include a computer-readable storage medium storing stored video data. In this way, the techniques described in the present disclosure may be applied to live, streaming, real-time audio and video data, or archived, pre-recorded audio and video data.

[0027]

[0033] An audio frame corresponding to a video frame is generally an audio frame that includes audio data captured (or generated) by an audio source 22 simultaneously with video data captured (or generated) by a video source 24 included within the video frame. For example, while a call participant generally generates audio data by speaking, the audio source 22 captures the audio data, and simultaneously, i.e., while the audio source 22 is capturing the audio data, the video source 24 captures the video data of the call participant. Thus, an audio frame can temporally correspond to one or more specific video frames. Accordingly, an audio frame corresponding to a video frame generally corresponds to a situation where audio data and video data are captured simultaneously, and a situation where the audio frame and the video frame each comprise audio data and video data captured simultaneously.

[0028]

[0034] In some examples, an audio encoder 26 can encode a time stamp in each encoded audio frame that represents the time at which the audio data of the encoded audio frame was recorded. Similarly, a video encoder 28 can encode a time stamp in each encoded video frame that represents the time at which the video data of the encoded video frame was recorded. In such examples, an audio frame corresponding to a video frame can comprise an audio frame with a certain time stamp and a video frame with the same time stamp. The content creation device 20 can include an internal clock from which the audio encoder 26 and / or the video encoder 28 can generate time stamps, or that the audio source 22 and the video source 24 can use to associate the audio data and the video data with time stamps, respectively.

[0029]

[0035] In some examples, the audio source 22 can send data corresponding to the time when the audio data was recorded to the audio encoder 26, and the video source 24 can send data corresponding to the time when the video data was recorded to the video encoder 28. In some examples, the audio encoder 26 can encode the sequence identifier in the encoded audio data to indicate the relative time order of the encoded audio data, without necessarily indicating the absolute time when the audio data was recorded. Similarly, the video encoder 28 can also use the sequence identifier to indicate the relative time order of the encoded video data. Similarly, in some examples, the sequence identifier can be mapped to a timestamp or may be correlated with a timestamp in some cases.

[0030]

[0036] The audio encoder 26 generally generates a stream of encoded audio data, while the video encoder 28 generates a stream of encoded video data. Each individual stream of data (regardless of whether it is audio or video) may be referred to as an elementary stream. An elementary stream is a single, digitally encoded (and possibly compressed) component of a representation. For example, the encoded video or audio portion of a representation can be an elementary stream. The elementary stream can be converted into a packetized elementary stream (PES) before being encapsulated within a video file. A stream ID can be used to distinguish PES packets belonging to one elementary stream from others within the same representation. The basic data unit of an elementary stream is a packetized elementary stream (PES) packet. Thus, the encoded video data generally corresponds to an elementary video stream. Similarly, the audio data corresponds to one or more respective elementary streams.

[0031]

[0037] Many video coding standards, such as ITU-T H.264 / AVC and the upcoming High Efficiency Video Coding (HEVC) standard, define syntax, semantics, and a decoding process for an error-free bitstream, all of which conform to a specific profile or level. Video coding standards usually do not specify an encoder, but the encoder is tasked with ensuring that the generated bitstream conforms to the decoder's standard. In the context of video coding standards, a "profile" corresponds to a subset of algorithms, features, or tools, and the constraints applied to them. For example, the "profile" defined by the H.264 standard is a subset of the entire bitstream syntax specified by the H.264 standard. A "level" corresponds to a limit on decoder resource consumption, such as decoder memory and calculations related to picture resolution, bitrate, and block processing rate. A profile can be signaled by a profile_idc (profile indicator) value, while a level can be signaled by a level_idc (level indicator) value.

[0032]

[0038] The H.264 standard acknowledges that, for example, within the bounds imposed by the syntax of a given profile, it may still require significant variations in the performance of the encoder and decoder depending on the values taken by syntax elements in the bitstream, such as the specified size of the decoded picture. The H.264 standard further acknowledges that in many applications, it is neither practical nor economical to implement a decoder that can handle all hypothetical uses of the syntax within a particular profile. Therefore, the H.264 standard defines a "level" as a defined set of constraints imposed on the values of syntax elements in the bitstream. These constraints can be simple restrictions on values. Alternatively, these constraints can take the form of constraints on combinations of values (e.g., picture width × picture height × number of pictures decoded per second). The H.264 standard further specifies that individual implementations may support different levels for each supported profile.

[0033]

[0039] A decoder compliant with a profile typically supports all the functions defined in the profile. For example, as a coding function, B-picture coding is not supported in the baseline profile of H.264 / AVC but is supported in other profiles of H.264 / AVC. A decoder compliant with a level should be able to decode any bitstream that does not require resources beyond the limits defined at the level. The definitions of profiles and levels can be useful for explainability. For example, during video transmission, a pair of profile definition and level definition can be negotiated and agreed upon for the entire transmission session. More specifically, in H.264 / AVC, the level can define the limit on the number of macroblocks that need to be processed, the size of the decoded picture buffer (DPB), the size of the coded picture buffer (CPB), the vertical motion vector range, the maximum number of motion vectors per two consecutive MBs, and whether a sub-macroblock partition with a B-block smaller than 8×8 pixels can be present. In this way, the decoder can determine whether it is possible for the decoder to properly decode the bitstream.

[0034]

[0040] In the example of FIG. 1, the encapsulation unit 30 of the content creation device 20 receives an elementary stream comprising the encoded video data from the video encoder 28 and an elementary stream comprising the encoded audio data from the audio encoder 26. In some examples, the video encoder 28 and the audio encoder 26 may each include a packetizer for forming PES packets from the encoded data. In other examples, the video encoder 28 and the audio encoder 26 may each interface with a respective packetizer for forming PES packets from the encoded data. In still other examples, the encapsulation unit 30 may include a packetizer for forming PES packets from the encoded audio data and the encoded video data.

[0035]

[0041] Video encoder 28 can encode video data of multimedia content in various ways to generate different representations of the multimedia content using various characteristics such as various bitrates, pixel resolution, frame rate, compliance with various coding standards, compliance with various profiles and / or levels of profiles for various coding standards, representations having one or more views (e.g., for 2D or 3D playback), or other such characteristics. The representations used in the present disclosure may comprise one of audio data, video data, text data (e.g., for closed captions), or other such data. The representation may include an elementary stream such as an audio elementary stream or a video elementary stream. Each PES packet may include a stream_id that identifies the elementary stream to which the PES packet belongs. Encapsulation unit 30 is responsible for assembling the elementary stream into video files (e.g., segments) of various representations.

[0036]

[0042] Encapsulation unit 30 receives PES packets of the elementary stream of the representation from audio encoder 26 and video encoder 28 and forms corresponding network abstraction layer (NAL) units from the PES packets. The encoded video segment can be organized into NAL units that provide a "network-friendly" video representation for applications such as video telephony, storage, broadcast, or streaming. The NAL units can be classified into video coding layer (VCL) NAL units and non-VCL NAL units. The VCL units may include a core compression engine and may include block, macroblock, and / or slice level data. The other NAL units can be non-VCL NAL units. In some examples, the coded picture during one time instance, which is usually presented as a primary coded picture, may be included in an access unit that may include one or more NAL units.

[0037]

[0043] Non-VCL NAL units may include, in particular, parameter set NAL units and SEI NAL units. Parameter sets may include sequence level header information (in the sequence parameter set (SPS)) and picture level header information that rarely changes (in the picture parameter set (PPS)). If there are parameter sets (e.g., PPS and SPS), the information that rarely changes does not need to be repeated for each sequence or picture, and thus, the coding efficiency can be improved. Further, the use of parameter sets enables out-of-band transmission of important header information and can avoid the need for redundant transmission for error resilience. In an example of out-of-band transmission, the parameter set NAL unit can be transmitted on a different channel from other NAL units, such as an SEI NAL unit.

[0038]

[0044] Supplementary enhancement information (SEI) may include information that is not necessary to decode the coded picture samples from the VCL NAL unit but may assist in processes related to decoding, display, error resilience, and other purposes. SEI messages may be included in non-VCL NAL units. SEI messages are a normative part of some standard specifications and thus are not always essential for decoder implementations that comply with the standard. SEI messages can be sequence level SEI messages or picture level SEI messages. Some sequence level information may be included in SEI messages, such as scalability information SEI messages in an example of SVC, view scalability information SEI messages in MVC, etc. These exemplary SEI messages can carry, for example, information related to the extraction of operation points and the characteristics of those operation points. In addition, the encapsulation unit 30 may form a manifest file, such as a media presentation descriptor (MPD) that describes the characteristics of the presentation. The encapsulation unit 30 may format the MPD according to the extensible markup language (XML).

[0039]

[0045] The encapsulation unit 30 can provide data for one or more representations of multimedia content to the output interface 32, together with a manifest file (e.g., MPD). The output interface 32 can include a universal serial bus (USB) interface, a network interface or an interface for writing to a storage medium such as a CD or DVD writer or burner, an interface to a magnetic or flash storage medium, or other interfaces for storing or transmitting media data. The encapsulation unit 30 can provide the data of each representation of the multimedia content to the output interface 32, and the output interface 32 can send the data to the server device 60 via network transmission or a storage medium. In the example of FIG. 1, the server device 60 includes a storage medium 62 for storing various multimedia contents 64, and each multimedia content 64 includes its respective manifest file 66 and one or more representations 68A-68N (representations 68). In some examples, the output interface 32 can also send the data directly to the network 74.

[0040]

[0046] In some examples, the representations 68 can be separated into adaptation sets. That is, various subsets of the representations 68 can include a common set of characteristics for each, such as codec, profile and level, resolution, number of views, file format of segments, text type information that can identify the text to be displayed using the representation and / or the language or other characteristics of the audio data to be decoded and presented, for example, by a speaker, camera angle information that can describe the camera angle of the scene for the representation in the adaptation set or the real-world camera perspective, rating information that describes the content suitability for a particular viewer, etc.

[0041]

[0047] The manifest file 66 can include data indicating a subset of the representations 68 corresponding to a particular adaptation set and common characteristics about the adaptation set. The manifest file 66 can also include data representing individual characteristics about the individual representations of the adaptation set, such as bitrate. In this way, the adaptation set can enable simplified network bandwidth adaptation. The representations within the adaptation set can be indicated using child elements of the adaptation set elements of the manifest file 66.

[0042]

[0048] The server device 60 includes a request processing unit 70 and a network interface 72. In some examples, the server device 60 can include multiple network interfaces. Further, any or all of the functions of the server device 60 can be implemented on other devices of the content delivery network, such as a router, bridge, proxy device, switch, or other device. In some examples, an intermediate device of the content delivery network can cache data of the multimedia content 64 and can include components that substantially conform to the components of the server device 60. Generally, the network interface 72 is configured to transmit and receive data via the network 74.

[0043]

[0049] The request processing unit 70 is configured to receive network requests for the data in the storage medium 62 from a client device such as the client device 40. For example, the request processing unit 70 may implement the Hypertext Transfer Protocol (HTTP) version 1.1 described in RFC 2616, "Hypertext Transfer Protocol - HTTP / 1.1", Network Working Group, IETF, June 1999, by R. Fielding et al. That is, the request processing unit 70 may be configured to receive HTTP GET requests or partial GET requests and provide the data of the multimedia content 64 in response to these requests. The request may specify one of the segments of the representation 68, for example, using the URL of the segment. In some examples, the request may also specify one or more byte ranges of the segment and thus may include a partial GET request. The request processing unit 70 may be further configured to service HTTP HEAD requests for providing the header data of one of the segments of the representation 68. In any case, the request processing unit 70 may be configured to process requests to provide the requested data to the requesting device such as the client device 40.

[0044]

[0050] Additionally or alternatively, the request processing unit 70 may be configured to deliver media data via a broadcast or multicast protocol such as eMBMS. The content creation device 20 may create DASH segments and / or sub - segments in substantially the same way as described, but the server device 60 may use eMBMS or another broadcast or multicast network transport protocol to deliver these segments or sub - segments. For example, the request processing unit 70 may be configured to receive a multicast group participation request from the client device 40. That is, the server device 60 can advertise an Internet Protocol (IP) address associated with a multicast group to client devices associated with a particular media content (e.g., a live event broadcast) including the client device 40. Then, the client device 40 can submit a request to join the multicast group. This request can be propagated across the network 74, for example, to the routers that make up the network 74, such that the routers direct traffic destined for the IP address associated with the multicast group to joined client devices such as the client device 40.

[0045]

[0051] As shown in the example of FIG. 1, the multimedia content 64 may include a manifest file 66 that corresponds to a Media Presentation Description (MPD). The manifest file 66 may include descriptions of different alternative representations 68 (e.g., video services with different qualities), and the descriptions may include, for example, codec information of the representation 68, profile values, level values, bitrates, and other descriptive characteristics. The client device 40 may retrieve the MPD of the media presentation to determine how to access the segments of the representation 68.

[0046]

[0052] Specifically, the extraction unit 52 can extract configuration data (not shown) of the client device 40 to determine the decoding ability of the video decoder 48 and the rendering ability of the video output 44. The configuration data may also include any or all of the language preference selected by the user of the client device 40, one or more camera perspectives corresponding to the depth preference set by the user of the client device 40, and / or the rating preference selected by the user of the client device 40. The extraction unit 52 may comprise, for example, a web browser or media client configured to submit HTTP GET requests and partial GET requests. The extraction unit 52 may correspond to software instructions executed by one or more processors or processing units (not shown) of the client device 40. In some examples, all or part of the functions described with respect to the extraction unit 52 may be implemented in hardware, or a combination of hardware, software, and / or firmware, and hardware essential for executing software or firmware instructions may be provided.

[0047]

[0053] The extraction unit 52 can compare the decoding and rendering capabilities of the client device 40 with the characteristics of the representation 68 indicated by the information in the manifest file 66. The extraction unit 52 can first extract at least a portion of the manifest file 66 to determine the characteristics of the representation 68. For example, the extraction unit 52 may request a portion of the manifest file 66 that describes the characteristics of one or more adaptation sets. The extraction unit 52 can select a subset (e.g., an adaptation set) of the representation 68 having characteristics that can be satisfied by the coding and rendering capabilities of the client device 40. Next, the extraction unit 52 can determine the bitrate of the representation in the adaptation set, determine the currently available amount of network bandwidth, and extract a segment from one of the representations having a bitrate that can be satisfied by the network bandwidth.

[0048]

[0054] Generally, higher bitrate representations can result in higher quality video playback, but when the available network bandwidth decreases, lower bitrate representations can provide sufficient quality video playback. Thus, when the available network bandwidth is relatively high, the retrieval unit 52 can retrieve data from a relatively high bitrate representation, but when the available network bandwidth is low, the retrieval unit 52 can retrieve data from a relatively low bitrate representation. In this way, the client device 40 can adapt to the changing network bandwidth availability of the network 74 while streaming multimedia data over the network 74.

[0049]

[0055] Additionally or alternatively, the retrieval unit 52 can be configured to receive data according to a broadcast or multicast network protocol such as eMBMS or IP multicast. In such an example, the retrieval unit 52 can submit a request to join a multicast network group associated with a particular media content. After joining the multicast group, the retrieval unit 52 can receive the data of the multicast group without further requests issued to the server device 60 or the content creation device 20. The retrieval unit 52 can submit a request to leave the multicast group when the data of the multicast group is no longer needed, for example, to stop playback or to change channels to a different multicast group.

[0050]

[0056] The network interface 54 can receive the data of the segment of the selected representation and provide it to the extraction unit 52, and the extraction unit 52 can in turn provide the segment to the decapsulation unit 50. The decapsulation unit 50 decapsulates the elements of the video file into a PES stream, depacketizes the PES stream to extract the encoded data, and sends the encoded data to either the audio decoder 46 or the video decoder 48 depending on whether the encoded data is part of an audio stream or part of a video stream as indicated by, for example, the PES packet header of the stream. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, and the video decoder 48 decodes the encoded video data and sends the decoded video data, which may include multiple views of the stream, to the video output 44.

[0051]

[0057] Video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, extraction unit 52, and decapsulation unit 50 can each be implemented as any of a variety of suitable processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof, where applicable. Each of video encoder 28 and video decoder 48 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined video encoder / decoder (codec). Similarly, each of audio encoder 26 and audio decoder 46 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined codec. An apparatus including video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, extraction unit 52, and / or decapsulation unit 50 can include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular phone.

[0052]

[0058] Client device 40, server device 60, and / or content creation device 20 can be configured to operate in accordance with the techniques of the present disclosure. By way of example, the present disclosure will describe these techniques with respect to client device 40 and server device 60. However, it should be understood that content creation device 20 can be configured to perform these techniques instead of (or in addition to) server device 60.

[0053]

[0059] The encapsulation unit 30 may form a NAL unit that includes a header identifying the program to which the NAL unit belongs, as well as a payload, for example, audio data, video data, or data describing the transport or program stream corresponding to the NAL unit. For example, in H.264 / AVC, a NAL unit includes a 1-byte header and a payload of variable size. The NAL unit including video data in its payload may include video data at various granularity levels. For example, a NAL unit may include a block of video data, a plurality of blocks, a slice of video data, or an entire picture of video data. The encapsulation unit 30 may receive encoded video data from the video encoder 28 in the form of PES packets of an elementary stream. The encapsulation unit 30 may associate each elementary stream with a corresponding program.

[0054]

[0060] The encapsulation unit 30 may also assemble an access unit from a plurality of NAL units. Generally, an access unit may include one or more NAL units for representing a frame of video data and, when audio data corresponding to the frame is available, such audio data. An access unit generally includes all NAL units over one output time instance, for example, all audio data and video data over one time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, the unique frames of all views of the same access unit (same time instance) may be rendered simultaneously. In one example, an access unit may include an encoded picture during one time instance that may be presented as a primary encoded picture.

[0055]

[0061] Therefore, an access unit may comprise all audio frames and video frames of a common time instance, for example, all views corresponding to time X. The present disclosure also refers to the encoded pictures of a specific view as "view components". That is, a view component may comprise an encoded picture (or frame) of a specific view at a specific time. Therefore, an access unit may be defined as comprising all view components of a common time instance. The decoding order of the access unit does not necessarily have to be the same as the output order or the display order.

[0056]

[0062] A media presentation may include a Media Presentation Description (MPD) that may contain descriptions of different alternative representations (e.g., video services having different qualities), the description may include, for example, codec information, profile values, and level values. The MPD is an example of a manifest file such as manifest file 66. To determine how to access the movie fragments of various presentations, the client device 40 may retrieve the MPD of the media presentation. The movie fragments may be arranged in the movie fragment box (moof box) of the video file.

[0057]

[0063] The manifest file 66 (which may comprise an MPD, for example) may advertise the availability of segments of the representation 68. That is, the MPD may include information indicating the wall clock time when the first segment of one of the representations 68 becomes available, and information indicating the duration of the segments within the representation 68. In this way, the retrieval unit 52 of the client device 40 may determine when each segment is available based on the start time and the duration of the segments preceding a specific segment.

[0058]

[0064] After the encapsulation unit 30 assembles NAL units and / or access units into a video file based on the received data, the encapsulation unit 30 passes the video file to the output interface 32 for output. In some examples, instead of storing the video file locally or sending the video file directly to the client device 40, the encapsulation unit 30 can send the video file to a remote server via the output interface 32. The output interface 32 can include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium such as an optical drive, a magnetic media drive (e.g., a floppy (registered trademark) drive), a universal serial bus (USB) port, a network interface, or other output interfaces. The output interface 32 outputs the video file to a computer-readable medium such as, for example, a transmission signal, magnetic media, optical media, memory, a flash drive, or other computer-readable media.

[0059]

[0065] The network interface 54 can receive NAL units or access units via the network 74 and provide the NAL units or access units to the decapsulation unit 50 via the extraction unit 52. The decapsulation unit 50 decapsulates the elements of the video file into a PES stream, depacketizes the PES stream to extract the encoded data, and sends the encoded data to either the audio decoder 46 or the video decoder 48 depending on whether the encoded data is part of an audio stream or a video stream as indicated, for example, by the PES packet header of the stream. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, and the video decoder 48 decodes the encoded video data and sends the decoded video data, which can include multiple views of the stream, to the video output 44.

[0060]

[0066] According to the techniques of the present disclosure, the content creation device 20 and / or the server device 60 may add additional random access points in the DASH / CMAF segments. Random access includes clean random access and open or progressive decoder refresh until it only provides resynchronization during file format parsing. This can be addressed by providing a chunk boundary that provides information that resynchronization and decoding can start at this point, and signaling regarding the type of the following random access points. The availability of tfdt enables time resynchronization at the presentation time level, together with the use of moof header information and optionally the initialization segment. In the present disclosure, this new point is referred to as a "resynchronization point". That is, the resynchronization point represents a point at which a file-level container (e.g., a box in ISO BMFF) can be appropriately parsed, followed by the occurrence of a random access point (e.g., an I-frame) in the media data. Thus, the client device 40 can randomly access the multimedia content 64, for example, at one of these random access points.

[0061]

[0067] The content creation device 20 and / or the server device 60 may also add appropriate signaling in the manifest file 66 (e.g., MPD) that indicates the availability of random access points and resynchronization in each DASH segment, and provides information regarding the location, type, and timing of the random access points. The content creation device 20 and / or the server device 60 may optionally add characteristics regarding the location, timing, type of random access, and whether the information is accurate or estimated, to provide signaling in the manifest file 66 (MPD) indicating that additional resynchronization points are available in the segment. Thus, the client device 40 can use this signaled data to determine whether such random access points are available and resynchronize extraction and playback accordingly.

[0062]

[0068] For any starting point, the client device 40 can be configured to resynchronize to decapsulation, decoding, and deciphering by finding a resynchronization point. The content creation device 20 and / or the server device 60 can provide appropriate chunks that meet the above requirements addressed in the CMAF TUC. Different types can be defined later.

[0063]

[0069] The client device 40 can be configured to start processing in a restricted receiver environment, for example, to be available in HTML-5 / MSE-based playback. This problem can be addressed through the receiver implementation form. However, it is appropriate to provide resynchronization triggers and information to a receiver pipeline that has the ability to obtain a map of resynchronization points in terms of data structure, timing, and type, which enables the user of the decoding pipeline to initialize playback at a random access point.

[0064]

[0070] Furthermore, signaling in the backward-compatible manifest file 66 can be provided. Client devices not configured with the ability to analyze the signaling can ignore the signaling and execute the methods described above. Additionally, the manifest file 66 can include signaling that associates the position with the value of @bandwidth to enable signaling at the adaptation set level.

[0065]

[0071] FIG. 2 is a block diagram showing in more detail an exemplary set of components of the extraction unit 52 of FIG. 1. In this example, the extraction unit 52 includes an eMBMS middleware unit 100, a DASH client 110, and a media application 112.

[0066]

[0072] In this example, the eMBMS middleware unit 100 further includes an eMBMS receiving unit 106, a cache 104, and a proxy server unit 102. In this example, the eMBMS receiving unit 106 is configured to receive data via eMBMS, for example, according to File Delivery over Unidirectional Transport (FLUTE) described in T. Paira et al., "FLUTE - File Delivery over Unidirectional Transport", Network Working Group, RFC6726, November 2012, which is available at tools.ietf.org / html / rfc6726. That is, the eMBMS receiving unit 106 can receive a file via broadcast from a server device 60 that can act as a Broadcast / Multicast Service Center (BM-SC), for example.

[0067]

[0073] When the eMBMS middleware unit 100 receives data related to a file, the eMBMS middleware unit may store the received data in the cache 104. The cache 104 may comprise a computer-readable storage medium such as a flash memory, a hard disk, a RAM, or any other suitable storage medium.

[0068]

[0074] The proxy server unit 102 can act as a server for the DASH client 110. For example, the proxy server unit 102 can provide an MPD file or other manifest file to the DASH client 110. The proxy server unit 102 can advertise the available time for segments in the MPD file and the hyperlinks from which the segments can be retrieved. These hyperlinks can include the local host address prefix corresponding to the client device 40 (e.g., 127.0.0.1 in the case of IPv4). In this way, the DASH client 110 can request segments from the proxy server unit 102 using an HTTP GET or partial GET request. For example, in the case of a segment available from the link http: / / 127.0.0.1 / rep1 / seg3, the DASH client 110 can construct an HTTP GET request including a request for http: / / 127.0.0.1 / rep1 / seg3 and submit that request to the proxy server unit 10 to. The proxy server unit 102 can retrieve the requested data from the cache 104 and provide that data to the DASH client 110 in response to such a request.

[0069]

[0075] Figure 3 is a conceptual diagram showing elements of an exemplary multimedia content 120. The multimedia content 120 can correspond to the multimedia content 64 (FIG. 1) or another multimedia content stored in the storage medium 62. In the example of FIG. 3, the multimedia content 120 includes a media presentation description (MPD) 122 and a plurality of representations 124A - 124N (representations 124). Representation 124A includes any header data 126 and segments 128A - 128N (segments 128), and representation 124N includes any header data 130 and segments 132A - 132N (segments 132). The letter N is used for convenience to designate the last movie fragment in each of the representations 124. In some examples, there may be a different number of movie fragments between the representations 124.

[0070]

[0076] MPD 122 may have a data structure different from that of representation 124. MPD 122 may correspond to the manifest file 66 in FIG. 1. Similarly, representation 124 may correspond to the representation 68 in FIG. 1. Generally, MPD 122 may include data that generally describes the characteristics of representation 124, such as coding characteristics and rendering characteristics, adaptation sets, the profile to which MPD 122 corresponds, text type information, camera angle information, rating information, trick mode information (e.g., information indicating a representation including a time subsequence), and / or information for extracting a remote period (e.g., for targeted advertisement insertion into media content during playback).

[0071]

[0077] When present, the header data 126 can describe the characteristics of segment 128, such as the time location of a random access point (RAP, also called a stream access point (SAP)), and the random access point of segment 128 can include a random access point, a byte offset to the random access point within segment 128, a uniform resource locator (URL) of segment 128, or other aspects of segment 128. When present, the header data 130 can describe similar characteristics regarding segment 132. Additionally or alternatively, such characteristics can be fully included within MPD 122.

[0072]

[0078] Segments 128, 132 include one or more coded video samples, and each of the coded video samples can include a frame or slice of video data. Each of the coded video samples of segment 128 can have similar characteristics, such as height, width, and bandwidth requirements. Such characteristics can be described by the data of MPD 122, but such data is not shown in the example of FIG. 3. In addition to any or all of the signaling information described in this disclosure, MPD 122 can include characteristics described by 3GPP specifications.

[0073]

[0079] Each of segments 128, 132 can be associated with a unique Uniform Resource Locator (URL). Thus, each of segments 128, 132 can be independently retrievable using a streaming network protocol such as DASH. In this way, a destination device such as client device 40 can use an HTTP GET request to retrieve segment 128 or 132. In some examples, client device 40 can use an HTTP partial GET request to retrieve a specific byte range of segment 128 or 132.

[0074]

[0080] According to the techniques of the present disclosure, the MPD 122 (which may also correspond to the manifest file 66 of FIG. 1) can include signaling to address the problems described above. For example, for each of the representations 124 (and optionally default set at the adaptation set level), the MPD 122 can include one or more resynchronization elements (which may enable backward compatibility). Each resynchronization element can indicate that the following holds for each of segments 128, 132 within the corresponding one of the representations 124. · A resynchronization point of type @type or smaller (but greater than 0) stream access point (SAP) exists within each segment having a maximum delta T signaled by @dT, a maximum byte offset difference signaled by @dImax, and a minimum byte offset difference signaled by @dImin, and the values of both @dImax and @dImin need to be multiplied by the value of the @bandwidth attribute assigned to this representation to obtain the values. "Delta T" refers to the difference in the earliest presentation time of any data following the resynchronization point in units of the @timescale of the representation. If @type is set to 0, only resynchronization regarding the container and decoding levels is guaranteed. · The resynchronization marker flag @marker can be set to indicate that a resynchronization point is included for each resynchronization point using a resynchronization pattern defined by the segment format in use. · Multiple resynchronization elements may exist for different SAP types. · The resynchronization point requires that media processing can occur on the ISO BMFF and decoding information in combination with the CMAF header / initialization segment.

[0075]

[0081] An exemplary use of the resynchronization point is described with respect to Figure 8 below.

[0076]

[0082] When the resynchronization element is provided in MPD122, the following may hold for segments with respect to either ISO BMFF - based segment 128 or CMAF - based segment 132. · The segment can be a sequence of one or more chunks as defined below. Further, for any two consecutive resynchronization points in a segment of the @type specified in the element, the following may hold. ○ The difference in the earliest presentation times of the two is at most the value of @dT.

[0077] ○ The difference in byte offsets from the start is at most @dImax normalized by the @bandwidth value.

[0078] ○ The difference in byte offsets from the start is at least @dImin normalized by the @bandwidth value. · When the resynchronization marker flag is set, each resynchronization point may include a resynchronization box / styp.

[0079]

[0083] FIG. 4 is a block diagram showing elements of an exemplary video file 150 that may correspond to segments of a representation such as one of segments 128, 132 of FIG. 3. Each of segments 128, 132 may include data that substantially matches an array of data shown in the example of FIG. 4. The video file 150 is sometimes said to encapsulate the segments. As described above, video files according to the ISO base media file format and its extensions store data in a series of objects called “boxes”. In the example of FIG. 4, the video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a segment index (sidx) box 162, a movie fragment (MOOF) box 164, and a movie fragment random access (MFRA) box 166. Although FIG. 4 represents an example of a video file, it should be understood that other media files may include other types of media data (e.g., audio data, timed text data, etc.) structured similarly to the data of video file 150 according to the ISO base media file format and its extensions.

[0080]

[0084] The file type (FTYP) box 152 generally describes the file type of the video file 150. The file type box 152 may include data that identifies a specification describing the best use of the video file 150. The file type box 152 may alternatively be placed before the MOOV box 154, the movie fragment box 164, and / or the MFRA box 166.

[0081]

[0085] In some examples, a segment such as video file 150 may include an MPD update box (not shown) before the FTYP box 152. The MPD update box may include information for updating the MPD, along with information indicating that the MPD corresponding to the representation including video file 150 should be updated. For example, the MPD update box may provide a URI or URL for the resources used to update the MPD. As another example, the MPD update box may include data for updating the MPD. In some examples, the MPD update box can come immediately after the segment type (STYP) box (not shown) of video file 150, where the STYP box may define the segment type of video file 150.

[0082]

[0086] In the example of FIG. 4, the MOOV box 154 includes a movie header (MVHD) box 156, a track (TRAK) box 158, and one or more movie extension (MVEX) boxes 160. Generally, the MVHD box 156 may describe the general characteristics of video file 150. For example, the MVHD box 156 may include data describing when video file 150 was first generated, when video file 150 was last modified, the time axis of video file 150, the duration of playback of video file 150, or other data generally describing video file 150.

[0083]

[0087] The TRAK box 158 may include data about the tracks of video file 150. The TRAK box 158 may include a track header (TKHD) box that describes the characteristics of the track corresponding to the TRAK box 158. In some examples, the TRAK box 158 may include encoded video pictures, while in other examples, the encoded video pictures of the track may be included in movie fragment 164 that can be referenced by the data of the TRAK box 158 and / or the sidx box 162.

[0084]

[0088] In some examples, video file 150 may include two or more tracks. Thus, MOOV box 154 may include several TRAK boxes equal to the number of tracks in video file 150. TRAK box 158 may describe the characteristics of the corresponding track of video file 150. For example, TRAK box 158 may describe time and / or spatial information about the corresponding track. When encapsulation unit 30 (FIG. 3) includes a parameter set track in a video file such as video file 150, a TRAK box similar to TRAK box 158 of MOOV box 154 may describe the characteristics of the parameter set track. Encapsulation unit 30 may signal that a sequence level SEI message exists in the parameter set track within the TRAK box that describes the parameter set track.

[0085]

[0089] MVEX box 160 may describe the characteristics of the corresponding movie fragment 164, for example, to signal that video file 150 includes movie fragment 164 in addition to the video data contained within MOOV box 154 if any. In the context of streaming video data, the encoded video picture may be contained in movie fragment 164 rather than MOOV box 154. Thus, all encoded video samples may be contained in movie fragment 164 rather than MOOV box 154.

[0086]

[0090] MOOV box 154 may include several MVEX boxes 160 equal to the number of movie fragments 164 in video file 150. Each of MVEX boxes 160 may describe the characteristics of the corresponding one of movie fragments 164. For example, each MVEX box may include a movie extension header box (MEHD) box that describes the duration of the corresponding one of movie fragments 164.

[0087]

[0091] As described above, the encapsulation unit 30 may store the sequence data set in video samples that do not include actual encoded video data. The video samples may generally correspond to access units, which are representations of the encoded pictures at specific time instances. In the context of AVC, an encoded picture includes one or more VCL NAL units containing information for constructing all the pixels of the access unit, and other related non-VCL NAL units such as SEI messages. Thus, the encapsulation unit 30 may include the sequence data set, which may include the sequence level SEI message, in one of the movie fragments 164. The encapsulation unit 30 may further signal the presence of the sequence data set and / or the sequence level SEI message as being present in one of the movie fragments 164 within one of the MVEX boxes 160 corresponding to one of the movie fragments 164.

[0088]

[0092] The SIDX box 162 is an optional element of the video file 150. That is, video files conforming to the 3GPP file format, or other such file formats, do not necessarily include the SIDX box 162. According to an example of the 3GPP file format, the SIDX box can be used to identify sub - segments of a segment (e.g., a segment included in the video file 150). The 3GPP file format defines a sub - segment as "a self - contained set of one or more consecutive movie fragment boxes, having a corresponding media data box and a media data box containing data referenced by the movie fragment box, which must follow the movie fragment box and precede the next movie fragment box containing information about the same track." The 3GPP file format also indicates that the SIDX box "contains a sequence of references to sub - segments of the (sub) segment documented by the box. The referenced sub - segments are consecutive in presentation time. Similarly, the bytes referenced by the segment index box are always consecutive within the segment. The referenced size gives a count of the number of bytes in the referenced material."

[0089]

[0093] The SIDX box 162 generally provides information representing one or more sub - segments of a segment included in the video file 150. For example, such information can include the playback time at which the sub - segment starts and / or ends, the byte offset of the sub - segment, whether the sub - segment includes a stream access point (SAP) (e.g., starts with it), the type of the SAP (e.g., whether the SAP is an instantaneous decoder refresh (IDR) picture, a clean random access (CRA) picture, a broken link access (BLA) picture, etc.), the position of the SAP within the sub - segment (with respect to playback time and / or byte offset), etc.

[0090]

[0094] Movie fragment 164 may include one or more encoded video pictures. In some examples, movie fragment 164 may include one or more picture groups (GOPs), each of which may include some encoded video pictures, such as frames or pictures. Further, as described above, movie fragment 164 may include a sequence data set in some examples. Each of movie fragments 164 may include a movie fragment header box (MFHD, not shown in FIG. 4). The MFHD box may describe characteristics of the corresponding movie fragment, such as the sequence number of the movie fragment. Movie fragments 164 may be included in the order of the sequence numbers in video file 150.

[0091]

[0095] MFRA box 166 may describe random access points within movie fragment 164 of video file 150. This may assist in performing trick modes, such as performing a seek to a specific time location (i.e., playback time) within a segment encapsulated by video file 150. MFRA box 166 is generally optional and in some examples need not be included in the video file. Similarly, a client device, such as client device 40, need not necessarily refer to MFRA box 166 in order to correctly decrypt and display the video data of video file 150. MFRA box 166 may include a number of track fragment random access (TFRA) boxes (not shown), equal to the number of tracks of video file 150 or, in some examples, equal to the number of media tracks (e.g., non-hint tracks) of video file 150.

[0092]

[0096] In some examples, movie fragment 164 may include one or more stream access points (SAPs), such as an IDR picture. Similarly, MFRA box 166 may provide an indication of the location within video file 150 of the SAP. Thus, a temporal subsequence of video file 150 may be formed from the SAP of video file 150. The temporal subsequence may also include other pictures, such as P-frames and / or B-frames that are dependent on the SAP. The frames and / or slices of the temporal subsequence may be configured within the segment such that frames / slices of the temporal subsequence that are dependent on other frames / slices of the subsequence can be properly decoded. For example, in a hierarchical structure of data, data used for prediction of other data may also be included within the temporal subsequence.

[0093]

[0097] In one example, the present disclosure defines “chunk” as follows in the table below. This table provides exemplary definitions of both the cardinality and the sequentiality of the chunks.

[0094]

Table 1

[0095]

[0098] In one example, the present disclosure defines a resynchronization point as the start of a chunk. Further, the resynchronization point may be assigned the following properties. · It has a byte offset from the start of the segment that points to the first byte of the chunk. · It has the earliest presentation time derived from the information in the moof and optionally the movie header assigned to it. · It has the SAP type assigned to it as defined in ISO / IEC 14496-12. · There is an indication of whether the chunk contains a resynchronization box (the styp above). · Starting from the resynchronization point, file format parsing and decoding can be performed together with the information in the movie header.

[0096]

[0099] In some examples, a resynchronization marker box may be defined that enables synchronization to the start of a segment (e.g., video file 150) by scanning the byte stream of the marker. This resynchronization box may have the following properties. · It defines a unique pattern for resynchronization with a very high likelihood. · It defines the SAP type.

[0097]

[0100] The resynchronization marker box may be a new box or may reuse an existing box such as the styp box. In the present disclosure, it is assumed that a styp with specific limitations can be used as the resynchronization marker box. Research on the robustness of this approach is underway.

[0098]

[0101] The following FIGS. 5 to 7 are used to illustrate some use cases where random access to a DASH segment at points other than the start of the segment may be useful.

[0099]

[0102] FIG. 5 is a conceptual diagram showing an exemplary low-latency architecture 200 that can be used in a first use case according to the present disclosure. That is, FIG. 5 shows the basic flow of information for operating a low-latency DASH service by the DASH-IF IOP. The low-latency architecture 200 includes a DASH packager 202, an encoder 216, a content delivery network (CDN) 220, a normal DASH client 230, and a low-latency DASH client 232. The encoder 216 generally may correspond to either or both of the audio encoder 26 and the video encoder 28 of FIG. 1, while the DASH packager 202 may correspond to the encapsulation unit 30 of FIG. 1.

[0100]

[0103] In this example, the encoder 216 encodes the received media data to form CMAF headers (CH), such as CH208, CMAF initial chunks 206A, 206B (CIC206), and CMAF non-initial chunks 204A to 204D (CNC204). The encoder 216 provides the CH208, the CIC206, and the CNC204 to the DASH packager 202. The DASH packager 202 also receives a service description including general descriptions of the service and information about the encoder configuration of the encoder 216.

[0101]

[0104] The DASH packager 202 uses the service description, the CH208, the CIC206, and the CNC204 to form a Media Presentation Description (MPD) 210 and an initialization segment 212. The DASH packager 202 also generates and maps the CH208, the CIC206, and the CNC204 within segments 214A, 214B (segments 214), and provides the segments 214 to the CDN 220 in an incremental manner. The DASH packager 202 can deliver the segments 214 in the form of chunks when they are generated. The CDN 220 includes a segment storage 222 for storing the MPD 210, the IS 212, and the segments 214. The CDN 220 delivers complete segments to the normal DASH client 230 in response to HTTP Get or partial Get requests from, for example, the normal DASH client 230 and the low-latency DASH client 232, but delivers individual chunks (e.g., CH208, CIC206, and CNC204) to the low-latency DASH client 232.

[0102]

[0105] FIG. 6 is a conceptual diagram showing in more detail an example of the use case described with respect to FIG. 5. The example of FIG. 6 shows segments 250A-250E (segment 250) each including a respective set of chunks 252A-252E (chunk 252). A client device such as client device 40 of FIG. 1 can retrieve either a complete segment 250 or an individual chunk 252. For example, as shown in FIG. 5, a normal DASH client 230 can retrieve a segment 250, while a low-latency DASH client 232 can retrieve individual chunks 252 (at least initially).

[0103]

[0106] FIG. 6 further shows a way in which latency can be reduced by retrieving individual chunks 252 rather than complete segments 250. For example, retrieving a complete segment at the current time can cause higher latency. Simply retrieving the most recently fully available segment reduces latency, but can still result in relatively high latency.

[0104]

[0107] By retrieving chunks instead, these latencies can be significantly reduced. For example, at the current time indicated by "now" in FIG. 6, segment 250E is not fully formed. Nevertheless, assuming that chunks 252E-1 and 252E-2 are formed and available for retrieval, the client device can retrieve the formed chunks such as chunks 252E-1 and 252E-2 of segment 250E even before segment 250E is fully formed.

[0105]

[0108] When participating in a live stream, typically both low latency and fast startup should be achieved. However, this is not obvious, and several strategies are described below with reference to FIG. 6. · In the first case, at the live edge, the segment that is three segments behind in the time history (i.e., segments 250B, 250C, and 250D) is loaded into the buffer. When one segment becomes available, playback starts. This results in a significant latency, but playback can start relatively quickly because the random access at the start of the segment is loaded. · In the second case, instead of the segment that is three segments old, segment 250D, which is the latest available segment, is selected. The playback latency in this case is at least the segment duration, but it could be longer. The start may be the same as in the above case. · In the other three cases, a segment containing multiple chunks (e.g., segment 250E) is played while it is still being generated. This reduces latency, but there is a problem that the start of playback can be affected, especially when the difference between the segment availability start time of the latest published segment and the wall clock time is greater than the target latency. In this case, the client device may have to wait until the next second is emitted. In the case of a 6 - second segment, this can result in a startup latency of 4 - 5 seconds. · There are other techniques and use cases. For example, the client can access the old segment at startup, download all of it, accelerate playback, and perform fast - forward decoding. However, such an approach has the drawback that significant data needs to be downloaded before accelerated decoding can occur. Additionally, it is not widely supported in the decoder interface.

[0106]

[0109] A suitable solution could be the following. · At least one representation of the adaptive set includes more frequent random access points and non - initial chunks in the segment / fragment. ·The DASH client may determine, using the information from the MPD, that such a random access method exists, but the location / byte offset of the random access point may not be signaled precisely. ·The DASH client can access this representation at startup, but only starts downloading from the byte range of the latest available non-initial chunk or at least a byte range close thereto. ·Once downloaded, the DASH client may determine the random access points and also start processing the data together with the initialization segment / CMAF header of the same representation that was also downloaded. The location of the random access points will be described below.

[0107]

[0110] However, the latter approach may encounter various problems, as summarized below.

[0108]

[0111] Thus, as shown in the example of FIG. 6 and described below, the use of chunks as described in the present disclosure can substantially reduce latency. Signaling the start of the chunks in advance allows for the generation in advance of a manifest file that does not require frequent updates but can still indicate the approximate location of the stream access points (SAPs) within the chunks. In this way, the client device does not need continuous manifest file updates and can use the manifest file to determine the location of the chunk boundaries, but still allows the client device to start media streaming at the beginning of the chunk boundaries, for example at resynchronization points. That is, the client device can determine the byte range of the segment including the resynchronization point from the manifest file even before the segment is fully formed, because the manifest file can signal the byte range or other data representing the approximate location of the resynchronization point within the segment.

[0109]

[0112] FIG. 7 is a conceptual diagram showing an exemplary second specification case using DASH and CMAF random access in the context of a broadcast protocol. FIG. 7 shows an example including a media encoder 280, a CMAF / File Format (FF) packager 282, a DASH packager 284, a ROUTE sender 286, a CDN origin server 288, a ROUTE receiver 290, a DASH client 292, a CMAF / FF parser 294, and a media decoder 296. The media encoder 280 encodes media data such as audio data or video data. The media encoder 280 may correspond to the audio encoder 26 or the video encoder 28 of FIG. 1, or the encoder 216 of FIG. 5. The media encoder 280 provides the encoded media data to the CMAF / FF packager 282, and the CMAF / FF packager 282 formats the encoded media data in a file according to CMAF and a specific file format such as ISO BMFF or an extension thereof.

[0110]

[0113] The CMAF / FF packager 282 provides these files (e.g., chunks) to the DASH packager 284, and the DASH packager 284 aggregates the files / chunks into DASH segments. The DASH packager 284 may also form a manifest file such as an MPD that includes data describing the files / chunks / segments. Further, according to the techniques of the present disclosure, the DASH packager 284 may determine an approximate location of a future stream access point (SAP) or random access point (RAP) and signal the approximate location in the MPD. The CMAF / FF packager 282 and the DASH packager 284 may correspond to the encapsulation unit 30 of FIG. 1 or the DASH packager 202 of FIG. 5.

[0111]

[0114] The DASH packager 284 provides segments, along with the MPD, to the ROUTE sender 286 and the CDN origin server 288. The ROUTE sender 286 and the CDN origin server 288 may correspond to the server device 60 of FIG. 1 or the CDN 220 of FIG. 5. Generally, the ROUTE sender 286 can send media data to the ROUTE receiver 290 according to ROUTE in this example. In other examples, other file-based delivery protocols, such as FLUTE, can be used for broadcast or multicast. Additionally or alternatively, the CDN origin server 288 can send media data to the ROUTE receiver 290 and / or directly to the DASH client 292, for example, according to HTTP.

[0112]

[0115] The ROUTE receiver 290 can be implemented in middleware such as the eMBMS middleware unit 100 of FIG. 2. The ROUTE receiver 290 can buffer received media data, for example, in the cache 104 shown in FIG. 2. The DASH client 292 (which may correspond to the DASH client 110 of FIG. 2) can retrieve the cached media data from the ROUTE receiver 290 using HTTP. Alternatively, the DASH client 292 can directly retrieve media data from the CDN origin server 288 according to HTTP as described above.

[0113]

[0116] Furthermore, according to the techniques of the present disclosure, the DASH client 292 may use a manifest file such as an MPD to determine the location of the SAP or RAP, for example, following a resynchronization point signaled in the manifest file. The DASH client 292 may begin retrieving a media presentation starting from the next earliest resynchronization point. A resynchronization point may generally indicate the location in the bitstream where file container level data can be correctly parsed. Thus, the DASH client 292 may begin streaming starting at the resynchronization point and deliver the received media data starting from the resynchronization point to the CMAF / FF parser 294.

[0114]

[0117] The CMAF / FF parser 294 can begin parsing the media data starting from the resynchronization point. The CMAF / FF parser 294 may correspond to the decapsulation unit 50 of FIG. 1. Further, the CMAF / FF parser 294 may extract decodable media data from the parsed data and deliver the decodable media data to a media decoder 296 that may correspond to the audio decoder 46 or the video decoder 48 of FIG. 1. The media decoder 296 may decode the media data and deliver the decoded media data to a corresponding output device such as the audio output 42 or the video output 44 of FIG. 1.

[0115]

[0118] An example of the combination of DASH / CMAF and ROUTE in a broadcast scenario is shown in FIG. 7. In a combination of the low-latency DASH mode and ROUTE (e.g., as considered for the DVB TM-IPI task force in ABR multicast and the ATSC profile), the following problem may occur. If the ROUTE receiver 290 participates in the middle of a DASH / CMAF low-latency segment, synchronization is not available and there is no random access for other purposes, so data processing cannot be started. Thus, even if more frequent random access is provided in the middle of the segment, the start-up is delayed.

[0116]

[0119] Suitable solutions may be as follows. · Broadcast / multicast representations include more frequent random access points and non-initial chunks within segments / fragments. · The DASH client 292 determines that such a random access method exists using the information in the MPD and / or, in some cases, information from the ROUTE receiver 290. The DASH client 292 may or may not accurately use such information to locate the random access points. · The DASH client 292 can access this representation at startup, but may not be able to access all information from the start. · When starting to access the received portion of a segment, the DASH client 292 can find a random access point and start processing the data together with the same downloaded initialization segment / CMAF header of the same representation. The location of the random access point will be described below.

[0117]

[0120] However, the latter approach may encounter various problems, as summarized below.

[0118]

[0121] In a case similar to that described in the second use case above, not only random access resynchronization but also packet loss can be a problem. In this exemplary third use case, the same procedures as those described above may be applied. In addition to attempting a clean random access, after sufficient box parsing becomes possible, events, decoding, and presentation in non-random access chunks (e.g., without IDR frames) may also be attempted. Therefore, resynchronization not only to a clean random access but also to a random access for file format parsing is important.

[0119]

[0122] Yet another fourth use case can occur typically when live media content is delivered with low latency and then the same media content is used with time shifting for delayed playback. A client may want to access a media presentation at a specific time which may not coincide (generally does not coincide) with the start of a segment / CMAF fragment.

[0120]

[0123] Appropriate solutions can be the following. · At least one representation of an adaptation set may include more frequent random access points and non-initial chunks in a segment / fragment. · The DASH client 292 may use the information from the MPD to determine that such a random access method exists, but the location / byte offset of the random access points may not be exactly known. · The DASH client 292 can access this representation during seeking, but only downloads starting from the byte range of the latest available non-initial chunk or at least a byte range close to it may be permitted. · Once downloaded, the DASH client 292 may find a random access point and start processing the data together with the initialization segment / CMAF header of the same representation that was also downloaded. The location identification of the random access point will be described below.

[0121]

[0124] However, the latter approach may encounter various problems as summarized below.

[0122]

[0125] Resynchronization of ISO BMFF / DASH / CMAF segment cases generally involves multiple processes summarized as follows. 1) Finding the box structure. 2) Finding the CMAF chunk / fragment with all relevant information. 3) Finding the timing via mdat and tfdt. 4) If applicable, obtain all decoding-related information. 5) In some cases, process event messages. 6) Start decoding at the elementary stream level.

[0123]

[0126] An exemplary method of finding resynchronization points in a box structure at a specific time is summarized below. · If there is a segment index (SIDX box), such resynchronization points are provided as presentation time and byte offset. However, since the segment is not fully formed in advance, the segment index is typically not available for low-latency live. · If the start of the segment is available, the client can download the minimum set of byte ranges so that the box structure can be processed. · Resynchronization is provided by a protocol that, for example, provides that chunk boundaries are signaled and the client can start parsing. · If the start of the segment cannot be easily determined through the signaled data, the client can find a synchronization pattern that allows the client to access the data randomly. The client can then start parsing and find an appropriate box structure that allows processing, for example, of emsg, prft, mdat, moof, and / or mdat.

[0124]

[0127] This disclosure describes techniques applicable to the fourth example above. The first three represent exemplary simplifications when the corresponding information is available.

[0125]

[0128] This disclosure recognizes the following problems based on the above description and that these problems require solutions. 1) Adding additional random access points in DASH / CMAF segments. Random access can include clean random access and open or progressive decoder refreshes until it only provides resynchronization during file format parsing. 2) Adding appropriate signaling in the MPD (or other manifest file) that indicates the availability of random access points and resynchronization in each DASH segment, and provides information about the location, type, and timing of the random access points. The information can be accurate or within a range. 3) The ability to resynchronize decapsulation, decoding, and deciphering by finding resynchronization points for any starting point. 4) The ability to start processing in a limited receiver environment, for example, as available in HTML-5 / MSE-based playback.

[0126]

[0129] FIG. 8 is a conceptual diagram showing exemplary signaling of stream access points (SAPs) in a manifest file. In particular, FIG. 8 shows a bitstream 300 including SAPs 302A-302D (SAP 302) and segments 304A-304D (segment 304), and a bitstream 310 including SAPs 312A-312D (SAP 312), SAPs 316A-316D (SAP 316), and segments 314A-314D (segment 314). That is, in this example, segment 314 of bitstream 310 includes more frequent SAPs 312, 316 than segment 304 of bitstream 300. Each of SAPs 302, 312 can correspond to both the start of a corresponding one of segments 304, 314 and the first chunk of these segments. SAP 316 can correspond to the start of a chunk within the corresponding segment 316, but not to the start of the corresponding segment 316.

[0127]

[0130] To provide a simple technique for achieving a constant bitrate representation using equidistant chunks of 1000 samples (and @timescale=1000 in the sample duration) and an SAP type 1, which can be, for example, an audio representation, a resynchronization element can be added as follows.

[0128]

Number

[0129]

[0131] A client that receives such information, for example, the client device 40 in FIG. 1, may not be able to identify that for a segment with @duration=10000, the random access points can be accessed per second within the exact byte range. When the bitrate is variable, the receiver (e.g., the client device 40) can use @dIMin and @dIMax to identify the range within which it should look for random access points. As an alternative to @dT signaling the maximum value, it can also signal the nominal chunk duration.

[0130]

[0132] The resynchronization element of the manifest file may also include a URL @index that points to the binary resynchronization index of the resynchronization points in each segment, using the same template function as for normal segments. This resynchronization, if it exists, can provide the exact positions of all resynchronization points in the segment, similar to the segment index. If this index exists, the resynchronization index can be available for all segments during the period that is available at the issuance time of the manifest file / MPD.

[0131]

[0133] In one approach, the resynchronization index can be the same as the segment index, but it may be modified.

[0132]

[0134] The client device 40 of FIG. 1 may use the ISO BMFF 4-character box type as a basis for resynchronizing to a media file (e.g., video file 150 of FIG. 4 which may be a segment). In one example, the selected box type is the "styp" box, although it may also be the "moof" box itself. Random emulation of box sequence types is extremely rare. A test report of styp emulation is described below. This emulation is then avoided by checking against known expected box types. The client device 40 may execute a resynchronization mechanism outlined as follows. 1) For example, at byte offset B1, find the occurrence of the "styp" byte sequence in the segment. 2) Verify against random emulation as follows. The next box type is compared to a list of expected box types, namely, "styp", "sidx", "ssix", "prft", "moof", "mdat", "free", "mfra", "skip", "meta", "meco".

[0133] a. If one of the known box types is found, the byte offset B1 - 4 bytes is the byte offset of the resynchronization point.

[0134] b. If this is not one of the aforementioned known box types, this occurrence of the styp box is considered an invalid synchronization point and is ignored. Resume from step 1 above.

[0135]

[0135] The techniques of this disclosure were tested on 30,282 segments from scanned DASH-IF test assets. This scan revealed 28,408 occurrences of the "styp" column in the file, and only 10 of these 28,408 occurrences (about 1 out of 2,840 occurrences) were determined to be emulations that were discarded if the next box was not of the box types expected, namely, one of "styp", "sidx", "ssix", "prft", "moof", "mdat", "free", "mfra", "skip", "meta", "meco".

[0136]

[0136] Based on these results, it is considered sufficient to use styp resynchronization detection along with the chunk structure. It is appropriate to limit this to only a subset of the boxes that can follow styp, such as prft, emsg, free, skip, and moof.

[0137]

[0137] The remaining problem is the determination of the SAP type and the earliest presentation time. The latter is easily achieved by using the tfdt and other information in the movie fragment header. It is appropriate to document the algorithm.

[0138]

[0138] As follows, there are several options for determining the SAP type. · Detection based on information in moof. A simple technique can be documented and implemented. · Use of compatibility brands in the SAP type. By already using CMAF, the following can be inferred.

[0139] ○ cmff: Indicates that the SAP is 1 or 2 ○ cmfl: Indicates that the SAP is 0 (Is this correct regarding decoding?) ○ cmfr: Indicates that the SAP is 1, 2, or 3 · This signaling, if used consistently, may be sufficient. Compatibility brands for other SAP types can be defined. · Other techniques may be used to indicate the SAP type.

[0140]

[0139] Existing options may be used to determine the SAP type.

[0141]

[0140] In this way, the techniques of the present disclosure may be summarized as follows and may be executed by devices such as the content creation device 20, the server device 60, and / or the client device 40 of FIG. 1 as described above.

[0142]

[0141] In the DASH context, in some cases, a segment is treated as a single unit for downloading and accessing a media presentation and is also addressed by a specified URL. However, a segment may be structured to allow for resynchronization at the container level and random access to each representation within the segment. The resynchronization mechanism is supported and signaled by resynchronization elements.

[0143]

[0142] A resynchronization element signals a resynchronization point in a segment. A resynchronization point is the start of a chunk (at a byte position), where a chunk is defined as a structured consecutive byte range within a segment that contains media data for a specific presentation duration and can be independently accessed on a container format including the possibility of decoding. The resynchronization points in a segment may be defined as follows. · A resynchronization point is the start of a chunk. · Further, a resynchronization point assigns the following properties.

[0144] ○ It has a byte offset or index value from the start of the segment that points to the first byte of the chunk.

[0145] ○ It has the earliest presentation time assigned in the representation.

[0146] ○ It has an assigned SAP type, which is defined, for example, by the SAP type in ISO / IEC 14496-12.

[0147] ○ It has an assigned marker property indicating whether a resynchronization point can be detected during segment analysis through a specific marker, or whether a resynchronization point needs to be signaled by external means. · Starting processing from a resynchronization point enables container parsing and decoding, along with the information in the initialization segment, if present. The ability to access the contained elementary video stream, and how to access it, is defined by the SAP type.

[0148]

[0143] Signaling each resynchronization point in the MPD can be difficult for causal reasons, as resynchronization points can be added by the segment packager independently of MPD updates. For example, resynchronization points can be generated by the encoder and packager independently of the MPD. Also, in low latency, MPD signaling may not be available to DASH clients, such as the DASH client 110 in Figure 2 or the DASH client 292 in Figure 7. Therefore, there are two ways to signal resynchronization points provided in segments in the MPD. · By providing a binary map regarding the resynchronization points in the resynchronization index segment of each segment. This is most easily used for segments that are fully available on the network. · By signaling the presence of resynchronization points in the segment, along with some additional information that enables easy finding of the resynchronization points regarding byte position and presentation time.

[0149]

[0144] To signal the above characteristics, the resynchronization element has different attributes, which are described in more detail in Section 5.3.12.2 of the DASH specification.

[0150]

[0145] Random access, if present, initializes the representation using the initialization segment and starts processing, decrypting, and presenting the representation from the random access point after the signaled segment by decrypting and presenting the representation from the random access point after time t. The random access point can be signaled using the RandomAccess element as defined in Table 10 below.

[0151]

Table 2

[0152]

[0146] Table 11 provides various random access point types.

[0153]

Table 3

[0154]

[0147] The resynchronization index segment contains information related to the media segment. The resynchronization index segment, similar to the segment index, provides the exact positions of all resynchronization points in the segment. The resynchronization points are defined in Section 5.3.12.1 of the DASH specification.

[0155]

[0148] The resynchronization points of ISO BMFF can be defined as the start of an ISO BMFF segment having the following restrictions with respect to both cardinality and ordinality.

[0156]

Table 4

[0157]

[0149] For ISO BMFF-based resynchronization points, the properties can be defined as follows. · The index Index is defined as the offset of the first byte of the constrained ISO BMFF segment as described above. · The earliest presentation time, Time, is defined as the minimum of the decoding time of any sample in the chunk, the composition offset, and the combination of the edit list. · The SAP type is defined according to Section 4.5.2 of the DASH specification. · If "styp" exists as the main compatibility brand together with "cmfl", the marker exists.

[0158]

[0150] The resynchronization index segment can index one media segment of one representation and can be defined as follows. · Each representation index segment should start with a "styp" box, and the brand "risg" should exist in the "styp" box. The compliance requirements for the brand "risg" are defined by this subclause. · Each media segment is indexed by one or more segment index boxes, and the boxes for a given media segment are consecutive.

[0159]

[0151] FIG. 9 is a flowchart showing an exemplary method of retrieving media data according to the techniques of the present disclosure. The method of FIG. 9 is described with respect to the client device 40 of FIG. 1. However, the client device including the low-latency DASH client 232 of FIG. 5, or the media decoder 296, the CMAF / FF parser 294, the DASH client 292, and the ROUTE receiver 290 of FIG. 7 may also be configured to perform this method or a similar method.

[0160]

[0152] First, client device 40 can retrieve a manifest file, such as an MPD for a media presentation (350). The client device 40 can retrieve the manifest file from, for example, the server device 60. The manifest file can include data indicating that the media presentation includes resynchronization points at chunk boundaries within segments of the media presentation. Thus, the client device 40 can determine a resynchronization point for the media presentation, for example, the most recently available resynchronization point (352). Generally, a resynchronization point can indicate the start of a chunk boundary and is a randomly accessible point in the representation where a file-level container (e.g., a data structure such as a box as described above) can be properly parsed.

[0161]

[0153] In particular, the manifest file can indicate the location of a resynchronization point, such as a byte offset from the start of a segment. This information cannot precisely identify the location of the resynchronization point within the segment, but can guarantee that the resynchronization point is available within a byte range from the byte offset. Thus, the client device 40 can form a request, such as an HTTP partial Get request, that specifies the indicated byte offset to begin retrieval at the resynchronization point (354). The client device 40 can then send the request to the server device 60 (356).

[0162]

[0154] In response to the request, the client device 40 can receive the requested media data including resynchronization points (358). As described above, the byte offset may not accurately identify the location of the resynchronization point, and thus, the client device 40 may analyze the data until it detects the actual location of the resynchronization point. The client device 40 may start at the resynchronization point and analyze file-level data structures such as file format boxes to determine the location of the corresponding chunk of the retrieved media data. Specifically, the client device 40 may identify the resynchronization point as the start of a chunk by detecting, for example, a segment type value, a generator reference time value, an event message, a movie fragment, and a media data container box. The movie fragment may include encoded media data.

[0163]

[0155] The decapsulation unit 50 can extract the encoded media data of the corresponding chunk from, for example, the movie fragment (360) and provide the encoded media data to, for example, the video decoder 48. The chunk may start with a random access point (RAP) such as an intra prediction frame (I-frame) of the video data. The manifest file further indicates whether the RAP is the start of a closed picture group (GOP) or an open GOP, thereby indicating the type of random access that can be performed starting at the RAP (e.g., whether the leading picture of the I-frame is decodable or not). The video decoder 48 can then decode the encoded media data (362) and send the decoded media data to, for example, the video output 44 to present the decoded media data (364).

[0164]

[0156] In this way, the method of FIG. 9 involves retrieving a media presentation manifest file indicating that container parsing of the media data of the bitstream can be started at the resynchronization point of the segments of the media presentation representation, where the resynchronization point is at a position other than the start of the segment and represents the point at which container parsing of the media data of the bitstream can be started, using the manifest file to form a request to retrieve the media data of the representation starting at the resynchronization point, sending a request to start retrieving the media data of the media presentation starting at the resynchronization point, and presenting the retrieved media data, which represents an example of a method for retrieving media data.

[0165]

[0157] Some techniques of the present disclosure are summarized in the following examples.

[0166]

[0158] Example 1: A method for retrieving media data, comprising retrieving a media presentation manifest file indicating that resynchronization and decoding can be started at the resynchronization point of the representation of the media presentation, retrieving the media data of the representation starting at the resynchronization point, and presenting the retrieved media data.

[0167]

[0159] Example 2: The method of Example 1, wherein the resynchronization point comprises the start of a chunk boundary.

[0168]

[0160] Example 3: The method of Example 2, wherein the chunk boundary comprises the start of a chunk comprising zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.

[0169]

[0161] Example 4: The manifest file of any of Examples 1 to 3, which indicates the availability of the resynchronization point in the segment of the representation.

[0170]

[0162] Example 5: The resynchronization point is at a position other than the start of the segment, the method of Example 4.

[0171]

[0163] Example 6: The manifest file indicates the type of random access that can be performed at the resynchronization point, the method of either Example 4 or 5.

[0172]

[0164] Example 7: The manifest file indicates the position and timing of the resynchronization point and whether the position and timing information is accurate or estimated, the method of any one of Examples 4 to 6.

[0173]

[0165] Example 8: The manifest file includes a media presentation description (MPD), the method of any one of Examples 1 to 7.

[0174]

[0166] Example 9: A device for retrieving media data, comprising one or more means for performing the method of any one of Examples 1 to 8.

[0175]

[0167] Example 10: The one or more means comprise one or more processors implemented in a circuit and a memory configured to store media data, the device of Example 9.

[0176]

[0168] Example 11: The device of Example 9, comprising at least one of an integrated circuit, a microprocessor, or a wireless communication device.

[0177]

[0169] Example 12: A computer-readable storage medium storing instructions that, when executed, cause a processor to execute the method of any one of Examples 1 to 8.

[0178]

[0170] Example 13: A device for retrieving media data, comprising means for retrieving a manifest file of a media presentation indicating that resynchronization and decoding can be started at a resynchronization point of the presentation of the media presentation, means for retrieving media data of the presentation starting at the resynchronization point, and means for presenting the retrieved media data.

[0179]

[0171] Example 14: A method for sending media data, comprising sending a manifest file of a media presentation indicating that resynchronization and decoding can be started at a resynchronization point of the presentation of the media presentation to a client device, receiving a request for media data starting at the resynchronization point from the client device, and in response to the request, sending the requested media data of the presentation starting at the resynchronization point to the client device.

[0180]

[0172] Example 15: The method of Example 14, further comprising generating a manifest file.

[0181]

[0173] Example 16: The method according to any one of Examples 14 and 15, wherein the resynchronization point comprises the start of a chunk boundary.

[0182]

[0174] Example 17: The method of Example 16, wherein the chunk boundary comprises the start of a chunk comprising zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.

[0183]

[0175] Example 18: The method according to any one of Examples 14 to 17, wherein the manifest file indicates the availability of resynchronization points in segments of the presentation.

[0184]

[0176] Example 19: The method of Example 18, wherein the resynchronization point is at a position other than the start of the segment.

[0185]

[0177] Example 20: The manifest file is any of the methods of Examples 18 and 19 that indicate the type of random access that can be performed at the resynchronization point.

[0186]

[0178] Example 21: The manifest file is any of the methods of Examples 18 to 20 that indicate the position and timing of the resynchronization point and whether the position and timing information is accurate or estimated.

[0187]

[0179] Example 22: The manifest file comprises a Media Presentation Description (MPD), any of the methods of Examples 14 to 21.

[0188]

[0180] Example 23: A device for sending media data, comprising one or more means for performing any of the methods of Examples 14 to 22.

[0189]

[0181] Example 24: The one or more means comprise one or more processors implemented in a circuit and a memory configured to store media data, the device of Example 23.

[0190]

[0182] Example 25: The device of Example 23, comprising at least one of an integrated circuit, a microprocessor, or a wireless communication device.

[0191]

[0183] Example 26: A computer-readable storage medium storing instructions that, when executed, cause a processor to execute any of the methods of Examples 1 to 8.

[0192]

[0184] Example 27: A device for sending media data, comprising means for sending to a client device a manifest file of a media presentation indicating that resynchronization and decoding can be started at a resynchronization point of the presentation of the media presentation; means for receiving from the client device a request for media data starting at the resynchronization point; and means for sending, in response to the request, the requested media data of the presentation starting at the resynchronization point to the client device.

[0193]

[0185] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or may include a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, the computer-readable medium generally corresponds to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0194] By way of example and not limitation, such a computer-readable storage medium can comprise RAM, ROM, EEPROM (registered trademark), CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Further, any connection can appropriately be called a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that the computer-readable storage medium and data storage medium are directed to non-transitory, tangible storage media instead of including connections, carrier waves, signals, or other transitory media. As used herein, disk and disc include compact disc (CD), laser disc (registered trademark), optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray (registered trademark) disc, where disk typically magnetically reproduces data and disc optically reproduces data with a laser. Combinations of the above should also be included within the scope of computer-readable media.

[0195]

[0187] The commands can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated circuits or discrete logic circuits. Thus, the term "processor" as used herein may refer to either the foregoing structures, or any other structure suitable for implementation of the techniques described herein. Further, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Also, the techniques may be fully implemented with one or more circuits or logic elements.

[0196]

[0188] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or sets of ICs (e.g., chip sets). Although various components, modules, or units have been described herein to emphasize the functional aspects of devices configured to execute the disclosed techniques, those components, modules, or units need not necessarily be realized by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit, including one or more of the processors described above, along with suitable software and / or firmware, or provided by a set of interoperable hardware units.

[0197]

[0189] A variety of examples have been described. These and other examples fall within the scope of the following claims.

Claims

1. A method for retrieving media data, comprising: retrieving a manifest file of the media presentation, indicating that container analysis of the media data of the bitstream can be started at a resynchronization point of a segment of the representation of the media presentation, wherein the resynchronization point is at a position other than the start of the segment, and represents a point at which the container analysis of the media data of the bitstream can be started; using the manifest file to form a request to retrieve the media data of the representation starting at the resynchronization point; sending the request to start retrieving the media data of the media presentation starting at the resynchronization point; and presenting the retrieved media data. A method comprising the above steps.

2. The method according to claim 1, wherein presenting the retrieved media data comprises analyzing a file-level media data container of the retrieved media data at the resynchronization point.

3. Analyzing comprises: analyzing the file-level media data container until a random access point (RAP) of the media presentation is detected; and sending the RAP to a media decoder. The method according to claim 2, comprising the above steps.

4. The method according to claim 1, wherein the resynchronization point comprises the start of a chunk boundary.

5. The method according to claim 2, wherein the chunk boundary comprises the start of a chunk comprising zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.

6. The method according to claim 1, wherein the manifest file indicates the availability of the resynchronization point in the segment of the representation.

7. The method according to claim 6, wherein the resynchronization point is at a position other than the start of the segment.

8. The method according to claim 6, wherein the manifest file indicates the type of random access that can be performed at the resynchronization point.

9. The method according to claim 6, wherein the manifest file indicates the position and timing of the resynchronization point, and whether the position and timing information is accurate or estimated.

10. The method according to claim 1, wherein the manifest file comprises a Media Presentation Description (MPD).

11. A device for retrieving media data, comprising: a memory configured to store media data of a media presentation; retrieving the manifest file of the media presentation indicating that container analysis of the media data of the bitstream can be started at a resynchronization point of a segment of the representation of the media presentation, wherein the resynchronization point is at a position other than the start of the segment and represents a point at which the container analysis of the media data of the bitstream can be started; using the manifest file to form a request to retrieve the media data of the representation starting at the resynchronization point; sending the request to start retrieving the media data of the media presentation starting at the resynchronization point; and presenting the retrieved media data; one or more processors implemented in a circuit and configured to perform the above; A device comprising the above.

12. The device according to claim 11, wherein, to present the retrieved media data, the one or more processors are configured to analyze a file-level media data container of the retrieved media data at the resynchronization point.

13. To analyze the file-level media data container, the one or more processors are configured to: analyze the file-level media data container until a Random Access Point (RAP) of the media presentation is detected; send the RAP to a media decoder; The device according to claim 12, configured to perform the above.

14. The device according to claim 11, wherein the resynchronization point comprises the start of a chunk boundary.

15. The device according to claim 14, wherein the chunk boundary comprises the start of a chunk comprising zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.

16. The device according to claim 11, wherein the manifest file indicates the availability of the resynchronization point in the segment of the representation.

17. The device according to claim 16, wherein the resynchronization point is at a position other than the start of the segment.

18. The device according to claim 16, wherein the manifest file indicates the type of random access that can be performed at the resynchronization point.

19. The device according to claim 16, wherein the manifest file indicates the position and timing of the resynchronization point and whether the position and timing information is accurate or estimated.

20. The device according to claim 11, wherein the manifest file comprises a media presentation description (MPD).

21. A computer-readable storage medium storing instructions that, when executed, cause a processor to retrieve the manifest file of the media presentation indicating that container parsing of the media data of the bitstream can be started at a resynchronization point of a segment of the representation of the media presentation, wherein the resynchronization point is at a position other than the start of the segment and represents a point at which the container parsing of the media data of the bitstream can be started, use the manifest file to form a request to retrieve the media data of the representation starting at the resynchronization point, send the request to start retrieving the media data of the media presentation starting at the resynchronization point, and present the retrieved media data. A computer-readable storage medium.

22. The computer-readable storage medium according to claim 21, wherein the resynchronization point comprises the start of a chunk boundary.

23. The computer-readable storage medium according to claim 22, wherein the chunk boundary comprises the start of a chunk comprising zero or one segment type value, zero or one generator reference time value, zero or more event messages, at least one movie fragment box, and at least one media data container box.

24. The computer-readable storage medium according to claim 21, wherein the manifest file indicates the availability of the resynchronization point in the segment of the representation.

25. The computer-readable storage medium according to claim 24, wherein the resynchronization point is located at a position other than the start of the segment.

26. The computer-readable storage medium according to claim 24, wherein the manifest file indicates the type of random access that can be executed at the resynchronization point.

27. The computer-readable storage medium according to claim 24, wherein the manifest file indicates the position and timing of the resynchronization point and whether the position and timing information is accurate or estimated.

28. The computer-readable storage medium according to claim 21, wherein the manifest file includes a media presentation description (MPD).

29. A device for retrieving media data, means for retrieving the manifest file of the media presentation, which indicates that container analysis of the media data of the bitstream can be started at a resynchronization point of a segment of the representation of the media presentation, wherein the resynchronization point is at a position other than the start of the segment and represents a point at which the container analysis of the media data of the bitstream can be started; means for using the manifest file to form a request to retrieve the media data of the representation starting at the resynchronization point; means for sending the request to start retrieving the media data of the media presentation starting at the resynchronization point; and means for presenting the retrieved media data A device comprising.

Citation Information

Patent Citations

  • Media representation groups for network streaming of coded video data

    JP2015111898A

  • Content transmission device, content reproduction device, content distribution system, control method for content transmission device, control method for content reproduction device, control program, and recording medium

    JP2015208018A

  • Locating and accessing segment chunks for media streaming

    JP2019523600A

  • System-Level Signaling of SEI Tracks for Media Data Streaming

    JP2019525677A