Streaming includes media data with addressable resource index tracks with switching sets

By introducing Addressable Resource Index (ARI) track technology, the problem of inaccurate addressable resource description in video data transmission is solved, enabling more flexible and efficient streaming transmission, adapting to different network conditions and operating modes, and optimizing media data acquisition and switching on client devices.

CN115943631BActive Publication Date: 2026-01-23QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180045138.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-29
Filing Date
2021-06-30
Publication Date
2026-01-23
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

Existing video data transmission technologies are not accurate or flexible enough in describing the duration and size of addressable resources, making it difficult for client devices to effectively schedule the download and switching of media data when network conditions change, thus affecting the quality and efficiency of streaming.

Method used

By employing Addressable Resource Index (ARI) track technology, more accurate duration and size information is provided by describing the details of the switching set and addressable resources presented in the media presentation, enabling client devices to dynamically switch and schedule media data under different network conditions.

Benefits of technology

It improves the flexibility and efficiency of streaming transmission, ensuring that client devices can accurately acquire and switch media data in different operating modes such as low-latency live streaming, live broadcasting, and video-on-demand, optimizing network bandwidth usage and reducing latency and data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115943631B_ABST
    Figure CN115943631B_ABST
Patent Text Reader

Abstract

An example device for retrieving media data includes a memory configured to store media data; and one or more processors implemented in circuitry and configured to retrieve data of an addressable resource information (ARI) track of a media presentation, the data of the ARI track describing addressable resources and subsets of a switching set of the media presentation, the switching set comprising a plurality of media tracks, the media tracks comprising addressable resources, the ARI track being a single index track of the media presentation, the addressable resources comprising retrievable media data; determine, from the data of the ARI track, durations and sizes of the addressable resources; determine, using the data of the ARI track including the durations and sizes of the addressable resources, one or more addressable resources to retrieve; retrieve the determined addressable resources; and store the retrieved addressable resources in the memory.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefits of U.S. Application No. 17 / 362,673, filed June 29, 2021, and U.S. Provisional Application No. 63 / 047,153, filed July 1, 2020, the entire contents of each of which are incorporated herein by reference. U.S. Application No. 17 / 362,673 claims the benefit of U.S. Provisional Application No. 63 / 047,153, filed July 1, 2020. Technical Field

[0003] This disclosure relates to the storage and transmission of coded video data. Background Technology

[0004] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones, and video conferencing equipment. Digital video devices implement video compression technologies, such as those defined in MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4 Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)), and extensions to these standards, to enable more efficient transmission and reception of digital video information.

[0005] Video compression techniques perform spatial and / or temporal prediction to reduce or remove inherent redundancy in video sequences. For block-based video decoding, video frames or slices can be divided into macroblocks. Each macroblock can be further subdivided. Macroblocks in intra-frame decoded (I) frames or slices are encoded using spatial prediction relative to adjacent macroblocks. Macroblocks in inter-frame decoded (P or B) frames or slices can use spatial prediction with respect to adjacent macroblocks in the same frame or slice, or temporal prediction with respect to other reference frames.

[0006] After video data has been encoded, it can be packetized for transmission or storage. Video data can be assembled into video files conforming to any of several standards, such as the International Organization for Standardization (ISO) Basic Media File Format and its extensions, such as AVC. Summary of the Invention

[0007] In summary, this disclosure describes techniques for streaming media data including addressable resource index (ARI) tracks. Streaming media data typically involves streaming multiple files that together form a media presentation (e.g., a full-length movie). Files may include multiple tracks, each of which may include, for example, video data, signaling data, cue data, etc. According to the techniques of this disclosure, an ARI track can describe details of addressable resources and / or subsets of a Common Media Application Format (CMAF) switch set within a single index track. Addressable resources can typically be individual, retrievable media datasets, such as track files, segments, or chunks in a CMAF.

[0008] In one example, a method for retrieving media data includes: retrieving data from an Addressable Resource Information (ARI) track of media presentation, the ARI track data describing a subset of a switchset of media presentation and addressable resources, the switchset including multiple media tracks, each media track including the addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; determining the duration and size of the addressable resources based on the ARI track data; using the ARI track data including the duration and size of the addressable resources to determine one or more of the addressable resources to be retrieved; and retrieving the determined addressable resources.

[0009] In another example, an apparatus for retrieving media data includes: a memory configured to store the media data; and one or more processors implemented in circuitry and configured to: retrieve data of an Addressable Resource Information (ARI) track for media presentation, the ARI track data describing a subset of a switch set for media presentation and addressable resources, the switch set including multiple media tracks, each media track including addressable resources, the ARI track being a single index track for media presentation, the addressable resources including retrievable media data; determine the duration and size of the addressable resources based on the ARI track data; determine one or more addressable resources to be retrieved using the ARI track data including the duration and size of the addressable resources; retrieve the determined addressable resources; and store the retrieved addressable resources in the memory.

[0010] In another example, a computer-readable storage medium stores instructions thereon that, when executed, cause the processor to: retrieve data from an Addressable Resource Information (ARI) track of media presentation, the data of which describes a subset of a switching set of media presentation and addressable resources, the switching set comprising multiple media tracks, each media track comprising addressable resources, the ARI track being a single index track of media presentation, and the addressable resources comprising retrievable media data; determine the duration and size of the addressable resources based on the data from the ARI tracks; determine one or more addressable resources to be retrieved using the data from the ARI tracks, which includes the duration and size of the addressable resources; and retrieve the determined addressable resources.

[0011] In another example, the device for retrieving media data includes: a unit for retrieving data of an Addressable Resource Information (ARI) track for media presentation, the data of the ARI track describing a subset of a switchset and addressable resources of the media presentation, the switchset including multiple media tracks, each media track including addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; a unit for determining the duration and size of the addressable resources based on the data of the ARI track; a unit for determining one or more of the addressable resources to be retrieved using the data of the ARI track including the duration and size of the addressable resources; and a unit for retrieving the determined addressable resources.

[0012] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, the drawings, and the claims. Attached Figure Description

[0013] Figure 1 This is a block diagram illustrating an example system for implementing a technology for streaming media data over a network.

[0014] Figure 2 This explains Figure 1 A block diagram of an example set of components for the retrieval unit.

[0015] Figure 3 This is a conceptual diagram illustrating the elements of example multimedia content.

[0016] Figure 4 It is a block diagram illustrating the elements of the example video file, which can correspond to the segments represented.

[0017] Figure 5 This is a conceptual diagram illustrating an example system for performing low-latency streaming using chunking and addressable resource index (ARI) tracks, based on the technology disclosed herein.

[0018] Figure 6 This is a conceptual diagram illustrating an example dataset containing ARI orbitals according to the technology disclosed herein.

[0019] Figure 7 This is a conceptual diagram illustrating examples of boxes that can be included in HTTP-based Dynamic Adaptive Streaming (DASH) segments and chunks.

[0020] Figure 8 This is a flowchart illustrating an example method for providing addressable resource information (ARI) tracks from a server device to a client device according to the technology of this disclosure.

[0021] Figure 9 This is a flowchart illustrating an example method for retrieving media data using Addressable Resource Information (ARI) tracks according to the technology disclosed herein. Detailed Implementation

[0022] In summary, this disclosure describes techniques for streaming media data using Addressable Resource Index (ARI) tracks. Adaptive streaming clients (such as HTTP-based Dynamic Adaptive Streaming (DASH) clients may require information accurately describing the duration and size of addressable resources on a server device, as well as possible subsets of addressable resources. Addressable resources can be track files, segments, or chunks of Common Media Application Format (CMAF) or other similar elements of DASH, High-Level Streaming (HLS) or other streaming protocols and media formats.

[0023] The Segment Indexing (SIDX) box for media files can provide a precise mapping of information similar to that discussed above for on-demand services. For live streaming services, it can provide minimal buffer and bandwidth pairings for signaling. Resync index segments can be used to provide a localized index for each segment.

[0024] This disclosure recognizes the potential benefits of creating additional information about the estimated or accurate bit rate (duration and size) of addressable resources to optimize client operation. Client optimization may vary depending on the client's operating mode. Such operating modes may include low-latency live streaming, live broadcasting, time-shifting, video-on-demand (VoD), etc. Other considerations may include the client device's target latency, network conditions, and desired content quality.

[0025] The minimum buffer / bandwidth pair signaling discussed above can reflect inaccurate bit rates that the client device cannot use for scheduled downloads. While resynchronizing index segments is generally good practice, the signaling notification information is individual for specific segments and representations, resulting in numerous requests for all information across all segments. SIDX box information is typically only available for VoD because it records the entire file including the SIDX boxes.

[0026] The techniques disclosed herein regarding the use of ARI orbitals can provide a more general and flexible approach to addressing issues such as signaling the duration and size of addressable resources, and others.

[0027] The technology disclosed herein can be applied to video files containing video data encapsulated according to any of the ISO Basic Media File Format, Scalable Video Decoding (SVC) file format, Advanced Video Decoding (AVC) file format, 3rd Generation Partnership Project (3GPP) file format and / or Multi-View Video Decoding (MVC) file format or other similar video file formats.

[0028] In HTTP streaming, commonly used operations include HEAD, GET, and partial GET. The HEAD operation retrieves the header of a file associated with a given Uniform Resource Locator (URL) or Uniform Resource Name (URN), but not the payload associated with the URL or URN. The GET operation retrieves the entire file associated with a given URL or URN. The partial GET operation takes a range of bytes as input and retrieves the number of consecutive bytes in the file, where the number of bytes corresponds to the received range of bytes. Therefore, movie fragments can be provided for HTTP streaming, as a partial GET operation can obtain one or more individual movie fragments. A movie fragment can consist of multiple track segments from different tracks. In HTTP streaming, the media representation can be a structured collection of data accessible to the client. Clients can request and download media data information to present the streaming service to the user.

[0029] In examples of using HTTP streaming to stream 3GPP data, the video and / or audio data of multimedia content may have multiple representations. As described below, different representations may correspond to different decoding characteristics (e.g., different profiles or levels of the video decoding standard), different decoding standards or extensions to those standards (e.g., multi-view and / or scalable extensions), or different bitrates. This list of representations can be defined in a Media Presentation Description (MPD) data structure. A media presentation can correspond to a structured set of data accessible to the HTTP streaming client device. The HTTP streaming client device can request and download media data information to present the streaming service to the user of the client device. The media presentation can be described in the MPD data structure, which may include updates to the MPD.

[0030] A media presentation may consist of one or more periods. Each period may extend until the beginning of the next period, or, in the case of the last period, until the end of the media presentation. Each period may contain one or more representations of the same media content. A representation may be one of several alternative encoded versions of audio, video, timed text, or other such data. Representations may vary depending on the encoding type, such as the bitrate, resolution, and / or codec of video data, and the bitrate, language, and / or codec of audio data. The term "representation" can be used to refer to a portion of encoded audio or video data that corresponds to a specific period of multimedia content and is encoded in a specific manner.

[0031] A representation for a specific period can be assigned to a group indicated by an attribute in the MPD, which indicates the adaptation set to which the representation belongs. Representations within the same adaptation set are generally considered alternatives to each other because client devices can dynamically and seamlessly switch between these representations, for example, to perform bandwidth adaptation. For example, each representation of video data for a specific period can be assigned to the same adaptation set, allowing any representation to be selected to decode media data, such as video or audio data, for the corresponding period's multimedia content. In some examples, the media content within a period can be represented by one representation from group 0 (if present) or a combination of at most one representation from each non-zero group. Timing data for each representation of a period can be expressed relative to the start time of that period. According to the techniques of this disclosure, as discussed in more detail below, one or more adaptation sets may correspond to switch sets, such as Common Media Application Format (CMAF) switch sets. Similarly, each representation of one or more adaptation sets may correspond to a CMAF track.

[0032] A representation may include one or more segments. Each representation may include an initialization segment, or each segment of a representation may be self-initializing. When present, the initialization segment may contain initialization information for accessing the representation. Typically, the initialization segment does not contain media data. Segments can be uniquely referenced by identifiers such as Uniform Resource Locators (URLs), Uniform Resource Names (URNs), or Uniform Resource Identifiers (URIs). The MPD may provide an identifier for each segment. In some examples, the MPD may also provide byte ranges in the form of range attributes, which may correspond to data in segments within a file accessible via URL, URN, or URI.

[0033] Furthermore, each segment may include a corresponding plurality of media data chunks. Chunks may be retrievable, for example, using a partial HTTP GET request with a specified byte range for a given chunk. These chunks may be time-aligned between segments of different representations. Additionally, according to the techniques of this disclosure, an ARI track may include samples time-aligned with chunks of various segments and representations. Each sample of an ARI track may describe characteristics of the corresponding chunk, such as: whether the chunk forms the start of a new segment, the track identifier of the corresponding representation in the representation, and the tag signaling flags of each representation, the Stream Access Point (SAP) type, the offset of the corresponding segment start, and multiple pairs for minimum buffer / bandwidth prediction.

[0034] In this way, the client device can retrieve samples of ARI tracks and use the data from these samples to retrieve corresponding chunks of various representations and segments. For example, as described above, the client device can formulate an HTTP partial GET request to retrieve individual chunks of a segment. The client device can use data from samples of ARI tracks representing offsets relative to the start of the corresponding segment to determine the start and end bytes of the segment to include in the byte range of the HTTP partial GET request to retrieve a specific chunk. That is, for a specific chunk, the client device can determine the offset from the sample of the ARI track corresponding to the specific chunk to the start of the corresponding segment as the start of the byte range, and the offsets of adjacent samples of the ARI track in the same segment as the end of the byte range.

[0035] Different representations can be selected to substantially retrieve different types of media data simultaneously. For example, a client device can choose to retrieve segmented audio representations, video representations, and timed text representations. In some examples, the client device can select a specific set of adapters to perform bandwidth adaptation. That is, the client device can select an adapter set that includes video representations, an adapter set that includes audio representations, and / or an adapter set that includes timed text. Alternatively, the client device can select an adapter set for certain types of media (e.g., video) and directly select a representation for other types of media (e.g., audio and / or timed text).

[0036] Figure 1 This is a block diagram illustrating an example system 10 for implementing techniques for streaming media data over a network. In this example, system 10 includes a content preparation device 20, a server device 60, and a client device 40. Client device 40 and server device 60 are communicatively coupled via a network 74, which may include the Internet. In some examples, content preparation device 20 and server device 60 may also be coupled via network 74 or another network, or they may be directly communicatively coupled. In some examples, content preparation device 20 and server device 60 may include the same device.

[0037] Content preparation equipment 20 (in Figure 1 In the example, the audio source 22 and video source 24 are included. The audio source 22 may include, for example, a microphone that generates electrical signals representing captured audio data to be encoded by the audio encoder 26. Alternatively, the audio source 22 may include a storage medium storing previously recorded audio data, an audio data generator (e.g., a computerized synthesizer), or any other audio data source. The video source 24 may include a camera that generates video data to be encoded by the video encoder 28, a storage medium encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other video data source. The content preparation device 20 is not necessarily communicatively coupled to the server device 60 in all examples, but may instead store multimedia content on a separate medium that is read by the server device 60.

[0038] The raw audio and video data may include analog or digital data. Analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. Audio source 22 may acquire audio data from the speaker while the speaker is speaking, and video source 24 may acquire video data of the speaker simultaneously. In other examples, audio source 22 may include a computer-readable storage medium comprising stored audio data, and video source 24 may include a computer-readable storage medium comprising stored video data. In this way, the techniques described in this disclosure can be applied to live, streaming, real-time audio and video data or to archived, pre-recorded audio and video data.

[0039] An audio frame corresponding to a video frame is typically an audio frame containing audio data captured (or generated) by audio source 22 and video data captured (or generated) simultaneously by video source 24, which is included within the video frame. For example, while a speaker is typically generating audio data through speaking, audio source 22 captures the audio data, and video source 24 simultaneously captures the speaker's video data; that is, while audio source 22 is capturing audio data. Therefore, an audio frame can temporally correspond to one or more specific video frames. Thus, an audio frame corresponding to a video frame typically corresponds to the following situation: where audio and video data are captured simultaneously, and in this case, the audio frame and video frame respectively include the simultaneously captured audio and video data.

[0040] In some examples, audio encoder 26 may encode a timestamp in each encoded audio frame, representing the time at which the audio data of the encoded audio frame was recorded, and similarly, video encoder 28 may encode a timestamp in each encoded video frame, representing the time at which the video data of the encoded video frame was recorded. In such examples, the audio frame corresponding to the video frame may include an audio frame containing a timestamp and a video frame containing the same timestamp. Content preparation device 20 may include an internal clock from which audio encoder 26 and / or video encoder 28 may generate timestamps, or audio source 22 and video source 24 may use the internal clock to associate audio and video data with timestamps respectively.

[0041] In some examples, audio source 22 may send data to audio encoder 26 corresponding to the time the audio data was recorded, and video source 24 may send data to video encoder 28 corresponding to the time the video data was recorded. In some examples, audio encoder 26 may encode sequence identifiers in the encoded audio data to indicate the relative temporal order of the encoded audio data, but not necessarily the absolute time of the recorded audio data; similarly, video encoder 28 may use sequence identifiers to indicate the relative temporal order of the encoded video data. Similarly, in some examples, sequence identifiers may be mapped to or otherwise associated with timestamps.

[0042] Audio encoder 26 typically produces encoded audio data streams, while video encoder 28 produces encoded video data streams. Each individual data stream (whether audio or video) can be called an elementary stream. An elementary stream is a single digitally decoded (potentially compressed) component of a representation. For example, a represented decoded video or audio portion can be an elementary stream. Elementary streams can be converted into packetized elementary streams (PES) before being encapsulated within a video file. In the same representation, a stream ID can be used to distinguish PES packets belonging to one elementary stream from those belonging to another. The basic data unit of an elementary stream is a packetized elementary stream (PES). Therefore, decoded video data typically corresponds to an elementary video stream. Similarly, audio data corresponds to one or more corresponding elementary streams.

[0043] Many video decoding standards (such as ITU-T H.264 / AVC, ITU-T H.265 / High-Efficiency Video Decoding (HEVC), and the upcoming Universal Video Decoding (VVC) standard) define the syntax, semantics, and decoding process for error-free bitstreams, each conforming to a specific profile or level. Video decoding standards typically do not specify the encoder, but the encoder's task is to ensure that the generated bitstream conforms to the decoder's standard. In the context of video decoding standards, a "profile" corresponds to a subset of the algorithms, features, or tools and constraints applicable to them. For example, according to the H.264 standard, a "profile" is a subset of the entire bitstream syntax specified by the H.264 standard. A "level" corresponds to limitations on decoder resource consumption, such as decoder memory and computation, which are related to image resolution, bitrate, and block processing rate. A profile can be indicated by a `profile_idc` (profile indicator) value, while a level can be indicated by a `level_idc` (level indicator) value.

[0044] For example, the H.264 standard recognizes that, within the range of syntax imposed on a given profile, the performance of the encoder and decoder may still vary considerably depending on the values ​​of the syntactic elements in the bitstream, such as the specified size of the decoded images. The H.264 standard further recognizes that, in many applications, implementing a decoder capable of handling all assumptions about the syntax within a particular profile is neither practical nor economical. Therefore, the H.264 standard defines a “level” as a specified set of constraints imposed on the values ​​of syntactic elements in the bitstream. These constraints may be simple restrictions on the values. Alternatively, these constraints may take the form of constraints on an arithmetic combination of values ​​(e.g., image width multiplied by image height multiplied by the number of images decoded per second). The H.264 standard further specifies that different implementations may support different levels for each supported profile.

[0045] A profile-compliant decoder typically supports all features defined in the profile. For example, B-picture decoding, as a decoding feature, is not supported in the baseline H.264 / AVC profile but is supported in other H.264 / AVC profiles. A level-compliant decoder should be able to decode any bitstream that does not require exceeding the limits defined in the level. The definitions of the profile and level can contribute to interpretability. For example, during video transmission, a pair of profile and level definitions can be negotiated and agreed upon for the entire transmission session. More specifically, in H.264 / AVC, a level can define limits on the number of macroblocks to be processed, the size of the decoded picture buffer (DPB), the size of the decoded picture buffer (CPB), the vertical motion vector range, the maximum number of motion vectors per two consecutive MB, and whether a B-block can have sub-macroblock partitions smaller than 8x8 pixels. In this way, the decoder can determine whether it can correctly decode the bitstream.

[0046] exist Figure 1 In one example, the encapsulation unit 30 of the content preparation device 20 receives a base stream comprising decoded video data from the video encoder 28 and a base stream comprising decoded audio data from the audio encoder 26. In some examples, the video encoder 28 and the audio encoder 26 may each include a blocker for forming PES packets from the encoded data. In other examples, the video encoder 28 and the audio encoder 26 may each be coupled with a corresponding blocker to form PES packets from the encoded data. In still other examples, the encapsulation unit 30 may include a blocker for forming PES packets from the encoded audio and video data.

[0047] Video encoder 28 can encode video data of multimedia content in various ways to produce different representations of multimedia content with various bit rates and characteristics, such as pixel resolution, frame rate, conformity to various decoding standards, conformity to various profiles and / or profile levels for various decoding standards, representations with one or more views (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. Representations used in this disclosure may include audio data, video data, text data (e.g., for closed captions), or one of these types of data. The representation may include a primary stream, such as an audio primary stream or a video primary stream. Each PES packet may include a stream_id that identifies the primary stream to which the PES packet belongs. Encapsulation unit 30 is responsible for assembling the primary streams into video files of various representations (e.g., segments).

[0048] Encapsulation unit 30 receives PES packets for representing the basic stream from audio encoder 26 and video encoder 28 and forms corresponding Network Abstraction Layer (NAL) units from the PES packets. Decoded video segments can be organized into NAL units, which provide a “network-friendly” video representation for applications such as video telephony, storage, broadcasting, or streaming. NAL units can be classified as Video Decoding Layer (VCL) NAL units and non-VCL NAL units. VCL units may contain the core compression engine and may include block, macroblock, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some examples, a decoded picture in a time instance (typically presented as a primary decoded picture) may be included in an access unit, which may include one or more NAL units.

[0049] Non-VCL NAL units can include parameter set NAL units and SEI NAL units, etc. Parameter sets may contain sequence-level header information (in the Sequence Parameter Set (SPS)) and infrequently changing picture-level header information (in the Picture Parameter Set (PPS)). Using parameter sets (e.g., PPS and SPS) eliminates the need to repeat infrequently changing information for each sequence or picture; therefore, decoding efficiency can be improved. Furthermore, the use of parameter sets enables out-of-band transmission of important header information, thus avoiding the need for redundant transmission for fault tolerance. In out-of-band transmission examples, parameter set NAL units can be transmitted on different channels than other NAL units (e.g., SEI NAL units).

[0050] Supplemental Enhancement Information (SEI) may contain information that is not essential for decoding decoded image samples from VCL NAL units, but can be helpful for processes related to decoding, display, fault tolerance, and other purposes. SEI messages may be contained in non-VCL NAL units. SEI messages are specification parts of certain standards and are therefore not always mandatory for standards-compliant decoder implementations. SEI messages can be sequence-level or picture-level. SEI messages may contain some sequence-level information, such as the extensibility information SEI message in the SVC example and the view extensibility information SEI message in MVC. These example SEI messages can convey information about, for example, the extraction and characteristics of operation points. Furthermore, the encapsulation unit 30 can form a manifest file, such as a Media Rendering Descriptor (MPD) describing the characteristics of the representation. The encapsulation unit 30 can format the MPD according to Extensible Markup Language (XML).

[0051] The encapsulation unit 30 can provide data of one or more representations of multimedia content, along with a manifest file (e.g., MPD), to the output interface 32. The output interface 32 may include a network interface or an interface for writing to storage media, such as a Universal Serial Bus (USB) interface, a CD or DVD burner or drive, an interface for magnetic or flash storage media, or other interfaces for storing or transmitting media data. The encapsulation unit 30 can provide data of each representation of the multimedia content to the output interface 32, which can then transmit the data to the server device 60 via a network or storage medium. Figure 1 In the example, server device 60 includes storage medium 62 for storing various multimedia content 64, each multimedia content 64 including a corresponding manifest file 66 and one or more representations 68A-68N (representations 68). In some examples, output interface 32 may also send data directly to network 74.

[0052] In some examples, representation 68 can be divided into adaptation sets. That is, each subset of representation 68 may include a corresponding set of common characteristics, such as codecs, profiles and levels, resolution, number of views, segmented file formats, text type information that can identify the language or other characteristics of the text to be displayed along with the representation and / or audio data to be decoded and presented (e.g., via a speaker), camera angle information that can describe the real-world camera perspective of the camera angle or scene used for the representation in the adaptation set, rating information describing the suitability of the content for a particular audience, and so on.

[0053] The manifest file 66 may include data indicating a subset of representations 68 corresponding to a particular adaptation set, as well as common characteristics of the adaptation set. The manifest file 66 may also include data representing individual characteristics of the individual representations of the adaptation set, such as bit rate. In this way, the adaptation set can provide simplified network bandwidth adaptation. Representations in the adaptation set can be indicated using sub-elements of the adaptation set element in the manifest file 66.

[0054] Encapsulation unit 30 can further signal addressable resource index (ARI) tracks in a media file including media data according to the techniques of this disclosure. An ARI track can describe all details of a subset of the Common Media Application Format (CMAF) switch set and addressable resources within a single index track. These techniques presuppose the existence of a CMAF switch set with the same (i.e., common) segmentation, fragment, and chunking structure as all tracks, and that encapsulation unit 30 can assign a single ARI track to the CMAF switch set. Encapsulation unit 30 can be configured to apply certain principles. For example, encapsulation unit 30 can time-align the ARI track with the CMAF switch set. Encapsulation unit 30 can construct ARI tracks to record the attributes of all tracks in the CMAF switch set. Encapsulation unit 30 can define header information for metadata tracks. Encapsulation unit 30 can define samples of ARI tracks for each CMAF chunk in a time-aligned manner, such that the samples contain detailed information for each sample of the CMAF chunk.

[0055] Segmentation of ARI tracks can be independent of the segmentation structure. For example, client device 40 can issue a single HTTP request to retrieve an ARI track, and a service performed by server device 60 can provide updates based on HTTP chunks. Random access is not necessarily relevant, as each sample may be a synchronized sample. Movie clips can be used. The header information of the ARI tracks discussed above may include the track number, switchset identifier, timescale identical to the timescale of the switchset track, and additional information.

[0056] Encapsulation unit 30 can form an ARI track to include samples aligned with the chunk time of the represented segment. Each sample may include a CMAF track identifier, a segment boundary flag (indicating whether a chunk corresponds to the start of a new segment), and for each CMAF track id in the switching set (e.g., one or more adaptation sets, including one or more representations, where each representation may correspond to a CMAF media track), a flag signaling flag, a stream access point (SAP) type (e.g., SAP_type) value, an offset to the start of the segment, and multiple pairs of minimum buffer / bandwidth predictions (for each sample or chunk at decoding time t).

[0057] The techniques disclosed herein can be applied to use cases involving live streaming with latency of 6 to 8 seconds. In this use case, an adapter set may have multiple representations, each with a capped variable bitrate (VBR). The bitrate typically defines the content quality. Video encoder 28 can encode the data representing the capped VBR into parameters in the hypothetical reference decoder (HRD) information of the video bitstream (e.g., in the sequence parameter set (SPS)). Encapsulation unit 30 can further record the capped VBR in manifest file 66 (e.g., DASH MPD) using a minimum buffer time / bandwidth pair.

[0058] Based on these techniques, client device 40 can appropriately handle segments with peak values ​​significantly higher than the average value. Specifically, client device 40 can be configured to persist the use of a bitstream with a capped VBR, even when retrieving segments or chunks with capped peak values ​​and delivering such segments or chunks through the access pipeline. If the access network bandwidth changes, client device 40 can switch to a relevant representation (e.g., one of representations 68) with new capped VBR parameters.

[0059] This approach avoids downtime, ensuring client devices consistently retrieve good quality media data, optimizes network bandwidth usage, and is directly applicable to DASH MPD (or other manifest files). Typically, service providers want client devices to select a specific quality representation and use only switching to achieve service continuity. Constant switching to fill the delivery pipeline can have negative consequences. Fine-grained switches used to fill access bandwidth can over-consume processing resources and excessively consume available network bandwidth. Always filling the access pipeline with the highest possible bitrate representation can waste bits.

[0060] Traditional client devices can operate in various ways. For example, a traditional client device can switch to different operating modes for certain content (e.g., time-shifted consumption, live-to-VoD conversion, etc.). In this case, the latency target may increase, but the capped VBR may remain unchanged, so the new operating mode may not actually change the operation. Alternatively, traditional client devices can operate on highly volatile networks. However, signaling data may not be able to handle highly volatile networks.

[0061] It has been asserted that multiple minimum buffer time / bandwidth pairs are required, along with signaling for each segment size. However, the generation and client use of this data are questionable. Based on these assertions and issues, the techniques disclosed herein can be used to address how to map multiple minimum buffer / bandwidth pairs to encoding parameters, such as whether it is possible to run an encoder that produces outputs with multiple meaningful minimum buffer / bandwidth pairs, i.e., whether multiple such pairs provide more information than a single minimum buffer / bandwidth pair. These techniques can also be used to address whether client device 40 has sufficient information to switch representations (e.g., between representations 68) and handle segments with peak values ​​significantly greater than the average bandwidth of the selected representation in representation 68, if the segment size and duration of each segment are signaled in the live stream.

[0062] The following aspects are based on the aforementioned considerations. Currently, video encoders do not signal multiple HRD parameters, including multiple minimum buffer / bandwidth values. For example, in FFMPEG, there are three parameters:

[0063] ●--bitrate <integer>Enable single-channel ABR rate control. Specify the target bit rate in kbps.

[0064] ●--vbv-bufsize <integer>: Specifies the size of the VBV buffer (kbits). Enables VBV in ABR mode. In CRF mode, --vbv-maxrate must also be specified.

[0065] ●--vbv-maxrate <integer>

[0066] Maximum local bit rate (kbits / sec). Only used if vbv-bufsize is also not zero. Enabling VBV in CRF mode requires both vbv-bufsize and vbv-maxrate. Default: 0 (disabled).

[0067] Please note that VBV emergency noise reduction is enabled when VBV is enabled (using a valid `--vbv-bufsize`). When frame QP >

[0068] When QP_MAX_SPEC(51), this enables aggressive denoising at the frame level, significantly reducing the bit rate and allowing rate control to assign lower QPs to subsequent frames. The visual effect is blurred, but noticeable block / shift artifacts are removed.

[0069] The VBV (Video Buffer Verifier) ​​can directly respond to HRD parameters, but it cannot set multiple HRDs. What's missing is how to set VBV parameters.

[0070] The bitrate is also not clearly defined in terms of window size and compliance. It may be completely unusable for real-time services. Real-time services require running CRF plus VBV parameters. This would be an improvement on the reference scenario to run encoding using three parameters set for each representation:

[0071] ●vbv-maxrate: Set to the desired maximum bitrate and @bandwidth value.

[0072] ●vbv-bufsize: The number of bits adjusted based on @minBufferTime multiplied by vbv-maxrate.

[0073] ●CRF Parameter: A reasonable CRF parameter for the content settings. Note that this CRF parameter is not signaled (unless possibly through quality ranking) and may not be.

[0074] Well-defined encoder settings and the resulting compliance aspects are important. More parameters than VBV parameters with explicit semantics can be defined. Subsequent reporting of segment size may be helpful, but it is not critical for operations at the live edge of streaming, as client device 40 may not have access to this information. Therefore, the use of this information by client device 40 and under what operating modes it is used can be defined. Furthermore, apart from VBV parameters for live operation, there are no agreed-upon parameter compliances.

[0075] In DASH, as an example, several tools exist for signaling bandwidth information. Manifestation file 60 can be implemented as an MPD, which includes a minimum buffer size and bandwidth value representing the data signaling of 68. That is, a DASH MPD can include a pair of values: a bandwidth value and a buffer description, namely the minimum buffer time (MBT), represented as MPD@minBufferTime, and the bandwidth (BW) represented by the value of Representation@bandwidth. The following holds true:

[0076] The minimum buffer time (MBT) value does not provide the client with any indication of how long the buffered media should last. However, it describes how much buffer the client should have under ideal network conditions. Therefore, MBT does not describe burstiness or jitter in the network; it describes burstiness or jitter in the content encoding. Along with the BW value, it is a property of the content. Using the "leaky bucket" model, given the content encoding method, it is the size of the bucket that makes BW true.

[0077] ●Minimum buffer time provides information for each representation, the following should be true: if the representation (starting from any segment) is transmitted via a constant bit rate channel with a bit rate equal to the BW attribute value, then each access unit with a presentation time PT is available at the client at the latest after a delay of at most PT+MBT.

[0078] ● In the absence of any other guidelines, MBT should be set to the maximum GOP size (decoded video sequence) of the content, which is typically the same as the maximum segment duration of a live profile or the maximum sub-segment duration of a video-on-demand profile. MBT can be set to a value less than the maximum (sub)segment duration, but should not be set to a higher value.

[0079] The following additional information is provided in DASH-IF IOP:

[0080] In a simple and straightforward implementation, the DASH client determines which segment to download based on the following status information:

[0081] ● Currently available buffers in the media pipeline

[0082] ●Current estimated download speed, rate

[0083] ●The value of the @minBufferTime attribute, MBT

[0084] ● Each set of @bandwidth attribute values ​​representing i, BW[i]

[0085] The client's task is to select a suitable representation i.

[0086] The relevant issue is that, starting from SAP, the DASH client can continue playing data. This means that at the current time, its buffer does indeed contain buffered data. Based on this model, the client can download a representation i where BW[i] ≤ rate*buffer / MBT without emptying the buffer.

[0087] Please note that some idealizations in this model generally do not hold true in practice, such as constant bitrate channels, segmented progressive downloads and playback, and non-blocking and congested HTTP requests. Therefore, DASH clients should use these values ​​with caution to compensate for such realities; in particular, variations in download speed, latency, jitter, media component request scheduling, and other practical considerations. One example is whether the DASH client operates at a segment granularity. In this case, not only a portion of the segment (i.e., MBT) needs to be downloaded, but the entire segment also needs to be downloaded, and if the MBT is less than the segment duration, the segment duration should be used instead of the MBT for the required buffer size and download schedule, i.e., downloading representations i where BW[i] ≤ rate*buffer / max_segment_duration.

[0088] The DASH-IF IOP v5 draft provides details regarding subsegment information and the segment index (SIDX). The parameters related to subsegment information in the MPD are documented in Table 6 of DASH-IF IOP v6, reproduced below. These parameters are contained in one or more BaseURLs and a SegmentBase element, as defined in Clause 5.3.9.4 of ISO / IEC 23009-1.

[0089] Table 6 - Follow-up Information

[0090]

[0091] The `SegmentBase` element is sufficient to describe the subsegment information, and the media segment URL is included in the `BaseURL` element. This option is referred to as Option 4 in terms of addressing modes. This subsegment information is provided in a single segment index box, describing all subsegments (i.e., movie clips) in the representation. To address complexity and interoperability issues, DASH-IF IOP limits the on-demand profile to a single segment index box.

[0092] Table 6 also indicates which elements are mandatory (M), optional (O), or optional. Default values ​​are also provided. For example, M4 indicates that option 4 requires the attribute / element to exist.

[0093] All other elements not recorded in Table 6 are either introduced elsewhere in the specification or are not expected to appear, and if they do appear, are expected to be ignored by the client.

[0094] Based on the information in the MPD table 6, a list of segments contained in the representation of period i with period duration PD[i] can be calculated. Based on the above information, for each representation r in period i, the following information can be derived:

[0095] ● The total number of sub-segments in the periodicity, N[i,r],

[0096] ● The MPD start time MST[k,i,r], k=1,...,N[i,r], relative to the start of the cycle for each media segment.

[0097] ● The MPD-based segment duration MSD[k,i,r], k=1,...,N[i,r] for each media segment.

[0098] ● The URL of each sub-segment, URL[k,i,r], is the URLByteRange(BaseURL,first,last) defined in Section 6.4.3.3.

[0099] The MPD information used and the results therein are documented below. A single segmented index is provided, using syntax conforming to ISO / IEC 14496-12:

[0100]

[0101] The definition is as follows:

[0102] ●ts[i,r] is the value of the time scale field and is the same as the value of the @timescale attribute.

[0103] ●e[i,r] is the value of the @eptDelta attribute.

[0104] ●o[i,r] is the value of the @presentationTimeOffset property.

[0105] ●pd[i,r] is the value of the @presentationDuration attribute (if it exists), otherwise it is PD[i]*ts

[0106] ●rc is the value of the reference_count field.

[0107] ●pt is the value of the earliest_presentation_time field.

[0108] ● os is the value of the first_offset field

[0109] ● d[j] is the value of the subsegment_duration field of the j-th entry, where j = 1, … rc

[0110] ● b[j] is the value of the referenced_size field of the j-th entry, where j = 1, … rc

[0111] Then, the client device 40 can derive the segment information for each segment k = 1, …, N[i,r] as follows:

[0112] ● j = 1

[0113] ● while(pt + d[j] < o[i,r]) / * find the first subsegment * /

[0114] ○ pt = pt + d[j]

[0115] ○ os = os + b[j]

[0116] ○ j++

[0117] ● k = 1

[0118] ● e[i,r] = pt - o[i,r] / * overlap at the start - negative or 0 * /

[0119] ● while(pt + d[j] - o[i,r] < PD[i] * ts[i,r]) / * find the last subsegment * /

[0120] ○ MST[k,i,r] = (pt - o[i,r]) / ts[i,r]

[0121] ○ MSD[k,i,r] = d[j] / ts[i,r]

[0122] ○ URL[k,i,r] = URLByteRange(BaseURL, os, os + b[j] - 1)

[0123] ○ pt = pt + d[j]

[0124] ○ os = os + b[j]

[0125] ○ j++

[0126] ○ k++

[0127] ● pd[i,r] = pt

[0128] ● N[i,r] = k

[0129] Server device 60 can receive and process byte-range requests to exchange media data within byte ranges, while client device 40 can send byte-range requests. The URLByteRange(URL, first, last) call maps to a partial HTTP request with a byte range, as shown below.

[0130] URL URL requested by the HTTP Request first The value is mapped to the value of 'first-byte-pos' in 'byte-range-spec' of IETF RFC7231:2014,2.1. Last The value maps to the value of 'last-byte-pos' in 'byte-range-spec' of IETF RFC7231:2014,2.1.

[0131] For media presentations that conform to the DASH-IF core profile, the following requirements and recommendations apply to sub-segment information:

[0132] ●A precise segment index should exist, with its value set as follows:

[0133] ○reference_ID should be set to the segment's track_ID.

[0134] ○ The time scale should be set to the time scale field of the media header box of the track.

[0135] ○earliest_presentation_time should be set to 0.

[0136] ○reference_type should be set to 0 for all values, because only movie clip header boxes are referenced.

[0137] SAP_delta_time should be set to 0.

[0138] ● If the value of e[i,r] is not 0, then @eptDelta should exist, and if it exists, it should be set to the value of e[i,r] as defined in Clause 6.4.3.4.

[0139] ● If the value of (pd[i,r]-o[i,r])-PD[i]*ts[i,r] is not 0, then @presentationDuration should exist, and if it exists, it should be set to the value of pd[i,r] as defined in Clause 6.4.3.4.

[0140] For DASH-IF clients that support DASH-IF media rendering, the following requirements and recommendations apply to sub-segment information:

[0141] ● The client shall support the services provided in accordance with the requirements of Clause 6.4.3.4. Specifically, this includes:

[0142] ○ Use URLByteRange(URL, first, last) to download the index range, including the CMAF header and segment index.

[0143] ○ Download sidx as a byte range

[0144] Download sub-segments as byte ranges

[0145] Given a client that supports live profiles, the above requirements can be met as follows:

[0146] ● The client has already sent requests for the moov box—these requests also need to be expanded to cover the segmented index (sidx).

[0147] ● Parse the segmented index (<150 lux, in Javascript, open source, and DASH-IF reference clients), outputting a list of (URL, byte range) pairs that will replace the list of commonly used URLs in the live profile.

[0148] ● Create a list of request (URL, byte range) pairs, instead of a list of URLs.

[0149] Only movie clips are referenced. Therefore, the entire switching timeline of the representation can be constructed by downloading SIDX first. However, a smart client might choose to download only the beginning of the SIDX box of the representation for a quick start, and then download the rest after the initial media sub-segments begin streaming.

[0150] Client device 40 can also determine its request size independently of the segment duration.

[0151] By resolving SIDX for each representation of interest, client device 40 can create a complete sub-segment mapping.

[0152] In the DASH context, segments are typically treated as individual units for downloading and randomly accessing media presentations, and they are also addressed by a single URL. However, segments may have internal structures that implement resynchronization at the container level, and may even allow random access to the corresponding representation within a segment. The resynchronization mechanism is supported by resynchronization elements and signaled notifications.

[0153] The resynchronization element signals the resynchronization point within a segment. The resynchronization point marks the beginning (at the byte position) of a well-structured, contiguous range of bytes within a segment that contains media data for a specific presentation duration and can be accessed independently at the container format level. Resynchronization points can provide additional functionality, such as access at the decryption and decoding levels.

[0154] Container formats that use the resynchronization feature must define the resynchronization point and associated properties.

[0155] The resynchronization point in a segment can be defined as follows:

[0156] 1. Resynchronization points enable parsing and processing to begin at the container level.

[0157] 2. The resynchronization point has been assigned the following properties:

[0158] a. It has a byte offset or index starting from the segment, pointing to the resynchronization point.

[0159] b. It has an earliest rendering time in the representation, which is the minimum rendering time of any sample contained in the representation when processing begins from the resynchronization pointer.

[0160] c. It assigns a type, such as the SAP type definition in ISO / IEC 14496-12.

[0161] d. It assigns a Boolean marker attribute, Marker, to determine whether a resynchronization point can be detected while parsing segments through a specific structure, or whether a resynchronization point needs to be signaled externally.

[0162] 3. Processing segments begins from the resynchronization point, along with information from the initialization segment (if present), allowing the container to parse it. Whether and how the contained, and possibly encrypted, underlying stream is accessed can be indicated by the resynchronization access point type.

[0163] Each resynchronization point can be signaled using all attributes in MPD by providing a side-car segment describing the resynchronization point within the segment. However, such side-car segments are not always available, or at least timely provision can be difficult, for example, in dynamic and live services, because resynchronization points are added by the segment packer independently of MPD updates. Resynchronization points can be generated by encoders and packers independent of MPD. Furthermore, in low-latency scenarios, MPD signaling may be unavailable to DASH clients.

[0164] Therefore, two non-mutually exclusive methods are specified to signal the resynchronization point provided in the segment of MPD:

[0165] 1. Provide a binary mapping for each resynchronization point in the segment by indexing the resynchronization of each media segment. This is most easily used for segments that are fully available on the network.

[0166] 2. The presence of a resynchronization point in a media segment is signaled by additional information that allows resynchronization to be easily located based on byte position and presentation time, and the type of resynchronization point is provided.

[0167] If the resynchronization element exists with the included @dImin and @dT attributes, and has adjustment values ​​dImin in bytes and dT in seconds respectively, and the @availabilityTimeComplete attribute is set to false, the following should be true:

[0168] ●The first chunk becomes available at the start time of the segmented adjusted availability.

[0169] ●The (i+1)th block becomes available at the sum of the adjusted availability start time and i*dT of the segment, where i = 1, ..., N, and N is the total number of blocks in the segment.

[0170] ● If the @rangeAccess attribute is set to true, available chunks can be accessed using a byte range. If set to false, the client cannot expect a response to an available byte range request to produce valid data.

[0171] When the DASH standard was written, requesting the available byte range for partially available segments (i.e., segments still in production) was not consistently supported in CDNs, but efforts were planned to provide consistent behavior. Therefore, content providers are encouraged to examine the capabilities of the CDNs on which they deploy services before allowing byte range access to the available portion of partially available segments by setting the @rangeAccess attribute to true.

[0172] To signal the aforementioned attributes, the resynchronization element is defined with different attributes, which are explained in more detail in Clause 5.3.12.2, Table X. XML syntax is provided in Clause 5.3.12.3.

[0173] The technology disclosed herein can be used to solve various use cases. According to the technology disclosed herein, content preparation device 20, server device 60, and client device 40 can be configured to perform the following.

[0174] The techniques disclosed herein can be used to improve signaling used for encoding parameters. These parameters may include the following:

[0175] 1. Mapping from VBV parameters to @minBufferTime and @bandwidth.

[0176] 2. Signal: CRF(VBR) specifies whether to encode with VBV parameters and specifies the quality of the signal.

[0177] 3. Provide guidance on how to generate content and how to signal that content.

[0178] 4. May provide instructions to the client to select and adhere to a representation based on VBV, regardless of content bitrate fluctuations.

[0179] 5. It might be possible to allow encoder suppliers to accurately signal their encoding parameters, but we want to avoid this.

[0180] 6. For live encoders, we do not need multiple bandwidths / minimum buffer times, and there is no way to control the encoder to do so.

[0181] For signaling of each segment size:

[0182] 1. First, the questions above still need to be answered. How can the reference model be changed to require signaling the size of each segment? How does the player use this information?

[0183] 2. We have a segmented index, which only applies to the entire track file. If we need to expand it, we can create a segmented index, which can also be used for segments. However, this can only be added after a cycle is completed. Adding a cycle at regular intervals is also possible.

[0184] 3. We have a resynchronization index that can create a size signaling for each segment.

[0185] 4. If we need anything between a segment index and a resync index, let us be very clear about what we need and why we need it.

[0186] This disclosure describes techniques that can address the above-mentioned concerns. To improve signal transmission of the encoded parameters, encapsulation unit 30 can be added as an option to each of representation 68, appended to the VBV parameters @minBufferTime and @bandwidth.

[0187] ●Signal: The content is a capped VBR encoded using the following method.

[0188] ○ VBV parameters used for capping

[0189] Additionally, a VBR flag with the following parameters

[0190] ■Window Size

[0191] ■Target Bit Rate

[0192] ■ The maximum overshoot percentage of the target bit rate within the window size.

[0193] ○ Constant mass parameters (if applicable).

[0194] The issues concerning Addressable Resource Index (ARI) tracks are summarized below:

[0195] ● In some cases, it may be desirable for adaptive streaming clients to have information about the duration and size of addressable resources, and possibly a subset of those resources on the server.

[0196] ● Addressable resources include track files, segments, or chunks in the CMAF context, but also apply to DASH or HLS.

[0197] ● For on-demand services, the precise mapping of this information is provided by segmented indexes.

[0198] ●Generally speaking, and especially for live services, minimum buffer and bandwidth signaling.

[0199] ● Resynchronization allows the use of Resync Index Segment to provide a localized index for each segment.

[0200] Even with these options, it is beneficial to create additional information about the addressable resources, such as estimates or accurate bit rates (duration and size), for optimized client operations. Client optimization can vary depending on how the client operates, for example:

[0201] ●Operation modes: Low-latency live streaming, live streaming, time-shift, VoD

[0202] ●Target latency of the client

[0203] ●Network status

[0204] ●Expected content quality

[0205] However, existing solutions have some problems:

[0206] ●Minimum buffer / bandwidth: Only provides an inaccurate bit rate instead of a usable client download schedule.

[0207] ●Resynchronization: Generally works well, but all information is separate for each segment and representation, resulting in a large number of requests to retrieve all information.

[0208] ● Segment indexes are typically only available for VoD because they record the entire file.

[0209] The techniques disclosed herein related to the use of ARI orbitals can provide a more general and flexible approach to these problems.

[0210] Addressable Resource Index (ARI) tracks describe all details of the addressable resources and subsets of a CMAF switch set within a single index track. These techniques are based on the assumption that a CMAF switch set exists with the same (i.e., common) segmentation, fragment, and chunking structure across all tracks and is assigned a single ARI track. The following principles can be applied:

[0211] ● The ARI Track is time-aligned with the CMAF switch set.

[0212] ●ARI Track records the properties of all tracks in the CMAF switching set.

[0213] ● Define header information for the metadata track.

[0214] ● A sample of tracks is defined for each CMAF block in a time-aligned manner.

[0215] ●The sample contains detailed information for each sample.

[0216] Track delivery and segmentation can be independent of the chunking / segmentation structure of the associated switch set. For example, this can be achieved by issuing a single HTTP request, with the service only providing updates based on HTTP chunks. Random access is not necessarily correlated, as each sample could be a synchronous sample. However, in this case, movie segments are required.

[0217] The CMAF Addressable Resource Index (ARI) can be defined as follows:

[0218] Example entry type: 'cari'

[0219] Container: Sample description box ('stsd')

[0220] Mandatory: No

[0221] Quantity: 0 or 1

[0222] This metadata describes all the details of the addressable resources and subsets of the CMAF switching set as defined in ISO / IEC 23000-19 in a single index track.

[0223] Assume that the same segmentation, fragmentation, and chunking structure is applied to all tracks in the CMAF switching set.

[0224] Apply the following principles:

[0225] ● The ARI track is time-aligned with the CMAF switch set track.

[0226] ●ARI tracks record the properties of all tracks in the CMAF switch set.

[0227] ● Define header information for the metadata track.

[0228] ● Track samples are defined for each CMAF chunk in a time-aligned manner. The association between the chunk and the metadata samples is completed, ensuring that the baseMediaDecodeTime of the chunk is the same as the sample time in the metadata track.

[0229] ●The sample contains detailed information for each sample.

[0230] This track can even be used to carry events or producer reference time, for example.

[0231] The syntax of CMAF ARI metadata can be defined for sample entries as follows:

[0232]

[0233] CMAF addressable resource index samples can use the following syntax:

[0234]

[0235]

[0236] The semantics of the above syntax can be defined as follows:

[0237] `switching_set_identifier` is a unique identifier for the switching set within the application context.

[0238] num_tracks is the number of tracks in the switching set.

[0239] track_ID provides the sorting order of track_IDs in the sample.

[0240] The quality_indicator_flag flag indicates whether a particular quality indicator is used to identify the quality of a chunk.

[0241] quality_identifier is an identifier that describes how to interpret the quality values ​​in a sample.

[0242] The segment_start_flag flag indicates whether this block is the same as the start of a segment; that is, if the block is the start of a segment.

[0243] The `starts_with_SAP` flag indicates whether this block begins with SAP.

[0244] If SAP_type is set to starts_with_SAP, it identifies the SAP type.

[0245] The marker indicates whether this block includes a marker containing styp.

[0246] The emsg_flag flag indicates whether this block provides an emsg box.

[0247] The prft_flag indicates whether this block contains a prft box.

[0248] first_offset indicates the offset of the block from the beginning of the sequence.

[0249] size provides the size of the chunk in octets.

[0250] The quality is provided by the quality scheme for the component. If there is no quality scheme, the quality is interpreted linearly with the increase in quality that has an increasing value.

[0251] Loss identifier data is lost.

[0252] num_prediction_pairs provides the number of prediction pairs for the expected bit rate.

[0253] The prediction_min_windows value provides the same minimum buffer time as the MPD value.

[0254] The predicted_max_bitrate provides a value for the bandwidth time that is the same as the MPD semantics maintained for the duration of the predicted_min_windows value.

[0255] Encapsulation unit 30 can construct manifest file 66 (e.g., DASH MPD) to include data related to:

[0256] 1. Metadata tracks are provided as a regular adapter set with a single track.

[0257] 2. The switching set is associated with this track.

[0258] 3. Streaming is proceeding normally, but any optimizations can be made:

[0259] a. Availability Time Offset

[0260] b. Chunking

[0261] c. Segmentation

[0262] Client device 40 can access the metadata track in advance, or at least along with the segmented available time. The DASH client of client device 40 can implement a metadata processor for the ARI track and utilize this information. This simplifies the overall addition because all existing streaming technologies can be applied.

[0263] If ARI orbital metadata is provided, resynchronization can be used solely to signal block commitments, but the resynchronization of indexed segment functionality can be removed. Further optimizations are possible.

[0264] Server device 60 includes a request processing unit 70 and a network interface 72. In some examples, server device 60 may include multiple network interfaces. Furthermore, any or all features of server device 60 may be implemented on other devices in the content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, intermediate devices in the content delivery network may cache data for multimedia content 64 and include components that substantially conform to those of server device 60. Typically, network interface 72 is configured to send and receive data via network 74.

[0265] Request processing unit 70 is configured to receive network requests for data on storage medium 62 from a client device, such as client device 40. For example, request processing unit 70 may implement Hypertext Transfer Protocol (HTTP) version 1.1, as described in RFC 2616, "Hypertext Transfer Protocol – HTTP / 1.1," by R. Fielding et al., Network Working Group, IETF, June 1999. That is, request processing unit 70 may be configured to receive HTTP GET or partial GET requests and, in response to the request, provide data for multimedia content 64. The request may specify a segment representing one of 68, for example, using a URL with a segment. In some examples, the request may also specify one or more byte ranges of the segment, thus including partial GET requests. Request processing unit 70 may also be configured to serve HTTP HEAD requests to provide header data representing one of 68. In any case, request processing unit 70 may be configured to process the request to provide the requested data to a requesting device, such as client device 40.

[0266] Additionally or alternatively, request processing unit 70 may be configured to deliver media data via a broadcast or multicast protocol such as eMBMS. Content preparation device 20 may create DASH segments and / or sub-segments in substantially the same manner as described, but server device 60 may use eMBMS or another broadcast or multicast network transport protocol to deliver these segments or sub-segments. For example, request processing unit 70 may be configured to receive multicast group join requests from client device 40. That is, server device 60 may advertise to client devices including client device 40 the Internet Protocol (IP) address associated with the multicast group and specific media content (e.g., broadcast of a live event). Client device 40 may then submit a request to join the multicast group. This request may be propagated throughout network 74, such as routers comprising network 74, causing routers to direct traffic destined for the IP address associated with the multicast group to subscribing client devices, such as client device 40.

[0267] like Figure 1 As shown in the example, multimedia content 64 includes a manifest file 66, which may correspond to a Media Presentation Description (MPD). The manifest file 66 may contain descriptions of different alternative representations 68 (e.g., video services with different qualities), and this description may include, for example, codec information, profile values, level values, bitrate, and other descriptive characteristics of representation 68. Client device 40 may retrieve the MPD of the media presentation to determine how to access segments of representation 68.

[0268] Specifically, retrieval unit 52 can retrieve configuration data (not shown) of client device 40 to determine the decoding capabilities of video decoder 48 and the rendering capabilities of video output 44. The configuration data may also include any one or all of the following: language preferences selected by the user of client device 40, one or more camera angles corresponding to depth preferences set by the user of client device 40, and / or rating preferences selected by the user of client device 40. Retrieval unit 52 may, for example, include a web browser or media client configured to submit HTTP GET and partial GET requests. Retrieval unit 52 may correspond to software instructions executed by one or more processors or processing units (not shown) of client device 40. In some examples, all or part of the functionality described with respect to retrieval unit 52 may be implemented in hardware, or a combination of hardware, software, and / or firmware, wherein the necessary hardware may be provided to execute the software or firmware instructions.

[0269] The retrieval unit 52 can compare the decoding and rendering capabilities of the client device 40 with the characteristics of the representation 68 indicated by the information in the manifest file 66. The retrieval unit 52 can initially retrieve at least a portion of the manifest file 66 to determine the characteristics of the representation 68. For example, the retrieval unit 52 can request a portion of the manifest file 66 that describes the characteristics of one or more adapter sets. The retrieval unit 52 can select a subset of representations 68 (e.g., an adapter set) that has characteristics that can be satisfied by the decoding and rendering capabilities of the client device 40. The retrieval unit 52 can then determine the bit rate of the representations in the adapter set, determine the amount of currently available network bandwidth, and retrieve segments from one of the representations with a bit rate that the network bandwidth can satisfy.

[0270] Generally, a higher bitrate indicates higher quality video playback, while a lower bitrate indicates sufficient quality video playback when available network bandwidth is reduced. Therefore, when available network bandwidth is relatively high, retrieval unit 52 can retrieve data from a relatively high bitrate representation, and when available network bandwidth is low, retrieval unit 52 can retrieve data from a relatively low bitrate representation. In this way, client device 40 can stream multimedia data over network 74 while adapting to varying network bandwidth availability.

[0271] Additionally or alternatively, the retrieval unit 52 may be configured to receive data according to a broadcast or multicast network protocol, such as eMBMS or IP multicast. In such an example, the retrieval unit 52 may submit a request to join a multicast network group associated with specific media content. After joining the multicast group, the retrieval unit 52 may receive the multicast group's data without issuing further requests to the server device 60 or content preparation device 20. When the multicast group's data is no longer needed, the retrieval unit 52 may submit a request to leave the multicast group, such as stopping playback or changing the channel to a different multicast group.

[0272] Network interface 54 can receive data from selected segments and provide it to retrieval unit 52, which in turn can provide the segments to decapsulation unit 50. Decapsulation unit 50 can decapsulate the elements of the video file into a PES stream, unpack the PES stream to retrieve encoded data, and send the encoded data to audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream, as indicated by the PES packet header of the stream. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data to video output 44. The decoded video data may include multiple views of the stream.

[0273] Each of the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, retrieval unit 52, and decapsulation unit 50 can be implemented as any of a variety of suitable processing circuits, such as, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. Each of the video encoder 28 and video decoder 48 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined video encoder / decoder (CODEC). Similarly, each of the audio encoder 26 and audio decoder 46 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined CODEC. The apparatus including the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, retrieval unit 52, and / or decapsulation unit 50 can include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.

[0274] Client device 40, server device 60, and / or content preparation device 20 may be configured to operate according to the techniques of this disclosure. For illustrative purposes, this disclosure describes these techniques with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 may be configured to perform these techniques, rather than (or attached to) server device 60.

[0275] Encapsulation unit 30 can form NAL units, including a header identifying the program to which the NAL unit belongs and a payload, such as audio data, video data, or data describing the transport or program stream corresponding to the NAL unit. For example, in H.264 / AVC, a NAL unit includes a 1-byte header and payloads of varying sizes. NAL units whose payloads include video data can include video data at various granularities. For example, a NAL unit can include video data blocks, multiple blocks, video data slices, or an entire picture of video data. Encapsulation unit 30 can receive encoded video data from video encoder 28 in the form of PES packets of the elementary stream. Encapsulation unit 30 can associate each elementary stream with its corresponding program.

[0276] Encapsulation unit 30 can also assemble access units from multiple NAL units. Typically, an access unit may include one or more NAL units to represent a frame of video data, and, when such audio data is available, the corresponding audio data for that frame. An access unit typically includes all NAL units of an output time instance, for example, all audio and video data for that time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, specific frames of all views within the same access unit (same time instance) can be rendered simultaneously. In one example, an access unit may include a decoded image within a time instance, which can be rendered as a primary decoded image.

[0277] Therefore, an access unit can include all audio and video frames of a common time instance, such as all views corresponding to time X. This disclosure also refers to the encoded image of a specific view as a "view component." That is, a view component can include the encoded image (or frame) of a specific view for a specific time. Therefore, an access unit can be defined as including all view components of a common time instance. The decoding order of the access units is not necessarily the same as the output or display order.

[0278] Media presentations may include a Media Presentation Description (MPD), which may contain descriptions of different alternative representations (e.g., video services with different qualities), and this description may include, for example, codec information, profile values, and level values. An MPD is an example of a manifest file, such as manifest file 66. Client device 40 can retrieve the MPD of a media presentation to determine how to access movie clips in various presentations. Movie clips may be located within a movie clip frame (moof frame) of a video file.

[0279] The manifest file 66 (which may include, for example, an MPD) can announce the availability of segments representing 68. That is, the MPD may include information indicating a clock time (at which time the first segment of one of 68 becomes available) and information indicating the duration of segments within 68. In this way, the retrieval unit 52 of the client device 40 can determine when each segment is available based on the start time and the duration of segments preceding that particular segment.

[0280] After the encapsulation unit 30 has assembled the NAL units and / or access units into a video file based on the received data, the encapsulation unit 30 passes the video file to the output interface 32 for output. In some examples, the encapsulation unit 30 may store the video file locally or send the video file to a remote server via the output interface 32, instead of sending the video file directly to the client device 40. The output interface 32 may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium, such as an optical drive, a magnetic media drive (e.g., a floppy disk drive), a Universal Serial Bus (USB) port, a network interface, or other output interfaces. The output interface 32 outputs the video file to a computer-readable medium, such as a transmission signal, magnetic media, optical media, memory, flash memory, or other computer-readable media.

[0281] Network interface 54 can receive NAL units or access units via network 74 and provide NAL units or access units to decapsulation unit 50 via retrieval unit 52. Decapsulation unit 50 can decapsulate the elements of the video file into a constitutive PES stream, unpack the PES stream to retrieve encoded data, and send the encoded data to audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream, for example, as indicated by the PES packet header of the stream. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data to video output 44. The decoded video data may include multiple views of the stream.

[0282] Figure 2 To explain in more detail Figure 1 A block diagram of an example set of components for retrieval unit 52. In this example, retrieval unit 52 includes eMBMS middleware unit 100, DASH client 110, and media application 112.

[0283] In this example, the eMBMS middleware unit 100 also includes an eMBMS receiving unit 106, a cache 104, and a proxy server unit 102. In this example, the eMBMS receiving unit 106 is configured to receive data via eMBMS, for example, according to file transfer over unidirectional transport (FLUTE), as described in T. Paila et al., "FLUTE—File Delivery over Unidirectional Transport," Network Working Group, RFC 6726, Nov. 2012, available at tools.ietf.org / html / rfc6726. That is, the eMBMS receiving unit 106 can receive files via broadcast from, for example, a server device 60 that can act as a broadcast / multicast service center (BM-SC).

[0284] When the eMBMS middleware unit 100 receives file data, it can store the received data in cache 104. Cache 104 may include computer-readable storage media, such as flash memory, hard disk, RAM, or any other suitable storage media.

[0285] Proxy server unit 102 can act as a server for DASH client 110. For example, proxy server unit 102 can provide DASH client 110 with an MPD file or other manifest file. Proxy server unit 102 can announce the availability time of segments in the MPD file, as well as hyperlinks from which segments can be retrieved. These hyperlinks may include a local host address prefix corresponding to client device 40 (e.g., 127.0.0.1 for IPv4). In this way, DASH client 110 can request segments from proxy server unit 102 using an HTTP GET or partial GET request. For example, for a segment obtainable from the link http: / / 127.0.0.1 / rep1 / seg3, DASH client 110 can construct an HTTP GET request including a request to http: / / 127.0.0.1 / rep1 / seg3 and submit the request to proxy server unit 102. Proxy server unit 102 can retrieve the requested data from cache 104 and provide the data to DASH client 110 in response to such a request.

[0286] Figure 3 This is a conceptual diagram illustrating the elements of example multimedia content 120. Multimedia content 120 may correspond to multimedia content 64 ( Figure 1 (or another multimedia content stored in storage medium 62.) Figure 3 In the example, multimedia content 120 includes a Media Presentation Description (MPD) 122, multiple representations 124A-124N (representations 124), and an Addressable Resource Information (ARI) track 140. Representation 124A includes optional header data 126 and segments 128A-128N (segments 128), while representation 124N includes optional header data 130 and segments 132A-132N (segments 132). For convenience, the letter N is used to specify the last movie segment in each representation 124. In some examples, there may be different numbers of movie segments between representations 124.

[0287] MPD 122 may include data structures separate from representation 124 and ARI track 140. MPD 122 may correspond to Figure 1 The manifest file 66. Similarly, it indicates that 124 can correspond to... Figure 1 The representation 68. Generally, MPD 122 may include data that generally describes the characteristics of representation 124, such as decoding and rendering characteristics, adaptation set, profile corresponding to MPD 122, text type information, camera angle information, rating information, trick pattern information (e.g., information indicating representations including time subsequences) and / or information for retrieving remote periods (e.g., for inserting targeted advertisements into media content during playback).

[0288] Header data 126 (if present) may describe characteristics of segment 128, such as the time position of the Random Access Point (RAP, also known as the Streaming Access Point (SAP)), which segment of 128 includes the RAP, the byte offset of the RAP within segment 128, the Uniform Resource Locator (URL) of segment 128, or other aspects of segment 128. Header data 130 (if present) may describe similar characteristics of segment 132. Alternatively, such characteristics may be entirely contained within MPD 122.

[0289] Segments 128 and 132 include one or more decoded video samples, each of which may include frames or slices of video data. Each decoded video sample in segment 128 may have similar characteristics, such as height, width, and bandwidth requirements. Such characteristics can be described by data from MPD 122, although such data is not explicitly stated in the original text. Figure 3 The example is shown. MPD 122 may include features as described in the 3GPP specifications, supplementing any or all signaled information described in this disclosure.

[0290] Furthermore, in this example, each of segments 128 and 132 comprises an independently retrievable media data chunk. Specifically, segments 128A-128N comprise corresponding chunks 134A-134N (chunks 134), and segments 132A-132N comprise corresponding chunks 136A-136N. In this example, segments 128 and 132 are time-aligned such that the boundaries of corresponding segments 128 and 132 correspond to a common playback time. Similarly, chunks 134 and 136 are time-aligned such that the boundaries of corresponding chunks 134 and 136 correspond to a common playback time. That is, the playback time at the boundary between segments 128A and 128B is the same as the playback time at the boundary between segments 132A and 132B. In this way, switching between segments or chunks can be performed at segment / chunk boundaries.

[0291] Each of segments 128 and 132 can be associated with a unique Uniform Resource Locator (URL). Therefore, each of segments 128 and 132 can be retrieved independently using a streaming network protocol such as DASH. In this way, a destination device, such as client device 40, can retrieve segment 128 or 132 using an HTTP GET request. In some examples, client device 40 can use a partial HTTP GET request to retrieve a specific range of bytes from segment 128 or 132.

[0292] In this example, ARI track 140 includes header data 144 and samples 142A-142N (samples 142). Each sample 142 corresponds to a block 134, 136 of the corresponding group. That is, ARI track 140 includes one of the samples 142 time-aligned with each time of blocks 134, 136 of segments 128, 132, such as... Figure 3 As shown. The header data 144 may include data indicating the number of orbits included in the switch set described by ARI orbit 140 (i.e., the number of 124), a switch set identifier indicating the switch set described by ARI orbit 140, and a time scale that is the same as the time scale of the switch set, and may include additional information.

[0293] Each of the samples 142 may include data corresponding to each of the blocks 134, 136, such as a segment boundary flag indicating whether one of the blocks 134, 136 has started a new segment (e.g., a new one of the segments 128, 132), a track identifier for the block (e.g., an identifier indicating the one corresponding to 124), and for each track identifier: a marking signaling flag, a Stream Access Point (SAP) type, an offset to the start of the corresponding segment, and multiple pairs for minimum buffer / bandwidth prediction.

[0294] Figure 4 This is a block diagram illustrating the elements of example video file 150, which may correspond to the represented segments, for example... Figure 3 One of segments 128 and 132. Each of segments 128 and 132 may include substantially conforming to Figure 4 The example shows the data arrangement. Video file 150 can be said to encapsulate a segment. As mentioned above, video files conforming to the ISO Basic Media File Format and its extensions store data in a series of objects (called "boxes"). Figure 4 In the example, video file 150 includes a File Type (FTYP) box 152, a Movie (MOOV) box 154, a Segment Index (sidx) box 162, a Movie Clip (MOOF) box 164, and a Movie Clip Random Access (MFRA) box 166. Although Figure 4 The example shown is a video file. It should be understood that other media files may include other types of media data (e.g., audio data, timed text data, etc.) that conform to the ISO Basic Media File Format and its extensions and have a structure similar to that of a video file 150.

[0295] The File Type (FTYP) box 152 typically describes the file type of the video file 150. The File Type box 152 may include data that identifies specifications describing the best use of the video file 150. The File Type box 152 may alternatively be placed before the MOOV box 154, the Movie Clip box 164, and / or the MFRA box 166.

[0296] In some examples, a segment such as video file 150 may be included in an MPD update box (not shown) preceding the FTYP box 152. The MPD update box may include information indicating that the MPD corresponding to the representation including video file 150 will be updated, as well as information for updating the MPD. For example, the MPD update box may provide a URI or URL of the resource used to update the MPD. As another example, the MPD update box may include data for updating the MPD. In some examples, the MPD update box may immediately follow the segment type (STYP) box (not shown) of video file 150, where the STYP box may define the segment type of video file 150.

[0297] MOOV box 154 (in Figure 4 In the example, this includes a Movie Header (MVHD) frame 156, a Track (TRAK) frame 158, and one or more Movie Extensions (MVEX) frames 160. Generally, the MVHD frame 156 can describe general characteristics of the video file 150. For example, the MVHD frame 156 may include data describing the initial creation time of the video file 150, the last modification time of the video file 150, the time scale of the video file 150, the playback duration of the video file 150, or other data generally describing the video file 150.

[0298] TRAK frame 158 may include data for a track in video file 150. TRAK frame 158 may include a Track Header (TKHD) frame describing the characteristics of the track corresponding to TRAK frame 158. In some examples, TRAK frame 158 may include a decoded video picture, while in other examples, the decoded video picture of the track may be included in movie clip 164, which may be referenced by data from TRAK frame 158 and / or sidx frame 162.

[0299] In some examples, video file 150 may include more than one track. Therefore, MOOV frame 154 may include a number of TRAK frames equal to the number of tracks in video file 150. TRAK frames 158 may describe the characteristics of the corresponding track in video file 150. For example, TRAK frames 158 may describe the temporal and / or spatial information of the corresponding track. When encapsulation unit 30 ( Figure 3 When a parameter set track is included in a video file such as video file 150, a TRAK box, similar to MOOV box 154 and TRAK box 158, can describe the characteristics of the parameter set track. Encapsulation unit 30 can signal the presence of a sequence-level SEI message in the parameter set track within the TRAK box describing the parameter set track.

[0300] MVEX frame 160 can describe the characteristics of the corresponding movie clip 164, for example, to signal that video file 150 includes movie clip 164, and video data (if any) appended to MOOV frame 154. In the context of streaming video data, decoded video images can be included in movie clip 164 instead of MOOV frame 154. Therefore, all decoded video samples can be included in movie clip 164 instead of MOOV frame 154.

[0301] MOOV boxes 154 may include an MVEX box 160 equal to the number of movie segments 164 in the video file 150. Each MVEX box 160 may describe the characteristics of the corresponding movie segment in the movie segment 164. For example, each MVEX box may include a Movie Extended Header (MEHD) box describing the duration of the corresponding movie segment in the movie segment 164.

[0302] As described above, encapsulation unit 30 can store the sequence dataset in a video sample that does not include the actual decoded video data. The video sample typically corresponds to an access unit, which is a representation of the decoded image at a specific time instance. In the context of AVC, the decoded image includes one or more VCL NAL units containing information, such as SEI messages, about all pixels used to construct the access unit and other associated non-VCL NAL units. Therefore, encapsulation unit 30 can include the sequence dataset in one of the movie clips 164, which may include sequence-level SEI messages. Encapsulation unit 30 can further signal the presence of the sequence dataset and / or sequence-level SEI messages within one of the movie clips 164, corresponding to one of the MVEX frames 160.

[0303] SIDX frame 162 is an optional element of video file 150. That is, video files conforming to 3GPP file formats or other such file formats do not necessarily include SIDX frame 162. According to examples of 3GPP file formats, SIDX frames can be used to identify sub-segments of a segment (e.g., segments contained within video file 150). The 3GPP file format defines a sub-segment as "a self-contained set of one or more consecutive movie clip frames with corresponding media data frames and media data frames containing data referenced by the movie clip frames that must follow the movie clip frame and precede the next movie clip frame containing information about the same track." The 3GPP file format also states that a SIDX frame "contains a sequence of references to sub-segments within the (sub)segment recorded by the frame. The referenced sub-segments are sequential in presentation time. Similarly, the bytes referenced by the segment index frame are always sequential within the segment. The size of the reference gives a count of the number of bytes in the referenced material."

[0304] SIDX frame 162 typically provides information representing one or more sub-segments of a segment included in video file 150. Such information may include, for example, the playback times at the start and / or end of the sub-segment, the byte offset for the sub-segment, whether the sub-segment includes (e.g., begins at) a Stream Access Point (SAP), the type of SAP (e.g., whether the SAP is an Instant Decoder Refresh (IDR) picture, a Clean Random Access (CRA) picture, a Broken Link Access (BLA) picture, etc.), the position of the SAP within the sub-segment (in terms of playback time and / or byte offset), etc.

[0305] Movie clip 164 may include one or more decoded video images. In some examples, movie clip 164 may include one or more groups of images (GOPs), each group of images may include multiple decoded video images, such as frames or pictures. Furthermore, as mentioned above, in some examples, movie clip 164 may include a sequence dataset. Each movie clip 164 may include a Movie Clip Header Box (MFHD). Figure 4 (Not shown in the image). The MFHD frame can describe the characteristics of the corresponding movie clip, such as the movie clip's sequence number. Movie clips 164 can be included in video file 150 in the order of their sequence numbers.

[0306] MFRA frame 166 can describe random access points within movie segments 164 of video file 150. This can help perform trick patterns, such as searching for specific time positions (i.e., playback times) within segments encapsulated by video file 150. In some examples, MFRA frame 166 is typically optional and does not need to be included in the video file. Similarly, client devices such as client device 40 do not necessarily need to refer to MFRA frame 166 to correctly decode and display the video data of video file 150. MFRA frame 166 may include a number of Track Segmented Random Access (TFRA) frames (not shown) equal to the number of tracks in video file 150, or in some examples, a number of Track Segmented Random Access (TFRA) frames (not shown) equal to the number of media tracks (e.g., non-cue tracks) in video file 150.

[0307] In some examples, movie clip 164 may include one or more streaming access points (SAPs), such as IDR pictures. Similarly, MFRA frame 166 may provide a positional indication within the SAP video file 150. Therefore, a temporal subsequence of video file 150 can be formed from the SAP of video file 150. The temporal subsequence may also include other pictures, such as SAP-dependent P-frames and / or B-frames. Frames and / or segments of the temporal subsequence can be arranged within segments such that frames / slices of the temporal subsequence that depend on other frames / slices of the subsequence can be appropriately decoded. For example, in a hierarchical arrangement of data, data used to predict other data may also be included in the temporal subsequence.

[0308] Figure 5 This is a conceptual diagram illustrating an example system 180 for performing low-latency streaming using chunking and Addressable Resource Index (ARI) tracks according to the technology of this disclosure. In this example, system 180 includes a DASH packetizer 182, a Content Delivery Network (CDN) 184, a regular DASH client 186, and a low-latency DASH client 188. In this example, the DASH packetizer 182 forms DASH segments and chunks from CMAF format media data. The DASH packetizer 182 may correspond to... Figure 1 Content preparation equipment 20 or packaging unit 30.

[0309] DASH Packer 182 can create all metadata and add metadata to stream ARI metadata tracks, such as Figure 6 As shown, this will be discussed in more detail below. Each set of chunks at the same media time (e.g., presentation time) generates a sample in the metadata track. An interesting aspect is that the release latency of metadata can be handled in a flexible way, and then there are issues related to general streaming optimization:

[0310] ● More or fewer segments

[0311] ● More or fewer requests

[0312] ● More or fewer chunks

[0313] ●Scalability

[0314] ● etc.

[0315] This information can also be used in the exact same way for live streaming, video-on-demand, and time-shift streaming operations. Content can also be removed based on a custom time-shift buffer.

[0316] Figure 6 This is a conceptual diagram illustrating an example dataset containing ARI orbitals according to the technology disclosed herein. Figure 6 This demonstrates basic concepts regarding the flexibility of designing ARI Track 190 (blocked, segmented, or long segments). The trade-offs are the latency and delay of the live service that can utilize this information, as well as potential overhead.

[0317] Additional information can be added to the metadata track.

[0318] If the information is provided appropriately, it can be used not only by the DASH client but also by network nodes, for example, for:

[0319] ●Resource allocation in the context of 5G streaming

[0320] ●Prefilled cache

[0321] ● etc.

[0322] This data can be used, for example, in fifth-generation (5G) media. Media session handlers and 5GMSd AF coordinate support for media streaming sessions, for example, by allocating the necessary QoS to the PDU session and recommending available bitrates to the media player for appropriate adaptation.

[0323] Currently, QoS allocation is performed statically, where information about session requirements is read once from a manifest file (such as DASH MPD) and then shared with the media session handler.

[0324] However, as mentioned above, the content is typically encoded in a capped VBR manner, and the bit rate varies greatly depending on the complexity of the current scenario.

[0325] To provide more dynamic QoS allocation, 5GMSd can receive the location of CMAF addressable resource index metadata and stream it to obtain current information about the actual and predicted media bitrates and adjust QoS allocation, for example, by coordinating with the PCF or sending bitrate recommendation queries to the RAN via MSH. For this purpose, the player needs to provide the following information to the MSH and then to the 5GMSd AF: ● The URL used to access the CMAF addressable resource index metadata.

[0326] ● Current playback position in the media timeline

[0327] ● Current operation point, i.e., which representations are being consumed.

[0328] The media session handler uses this information to perform RAN-based assistance and also passes it to 5GMSdAF via M5.

[0329] CMAF addressable resource index metadata must be accessible to external applications such as MSH and 5GMSd AF, without an authorized media streaming session (an authorized media streaming session has the authority to access the actual content itself). Even without access to the actual media segments, index blocks can reveal content-related information that content providers and users may be interested in.

[0330] The URL signature is used by the player application to authorize the MSH and AF to access the metadata only during the lifetime of that streaming session. Access verification is performed by the content provider. The URL may contain a shared secret between the trusted player and the content provider, as well as other information such as the client_id or IP address of the UE or 5GMSd that is allowed to access the metadata.

[0331] Figure 7 This is a conceptual diagram illustrating examples of boxes that can be included in DASH segments and blocks. For example... Figure 7 As shown, a DASH segment can include a MOOF box and an MDAT box. When a segment is divided into chunks, each chunk can include a MOOF box and an MDAT box.

[0332] Figure 8 This is a flowchart illustrating an example method for providing Addressable Resource Information (ARI) tracks from a server device to a client device according to the technology of this disclosure. For example, it can be done through... Figure 1 Content preparation device 20 and / or server device 60 and client device 40 to perform Figure 8 The method, as described above, allows a single device to be configured to perform functions attributed to the content preparation device 20 and the server device 60.

[0333] Initially, the content preparation device 20 can form time-aligned segments and chunks (200) for media presentation. For example, the content preparation device 20 can form such as Figure 3 The content preparation device 20 can form multiple representations, each including corresponding segments and chunks. Each representation can correspond to a media track of Common Media Application Format (CMAF) data. Various media tracks can form switch sets, such as CMAF switch sets. The content preparation device 20 can also form ARI tracks describing chunks (202) and manifest files (204) describing media presentation. In particular, an ARI track can be a single index track, which can be a single representation of an adaptation set. The manifest file can signal the data for each representation and ARI track. The manifest file can be a DASH Media Presentation Description (MPD). The content preparation device 20 can provide the manifest file, ARI tracks, and media data to the server device 60.

[0334] Client device 40 may send a request for a manifest file to server device 60 (206). Server device 60 may receive the request for the manifest file and, in response, send the manifest file to client device 40 (208). Client device 40 may then receive the manifest file (210). Client device 40 may process the manifest file to determine the address of the ARI track and then request ARI track data from server device 60 (212).

[0335] Server device 60 may receive an ARI orbit data request from client device 40 and, in response, send ARI orbit data to client device 40 (214). Client device 40 may then receive the requested ARI orbit data (216).

[0336] Figure 9 This is a flowchart illustrating an example method for retrieving media data using Addressable Resource Information (ARI) tracks according to the technology of this disclosure. For example, it can be done through... Figure 1 Content preparation device 20 and / or server device 60 and client device 40 to perform Figure 9 The method, as described above, allows a single device to be configured to perform functions attributed to the content preparation device 20 and the server device 60. Figure 9 The method can be executed Figure 8 The method will be executed after that.

[0337] After retrieving the manifest file (e.g., DASH MPD) and ARI tracks, client device 40 can determine the operating mode (220). The operating mode can be, for example, low-latency live streaming, live streaming, time-shifted streaming, or video-on-demand (VoD). Alternatively, client device 40 can determine other streaming characteristics, such as the target latency of the media, network conditions, and / or the desired content quality. Client device 40 can also determine the playback time (222) during which media data will be retrieved.

[0338] Client device 40 can then process the manifest file and ARI tracks (224), for example, to determine the characteristics of chunks corresponding to playback times. For example, client device 40 can determine the duration and size of each chunk corresponding to a playback time from a sample of ARI tracks corresponding to playback times (226). Client device 40 can then select one of the chunks for playback time retrieval based on the determined characteristics of the chunks and the operating mode and / or streaming characteristics (228). For example, for low-latency live operation, client device 40 may prioritize small chunks with short durations, while for time-shift or VoD operating modes, longer and / or larger chunks may be preferred.

[0339] Client device 40 can then construct a request (230) for the selected chunk. For example, client device 40 can determine the byte range of the selected chunk based on the data of samples from the ARI track corresponding to the chunk and adjacent samples from the ARI track. Each sample from the ARI track can indicate a byte offset relative to the start of the corresponding segment. Therefore, client device 40 can determine the byte range based on the byte offset of the samples from the ARI track. Client device 40 can further determine the URL of the segment that includes the chunk from the manifest file. Therefore, client device 40 can construct an HTTP partial GET request specifying the URL of the segment and the byte range determined based on the byte offset.

[0340] Client device 40 can then send the request to server device 60 (232). Server device 60 can receive the request and send the requested chunk to client device 40 (234). Client device 40 can receive the chunk (236), decode the media data of the chunk (238), and present the media data (240).

[0341] In this way, Figure 8 and Figure 9 The method combination represents an example of a method that includes the following operations: retrieving data from an addressable resource information (ARI) track of media presentation, the ARI track data describing a subset of a switch set of media presentation and addressable resources, the switch set including multiple media tracks containing addressable resources, an ARI track being a single index track of media presentation, and addressable resources including retrievable media data; determining the duration and size of addressable resources based on the ARI track data; using the ARI track data including the duration and size of addressable resources to determine one or more addressable resources to retrieve; and retrieving the determined addressable resources.

[0342] The following clauses describe various examples of the technology disclosed herein.

[0343] Clause 1: A method for retrieving media data, the method comprising: retrieving data from an Addressable Resource Information (ARI) track of the media data, the data from the ARI track describing a subset of a switch set and addressable resources, and the ARI track being a single index track, the addressable resources including retrievable media data; and retrieving one or more addressable resources using the data from the ARI track.

[0344] Clause 2: The method of Clause 1, wherein the switch set includes the Common Media Application Format (CMAF) switch set.

[0345] Clause 3: The method of any of Clauses 1 and 2, wherein the addressable resource includes one or more segments, fragments, or blocks.

[0346] Clause 4: The method of any one of Clauses 1-3, wherein the ARI track is time-aligned with the switching set.

[0347] Clause 5: The method of any one of Clauses 1-4, wherein the ARI track data includes the attributes of all tracks in the switch set.

[0348] Clause 6: The method of any one of Clauses 1-5, wherein the ARI track includes header information, the method further including processing the header information of the ARI track.

[0349] Clause 7: The method of Clause 6, wherein the header information includes one or more of the following: the number of tracks in the switch set, the switch set identifier value, or a time scale that is the same as the time scale of the tracks in the switch set.

[0350] Clause 8: The method of any one of Clauses 1-7, wherein the ARI track comprises samples of each block of a time-aligned switching set.

[0351] Clause 9: The method of any one of Clauses 1-8, wherein for at least one addressable resource, the data of the ARI track includes one or more of the following: an indication of whether the resource is the start of a segment, whether the resource begins with a stream access point, the type of the stream access point, the offset of the resource, the size of the resource, the quality of the resource, the number of prediction pairs signaled, or a set of one or more prediction pairs, each prediction pair including a window value and a bit rate value.

[0352] Clause 10: The method of any one of Clauses 1-9 further includes retrieving a manifest file of media data, the manifest file including one or more of the following: an indication that an ARI track is provided as an adapter set having a single track or an indication of one or more of a switch set associated with said ARI track.

[0353] Clause 11: The method of Clause 10, wherein the manifest document includes a Media Presentation Description (MPD).

[0354] Clause 12: An apparatus for retrieving media data, the apparatus comprising one or more units for performing the methods of any one of Clauses 1-11.

[0355] Clause 13: The device of Clause 12, wherein the device includes at least one of the following: an integrated circuit; a microprocessor; or a wireless communication device.

[0356] Clause 14: A computer-readable storage medium having instructions stored thereon that, when executed, cause a processor to perform any of the methods described in Clauses 1-11.

[0357] Clause 15: An apparatus for retrieving media data, the apparatus comprising: a unit for retrieving data of an Addressable Resource Information (ARI) track of the media data, a subset of a data description switching set of the ARI track and addressable resources, wherein the ARI track is a single index track and the addressable resources include retrievable media data; and a unit for retrieving one or more addressable resources using the data of the ARI track.

[0358] Clause 16: A method for retrieving media data, the method comprising: retrieving data of an addressable resource information (ARI) track for media presentation, the data of the ARI track describing a subset of a switch set of the media presentation and addressable resources, the switch set including a plurality of media tracks, the media tracks including the addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; determining the duration and size of the addressable resources based on the data of the ARI track; using the data of the ARI track including the duration and size of the addressable resources to determine one or more of the addressable resources to be retrieved; and retrieving the determined addressable resources.

[0359] Clause 17: The method of Clause 16, wherein the switch set includes the Common Media Application Format (CMAF) switch set.

[0360] Clause 18: The method of Clause 16, wherein each of the media tracks has a common segmentation, fragmentation, and chunking structure.

[0361] Clause 19: The method of Clause 16, wherein the addressable resource includes one or more of segments, fragments or blocks.

[0362] Clause 20: The method of Clause 16, wherein the ARI track is time-aligned with the switching set.

[0363] Clause 21: The method of Clause 16, wherein the media track has a common chunk structure such that each media track includes a playback time-aligned chunk, and wherein the ARI track includes a sample of each of the playback time-aligned chunks.

[0364] Clause 22: The method of Clause 21, wherein each sample of the ARI track describes a corresponding playback time-aligned block for each of the media tracks.

[0365] Clause 23: The method of Clause 22, wherein determining one or more of the addressable resources to be retrieved comprises: determining a playback time of an addressable resource among the one or more of the addressable resources to be retrieved during the period; determining a sample of the samples of the ARI track corresponding to the playback time; determining the duration and size of the addressable resources of the playback time-aligned chunks of each media track based on the determined sample of the samples of the ARI track; and determining at least one of the addressable resources to be retrieved during the determined playback time based on the determined duration and size.

[0366] Clause 24: The method of Clause 16, wherein determining the one or more addressable resources includes determining the one or more addressable resources having a duration and size that satisfy a target duration and size according to an operating mode, the operating mode including one of: low-latency live streaming, live streaming, time-shifted, or video on demand (VoD).

[0367] Clause 25: The method of Clause 16, wherein the data of the ARI track includes the attributes of all tracks in the switch set.

[0368] Clause 26: The method of Clause 16, wherein the ARI track includes header information, the method further comprising processing the header information of the ARI track.

[0369] Clause 27: The method of Clause 16, wherein the header information includes one or more of the following: the number of tracks in the switching set, the switching set identifier value, or a time scale that is the same as the time scale of the tracks in the switching set.

[0370] Clause 28: The method of Clause 16, wherein the ARI track includes samples for each block of the switch set in a time-aligned manner.

[0371] Clause 29: The method of Clause 16, wherein for at least one of the addressable resources, the data of the ARI track includes one or more of the following: an indication of whether the resource is the start of a segment, whether the resource begins with a stream access point, the type of the stream access point, the offset of the resource, the size of the resource, the quality of the resource, the number of prediction pairs notified by signaling, or a group of one or more prediction pairs, wherein each prediction pair includes a window value and a bit rate value.

[0372] Clause 30: The method of Clause 16 further includes retrieving a manifest file of the media data, the manifest file including one or more of an indication that the ARI track is provided as an adapter set having a single track or an indication of one or more of a switch set associated with the ARI track.

[0373] Clause 31: The method of Clause 30, wherein the manifest file includes a Media Presentation Description (MPD).

[0374] Clause 32: An apparatus for retrieving media data, the apparatus comprising: a memory configured to store media data; and one or more processors implemented in circuitry and configured to: retrieve data of an Addressable Resource Information (ARI) track for media presentation, the data of the ARI track describing a subset of a switch set of media presentation and addressable resources, the switch set including a plurality of media tracks, each media track including addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; determine the duration and size of the addressable resources based on the data of the ARI track; determine one or more of the addressable resources to be retrieved using the data of the ARI track including the duration and size of the addressable resources; retrieve the determined addressable resources; and store the retrieved addressable resources in the memory.

[0375] Clause 33: The device of Clause 32, wherein the switching set includes the Common Media Application Format (CMAF) switching set.

[0376] Clause 34: The device of Clause 32, wherein each of the media tracks has a common segmentation, fragmentation and chunking structure.

[0377] Clause 35: The device of Clause 32, wherein the addressable resource includes one or more segments, fragments or blocks.

[0378] Clause 36: The device of Clause 32, wherein the ARI track is time-aligned with the switching set.

[0379] Clause 37: The device of Clause 32, wherein the media track has a common chunk structure such that each media track includes a playback time-aligned chunk, and wherein the ARI track includes a sample of each of the playback time-aligned chunks.

[0380] Clause 38: The device of Clause 37, wherein each sample of the ARI track describes a corresponding playback time-aligned block for each of the media tracks.

[0381] Clause 39: The apparatus of Clause 38, wherein, in order to determine the one or more addressable resources to be retrieved, the one or more processors are configured to: determine a playback time of an addressable resource among the one or more addressable resources to be retrieved during the period; determine a sample of the samples of the ARI track corresponding to the playback time; determine the duration and size of the addressable resources of the playback time-aligned chunks of each media track based on the determined sample of the ARI track; and determine at least one of the addressable resources to be retrieved during the determined playback time based on the determined duration and size.

[0382] Clause 40: A device according to Clause 32, wherein the one or more processors are configured to determine one or more addressable resources having a duration and size that satisfy a target duration and size, depending on an operating mode, the operating mode including one of low-latency live streaming, live streaming, time-shifted streaming, or video on demand (VoD).

[0383] Clause 41: The device of Clause 32, wherein the data of the ARI track includes the attributes of all tracks of the switch set.

[0384] Clause 42: The device of Clause 32, wherein the ARI track contains header information, and wherein the one or more processors are further configured to process the header information of the ARI track.

[0385] Clause 43: The device of Clause 42, wherein the header information includes one or more of the following: the number of tracks in the switch set, the switch set identifier value, or a time scale that is the same as the time scale of the tracks in the switch set.

[0386] Clause 44: The device pursuant to Clause 32, wherein the ARI track comprises samples for each block of the switching set in a time-aligned manner.

[0387] Clause 45: The device of Clause 32, wherein for at least one of the addressable resources, the data of the ARI track includes one or more of the following: an indication of whether the resource is the start of a segment, whether the resource begins with a stream access point, the type of stream access point, the offset of the resource, the size of the resource, the quality of the resource, the number of prediction pairs notified by signaling, or a group of one or more prediction pairs, wherein each prediction pair includes a window value and a bit rate value.

[0388] Clause 46: The device of Clause 32, wherein the one or more processors are further configured to retrieve a manifest file of the media data, the manifest file containing one or more of an indication that the ARI track is provided as an adapter set having a single track or an indication of one or more of the switch sets associated with the ARI track.

[0389] Clause 47: Devices specified in Clause 46, wherein the manifest file includes a Media Presentation Description (MPD).

[0390] Clause 48: The equipment of Clause 32, wherein the equipment includes at least one of the following: an integrated circuit; a microprocessor; or a wireless communication device.

[0391] Clause 49: A computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor to: retrieve data of an addressable resource information (ARI) track for media presentation, the data of the ARI track describing a subset of a switching set of media presentation and addressable resources, the switching set including a plurality of media tracks, the media tracks including the addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; determine the duration and size of the addressable resources based on the data of the ARI track; determine one or more of the addressable resources to be retrieved using the data of the ARI track, including the duration and size of the addressable resources; and retrieve the determined addressable resources.

[0392] Clause 50: Computer-readable storage media of Clause 49, wherein the switching set includes the Common Media Application Format (CMAF) switching set.

[0393] Clause 51: A computer-readable storage medium of Clause 49, wherein each of the media tracks has a common segmentation, fragmentation, and chunking structure.

[0394] Clause 52: The computer-readable storage medium of Clause 49, wherein the addressable resource includes one or more segments, fragments, or blocks.

[0395] Clause 53: A computer-readable storage medium of Clause 49, wherein the ARI track is time-aligned with the switching set.

[0396] Clause 54: A computer-readable storage medium of Clause 49, wherein the media tracks have a common chunking structure such that each media track includes a playback time-aligned chunk, and wherein the ARI track includes a sample for each playback time-aligned chunk.

[0397] Clause 55: The computer-readable storage medium of Clause 54, wherein each sample of the ARI track describes a corresponding playback time-aligned block for each of the media tracks.

[0398] Clause 56: A computer-readable storage medium of Clause 55, wherein instructions for causing a processor to determine one or more of the addressable resources to be retrieved include instructions for causing the processor to perform the following operations: determining a playback time of an addressable resource among the one or more of the addressable resources to be retrieved during the period; determining a sample of the samples of the ARI track corresponding to the playback time; determining, based on the sample of the determined ARI track, the duration and size of the addressable resources of the playback time-aligned chunks of each media track; and determining, based on the determined duration and size, at least one of the addressable resources to be retrieved during the determined playback time.

[0399] Clause 57: A computer-readable storage medium of Clause 49, wherein instructions for causing the processor to determine the one or more addressable resources include instructions for causing the processor to determine, according to an operating mode, the one or more addressable resources having a duration and size that satisfy a target duration and size, the operating mode including one of low-latency live streaming, live streaming, time-shifted streaming, or video on demand (VoD).

[0400] Clause 58: The computer-readable storage medium of Clause 49, wherein the data of the ARI track includes attributes of all tracks of the switching set.

[0401] Clause 59: The computer-readable storage medium of Clause 49, wherein the ARI track includes header information, the method further comprising: processing the header information of the ARI track.

[0402] Clause 60: A computer-readable storage medium of Clause 59, wherein the header information includes one or more of the following: the number of tracks in the switching set, a switching set identifier value, or a time scale that is the same as the time scale of the tracks in the switching set.

[0403] Clause 61: A computer-readable storage medium of Clause 49, wherein the ARI track comprises samples for each block of the switching set in a time-aligned manner.

[0404] Clause 62: A computer-readable storage medium of Clause 49, wherein, for at least one of the addressable resources, the data of the ARI track includes one or more of the following: an indication of whether the resource is the start of a segment, whether the resource begins with a stream access point, the type of the stream access point, the offset of the resource, the size of the resource, the quality of the resource, the number of prediction pairs notified by signaling, or a group of one or more prediction pairs, wherein each prediction pair includes a window value and a bit rate value.

[0405] Clause 63: The computer-readable storage medium of Clause 49 further includes instructions for causing the processor to retrieve a manifest file of the media data, the manifest file including one or more of an indication that the ARI track is provided as an adapter set having a single track or an indication of one or more switch sets associated with the ARI track.

[0406] Clause 64: Computer-readable storage media of Clause 63, wherein the manifest file includes a media presentation description (MPD).

[0407] Clause 65: An apparatus for retrieving media data, the apparatus comprising: unit for retrieving data of an addressable resource information (ARI) track for media presentation, the data of the ARI track describing a subset of a switch set and addressable resources of the media presentation, the switch set including a plurality of media tracks, the media tracks including addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; unit for determining the duration and size of the addressable resources based on the data of the ARI track; unit for determining one or more of the addressable resources to be retrieved using the data of the ARI track including the duration and size of the addressable resources; and unit for retrieving the determined addressable resources.

[0408] Clause 66: A method for retrieving media data, the method comprising: retrieving data of an addressable resource information (ARI) track for media presentation, the data of the ARI track describing a subset of a switch set and addressable resources of the media presentation, the switch set including a plurality of media tracks, the media tracks including the addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; determining the duration and size of the addressable resources based on the data of the ARI track; using the data of the ARI track including the duration and size of the addressable resources to determine one or more of the addressable resources to be retrieved; and retrieving the determined addressable resources.

[0409] Clause 67: The method of Clause 66, wherein the switch set includes the Common Media Application Format (CMAF) switch set.

[0410] Clause 68: The method of any one of Clauses 66 and 67, wherein each of the media tracks has a common segmentation, fragmentation and chunking structure.

[0411] Clause 69: The method of any one of Clauses 66-68, wherein the addressable resource includes one or more of segments, fragments or blocks.

[0412] Clause 70: The method of any one of Clauses 66-69, wherein the ARI track is time-aligned with the switching set.

[0413] Clause 71: The method of any one of Clauses 66-70, wherein the media track has a common chunk structure such that each media track includes a playback time-aligned chunk, and wherein the ARI track includes a sample of each of the playback time-aligned chunks.

[0414] Clause 72: The method of Clause 71, wherein each sample of the ARI track describes a corresponding playback time-aligned block for each of the media tracks.

[0415] Clause 73: The method of Clause 72, wherein determining one or more of the addressable resources to be retrieved comprises: determining a playback time of an addressable resource among the one or more of the addressable resources to be retrieved during the period; determining a sample of the samples of the ARI track corresponding to the playback time; determining the duration and size of the addressable resources of the playback time-aligned chunks of each media track based on the determined sample of the samples of the ARI track; and determining at least one of the addressable resources to be retrieved during the determined playback time based on the determined duration and size.

[0416] Clause 74: The method of any one of Clauses 66-73, wherein determining the one or more addressable resources includes determining the one or more addressable resources having a duration and size that satisfy a target duration and size according to an operating mode, the operating mode including one of: low-latency live streaming, live streaming, time-shifted, or video on demand (VoD).

[0417] Clause 75: The method of any one of Clauses 66-74, wherein the data of the ARI track includes the attributes of all tracks in the switch set.

[0418] Clause 76: The method of any one of Clauses 66-75, wherein the ARI track includes header information, the method further comprising processing the header information of the ARI track.

[0419] Clause 77: The method of Clause 76, wherein the header information includes one or more of the following: the number of tracks in the switch set, the switch set identifier value, or a time scale that is the same as the time scale of the tracks in the switch set.

[0420] Clause 78: The method of any one of Clauses 66-77, wherein the ARI track comprises samples for each block of the switch set in a time-aligned manner.

[0421] Clause 79: The method of any one of Clauses 66-78, wherein for at least one of the addressable resources, the data of the ARI track includes one or more of the following: an indication of whether the resource is the start of a segment, whether the resource begins with a stream access point, the type of the stream access point, the offset of the resource, the size of the resource, the quality of the resource, the number of prediction pairs notified by signaling, or a group of one or more prediction pairs, wherein each prediction pair includes a window value and a bit rate value.

[0422] Clause 80: The method of any one of Clauses 66-79 further includes retrieving a manifest file of the media data, the manifest file including one or more of an indication that the ARI track is provided as an adapter set having a single track or an indication of one or more of a switch set associated with the ARI track.

[0423] Clause 81: The method of Clause 80, wherein the manifest document includes a Media Presentation Description (MPD).

[0424] Clause 82: An apparatus for retrieving media data, the apparatus comprising: a memory configured to store media data; and one or more processors implemented in circuitry and configured to: retrieve data of an Addressable Resource Information (ARI) track for media presentation, the data of the ARI track describing a subset of a switch set of media presentation and addressable resources, the switch set including a plurality of media tracks, each media track including addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; determine the duration and size of the addressable resources based on the data of the ARI track; determine one or more of the addressable resources to be retrieved using the data of the ARI track including the duration and size of the addressable resources; retrieve the determined addressable resources; and store the retrieved addressable resources in the memory.

[0425] Clause 83: The device of Clause 82, wherein the switching set includes the Common Media Application Format (CMAF) switching set.

[0426] Clause 84: A device of any of Clauses 82 and 83, wherein each of the media tracks has a common segmentation, fragmentation and chunking structure.

[0427] Clause 85: A device of any one of Clauses 82-84, wherein the addressable resource includes one or more segments, fragments or blocks.

[0428] Clause 86: A device of any of Clauses 82-85, wherein the ARI track is time-aligned with the switching set.

[0429] Clause 87: A device of any one of Clauses 82-86, wherein the media track has a common chunking structure such that each media track includes a playback time-aligned chunk, and wherein the ARI track includes a sample of each of the playback time-aligned chunks.

[0430] Clause 88: The device of Clause 87, wherein each sample of the ARI track describes a corresponding playback time-aligned block for each of the media tracks.

[0431] Clause 89: The apparatus of Clause 88, wherein, in order to determine the one or more addressable resources to be retrieved, the one or more processors are configured to: determine a playback time of an addressable resource among the one or more addressable resources to be retrieved during the period; determine a sample of the samples of the ARI track corresponding to the playback time; determine, based on the determined sample of the samples of the ARI track, the duration and size of the addressable resources of the playback time-aligned chunks of each media track; and determine, based on the determined duration and size, at least one of the addressable resources to be retrieved during the determined playback time.

[0432] Clause 90: A device of any one of Clauses 82-89, wherein the one or more processors are configured to determine, according to an operating mode, one or more addressable resources having a duration and size that satisfy a target duration and size, the operating mode including one of low-latency live streaming, live streaming, time-shifted streaming, or video on demand (VoD).

[0433] Clause 91: A device of any one of Clauses 82-90, wherein the data for the ARI orbital includes attributes of all orbitals in the switch set.

[0434] Clause 92: A device of any one of Clauses 82-91, wherein the ARI track contains header information, and wherein the one or more processors are further configured to process the header information of the ARI track.

[0435] Clause 93: The device of Clause 92, wherein the header information includes one or more of the following: the number of tracks in the switch set, the switch set identifier value, or a time scale that is the same as the time scale of the tracks in the switch set.

[0436] Clause 94: A device of any of Clauses 82-93, wherein the ARI track comprises samples for each block of the switching set in a time-aligned manner.

[0437] Clause 95: A device of any one of Clauses 82-94, wherein for at least one of the addressable resources, the data of the ARI track includes one or more of the following: an indication of whether the resource is the start of a segment, whether the resource begins with a stream access point, the type of the stream access point, the offset of the resource, the size of the resource, the quality of the resource, the number of prediction pairs notified by signaling, or a group of one or more prediction pairs, wherein each prediction pair includes a window value and a bit rate value.

[0438] Clause 96: A device of any one of Clauses 82-95, wherein the one or more processors are further configured to retrieve a manifest file of the media data, the manifest file containing one or more of an indication that the ARI track is provided as an adapter set having a single track or an indication of one or more of the switch sets associated with the ARI track.

[0439] Clause 97: Devices specified in Clause 96, wherein the manifest file includes a Media Presentation Description (MPD).

[0440] Clause 98: Equipment of any one of Clauses 82-97, wherein the equipment comprises at least one of: an integrated circuit; a microprocessor; or a wireless communication device.

[0441] Clause 99: A computer-readable storage medium having instructions stored thereon, which, when executed, cause a processor to: retrieve data of an addressable resource information (ARI) track for media presentation, the data of the ARI track describing a subset of a switching set of media presentation and addressable resources, the switching set including a plurality of media tracks, the media tracks including the addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; determine the duration and size of the addressable resources based on the data of the ARI track; determine one or more of the addressable resources to be retrieved using the data of the ARI track, including the duration and size of the addressable resources; and retrieve the determined addressable resources.

[0442] Clause 100: Computer-readable storage media of Clause 99, wherein the switching set includes the Common Media Application Format (CMAF) switching set.

[0443] Clause 101: A computer-readable storage medium of any of Clauses 99 and 100, wherein each of the media tracks has a common segmentation, fragmentation, and chunking structure.

[0444] Clause 102: A computer-readable storage medium of any of Clauses 99-101, wherein the addressable resource includes one or more segments, fragments, or blocks.

[0445] Clause 103: A computer-readable storage medium of any one of Clauses 99-102, wherein the ARI track is time-aligned with the switching set.

[0446] Clause 104: A computer-readable storage medium of any one of Clauses 99-103, wherein the media tracks have a common chunking structure such that each media track includes a playback time-aligned chunk, and wherein the ARI track includes a sample for each playback time-aligned chunk.

[0447] Clause 105: A computer-readable storage medium of Clause 104, wherein each sample of the ARI track describes a corresponding playback time-aligned block for each of the media tracks.

[0448] Clause 106: A computer-readable storage medium of Clause 105, wherein instructions for causing a processor to determine one or more of the addressable resources to be retrieved include instructions for causing the processor to perform the following operations: determining a playback time of an addressable resource among the one or more of the addressable resources to be retrieved during the period; determining a sample of the samples of the ARI track corresponding to the playback time; determining, based on the sample of the determined ARI track, the duration and size of the addressable resources of the playback time-aligned chunks of each media track; and determining, based on the determined duration and size, at least one of the addressable resources to be retrieved during the determined playback time.

[0449] Clause 107: A computer-readable storage medium of any one of Clauses 99-106, wherein instructions for causing the processor to determine the one or more addressable resources include instructions for causing the processor to determine, according to an operating mode, the one or more addressable resources having a duration and size that satisfy a target duration and size, the operating mode including one of low-latency live streaming, live streaming, time-shifted streaming, or video on demand (VoD).

[0450] Clause 108: A computer-readable storage medium of any one of Clauses 99-107, wherein the data of the ARI track includes attributes of all tracks in the switching set.

[0451] Clause 109: A computer-readable storage medium of any one of Clauses 99-108, wherein the ARI track includes header information and instructions for causing the processor to process the header information of the ARI track.

[0452] Clause 110: A computer-readable storage medium of Clause 109, wherein the header information includes one or more of the following: the number of tracks in the switching set, a switching set identifier value, or a time scale that is the same as the time scale of the tracks in the switching set.

[0453] Clause 111: A computer-readable storage medium of any one of Clauses 99-110, wherein the ARI track comprises samples for each block of the switching set in a time-aligned manner.

[0454] Clause 112: A computer-readable storage medium of any one of Clauses 99-111, wherein, for at least one of the addressable resources, the data of the ARI track includes one or more of the following: an indication of whether the resource is the start of a segment, whether the resource begins with a stream access point, the type of the stream access point, the offset of the resource, the size of the resource, the quality of the resource, the number of prediction pairs notified by signaling, or a group of one or more prediction pairs, wherein each prediction pair includes a window value and a bit rate value.

[0455] Clause 113: A computer-readable storage medium of any one of Clauses 99-112 further includes instructions for causing the processor to retrieve a manifest file of the media data, the manifest file including one or more of an indication that the ARI track is provided as an adapter set having a single track or an indication of one or more switch sets associated with the ARI track.

[0456] Clause 114: Computer-readable storage media of Clause 113, wherein the manifest file includes a media presentation description (MPD).

[0457] Clause 115: An apparatus for retrieving media data, the apparatus comprising: a unit for retrieving data of an addressable resource information (ARI) track for media presentation, the data of the ARI track describing a subset of a switch set and addressable resources of the media presentation, the switch set including a plurality of media tracks, the media tracks including addressable resources, the ARI track being a single index track of the media presentation, the addressable resources including retrievable media data; a unit for determining the duration and size of the addressable resources based on the data of the ARI track; a unit for determining one or more of the addressable resources to be retrieved using the data of the ARI track including the duration and size of the addressable resources; and a unit for retrieving the determined addressable resources.

[0458] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, including any medium that facilitates, for example, the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures to implement the techniques described in this disclosure. Computer program products may include computer-readable media.

[0459] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and is accessible by a computer. Furthermore, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, the definition of medium includes coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient, tangible storage media. Disks and optical discs as used herein include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0460] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques may be fully implemented in one or more circuit or logic elements.

[0461] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be combined in a codec hardware unit or provided by a collection of interoperable hardware units, including one or more processors as described above, along with suitable software and / or firmware.

[0462] Various examples have been described. These and other examples are within the scope of the appended claims.< / integer> < / integer> < / integer>

Claims

1. A method of retrieving media data, the method comprising: retrieving data of an addressable resource information (ARI) track of a media presentation, the ARI track being separate from a manifest file for the media presentation, the ARI track being divided into samples, each sample describing a subset of a switching set of the media presentation and a corresponding addressable resource, the switching set comprising a plurality of media tracks, the plurality of media tracks comprising the addressable resource, the plurality of media tracks being alternative representations of one another switchable for bandwidth adaptation, each of the plurality of media tracks having a respective bitrate, the ARI track being a single track of the media presentation separate from the plurality of media tracks, the addressable resource comprising retrievable media data; determining a duration and a size of the addressable resource from data of a respective sample of the ARI track; determining one or more of the addressable resources to retrieve using data of the respective sample of the ARI track including the duration and the size of the addressable resource; and retrieving the determined addressable resource.

2. The method of claim 1, wherein, the switching set comprises a common media application format (CMAF) switching set.

3. The method of claim 1, wherein, each of the media tracks has a common segment, fragment and chunk structure.

4. The method of claim 1, wherein, the addressable resource comprises one or more of a segment, a fragment or a chunk.

5. The method of claim 1, wherein, the samples of the ARI track are time-aligned with respective addressable resources of the switching set.

6. The method of claim 1, wherein, the media tracks have a common chunk structure such that each of the media tracks comprises playback time-aligned chunks, and wherein the ARI track comprises a sample of each of the playback time-aligned chunks.

7. The method of claim 6, wherein, each of the samples of the ARI track describes a corresponding playback time-aligned chunk of each of the media tracks.

8. The method of claim 7, wherein, determining one or more of the addressable resources to retrieve comprises: determining a playback time during which to retrieve an addressable resource of one or more of the addressable resources; determining one of the samples of the ARI track corresponding to the playback time; determining a duration and a size of the addressable resource of the playback time-aligned chunk of each of the media tracks from the determined one of the samples of the ARI track; and determining at least one of the addressable resources to retrieve within the determined playback time according to the determined duration and size.

9. The method of claim 1, wherein, determining one or more of the addressable resources to retrieve comprises determining one or more of the addressable resources having a duration and a size satisfying a target duration and size according to an operational mode, the operational mode comprising one of: low latency live, time-shifted or video on demand (VoD).

10. The method of claim 1, wherein, the data of the ARI track comprises properties of all tracks of the switching set.

11. The method of claim 1, wherein, the ARI track comprises header information, the method further comprising processing the header information of the ARI track.

12. The method of claim 11, wherein, the header information comprises one or more of: a number of tracks in the switching set, a switching set identifier value, or a same time scale as a time scale of tracks of the switching set.

13. The method of claim 1, wherein, The ARI track includes a sample of each chunk of the switching set in time-aligned fashion.

14. The method of claim 1, wherein, For at least one of the addressable resources, data of the ARI track includes one or more of an indication of whether the resource is a start of a segment, whether the resource starts with a stream access point, a type of the stream access point, an offset of the resource, a size of the resource, a quality of the resource, a number of signaled prediction pairs, or a set of one or more of the prediction pairs, wherein each of the prediction pairs includes a window value and a bitrate value.

15. The method of claim 1, further comprising retrieving the manifest file for the media data, the manifest file including one or more of an indication that the ARI track is provided as an adaptation set with the single track or an indication of one or more of the switching sets associated with the ARI track.

16. The method of claim 15, wherein, The manifest file includes a media presentation description (MPD).

17. A device for retrieving media data, the device comprising: a memory; and one or more processors coupled to the memory and configured to: retrieve data of an addressable resource information (ARI) track for a media presentation, the ARI track separate from a manifest file for the media presentation, the ARI track divided into samples, each sample describing a subset of a switching set of the media presentation and a corresponding addressable resource, the switching set including a plurality of media tracks, the plurality of media tracks including the addressable resource, the plurality of media tracks being alternative representations of one another switchable for bandwidth adaptation, each of the plurality of media tracks having a respective bitrate, the ARI track being a single track of the media presentation separate from the plurality of media tracks, the addressable resource including retrievable media data; determine a duration and a size of the addressable resource from data of a respective sample of the ARI track; determine one or more of the addressable resources to retrieve using data of the respective sample of the ARI track including the duration and the size of the addressable resource; retrieve the determined addressable resources; and store the retrieved addressable resources in the memory.

18. The apparatus of claim 17, wherein, The switching set includes a common media application format (CMAF) switching set.

19. The apparatus of claim 17, wherein, Each of the media tracks has a common segment, fragment, and chunk structure.

20. The apparatus of claim 17, wherein, The addressable resource includes one or more of a segment, a fragment, or a chunk.

21. The apparatus of claim 17, wherein, The samples of the ARI track are time-aligned with respective addressable resources of the switching set.

22. The apparatus of claim 17, wherein, The media tracks have a common chunk structure such that each of the media tracks includes playback time-aligned chunks, and wherein the ARI track includes a sample of each of the playback time-aligned chunks.

23. The apparatus of claim 22, wherein, Each of the samples of the ARI track describes a corresponding playback time-aligned chunk of each of the media tracks.

24. The apparatus of claim 23, wherein, To determine one or more of the addressable resources to retrieve, the one or more processors are configured to: determining a playback time at which to retrieve an addressable resource in one or more of the addressable resources; determining one of the samples of the ARI track corresponding to the playback time; determining, from the determined one of the samples of the ARI track, a duration and size of a chunk of the addressable resource to which the playback time of each of the media tracks is aligned; and determining, from the determined duration and size, at least one of the addressable resources to retrieve at the determined playback time.

25. The apparatus of claim 17, wherein, The one or more processors are configured to determine, from an operation mode, one or more of the addressable resources having a duration and size that satisfy a target duration and size, the operation mode comprising one of: low latency live, time-shifted, or video on demand, VoD.

26. The apparatus of claim 17, wherein, Data of the ARI track comprises properties of all tracks of the switching set.

27. The apparatus of claim 17, wherein, The ARI track comprises header information, and wherein the one or more processors are further configured to process the header information of the ARI track.

28. The apparatus of claim 27, wherein, The header information comprises one or more of: a number of tracks in the switching set, a switching set identifier value, or a same time scale as a time scale of tracks of the switching set.

29. The apparatus of claim 17, wherein, The ARI track comprises a sample of each chunk of the switching set in time-aligned manner.

30. The apparatus of claim 17, wherein, For at least one of the addressable resources, data of the ARI track comprises one or more of: an indication of whether the resource is a start of a segment, whether the resource starts with a stream access point, a type of the stream access point, an offset of the resource, a size of the resource, a quality of the resource, a number of signaled prediction pairs, or a set of one or more of the prediction pairs, wherein each of the prediction pairs comprises a window value and a bitrate value.

31. The apparatus of claim 17, wherein, The one or more processors are further configured to retrieve the manifest file of the media data, the manifest file comprising one or more of: an indication that the ARI track is provided as an adaptation set with the single track, or an indication of one or more of the switching sets associated with the ARI track.

32. The apparatus of claim 31, wherein, The manifest file comprises a media presentation description, MPD.

33. The apparatus of claim 17, wherein, The device comprises at least one of: an integrated circuit; a microprocessor; or a wireless communication device.

34. A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to: retrieving data of an addressable resource information, ARI, track of a media presentation, the ARI track being separate from a manifest file for the media presentation, the ARI track being divided into samples, each sample describing a subset of a switch set of the media presentation and a corresponding addressable resource, the switch set comprising a plurality of media tracks, the plurality of media tracks comprising the addressable resource, the plurality of media tracks being alternative representations of each other switchable for bandwidth adaptation, each of the plurality of media tracks having a respective bitrate, the ARI track being a single track of the media presentation separate from the plurality of media tracks, the addressable resource comprising retrievable media data; determining a duration and a size of the addressable resource from data of a respective sample of the ARI track; determining one or more of the addressable resources to retrieve using data of the respective sample of the ARI track including the duration and the size of the addressable resource; and retrieving the determined addressable resource.

35. The computer-readable storage medium of claim 34, wherein, the switch set comprises a common media application format, CMAF, switch set.

36. The computer-readable storage medium of claim 34, wherein, each of the media tracks has a common segment, fragment, and chunk structure.

37. The computer-readable storage medium of claim 34, wherein, the addressable resource comprises one or more of a segment, a fragment, or a chunk.

38. The computer-readable storage medium of claim 34, wherein, the samples of the ARI track are time-aligned with respective addressable resources of the switch set.

39. The computer-readable storage medium of claim 34, wherein, the media tracks have a common chunk structure such that each of the media tracks comprises playback time-aligned chunks, and wherein the ARI track comprises a sample of each of the playback time-aligned chunks.

40. The computer-readable storage medium of claim 39, wherein, each of the samples of the ARI track describes a corresponding playback time-aligned chunk of each of the media tracks.

41. The computer-readable storage medium of claim 40, wherein, the instructions to cause the processor to determine one or more of the addressable resources to retrieve comprise instructions to cause the processor to: determine a playback time of an addressable resource during which one or more of the addressable resources are to be retrieved; determine one of the samples of the ARI track corresponding to the playback time; determine a duration and a size of the addressable resource of a playback time-aligned chunk of each of the media tracks from the determined one of the samples of the ARI track; and determine at least one of the addressable resources to retrieve within the determined playback time according to the determined duration and size. the instructions to cause the processor to determine one or more of the addressable resources comprise instructions to cause the processor to determine one or more of the addressable resources having a duration and a size satisfying a target duration and size according to an operating mode, the operating mode comprising one of: low latency live, time-shifted, or video on demand, VoD.

42. The computer-readable storage medium of claim 34, wherein, the data of the ARI track comprises properties of all tracks of the switch set.

43. The computer-readable storage medium of claim 34, wherein, the ARI track comprises header information, the instructions further comprising instructions to cause the processor to process the header information of the ARI track.

44. The computer-readable storage medium of claim 34, wherein, ​ 45. The computer-readable storage medium of claim 44, wherein, The header information includes one or more of a number of tracks in the switching set, a switching set identifier value, or a time scale that is the same as a time scale of tracks in the switching set.

46. The computer readable storage medium of claim 34, wherein, The ARI track includes samples of each chunk of the switching set in time alignment.

47. The computer readable storage medium of claim 34, wherein, For at least one of the addressable resources, the data of the ARI track includes one or more of an indication of whether the resource is a start of a segment, whether the resource starts with a stream access point, a type of the stream access point, an offset of the resource, a size of the resource, a quality of the resource, a number of signaled prediction pairs, or a set of one or more of the prediction pairs, wherein each prediction pair includes a window value and a bitrate value.

48. The computer-readable storage medium of claim 34, further comprising instructions to cause the processor to retrieve the manifest file for the media data, the manifest file including one or more of an indication that the ARI track is provided as an adaptation set with the single track, or an indication of one or more of the switching sets associated with the ARI track.

49. The computer-readable storage medium of claim 48, wherein, The manifest file includes a media presentation description (MPD).

50. An apparatus for retrieving media data, the apparatus comprising: a unit for retrieving data of an addressable resource information (ARI) track for a media presentation, the ARI track being separate from a manifest file for the media presentation, the ARI track being divided into samples, each sample describing a subset of a switching set of the media presentation and a corresponding addressable resource, the switching set including a plurality of media tracks, the media tracks including the addressable resource, the plurality of media tracks being alternative representations of one another switchable for bandwidth adaptation, each of the plurality of media tracks having a respective bitrate, the ARI track being a single track of the media presentation separate from the plurality of media tracks, the addressable resource including retrievable media data; a unit for determining a duration and a size of the addressable resource from the data of the respective sample of the ARI track; a unit for determining one or more of the addressable resources to retrieve using the data of the respective sample of the ARI track including the duration and the size of the addressable resource; and a unit for retrieving the determined addressable resource.

Citation Information

Patent Citations

  • Content transmission device, content playback device, content distribution system, method for controlling content transmission device, method for controlling content playback device, control program, and recording medium

    EP2908535A1