Method, apparatus and computer program for dynamically encapsulating media content data - Patents.com
Patent Information
- Application Number
- JP2023563113
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-28
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-30
AI Technical Summary
Existing ISOBMFF file formats face limitations in supporting dynamic session changes, requiring complete media data description before encapsulation and analysis, which hinders efficient handling of changes in media structure and organization.
The method allows for dynamic encapsulation and analysis of media data by introducing metadata portions that describe tracks and sample entries within media fragments, enabling changes to be signaled during the presentation without prior complete description.
Enables efficient handling of dynamic session changes by allowing media data to be encapsulated and analyzed without requiring a complete description upfront, supporting changes in encoding formats, protection schemes, and adding new tracks or sample entries during the presentation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a method, apparatus and computer program for improving the encapsulation and parsing of media data, allowing for improved handling of changes in the structure and organization of the encapsulated media data. [Background technology]
[0002] The International Organization for Standardization Base Media File Format (ISOBMFF, ISO / IEC 14496-12) is a well-known, flexible and extensible format that describes encoded timed or non-timed media data or bitstreams for local storage, transmission over a network or other bitstream distribution mechanisms. There are several extensions to this file format, for example Part 15 of ISO / IEC 14496-15, which describes encapsulation tools for various NAL (Network Abstraction Layer) unit-based video coding formats. Examples of such coding formats include AVC (Advanced Video Coding), SVC (Scalable Video Coding), HEVC (High Efficiency Video Coding), L-HEVC (Layered HEVC), and VVC (Versatile Video Coding). Another example of a file format extension is ISO / IEC 23008-12, which describes encapsulation tools for still images or still image sequences, such as HEVC still images. Yet another example of a file format extension is ISO / IEC 23090-2, which defines the Omnidirectional Media Application Format (OMAF). Still other examples of file format extensions are ISO / IEC 23090-10 and ISO / IEC 23090-18, which define the transport of Visual Volumetric Video-based Coding (V3C) media data and Geometry-based Point Cloud Compression (G-PCC) media data.
[0003] The file format is object-oriented. It is composed of building blocks called boxes (or data structures where each box is identified by a four-letter code, also written as FourCC or 4CC). A full box is a data structure similar to a box, and additionally contains version and flag value attributes. In the following, the term box may refer to both full boxes or boxes. These boxes or full boxes are organized sequentially or hierarchically. They define parameters that describe the encoded timed or untimed media data or bitstreams, their structure, and associated timing (if any). In the following, encapsulated media data is considered to specify encapsulated data that contains metadata and media data (the latter refers to an encapsulated bitstream). All data in an encapsulated media file (media data and metadata describing the media data) is contained within boxes. There is no other data in the file. A file-level box is a box that is not contained within another box.
[0004] According to the file format, the overall presentation (or session) is called a movie. A movie is described in a movie box (identified by the four-letter code 'moov') at the top level of a media or presentation file. This movie box represents an initialization information container that contains a set of different boxes that describe the presentation. It may be logically divided into tracks represented by track boxes (identified by the four-letter code 'trak'). Each track (uniquely identified by a track identifier (track_ID)) represents a timed sequence of media data related to the presentation (e.g. a sequence of video frames or a sequence of subparts of video frames). Within each track, each timed unit of media data is called a sample. Such a timed unit may be a video frame or a subpart of a video frame, an audio sample, or a set of timed metadata. The samples are numbered in increasing order of decoding, implicitly. Each track box contains a hierarchy of boxes describing the samples of the corresponding track. Among the boxes in this hierarchy, the sample table box (identified by the four-letter code 'stbl') contains all the items of temporal information and data indexes of the media samples in the track. Each sample entry gives the necessary information about the encoding configuration of the media data in the sample (including the encoding type, which identifies the encoding format, various encoding parameters that characterize the encoding format), as well as the initialization information required to decode the sample. The actual sample data is stored in a box called a Media Data Box (identified by the four-letter code 'mdat') or in a box called an Identified Media Data Box (identified by the four-letter code 'imda' and similar to a Media Data Box but containing an additional identifier). Media Data Boxes and Identified Media Data Boxes are placed at the same level as the Movie Box.
[0005] Movies may also be fragmented, i.e. organised temporally as a Movie Box containing information about the overall presentation followed by a list of Movie Fragments, i.e. a list of couples containing a Movie Fragment Box (identified by the four letter code 'moof') and a Media Data Box ('mdat'), or a list of couples containing a Movie Fragment Box ('moof') and an identified Media Data Box ('imda').
[0006] FIG. 1 shows an example of encapsulated media data temporally organized as fragmented presentations within one or more media files according to the ISO Base Media File Format.
[0007] Media data encapsulated in one or more media files 100 begins with a FileTypeBox ('ftyp') box (not shown) that provides a set of brands that identify the exact specifications to which the encapsulated media data conforms, which a reading device uses to determine whether it can process the encapsulated media data. The 'ftyp' box is followed by a MovieBox ('moov') box, referenced 105. The MovieBox box provides initialization information necessary for a reading device to start processing the encapsulated media data. In particular, it provides information about a description of the presentation content, the number of tracks, and their respective timelines and characteristics. For illustrative purposes, a MovieBox box may indicate that the presentation contains one track with identifier track_ID equal to 1.
[0008] As shown, the MovieBox box 105 is followed by one or more movie fragments (also called media fragments), each of which contains metadata stored in a MovieFragmentBox ('moof') box and media data stored in a MediaDataBox ('mdat') box. For purposes of illustration, the one or more media files 100 contain a first movie fragment that contains and describes samples 1 through N of a track with track_ID equal to 1. This first movie fragment is comprised of a 'moof' box 110 and an 'mdat' box 115. This second movie fragment is comprised of a 'moof' box 120 and an 'mdat' box 125.
[0009] When the encapsulated media data is fragmented into multiple files, the FileTypeBox and MovieBox boxes (hereafter also referred to as initialization fragments) are included in the first media file whose track does not contain any samples (hereafter also referred to as initialization segment). Subsequent media files (hereafter also referred to as media segments) contain one or more movie fragments.
[0010] Among other information, the 'moov' box 105 may contain a MovieExtendsBox ('mvex') box 130. If present, the information contained in this box alerts the reading device that subsequent movie fragments may exist and that these must be found and scanned in a predefined order to obtain all samples of the track. To that end, the information contained in this box should be combined with other information in the MovieBox box. The MovieExtendsBox box 130 may contain an optional MovieExtendsHeaderBox ('mehd') box and one TrackExtendsBox ('trex') box for each track defined in the MovieBox box 105. When present, the MovieExtendsHeaderBox box provides the overall duration of the fragmented movie. Each TrackExtendsBox box defines default parameter values to be used for the associated track in the movie fragment.
[0011] As shown, the 'moov' box 105 also contains one or more TrackBox ('trak') boxes 135 that describe each track of the presentation. The TrackBox boxes 135 contain in their box hierarchy SampleTableBox ('stbl') boxes which in turn contain descriptive and timing information for the media samples of the track. In particular, they contain SampleDescriptionBox ('stsd') boxes which contain descriptive information about the encoding format of the samples (the encoding format is identified in the 4CC as shown by the 'xxxx' letters), and one or more sample entry boxes which provide initialization information needed to configure decoding according to the encoding format.
[0012] For example, a SampleEntry box with a four-character type set to 'vvc1' or 'vvi1' signals that the associated sample contains media data encoded according to the Versatile Video Coding (VVC) format, and a SampleEntry box with a four-character type set to 'hvc1' or 'hev1' signals that the associated sample contains media data encoded according to the High Efficiency Video Coding (HEVC) format. A SampleEntry box may contain other boxes containing information that applies to all samples associated with this SampleEntry box.
[0013] A sample is associated with a SampleEntry box via the sample_description_index parameter in the SampleToChunkBox('stsc') box in a SampleTableBox('stbl') box if the media file is a non-fragmented media file, in the TrackFragmentHeaderBox('tfhd') box in the TrackFragmentBox('traf') box of a MovieFragmentBox('moof') box, or in the TrackExtendsBox('trex') box of a MovieExtendsBox('mvex') box if the media file is fragmented.
[0014] According to the ISO Base Media File Format, all tracks and all sample entries of a presentation are defined in the 'moov' box 105 and cannot be declared later during the presentation.
[0015] It is observed that a movie fragment may contain samples of one or more of the tracks declared in the 'moov' box, but not necessarily all of the tracks. The MovieFragmentBox box 110 or 120 contains a TrackFragmentBox ('traf') box which contains a TrackFragmentHeaderBox ('tfhd') box (not shown) that provides an identifier (e.g., Track_ID=1) that identifies each track whose samples are contained in the movie fragment's 'mdat' box 115 or 125. Among other information, the 'traf' box contains one or more TrackRunBox ('trun') boxes that document a contiguous set of samples of the tracks within the movie fragment.
[0016] An ISOBMFF file or segment may contain multiple sets of encoded timed media data (also denoted as bitstreams or streams) or sub-parts of sets of encoded timed media data (also denoted as sub-bitstreams or sub-streams), forming multiple tracks. If a sub-part corresponds to one or a contiguous spatial portion of a video source captured over time (e.g. at least one rectangular region also known as a 'tile' or 'sub-picture' captured over time), the corresponding multiple tracks may be referred to as tile tracks or sub-picture tracks.
[0017] It should also be noted that ISOBMFF and its extensions consist of several grouping mechanisms to group tracks, static items, or samples and associate groups with group descriptions. Groups typically share common semantics and / or characteristics. For example, the MovieBox box 105 and / or the MovieFragmentBox boxes 110 and 120 may contain sample groups that associate properties with groups of samples of a track. A sample group, characterized by a grouping type, may be defined by two linked boxes: a SampleToGroupBox ('sbgp') box, which represents the assignment of samples to a sample group, and a SampleGroupDescriptionBox ('sgpd') box, which contains a sample group entry for each sample group that describes the properties of the group.
[0018] Although these ISOBMFF file formats have proven to be efficient, they have some limitations regarding dynamic session support for fragmented ISOBMFF files, and therefore require signaling some functionality of the encapsulation mechanism that a reading device should understand in order to parse and decode the encapsulated media data.
[0019] An example is the core definition of presentation. Another example concerns improvements to the signaling of dynamic session support.
[0020] Therefore, in order to parse and decode the encapsulated media data, it is necessary to signal some functionality of the encapsulation mechanism that the reading device should understand and help the reading device select the data to process. Summary of the Invention
[0021] The present invention is devised to address one or more of the above-mentioned concerns.
[0022] According to a first aspect of the present invention, there is provided a method for encapsulating media data, the encapsulated media data comprising media data and a metadata portion associated with a media fragment, the method being performed by a server, the method comprising: obtaining a first portion of media data, the first portion being organized into a first set of one or more media data tracks; encapsulating metadata describing one or more tracks of a first set of one or more tracks in a metadata portion and encapsulating a first portion of the media data in one or more media fragments; obtaining a second portion of the media data, the obtained second portion being organized into a second set of one or more media data tracks; encapsulating in one media fragment the metadata describing the second portion of the media data and at least one track of the second set of one or more tracks if at least one track of the second set of one or more tracks differs from the one or more tracks of the first set; Includes.
[0023] Thus, the method of the present invention allows for dynamic encapsulation of media data without requiring a complete description of the media data before encapsulation can begin.
[0024] According to some embodiments, the method further comprises: obtaining a third portion of the media data, the obtained third portion being organized into a third set of one or more media data tracks; if at least one track of the second set of one or more tracks belongs to a third set of one or more media data tracks, encapsulating in one media fragment metadata describing the third portion of the media data and at least one track of the second set of one or more tracks; Includes.
[0025] According to some embodiments, encapsulating the metadata describing the third portion of the media data and at least one track of the second set of one or more tracks in one media fragment comprises copying the metadata describing the at least one track of the second set of one or more tracks from the media fragment encapsulating the second portion to the media fragment encapsulating the third portion.
[0026] According to some embodiments, the metadata portion includes an indication that the media fragment may include one or more tracks different from the first set of tracks.
[0027] According to some embodiments, the metadata of the media fragment encapsulating the second portion of the media data includes an indication signaling that the media fragment encapsulating the second portion of the media data includes a track that is different from one or more tracks of the first set of tracks.
[0028] According to some embodiments, the first portion of media data is encoded according to at least one first encoding configuration defined in a set of one or more first sample entries, the method further comprising: encapsulating metadata describing one or more first sample entries in a metadata portion; obtaining a fourth portion of media data, the fourth portion of media data being encoded according to at least one second encoding configuration; encapsulating in one media fragment the fourth portion of the media data and metadata describing the at least one second sample entry defining the at least one second encoding configuration if the at least one second encoding configuration differs from an encoding configuration defined in the set of one or more first sample entries; Includes.
[0029] According to some embodiments, the fourth portion corresponds to the second portion.
[0030] According to some embodiments, the metadata portion includes an indication signaling that the media fragment may include a different sample entry than the sample entry in the set of one or more first sample entries, or the metadata of the media fragment encapsulating the fourth portion of the media data includes an indication signaling that the metadata of the media fragment encapsulating the fourth portion of the media data includes a different sample entry than the sample entry in the set of one or more first sample entries.
[0031] According to some embodiments, at least one track of the second set of one or more tracks comprises a reference to at least one other track, the at least one other track being described in the metadata part or in the metadata of the media fragment.
[0032] According to a second aspect of the present invention, there is provided a method for parsing encapsulated media data, the encapsulated media data comprising a metadata portion associated with media data and a media fragment, the method being performed by a client, the method comprising: obtaining metadata describing one or more tracks of the first set of one or more tracks from the metadata portion; obtaining a media fragment, called a first media fragment; Parsing the first media fragment to obtain metadata; parsing the first media fragment to obtain a portion of media data if the metadata obtained from the first media fragment describes at least one track different from the one or more tracks of the first set, the portion of media data being organized into a second set of one or more media data tracks comprising the at least one track; Includes.
[0033] Thus, the method of the present invention allows for analysis of dynamically encapsulated media data without requiring a complete description of the media data before encapsulation begins.
[0034] According to some embodiments, the method further comprises: Obtaining a second media fragment; Parsing the second media fragment to obtain metadata; and parsing the second media fragment to obtain a portion of media data if the metadata of the second media fragment comprises a description of at least one track different from the same one or more tracks of the first set as the metadata of the first media fragment, where the media data obtained from the first portion of the media data and the media data obtained from the second portion of the media data belong to the same at least one track; Includes.
[0035] According to some embodiments, the method further comprises obtaining, from the metadata portion, an indication that the media fragment may include a track different from one or more tracks of the first set of tracks.
[0036] According to some embodiments, the method further includes obtaining an indication from metadata of the obtained media fragment indicating that the media fragment including the indication includes a track different from the one or more tracks of the first track set.
[0037] According to some embodiments, the method further comprises: obtaining a third media fragment; parsing the third media fragment to obtain metadata; if the metadata of the third media fragment includes metadata describing at least one sample entry defining at least one encoding configuration, parsing at least a portion of the media data of the third media fragment according to the at least one encoding configuration to obtain the media data; Includes.
[0038] According to some embodiments, the third portion corresponds to the first portion.
[0039] According to some embodiments, the method further comprises obtaining from the metadata portion an indication indicating that the media fragment may describe a different sample entry than the sample entry described in the metadata portion, or obtaining from the metadata of the obtained media fragment an indication indicating that the media fragment including the indication describes a different sample entry than the sample entry described in the metadata portion.
[0040] According to a third aspect of the present invention there is provided a method of encapsulating media data, the method being performed by a server, comprising: identifying a portion of the media data or a set of information items related to the portion of the media data according to an encapsulation independent parameter; encapsulating a set of media data or items of information as entities within a media file, the entities being grouped into a set of entities associated with a first representation representative of the parameters; Including, The media file includes a second representation that signals to the client that the set of entities will be parsed only if the client has knowledge of the first representation.
[0041] Thus, the method of the present invention makes it possible to signal certain functionality of the encapsulation mechanism that a reading device must understand in order to parse and decode the encapsulated media data, and helps the reading device select the data to process.
[0042] According to a fourth aspect of the present invention there is provided a method of analysing encapsulated media data, the method being performed by a client comprising: determining that the encapsulated media data includes a second instruction signaling that the set of entities is to be parsed only if the client has knowledge of the first instruction related to the set of entities to be parsed; Obtaining a reference to a set of entities to be parsed; Obtaining a first indication associated with the set of entities from which the reference was obtained; if the client does not have knowledge of a retrieved first representation associated with the set of entities for which the references were retrieved, then ignoring the set of entities for which the references were retrieved when parsing the encapsulated media data; Includes.
[0043] Thus, the method of the present invention enables a reading device to understand some of the functionality of the encapsulation mechanism used to generate encapsulated media data, to parse and decode the encapsulated media data, and to select the data to be processed.
[0044] According to some embodiments, the second representation further indicates that the media data can only be rendered if the client has knowledge of the first representation.
[0045] According to some embodiments, an entity is a sample or supplemental information describing the media data, for example a set of samples, such as a sample group, an entity group, a track, a sample entry, etc.
[0046] According to some embodiments, the entities are samples and the set of entities is a set of corrupted samples.
[0047] According to another aspect of the present invention, there is provided a processing device comprising a processing unit configured to perform each of the steps of the above-mentioned method. The other aspects of the present disclosure have any of the features and advantages similar to the first, second, third and fourth aspects described above.
[0048] At least some of the methods according to the invention may be computer implemented. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generally herein as a "circuit," "module," or "system." Furthermore, the present invention may take the form of a computer program product embodied in any tangible medium of expression having computer usable program code embodied in the medium.
[0049] Since the present invention can be implemented in software, the present invention can be embodied as computer readable code for provision to a programmable device on any suitable carrier medium. The tangible carrier medium can consist of a storage medium such as a floppy disk, a CD-ROM, a hard disk drive, a magnetic tape device, or a solid-state memory device. The transitory carrier medium can include signals such as electrical, electronic, optical, acoustic, magnetic, or electromagnetic signals, e.g., microwave and RF signals. [Brief description of the drawings]
[0050] Embodiments of the invention will now be described, by way of example only, with reference to the following drawings in which:
[0051] [Figure 1] 1 shows an example of a structure of fragmented media data encapsulated according to the ISO Base Media File Format. [Diagram 2] 1 illustrates an example of a system in which some embodiments of the present invention may be implemented. [Diagram 3] FIG. 1 illustrates an example of a fragmented presentation encapsulated within one or more media files, where new tracks are defined within a movie fragment in accordance with some embodiments of the present invention. [Figure 4] FIG. 1 illustrates an example of a fragmented presentation encapsulated within one or more media files, where sample entries and / or tracks are defined within a movie fragment in accordance with some embodiments of the present invention. [Diagram 5] FIG. 2 is a block diagram illustrating an example of steps performed by a server or writing device to encapsulate encoded media data according to some embodiments of the present invention. [Figure 6] 2 is a block diagram illustrating an example of steps performed by a client or reading device to process encapsulated media data according to some embodiments of the present invention. [Figure 7] 2 is a block diagram illustrating example steps performed by a client or reader to obtain data according to some embodiments of the present invention. [Figure 8] 1 illustrates a schematic diagram of a processing device configured to implement at least one embodiment of the present invention; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0052] According to some embodiments of the present invention, track and / or sample entries can be dynamically signaled in a movie fragment, without being pre-described in an initialization fragment.
[0053] FIG. 2 illustrates an example of a system in which some embodiments of the present invention may be implemented.
[0054] As shown, a server or writing device referenced at 200 is connected via a network interface (not shown) to a communications network 230 to which a client or reading device 250 is also connected via a network interface (not shown), enabling the server or writing device 200 and the client or reading device 250 to exchange media files referenced at 225 via the communications network 230.
[0055] According to another embodiment, the server or writing device 200 can exchange media files 225 with the client or reading device 250 via a storage means, e.g., the storage means referenced 240. Such a storage means can be, e.g., a memory module (e.g., a random access memory (RAM)), a hard disk, a solid state drive, or any removable digital medium, e.g., a disk or a memory card.
[0056] According to the illustrated embodiment, the server or writing device 200 is intended to process media data, such as media data referenced at 205, such as video data, audio data and / or descriptive metadata, for streaming or for storage purposes. To that end, the server or writing device 200 acquires or receives original media data or bitstream 205, such as media content comprising one or more timed sequences of images, timed sequences of audio samples or timed sequences of descriptive metadata, and encodes the acquired media data into encoded media data referenced at 215 using an encoding module referenced at 210 (e.g. video encoding or audio encoding). The server or writing device 200 then encapsulates the encoded media data into one or more media files referenced at 225, including the encapsulated media data, using an encapsulation module referenced at 220. According to the illustrated embodiment, the server or writing device 200 comprises at least one encapsulation module 220 for encapsulating the encoded media data. The encoding module 210 may be implemented within the server or writing device 200 to encode the received media data, or may be separate from the server or writing device 200. The encoding module 210 is optional, as the server or writing device 200 may encapsulate media data previously encoded by another device or may encapsulate raw media data.
[0057] The encapsulation module 220 can generate a media file or multiple media files that correspond to the encapsulated media data including encapsulations of alternative versions of the media data and / or contiguous fragments of the encapsulated media data.
[0058] It should be noted that, according to the illustrated embodiment, a client or reader device 250 is used to process the encapsulated media data in order to display or output the media data to a user.
[0059] As shown, a client or reading device 250 obtains or receives one or more media files, such as media file 225, via a communication network 230 or from a storage means 240. Upon obtaining or receiving the media file, the client or reading device 250 parses and decapsulates the media file using a decapsulation module referenced at 260 to retrieve encoded media data referenced at 265. The client or reading device 250 then decodes the encoded media data 265 using a decoding module referenced at 270 to obtain media data referenced at 275 representing audio and / or video content (signals) that may be processed by the client or reading device 250 (e.g., rendered or displayed to a user by a dedicated module, not shown). It is noted that the decoding module 270 may be implemented within the client or reading device 250 to decode the encoded media data, or may be separate from the client or reading device 250. The decryption module 270 is optional since the client or reading device 250 may receive a media file that corresponds to the encapsulated raw media data.
[0060] It should be noted here that the media file or files, e.g., media file 225, may be communicated to the decapsulation module 260 of the client or reading device 250 in many ways. For example, it may be pre-generated by the encapsulation module 220 of the server or writing device 200 and stored as data in a remote storage device (e.g., on a server or cloud storage) in the communication network 230 or in a local storage device such as the storage means 240 until a user requests the media file encoded therein from the remote or local storage device. Upon requesting the media file, the data is read, communicated, or streamed from the storage device to the decapsulation module 260.
[0061] The server or writing device 200 may also include a content providing device for providing or streaming content information directed to media files stored in the storage device to a user (e.g., the content information may be described via a manifest file (e.g., a Media Presentation Description (MPD) conforming to the ISO / IEC MPEG-DASH standard, or an HTTP Live Streaming (HLS) manifest) containing, for example, the title of the content and other descriptive metadata and storage location data for identifying, selecting, and requesting the media files). The content providing device may also be adapted to receive and process user requests for media files to be delivered or streamed from the storage device to the client or reading device 250. Alternatively, the server or writing device 200 may generate a media file or files using the encapsulation module 220 and communicate or stream them directly to the client or reading device 250 and / or the decapsulation module 260 when a user requests the content.
[0062] A user can access the audio / video media data (signals) through a user interface of a user terminal constituting a client or reading device 250 or having means for communicating with the client or reading device 250. Such a user terminal can be a computer, a mobile phone, a tablet, or any other type of device capable of providing / displaying media data to a user.
[0063] For purposes of illustration, a media file or media files such as media file 225 represent encapsulated encoded media data (e.g., one or more timed sequences of encoded audio or video data) in boxes according to the ISO Base Media File Format (ISOBMFF, ISO / IEC 14496-12 and ISO / IEC 14496-15 standards). A media file or media files may correspond to one single media file (preceded by a FileTypeBox'ftyp' box) or one initialization segment file (preceded by a FileTypeBox'ftyp' box) followed by one or more media segment files (possibly preceded by a SegmentTypeBox'styp' box). According to ISOBMFF, a media file (and segment files, if present) may contain two types of boxes: 'media data' boxes ('mdat' or 'imda') that contain the encoded media data, and 'metadata boxes' ('moov' or 'moof' or 'meta' box hierarchy) that contain metadata that define the placement and timing of the encoded media data.
[0064] The encoding or decoding modules (see 210 and 270, respectively, in FIG. 2) use an image or video standard to encode and decode image or video content. For example, image or video encoding / decoding (codec) standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 (ISO / IEC MPEG-2 Visual), ITU-T H.263 (ISO / IEC MPEG-4 Visual), ITU-T H.264 (ISO / IEC MPEG-4 AVC), ITU-T H.265 (HEVC) with Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions, and ITU-T H.VVC (ISO / IEC MPEG-I Versatile Video Coding (VVC)) with Scalable (SHVC) and Multiview (MV-HEVC) extensions. The techniques and systems described herein are also applicable to other encoding standards already available or yet to be available or developed.
[0065] FIG. 3 is a diagram illustrating an example of a fragmented presentation encapsulated within one or more media files, where new tracks are defined within a movie fragment in accordance with some embodiments of the present invention.
[0066] According to a particular embodiment, ISOBMFF is extended to allow defining new tracks in a movie fragment (also called media fragment), where the new tracks have not been previously defined in an initialization fragment, e.g. the MovieBox ('moov') box 300. Such new tracks are sometimes called "dynamic tracks", i.e. tracks that appear along the presentation and were not present at the start of the presentation (e.g. were not declared in a MovieBox box). Dynamic tracks may have the duration of one movie fragment or may span several movie fragments. In contrast to dynamic tracks, tracks declared in a MovieBox box are sometimes called "static tracks".
[0067] The MovieBox box 300 provides a description of a presentation which is initially composed of at least one track defined by one TrackBox ('trak') box 305 having a track identifier, e.g. track_ID equal to 1, which indicates a sequence of samples having an encoding format described by sample entries depicted in 4CC 'xxxx'.
[0068] According to the illustrated embodiment, the MovieBox box 300 is followed by two movie fragments (or media fragments). The first movie fragment is composed of a MovieFragmentBox ('moof') box referenced in 310 containing metadata and a MediaDataBox ('mdat') box 320 containing media data (or encoded media data). More precisely, as shown, this first movie fragment contains N samples (samples 1 to N) stored in the 'mdat' box 320 that belong to a track fragment described in a TrackFragmentBox ('traf') box with a track identifier track_ID equal to 1 in the MovieFragmentBox ('moof') box referenced in 310. Samples 1 to N are associated with a sample entry denoted by 4CC 'xxxx' via the parameter denoted sample_description_index in the TrackFragmentHeaderBox('tfhd') box, or by default via the parameter denoted default_sample_description_index in the TrackExtendsBox('trex') box.
[0069] As shown, a new or dynamic track (with track identifier track_ID equal to 2) is declared in the second movie fragment, in the MovieFragmentBox('moof') box referenced at 330. This new or dynamic track is not defined in the MovieBox box 300 (i.e., in the initialization fragment). According to an embodiment, the second movie fragment contains samples of two tracks in a MediaDataBox('mdat') 340: samples N+1 to N+M belonging to the track with track identifier track_ID equal to 1 are associated to the sample entry 'xxxx' via the parameter indicated by sample_description_index in the TrackFragmentHeaderBox('tfhd') box with track identifier equal to 1 or by default via the parameter indicated by default_sample_description_index in the TrackExtendsBox('trex') box with track identifier equal to 1, and samples N+M+1 to N+M+Z belonging to the new track with track identifier equal to 2 are associated to the sample entry 'yyyy' via the parameter indicated by sample_description_index in the TrackFragmentHeaderBox('tfhd') box with track identifier equal to 2.
[0070] A new or dynamic track in the second movie fragment is declared by defining a TrackBox('trak') box within the 'moof' box 330 , here 'trak' box 345 .
[0071] New or dynamic tracks can be defined in any movie fragment of the presentation using the same mechanism.
[0072] According to some embodiments, the definition of the 'trak' box is modified to allow for the definition in the MovieFragmentBox ('moof') box as follows: Box Type: 'trak' Container: MovieBox or MovieFragmentBox Mandatory: Yes Quantity: One or more
[0073] If a TrackBox('trak') box is present in a MovieFragmentBox('moof') box, it is preferably placed before a TrackFragmentBox('traf') box, e.g. before the 'traf' boxes 350 and 355. The tracks defined in this case are valid only for the life of the movie fragment. A track defined in such a way should use a track_ID different from that of the track defined in the MovieBox box (i.e. in the initialization fragment or segment), but may have the same track_ID as that of the track defined in the previous movie fragment, to identify the continuation of the same track. In such cases, the corresponding TrackBox boxes are (bit-for-bit) identical. To span multiple fragments, multiple declarations of new tracks, or dynamic tracks, are necessary from one fragment to another to ensure random access in each track fragment. The continuity of one or more new tracks or dynamic tracks may pertain to some movie fragments corresponding to a period, but not necessarily to all movie fragments covering this period. In a variant, to determine the continuity of a new or dynamic track on two movie fragments (not necessarily consecutive), the parser or reader can simply compare the TrackBox payloads: if they are equal (bit for bit), the parser or reader can safely consider it to be the same track and avoid some decoding reinitialization. If they are not equal (even if the same track identifier is used), the parser or reader should consider it to be a different track from the previous track. Preferably, it may be recommended to assign different track identifiers to new or dynamic tracks in case there is no continuity between these tracks across different movie fragments.More generally, track identifiers for new or dynamic tracks should not collide with other track identifiers in the MovieBox box (for that, writers may use the following track identifier values): When a file is encapsulated using the 'unif' brand, track identifiers used for new or dynamic tracks must not collide with identifiers of tracks, track groups, entity groups, or items. Preferably, a reserved range of 32-bit identifiers allocated for track identifiers may be used. For example, track identifiers above 0x10000000 can be considered reserved for new or dynamic tracks to avoid collisions with static tracks, i.e. tracks declared in the MovieBox ('moov') box.
[0074] The 'trak' box defined in the movie fragment should be empty (i.e. the 'trak' box does not define any samples), and any mandatory boxes defining sample timings or offsets such as the TimeToSampleBox ('stts') box, the SampleToChunkBox ('stsc') box and the ChunkOffsetBox ('stco') box have no entries (entry_count=0), and it must contain at least one sample entry describing the encoding format of the samples that belong to this track (indicated by a sample entry with 4CC 'yyyy') and are contained in the movie fragment, and it must have a SampleGroupDescriptionBox ('sgpd') box defined. Samples in the track are implicitly fragmented. The TrackHeaderBox ('tkhd') box of any TrackBox ('trak') box in the movie fragment must have a duration of 0.
[0075] According to some embodiments, the MovieFragmentBox('moof') box contains a TrackFragmentBox('traf') box for each track for which there are associated samples stored in an associated MediaDataBox('mdat') box, i.e. for the tracks defined in the initialization fragment and for the tracks defined in the considered movie fragment, e.g. for the new track defined in the 'trak' box 345. Thus, according to the illustrated example, the 'moof' box 330 contains two TrackFragmentBox('traf') boxes, one for the track with track identifier track_ID equal to 1 ('traf' box 350) and one for the new track with track identifier track_ID equal to 2 ('traf' box 355).
[0076] A TrackFragmentHeaderBox('tfhd') box of a 'traf' box referencing a new track not previously defined in a MovieBox('moov') box should have the following flags set in its tf_flags parameter indicating information that should be present in the new track's TrackFragmentHeaderBox('tfhd') box (because this new track's MovieBox('moov') box has not declared an associated TrackExtendsBox('trex') box providing the corresponding default values): sample-description-index-present, default-sample-duration-present, default-sample-size-present, default-sample-flags-present.
[0077] Note that any new track defined in a MovieFragment can benefit from all the usual track mechanisms, e.g. can reference another track or belong to a group of tracks, by defining a corresponding box in the 'trak' box, e.g. in the 'trak' box 345, e.g. by using a TrackReferenceBox ('tref') box or a TrackGroupBox ('trgr') box, respectively. According to some embodiments, a static track (i.e. a track defined initially in a MovieBox box) cannot reference a dynamic track (i.e. a track defined in a MovieFragmentBox box) through a track reference box, but a dynamic track can reference a static track. A dynamic track can reference another dynamic track through a track reference box, provided they are declared in the same MovieFragment (if the referenced track does not exist, the reference is ignored).
[0078] In a variant, new or dynamic tracks may be defined using dedicated boxes with new 4CCs (e.g. a TemporaryTrackBox box with 4CC 'ttrk' or a DynamicTrackBox box with 4CC ('dntk'). Such a dedicated box may be a light version of the TrackBox hierarchy of boxes. It may contain at least a track identifier, a media handler type specifying whether the track is a video, audio or metadata track, a description of the content of the samples that make up the track, and a data reference indicating the location of the samples (i.e. whether the samples are placed in a MediaDataBox('mdat') box or an IdentifiedMediaDataBox('imda') box). For example, the specialized boxes may include a TrackHeaderBox ('tkhd') box that provides the track identifier, a HandlerBox ('hdlr') box that provides the media handler type, a SampleDescriptionBox ('stsd') box that provides a description of the sample content, and a DataReferenceBox ('dref') box that indicates the location of the sample (e.g. in a MediaDataBox ('mdat') box or an IdentifiedMediaDataBox ('imda') box).
[0079] In another variant, a new or dynamic track can be defined by extending the TrackFragmentBox('traf') box. A new version (e.g. version=1) or a new flag value (e.g. a new flag value trak_in_moof in the tf_flags parameter) of that TrackFragmentHeaderBox('tfhd') box can be defined to inform the reading device that this TrackFragmentBox('traf') box corresponds to a new track not previously defined in the MovieBox box. When this new version of the 'tfhd' box is used or when this new flag value trak_in_moof is set in the 'tfhd' box, the 'traf' box contains a MediaInformationBox('mdia') box that provides at least the type of media handler, a description of the samples that make up the track. That is, it does not define the timing and offsets of the samples, but rather the samples are defined within a TrackRunBox ('trun') box), and it provides a data reference indicating the location of the samples (e.g. within a MediaDataBox ('mdat') box or an IdentifiedMediaDataBox ('imda') box).
[0080] In yet another variation, when this new version of the 'tfhd' box is used, or when this new tf_flags value trak_in_moof is set in the 'tfhd' box, the 'traf' box can contain some of the following boxes: MediaHeaderBox ('mdhd'), HandlerBox ('hdlr'), DataReferenceBox ('dref'), SampleDescriptionBox ('stsd').
[0081] Nevertheless, according to the above embodiment and its variants, the possibility of defining a new track or dynamic track in a movie fragment not previously defined in a MovieBox('moov') box can be indicated to a reading device by defining a new DynamicTracksConfigurationBox('dytk') box, e.g. the 'dytk' box 330, in a MovieExtendsBox('mvex') box, e.g. the 'mvex' box 360 of the 'movex' box 300.
[0082] A 'dytk' box is defined as follows: BoxType: 'dytk' Container: MovieExtendsBox Mandatory: Yes if TrackBox is present in movie fragments Quantity: Zero or one aligned(8) class DynamicTracksConfigurationBox extends FullBox('dytk', 0, 0){}
[0083] According to some embodiments, the presence of this box indicates that tracks not previously defined in a MovieBox box may be or are allowed to be declared inside a MovieFragment. Conversely, the absence of this box indicates that no tracks other than those previously defined in a MovieBox box can be present in any MovieFragmentBox('moof') box, i.e., it is not allowed to declare new tracks in any MovieFragmentBox('moof') box.
[0084] In a variant, the possibility to define a new track (e.g. a dynamic track) in a movie fragment not previously defined in a MovieBox box is signaled to the reading device by defining a new flag value, such as the track_in_moof flag, in the MovieExtendsHeaderBox('mehd') box of the MovieExtendsBox('mvex') box, e.g. in the 'mvex' box 360, e.g. as follows: track_in_moof: flag mask is 0x000001: If set, indicates that tracks can be defined within a movie fragment that have not been previously defined in a MovieBox box (i.e., a TrackBox box can be declared within a MovieFragmentBox('moof') box), If not set, no tracks other than those defined in the MovieBox box (static tracks) may be defined in any MovieFragmentBox('moof') box (e.g. TrackBox boxes must not be present in any MovieFragmentBox('moof') box). In another variant, the flag value track_in_moof is preferably defined in a MovieHeaderBox('mvhd') box within a MovieBox('moov') box (not shown in FIG. 3) instead of a MovieExtendsHeaderBox('mehd') box, to avoid indicating an optional MovieExtendsHeaderBox('mehd') box that may not be available in derived specifications (e.g. in Common Media Application Format (CMAF) when the duration of the movie fragment is unknown).
[0085] An advantage of the above described embodiment and its variants is that the writer does not need to know in advance, i.e. in the worst case, all possible tracks and their configurations in order to generate the initial MovieBox ('moov') box: it can introduce new or dynamic tracks as they become available during the life of a movie fragment (e.g. adding one or more media streams corresponding to additional languages or subtitles, or additional camera views).
[0086] FIG. 4 is a diagram illustrating an example of a fragmented presentation encapsulated within one or more media files, where sample entries and / or tracks are defined within a movie fragment in accordance with some embodiments of the present invention.
[0087] According to the embodiment described with reference to Figure 4, the ISOBMFF, in addition to allowing new or dynamic tracks to be defined within a movie fragment as described with reference to Figure 3, is extended to allow sample entries to be defined within a movie fragment of a track defined within a movie fragment ('moov') box, where the new sample entries have not been previously defined within a movie fragment ('moov') box. Such sample entries present in a track fragment may be referred to as dynamic sample descriptions or dynamic sample entries.
[0088] As shown, a MovieBox ('moov') box 400 provides a description of a presentation that is initially composed of at least one track defined by a TrackBox ('trak') box 405 which declares a sequence of samples with a track identifier track_ID equal to 1 and an encoding format described by sample entries illustrated with 4CC 'xxxx'.
[0089] According to the illustrated example, the 'moov' box 400 is followed by two movie fragments. The first movie fragment consists of a MovieFragmentBox('moof') box referenced at 410 and a MediaDataBox('mdat') box referenced at 420. This first movie fragment defines a new sample entry, denoted by 4CC 'zzz', for the track with track identifier track_ID equal to 1. This movie fragment stores in the 'mdat' box 420 within the 'moof' box 410 a sample of the track fragment described by the TrackFragmentBox('traf') with track identifier track_ID equal to 1.
[0090] The second movie fragment defines in its MovieFragmentBox('moof') box referenced at 430 a new or dynamic track (i.e. a track with a track identifier track_ID equal to 2) that is not defined in the 'moov' box 400. It stores samples of two tracks in its MediaDataBox('mdat') box referenced at 440: samples N+1 to N+M for the track with track identifier track_ID equal to 1 and samples N+M+1 to N+M+Z for the new track with track identifier track_ID equal to 2.
[0091] According to some embodiments of the invention, new or dynamic tracks or new (or dynamic) sample entries for tracks already defined in a MovieBox box can be defined in any movie fragment in the presentation. New tracks and new sample entries may be defined in the same movie fragment or in different movie fragments. Furthermore, the same new track and / or the same new sample entries may be defined in different movie fragments.
[0092] ISOBMFF is extended to allow the declaration of sample description boxes in the SampleTableBox('stbl') box of a TrackBox('trak') box of a MovieFragmentBox('moof') box, as well as in the TrackFragmentBox('traf') box of a MovieFragmentBox('moof') box, as follows: Box Types: 'stsd' Container: SampleTableBox or TrackFragmentBox Mandatory: Yes Quantity: Exactly one
[0093] It is therefore possible to declare new sample entries with different encoding format parameters at the movie fragment level. This is useful, for example, when there is a change in the codec configuration of the encoded media data or stream (e.g., an unexpected switch from an HD format in AVC to an UHD format in HEVC) or a change in the content protection information (e.g., an unexpected switch from a clear format to a protected format). Indeed, some application profiles require the support of multiple codecs, for example the DVB specification. This means that codec changes can occur during a multimedia presentation or program. If this happens (and is initially unknown) during a live program, such a codec change can be signaled by a dynamic sample entry. More generally, even if the codec does not change, a change in the encoding parameters can be signaled by a dynamic sample entry.
[0094] Samples (e.g. samples 1 to N belonging to a track with track identifier track_ID equal to 1) are associated to a sample entry via a sample description index value. The range of values of the sample description index can be split into several ranges to allow using one sample entry or another. For example, to use sample entries defined in the 'trak' box 405 and sample entries defined in the 'traf' box 415, the range of values of the sample description index can be split into two ranges: - Values from 0x0001 to 0x10000: These values indicate the index of the sample entry (or sample description) in the SampleDescriptionBox ('stsd') box in the TrackBox ('trak') box that corresponds to the indicated track_ID in the TrackFragmentBox ('traf') box. - Values from 0x10001 to 0xFFFFFFFF: These values indicate the index of the sample entry (or sample description) in the SampleDescriptionBox ('stsd') box contained in the TrackFragmentBox ('traf') box corresponding to the specified track_ID, incremented by 0x10000.
[0095] Thus, depending on the associated sample description index value, samples 1 to N stored in the 'mdat' box 420 are associated with either a sample entry having the 4CC 'xxxx' defined in the 'trak' box 405, or a sample entry having the 4CC 'zzz' defined in the 'traf' box 415.
[0096] If a SampleDescriptionBox('stsd') box is present within a TrackFragmentBox('traf') box, it is preferably placed immediately after the TrackFragmentHeaderBox('tfhd') box. According to some embodiments, the sample entries given in a SampleDescriptionBox('stsd') box defined in a movie fragment's TrackFragmentBox('traf') box are only valid for the lifetime of the movie fragment.
[0097] Further, according to some embodiments in which dynamic sample entries may be declared in a movie fragment, a new track or dynamic track in a second movie fragment (or any new track in any movie fragment) may be declared as described with reference to FIG. 3.
[0098] In the embodiment shown in FIG. 4, a new track is declared by defining a TrackBox('trak') box, referenced at 445, within the MovieFragmentBox('moof') box, referenced at 430.
[0099] Therefore, the TrackBox('trak') box definition can be changed to allow it to be defined in a MovieFragmentBox, as follows: Box Type: 'trak' Container: MovieBox or MovieFragmentBox Mandatory: Yes Quantity: One or more
[0100] If a TrackBox('trak') box (e.g., 'trak' box 445) is present within a MovieFragmentBox('moof') box (e.g., 'moof' box 430), it is preferably placed before any TrackFragmentBox('traf') boxes (e.g., 'trak' box 445 is placed before 'traf' boxes 450 and 455). Tracks defined in this case are only valid for the life of the corresponding movie fragment. Such new or dynamic tracks should not have the same track identifier (track_ID) as any track defined in the MovieBox box, but may have the same track identifier as tracks defined in one or more other movie fragments, in order to identify the continuation of the same track, in which case these TrackBox boxes are identical (bit-for-bit).
[0101] Again, spanning multiple fragments requires multiple declarations of new or dynamic tracks from one fragment to another to guarantee random access in each track fragment. The continuity of one or more new or dynamic tracks may concern several movie fragments corresponding to a period, but not necessarily all movie fragments covering this period. In a variant, to determine the continuity of a new or dynamic track on two movie fragments (not necessarily consecutive), the parser or reader can simply compare the TrackBox payloads. If they are equal (bit-for-bit), the parser or reader can safely consider it to be the same track and avoid some decoding reinitialization. If they are not equal (even if the same track identifier is used), the parser or reader should consider it to be a different track from the previous track. Preferably, it may be recommended to assign different track identifiers to the new or dynamic tracks in case there is no continuity between these tracks across different movie fragments. More generally, track identifiers for new or dynamic tracks should not collide with other track identifiers in the MovieBox box (for that, writers may use the following track identifier values): When a file is encapsulated using the 'unif' brand, track identifiers used for new or dynamic tracks must not collide with identifiers of tracks, track groups, entity groups, or items. Preferably, a reserved range of 32-bit identifiers allocated for track identifiers may be used. For example, track identifiers above 0x10000000 are considered to be reserved for new or dynamic tracks to avoid collisions with static tracks, i.e. tracks declared in the MovieBox ('moov') box.
[0102] Note again that TrackBox('trak') boxes, such as 'trak' box 445 defined in a movie fragment, should be empty (i.e., the 'trak' box does not define any samples). The 'trak' box does not define any samples, and required boxes, such as a TimeToSampleBox('stts') box, a SampleToChunkBox('stsc') box, or a ChunkOffsetBox('stco') box that defines sample timing or offsets, should have no entries (entry_count=0), should have an empty SampleDescriptionBox('stsd') box (e.g., 'stsd' box 460), and preferably have no SampleGroupDescriptionBox('sgpd') defined. Track samples are implicitly fragmented. The duration of the TrackHeaderBox('tkhd') box of a movie fragment's TrackBox('trak') box (e.g., 'trak' box 445) should be 0.
[0103] According to some embodiments, the MovieFragment('moof') box contains a TrackFragmentBox('traf') box for each track for which there are associated samples stored in an associated MediaDataBox('mdat') box, i.e. for the tracks defined in the initialization fragment and for the tracks defined in the considered movie fragment, e.g. for the new track defined in the 'trak' box 445. Thus, according to the illustrated example, the 'moof' box 430 contains two TrackFragmentBox('traf') boxes, one for the track with track identifier track_ID equal to 1 ('traf' box 450) and one for the new or dynamic track with track identifier track_ID equal to 2 ('traf' box 455).
[0104] Again, the TrackFragmentHeaderBox('tfhd') box of a 'traf' box referencing a new track not previously defined in a MovieBox box preferably has the following flags set in tf_flags (description of track fragment properties through a list of flags): sample-description-index-present, default-sample-duration-present, default-sample-size-present, default-sample-flags-present. These flags appear because there is no TrackExtendsBox('trex') box associated with this new or dynamic track, providing default values for samples (N+M+1 to N+M+Z) in the track fragment.
[0105] If a TrackBox('traf') box (e.g., the 'traf' box 445) is declared within a MovieFragmentBox('moof') box (e.g., the 'moof' box 430), then a SampleDescriptionBox('stsd') box (e.g., the 'stsd' box 470) should be declared within a TrackFragmentBox('traf') box (e.g., the 'traf' box 455) with the same track identifier as this TrackBox.
[0106] Note again that any new track defined in a MovieFragment can benefit from all the usual track mechanisms, e.g., can reference another track or belong to a group of tracks, by defining a corresponding box in the TrackBox('trak') box (e.g., in the 'trak' box 445), e.g., by using a TrackReferenceBox('tref') box or a TrackGroupBox('trgr') box, respectively. According to some embodiments, static tracks (i.e., tracks defined initially in a MovieBox box) cannot reference dynamic tracks (i.e., tracks defined in a MovieFragmentBox box) through a track reference box, but dynamic tracks can reference static tracks. Dynamic tracks can also reference another dynamic track through a track reference box, provided that they are declared in the same MovieFragment (the reference is ignored if the referenced track does not exist).
[0107] In a variant, a new or dynamic track may be defined using a dedicated box with a new 4CC (e.g. a TemporaryTrackBox box with 4CC 'ttrk', or a DynamicTrackBox box with 4CC ('dntk'). Such a dedicated box may be a light version of the TrackBox hierarchy of boxes. It may contain at least a track identifier, a media handler type specifying whether the track is a video, audio or metadata track, and a data reference indicating the location of the samples (i.e. whether the samples are placed in a MediaDataBox('mdat') box or an IdentifiedMediaDataBox('imda') box). In this variant, it is not necessary to contain a description of the samples that make up the track, since the description of the samples that make up the track is provided by a sample entry (e.g. sample entry 'yyyy') declared in a SampleDescriptionBox('stsd') box (e.g. 'stsd' box 470) of a TrackFragmentBox('traf') box (e.g. 'traf' box 455) with the same track identifier.
[0108] As will be apparent to those skilled in the art, other embodiments or variations of defining the new or dynamic tracks described with reference to FIG. 3 also apply here.
[0109] Nevertheless, according to the above embodiment and its variants, it may be signalled to the reading device at the beginning of the file to provide the possibility to define new tracks and / or new sample entries in the movie fragment that have not been previously defined in the MovieBox box. This indication may help the reading device to decide whether it can support the media file or only some of the tracks in the file. This can be done, for example, by defining a new DynamicTracksConfigurationBox('dytk') box (e.g. 'dytk') box 480) in a MovieExtendsBox('mvex') box (e.g. 'mvex' box 490) in a MovieBox('moov') box (e.g. 'moov' box 400). (Name and 4cc are an example)
[0110] A 'dytk' box is defined as follows: BoxType: 'dytk' Container: MovieExtendsBox Mandatory: Yes if TrackBox or SampleDescriptionBox are present in movie fragments Quantity: Zero or one aligned(8) class DynamicTracksConfigurationBox extends FullBox('dytk', 0, flags ){}
[0111] According to some embodiments, this box can be used to signal the presence of a TrackBox ('trak') box (or a similar box, depending on the variant, e.g. a TemporaryTrackBox ('ttrk') box) or a SampleDescriptionBox ('stsd') box in a movie fragment. For illustration purposes, the following flag values are available and are defined as follows: -track_in_moof: flag mask is 0x000001: If set, indicates that TrackBox ('trak') boxes can be declared in movie fragments. If not set, the MovieFragmentBox box should not contain a TrackBox ('trak') box; -stsd_in_traf: flag mask is 0x000002: If set, indicates that a SampleDescriptionBox ('stsd') box may be declared inside a TrackFragmentBox ('traf') box. If not set, the TrackFragmentBox('traf') box should not contain a SampleDescriptionBox('stsd') box.
[0112] In a variant, the possibility to define new or dynamic tracks (i.e. tracks not previously defined in a MovieBox box) in a MovieFragment is signaled to the reading device by defining a new flag value in a MovieExtendsHeaderBox('mehd') box of a MovieExtendsBox('mvex') box (e.g. 'mvex') box 490), e.g. a new flag value track_in_moof as follows: track_in_moof: flag mask is 0x000001: If set, indicates that tracks can be defined within a movie fragment that have not been previously defined in a MovieBox box (i.e., a TrackBox box can be declared within a MovieFragmentBox('moof') box). If not set, no tracks may be defined in any MovieFragmentBox('moof') box other than tracks previously defined in the MovieBox box (static tracks) (i.e., for example, a TrackBox box may not be present in any MovieFragmentBox('moof') box).
[0113] In another variant, the flag value track_in_moof is preferably defined in a MovieHeaderBox('mvhd') box within a MovieBox('moov') box (not shown in FIG. 3) instead of a MovieExtendsHeaderBox('mehd') box, to avoid indicating an optional MovieExtendsHeaderBox('mehd') box that may not be available in derived specifications (e.g. in Common Media Application Format (CMAF) when the duration of the movie fragment is unknown).
[0114] Similarly, a reading device can be informed of the possibility of defining a new sample entry in a movie fragment (i.e., a sample entry not previously defined in a MovieBox box) by defining a new flag in a TrackExtendsBox('trex') box of a MovieExtendsBox('mvex') box (e.g., 'mvex') box 490), e.g., a flag value stsd_in_traf, as follows: stsd_in_traf: flag mask is 0x000001: If set, signals that a SampleDescriptionBox ('stsd') box may be declared inside a TrackFragmentBox ('traf') box with the same track identifier as the track identifier set in the associated TrackExtendsBox ('trex') box. If not set, a TrackFragmentBox ('traf') box must not constitute a SampleDescriptionBox ('stsd') box.
[0115] In another variant, the flag value stsd_in_traf is preferably defined in the SampleDescriptionBox('stsd') box of the MovieBox('moov') box instead of the TrackExtendsBox('trex') box. With this variant, a reading device can directly know whether the sample description is likely to be updated during the presentation when parsing the SampleDescriptionBox('stsd') box, without having to parse any extra boxes in the MovieExtendBox('mvex') box. For example, if the flag value stsd_in_traf is not set in the SampleDescriptionBox('stsd') box, the parser is guaranteed that all possible sample entries for a given track are declared in the MovieBox('moov') box part of the file, i.e. in the initialization fragment. Conversely, if the flag value stsd_in_traf in a SampleDescriptionBox('stsd') box is set, this indicates to the parser that some additional, new or dynamic sample entries for the corresponding track may be defined later in subsequent movie fragments.
[0116] In another variant, the possibility to define either new tracks and / or new sample entries within a MovieFragment is signalled to the reading device by defining a new DynamicTracksConfigurationBox ('dytk') box within a MovieExtendsBox ('mvex') box within a MovieBox ('moov') box.
[0117] The DynamicTracksConfigurationBox('dytk') box is defined as follows: BoxType: 'dytk' Container: MovieExtendsBox Mandatory: Yes if TrackBox or SampleDescriptionBox are present in movie fragments Quantity: Zero or one aligned(8) class DynamicTracksConfigurationBox extends FullBox('dytk', 0, flags){ if ( ! (flags & all_tracks_dynamic_stsd)) { unsigned int(32) nb_tracks; unsigned int(32) trackIDs[nb_tracks]; } }
[0118] This box can be used to signal the presence of TrackBox ('trak') and SampleDescriptionBox ('stsd') boxes in a movie fragment. To do so, the following flags can be defined: track_in_moof: flag mask is 0x000001: If set, indicates that track boxes ('trak') boxes may be declared in movie fragments. If not set, the MovieFragmentBox box should not contain a TrackBox ('trak') box; all_tracks_dynamic_stsd: flag mask is 0x000002: If set, indicates that a SampleDescriptionBox ('stsd') box may be declared within a TrackFragmentBox ('traf') box for any track defined in the movie. If not set, no TrackFragmentBox ('traf') box that is not signaled in the considered 'dytk' box must constitute any SampleDescriptionBox ('stsd').
[0119] Additionally, the parameters of the DynamicTracksConfigurationBox('dytk') box can have the following semantics defined: nb_tracks indicates the number of trackIDs listed; trackIDs indicates the track identifiers of tracks for which a SampleDescriptionBox('stsd') box is declared in a TrackFragmentBox('traf') box. If the all_tracks_dynamic_stsd flag is not set and no tracks are listed in this box, then a SampleDescriptionBox('stsd') box must not be present in any TrackFragmentBox('traf') box for this track.
[0120] Further according to the above embodiment, the presence of either new or dynamic tracks, or new or dynamic sample entries in a movie fragment can be further signalled to a reading device by defining the following flag values in a MovieFragmentHeaderBox ('mfhd') box in a MovieFragmentBox ('moof') box: new-track-present (or dynamic-track-present): flag mask is 0x000001: If set, indicates that a new (or dynamic) track is being declared inside this movie fragment. If not set, it indicates that no new (or dynamic) tracks have been declared inside this MovieFragment. stsd-present: Flag mask is 0x000001: If set, indicates that a SampleDescriptionBox ('stsd') box has been declared in this movie fragment. If not set, it indicates that no SampleDescriptionBox ('stsd') box has been declared in this movie fragment.
[0121] While the signaling in the MovieBox('moov') box of a new (or dynamic) track or sample entry informs the reading device that it may have to process a new track or a new sample description in a subsequent media fragment during the presentation, this signaling in the movie fragment allows the reading device to know if the current movie fragment actually contains a new track or sample description definition. Furthermore, the MovieFragmentHeaderBox('mfhd') box in the MovieFragmentBox('moof') box can be extended by declaring a new version of the box (e.g. version=1) and adding the parameter next_track_ID. This parameter contains the value of the next track identifier that can be used to create a new or dynamic track in the next media fragment. It typically contains a value one higher than the highest identifier value used at the file level found in the presentation up to the current media fragment. This allows for easy generation of unique track identifiers without knowing all previous media fragments between the MovieBox('moov') box and the last generated media fragment.
[0122] According to some embodiments, the encapsulation module may use the brand (defined at the file level in a FileTypeBox('ftyp') box, at the segment level in a SegmentTypeBox('styp') box, or at the track level in a TrackTypeBox('ttyp') box), the existing or new box, or a field or flag value within these existing or new box(es) to indicate that new tracks or new sample entries may be declared in a media file or media segment. If such an indication is not used by the encapsulation module, the parser does not need to check for the presence of new tracks or sample entries. Conversely, if such an indication is present, the parser according to the invention should check for the presence of a new track or dynamic track or sample entry at the beginning of a movie fragment and finally check for the continuation of a new track or dynamic track.
[0123] FIG. 5 is a block diagram illustrating an example of steps performed by a server or writer to encapsulate encoded media data according to some embodiments of the present invention.
[0124] Such steps may be performed, for example, in encapsulation module 220 of FIG.
[0125] As shown, a first step (step 500) is directed to obtaining a first portion of encoded media data, which may consist of one or more bitstreams representing encoded timed sequences of video, audio, and / or metadata, including one or more bitstream features (e.g. scalability layers, temporal sub-layers, and / or spatial sub-parts such as HEVC tiles or VVC sub-pictures). Possibly, multiple options of the encoded media data are obtained, for example in terms of quality and resolution. The encoding is optional and the encoded media data may be raw media data.
[0126] From this first portion of the encoded media data, the encapsulation module determines a first set of tracks to be used to encapsulate the first portion of the encoded media data (step 505). It may decide to use one track per bitstream (e.g., one track for one video bitstream, one track for one audio bitstream, or two tracks for two video bitstreams, one track per video bitstream). It may also multiplex multiple bitstreams onto one track (e.g., multiplexed audio-video tracks). It may also split the bitstream into multiple tracks (e.g., one track per layer, or one track per spatial subpicture). It may also define additional tracks that provide instructions for combining other tracks (e.g., a VVC base track used to describe the configuration of multiple VVC subpicture tracks).
[0127] Then, in step 510, a description of the determined first set of tracks and associated sample entries is generated and encapsulated in an initialization fragment including a MovieBox box, such as 'moov' box 300 or 'moov' box 400. Such a description consists of defining one TrackBox box, such as 'trak' box 305 or 'trak' box 405, for each track of the first set of tracks, including one or more sample entries in the MovieBox box. Furthermore, the encapsulation module signals in a MovieExtendsBox box that some tracks of the first set of tracks are fragmented and that new tracks and / or new sample entries may be defined later in the movie fragment according to the aforementioned embodiment. Each time unit of the encoded media data corresponding to the first set of tracks is encapsulated into samples in each corresponding track in one or more movie fragments (also denoted as first media fragments), each media fragment consisting of one MovieFragmentBox ('moof') box and one MediaDataBox ('mdat') box.
[0128] Then, during step 515, the initialization fragment and the first media fragment are possibly output or stored as a set of media segment files following the initialization segment file, each media segment file containing one or more media fragments. It is noted that such a step (step 515) is optional, since for example the initialization fragment and the first media fragment may be output later together with the second media fragment.
[0129] Next, in step 520, the encapsulation module obtains a second portion of the encoded media data. The second portion of the encoded media data may consist of the same temporal continuation of bitstream as the first portion of the encoded media data. The temporal continuation of bitstream may be encoded with a different encoding format or parameters (e.g., changing from an HD encoding format of AVC to a 4K encoding format of HEVC), a different protection or encryption scheme, or a different packing organization (e.g., for stereoscopic or region-wise packing of omnidirectional media). The second encoded media data may also consist of additional bitstreams not present in the first encoded media data (e.g., new subtitle or audio language bitstreams, new video bitstreams corresponding to new cameras or viewpoints, detected objects or regions of interest).
[0130] Next, in step 525, the encapsulation module determines a second set of tracks and associated sample entries to be used to encapsulate the second portion of the encoded media data. If the second portion of the encoded media data is a simple temporal continuation of the first portion of the encoded media data, it may decide to maintain the same set of tracks for the second portion of the encoded media data. It may also decide to change the number of tracks, for example, to encapsulate the temporal continuation of the first portion of the encoded media data differently (e.g., a VVC bitstream including multiple sub-pictures is encapsulated in one single track for a first period, and then in multiple tracks, e.g., a VVC base track and multiple VVC sub-picture tracks, for a second period). It should also be noted that new tracks may be added for encapsulating additional bitstreams of the second portion of the encoded media data, and that some of the additional bitstreams may be multiplexed with some of the previous existing bitstreams. New sample entries may also be defined for the bitstream of the second part of the encoded media data if changes occur in this second part, for example if any of the information in the encoding format or codec profile and / or protection scheme or packing or sample description is changed.
[0131] It is then determined whether the description of the second set of tracks is the same as the description of the first set of tracks declared in the MovieBox box, step 530. If the description of the second set of tracks is the same as the description of the first set of tracks, then the second portion of the encoded media data is encapsulated in a movie fragment (i.e., a second media fragment) on a standard basis using the track and sample entries previously defined in the MovieBox box (step 535).
[0132] Otherwise, if there are new tracks in the second set of tracks compared to the first set of tracks, new tracks are defined in the movie fragment. Similarly, if there are new sample entries associated with the second set of tracks compared to the sample entries associated with the first set of tracks, new sample entries are defined in the movie fragment. The definition of new tracks and / or new sample entries may be performed as described with reference to Figures 3 or 4.
[0133] A description of the new track and / or the new sample entries is encapsulated in a second media fragment together with the second portion of the encoded media data (step 540).
[0134] The second media fragment is then output, e.g. transmitted to a client via a communications network or stored in a storage means, in step 545. The media fragment may be stored or transmitted as a segment file or appended together with the first media fragment in an ISO base media file.
[0135] As indicated by the dashed arrow, if there is more encoded media data to encapsulate, the process loops at step 520 until there is no more encoded media data to process.
[0136] FIG. 6 is a block diagram illustrating example steps performed by a client, parser or reader to process encapsulated media data according to some embodiments of the present invention.
[0137] Such steps may be performed, for example, by decapsulation module 260 of FIG.
[0138] As shown, the first step (step 600) is directed to obtaining an initialization fragment corresponding to the MovieBox box, which can be obtained by parsing an initialization segment file received from a communication network or by reading a file on a storage means.
[0139] The resulting MovieBox box is then parsed in step 605 to obtain a first set of tracks corresponding to a description of all tracks defined for the presentation, and to obtain associated sample entries describing the encoding format of the samples in each track. In this same step, the de-encapsulation module 260 can determine from branding information or specific boxes (e.g. 'dytk' 480) in the initialization fragment that the file or some segments may contain dynamic (or new) tracks or dynamic (or new) sample entries.
[0140] Next, in step 610, processing and decoding of the encoded media data encapsulated in the track is initialized using items of information from the first set of tracks and using the associated sample entries. Typically, the media decoder is initialized using the decoder configuration information present in the sample entries.
[0141] Next, in step 615, the decapsulation module obtains movie fragments (also written as media fragments) by parsing the media segment files received from the communication network or by reading the media segment files on the storage means.
[0142] Next, in step 620, the de-encapsulation module determines a second set of tracks from information obtained when parsing the MovieFragmentBox and TrackFragmentBox boxes present in the obtained movie fragment. In particular, it determines whether one or more new tracks and / or one or more new sample entries have been signaled and defined in the MovieFragmentBox and / or TrackFragmentBox boxes according to the above-described embodiment.
[0143] It is then determined in step 615 whether the second set of tracks and associated sample entries differ from the first set of tracks and associated sample entries. If the second set of tracks and associated sample entries differ from the first set of tracks and associated sample entries, the items of information retrieved from the MovieFragmentBox and / or TrackFragmentBox boxes of the retrieved media fragments are used to update the processing and decoding configuration of the encoded media data (step 630). For example, a new decoder may be instantiated to process a new bitstream, or a decoder may be reconfigured to process a bitstream with a changed encoding format, codec profile, encoding parameters, protection scheme, or packing organization.
[0144] After updating the configuration for processing and decoding the encoded media data, or if the second set of tracks and associated sample entries are the same as the first set of tracks and associated sample entries, the encoded media data is de-encapsulated from the movie fragment samples and processed (step 670), e.g., decoded and displayed or rendered to a user).
[0145] As indicated by the dotted arrow, if there are more media fragments to process, the process loops to step 615 until there are no more media fragments to process.
[0146] According to another aspect of the invention, in order to signal some features of the encapsulation mechanism that a reading device should understand in order to parse and decode the encapsulated media data, the data of a sample or of a NALU (Network Abstraction Layer (NAL) Unit) within a sample that is actually corrupted is signaled in order to help the reading device select the data to process. Data corruption can occur, for example, when data is received over error-prone communication means. To signal corrupted data in the encapsulated bitstream, a new sample group description with grouping_type 'corr' (or other predefined name) can be defined. This sample group 'corr' can be defined in any kind of track (e.g. video, audio or metadata) and signals the set of samples within the track that are corrupted or lost. For illustration purposes, an entry in this sample group description can be defined as follows: class CorruptedSampleInfoEntry() extends SampleGroupDescriptionEntry ('corr') { bit(2) corrupted; bit(6) reserved; } Here, corrupted is a parameter that indicates the corrupted state of the associated data.
[0147] According to some embodiments, a value of 1 means that the entire data set has been lost. In this case, the associated data size (sample size, or NAL size) should be set to 0. A value of 2 means that the data has been corrupted in such a way that it cannot be recovered by a resilient decoder (e.g. loss of a NAL slice header). A value of 3 means that the data has been corrupted but may be processed by an error-resilient decoder. A value of 0 is reserved.
[0148] According to some embodiments, the CorruptedSampleInfoEntry does not have an associated grouping_type_parameter defined. If some data is not associated with an entry in the CorruptedSampleInfoEntry, it means that these data are not corrupted.
[0149] The SampleToGroup('sbgp') box with grouping_type 'corr' allows you to associate a CorruptedSampleInfoEntry with each sample, indicating whether the sample contains corrupted data.
[0150] This sample group description with grouping_type 'corr' can also be advantageously combined within a NALU mapping mechanism consisting of a sampletogroup ('sbgp') box with grouping_type 'nalm', a sample group description ('sgpd') box, and a sample group description entry NALUMapEntry. A NALU mapping mechanism with grouping_type_parameter set to 'corr' can signal corrupted NALUs within a sample. The groupID in the NALUMapEntry map entry indicates the 1-based index of the sample group description in the CorruptedSampleInfoEntry. If groupID is set to 0, this indicates that no entry is associated here (the identified data is present and not corrupted).
[0151] The use of sample groups to indicate whether a sample is corrupted or not is important in order to be able to inform a reading device whether the sample group should be supported (analyzed and understood) in order to process the sample. Indeed, in such cases the sample entries alone may not be sufficient to determine whether a track is supported or not.
[0152] According to a particular embodiment, a new version of SampleGroupDescriptionBox is defined (e.g. version=3): If the version ('sgpd') of SampleGroupDescriptionBox is equal to 3, the sample group description describes mandatory information for the associated sample, and a parser, player or reader must not attempt to decode a track in which an unrecognized sample group description marked as mandatory is present.
[0153] In a simple version 3 variant of the SampleGroupDescriptionBox ('sgpd') box, the essentiality indication can be indicated by a new parameter called, for example, "essential" in the 'sgpd' box, like this: aligned(8) class SampleGroupDescriptionBox () extends FullBox('sgpd', version, flags){ unsigned int(32) grouping_type; if (version>=1) { unsigned int(32) default_length;} if (version>=2) {unsigned int(32) default_group_description_index;} if (version >= 3) { unsigned int (1) essential; unsigned int (7) reserved, / / =0 } unsigned int(32) entry_count; / / remaining parts are unchanged… …. }
[0154] If the essential parameter takes the value 0, the sample group entries declared in this SampleGroupDescriptionBox ('sgpd') box are descriptive and will not be exposed to parsers, player devices, or readers. If the essential parameter takes the value 1, the sample group entries declared in this SampleGroupDescriptionBox ('sgpd') box are mandatory to support in order to correctly process samples mapped to these sample group entries and should be exposed to parsers, player devices, or readers.
[0155] The essentiality of sample properties defined in a file by essential sample groups should be advertised through a MIME subparameter. This informs parsers and readers about the additional requirements expected to support the file. For example, if a corrupted sample group is declared as essential, a player with a basic decoder may not support a track with such an essential sample group. Conversely, a media player with a robust decoder (e.g. with concealment capabilities) may support a track with such an essential sample group. This new subparameter may be called, for example, "essential", and take as value a comma-separated list of four-letter codes corresponding to the grouping_types of sample groups declared as essential. For example: codecs= "avc1.420034", essential="4CC0" indicates an AVC track with one mandatory sample group of grouping type (e.g., 4CC0).
[0156] In another example: codecs="hvc1.1.6.L186.80", essential="4CC1,4CC2" indicates an HEVC track with two mandatory sample groups of grouping type, e.g., 4CC1 and 4CC2.
[0157] Alternatively, instead of defining a new version of the SampleGroupDescriptionBox('sgdb') box, you could define a new value for the flags parameter of the SampleGroupDescriptionBox('sgdb') box, as follows: essential: 4 (0x000004). When set to 1, this flag indicates that the Sample Group Description describes information that is essential to the associated samples, and file processors and file readers must not attempt to decode a track in which an unrecognized Sample Group Description marked as essential is present.
[0158] As another variant, a new box name SampleGroupEssentialPropertyBox with the four-letter code 'sgep' (or any other name or non-conflicting 4CC) can be defined with the same syntax and semantics as SampleGroupDescriptionBox ('sgpd'), except that it indicates the essential properties of the sample group that a reading device must support in order to process the track.
[0159] Defining a new brand in the FileTypeBox('ftyp'), SegmentTypeBox('styp'), or TrackTypeBox('ttyp') box forces the reader to support a new version of the SampleGroupDescriptionBox box (or its variants).
[0160] Thus, one possible mechanism for indicating the required properties of a sample could be to use two boxes: - A SampleToGroupBox('sbgp') box that describes the assignment of each sample to a sample group and a description of that sample group. - A SampleGroupDescriptionBox('sgpd') box (or SampleGroupEssentialPropertyBox('sgep') box) containing important flag values (or parameters) describing important properties of the samples in a particular sample group. The SampleGroupDescriptionBox('sgpd') box (or SampleGroupEssentialPropertyBox('sgep') box) contains a list of SampleGroupEntry (or VisualSampleGroupEntry in the case of video content), where each instance of SampleGroupEntry provides different values for the essential properties defined for a particular sample group (identified by its 'grouping_type').
[0161] A required sample grouping of a particular type is defined, via the type field ('grouping_type'), by the combination of one SampleToGroupBox ('sbgp') box and one SampleGroupDescriptionBox ('sgpd') box with required flag values or parameters (or one SampleGroupEssentialPropertyBox 'sgep') box.)
[0162] Similarly, mandatory properties can be associated with one or more Network Abstraction Layer (NAL) units in a sample by using a sample group of type 'nalm' as described in ISO / IEC 14496-15. Mandatory properties are associated with one or more NAL units in one or more samples by associating a mandatory sample group description box with a sample group of type 'nalm'. A SampletoGroupBox('sbgp') box of grouping type 'nalm' provides the index of the NALUMapEntry assigned to each group of samples and a grouping_type_parameter that identifies the grouping_type of the associated mandatory sample group description box. The NALUMapEntry associates a groupID with each NAL unit of the sample, and the associated groupID provides the index of the entry of the mandatory sample group description that provides the mandatory characteristics of the associated NAL unit.
[0163] For example, the grouping_type_parameter of a SampletoGroupBox('sbgp') box with grouping type 'nalm' may be equal to 'corr' to associate a sample group of type 'nalm' with a sample group description of type 'corr'.
[0164] Preferably, when a presentation is fragmented, a sample group with a given grouping_type is the first to be marked as required (e.g. in a MovieBox 'moov' box), and if a sample group with the same grouping_type is defined in a subsequent media fragment, it should also be marked as required.
[0165] According to some embodiments, the principle of mandatory sample groups described above can be used as an extensible way to mandate support for sample group descriptions in tracks. In particular, mandatory sample groups can be used to describe transformations applied to samples: before decryption, i.e. if the content of the sample has been transformed (for example by encryption) in such a way that it can no longer be decrypted by a normal decoder, or if the content should only be decrypted if the protection system or scrambling operation applied to the sample is understood and implemented by the playback device, - Post-decoding, i.e. if the file creator requires that certain actions be performed on the decoded samples before playback or rendering (e.g. if the content of the decoded samples needs to be unpacked before rendering, e.g. in the case of stereoscopic images where left and right views are packed into the same picture).
[0166] Compared to transformations that are usually described using restricted or protected sample entries, mandatory sample groups offer several advantages: - The 4-character code (4CC) of the sample entries, which indicates the coding format of the samples, does not need to be modified, thus not hiding the original nature of the data in the track, and a playback device does not need to parse a whole hierarchy of restricted or protected sample entries to find out the original format of the track. - Mandatory Sample Groups allow defining sample granularity transformations more efficiently (less impact on file size and fragmentation). -Required sample groups can easily support many, potentially nested transformations. -Required sample groups can easily accommodate new transformations whenever new properties are defined.
[0167] Another benefit of using a generic mechanism like mandatory sample groups to signal the transformations applied to samples is that it allows several classes of file processors to operate on files with transformations unknown to the file processor. With approaches that rely on restricted or protected sample entries, introducing a new transformation through a sample entry would require updating the code of dashers (e.g., devices preparing content to be streamed according to the Dynamic HTTP Adaptive Streaming (DASH) standard) or transcoders, which would always be error-prone. With a generic mechanism like mandatory sample groups, dashers and transcoders can process files without understanding the nature of the transformations.
[0168] According to some embodiments, the transform is a grouping_type value that identifies the type of transform (which may correspond, for example, to some known transformation scheme type, e.g., 'stvi' for stereoscopic packing, 'cenc' or 'cbc1' for some common encryption 4CCs, or some known transformation property 4CCs such as 'clap' for cropping / clean aperture or 'irot' for rotation, or any 4CC corresponding to a new defined transform). The properties of the transform are declared in one or more SampleGroupDescriptionEntry() of the mandatory SampleGroupDescriptionBox, each entry corresponding to an alternative value of the property. One of the one or more SampleGroupDescriptionEntry() is associated with each sample of the track using the default_group_description_index parameter of the SampleToGroupBox or SampleGroupDescriptionBox.
[0169] Alternatively, a transformation can be defined as a mandatory sample group if the transformation is mandatory (e.g., using a SampleGroupDescriptionBox with version=3 or using the alternatives mentioned above), or as a regular sample group (i.e., non-mandatory) if the transformation is optional (e.g., using a SampleGroupDescriptionBox with a version less than 3 without mandatory signaling).
[0170] When multiple sample groups defining transforms and / or description properties are declared and associated with samples of a track, it is necessary to define the order in which these sample groups are processed by the parser or reader in order to correctly decode or render each sample.
[0171] According to some embodiments, a new essential sample group grouping_type, e.g., 'esgh', is defined in the SampleGroupDescriptionEntry EssentialDescriptionsHierarchyEntry as follows: -EssentialDescriptionsHierarchyEntry A sample group description indicates the processing order of essential sample group descriptions that apply to a given sample. - The sample group description in the EstessentialDescriptionsHierarchyEntry is a mandatory sample group description, using version 3 (or other alternatives as mentioned above) of the SampleGroupDescriptionBox. This sample group exists if at least one mandatory sample group description is present.
[0172] Each essential sample group description other than EssentialDescriptionsHierarchyEntry is described in an EssentialDescriptionsHierarchyEntry sample group description.
[0173] The grouping_type_paramater in the EssentialDescriptionsHierarchyEntry sample group description is not defined and its value is set to 0.
[0174] The syntax and semantics of EssentialDescriptionsHierarchyEntry are defined as follows: class EssentialDescriptionsHierarchyEntry () extends SampleGroupDescriptionEntry ('esgh') { bit(8) num_groupings; unsigned int(32) sample_group_description_type[num_groupings]; } Where: num_groupings indicates the number of required sample group description types listed in the entry. sample_group_description_type indicates the four-letter code of the mandatory sample group description that is applied to the associated sample. That is, sample processing that may be described by a sample group of type sample_group_description_type[i] is applied before sample processing that may be described by a sample group of type sample_group_description_type[i+1]. The reserved value 'stsd' indicates the position of the decoding processing in the transform chain.
[0175] If 'stsd' is not present in the list of sample_group_description_types, then all listed mandatory sample groups apply to the decoded samples.
[0176] As an example, a sample that is encrypted before being encoded is signaled through a mandatory sample group of grouping_type 'vcne'. If the same sample also needs to have a post-processing filter applied after it is decoded, this post-processing may be signaled through a mandatory sample group of grouping_type 'ppfi'. In this example, an EssentialDescriptionsHierarchyEntry is defined. It lists the transforms in the following order ['vcne','stsd','ppfi'] to signal the order of nested transforms. From this signaling and the order of the transforms, a parser or reader can determine that a sample must be decoded before being decoded by the decoder identified in the sample entry, and that it must be processed by a post-processing filter before being rendered or displayed.
[0177] In a variant, instead of using the reserved value 'stsd' in the sample_group_description_type[] array of the EssentialDescriptionsHierarchyEntry to distinguish between essential sample groups before and after decoding, the nature of the essential sample group (pre-decoding or post-decoding) is determined from the definition associated with the four-character code (4CC) stored in the grouping_type parameter of the essential SampleGroupDescriptionBox, which identifies the transformation or description property. For example, 'clap' may actually identify a crop / clean aperture transformation, and this transformation may be defined to be a post-decoding transformation. As another example, 'cenc' may actually identify a generic encryption method (e.g., encryption of full samples and video NAL sub-samples in AES-CTR mode), and this transformation may be defined to be a pre-decoding transformation.
[0178] Thus, sample_group_description_type is defined as follows: It indicates the four-letter code of the mandatory sample group description that applies to the associated sample. That is, sample processing that may be described by a sample group of type sample_group_description_type[i] is applied before sample processing that may be described by a sample group of type sample_group_description_type[i+1]. All pre-decoding mandatory sample group descriptions are listed first, followed in order by all post-decoding mandatory sample group descriptions.
[0179] In another variant, the nature of the essential sample group (pre-decoding or post-decoding) is explicitly indicated using a parameter of the essential SampleGroupDescriptionBox. For example, this parameter corresponds to the 'flags' parameter of the essential SampleGroupDescriptionBox box header. A new flag value pre-decoding_group_description can be defined. Setting this flag to 1 indicates that the essential sample group is a pre-decoding essential sample group. Otherwise, setting this flag to 0 indicates that the essential sample group is a post-decoding essential sample group.
[0180] Therefore, in the sample_group_description_type[] array of EssentialDescriptionsHierarchyEntry, all descriptions of the essential sample groups before decoding are listed first, followed by descriptions of the essential sample groups after decoding, in order.
[0181] In another variant, it is possible to distinguish three different natures of mandatory sample groups: pre-decoding transformation, post-decoding transformation, and descriptive nature. The nature of the mandatory sample group can be indicated using a similar method as in either variant above (i.e. as part of the semantics of the four-letter code identifying the type of sample group, or using a parameter of the mandatory SampleGroupDescriptionBox). When indicating the nature of the mandatory sample group using the 'flags' parameter in the box header of the mandatory SampleGroupDescriptionBox, the following two-bit flag value can be defined (corresponding to bits 3 and 4 of the 'flags' parameter): - descriptive_group_description, which has the value 0 (0x00), if set, this 2-bit flag value indicates that the mandatory sample group is a descriptive mandatory sample group. - pre-decoding_group_description, with value 1 (0x01), if set, this 2-bit flag value indicates that the mandatory sample group is a pre-decoding mandatory sample group. - post-decoding_group_description, with value 2 (0x10), if set, this 2-bit flag value indicates that the mandatory sample group is a post-decoding mandatory sample group. - Value 3 (0x11) is reserved.
[0182] Thus, according to this variation, the sample_group_description_type[ ] array of EssentialDescriptionsHierarchyEntry lists all pre-decoding essential sample group descriptions first, followed by all description essential sample groups, followed by all post-decoding essential sample group descriptions, in that order.
[0183] Alternatively, the sample_group_description_type[] array of EssentialDescriptionsHierarchyEntry may list all essential sample group descriptions before decoding first, followed by all essential sample group descriptions after decoding, followed by all essential sample group descriptions of descriptive type.
[0184] The essentiality of sample properties defined in a file by mandatory sample groups, and the order of these mandatory sample groups, should be advertised through MIME subparameters. This informs parsers and readers of the expected additional requirements to support the file. Mandatory sample groups to be applied before decoding are listed in the order in which they will be applied (as indicated by the 'esgh' sample group description) in the 'codec' subparameter, ensuring maximum compatibility with existing practices. Other mandatory sample groups may be listed in a new subparameter called, for example, "essential", whose value is a comma-separated list of four-letter codes corresponding to the grouping_types of mandatory sample groups to be applied after decoding.
[0185] In other words, if a track has an essential sample group description: -Before any codec configuration, the 'codecs' subparameter lists the four-letter codes of the mandatory sample group descriptions to be applied before the decoding process, in the order in which they will be applied. Each mandatory sample group description is separated by a 'dot'. - The 'essential' subparameter can list four-letter codes of essential sample group descriptions to be applied after the decoding process in the order they are to be applied. Multiple values of essential sample group descriptions are separated by dots.
[0186] As an example, a sample that is encrypted before encoding is signaled through a mandatory sample group of grouping_type 'vcne'. If the same sample also needs to have a post-processing filter applied after it is decoded, this post-processing may be signaled through a mandatory sample group of grouping_type 'ppfi'. According to this example, as mentioned above, an EssentialDescriptionsHierarchyEntry is defined, which lists the transforms in the following order ['vcne', 'stsd', 'ppfi'] to signal the order of nested transforms.
[0187] So the 'codecs' mime type subparameter would be: codecs=vcne.hvc1.1.6.L186.80
[0188] This informs a reading device that the sample has been encrypted and must be decoded according to the transform identified in 4CC 'vcne' before being decoded by a decoder conforming to the codec and profile hierarchical level identified in 'hvc1.1.6.L186.80'.
[0189] And the "required" mime type subparameters are: essential=ppfi
[0190] This is signaling to the reading device that after decoding and before rendering, it should apply the post-processing filter identified by the mandatory sample group 'ppfi' to the samples.
[0191] Alternatively, an EstecalLDescriptionsHierarchyEntry sample group description indicates the processing order of mandatory and non-mandatory sample group descriptions that apply to a given sample. Non-mandatory sample group descriptions are either post-decoding transformations or description information.
[0192] In a variant, rather than describing the processing order of sample groups in the sample group descriptions of the EstenalLDescriptionsHierarchyEntry, the EstenalLDescriptionsHierarchyEntry can be declared as a SampleGroupDescriptionsHierarchyBox defined in a SampleTableBox or TrackFragmentBox, as follows: class SampleGroupDescriptionsHierarchyBox () extends FullBox ('esgh', version, flags = 0) { bit(8) num_groupings; unsigned int(32) sample_group_description_type[num_groupings]; } Has the same syntax and semantics as EssentialDescriptionsHierarchyEntry.
[0193] In such cases, the declaration of the processing order of sample groups in a SampleGroupDescriptionsHierarchyBox of a TrackFragmentBox replaces the processing order declared in the SampleGroupDescriptionsHierarchyBox of any preceding TrackFragmentBox with the same trackID and in any preceding SampleTableBox of any TrackBox with the same trackID.
[0194] According to some embodiments of the invention, as an alternative to the embodiment shown in Figures 3 and 4, codec switching within a track and dynamic sample descriptions or sample entries within a track are declared using intrinsic sample groups rather than defining a sample description box ('stsd') within a track fragment box ('traf').
[0195] A new mandatory sample group description with grouping_type 'stsd' is defined, whose entries consist of sample description child boxes (i.e. SampleEntry boxes providing sample descriptions for the specified encoding format), and assign samples to the appropriate entries using either a SampleToGroupBox with the same grouping_type or a default entry using the default_group_description_index parameter of the mandatory SampleGroupDescriptionBox.
[0196] According to a first variant, the sample description box 'stsd' of the track's SampleTableBox is either an empty box (no sample entry boxes defined inside it) or no sample description box 'stsd' is defined in the track's SampleTableBox. All sample entries are declared with a mandatory sample group description box with grouping_type 'stsd'. This sample group description box may be defined in the track's SampleTableBox (i.e. in the trackBox of a MovieBox) or in any movie fragment, or both.
[0197] To allow the parser to retrieve the sample entry definition associated with a sample, the semantics of the sample_description_index in the track fragment header 'tfhd' and the default_sample_description_index in the TrackExtendsBox 'trex' can be redefined as follows: A value of 0 means that the mapping from sample to sample description index is done via a SampleToGroupBox of grouping type 'stsd'. A value greater than -0 and less than or equal to 0x10000 indicates the index of the 'stsd' entry in the sample group description of grouping type 'stsd' defined in the track box, where 1 is the first entry. A value strictly greater than -0x10000 gives the index (value -0x10000) of the sample group description 'sgpd' of grouping type 'stsd' defined in the track fragment 'traf', where 1 is the first entry.
[0198] According to a second variant, a sample entry can be declared with a sample description box 'stsd' in the track's SampleTableBox and / or a mandatory sample group description box with grouping_type 'stsd' defined in any movie fragment. A mandatory sample group description box with grouping_type 'stsd' cannot be defined in the track's SampleTableBox.
[0199] To allow the parser to retrieve the sample entry definition associated with a sample, the semantics of the sample_description_index in the track fragment header 'tfhd' and the default_sample_description_index in the TrackExtendsBox 'trex' can be redefined as follows: A value of 0 means that the mapping from sample to sample description index is done via a SampleToGroupBox of grouping type 'stsd'. A value between -0 and 0x10000, inclusive, indicates the index of the 'stsd' entry within the track defined in the MovieBox, where 1 is the first entry. A value strictly greater than -0x10000 gives the index (value -0x10000) of the sample group description 'sgpd' of grouping type 'stsd' defined in the track fragment 'traf', where 1 is the first entry.
[0200] According to a third variant, a sample entry may be declared with a sample description box 'stsd' in the track's SampleTableBox, or with a mandatory sample group description box with grouping_type 'stsd', or both. Also, a mandatory sample group description box with grouping_type 'stsd' may be defined in the track's SampleTableBox, or in any movie fragment, or both.
[0201] To allow the parser to retrieve the sample entry definition associated with a sample, the semantics of the sample_description_index in the track fragment header 'tfhd' and the default_sample_description_index in the TrackExtendsBox 'trex' can be redefined as follows: A value of 0 means that the mapping from sample to sample description index is done via a SampleToGroupBox of grouping type 'stsd'. A value between -0 and 0x10000, inclusive, indicates the index of the stsd entry within the track defined in the MovieBox, where 1 is the first entry. -0x10000 strictly greater than or equal to 0x20000 indicates the index (value -0x10000) of the 'stsd' entry in the sample group description of grouping type 'stsd' defined in the track box, where 1 is the first entry. A value strictly greater than -0x20000 gives the index (value -0x20000) of the sample group description 'sgpd' of grouping type 'stsd' defined in the track fragment 'traf', where 1 is the first entry.
[0202] Thus, according to these embodiments, the invention provides a method for encapsulating a media data unit, the method being performed by a server, comprising: identifying a set of media data units according to an encapsulation independent parameter; obtaining an indication of a transformation to be applied to each media data unit of the set of media data units before or after parsing or decoding the encapsulated media data unit; encapsulating the media data units into one or more media files, where the media data units of the set of media data units are grouped into groups associated with an indication of transformation; Includes.
[0203] The media data units may be samples.
[0204] According to some embodiments, the one or more media files further include additional instructions signaling to the client that the media data unit group will be parsed, decoded or rendered only if the client has knowledge of the conversion instructions.
[0205] According to further embodiments, identifying a set of media data units, obtaining an indication of the transformation and encapsulating the media data units is repeated, the method further comprising obtaining an order for applying the transformation and encapsulating an item of information representative of the order.
[0206] Additionally, according to some embodiments, one or more of the transformations and orders may be provided as properties of one or more media files.
[0207] According to a further embodiment, the present invention provides a method for parsing an encapsulated media data unit, the method being performed by a client, the method comprising: obtaining an indication of a transformation to be applied to each media data unit of the media data unit group before or after parsing or decoding the encapsulated media data unit; obtaining an encapsulated media data unit of a media data unit group; parsing the encoded media data units obtained from the one or more media files while applying a transformation to each media data unit of the obtained encoded media data units before parsing the obtained encoded media data units or after parsing or decoding the obtained encoded media data units; Includes.
[0208] Again, the media data units may be samples.
[0209] According to some embodiments, the one or more media files further include additional instructions signaling to the client that encapsulated media data units of the media data unit group are parsed, decoded or rendered only if the client has knowledge of the conversion instructions.
[0210] According to further embodiments, obtaining an indication of the transformation, obtaining an encapsulated media data unit, parsing the obtained encapsulated media data unit, and applying the transformation are repeated, the method further comprising obtaining an order for applying the transformation from the one or more media files, the transformations being applied according to the order obtained.
[0211] Moreover, according to some embodiments, the one or more transformations and orders are derived from properties of one or more media files.
[0212] The advantage of using mandatory sample group descriptions in sample entries is that it is possible to decorate track fragments from the index of the sample description (given in the track fragment header) and to avoid creating new track fragments when the sample description is changed.
[0213] More generally, signaling required features may be useful to signal to a reading device other characteristics of a presentation applied to various sets of samples (tracks, groups of tracks, groups of entities) that are essential to be supported (analyzed and understood) in order to render or process the set of samples or the entire presentation.
[0214] According to a particular embodiment, tracks may be grouped together to form one or more groups of tracks, where each group shares certain characteristics or the tracks within a group have a particular relationship indicated by a particular 4CC value of the track_group_type parameter. A group of tracks is declared by defining a TrackGroupTypeBox box with the same track_group_type parameter and the same group identifier, track_group_id, in each TrackBox('trak') box of the tracks that belong to the group of tracks.
[0215] A required flag value (e.g., value=0x2) may be defined in the flags of the TrackGroupTypeBox box to inform a reading device that the semantics of a particular track group with particular values of the track_group_type parameter and track_group_id should be supported (parsed and understood) in order to render or process the set of samples formed by this track group. If the required flag is set for a track group and the semantics of this track group are not understood by the parser, the parser should not render or process any of the tracks belonging to this track group.
[0216] If this required flag is set in a TrackGroupTypeBox box with a particular track_group_type parameter value for a track, then it SHOULD also be set in all TrackGroupTypeBox boxes with the same particular track_group_type parameter value for all other tracks belonging to the same track group.
[0217] For ease of explanation, track groups of 'ster' type signaling tracks that form a stereo pair suitable for playback on a stereoscopic display, or track groups of '2dsr' type signaling tracks that have a two-dimensional spatial relationship (e.g. corresponding to a spatial portion of a video source), are examples of track groups that can benefit from this mandatory signaling.
[0218] In a variant, if the required flag is set for a group of tracks in a presentation and the semantics of the group of tracks indicated by the track_group_type parameter are not understood by the parser, the parser should not render or process the entire presentation.
[0219] In another variant, a mandatory group of tracks can be declared by defining a specific track group with a specific track_group_type, e.g. equal to 'etial', to inform a reading device that all tracks belonging to this track group are mandatory and should be supported (parsed and understood) to be rendered or processed together, i.e. all sample entries defined in the SampleDescriptionBox('stsd') boxes of each TrackBox('trak') of each track belonging to the track group should be supported (parsed and understood) by the reading device.
[0220] According to another embodiment, entities (i.e. tracks and / or items) may be grouped together to form one or more groups of entities where each group shares certain characteristics or the entities within a group have a certain relationship as indicated by a particular 4CC value of the grouping_type parameter.As opposed to tracks, which represent a timed sequence of samples primarily described by a set of MovieBox ('moov') and TrackBox ('trak') boxes, items represent non-timed media data described by a MetaBox ('meta') box and its hierarchy of boxes (e.g. including ItemInfoBox ('iinf') boxes, ItemLocationBox ('iloc') boxes, etc.).
[0221] A group of entities is declared by defining an EntityToGroupBox box with a particular grouping_type parameter, a group identifier group_id, and a list of entity identifiers (track_ID or item_ID) within a MetaBox('meta') box. This MetaBox('meta') box can be located at different levels, for example at the file level (i.e. at the same level as the MovieBox('moov') box), at the track level (i.e. in a TrackBox('trak') box), at the movie fragment level (i.e. in a MovieFragmentBox('moof') box), or at the track fragment level (i.e. in a TrackFragmentBox('traf') box).
[0222] A required flag value (e.g. value=0x1) can be defined in the flags of an EntityToGroupBox box to inform a reading device that the semantics of a group of entities declared by an EntityToGroupBox box with a particular value of the grouping_type parameter and group_id should be supported (parsed and understood) in order to render or process the set of samples and items formed by this group of entities. If the required flag is set for an entity group and the semantics of this entity group are not understood by the parser, the parser should not render or process any of the tracks or items belonging to this entity group.
[0223] In a variant, if the required flag is set for a group of entities in a presentation and the semantics of the group of entities indicated by the grouping_type parameter are not understood by the parser, then the parser should not render or process the entire presentation.
[0224] In another variant, a mandatory group of entities can be declared by defining a specific group of entities with a specific grouping_type, e.g. equal to 'etial', to inform a reading device that all tracks and / or items belonging to this group of entities are mandatory and should be supported (parsed and understood) to be rendered or processed together, i.e. all sample entries defined in the SampleDescriptionBox ('stsd') box of the TrackBox ('trak') box of each track belonging to the entity group, and all items defined in their ItemInfoBox ('iinf') boxes belonging to the entity group, should be supported (parsed and understood) by a reading device.
[0225] More generally, signaling of required capabilities may be useful to signal to a reading device which data structures or boxes are mandatory to be supported (parsed and understood) in order to render or process a presentation or part of a presentation.
[0226] According to another embodiment, a track may be associated with an EditList that provides an explicit timeline map for the track. Each EditList entry can define a part of the track timeline: by mapping a part of the Composition Timeline, or by indicating an 'empty' time (a part of the Presentation Timeline that is mapped to no media, an 'empty' edit), or by defining a 'dwell' where a single timepoint in the media is held for a period of time.
[0227] In the flags of an EditListBox('elst') box, a required flag value (e.g. value=0x2) can be defined to inform the reading device that the explicit timeline map defined in this edit list should be supported (parsed and applied) for rendering or processing the associated track and should not be ignored by the parser.
[0228] According to yet another embodiment, a mandatory flag value (e.g., value=0x800000) can be defined in the flags of the TrackHeaderBox('tkhd') box of the TrackBox('trak') box to indicate that the corresponding track should be supported (parsed and understood) for rendering or processing the presentation, i.e., all sample entries defined in the SampleDescriptionBox('stsd') box of the TrackBox('trak') of this track should be supported (parsed and understood) by a reading device for processing the presentation. In other words, all sample entries defined in the SampleDescriptionBox('stsd') box of the TrackBox('trak') of this track should be supported (parsed and understood) by a reading device for processing the presentation.
[0229] According to another embodiment, a required flag value (e.g., value = 0x1) can be defined in the flags of an ItemInfoEntry('infe') box within an ItemInfoBox('inf') box to inform a reading device that the corresponding item and all its required item properties should be supported (parsed and understood) in order to render or process the presentation.
[0230] More generally, full boxes (i.e., boxes with a flags parameter) in ISOBMFF and derived specifications can define a required flag value to signal to a reading device that this box should be supported (understood by a parser) and should not be ignored if not supported.
[0231] FIG. 7 is a block diagram illustrating example steps performed by a client or reader to obtain data according to some embodiments of the present invention.
[0232] In a variation to the above embodiment, instead of prohibiting a reading device from decoding tracks in which descriptions of unrecognized sample groups marked as mandatory are present, the following steps describe an alternative where only samples related to the mandatory sample group properties may be ignored by the reading device.
[0233] In step 700, the reader parses the ISOBMFF file or the segment file to obtain samples.
[0234] Next, in step 705, the reader checks whether the sample belongs to a sample group that has been signaled as a mandatory sample group according to some of the embodiments described above. If the sample group is mandatory, the reader checks whether the grouping type of the mandatory sample group is known (step 710).
[0235] If the grouping type of a required sample group is known, then that sample (and other samples in the group) can be processed (step 715), otherwise it is ignored (step 720).
[0236] For example, in certain use cases, supplemental enhancement information (SEI) that is typically transmitted in coded video bitstreams (e.g., for transmitting high dynamic range (HDR), virtual reality (VR), or film grain type media data) may be conveyed in the description of a sample group marked as mandatory.
[0237] It is recalled that media bitstreams may be accompanied by additional information that can be used to assist a playback device in processing the media file, for example, for decoding, display, or other purposes. For example, a video bitstream may be accompanied by SEI messages defined as a standard specification (e.g., ISO / IEC 23002-7). This additional information may be used by applications and may be specified by other standards, guidelines, or interoperability points in some consortia (e.g., ATSC, DVB, ARIB, DVB) that use MPEG specifications such as compression, encapsulation, and description specifications. For example, some DVB specifications mandate an alternative transfer characteristic SEI for HDR applications. Some other SCTE specifications mandate additional information that allows closed caption data in the video stream. With the prevalence of such additional information and the increasing complexity of media streams, some SEI messages and additional information may become mandatory or necessary for the rendering of the media presentation.
[0238] According to a particular embodiment, such SEI information is provided as a VisualSampleGroupEntry and associated with a group of samples, which are signaled as mandatory. One grouping type value can be defined and reserved for each important or mandatory SEI, for example, one value for an alternative transfer characteristics SEI for HDR, one value for a picture timing SEI, one value for a frame packing SEI for stereo applications, and one value for film grain characteristics to improve the decoded image. The payload of these sample group entries corresponds to the payload of additional information. For example, for an SEI message, the NAL units corresponding to the SEI message are provided in a VisualSampleGroupEntry in a SampleGroupDescriptionBox ('sgpd') box. With mandatory sample groups and MIME subparameters, the application knows that it needs the codec to support some additional features in order to process the media file. If the specified sample groups share the same additional information, a single VisualSampleGroupEntry can be used to provide all the NAL units corresponding to these SEIs at once. This allows content creators to indicate as mandatory which SEI messages are expected to be processed by a playback device. This is especially relevant for SEI messages that are not listed in the sample description (which might only be present as a NAL unit array in the decoder configuration record in the sample entry). Having SEI messages in a sample group rather than in a NAL unit array in the decoder configuration makes it easier to handle persistence of SEIs, for example, when they apply to a few samples and not systematically to the entire sequence. Additionally, having SEIs in a sample group allows content creators to indicate on the encapsulation side what is important or required that a playback device, client, or application is expected to support in order to render the media file.Additionally, when exposed through MIME subparameters, parsers and applications can determine whether they can support the media presentation.
[0239] 8 is a schematic block diagram of a computing device 800 for implementing one or more embodiments of the present invention. The computing device 800 may be a device such as a microcomputer, a workstation, or a lightweight portable device. The computing device 800 comprises a communication bus 802 connected to: - a central processing unit (CPU) 804, such as a microprocessor; a Random Access Memory (RAM) 808 for storing executable code of the methods according to embodiments of the invention, as well as registers for recording variables and parameters required for implementing the methods of encapsulation, indexing, decapsulation and / or data access, the memory capacity of which may be expanded, for example, by an optional RAM connected to an expansion port; - a read only memory (ROM) 806 for storing a computer program for implementing an embodiment of the present invention; - a network interface 812, typically connected to a communications network 814 over which the digital data to be processed is sent and received. The network interface 812 may be a single network interface or may consist of a set of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Data is written to the network interface for transmission or read from the network interface for reception under the control of software applications executing on the CPU 804; - A user interface (UI) 816 for receiving input from a user or displaying information to a user; -Hard Disk (HD) 810, and / or - An I / O module 818 for receiving / sending data from / to external devices such as video sources and displays.
[0240] The executable code may be stored either in the read-only memory 806, in the hard disk 810 or on a removable digital medium such as a disk. According to a variant, the executable code of the program may be received by means of a communications network, via the network interface 812, to be stored in one of the storage means of the communications device 800, such as the hard disk 810, before being executed.
[0241] The central processing unit 804 is adapted to control and direct the execution of programs or program instructions or parts of software code according to embodiments of the present invention, these instructions being stored in one of the aforementioned storage means. After power-on, the CPU 804 is capable of executing instructions from the main RAM memory 808 associated with software applications, after those instructions have been loaded, for example, from the program ROM 806 or the hard disk (HD) 810. Such software applications, when executed by the CPU 804, cause the execution of the steps of the flowcharts shown in the previous figures.
[0242] In this embodiment, the device is a programmable device that uses software to implement the invention, but the invention could alternatively be implemented in hardware (e.g. in the form of an application specific integrated circuit or ASIC).
[0243] Although the present invention has been described with reference to specific embodiments, it is understood that the present invention is not limited to those embodiments, and modifications which fall within the scope of the present invention will be apparent to those skilled in the art.
[0244] Many further modifications and variations will be suggested to those skilled in the art by reference to the exemplary embodiments described above, however, these embodiments are given by way of example only and are not intended to limit the scope of the invention, which is determined solely by the appended claims. In particular, different features from different embodiments may be interchanged where appropriate.
[0245] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
Claims
Claim 1 A method for generating a file, comprising: generating at least one track including a plurality of samples; generating a description entry including at least one transformation and / or an order of at least one sample group description, the order being applicable to processing of samples included in the at least one track; generating a file including the at least one track and associating the description entry therewith; A method characterized by including the above steps. Claim 2 The method according to claim 1, characterized in that the file includes information for notifying a file reader that at least one track will not be processed if the file reader fails to recognize a sample group description. The method according to claim 1, characterized by the above. Claim 3 The method according to claim 2, characterized in that in the file, the description entry is associated with the information. The method according to claim 2, characterized by the above. Claim 4 The method according to claim 1, characterized in that the order is described as a list of values including type values of the at least one transformation and / or the at least one sample group description. The method according to claim 1, characterized by the above. Claim 5 The method according to claim 4, characterized in that the list further includes a value indicating a position of a decoding process in a conversion chain including the at least one transformation and / or the at least one sample group description. The method according to claim 4, characterized by the above. Claim 6 The method according to claim 4, characterized in that when the list does not include a value indicating a position of a decoding process, all of the at least one transformation and / or the at least one sample group description are applied to the decoded samples. The method according to claim 4, characterized by the above. Claim 7 The method according to claim 4, characterized in that the value including the type value of the at least one transformation and / or the sample group description in the list is a four-character code. The method according to claim 4, characterized by the above. Claim 8 The method according to claim 1, characterized in that the order is described by a MIME type having a parameter including at least one type value of the transformation and / or the sample group description. The method according to claim 1, characterized by the above. Claim 9 The method according to claim 1, characterized in that the at least one transformation and / or the at least one sample group description are applied before or after processing or decoding at least one sample included in the at least one track. The method according to claim 1, characterized by the above. Claim 10 The file is generated according to the ISO base media file format The method according to claim 1, characterized in that
11. A method for processing a file of media data including at least one track, comprising: obtaining the at least one track from the file; obtaining, from the file, a description entry associated with the at least one track; processing samples included in the at least one track based on the description entry; and the description entry is an order of at least one conversion and / or at least one sample group description, and includes an order applied to processing of samples included in the at least one track, wherein the at least one conversion and / or the at least one sample group description is applied to processing of the samples according to the order, so that the samples are processed A method characterized by this
12. The file includes information for notifying the file reader that at least one track is not processed when the file reader does not recognize a sample group description, Based on the information, not processing at least one track The method according to claim 11, characterized in that
13. In the file, the description entry is associated with the information The method according to claim 12, characterized in that
14. The order is described as a list of values including type values of the at least one conversion and / or sample group description The method according to claim 11, characterized in that
15. The list further includes a value indicating a position of a decoding process in a conversion chain including the at least one conversion and / or the at least one sample group description The method according to claim 14, characterized in that
16. When the list does not include a value indicating a position of a decoding process, all of the at least one conversion and / or at least one sample group description are applied to the decoded samples The method according to claim 14, characterized in that
17. A value including a type value of the at least one conversion and / or sample group description in the list is a four-character code The method according to claim 14, characterized in that
18. The order is described in a MIME type having parameters including at least one type value of the conversion and / or sample group description The method according to claim 11, characterized in that
19. The at least one conversion and / or the at least one sample group description is applied before or after processing or decrypting at least one sample included in the at least one track The method according to claim 11, characterized in that
20. The file is generated according to the ISO base media file format The method according to claim 11, characterized in that
21. A computer program for causing a computer to execute the method according to any one of claims 1 to 20
22. A processing device configured to execute the method according to any one of claims 1 to 20