Method, apparatus, and computer program for encapsulating media data into a media file

By introducing an extractor track that references the first track media sample data in the media file and defining a default constructor and a reference constructor, the cost and complexity of signaling media data in the prior art is solved, and more efficient data transmission is achieved.

CN114503601BActive Publication Date: 2025-05-30CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080067104.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2020-09-22
Publication Date
2025-05-30
Estimated Expiration
2040-09-22

AI Technical Summary

Technical Problem

The prior art has high cost and complex problems in describing and signaling media data to be sent, especially in cases of multiple track references, especially when signaling includes most of the repeated values ​​on all tracks.

Method used

A method is proposed by including a first track and a second track in a media file, the first track contains a media sample, and the second track contains an extractor, the extractor references the media sample data in the first track, and defines a default constructor and a reference constructor in the second track to optimize signaling costs.

Benefits of technology

By reducing the description size of the extractor rail, the cost of signaling is reduced, especially when the constructor attributes are repeated on all rails, the efficiency of data transmission is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114503601B_ABST
    Figure CN114503601B_ABST
Patent Text Reader

Abstract

The present invention relates to providing a default constructor for use by reference in an extractor of an extractor track to reduce the size of the extractor track when using the same constructor in different extractors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods and apparatuses for encapsulating and transmitting media data. Background Art

[0002] The International Organization for Standardization's Base Media File Format (ISO BMFF, ISO / IEC 14496-12) is a well-known flexible and extensible file format that encapsulates and describes encoded timed or untimed media data for local storage or for transmission via a network or via another bitstream delivery mechanism. Examples of extensions are ISO / IEC 14496-15, which describes encapsulation tools for various NAL (Network Abstraction Layer) unit-based video coding formats. Examples of such coding formats are AVC (Advanced Video Coding), SVC (Scalable Video Coding), HEVC (High Efficiency Video Coding), and L-HEVC (Layered HEVC). Another example of a file format extension is ISO / IEC 23008-12, which describes encapsulation tools for still images or sequences of still images (such as HEVC still images, etc.). Another example of a file format extension is ISO / IEC 23090-2, which defines the Omnidirectional Media Application Format (OMAF). The ISO Base Media File Format is object-oriented. It consists of building blocks called boxes (or data structures characterized by a unique type identifier, typically a four-character code, also known as FourCC or 4CC). A full box is a data structure similar to a box that additionally includes version and flag value attributes. Hereinafter, the term "box" may specify either a full box or a box. These boxes or full boxes are hierarchically or sequentially organized in an ISO BMFF file and define parameters that describe the encoded timed or untimed media data, its structure, and its timing (if any). All data in the encapsulated media file (media data and metadata describing the media data) is contained in boxes. There is no other data in the file. A file-level box is a box that is not contained within other boxes.

[0003] In a file format, an entire media presentation is called an animation. The animation is described by an animation box (with four-character code "moov") at the top level of the file. This animation box represents an initialization information container that contains a collection of various boxes describing the media presentation. It is logically divided into tracks represented by track boxes (with four-character code "trak"). Each track (uniquely identified by a track identifier (track_id)) represents a timing sequence of media data belonging to the presentation (e.g., frames of video or audio samples). Within each track, each timing data unit is called a sample; this can be a frame of video, audio, or timing metadata. Samples are implicitly numbered in the decoding order sequence. Each track box contains a hierarchy of boxes that describe the samples of the track. For example, a sample table box ("stbl") contains all the time and data indices of the media samples in the track. The actual sample data is stored in a box at the same level as the animation box called a media data box (with four-character code "mdat") or an identified media data box (with four-character code "imda", similar to a media data box but containing additional identifiers). An animation can be organized in time as an animation box containing information for the entire presentation, followed by a list of media fragments, i.e., a list that couples animation fragments and media data boxes ("mdat" or "imda"). Within an animation fragment (a box with four-character code "moof"), there is a collection of track fragments (boxes with four-character code "traf") that describe the tracks within the media fragment, with each animation fragment having zero or more. A track fragment in turn contains zero or more track run boxes ("trun"), and each track run box records the consecutive runs of the samples of that track fragment.

[0004] An ISOBMFF file can contain multiple encoded timed media data or sub-parts of encoded timed media data that form multiple tracks. When the sub-parts correspond to one or consecutive spatial parts of a video source taken over time (e.g., at least one rectangular region taken over time, sometimes called a "tile" or "sub-picture"), the corresponding multiple tracks can be called tile tracks or sub-picture tracks. ISOBMFF and its extensions include several grouping mechanisms to group tracks, static items, or samples together. Groups typically share common semantics and / or characteristics.

[0005] The inventors have noticed several problems in describing and signaling information about media data to be sent, especially for multiple tracks when one track references another.

[0006] Examples relate to reducing the cost of signaling data entities referenced in another track, especially when signaling mostly repeated values across all tracks.

[0007] Another example relates to optimizing the signaling of the NAL unit length of data entities obtained by an extractor.

[0008] Existing solutions are complex or not well - defined. Summary of the Invention

[0009] The present invention aims to solve one or more of the above - mentioned problems.

[0010] In this context, a solution is provided for streaming media content (e.g., omnidirectional media content) over an IP network such as the Internet using, for example, the HTTP protocol.

[0011] According to a first aspect of the present invention, a method for encapsulating media data into a media file is proposed. The method includes: including a first track in the media file, the first track including media samples; including a second track in the media file, the second track including an extractor, the extractor being a structure that references data in the media samples included in the first track, the extractor including at least one constructor, and wherein the method further includes: including a default constructor in the second track; wherein the constructor is a reference constructor that references the default constructor.

[0012] In an embodiment, the default constructor is included in the metadata section of the second track.

[0013] In an embodiment, the default constructor in the second track is included as a list of default constructors, and the reference constructor includes an index in the list.

[0014] In an embodiment, the default constructor is included in the sample entry of the metadata section of the second track.

[0015] In an embodiment, the default constructor is included in the sample group entry that describes the sample group of the second track, and the extractor is included in the samples of the sample group.

[0016] According to another aspect of the present invention, a method for encapsulating media data into a media file is proposed. The method includes: including a first track in the media file, the first track including media samples; including a second track in the media file, the second track including an extractor, the extractor being a structure that references data in the media samples included in the first track, and wherein the method further includes: including a default extractor in the second track; the extractor included in the second track is a reference extractor that references the default extractor.

[0017] According to another aspect of the present invention, a method for encapsulating media data into a media file is provided. The method includes: including a first track in the media file, the first track including media samples, each media sample containing a set of one or more NAL units; including a second track in the media file, the second track including an extractor, the extractor being a structure that references data in the media samples included in the first track, the extractor including at least one constructor, and wherein the constructor includes inlined data and information indicating that the inlined data does not include any NAL unit length fields.

[0018] According to another aspect of the present invention, a method for parsing media data into a media file is provided. The method includes: obtaining a first track including media samples in the media file; obtaining a second track including an extractor in the media file, the extractor being a structure that references data in the media samples included in the first track, the extractor including at least one constructor, and wherein the method further includes: obtaining a default constructor in the second track; wherein the constructor is a reference constructor that references the default constructor; obtaining, based on the default constructor, the data referenced by the extractor from the first track.

[0019] According to another aspect of the present invention, a method for parsing media data into a media file is provided. The method includes: obtaining a first track including media samples in the media file; obtaining a second track including an extractor in the media file, the extractor being a structure that references data in the media samples included in the first track, and wherein the method further includes: obtaining a default extractor in the second track; the extractor included in the second track being a reference extractor that references the default extractor; obtaining, based on the default extractor, the data referenced by the extractor from the first track.

[0020] According to another aspect of the present invention, a method for parsing media data into a media file is provided. The method includes: obtaining a first track including media samples in the media file, each media sample containing a set of one or more NAL units; obtaining a second track including an extractor in the media file, the extractor being a structure that references data in the media samples included in the first track, the extractor including at least one constructor; and wherein the constructor includes inlined data and information indicating that the inlined data does not include any NAL unit length fields.

[0021] According to another aspect of the present invention, there is provided a computer program product for a programmable device, the computer program product comprising a sequence of instructions for implementing the method according to the present invention when loaded into and executed by the programmable device.

[0022] According to another aspect of the present invention, there is provided a computer-readable storage medium storing instructions of a computer program for implementing the method according to the present invention.

[0023] According to another aspect of the present invention, there is provided a computer program which, when executed, causes the method according to the present invention to be performed.

[0024] According to another aspect of the present invention, there is provided a device for encapsulating media data into a media file, the device comprising a processor configured to: include a first track in the media file, the first track including media samples; include a second track in the media file, the second track including an extractor which is a structure that references data in the media samples included in the first track, the extractor including at least one constructor, and wherein the method further includes: including a default constructor in the second track; including a reference constructor in the extractor that references the default constructor.

[0025] According to another aspect of the present invention, there is provided a device for encapsulating media data into a media file, the device comprising a processor configured to: include a first track in the media file, the first track including media samples; include a second track in the media file, the second track including an extractor which is a structure that references data in the media samples included in the first track, and wherein the method further includes: including a default extractor in the second track; the extractor included in the second track being a reference extractor that references the default extractor.

[0026] According to another aspect of the present invention, there is provided a device for encapsulating media data into a media file, the device comprising a processor configured to: include a first track in the media file, the first track including media samples, each media sample containing a set of one or more NAL units; include a second track in the media file, the second track including an extractor which is a structure that references data in the media samples included in the first track, the extractor including at least one constructor; and wherein the constructor includes inline data and information indicating that the inline data does not include any NAL unit length fields.

[0027] According to another aspect of the present invention, there is provided an apparatus for parsing media data into media files. The apparatus includes a processor configured to: obtain a first track including media samples in the media file; obtain a second track including an extractor in the media file, the extractor being a structure that references data in the media samples included in the first track, the extractor including at least one constructor; and wherein the method further includes: obtaining a default constructor in the second track; obtaining a reference constructor that references the default constructor in the extractor; and obtaining, based on the default constructor, the data referenced by the extractor from the first track.

[0028] According to another aspect of the present invention, there is provided an apparatus for parsing media data into media files. The apparatus includes a processor configured to: obtain a first track including media samples in the media file, each media sample including a set of one or more NAL units; obtain a second track including an extractor in the media file, the extractor being a structure that references data in the media samples included in the first track; and wherein the method further includes: obtaining a default extractor in the metadata portion of the second track; the extractor included in the second track being a reference extractor that references the default extractor; and obtaining, based on the default extractor, the data referenced by the extractor from the first track.

[0029] According to another aspect of the present invention, there is provided an apparatus for parsing media data into media files. The apparatus includes a processor configured to: obtain a first track including media samples in the media file, each media sample including a set of one or more NAL units; obtain a second track including an extractor in the media file, the extractor being a structure that references data in the media samples included in the first track, the extractor including at least one constructor; and wherein the constructor includes inline data and information indicating that the inline data does not include any NAL unit length fields.

[0030] Other aspects of the present invention relate to computing devices and corresponding computer programs for encapsulating media data and parsing media files. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Other advantages of the present invention will become apparent to those skilled in the art after studying the drawings and the detailed description. The present invention aims to incorporate any additional advantages.

[0032] Embodiments of the present invention will be described below only by way of example and with reference to the following drawings, wherein:

[0033] Figure 1Shows an exemplary system including an encapsulation / de-encapsulation module suitable for embodying embodiments of the present invention;

[0034] Figure 2 Shows an example structure of a NAL (Network Abstraction Layer) unit;

[0035] Figure 3a Shows an example structure of a video media sample according to the ISO base media file format;

[0036] Figure 3b Shows another example structure of a video media sample according to the ISO base media file format;

[0037] Figure 4 Shows an example extractor and aggregator structure according to ISO / IEC 14496-15;

[0038] Figure 5a and 5b Shows an encapsulation and de-encapsulation process according to embodiments of the present invention;

[0039] Figure 6 Shows an encapsulation file format according to embodiments of the present invention;

[0040] Figure 7 Is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. Detailed Description of the Invention

[0041] Figure 1 Shows exemplary systems 191 and 195 suitable for embodying embodiments of the present invention. System 191 includes an encapsulation module 150 connected to a communication network 199. System 195 includes a de-encapsulation module 100 connected to the communication network 199.

[0042] According to an embodiment, system 191 is used to process content, such as video, still images, and / or audio content, for streaming or storage. System 191 obtains / receives content including a set of raw non-timed images or a timed image sequence 151, encodes the set of non-timed images or the timed image sequence into encoded media data using a media encoder (e.g., an image or video encoder), and encapsulates the encoded media data in a media file 101 using the encapsulation module 150. The encapsulation module 150 includes at least one of a writer or a packer for encapsulating the encoded media data. The media encoder can be implemented within the encapsulation module 150 to encode the received content, or can be separated from the encapsulation module 150. Thus, the encapsulation module 150 can be dedicated only to encapsulating already encoded content (encoded media data). The encoding step is optional, and the encoded media data can correspond to raw media data.

[0043] According to an embodiment, system 195 is used to process the encapsulated encoded media data for display / output to a user. System 195 obtains / receives media file 101 via communication network 199 or by reading a storage device, uses demultiplexing module 100 to demultiplex media file 101 to retrieve the encoded media data, and uses a media decoder to decode the encoded media data into audio and / or video content (signals). The demultiplexing module 100 includes at least one of a parser or a player. The media decoder may be implemented within the demultiplexing module 100 to decode the encoded media data, or may be separated from the demultiplexing module 100.

[0044] Media file 101 is communicated to the parser or player of module 100 in various ways. For example, it may be pre-generated by the writer or packager of encapsulation module 150 and stored as data in a storage device in communication network 199 (e.g., on a server or cloud storage device) until the user requests the content encoded therein from the storage device. When the content is requested, the data is communicated / streamed from the storage device to demultiplexing module 100.

[0045] System 191 may also include a content providing device for providing / streaming content information for the content stored in the storage device to the user (e.g., the content information may be described via a manifest file that includes the title of the content and other descriptive metadata and storage location data for identifying, selecting, and requesting the content). The content providing device may also be adapted to receive and process user requests for the content to be communicated / streamed from the storage device to the user terminal.

[0046] Alternatively, encapsulation module 150 may generate media file 101 and directly communicate / stream it to demultiplexing module 100 when the user requests the content. Then, demultiplexing module 100 receives media file 101 and demultiplexes and decodes the media data according to an embodiment of the present invention to obtain / generate video signal 109 and / or audio signal, and then the user terminal uses video signal 109 and / or audio signal to provide the requested content to the user.

[0047] The user may access the audio / video content (signals) through a user interface of a user terminal including module 100 or a user terminal having components communicating with module 100. Such a user terminal may be a computer, a mobile phone, a tablet, or any other type of device capable of providing / displaying content to the user.

[0048] According to one embodiment, the media file 101 encapsulates the encoded media data (e.g., encoded audio or video) into boxes according to the ISO Base Media File Format (ISOBMFF, ISO / IEC 14496-12, and ISO / IEC 14496-15 standards). The media file 101 may correspond to a single media file (prefixed with FileTypeBox "ftyp") or one or more segment files (prefixed with SegmentTypeBox "styp") following a media file (prefixed with FileTypeBox "ftyp"). According to ISOBMFF, the media file 101 may include two types of boxes; "media data" boxes ("mdat" or "imda") containing media data and "metadata boxes" ("moov" or "moof" or "meta" box hierarchies) containing metadata defining the arrangement and timing of the media data.

[0049] An image or video encoder encodes image or video content using an image or video standard to generate encoded media data. For example, image or video coding / decoding (codec) standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 (ISO / IEC MPEG-2 Visual), ITU-T H.263 (ISO / IEC MPEG-4 Visual), ITU-T H.264 (ISO / IEC MPEG-4 AVC) (including its Scalable Video Coding (SVC) and Multi-View Video Coding (MVC) extensions), ITU-T H.265 (HEVC) (including its Scalable (SHVC) and Multi-View (MV-HEVC) extensions).

[0050] Many of the embodiments described herein use the HEVC standard or its extensions to describe examples. However, the techniques and systems described herein may also be applicable to other coding standards that are already available (such as AVC) or other coding standards that are not yet available or developed (such as ITU-T H.VVC (ISO / IEC MPEG-I VVC) (Versatile Video Coding) under specification).

[0051] Figure 2 A structural example of a NAL (Network Abstraction Layer) unit 200 used in a video codec such as H.264 / AVC, HEVC / H.265, or VVC is shown.

[0052] The NAL unit contains an NAL unit header 201 and an NAL unit payload 202. The NAL unit header 201 has a fixed length and provides general information about the NAL unit. For example, in HEVC, the NAL unit header 201 indicates the type of each NAL unit, the identifier of the layer, and the identifier of the temporal sublayer. There are two main types of NAL units 200: Video Coding Layer NAL units (VCL-NAL) and non-VCL NAL units. VCL NAL units typically contain a coded slice segment 205 in their payload. Non-VCL NAL units typically contain parameter sets (i.e., configuration information) or supplementary enhancement information messages.

[0053] The coded slice segment 205 is encoded in the HEVC bitstream (coded media data) as a slice_segment_header or "slice header" 206, followed by slice_segment_data or "slice data" 207, and then followed by rbsp_slice_segment_trailing_bits or "slice trailing bits" 208 to ensure byte alignment. The slice segment contains an integral number of consecutive (in raster scan order) coding tree units (i.e., blocks in the picture). The slice does not necessarily have a rectangular shape (and thus is less suitable compared to blocks for spatial subpart representation). The video compression format defines an access unit as a set of consecutive NAL units in decoding order corresponding to a coded picture.

[0054] Figure 3a An example of the structure of a media (video) sample 300 according to the ISO base media file format is shown.

[0055] A media sample is an audio / video data unit having a single time (e.g., an audio sample or a video frame). According to ISO / IEC 14496-15, a sample is a collection of one or more NAL units 302 (similar to the NAL unit 200 shown Figure 2 in) corresponding to an access unit or a part of an access unit. In front of each NAL unit 302 is an NAL unit length field that provides the length of the NAL unit 302 in bytes. For example, for single-layer video, a sample corresponds to a coded picture. For layered video, a sample can correspond to a part of an access unit, such as NAL units corresponding to the base layer.

[0056] The sample size in bytes is described in the sample size boxes "stsz" or "stsz2". Given the sample size and the NAL unit length, the ISOBMFF parser (e.g., module 100) can determine the number of NAL units in the sample. ISO / IEC 14496-15 defines specific NAL units, extractors, and aggregators, which are specific ISOBMFF structures embedded within the media data ("mdat" box). They are also referred to as "in-stream structures". An extractor is an in-stream structure that uses the NAL unit header to extract data from other tracks. The extractor contains instructions on how to extract data from other tracks. Logically, the extractor can be regarded as an indicator pointing to the data. When reading the track containing the extractor, the extractor is replaced by the data it points to. The extractor can also directly contain the data to be used when replacing the extractor (inline data). An aggregator is an in-stream structure that groups NAL units belonging to the same sample using the NAL unit header.

[0057] Figure 3b Another example of the structure of a media (video) sample 310 according to the ISO Base Media File Format is shown.

[0058] According to ISO / IEC 14496-15, the sample format of a track of type "hvt2" can consist of one and only one instance of the HEVC syntax elements slice_segment_data() 311 and rbsp_slice_segment_trailing_bits() 312 of an independent slice segment. No other data is present in the sample. Typically, this particular kind of sample format intentionally does not include HEVC syntax elements for applications that may need to rewrite these syntax elements, such as the NAL unit header and slice_segment_header(). This occurs when samples from different tracks contain NAL units that can be combined and merged to obtain a single bitstream of encoded media data to be decoded. For example, when different tracks contain NAL units corresponding to different spatial sub-parts of the video content. In this case, it will be necessary to rewrite the NAL unit header and slice_segment_header() during the merge operation to obtain a consistent bitstream of encoded media data, e.g., the slice addresses need to be recalculated and changed in the resulting bitstream. To avoid this rewriting, these headers are not encapsulated in the track containing slice_segment_data() as usual, but instead they are encapsulated in a separate track (common to all tracks to be merged). This track itself contains an extractor and a slice segment header. The extractor is a specific structure that allows referencing and extracting data from the samples of another referenced track. By resolving as Figure 4The extractor shown is used to obtain samples of the bitstream to be decoded (including the NAL unit length field, the NAL unit header, the slice header, the slice data, and the slice trailer bits).

[0059] Figure 4 An example extractor structure according to ISO / IEC 14496-15 is shown.

[0060] The first (media) track 400 corresponding to the encoded media data (e.g., the encoded video bitstream) includes media (e.g., video) samples 401, each media sample containing a set of one or more NAL units (preceded by their respective NAL unit length fields), as Figure 3a shown. The number of NAL units in the samples can be different from each other. The second track 420 (referred to as, for example, the reconstruction, synthesis, reference, or extractor track) includes samples 421 for mixing the NAL units 423 and the extractor NAL units 422 to reference data from another track (here, the first track 400 as shown by the arrow 410). Then, the samples 421 from the current track 420 can be reconstructed by extracting data from the track 400 and concatenating it with the data 423 from the current track 420.

[0061] It can be noted that some samples 421 of the reconstruction track 420 can contain only extractors or only data. The number of extractors or NAL units can vary from one sample 421 to another. The extractors are NAL units, and they are also preceded by NAL unit length fields like any other NAL unit. The extractor 422 is a structure that enables the efficient extraction of NAL units from a track other than the track containing the extractor. The extractor NAL units are identified by a specific NAL unit type value (the specific value can depend on the codec in use so as not to conflict with the codec-specific type values assigned to VLC and non-VLC NAL units).

[0062] ISO / IEC 14496-15 defines extractors for different compression formats (SVC, MVC, HEVC, …). For HEVC, the extractor is defined as a set of constructors as follows:

[0063]

[0064] The extractor first consists of a NALUnitHeader(), followed by one or more constructors for extracting data from the current track or from another track linked to the track where the extractor resides through a track reference of type “scal”.

[0065] Several types of constructors are specified:

[0066] The sample constructor (constructor_type = 0) extracts NAL unit data from samples on another track by reference.

[0067] The inline constructor (constructor_type = 1) includes NAL unit data.

[0068] The sample constructor from a track group (constructor_type = 2) extracts NAL unit data (either the whole NAL unit or the NAL unit payload according to the copy mode) from samples on a track in another track or track group by reference.

[0069] For HEVC, NALUnitHeader() corresponds to the first two bytes of an ISO / IEC 23008-2 NAL unit, where the specific NAL unit type value (nal_unit_type) is equal to the reserved value of the extractor (= 49).

[0070] EndOfNALUnit() is a function that returns 0 (false) when more data follows in the extractor; otherwise it returns 1 (true).

[0071] Each constructor can contain multiple attributes to characterize the data to be extracted.

[0072] The extractor shall be declared in each sample of the extractor track that references NAL unit data or partial NAL unit data from samples on another track or track group. In typical uses of sub-picture / strip / block extraction, the extractor is mainly a repetitive structure, and most of its constructor attributes are simply repeated from one extractor to another. For example, the sample constructor from a track group and the sample constructor have a track_ref_index attribute that specifies the index of a reference to a track of type "scal" to find the track from which to extract data. Typically, the value of this attribute is not intended to change frequently from one sample to another. As another example, when the extraction of data only involves temporally aligned samples of other tracks, the attribute sample_offset is usually 0 in all extractors.

[0073] According to one aspect of the present invention, a mechanism for signaling costs of an extractor is proposed by declaring default constructors and defining new constructors that reference these default constructors. One advantage is that declaring an extractor that references a default constructor allows reducing the description size of each extractor in each sample of the extractor track when the attributes of the constructor are mostly repeated values across all tracks. A list of default constructors can be defined in the metadata section of the extractor track. In the media data section, in a sample, a new type of constructor that references a default constructor in the list of default constructors is proposed, namely a reference constructor. For example, a reference constructor can contain an index in the list of default constructors. Thus, a constructor that is used multiple times in the extractor track can be sent once in the list of default constructors.

[0074] When the extractor is not based on constructors, the same mechanism can be applied at the extractor level. A list of default extractors can be defined in the metadata section of the extractor track. A new type of extractor can be defined, namely a reference extractor, for referencing a default extractor. The reference extractor can be used in the media section of the extractor track in a sample to reference a default extractor in the list of default extractors.

[0075] According to another aspect of the present invention, a new type of inline constructor is proposed. The inline constructor can be used to provide a new NAL unit header. The occurring inline constructor contains a NAL unit length field and can rewrite this field according to other constructors used to provide the payload of the NAL unit. When the NAL unit length field has been provided by the inline constructor and when the NAL unit length field needs to be rewritten, its transmission in the inline constructor is useless. The proposed new type of inline constructor explicitly does not contain any NAL unit length field, thus saving its transmission when it can be generated during the parsing of the extractor track.

[0076] Figure 5a An encapsulation process according to an embodiment of the present invention is shown. In the embodiment, the process is performed by a writer or packer of the encapsulation module 150 shown Figure 1 to encapsulate media data.

[0077] At step 500, the encapsulation module is initialized to be able to appropriately read the bitstream of the encoded media data. The initialization can be performed by a user through a user interface or by an application. The initialization can involve identifying the syntactic structure of the bitstream (generally referred to as a data entity) and configuring encapsulation parameters. The configuration of the encapsulation can consist of deciding, for example, whether to generate the media file 101 as a single media file or as multiple media files constituting a media segment composed of one or more media shards; including one track or multiple tracks for video streams in the media file; setting to split the video track into parts, views, or layers, etc.

[0078] When multiple tracks are included, the encapsulation module may set references between tracks or define track groups during step 500. Thus, a track constructed by reference to one or more other tracks contains a track reference to those one or more tracks. Track references may have different types to describe the kind of relationship or dependency between the referencing track and the referenced track. Track reference types may be encoded using four-character codes. For example, according to ISO / IEC 14496-15, the type code "scal" designates a track that contains at least an extractor that extracts data from another track from which the reference is made.

[0079] The extractor implementations specified in ISO / IEC 14496-15 for H.264 / AVC and HEVC achieve a compact form of a track that extracts NAL unit data by reference. The extractor is a NAL-unit-like structure. A NAL-unit-like structure may be specified to include a NAL unit header and a NAL unit payload like any NAL unit, but start code emulation prevention (which is required for NAL units) may not be followed in a NAL-unit-like structure. For HEVC, the extractor contains one or more constructors. A sample constructor extracts NAL unit data from samples of another track by reference. An inline constructor includes the NAL unit data. When the extractor is processed by a file reader that requires it, the extractor is logically replaced by the bytes produced when the contained constructors are decomposed in presentation order. Nested extraction may not be allowed; for example, the bytes referenced by a sample constructor should not contain an extractor; an extractor should not directly or indirectly reference another extractor. The extractor may contain one or more constructors that are used to extract data from the current track or from another track linked to the track on which the extractor resides by a track reference of type "scal". The bytes of the decomposed extractor may represent one or more whole NAL units. The decomposed extractor starts with a valid length field and a NAL header. The bytes of a sample constructor are copied only from a single identified sample in the track referenced by the indicated "scal" track reference. Alignment is at decoding time, i.e., only the time-to-sample table is used, followed by a count offset in the sample number. An extractor track may be defined as a track that contains one or more extractors.

[0080] Once the encapsulation module is initialized, the bitstream is read NAL unit by NAL unit at step 501. Depending on the initialization at step 500 (in-band or out-of-band parameter sets), the first NAL unit corresponding to the parameter set can be embedded in the DecoderConfigurationRecord structure. These parameter sets can be examined by the writer or packager to learn more about the bitstream partitioning and to determine the number of individual tracks to create, one track per partition, plus one or more extractor tracks, each extractor track referring to one or more tracks containing the partition using a track reference of type "scal". Thus, each extractor track represents a possible combination of all or a subset of the partition tracks. For example, it can be determined whether it is a segmented HEVC bitstream by examining, for example, SEI (Supplemental Enhancement Information) messages for the presence of blocks in the Temporal Motion Constrained Block Set or Picture Parameter Set. A track can be defined for each HEVC block and one extractor track describing the composition of all HEVC blocks. When reading an NAL unit at step 501, the writer checks at step 502 whether the NAL unit corresponds to a new sample. This can be done, for example, by decoding the picture order count or by checking whether the slice corresponding to the NAL unit is the first slice in the picture. If so, the previous sample is completed at step 503 by setting the parameters of the sample description (size, position in the media data, attributes in some sample groups...). In particular, the sample size is reset to 0 and the NAL unit count is reset to 0. Then, at step 504, it is checked whether the current NAL unit should be included in the media part of an extractor track, or should be included in the media part of a partition track and referenced from an extractor track, or should be partially included in the media part of a partition track and partially modified and referenced by an extractor track. This is determined according to the track dependencies or relationships established during the initialization step 500. If the NAL unit is not referenced, the length of the NAL unit is first inserted into the media data frame of the extractor track, followed by the NAL unit header and the NAL unit payload (step 505). Then, the size of the current sample is incremented by the number of bytes of these three structures, and at step 506, the writer or packager checks the next NAL unit from the bitstream of the encoded media data. If this is not the last NAL unit, the process iterates to step 501 until all NAL units have been processed.

[0081] According to an embodiment of the present invention, if a NAL unit is to be included partially or completely in an extractor track by reference (test 504 is true), then at step 507 the writer or packager includes in the media data box of the partition track the NAL unit data to be referenced (if the NAL unit data is a complete NAL unit, with the preceding NAL unit length field) and includes in the extractor track an extractor having one or more constructors and associated attributes. Specifically, the process appends in bytes in the media data box of the extractor track a NAL unit length field having the size of the extractor structure and creates an extractor NAL unit. When the extractor includes a constructor corresponding to a repeating structure along the samples, a default constructor can be inserted into the default constructor list in the sample entry of the extractor track (if it does not already exist). The created extractor NAL unit includes a reference constructor that provides an index to the default constructor to be used from the default constructor list.

[0082] When the extractor is written into the media data box, the sample description (sample size, current NALU index in the sample, etc.) is updated. Then, at step 506, the writer or packager checks the next NAL unit. When the last NAL unit is reached, the writer terminates the media file at step 508, for example by writing the size of the last sample, the index table, user data, or any metadata about the media.

[0083] Note that when the initialization step 500 indicates encapsulation into segments, an additional test (not shown) is performed before starting a new sample to check if the segment duration has been reached. When the segment duration is reached, the segment is complete and ready to be used by the player or sent over the distribution network. When the segment duration is not reached, the writer or packager iterates over the samples and NAL units.

[0084] Figure 5b An example of a de-encapsulation process according to an embodiment of the present invention is shown. In an implementation, the process is performed by Figure 1 the parser or player of the de-encapsulation module 100 shown for de-encapsulating media data.

[0085] At step 510, the player first receives the media file 101 (as a single file or as consecutive segments). The file can be stored in the memory of the parser or player, or can be read from a network socket.

[0086] First, at step 511 the initialization data is parsed, typically the "moov" box and its sub-boxes, to know the parameters / settings of the media file: number of tracks, track relationships and dependencies, sample types, duration, location, and size, etc.

[0087] From the set of tracks determined at step 511, the player or parser selects, at step 512, one extractor track to be rendered. Reconstruction then begins by parsing the media data frames sample by sample.

[0088] The parser or player iterates over the samples until the end of the file is reached (test 513 is false). In the case of segments, when a segment is fully read, the parser reads the next segment sample by sample.

[0089] For a given sample, the process reads data from the position indicated by the chunk offset box plus the cumulative size of the previous samples parsed for that chunk. From that position, the parser locates the NAL unit length field. The parser then reads the number of bytes given by the NAL unit length field to obtain the complete NAL unit. If the NAL unit corresponds to an extractor (test 515), the parser reads, at step 516, the constructor type attribute of the next constructor in the extractor. If the constructor is a reference constructor according to an aspect of the present invention, the parser retrieves, at step 517, the index of the default constructor from the reference constructor, which will be used to decompose the constructor at step 518. The default constructor is then retrieved from a list of default constructors (index values representing positions in the list), which is defined in the sample entry or in the sample group description associated with the current sample via the sample grouping mechanism. The default constructor can be any type of constructor other than another reference constructor (e.g., inline, sample, …). At step 518, the constructor is decomposed according to the semantics associated with the type of constructor (retrieved directly in the extractor NALU or referenced via a reference constructor). At step 520, the data obtained by decomposing all the constructors of the extractor NAL unit (test 519) is appended to the reconstructed bitstream of the encoded media data.

[0090] If the NAL unit is not an extractor at step 515, the parser appends, at step 520, the bytes corresponding to the NAL unit header and the NAL unit payload (without the NAL unit length) to the reconstructed bitstream of the encoded media data that will be provided to the media decoder for decoding. After step 520, the process iterates for the next NAL unit (goes to step 514) until the size of the current sample is reached.

[0091] In the following, examples are provided to illustrate the constructors of the extractor and the default values of the extractor constructors proposed according to an embodiment of the present invention. The new constructor implements a reference index attribute, as discussed above in steps 507 and 516 of the encapsulation / decapsulation process in Figure 5a and 5b the encapsulation / decapsulation process.

[0092] The implementation for specifying the reference index for the default constructor or default extractor applies to both extractors without constructors (such as SVC, MVC extractors, etc.) and extractors with constructors (such as HEVC or L-HEVC extractors, etc.). For extractors with constructors, a new type of constructor (identified by "constructor_type") can be defined as follows:

[0093]

[0094] where constructor_type specifies the subsequent constructor. SampleConstructor, InlineConstructor, SampleConstructorFromTrackGroup, and ReferenceConstructor correspond to constructor_type equal to 0, 2, 3, and 4 respectively. Other values of constructor_type are reserved.

[0095] The name "ReferenceConstructor" of the new constructor is provided as an example. Additionally, the reserved "constructor_type" value "4" is provided as an example. Instead of indicating ("SampleConstructor") or providing ("InlineConstructor") byte ranges, the new constructor allows referencing the default constructor declared in the metadata section of the extractor track, typically in a list within the sample entry or sample group description box. Any reserved name or reserved value of "constructor_type" can be used. Generally, this new constructor_type value can exist and be used in samples of HEVC or L-HEVC tracks with sample entries of type "hvc3" or "hev3" or samples of any other type of track defined later. The new constructor can be defined, for example, as shown below:

[0096]

[0097] where the parameters, fields, or properties of the new constructor have the following semantics:

[0098] - ref_index specifies the index of the constructor to be used in the constructor list of the DefaultHevcExtractorConstructorBox in the sample entry of the track from which data is to be extracted. The value 0 indicates the first entry.

[0099] As Figure 6As shown in the example of the encapsulation file of the extractor track, the default values of the extractor constructors are declared in the sample entry 602 of the extractor track 600. The extractor track is a track with a track reference box ("tref") 601 that contains a TrackReferenceTypeBox with a reference_type equal to "scal" (the reference of which can be the track_ID (such as 603 and 604) of the track from which the extractor can extract data or the track_group_id of a track group (not shown)). Declaring the default values of the extractor constructors in the sample entry ensures that their declarations can be shared by all samples in the extractor track.

[0100] In an embodiment of an HEVC or L-HEVC extractor, a new box DefaultHevcExtractorConstructorBox 605 is defined in the sample entry 602 as follows (similar embodiments can be easily derived for extractors of other video formats other than HEVC, such as AVC, SVC, MVC, VVC, etc.):

[0101]

[0102] num_entries gives the number of default constructors defined in this box.

[0103] constructor_type gives the type of the constructor of this entry. The value 4 (i.e., ReferenceConstructor) should not exist.

[0104] The DefaultHevcExtractorConstructorBox provides a list of constructors for replacing the ReferenceConstructor present in the sample. The constructors given in this box should only be resolved as replacements for the ReferenceConstructor constructors in the sample. The DefaultHevcExtractorConstructorBox can contain multiple constructors with the same constructor_type. One advantage of being able to declare multiple constructors with the same constructor_type is the possibility of changing the default values to be applied over time in the extractor track (e.g., due to some Picture Parameter Set (PPS) changes that may require minor modifications to the extractor definition). In this case, the extractor can simply include ReferenceConstructors with different ref_index values to reference different default constructors in the DefaultHevcExtractorConstructorBox, and thus use different default attribute values for constructors of the same type, as Figure 6 shown. Two samples 606 and 607 of the extractor track 600 include extractors 608 and 609, which each contain a ReferenceConstructor 610 and 611 respectively, where the ref_index attribute references different alternatives of the default constructor.

[0105] To better maintain the specification, the extractor and the DefaultHevcExtractorConstructorBox can also be redefined as follows, with the same semantics as above:

[0106]

[0107] In the above embodiments, SampleConstructor, InlineConstructor, and SampleConstructorFromTrackGroup are defined as specified in ISO / IEC 14496-15, Edition 5 and its amendments.

[0108] In the above embodiments, the design of defining the default constructor in the DefaultHevcExtractorConstructorBox located in the sample entry does not allow the list of default constructors to be updated once the moov is generated, which may force the moov to be reloaded when switching qualities. However, when the default constructor is no longer applied, for example, after switching qualities, it is still possible to build a bitstream of the encoded media data that conforms to the above embodiments by no longer using the default constructor.

[0109] If it is desired to be able to update the default constructor list as soon as moov is generated, an alternative embodiment is to use a sample grouping mechanism and a new sample group definition to replace the declaration of the DefaultHevcExtractorConstructorBox in the sample entry. This alternative embodiment would require a double reference. First, each sample of the extractor track is associated with a visual sample group entry defined in a sample group description box ("sgpd") with a given grouping_type (this association is done via the group_description_index of the SampleToGroupBox with the same grouping_type), and second, the ref_index from the ReferenceConstructor should reference an entry in the constructor list in the visual sample group entry associated with the sample.

[0110] According to this embodiment, a new grouping_type is defined, for example "dhec", and the new VisualSampleGroupEntry is defined as follows:

[0111]

[0112] The name of the new visual sample group entry "DefaultHevcExtractorConstructorSampleGroupEntry" is provided as an example. In addition, the reserved 4CC value "dhec" is provided as an example. All attributes in this new class have the same semantics as the DefaultHevcExtractorConstructorBox.

[0113] Then, for example, the semantics of the ref_index in the ReferenceConstructor are defined as follows:

[0114] The ref_index specifies the index of the constructor to be used in the constructor list of the DefaultHevcExtractorConstructorSampleGroupEntry associated with the sample via the group_description_index of the SampleToGroupBox ("sbgp") with a grouping_type equal to "dhec". The value 0 indicates the first entry.

[0115] With such an embodiment, it is possible to use the default sample group mechanism defined in ISO / IEC 14496-12 to associate the default visual sample group entry "DefaultHevcExtractorConstructorSampleGroupEntry" with samples without defining a SampleToGroupBox with the same grouping_type. When the list of default constructors does not change over time, this allows saving some bytes (corresponding to the definition of the SampleToGroupBox).

[0116] With such an embodiment, it is also possible to update the list of default constructors by creating new ISOBMFF animation fragments (several boxes "moof" and "mdat"), where a new DefaultHevcExtractorConstructorSampleGroupEntry is defined in a new SampleGroupDescriptionBox ("sgpd") with grouping_type "dhec" in the track fragment ("traf") corresponding to the extractor track in the animation fragment box ("moof"). Additionally, a new SampleToGroupBox with the same grouping_type "dhec" can be defined in the animation fragment to associate the new DefaultHevcExtractorConstructorSampleGroupEntry with the samples of the extractor track of the animation fragment.

[0117] In a variant, the new visual sample group entry DefaultExtractorConstructorSampleGroupEntry() can also include a unique identifier groupID as follows:

[0118]

[0119] where:

[0120] groupID is the unique identifier of the default extractor constructor group described by this sample group entry. The value of groupID in the default extractor constructor group entry shall be greater than 0. The value 0 is reserved for special purposes.

[0121] When there is a SampleToGroupBox of type "nalm" and the grouping_type_parameter is equal to "dhec", there shall be a SampleGroupDescriptionBox of type "dhec", and the following applies:

[0122] - The value of groupID in the default extractor constructor group entry shall be equal to the groupID in one of the entries of NALUMapEntry.

[0123] - Mapping of an NAL unit to groupID 0 via NALUMapEntry implies that the NAL unit is not an extractor or that the NAL unit does not reference any default constructor.

[0124] All other attributes in this new class have the same semantics as DefaultHevcExtractorConstructorBox.

[0125] With this variant, the NAL unit mapping mechanism defined in ISO / IEC 14496-15 can be used to associate different lists of default constructors (instead of a single list of default constructors for all NAL units in the sample) with individual NAL units in the sample. In this case, NALUMapEntry can be used to assign an identifier called groupID to each NAL unit in the sample. NALUMapEntry is a visual sample group entry with a grouping_type equal to "nalm" and is defined as follows:

[0126]

[0127] where:

[0128] large_size indicates whether the number of NAL unit entries in the track sample is represented in 8 bits or 16 bits.

[0129] rle indicates whether run-length encoding is used (1) to assign groupID to NAL units or (0) not to assign groupID to NAL units.

[0130] entry_count specifies the number of entries in the mapping. Note that when rle is equal to 1, entry_count corresponds to the number of runs of consecutive NAL units associated with the same group. When rle is equal to 0, entry_count represents the total number of NAL units.

[0131] NALU_start_number is the 1-based NAL unit index in the sample of the first NAL unit in the current run associated with groupID.

[0132] The groupID specifies the unique identifier of the group. More information about the group is provided by the sample group description entry, where this groupID and grouping_type are equal to the grouping_type_parameter of the SampleToGroupBox of type "nalm".

[0133] The NALUMapEntry (when present) shall be linked to the sample group description providing the semantics of this groupID. This link shall be provided by setting the grouping_type_parameter of the SampleToGroupBox of type "nalm" to the four-character code of the associated sample grouping type.

[0134] According to this variant, a SampleToGroupBox is defined with a grouping_type equal to "nalm" and a grouping_type_parameter equal to "dhec" to associate each sample of the extractor track with a NALUMapEntry, where each NALUMapEntry is defined in a SampleGroupDescriptionBox ("sgpd") with a grouping_type equal to "nalm" and describes the association of the groupID with each NAL unit in the sample. Another SampleGroupDescriptionBox ("sgpd") with a grouping_type equal to "dhec" is defined using a set of DefaultExtractorConstructorSampleGroupEntry, each entry providing a list of default constructors corresponding to the groupID associated with the NAL unit. Finally, the ref_index in the ReferenceConstructor contained in the NAL unit provides the index of the constructor to be used in the list of default constructors.

[0135] Similar embodiments can be easily derived to associate different lists of default extractors (instead of constructors) with each NAL unit of the sample.

[0136] In an alternative embodiment, for video formats that do not use an extractor with constructors but only use a simple extractor, a new extractor can be defined, where a reserved NAL unit type is used to distinguish the new extractor that references the default extractor defined in the sample entry or sample group description box from the existing byte-based extractor. The new extractor, called for example "ReferenceExtractor", is defined as follows:

[0137]

[0138] wherein:

[0139] ref_index specifies the index of the extractor to be used in the list of extractors of the DefaultHevcExtractorBox in the sample entry of the track from which data is to be extracted. A value of 0 indicates the first entry.

[0140] The main difference is that here we have a specific "NALUnitHeader". The "NALUnitHeader" is the NAL unit header corresponding to the video coding format in use, but has reserved values that have not been reserved for any VCL, non-VCL NAL units or existing extractors or aggregators of the video coding format in use.

[0141] In this embodiment, the default extractor is defined in the sample entry in the DefaultHevcExtractorBox as follows:

[0142]

[0143] wherein:

[0144] num_entries gives the number of default extractors defined in this box.

[0145] nalUnitLength gives the length of the nalUnitExtractor in bytes.

[0146] nalUnitExtractor is the default extractor defined for this entry. The nal_unit_type of the nalUnitExtractor should be a valid extractor and the ReferenceExtractor should not exist.

[0147] According to another aspect of the present invention, the extractor was initially designed in AVC / SVC / MVC to extract a complete set (one or more than one) of NAL units, and each data reference (extracted byte range) starts with a NAL unit length field using 1 to 4 bytes to provide the length of the subsequent NAL unit (depending on the value of the LengthSizeMinusOne attribute in the associated decoder configuration record). By introducing an inline constructor (e.g., to rewrite the slice header) in combination with hvt1 or hvt2 sample data in HEVC, a constructor (e.g., SampleConstructor or SampleConstructorFromTrackGroup) can be used to extract sample data starting from an offset other than the NAL unit length field, and the inline constructor is used to rewrite the start of the NAL unit. In fact, for example, the hvt2 sample data only contains a subset of NAL units, usually only contains slice data and slice trailing parts without a NAL unit header and without a slice header. In this case, it is not clear where and how to define the NAL unit length after a series of sample data is extracted, especially when the extractor (composed of multiple constructors) is decomposed into multiple NAL units. Additionally, when dealing with sample constructors, it is not easy for the reader to know whether the data offset in the extracted samples corresponds to the first byte of the NAL unit length field; all previous NAL units in the extracted samples must be parsed to figure it out. It is also not clear whether the inline constructor should include the NAL unit length field, but if it does, this may require rewriting to match the bytes extracted from subsequent constructors. We can observe that defining the NALU length field in the inline constructor may result in wasting up to 4 bytes in the inline constructor, especially when rewriting is required.

[0148] To optimize both the size and complexity of the reader, a new explicit NAL start inline constructor is defined in a way that will avoid the file reader maintaining the state of "these extracted bytes require NALU length field rewriting or not / where to rewrite the field".

[0149] An extractor with an inline constructor that explicitly starts a NALU but does not embed a NALU length field is defined as follows:

[0150]

[0151] Provide the name of the new constructor "NALUStartInlineConstructor" as an example. Additionally, provide the reserved "constructor_type" value "5" as an example.

[0152] NALUStartInlineConstructor is used to indicate that the NAL unit starts at this constructor and extends to the immediately following constructor (if any). inline_data contains the start of the NAL unit (which may contain a NAL unit header, part of a NAL unit header, a NAL unit payload, or part of a NAL unit payload), but does not contain any NAL unit length field. This field should be inserted by the file reader based on the track LengthSizeMinusOne field and set to the full NAL unit size after processing this constructor and the immediately following constructor (if any).

[0153] NALUStartInlineConstructor is defined as follows:

[0154]

[0155] in:

[0156] length: is the number of bytes that belong to the NALUStartInlineConstructor that follows this field. The value of length should be greater than 0. A length value equal to 0 is reserved.

[0157] inline_data: are the data bytes to be returned when disassembling an inline constructor. If this is the last constructor in the extractor, these bytes should contain exactly one complete NAL unit, or the start of a NAL unit.

[0158] In addition, the existing InlineConstructor (for constructor_type=2) is modified as follows:

[0159]

[0160] in:

[0161] length: is the number of bytes that belong to the InlineConstructor that follows this field. The value of length should be greater than 0. A length value equal to 0 is reserved.

[0162] inline_data: are the data bytes to be returned when the inline constructor is disassembled. There should not be any bytes in these bytes that correspond to the NAL unit length field of the extractor's reconstructed NAL units.

[0163] Figure 7Schematic block diagram of a computing device 700 for implementing one or more embodiments of the present invention. The computing device 700 may be a device such as a microcomputer, a workstation, or a lightweight portable device. The computing device 700 includes a communication bus that is connected to:

[0164] - A central processing unit (CPU) 701, such as a microprocessor, etc.;

[0165] - A random access memory (RAM) 702 for storing executable code of the method of the embodiments of the present invention and registers that are adapted to record variables and parameters required to implement the method for reading and writing manifests and / or for encoding video and / or for reading or generating data in a given file format, and whose memory capacity can be expanded by an optional RAM connected to an expansion port, for example;

[0166] - A read-only memory (ROM) 703 for storing computer programs for implementing the embodiments of the present invention;

[0167] - A network interface 704, which is in turn generally connected to a communication network through which digital data to be processed is sent or received. The network interface 704 may be a single network interface or consist of a collection of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Under the control of a software application running in the CPU 701, data is written to the network interface for transmission or read from the network interface for reception;

[0168] - A user interface (UI) 705 for receiving input from a user or displaying information to a user;

[0169] - A hard disk (HD) 706;

[0170] - An I / O module 707 for receiving data from an external device such as a video source or a display and sending data to an external device such as a video source or a display.

[0171] The executable code may be stored in the read-only memory 703, on the hard disk 706, or on a removable digital medium (such as a disk). According to a variant, the executable code of the program may be received via the network interface 704 by means of a communication network for storage in one of the storage components of the communication device 700 (such as the hard disk 706, etc.) before being executed.

[0172] The central processing unit 701 is adapted to control and direct the execution of instructions or a portion of software code of one or more programs in accordance with an embodiment of the present invention, which instructions are stored in one of the aforementioned storage components. After power-on, the CPU 701 is capable of executing such instructions after, for example, loading instructions related to a software application from the program ROM 703 or the hard disk (HD) 706 into the main RAM memory 702. When executed by the CPU 701, such software application causes the steps of the flowchart shown in the previous figures to be performed.

[0173] In this embodiment, the device is a programmable device that uses software to implement the present invention. However, alternatively, the present invention may be implemented in hardware (e.g., in the form of an application specific integrated circuit or ASIC).

[0174] Although the present invention has been described above with reference to specific embodiments, the present invention is not limited to these specific embodiments, and modifications within the scope of the present invention will be apparent to those skilled in the art.

[0175] For example, the present invention may be embedded in devices such as cameras, smart phones, head-mounted displays, or tablet computers, and used as a remote control for a TV or multimedia display, e.g., to magnify a specific region of interest. It may also be used from the same device to have a personalized browsing experience of a multimedia presentation by selecting a specific region of interest. Another use by users of these devices and methods is to share some selected sub-portions of the user's preferred videos with other connected devices. If a surveillance camera supports the method for providing data according to the present invention, it may also be used with a smart phone or tablet computer to monitor what is happening in a specific area of a building being monitored.

[0176] In referring to the foregoing illustrative embodiments, those skilled in the art will envision many further modifications and variations. The foregoing illustrative embodiments are given by way of example only and are not intended to limit the scope of the present invention, which is determined only by the appended claims. In particular, where appropriate, different features from different embodiments may be interchanged.

Claims

1. A method for encapsulating media data into a media file, the method comprises: generating a first track including media samples; generating a second track including a metadata part and a data part, the data part including samples, the samples including extractors, the extractors being structures that reference data in the media samples included in the first track, the extractors including at least one constructor and at least one constructor type, the at least one constructor containing attributes for characterizing the data to be extracted, each constructor being associated with a constructor type; and generating a media file including the first track and the second track, wherein, sample entries in the metadata part of the second track include a list of default constructors, at least one default constructor containing attributes for characterizing the data to be extracted, the sample entries being data structures, wherein, at least one constructor included in the samples of the second track and associated with a predetermined constructor type is a reference constructor that references an index of a default constructor in the list of default constructors.

2. The method according to claim 1, wherein, the list of default constructors is included in a sample group entry, the sample group entry being a data structure describing a sample group of the second track, the extractors being included in the samples of the sample group.

3. The method according to claim 1, wherein, the extractor includes at least two constructors, each of the at least two constructors being a reference constructor that references the same index corresponding to the same default constructor in the list.

4. The method according to claim 1, wherein, the extractor includes at least two constructors, among the at least two constructors, a first constructor is a reference constructor that references a first index corresponding to a first default constructor in the list, and a second constructor is a reference constructor that references a second index corresponding to a second default constructor in the list.

5. The method according to claim 1, wherein, the constructor type associated with each constructor is one of the following constructor types: a sample constructor type that extracts NAL unit data from the samples of the first track; an inline constructor type that includes NAL unit data; and the predetermined constructor type.

6. A method for encapsulating media data into a media file, the method comprises: generating a first track including media samples, each media sample containing a set of one or more NAL units; generating a second track including extractors, the extractors being structures for extracting data in the media samples included in the first track, the extractors including at least one inline constructor and at least one constructor type, the at least one inline constructor including NAL unit data to be returned when processing the constructor, each constructor being associated with a constructor type; and generating a media file including the first track and the second track, wherein, the constructor type associated with the at least one inline constructor of the extractor of the second track indicates that the NAL unit starts at this constructor and the NAL unit includes NAL unit data starting from the NAL unit header without including any NAL unit length field, and indicates that the at least one inline constructor needs to insert the NAL unit length field after processing the at least one inline constructor.

7. A method for parsing media data in a media file, the method comprises: obtaining a first track including media samples in the media file; obtaining a second track in the media file, the second track including a metadata section and a data section, the data section including samples, the samples including extractors, the extractors being structures that reference data in the media samples included in the first track, the extractors including at least one constructor and at least one constructor type, the at least one constructor containing attributes characterizing the data to be extracted, each constructor being associated with a constructor type; wherein, the method further comprises: obtaining a list of default constructors in a sample entry of the metadata section of the second track, at least one default constructor containing attributes characterizing the data to be extracted, the sample entry being a data structure; wherein, at least one constructor included in the samples of the second track and associated with a predetermined constructor type by the extractor is a reference constructor that references an index of a default constructor in the list of default constructors; and wherein, the method further comprises: obtaining the data referenced by the extractor from the first track based on the default constructor referenced by the index of at least one reference constructor.

8. The method according to claim 7, wherein, the extractor includes at least two constructors, and each constructor of the at least two constructors is a reference constructor that references the same index corresponding to the same default constructor in the list.

9. The method according to claim 7, wherein, the extractor includes at least two constructors, among the at least two constructors, the first constructor is a reference constructor that references a first index corresponding to a first default constructor in the list, and the second constructor is a reference constructor that references a second index corresponding to a second default constructor in the list.

10. The method according to claim 7, wherein, the constructor type associated with each constructor is one of the following constructor types: a sample constructor type, which extracts NAL unit data from the samples of the first track; an inline constructor type, which includes NAL unit data; and the predetermined constructor type.

11. A method for parsing media data in a media file, the method comprises: obtaining a first track including media samples in the media file, each media sample containing a set of one or more NAL units; Obtain a second track including an extractor in the media file, where the extractor is a structure for extracting data from media samples included in the first track, the extractor includes at least one inline constructor and at least one constructor type, the at least one inline constructor includes NAL unit data to be returned when processing the constructor, and each constructor is associated with a constructor type. Wherein, the constructor type associated with the at least one inline constructor of the extractor of the second track indicates that the NAL unit starts at this constructor and the NAL unit includes NAL unit data starting from the NAL unit header without including any NAL unit length field, and indicates that the at least one inline constructor needs to insert a NAL unit length field after processing the at least one inline constructor.

12. A non-transitory computer-readable storage medium storing instructions of a computer program for implementing the method according to any one of claims 1 to 11.

13. An apparatus for encapsulating media data into a media file, the apparatus includes a processor configured to: Generate a first track including media samples; Generate a second track including a metadata portion and a data portion, the data portion includes samples, the samples include an extractor, the extractor is a structure that references data from media samples included in the first track, the extractor includes at least one constructor and at least one constructor type, the at least one constructor contains attributes characterizing the data to be extracted, and each constructor is associated with a constructor type; And Generate a media file including the first track and the second track, Wherein, the sample entry of the metadata portion of the second track includes a list of default constructors, at least one default constructor contains attributes characterizing the data to be extracted, and the sample entry is a data structure. Wherein, at least one constructor included in the samples of the second track by the extractor and associated with a predetermined constructor type is a reference constructor that references an index of a default constructor in the list of default constructors.

14. The apparatus according to claim 13, Wherein, The constructor type associated with each constructor is one of the following constructor types: A sample constructor type that extracts NAL unit data from samples of the first track; An inline constructor type that includes NAL unit data; and The predetermined constructor type.

15. An apparatus for encapsulating media data into a media file, the apparatus includes a processor configured to: Generate a first track including media samples, each media sample containing a set of one or more NAL units; Generate a second track including an extractor, the extractor is a structure for extracting data from media samples included in the first track, the extractor includes at least one inline constructor and at least one constructor type, the at least one inline constructor includes NAL unit data to be returned when processing the constructor, and each constructor is associated with a constructor type; and Generate a media file including the first track and the second track, wherein, the constructor type associated with the at least one inline constructor of the extractor of the second track indicates that the NAL unit starts at this constructor and the NAL unit includes NAL unit data starting from the NAL unit header without including any NAL unit length field, and indicates that the at least one inline constructor needs to insert the NAL unit length field after processing the at least one inline constructor.

16. An apparatus for parsing media data in a media file, the apparatus includes a processor configured to: Obtain a first track including media samples in the media file; Obtain a second track in the media file, the second track includes a metadata part and a data part, the data part includes samples, the samples include extractors, the extractors are structures that reference data in the media samples included in the first track, the extractors include at least one constructor and at least one constructor type, the at least one constructor contains attributes for characterizing the data to be extracted, and each constructor is associated with a constructor type, wherein, the processor is further configured to: Obtain a list of default constructors in the sample entry of the metadata part of the second track, at least one default constructor contains attributes for characterizing the data to be extracted, and the sample entry is a data structure, wherein, at least one constructor associated with a predetermined constructor type included in the samples of the second track by the extractor is a reference constructor that references the index of the default constructor in the list of default constructors, and wherein, the processor is configured to: Obtain the data referenced by the extractor from the first track based on the default constructor referenced by the index of at least one reference constructor.

17. The apparatus according to claim 16, wherein, the constructor type associated with each constructor is one of the following constructor types: Sample constructor type, which extracts NAL unit data from the samples of the first track; Inline constructor type, which includes NAL unit data; and The predetermined constructor type.

18. An apparatus for parsing media data in a media file, the apparatus includes a processor configured to: Obtain a first track including media samples in the media file, each media sample contains a set of one or more NAL units; Obtain a second track including an extractor in the media file, the extractor is a structure for extracting data in the media samples included in the first track, the extractor includes at least one inline constructor and at least one constructor type, the at least one inline constructor includes NAL unit data to be returned when processing this constructor, and each constructor is associated with a constructor type, wherein, The constructor type associated with the at least one in-line constructor of the extractor of the second track indicates that the NAL unit starts at this constructor and that the NAL unit includes NAL unit data starting from the NAL unit header and does not include any NAL unit length fields, and indicates that the at least one in-line constructor requires insertion of a NAL unit length field after processing the at least one in-line constructor.

19. A computer program product comprising instructions for a computer program for implementing the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Media extractor tracks for file format track selection

    US20110064146A1

  • Scene section and region of interest handling in video streaming

    US20190174161A1