Media file encapsulating an image, methods for encapsulating and de-encapsulating an image, devices for encapsulating and de-encapsulating an image
The method and device for encapsulating and de-encapsulating images in media files address inefficiencies in HEIF by using a compact description mode to reduce unnecessary decoder configuration parameters, enabling efficient parsing and processing of small images while maintaining compatibility with legacy systems.
Patent Information
- Application Number
- PCT/EP2025/058039
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-02
Smart Images

Figure EP2025058039_02102025_PF_FP_ABST
Abstract
Description
[0001] MEDIA FILE ENCAPSULATING AN IMAGE, METHODS FOR ENCAPSULATING AND DE-ENCAPSULATING AN IMAGE, DEVICES FOR ENCAPSULATING AND DE-ENCAPSULATING AN IMAGE TECHNICAL FIELD OF THE INVENTION The present invention relates to a media file encapsulating an image, to a method for encapsulating an image, to a method for de-encapsulating an image and a corresponding computer program product, computer-readable storage medium and device. The present invention is applicable to the compact description of codec or decoder configurations for small images in the High Efficiency Image Format (HEIF) standard. BACKGROUND OF THE INVENTION With the advent of increasing transmission capabilities across networks and of increasing computing power, there has been a gradual increase in the size of media transmitted and processed by computing devices. Along with the size of the media content as such, the size of metadata describing said media content has also increased. In current systems, there is a need to transmit and process comparatively small images, for which the current media content to metadata size ratio is unsatisfactory, given the static nature of the metadata structure used by image transmission and processing standards, such as the High Efficiency Image Format standard. A small image may be defined by the resolution of said image (such as 1024x1024 or equivalent and below) or by the memory space said image occupies (such as 16 kilobytes and below, for example). A small image may further be defined in a given use-case in which specific operational parameters, in terms of image sizes, are to be used. Such a need is made complex by the fact that removing metadata may generate parsing errors on the side of the media receiving and processing units. Such metadata correspond, for example, to decoder (codec) signaling information. Transitorily, there exists a need today to generate a compact description for small images in HEIF, while ensuring a simple parsing and extraction of the item described with the compact description. Furthermore, there exists a need today to allow generating a compact description of codec or decoder configuration for small images in HEIF while maintaining the same decoders initialization capabilities as with a legacy decoder configuration in a media file parser. The Image File Format standard (HEIF – ISO / IEC 23008-12) describes the encapsulation of images as image items in an ISOBMFF-based media file. The MPEG File Format group is considering a more compact version of HEIF for use cases with small images. SUMMARY OF THE INVENTION The present invention is intended to remedy all or part of these disadvantages. To this effect, according to a first aspect, the present invention aims at a method for encapsulating an image in a media file, comprising the step of: - generating, in a media file, an image description and / or image header comprising a set of decoder configuration parameters based on an encapsulation mode among one of at least a regular description mode or a compact description mode, wherein, with the compact description mode, less parameters and / or a smaller encoding of a parameter value for the set of decoder configuration parameters are used in comparison to the regular description mode. Such provisions allow for, during encapsulation, providing a limited number of metadata parameters and / or smaller-size metadata parameters in an image header and / or image description to represent the image content and, during de-encapsulation, smart parsing of the metadata parameters in an image header and / or image description. In particular embodiments, if the encapsulation mode indicates a compact description mode, only decoder configuration parameters not associated with a reserved or predetermined value in the regular description mode are used to generate the image description and / or image header during the step of generating. Such provisions allow for the removal of decoder configuration parameters associated with reserved or predetermined values. In particular embodiments, if the encapsulation mode indicates a compact description mode, only decoder configuration parameters not associated with multi- image content are used to generate the image description and / or image header during the step of generating Such provisions allow for the removal of decoder configuration parameters associated with multi-image (such as video) media. Such provisions allow for, during encapsulation, providing a limited number of metadata parameters and / or smaller-size decoder configuration parameters. Such decoder configuration parameters, for example in the case of the HEVC, VVC and AVC codecs, provided in property boxes containing the codec or decoder configurations may represent a large amount of the HEIF signaling. In particular embodiments, the image encapsulation format corresponds to an image encapsulation format according to the High Efficiency Image Format (HEIF) standard. According to a second aspect, the present invention aims at a method for encapsulating at least one image, comprising the steps of: - selecting an encapsulation mode between a regular description mode and a compact description mode for an image, and - generating an image description and / or image header, depending on the encapsulation mode selected, wherein, according to the compact description mode, less parameters and / or a smaller encoding of a parameter value are used in comparison to the regular description mode. Such provisions allow for the generation of smaller-size image description and / or image header in an encapsulated image. In particular embodiments, if the compact description mode is selected, a compact set of decoder configuration parameters is generated during the step of generating. Such provisions allow for the generation of smaller-size decoder configuration parameters in an image description and / or image header in an encapsulated image. In particular embodiments, if the encapsulation mode indicates a compact description mode, compressing at least one decoder configuration parameter if the length of said at least one decoder configuration parameter is above a length threshold value, said at least one compressed decoder configuration parameter being used to generate the image description and / or image header during the step of generating. Such provisions allow for the compression in size of decoder configuration parameters. In particular embodiments, if the encapsulation mode indicates a compact description mode, each network abstraction layer unit parameter associated with a likelihood of use estimation above a threshold value is used to generate the image description and / or image header during the step of generating. Such provisions allow for the removal of decoder configuration parameters which are unlikely to be used by a decoder. In particular embodiments, if the encapsulation mode indicates a compact description mode, each configuration parameter determined to be nondeductible is used to generate an image description and / or image header during the step of generating. Such provisions allow for the removal of decoder configuration parameters which can be deducted by a decoder. In particular embodiments, at least one decoder configuration parameter is a network abstraction layer unit type parameter present in a configuration record of the image. In particular embodiments, at least one decoder configuration parameter is a number and / or length of network abstraction layer units present in a configuration record of the image. In particular embodiments, the encapsulation mode further indicates one of a plurality of compact description modes representative of an intended degree of compaction, the step of generating being performed as a function of said degree of compaction. According to a third aspect, the present invention aims at a method for de- encapsulating an image from a media file, comprising the steps of: - parsing an image description and / or image header comprising a set of decoder configuration parameters based on an encapsulation mode among one of at least a regular description mode or a compact description mode, wherein, with the compact description mode, less parameters and / or a smaller encoding of a parameter value for the set of decoder configuration parameters are parsed in comparison to the regular description mode. These provisions present similar advantages to the encapsulation method object of the second aspect of the present invention. According to a fourth aspect, the present invention aims at a device for encapsulating an image in a media file, comprising: - means for generating, in a media file, an image description and / or image header comprising a set of decoder configuration parameters based on, an encapsulation mode among one of at least a regular description mode or a compact description mode, wherein with the compact description mode, less parameters and / or a smaller encoding of a parameter value for the set of decoder configuration parameters are used in comparison to the regular description mode. According to a fifth aspect, the present invention aims at a device for de- encapsulating an image from a media file, comprising: - means of parsing an image description and / or image header comprising a set of decoder configuration parameters based on, an encapsulation mode among one of at least a regular description mode or a compact description mode, wherein, with the compact description mode, less parameters and / or a smaller encoding of a parameter value for the set of decoder configuration parameters are parsed in comparison to the regular description mode. According to a sixth aspect, the present invention aims at a media file encapsulating an image comprising a parameter representative of an encapsulation mode among one of at least a regular description mode or a compact description mode, wherein, the compact description mode is used for generating an image description and / or image header comprising a set of decoder configuration parameters with less parameters and / or a smaller encoding of a parameter values for the set of decoder configuration parameters than a regular description mode. In particular embodiments, the media file is compliant with an image encapsulation format according to the High Efficiency Image Format (HEIF) standard. According to a seventh aspect, computer program product, characterized in that it comprises instructions which upon execution by a computer cause the computer to execute a method object of the present invention. According to an eighth aspect, the present invention aims at a computer- readable storage medium storing programming instructions which upon execution by a computer cause the computer to execute a method object of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS Other advantages, purposes and particular characteristics of the invention shall be apparent from the following non-exhaustive description of at least one particular embodiment of the present invention, in relation to the drawings annexed hereto, in which: [Figure 1] represents, schematically, a succession of steps of a first particular embodiment of the method for encapsulating object of the present invention, [Figure 2a] represents, schematically, a succession of steps of a second particular embodiment of the method for encapsulating object of the present invention, [Figure 2b] represents, schematically, a succession of steps of a third particular embodiment of the method for encapsulating object of the present invention, [Figure 2c] represents, schematically, a succession of steps of a fourth particular embodiment of the method for encapsulating object of the present invention, [Figure 3] represents, schematically, a succession of steps of a fifth particular embodiment of the method for encapsulating object of the present invention, [Figure 4] represents, schematically, a succession of steps of a particular embodiment of the method for de-encapsulating object of the present invention, and [Figure 5] represents, schematically, a particular embodiment of a device suitable for encapsulating and de-encapsulating an image file. DETAILED DESCRIPTION OF THE INVENTION This description is not exhaustive, as each feature of one embodiment may be combined with any other feature of any other embodiment in an advantageous manner. Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments. The indefinite articles ‘a’ and ‘an’, as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean ‘at least one’. The phrase ‘and / or’, as used herein in the specification and in the claims, should be understood to mean ‘either or both’ of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with ‘and / or’ should be construed in the same fashion, i.e. ‘one or more’ of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the ‘and / or’ clause whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to ‘A and / or B’, when used in conjunction with open-ended language such as ‘comprising’ can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc. As used herein in the specification and in the claims, ‘or’ should be understood to have the same meaning as ‘and / or’ as defined above. For example, when separating items in a list, ‘or’ or ‘and / or’ shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as ‘only one of’ or ‘exactly one of’, or, when used in the claims, ‘consisting of’, will refer to the inclusion of exactly one element of a number or list of elements. In general, the term ‘or’ as used herein shall only be interpreted as indicating exclusive alternatives (i.e. ‘one or the other but not both’) when preceded by terms of exclusivity, such as ‘either,’ ‘one of,’ ‘only one of’, or ‘exactly one of’. ‘Consisting essentially of,’ when used in the claims, shall have its ordinary meaning as used in the field of patent law. As used herein in the specification and in the claims, the phrase ‘at least one’, in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase ‘at least one’ refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, ‘at least one of A and B’ (or, equivalently, ‘at least one of A or B’, or, equivalently ‘at least one of A and / or B’) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc. In the claims, as well as in the specification above, all transitional phrases such as ‘comprising,’ ‘including,’ ‘carrying,’ ‘having,’ ‘containing,’ ‘involving,’ ‘holding,’ ‘composed of’, and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases ‘consisting of’ and ‘consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively. It should be noted at this point that the figures are not to scale. In the context of the present invention, “ISOBMFF” refers to ISO base media file format, where “ISO” refers to International Organization for Standardization. ISOBMFF is a container file format that defines a general structure for files that contain time- based multimedia data such as video and audio. According to ISOBMFF and its extensions, a media file is constituted by “boxes”. Boxes, also called containers, are hierarchical data structures provided to describe the data in the files. Boxes are object-oriented building blocks defined by a unique type identifier, typically a four-character code, also noted “FourCC” or “4CC”, and a length. All data in a file, media data and metadata describing the media data, is contained in boxes. There is no other data within the file. File-level boxes are boxes that are not contained in other boxes. In the context of the present invention, “HEIF” refers to High Efficiency Image File Format. HEIF is a container format for storing individual digital images and image sequences. The standard covers multimedia files that can also include other media streams, such as timed text, audio and video. According to particular embodiments, the present invention aims at a media file encapsulating an image comprising a parameter representative of an encapsulation mode selected from at least a regular description mode and a compact description mode, wherein, the compact description mode is used for generating an image description and / or image header comprising a set of decoder configuration parameters with less parameters and / or a smaller encoding of a parameter values for the set of decoder configuration parameters than a regular description mode. Such a parameter may correspond to an explicit parameter which can be associated with two values (“compaction mode active” or “compaction mode inactive”). In variants, this explicit parameter may be associated with more than two values, corresponding to different degrees or levels of compaction. In variants, this explicit parameter may be signaled as a brand in the media file or as a four-character code representing the type of a box in the media file, the type of the box being representative on whether the box comprises a regular description or a compact description. Such a parameter may correspond to the fact that, in an image file format expecting a predetermined number of image description and / or image header parameters, less parameters are present in the image description and / or image header parameters of the image file. Such a parameter may correspond to the inclusion of brand in the image file parameters, said brand being representative of the presence of a compact format for the image file or indicating that the compact description mode is active, and another brand being representative of the presence of a regular format for the image file or indicating that the regular description mode is active. A possible brand definition for the compact description of an image may consist in using the ‘mif3’ brand to indicate one of the following file requirements: presence of MinimizedImageBox box for example with ‘mini’ Box type as top-level box of the media file; or the presence of a MetaBox ‘meta’ box and its sub-boxes as for media file compatible with the ‘mif1’ brand (for backward compatibility). In particular embodiments, the image encapsulation format comprises a compact set of decoder configuration parameters. In particular embodiments, the image encapsulation format corresponds to an image encapsulation format according to the High Efficiency Image Format (HEIF) standard. Figure 1 illustrates, schematically, a particular embodiment of the method for encapsulating object of the present invention in its main processing steps. It should be understood that this encapsulation method may be executed according to an iterative process that applies successively to each image to be represented in an HEIF format file. The processing steps of the figure 1 represent the processing of one image. However, in some cases more than one image may be represented in the same media file. Some images may also be derived from previously represented images and thus may introduce dependencies between image items. In another example, the input of the encapsulation may consist in a sequence of one or more images to represent an animation. In that case, the encapsulation process described in figure 1 may be applied successively to each input image of the sequence. The encapsulation method may thus keep in memory the previously generated media file and update this previously generated media file instead of generated a new media file such as disclosed in relation to the steps 103 to 106. In a first step 101, the method for encapsulating determines an image configuration. Such a step 101 may be performed by a computer program executed upon a computing device. Such a computer program may be configured to read and interpret metadata stored in the image file. This image configuration corresponds to a set of parameters or properties associated with the image to be encapsulated. For example, this image configuration may comprise parameters indicating the codec associated with a profile, a tier and / or a level data information that is used to compress the image to be represented in HEIF file. The configuration parameters may also include a set of coded structures describing coding configuration that allows instantiating a decoder. Such a set of coded structures is also denoted in this document as configuration data or decoder configuration parameters. Other parameters representing the configuration of the image comprise for example the dimension or resolution of the image. The configuration parameters may also comprise the number of components present in the image. A component is defined as a set of data having the same semantics that is associated to the image pixels. For example, the pixel values of an image form one image component. The alpha plane or alpha channel can correspond to another component if such channels are present in the coded image. Some images may be associated with metadata that can be described with EXIF (for “Exchangeable image file format”) or XMP (for “Extensible Metadata Platform”) metadata format. In this document, each set of metadata is considered as one component. It is to be noted that in monochrome format, the pixels values data may consist in a single channel, but such a channel can still be considered as a component. In color formats, each pixel is associated with several pixel values, for example a red value, a green value and a blue value. Each color forms a channel, and all the color channels form a component. The configuration parameters of the image may also comprise the length in bits or bytes of the configuration data for each component and / or the length in bits or bytes of the compressed data for each component. The method for encapsulating further comprises a check step 102 which confirms depending on an encapsulation mode whether a compact (or slim or condensed or compacted) file format representation or a regular file format representation (i.e. without compaction) is to be used to describe the media file. Such a step 102 may be performed by a computer program executed upon a computing device. Such a computer program may be configured to read and interpret configuration parameters stored in the image file. Such a step 102 may execute a deterministic rule to determine the encapsulation mode, based for example on image size parameters which are predetermined or determined as an execution parameter of said step 102. If such a rule is determined as an execution parameter, such a rule may be set by a user on a graphic user interface (GUI) or automatically by another computer program. In other embodiments, this step 102 is performed on a computing system regardless of the configuration parameters of the image and the encapsulation mode is determined following algorithmic rules which belong to said computing system. A computing system might require an image to be encapsulated in a compact manner for reasons of speed of transmission or processing. The compact description mode and the corresponding compact description disclosed in the present document aims at providing a minimal description of the image that allows a simplified parsing of the HEIF description and parallel processing of the image components. The benefit of the compact description is also to have a compact description for compressed images that are represented with a small amount of data. The compactness of the compressed data may come either from an efficient compression (with possible loss of quality due to the compression) or from the small dimensions of the images resulting in a low definition. The criteria that would trigger such a compact representation are manifold. Such criteria may rely for example on the necessity for an application to render rapidly low- definition images for the generation of a graphical user interface or for rendering interactive entity in a webpage such as clickable images. Other criteria can be related to the image configuration like the dimensions, the number of components or the length in bits or bytes of the coded data, which may also prevent the image to be described with a compact description. When the image does not meet the criteria as checked in step 102, the encapsulation process uses a regular or standard encapsulation mode and the corresponding regular encapsulation (such as HEIF) by generating a media file header in step 103 and represents the image as an HEIF item using a MetaBox with version equal to 0 to describe the image data. The generation of the MetaBox information is illustrated by the step 104 in figure 1. The steps of generating 103 a media file header and generating 104 a meta information box are known to persons skilled in the art of image encapsulation and are not disclosed further here. On the contrary, if the compact description is selected in step 102, the steps 105 to 106 apply, starting with the generation of a media file header in step 105. The step 105 may consist of the generation of an HEIF media file header that starts by a Box, for example a FileTypeBox, which indicates a set of brands that applies to the media file. This box may indicate a set of one major brand, one or more minor brands or compatible brands. When compact description is used, a specific brand may be indicated in the list of brands to indicate the presence of boxes specific to the compact description. For example, these boxes may be generated in the step 106. In compact description, an optimized version of the FileTypeBox may be used. In step 106, the encapsulation processing generates the compact (or condensed or slim or compacted) image description for the image. The processing employed to generate the compact image description is described with reference to figure 3. Figure 3 illustrates a particular embodiment of a step of generating a compact encapsulation of a media or image file, which may be implemented in an embodiment of the method of encapsulating object of the present invention. In a first step 300, a compact (or minimized or condensed or compacted) image header is generated. This compact image header comprises a reduced number of parameters describing the image configuration. This reduced number of parameters specifies the properties of the image. The compact image header may comprise a version number indicating a value that makes it possible to distinguish different versions of the compact description syntax. The compact image header may comprise a parameter indicating the type of the color profile configuration (e.g., sRGB, ICC profiles, etc.). The compact image header may comprise the width and the height of the image described by the compact description. The compact image header may comprise the number of bits used to represent the pixel values of the Image described by the compact description. The compact image header may comprise an indication whether the pixel format uses float or integer representation. The compact image header may comprise an indication of the range of the pixel values. The compact image header may comprise an indication whether the Image described by the compact description is associated with an alpha plane. The compact image header may comprise an indication if the pixels values of the image described by the compact description are pre-multiplied by the alpha plane. The compact image header may comprise an indication if the image described by the compact description data includes an explicit list of codec types. In a step 301, a data structure is generated comprising the compressed data corresponding to the pixel values component, namely the pixel values of the image. This data structure may also comprise configuration data or decoder configuration parameters. The configuration data may comprise a decoder (or codec) configuration box or a decoder configuration record that may be provided in SampleEntry of an ISOBMFF file or as an item property associated to one or more items. For example, for an HEVC (for “High Efficiency Video Coding”) item, the configuration data may correspond to the content of the HEVCConfigurationBox (‘hvcC’) as per ISO / IEC 14496-15 or to a compact version of the content of this box according to an embodiment of this disclosure. Another example is VVCConfigurationBox (‘vvcC’) or AVCConfigurationBox (‘avcC’) as defined in the same specification for a VVC (for “Versatile Video Coding”) or an AVC (for “Advanced Video Coding”) Item, respectively, or to a compact version of the content of these boxes according to an embodiment of this disclosure. Yet another example is any decoder configuration of an MPEG (for “Moving Picture Experts Group”) codecs that can be used to compress still images. The generation of the codec or decoder configuration in step 302 is further detailed with reference to the figure 2a to 2c. In a step 303, when the image comprises an alpha plane, another data structure is generated comprising the auxiliary component such as alpha channel or depth data. The data structure is the same as for the pixel components of the image described using compact description. It is to be noted that several configuration parameters may be provided in the compact image description, for example one for the main image and another one for the alpha image. In that case, a step similar to 302 may be used also to generate the decoder configuration for the alpha plane. More generally, the generation of the decoder configuration in step 302 may be applied to and decoder configuration data described by the condensed description of the image. The pixel and auxiliary data can both be represented as coded items for which the data is provided in the data structure generated in steps 301 and 303. This data structure is denoted as Minimized Item data structure (also denoted Minimized Item Data data structure) and the items described in the compact description of the image may be denoted as Minimized Item. A minimized item is an item that is described in the compact description of an image. As a counter-example, an item described in a legacy HEIF MetaBox based on ItemInfoBox and ItemLocationBox is not a minimized item. In a step 304, when the image is associated with metadata, such as EXIF or XMP metadata for example, this metadata is represented as a coded item using the same minimized item data structure. In a step 305, information related to the color information such as the color primaries, the Transfer Characteristics and the Matrix coefficients are represented in a ColourData data structure. These data structures are not ISOBMFF boxes, they are all comprised in a single ISOBMFF box, for example a Minimized Image box. A Minimized Image may be defined as an Image described using compact description. As it is understood, decoder or codec configuration may be provided in the compact description. The content of the decoder configuration depends on the coding format of coded data in the image or more precisely on coded data of the item that is associated with the decoder configuration. For example, MPEG (for “Moving Picture Experts Group”) codecs can be used to compress still images. For example, VVC, HEVC, AVC provides decoder configurations that describe the coding parameters of the encoder. In ISO / IEC 14496- 15, the decoder configurations for these standards are made to support a wide range of videos and still images resulting in a verbose decoder configuration. In the context of a compact image description the size or length of the decoder configuration as specified in ISO / IEC 14496-15 represents a significant part of the description of small images. In this document, the encapsulation process may compress the size of the decoder configuration through different means to ensure compactness of the description. As an overview of a general solution, the encapsulation process may rely on the fact that a decoder configuration is applied to single images (not videos) when used in a compact description of an image. It can be observed that some parameters in legacy or regular decoder configuration (i.e. not compact decoder configuration) are only useful or even meaningful for video sequences. The encapsulation process may signal in the bitstream that these video-related parameters are not present in the decoder configuration. Based on this signaling the de-encapsulation process may parse the compact decoder configuration provided in the compact image description. In order to ensure same initialization capabilities as with legacy decoder configuration syntax, the parser may infer the parameters not provided in the compact decoder configuration to generate a legacy decoder configuration as per ISO / IEC 14496-15. In the following section several embodiments are disclosed which compress the decoder configuration focusing on examples for VVC, AVC and HEVC codec formats. The following embodiments may be applied to other decoder configuration formats that may be compressed using compression mechanisms equivalent or like the ones described for the different embodiments. In a first embodiment, the generation step 302 of the compact decoder configuration is represented in figure 2a. During this generation step, it is determined in a first step 200 the regular or legacy decoder configuration parameters as specified per ISO / IEC 14496-15. Then, the parameters of the regular or legacy decoder configuration are processed successively in the processing loop composed of the steps 201 to 205. During step 201, a check is performed to ensure that all parameters have been processed to end the processing loop by running the step 206 that is described below. For each parameter, during a step 202, the encapsulation process checks if the parameter or parts of the parameter are associated with a reserved or a predetermined value. For example, the legacy decoder configuration may include padding bits that have a value equal to 0. The parameters that meet this criterion are marked as “optional” in a step 203, because not relevant for a single image or an image sequence track (that has no strict timing indication). Such a step of determining if at least one decoder configuration parameter is associated with a reserved or predetermined value, may be performed by executing a computer program upon a computing device. During this step of determining, several configurations may be generated. In some configurations, as shown above, during the step of determining, the presence of padding bits may be detected and determined to correspond to a reserved size or value which can be marked as “optional”. During this step of determining, the presence of some parameters indicating the version of the decoder configuration syntax may be detected as well – most of the time such parameters are set to a predetermined value (for example equal to 0). Then, during step 204, it is determined if the parameter describes information inconsistent for single image content that are meaningful only for multiple images content. For example, it can be information related to video such as a frame rate or a temporal level. The parameters that meet this criterion are marked as “optional” in a step 203. Otherwise, the parameters that do not meet the previous criteria are marked as “essential” in the step 205. In a final step, the encapsulation module generates a compact decoder configuration that contains only the “essential” parameters; in other words, only the parameters that are relevant for still image. It is to be noted that the list of relevant parameters may depend on a profile or brand indication. For example, when the profile indicates a still picture profile, the de- encapsulation module can infer that there is only one image in the file and that the parameters of the decoder configuration describing video aspects are set to pre- determined value for example 0 or 1. As another example, the de-encapsulation may do the same if the image file contains brand contains as compatible brand, for example in ‘ftyp’ box (or a compact version of this box) the ‘1pic’ brand. In these cases, the encapsulation module may consider the parameters of the decoder configuration describing video aspects optional at step 203. In such a context, the de-encapsulation module infers the value of the optional parameters absent from the compact decoder configuration as equal to a predetermined value (for example 1 or 0 or unspecified) in order to build a decoder configuration as expected by legacy players. The de-encapsulation process in charge of re-generating the legacy decoder configuration needs to have knowledge of the inferred value. This is possible since the parser can be configured to generate the correct value which is standardized in ISO / IEC 14496-15, for example. For example, below there are two reserved parameters normatively defined equal to '11111'b and a third reserved set equal to 0. For example, the de-encapsulation process may regenerate optional (as marked in a step 203) constant_format_rate as value 0 as not applicable to image. For example, the syntax of the content of the decoder configuration can correspond to the following as per ISO / IEC 14496-15: aligned(8) class VvcDecoderConfigurationRecord { bit(5) reserved = '11111'b; unsigned int(2) LengthSizeMinusOne; unsigned int(1) ptl_present_flag; if (ptl_present_flag) { unsigned int(9) ols_idx; unsigned int(3) num_sublayers; unsigned int(2) constant_frame_rate; unsigned int(2) chroma_format_idc; unsigned int(3) bit_depth_minus8; bit(5) reserved = '11111'b; VvcPTLRecord(num_sublayers) native_ptl; unsigned_int(16) max_picture_width; unsigned_int(16) max_picture_height; unsigned int(16) avg_frame_rate; } unsigned int(8) num_of_arrays; for (j=0; j < num_of_arrays; j++) { unsigned int(1) array_completeness; bit(2) reserved = 0; unsigned int(5) NAL_unit_type; if (NAL_unit_type != DCI_NUT && NAL_unit_type != OPI_NUT) unsigned int(16) num_nalus; for (i=0; i< num_nalus; i++) { unsigned int(16) nal_unit_length; bit(8*nal_unit_length) nal_unit; } } } aligned(8) class VvcPTLRecord(num_sublayers) { bit(2) reserved = 0; unsigned int(6) num_bytes_constraint_info; unsigned int(7) general_profile_idc; unsigned int(1) general_tier_flag; unsigned int(8) general_level_idc; unsigned int(1) ptl_frame_only_constraint_flag; unsigned int(1) ptl_multilayer_enabled_flag; unsigned int(8*num_bytes_constraint_info - 2) general_constraint_info; for (i=num_sublayers - 2; i >= 0; i--) unsigned int(1) ptl_sublayer_level_present_flag[i]; for (j=num_sublayers; j<=8 && num_sublayers > 1; j++) bit(1) ptl_reserved_zero_bit = 0; for (i=num_sublayers-2; i >= 0; i--) if (ptl_sublayer_level_present_flag[i]) unsigned int(8) sublayer_level_idc[i]; unsigned int(8) ptl_num_sub_profiles; for (j=0; j < ptl_num_sub_profiles; j++) unsigned int(32) general_sub_profile_idc[j]; } In these records, the bits of the parameter “reserved” and ptl_reserved_zero_bit have a predetermined value equal to 1 (or 0). The encapsulation process thus marks these parameters as optional accordingly to the criterion checked in step 202. The parameters “num_sublayers”, “constant_frame_rate” and “avg_frame_rate” are related to video and are inconsistent for still image content. Thus, in the step 204 the encapsulation process determines them as optional. The compact decoder configuration contains the parameters defined as essential and not the optional parameter. It allows for obtaining a reduced or compact size of the decoder configuration and thus of the image description. An additional set of trailing bits (represented by the trailing_bits() function) may be provided at the end of the compact configuration record to ensure it is byte aligned. As a result, the encapsulation process may generate (in step 206) a compact VVC decoder configuration that contains the following parameters accordingly to first embodiment. This CompactVvcDecoderConfigurationRecord (or CondensedVvcDecoderConfigurationRecord) may be for example signaled in a CompactVvcDecoderConfigurationBox (or CondensedVvcDecoderConfigurationBox) with ‘vvcc’ as four-character code box type: aligned(8) class CompactVvcDecoderConfigurationBox extends Box('vvcc') { CompactVvcDecoderConfigurationRecord () vvcConfig; } aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) LengthSizeMinusOne; unsigned int(1) ptl_present_flag; if (ptl_present_flag) { unsigned int(9) ols_idx; unsigned int(2) chroma_format_idc; unsigned int(3) bit_depth_minus8; VvcPTLRecord(1) native_ptl; unsigned_int(16) max_picture_width; unsigned_int(16) max_picture_height; } unsigned int(8) num_of_arrays; for (j=0; j < num_of_arrays; j++) { unsigned int(1) array_completeness; unsigned int(5) NAL_unit_type; if (NAL_unit_type != DCI_NUT && NAL_unit_type != OPI_NUT) unsigned int(16) num_nalus; for (i=0; i< num_nalus; i++) { unsigned int(16) nal_unit_length; bit(8*nal_unit_length) nal_unit; } } trailing_bits() } aligned(8) class VvcPTLRecord(num_sublayers) { unsigned int(6) num_bytes_constraint_info; unsigned int(7) general_profile_idc; unsigned int(1) general_tier_flag; unsigned int(8) general_level_idc; unsigned int(1) ptl_frame_only_constraint_flag; unsigned int(1) ptl_multilayer_enabled_flag; unsigned int(8*num_bytes_constraint_info - 2) general_constraint_info; for (i=num_sublayers - 2; i >= 0; i--) unsigned int(1) ptl_sublayer_level_present_flag[i]; for (i=num_sublayers-2; i >= 0; i--) if (ptl_sublayer_level_present_flag[i]) unsigned int(8) sublayer_level_idc[i]; unsigned int(8) ptl_num_sub_profiles; for (j=0; j < ptl_num_sub_profiles; j++) unsigned int(32) general_sub_profile_idc[j]; } In another example, the syntax of the content of the HEVC decoder configuration is the following as per ISO / IEC 14496-15:aligned(8) class HEVCDecoderConfigurationRecord { unsigned int(8) configurationVersion = 1; unsigned int(2) general_profile_space; unsigned int(1) general_tier_flag; unsigned int(5) general_profile_idc; unsigned int(32) general_profile_compatibility_flags; unsigned int(48) general_constraint_indicator_flags; unsigned int(8) general_level_idc; bit(4) reserved = '1111'b; unsigned int(12) min_spatial_segmentation_idc; bit(6) reserved = '111111'b; unsigned int(2) parallelismType; bit(6) reserved = '111111'b; unsigned int(2) chroma_format_idc; bit(5) reserved = '11111'b; unsigned int(3) bit_depth_luma_minus8; bit(5) reserved = '11111'b; unsigned int(3) bit_depth_chroma_minus8; unsigned int(16) avgFrameRate; unsigned int(2) constantFrameRate; unsigned int(3) numTemporalLayers; unsigned int(1) temporalIdNested; unsigned int(2) lengthSizeMinusOne; unsigned int(8) numOfArrays; for (j=0; j < numOfArrays; j++) { unsigned int(1) array_completeness; bit(1) reserved = 0; unsigned int(6) NAL_unit_type; unsigned int(16) numNalus; for (i=0; i< numNalus; i++) { unsigned int(16) nalUnitLength; bit(8*nalUnitLength) nalUnit; } } } In these records, the bits of the parameters “reserved” and configurationVersion have a predetermined value equal to 1 (or 0). The encapsulation process thus marks these parameters as optional accordingly to the criterium checked in step 202. The parameters “avgFrameRate”, “constantFrameRate” and “numTemporalLayers” and temporalIdNested are related to video and are inconsistent for still image content. Thus, in the step 204 the encapsulation process determines them as optional. As a result, the encapsulation process may generate (in step 206) a compact HEVC decoder configuration that contains the following parameters accordingly to first embodiment. This CompactHEVCDecoderConfigurationRecord (or condensedHEVCDecoderConfigurationRecord) may be for example signaled in a CompactHEVCDecoderConfigurationBox (or CondensedHEVCDecoderConfigurationBox) with the ‘hvcc’ four-character code box type (presence of these four-character code indicating the compact description mode is active): aligned(8) class CompactHEVCDecoderConfigurationBox extends Box('hvcc') { CompactHEVCDecoderConfigurationRecord () hevcConfig; } aligned(8) class CompactHEVCDecoderConfigurationRecord { unsigned int(2) general_profile_space; unsigned int(1) general_tier_flag; unsigned int(5) general_profile_idc; unsigned int(32) general_profile_compatibility_flags; unsigned int(48) general_constraint_indicator_flags; unsigned int(8) general_level_idc; unsigned int(12) min_spatial_segmentation_idc; unsigned int(2) parallelismType; unsigned int(2) chroma_format_idc; unsigned int(3) bit_depth_luma_minus8; unsigned int(3) bit_depth_chroma_minus8; unsigned int(16) avgFrameRate; unsigned int(2) constantFrameRate; unsigned int(3) numTemporalLayers; unsigned int(1) temporalIdNested; unsigned int(2) lengthSizeMinusOne; unsigned int(8) numOfArrays; for (j=0; j < numOfArrays; j++) { unsigned int(1) array_completeness; unsigned int(6) NAL_unit_type; unsigned int(16) numNalus; for (i=0; i< numNalus; i++) { unsigned int(16) nalUnitLength; bit(8*nalUnitLength) nalUnit; } } trailing_bits() } In another example, the syntax of the content of the AVC decoder configuration is the following as per ISO / IEC 14496-15: aligned(8) class AVCDecoderConfigurationRecord { unsigned int(8) configurationVersion = 1; unsigned int(8) AVCProfileIndication; unsigned int(8) profile_compatibility; unsigned int(8) AVCLevelIndication; bit(6) reserved = '111111'b; unsigned int(2) lengthSizeMinusOne; bit(3) reserved = '111'b; unsigned int(5) numOfSequenceParameterSets; for (i=0; i< numOfSequenceParameterSets; i++) { unsigned int(16) sequenceParameterSetLength ; bit(8*sequenceParameterSetLength) sequenceParameterSetNALUnit; } unsigned int(8) numOfPictureParameterSets; for (i=0; i< numOfPictureParameterSets; i++) { unsigned int(16) pictureParameterSetLength; bit(8*pictureParameterSetLength) pictureParameterSetNALUnit; } if( AVCProfileIndication != 66 && AVCProfileIndication != 77 && AVCProfileIndication != 88 ) { bit(6) reserved = '111111'b; unsigned int(2) chroma_format; bit(5) reserved = '11111'b; unsigned int(3) bit_depth_luma_minus8; bit(5) reserved = '11111'b; unsigned int(3) bit_depth_chroma_minus8; unsigned int(8) numOfSequenceParameterSetExt; for (i=0; i< numOfSequenceParameterSetExt; i++) { unsigned int(16) sequenceParameterSetExtLength; bit(8*sequenceParameterSetExtLength) sequenceParameterSetExtNALUnit; } } } In these records, the bits of the parameter “configurationVersion” and reserved fields have a predetermined value equal to 1 (or 0). The encapsulation process thus marks these parameters as optional accordingly to the criterium checked in step 202. As a result, the encapsulation process may generate (in step 206) a compact AVC decoder configuration that contains the following parameters accordingly to first embodiment. This CompactAVCDecoderConfigurationRecord (or CondensedAVCDecoderConfigurationRecord) may be for example signaled in a CompactAVCDecoderConfigurationBox (or CondensedAVCDecoderConfigurationBox) with the ‘avcc’ four-character code box: aligned(8) class CompactAVCDecoderConfigurationBox extends Box('avcc') { CompactAVCDecoderConfigurationRecord () avcConfig; } aligned(8) class CompactAVCDecoderConfigurationRecord { unsigned int(8) AVCProfileIndication; unsigned int(8) profile_compatibility; unsigned int(8) AVCLevelIndication; unsigned int(2) lengthSizeMinusOne; unsigned int(5) numOfSequenceParameterSets; for (i=0; i< numOfSequenceParameterSets; i++) { unsigned int(16) sequenceParameterSetLength ; bit(8*sequenceParameterSetLength) sequenceParameterSetNALUnit; } unsigned int(8) numOfPictureParameterSets; for (i=0; i< numOfPictureParameterSets; i++) { unsigned int(16) pictureParameterSetLength; bit(8*pictureParameterSetLength) pictureParameterSetNALUnit; } if( AVCProfileIndication != 66 && AVCProfileIndication != 77 && AVCProfileIndication != 88 ) { unsigned int(2) chroma_format; unsigned int(3) bit_depth_luma_minus8; unsigned int(3) bit_depth_chroma_minus8; unsigned int(8) numOfSequenceParameterSetExt; for (i=0; i< numOfSequenceParameterSetExt; i++) { unsigned int(16) sequenceParameterSetExtLength; bit(8*sequenceParameterSetExtLength) sequenceParameterSetExtNALUnit; } } trailing_bits() } The semantics of the parameters in this box remains unchanged compared to the legacy AVC decoder configuration. Similar approaches may be also employed for SVCDecoderConfigurationRecord, MVCDecoderConfigurationRecord, LHEVCDecoderConfigurationRecord, EVCDecoderConfigurationRecord and LCEVCDecoderConfigurationRecord as per ISO / IEC 14496-15. Although this document describes mainly examples of decoder configuration for 2D video content coded using MPEG codecs, similar approaches may be employed for the any decoder or codec configurations that may have optional parameters. In addition, the embodiment may be also used for codec or decoder configuration of volumetric or 3D content or audio content. In another embodiment, the legacy decoder configuration comprises parameters that are coded using several bits. The encapsulation process may compress the parameters by limiting the maximum values of some parameters. The generation of the compact decoder configuration is modified as represented in the figure 2b. As in the previous embodiment, each parameter is successively processed. In addition to the criterion described in previous embodiment, the generation of the decoder configuration determines if the parameter can be compressed during a step 207. During this step 207, a check may be performed to assess if the parameter length in bits exceeds a predetermined threshold. When greater than the threshold, the parameter can be encoded using a more compact coding representation and is marked as compressible in a step 208. Otherwise, the parameter is marked as essential. The predetermined threshold may be set by a user, a computer program or may correspond to pre-determined value for example defined in a standard which is implemented in the encapsulation format. During such a step 207, a maximum parameter value for a specific parameter value may be calculated as a function of all values for a specific parameter value. If this maximum parameter value can be encoded using a smaller code than the code used in the regular encapsulation mode for the particular format, then such a code is used to encode all (or at least one) of the parameter values for this parameter. In the final step 206, the encapsulation module generates a compact decoder configuration that contains only the “essential” and “compressible” parameters. For example, a parameter that is coded using 16 bits but is known to have values that are often coded using less than 8 bits may be encoded using reduced number of bits (for example at most 8 bits). As a result, several bits are saved within the signaling of the compact description of the decoder configuration. In such embodiments, if the compact description mode is selected, the method for encapsulating comprises the steps of: - determining the length of at least one decoder configuration parameter, and - compressing at least one decoder configuration parameter if the length of said at least one decoder configuration parameter is above a length threshold value, said at least one compressed decoder configuration parameter being used to generate an image description and / or image header during the step of generating. In order to prevent the encapsulation process to use compact codec configuration when the value exceeds a predetermined threshold, it can be checked in step 102 if one of the coding configuration parameters exceeds its pre-determined threshold. In that case, the encapsulation may not use the compact description and generate this part of the image description using legacy (i.e. not compact) media file description. In a variant, when the parameter value size exceeds the pre-determined threshold, the marking step 208 is modified. All parameters of the decoder configuration are marked as essential, and the processing loop stops. It indicates that the encapsulation process cannot compress the decoder configuration. In that case, the compact decoder configuration generated in step 206 is identical to the legacy decoder configuration. This corresponds to a fall back to a legacy syntax for the decoder configuration. In that case, the decoder configuration should not be signaled as a compact decoder configuration. In another example, a variable length compression method is used for the parameter marked as compressible when the parameter has mostly small range of values. The variable length coding can be for example exp-golomb code or varint encoding. In the VVC decoder configuration record, some parameters such as the number of arrays allows specifying up to 256 arrays to carry initialization non-VCL NAL units. For compact image description of a single image, this number can be reduced to a lower number. For example, there is at most 19 types of non-VCL NAL units in VVC. In addition, only the DCI, OPI, VPS, SPS, PPS, prefix APS, and prefix SEI NAL units are commonly provided in the decoder configuration. For this reason, a maximum number of 7 arrays is needed in practical. For this reason, the length of the num_of_arrays parameter can be limited to 3 bits. Same principle is applicable to HEVC decoder configuration: the number of arrays can be up to 32 in a legacy HEVC decoder configuration while there are at most 5 NAL units that can be stored in these arrays (e.g. VPS, SPS, PPS, prefix SEI, or suffix SEI NAL unit). For this reason, a maximum number of 5 arrays is needed in practical for most of the bitstreams. For this reason, the length of the num_of_arrays parameter can be limited to 3 bits. In VVC decoder configuration, the NAL_unit_type length allows specifying up to 32 different NAL unit types while only 7 types of NAL units can be provided in the arrays in practical. For this reason, the NAL_unit_type length is limited to 3 bits and the following values are used to indicate the NAL unit type of the array: when NAL_unit_type equal to 0, it corresponds to OPI_NUT; equal to 1 for DCI_NUT, equal to 2 for VPS_NUT, equal to 3 for SPS_NUT, equal to 4 for PPS_NUT, equal to 5 for PREFIX_APS_NUT, equal to 6 for PREFIX_SEI_NUT. For HEVC decoder configuration, the NAL_unit_type field allows defining more than the 5 NAL units allowed to be declared in the decoder configuration. As for VVC decoder configuration, the NAL_unit_type field in the compact decoder configuration is coded using 3 bits at most and one value for this field is associated to each of these five NAL units. For each array, it is possible to specify up to 65536 NAL units in the VVC decoder configuration. In practical, at most one OPI, VPS, SPS, PPS are provided for single image. For this reason, the num_nalus parameter is inferred equal to 1 for all this NAL unit type. For prefix APS and prefix SEI NAL units, the number may be slightly bigger. For this reason, the num_nalus parameter is present for prefix APS and prefix SEI NAL unit but is limited to 3 bits to limit the number of SEI and APS NAL units provided to 8. For the HEVC decoder configuration, similar approach may be used to encode the numNalus parameter in the compact HEVC decoder configuration using 3 bits for example for prefix or suffix SEIs NAL units. The numNalus parameter is inferred equal to 1 for other NAL unit types. For the AVC decoder configuration, similar approach may be used to encode the numOfSequenceParameterSets, numOfPictureParameterSets and numOfSequenceParameterSetExt parameters in the compact HEVC decoder configuration that can be inferred equal to 1 for example. The length in bytes of each NAL unit is signaled in the VVC decoder configuration and is limited to 65536 bytes. In practical, the sizes of the NAL units are lower than this value and can be represented in most cases using 10 bits. For this reason, the length in bits of nal_unit_length is limited to a lower number of bits for example lower or equal to 10 bits. For AVC and HEVC decoder configurations, the fields indicating the length in bits of the NAL units (e.g. nalUnitLength for HEVC and sequenceParameterSetLength, pictureParameterSetLength and sequenceParameterSetExtLength for AVC) may be limited to 10 bits in the corresponding compact decoder configuration. The array_completeness flag allows indicating whether all NAL units of the given type are in the VVC or HEVC decoder configuration, and none are present in the data of the samples or of the items of the image. For a single image, it is very likely that when a value of array_completeness is set equal to 1 for a given NAL unit type, it is the same for all the NAL unit types described in the decoder configuration. For this reason, the array_completeness value is specified only for the first array entry and inferred equal to the same value for the other array entry in the compact decoder configuration. In a variant, the array_completeness value is not signaled in compact decoder configuration even for the first array entry and is inferred to a predetermined value. For example, it may be inferred equal to 1 to indicate that all configuration NAL units are provided in the compact decoder configuration and are not duplicated in the data of the image item. The VVC decoder configuration may provide information about the output layer set index of an output layer set represented by the referenced CVSs. In legacy VVC decoder configuration the maximum value of ols_idx is equal to 512. In practical, the number of output layer sets indicated in a VVC bitstream is lower. The number of output layer sets that can be present in a bitstream is proportional to the number of layers in the bitstream. In practical, the number of layers present in a still image can be equal to 1 for some VVC profiles. For this reason, the ols_idx may be removed if the VVC profile used is not allowing multiple layers. The profile signaling is made in the VvcPTLRecord. As a result, in the compact decoder configuration record, the ols_idx is signaled after the profile indication. The VvcPTLRecord also specifies a ptl_multilayer_enabled_flag which also indicates if multiple layers are used. Thus, the presence of the ols_idx is made conditional to ptl_multilayer_enabled_flag equal to 1. When the multiple layers can be used (i.e. ptl_multilayer_enabled_flag equal to 1) a total number of 128 layers is allowed. In practical, the number of layers can be limited to 8 for small images. As a result, a lower number of bits for example 3 bits is sufficient to describe the index of the output layer sets. The syntax of the compact decoder configuration may be the following, for example, according to combination of previous embodiments. aligned(8) class CompactVvcDecoderConfigurationBox extends Box('vvcc') { CompactVvcDecoderConfigurationRecord () vvcConfig; } aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) LengthSizeMinusOne; unsigned int(1) ptl_present_flag; if (ptl_present_flag) { unsigned int(2) chroma_format_idc; unsigned int(3) bit_depth_minus8; VvcPTLRecord(1) native_ptl; if (native_ptl.ptl_multilayer_enabled_flag == 1) unsigned int(3) ols_idx; unsigned_int(16) max_picture_width; unsigned_int(16) max_picture_height; } unsigned int(3) num_of_arrays; for (j=0; j < num_of_arrays; j++) { if (j == 0) unsigned int(1) array_completeness; unsigned int(3) NAL_unit_type_idc; if (NAL_unit_type_idc >= 5) unsigned int(3) num_nalus; else num_nalus = 1 for (i=0; i< num_nalus; i++) { unsigned int(10) nal_unit_length; bit(8*nal_unit_length) nal_unit; } } trailing_bits() } The semantics of the parameters are unchanged except for NAL_unit_type_idc and ols_idx. NAL_unit_type_idc indicates the type of the NAL units in the following array (which shall be all of that type); it may take the following values: - NAL_unit_type_idc equal to 0, the NAL unit is an OPI NAL unit; - NAL_unit_type_idc equal to 1, the NAL unit is a DCI NAL unit; - NAL_unit_type_idc equal to 2, the NAL unit is a VPS NAL unit; - NAL_unit_type_idc equal to 3, the NAL unit is an SPS NAL unit; - NAL_unit_type_idc equal to 4, the NAL unit is a PPS NAL unit; - NAL_unit_type_idc equal to 5, the NAL unit is a prefix APS NAL unit; - NAL_unit_type_idc equal to 6, the NAL unit is a prefix SEI NAL unit. ols_idx specifies the output layer set index of an output layer set represented by the referenced CVSs when present. When not present, ols_idx is inferred equal to 0. In another embodiment, the decoder configuration includes some Parameter Set or Essential SEI NAL unit(s) that may be useful for initializing the decoder. In order to further compress the bitstream, the encapsulation process may constrain the type and the order of the NAL units present in the compact decoder configuration. For example, the order of the NAL units may follow the presence probability of the NAL units in the decoder configuration. In most cases, the NAL units present in the decoder configuration record are the Parameter Sets in order of activation and usage by the VCL NAL units. For example, mandatory NAL units are placed first followed by optional NAL units. Generally, the presence of Parameter Sets NAL units such as VPS, SPS and PPS NAL unit are mandatory in the bitstream to be encapsulated as per codec specification. Other parameters set NAL units are optional since depends on activation of specific coding tools. However, they are mandatory when some coding tools is enabled. For example for VVC, it is the case of APS NAL units. In that case, they are placed immediately following the mandatory NAL units. Other NAL units are optional and not necessary for a correct decoding of the bitstream. This is the case for DCI, OPI and SEI NAL units. DCI and OPI NAL units can be used to indicate which layer should be decoded by the decoder and thus are more important than prefix SEI NAL units that generally contain supplemental information not mandatory for the decoding of the bitstream. As a result, the order of NAL units can be for example first a VPS, then SPS, PPS and APS (if any) NAL units, followed by DCI OPI and finally prefix SEI NAL units. Other NAL units such as DCI or OPI and prefix SEI are less likely to be used. For this reason, the encapsulation may indicate a parameter (for example) after the most important parameter sets to indicate whether additional NAL units are present in the decoder configuration. For example, when this parameter equal to 0, it indicates the absence of the additional NAL units in the decoder configuration. Otherwise, when equal to 1, it indicates the presence of the additional NAL units in the decode configuration. In such embodiments, if the compact description mode is selected, the method may further comprise a step of estimating a likelihood of use of at least one network abstraction layer unit parameter. Each network abstraction layer parameter associated with a likelihood of use estimation above a threshold value can be used to generate an image description and / or image header during the step of generating. Based on the predetermined order or the NAL units in the decoder configuration, the parser can determine the type of the NAL unit. The syntax of the compact decoder configuration may be the following for example according to combination of previous embodiments: aligned(8) class CompactVvcDecoderConfigurationBox extends Box('vvcc') { CompactVvcDecoderConfigurationRecord () vvcConfig; } aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) LengthSizeMinusOne; unsigned int(1) ptl_present_flag; if (ptl_present_flag) { unsigned int(2) chroma_format_idc; unsigned int(3) bit_depth_minus8; VvcPTLRecord(1) native_ptl; if (native_ptl.ptl_multilayer_enabled_flag == 1) unsigned int(3) ols_idx; unsigned_int(16) max_picture_width; unsigned_int(16) max_picture_height; } unsigned int(1) nal_units_present_flag; if (nal_units_present_flag) { unsigned int(1) array_completeness; unsigned int(8) vps_nal_unit_length; bit(8* vps_nal_unit_length) vps_nal_unit; unsigned int(8) sps_nal_unit_length; bit(8*sps_nal_unit_length) sps_nal_unit; unsigned int(8) pps_nal_unit_length; bit(8*pps_nal_unit_length) pps_nal_unit; unsigned int(1) additional_nal_unit_flag; if (additional_nal_unit_flag) { unsigned int(3) num_aps_nal_unit; unsigned int(3) num_sei_nal_unit; for (i=0; i< num_aps_nalus; i++) { unsigned int(8) aps_nal_unit_length; bit(8*aps_nal_unit_length) aps_nal_unit; } for (i=0; i< num_sei_nalus; i++) { unsigned int(8) sei_nal_unit_length; bit(8*sei_nal_unit_length) sei_nal_unit; } } } trailing_bits() } The semantics are unchanged except for the following parameters: nal_units_present_flag indicates that NAL units are present in the decoder configuration. vps_nal_unit_length indicates the length in bytes of the NAL unit. When equal to 0, the VPS NAL unit is not present. vps_nal_unit contains the VPS NAL unit as specified in ISO / IEC 23090-3. sps_nal_unit_length indicates the length in bytes of the NAL unit. When equal to 0, the SPS NAL unit is not present. sps_nal_unit contains the SPS NAL unit as specified in ISO / IEC 23090-3. pps_nal_unit_length indicates the length in bytes of the NAL unit. When equal to 0, the PPS NAL unit is not present. pps_nal_unit contains the PPS NAL unit as specified in ISO / IEC 23090-3. additional_nal_unit_flag equal to 1 indicates the presence of additional NAL units in the decoder configuration record. additional_nal_unit_flag equal to 0 indicates the absence of additional NAL units in the decoder configuration record. num_aps_nal_unit indicates the number of APS NAL units included in the configuration record for the referenced CVS. num_sei_nal_unit indicates the number of SEI NAL units included in the configuration record for the referenced CVS. aps_nal_unit_length indicates the length in bytes of the APS NAL unit. aps_nal_unit contains the APS NAL unit as specified in ISO / IEC 23090-3. sei_nal_unit_length indicates the length in bytes of the SEI NAL unit. sei_nal_unit contains the SEI NAL unit as specified in ISO / IEC 23090-3. In a variant, the type and / or the order of the NAL units is not constrained in the bitstream. The NAL units provided in the decoder or codec configuration already define the type of NAL units in the NAL unit header. As a result, the decoder can determine the type of the NAL unit by parsing the type specified in the NAL unit header of each NAL unit present in the configuration record. The number of NAL units present in the decoder configuration record is signaled. Then the length in bytes and the data of each NAL unit is indicated in the decoder configuration. In another variant to avoid having to specify the number and the length of NAL units signaled in the decoder configuration, the encapsulation process provides the NAL units as a set of NAL units separated by start codes. Emulation prevention bytes may be inserted to avoid emulated unwanted start code in the buffer storing the set of NAL units. For example, for VVC, HEVC or AVC codecs, the set of NAL units forms a bitstream conforming to Byte stream format of the corresponding codec specifications. In another embodiment, the encapsulation compact decoder configuration may constrain that the decoder configuration does not convey any of the NAL units. In that case, the decoder configuration parameters related to the NAL units are predetermined. For example, for VVC decoder configuration, the num_of_arrays is predetermined equal to 0; for HEVC decoder configuration the numOfArrays is predetermined equal to 0; and for AVC numOfSequenceParameterSets, numOfPictureParameterSets and numOfSequenceParameterSetExt are inferred equal to 0. In another embodiment, to efficiently code the parameter in the decoder configuration, the encapsulation processing may rely on coding representation used in the Minimized Image Box to limit the length of the parameters in the decoder configuration. The MinimizedImageBox may define parameters or indicators that limit the range of some parameters coding attributes of the image. For example, one indicator may limit the size of the parameters coding the width and the height of the image (i.e. the image dimension) to be represented on a predetermined maximum number of bits. The encapsulation process may use the same coding representation employed for the image dimensions parameters coded in the MinimizeImageBox to encode the image dimensions data coded in the codec or decoder configuration box. Additionally, the minimized image box may specify the length of the codec or decoder configuration data associated to an item. The length in bits of this length parameter limits the length of the NAL units present in the decoder or codec configuration. Indeed, the length of the NAL units present in the decoder configuration must be lower that the length of the decoder configuration. As a result, the number of bits used to encode the length of the NAL units present in the decoder configuration must be lower or equal to the number of bits used to encode the length of the decoder configuration. The encapsulation process may use the coding representation of the length of codec or decoder configuration used in the MinimizedImageBox to encode the NAL unit lengths in the codec or decoder configuration record. The syntax of CompactVvcDecoderConfigurationRecord generated by the encapsulation process may be the following for example: aligned(8) class CompactVvcDecoderConfigurationRecord(unsigned int bit_len_dim, unsigned int bit_len_coded_sizes) { unsigned int(2) LengthSizeMinusOne; unsigned int(1) ptl_present_flag; if (ptl_present_flag) { unsigned int(9) ols_idx; unsigned int(2) chroma_format_idc; unsigned int(3) bit_depth_minus8; VvcPTLRecord(1) native_ptl; unsigned int(bit_len_dim) max_picture_width; unsigned int(bit_len_dim) max_picture_height; } unsigned int(1) array_completeness; unsigned int(bit_len_coded_sizes) vps_nal_unit_length; bit(8*vps_nal_unit_length) vps_nal_unit; unsigned int(bit_len_coded_sizes ) sps_nal_unit_length; bit(8*sps_nal_unit_length) sps_nal_unit; unsigned int(bit_len_coded_sizes ) pps_nal_unit_length; bit(8*pps_nal_unit_length) pps_nal_unit; trailing_bits() } The bit_len_dim and bit_len_coded_sizes parameters of the class indicate respectively the length in bits of the image dimension parameters (i.e. max_picture_width and max_picture_height) and of the coded size of the NAL unit length parameters. The encapsulation process may initialize these values to the length in bits of the image dimension parameters used in the MinimizedImageBox and the length in bits of the configuration data length parameter. In another embodiment, the decoder configuration indicates profile, tier and level information that are coded using a pre-determined number of bits. In practical terms, the number of profile, tier and level (PTL) defined by the codec specification is limited. As a result, instead of defining a wide range of bits for these parameters, the encapsulation processing may reduce number of bits used to encode PTL parameters. The values of these new parameter may indicate a profile according to a predetermined set of values. For example, in VVC decoder configuration the parameter general_profile_idc may have the value 1 for Main 10 profile, the value 65 for Main 10 Still picture profile, the value 33 for Main 104:4:4 profile, the value 97 for Main 104:4:4 Still Picture, the value 17 for Multilayer Main 10 profile, the value 49 for Multilayer Main 104:4:4 profile, the value 2 for Main 12 profile, the value 10 for Main 12 profile Intra profile and the value 66 for Main 12 Still Picture profile, etc. This represents a total of 17 possible values. As a result, instead of using 8 bits to encode this parameter, the general_profile_idc is coded on 5 bits. Each value of the new compressed general_profile_idc parameter is associated to one of value of the non-compressed general_profile_idc defined in ISO / IEC 23090-3. In a variant, a flag may be coded in the compact decoder configuration that can be set to 1 to indicate that compressed version of general_profile_idc is not referencing a predetermined profile and in that case an additional 8 bits field is provided to encode the uncompressed (i.e. using the full range of possible value for the profile) value of general_profile_idc. Similar approaches may be used for other parameter of the PTL structure in the decoder configurations. In another embodiment, the encapsulation process may avoid specifying parameters in the compact decoder configuration box that can be derived from parameter indicated in the MinimizedImageBox. For example, parameters provided in the header of the MinimizedImageBox or provided at the beginning of the MinimizedImageBox. The figure 2c illustrates the generation of the compact configuration. As in previous embodiments, each parameter is successively processed. In addition, to the criteria of previous embodiment, it is checked in a step 209 if the parameter provided in the legacy decoder configuration box can be deduced from the parameters encoded before the configuration box in the MinimizedImageBox . When the step 209 indicates that the parameter can be deduced from parameters coded prior in the MinimizedImageBox, the parameter is marked as redundant in step 210. For example, the decoder configuration box may provide parameters specifying the image dimensions, the chroma format and the bit depth used for the pixel value of the image. Since similar parameters are also provided at the beginning of the mini box, there is a redundancy. In such embodiments, if the compact description mode is selected, the method for encapsulating further comprises a step of determining if a decoder configuration parameter is deductible, each configuration parameter determined to be nondeductible being used to generate an image description and / or image header during the step of generating. As it is understood, a parameter can be considered as deductible if a value for this parameter may be derived from a value in another parameter of the image file format metadata. As a result, the encapsulation process may generate compact decoder configuration without these parameters in step 206. The parser can infer the value of the legacy decoder configuration box based on the parameter provided in the Minimized Image Box. In the final step 206, the encapsulation module generates a compact decoder configuration that contains only the “essential” and “compressible” parameters and not optional or redundant parameters. For example, the VVC decoder configuration record as defined in ISO / IEC 14496-15 contains the chroma_format_idc, which indicates the chroma format that applies to the referenced CVSs. The minimized image box also defines the chroma subsampling parameters. The semantics of the two parameters are the same and thus the value of chroma_format_idc and the one in the MinimizedImageBox are redundant. The encapsulation process thus generates a compact VVC decoder configuration that is not indicating the chroma_format_idc parameters. The parser uses the subsampling information provided in the MinimizedImageBox to determine the value of chroma_format_idc in the legacy VVC decoder configuration. In another example, the HEVC decoder configuration record as defined in ISO / IEC 14496-15 contains the chroma_format_idc that has the same semantics as the one in the MinimizedImageBox and thus can be considered as redundant. The MinimizedImageBox also specifies the image dimensions i.e., the width and the height of the main item. This information is also provided in the VVC decoder configuration. In order to further compress the VVC decoder configuration, the encapsulation process generates a compact VVC decoder configuration record without indication of the image size. The parser then determines the value in the legacy decoder configuration record using the information of image dimension as specified in MinimizedImageBox. For example, the encapsulation may generate the following compact decoder configuration record for which the parameters chroma_format_idc, bit_depth_minus8, max_picture_width and max_picture_height are not present. The parser would infer the value of these parameters from the parameters coded in MinimizedImageBox. aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) LengthSizeMinusOne; unsigned int(1) ptl_present_flag; if (ptl_present_flag) { VvcPTLRecord(1) native_ptl; if (native_ptl.ptl_multilayer_enabled_flag == 1) unsigned int(3) ols_idx; } unsigned int(1) nal_units_present_flag; if (nal_units_present_flag) { unsigned int(1) array_completeness; unsigned int(8) vps_nal_unit_length; bit(8* vps_nal_unit_length) vps_nal_unit; unsigned int(8) sps_nal_unit_length; bit(8*sps_nal_unit_length) sps_nal_unit; unsigned int(8) pps_nal_unit_length; bit(8*pps_nal_unit_length) pps_nal_unit; unsigned int(1) additional_nal_unit_flag; if (additional_nal_unit_flag) { unsigned int(3) num_aps_nal_unit; unsigned int(3) num_sei_nal_unit; for (i=0; i< num_aps_nalus; i++) { unsigned int(8) aps_nal_unit_length; bit(8*aps_nal_unit_length) aps_nal_unit; } for (i=0; i< num_sei_nalus; i++) { unsigned int(8) sei_nal_unit_length; bit(8*sei_nal_unit_length) sei_nal_unit; } } } trailing_bits() } The compact decoder configuration is not necessarily constrained to be used within MinimizedImageBox and can be used in other contexts or boxes. In another embodiment, the compact decoder configuration replaces the regular or legacy decoder configuration in SampleEntry of track (for example a video track or an image sequence track). It may also replace a decoder configuration property associated to an image item, or a compact image property may be defined with its fields corresponding to the parameters of a compact decoder configuration. Same compacting approaches as described in the previous embodiments are still valid for the parameter of the compact decoder configuration in a SampleEntry of a track. The parameters marked as optional or compressible can be handled as in previous embodiment. Concerning the parameters deductible in step 209, the encapsulation and de-encapsulation may rely on the parameters provided in the track box, such as the track header box, to infer value of parameters of the decoder configuration, or by a dedicated sample entry type. For example, the image dimension specified in the track may be used to infer the image dimensions of the compact decoder configuration record. In another embodiment, the compact description may include an additional information generated by the encapsulation processing. This additional information may be used by a parser to determine the features used for the compression of the decoder configuration box. The information may be for example, a set of bits indicating the parameter present in the compact description. For example, a first bit indicates that predetermined parameters are not present in the compact decoder configuration. A second bit indicates that inconsistent parameters are not present in the compact decoder configuration. A third bit indicates that compressible parameters have been compressed in the compact decoder configuration. A fourth bit indicates that deductible parameters are not present in the compact decoder configuration. In a variant, the additional information indicates a level of compaction (or degrees of compaction) for the compact description. For example, a first level indicates that predetermined parameters are not present in the compact decoder configuration. A second level indicates that predetermined parameters and inconsistent parameters are not present in the compact decoder configuration. A third level indicates that predetermined parameters, inconsistent parameters are not present and compressible parameters have been compressed in the compact decoder configuration. A fourth level indicates that in addition to the constraints of the third level, the deductible parameters are not present in the compact decoder configuration. The level information may be for example an integer parameter coded on 2 bits. In a variant, any selection and / or combination of compaction features (i.e. removal of optional parameters or removal of inconsistent parameters or removal of redundant parameters or compression of compressible parameters) can be associated to one level. In another variant, the additional information may be inferred when one brand is present in the list of the compatible brands of the file. For example, a brand ‘ccnf’ may indicate that compact decoder configuration is used for the boxes (e.g. including sample entries and or MinimizedImageBox) present in the files. A brand may be associated to each level of compaction. For example, the brand ‘ccf1’ may indicate that compact decoder configuration is used for the boxes present in the files and that the level of compaction is set to 1. A brand ‘ccf2’ may indicate that compact decoder configuration is used for the boxes present in the files and that the level of compaction is set to 2. More generally the brand ‘ccfX’ where in X is a number between 0 and 9, indicates that compact decoder configuration is used for the boxes present in the files and that the level of compaction is equal to the X number. In another embodiment, the array of parameter set that are usually placed in the decoder configuration box are implicitly in the data. Then, all the syntax elements for the parameter arrays can be omitted. This may be indicated by specific item types when used for image items or sample entry types when used in tracks. Especially, video tracks distinguish tracks having parameter sets all in sample entry or potentially in both sample entry or data, but there is no sample entry type to indicate that they are all kept in the data and none in the decoder configuration box. In another embodiment, the parameters provided in the decoder configuration record or box are constrained to provide essential information for a single image (or picture). It comprises parameter sets NAL unit essential for a single image such as SPS or PPS NAL units in a context of VVC, HEVC, AVC or any other NAL unit based codecs. The information considered as non-essential for a single image corresponds to optional NAL units not necessary for initialization of a single image decoder. For example, for a VVC bitstream, it may correspond to APS NAL units or more generally for any MPEG codecs it may be the SEI NAL units. Video Parameters Sets (VPS) NAL unit provides parameters that describe multiple layers for which multiple images or pictures are encoded. As a result, VPS are not considered as essential information for a single image. In addition, multiple SPS and PPS NAL units may be present for multilayer images (e.g. one NAL unit for each layer). These NAL units specify parameters that apply to a given layer. For such kind of images only one set of SPS and PPS NAL units is considered as essential. For example, only the SPS and PPS NAL unit corresponding to the layer that should be displayed by default is provided in the decoder configuration record. In a variant, only the parameter sets NAL units describing the base layer (e.g. the layer with a layer identifier equal to 0) are provided. In this embodiment, the essential information for single image is signaled and non-essential information for single images is optionally provided. For example, the encapsulation may generate the following compact decoder configuration record. aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) lengthSizeMinusOne; unsigned int(6) num_additional_nal_units; / / default = 0 unsigned int(8) sps_nal_unit_length; bit(8*sps_nal_unit_length) sps_nal_unit; unsigned int(8) pps_nal_unit_length; bit(8*pps_nal_unit_length) pps_nal_unit; for (i=0; i< num_additional_nal_units; i++) { unsigned int(8) nal_unit_length[ i ]; bit(8*nal_unit_length) nal_unit[ i ]; } } The lengthSizeMinusOne parameter plus 1 indicates the length in bytes of the NALUnitLength field in an image item associated with this decoder configuration. num_additional_nal_units indicates the number of NAL units present in the decoder configuration in addition to the mandatory or essential SPS and PPS of the main image item. This parameter is by default equal to 0 when the decoder configuration record provides only essential information for a single layer. Otherwise, when greater than 1 it indicates that non-essential information for single image is provided. As a result, during de-encapsulation process the player may determine the presence of a multiple layer image when the num_additional_nal_units is greater than 0 and one VPS or multiple PPS or SPS NAL units are present in the list of additional units (it corresponds to the nal_unit[ i ] syntax element). The encapsulation process may constrain this parameter to be equal to 0 when no essential information is provided (for example when both sps_nal_unit_length and pps_nal_unit_length equal to 0). The objective of this optional constraint is to ensure that the compact decoder configuration is providing at least a piece of essential information for a single image and is not describing only non-essential information. sps_nal_unit_length indicates the length in bytes of the NAL unit. When equal to 0, the SPS NAL unit is not present. sps_nal_unit contains the SPS NAL unit as specified in ISO / IEC 23090-3, when present. pps_nal_unit_length indicates the length in bytes of the NAL unit. When equal to 0, the PPS NAL unit is not present. A constraint may be imposed during encapsulation process that either one of sps_nal_unit_length and pps_nal_unit_length is greater than 0. This constraint ensures that essential information for the single image is not empty. In a variant, since the information provided in the SPS is of higher level than the information provided in the PPS, the encapsulation may constrain that the length of the SPS is greater than 0 (which makes sure that at least one SPS NAL unit is provided) and allows that the length of the PPS is greater than or equal to 0. nal_unit_length[ i ] indicates the length in bytes of the i-th additional NAL unit stored in this property. nal_unit[ i ] contains the i-th additional NAL unit as specified in ISO / IEC 23090- 3. During the encapsulation, the order of the additional NAL units may be constrained to follow their importance order to initialize the decoder. For example, the order of the NAL additional units may be the following: zero or one VPS NAL unit, zero or more SPS NAL units (describing additional layer different than the one described by the SPS of sps_nal_unit syntax element), zero or more PPS NAL units (describing additional layers different than the one described by the PPS of pps_nal_unit syntax element), zero or more APS NAL units (for any layer of the item), zero or more SEI NAL units (for any layer of the item). The advantage is that de-encapsulation process may determine the presence of VPS NAL unit by parsing only the first additional unit provided in the nal_unit
[0000] parameter i.e. by parsing at most three NAL units (the essential SPS and PPS NAL units if present and the NAL unit of the nal_unit
[0000] parameter). In this embodiment, the presence of Profile Tier Level information (e.g. represented ptl_present_flag as per ISO / IEC 14496-15) parameters of the decoder configuration record is skipped (e.g. ptl_present_flag is not signaled in the compact decoder configuration and is inferred to be equal to 0). Indeed, this information is generally also present in the SPS or PPS NAL unit. As a result, the determination of the Profile Tier Level information from compact decoder configuration record is inferred from the SPS and PPS NAL unit provided as essential information for single-image. The parameters describing the presence of setup NAL units (e.g. num_of_arrays, num_nalus and NAL_unit_type as per ISO / IEC 14496-15) in the decoder configuration can be deduced from the parameters of the compact decoder configuration. For example, when the num_additional_nal_units parameter is equal to 0, there are two arrays (num_of_arrays): one for the essential SPS and one for the essential PPS NAL unit. The number of arrays may be equal to 1 when the length of either one of the essential SPS or PPS NAL unit is set to 0. The number of arrays may be equal to 0 when both length of the SPS and PPS NAL unit is set equal to 0. For each array the number of NAL units (num_nalus) is predetermined equal to 1. The NAL_unit_type of the first array indicates a SPS NAL unit type and a PPS NAL unit type for the second array when present. When num_additional_nal_units is greater than 0, the number of arrays is incremented for each new NAL unit type parsed in the nal_unit arrays and the number of NAL units (num_nalus) of that type is incremented by 1 for each NAL unit parsed. In this embodiment,CompactVvcDecoderConfigurationRecordis considered equivalent to VvcDecoderConfigurationRecord as defined in ISO / IEC 14496-15 with the following fields: •ptl_present_flagis set to 1 • num_sublayers is set to 1, • constant_frame_rate is set to 1 • if the codec configuration property is associated with the main image item: ochroma_format_idcis set to the value of thechroma_subsamplingfield from the MinimizedImageBox, • if the codec configuration property is associated with the alpha auxiliary image item: o chroma_format_idc is set to 0, • if the codec configuration property is associated with the main image item or with the alpha auxiliary image item: obit_depth_minus8is set to 0 if the value of thehigh_bit_depth_flagfield from theMinimizedImageBoxis 0, or to the value of thebit_depth_minus9field from the MinimizedImageBox plus 1, o max_picture_width is set to the value plus 1 of the width_minus1 field from the MinimizedImageBox, o max_picture_height is set to the value plus 1 of the height_minus1 field from theMinimizedImageBox,• if the codec configuration property is associated with the gain map image item: o bit_depth_minus8 is set to 0 if the value of the gainmap_high_bit_depth_flag field from the MinimizedImageBox is 0, or to the value of thegainmap_bit_depth_minus9field from the MinimizedImageBoxplus 1, o max_picture_width is set to the value plus 1 of the gainmap_width_minus1 field from the MinimizedImageBox, o max_picture_height is set to the value plus 1 of the gainmap_height_minus1 field from the MinimizedImageBox,• avg_frame_rateis set to 0,•native_ptl parameters are inferred from the profile_tier_level( 1, sps_max_sublayers_minus1 ) structure as defined the seq_parameter_set_rbsp( ) as per ISO / IEC 23090-3 of the SPS NAL unit provided insps_nal_unit.• num_of_arrays is determined as follows: o num_of_arrays is set equal to 2 o There is a SPS NAL unit array with: ▪ NAL_unit_type is set to 15 (SPS_NUT as defined in ISO / IEC 23090-3), ▪num_nalusis set to1andnal_unit_lengthis set to sps_nal_unit_length, and nal_unit is set to sps_nal_unit. o There is a PPS NAL unit array with: ▪ NAL_unit_type set to 16 (PPS_NUT as defined in ISO / IEC 23090- 3), ▪ num_nalus set to 1, nal_unit_length set to pps_nal_unit_length, and nal_unit set to pps_nal_unit. o • if num_additional_nal_units is greater than 0, each i-th additional unit nal_unit[ i ] is processed as follows: o The NAL unit type of nal_unit[ i ] (nal_unit_type as defined in ISO / IEC 23090-3) is determined ▪ There is NAL unit array with NAL_unit_type set to nal_unit_type of the nal_unit[ i ] •If thenal_unit_type of nal_unit[ i ] is not 15 (SPS_NUT as defined in ISO / IEC 23090-3) nor 16 (PPS_NUT as defined in ISO / IEC 23090-3) and if nal_unit[ i ] is the first NAL unit with this NAL unit type, num_arrays is incremented by 1 • num_nalus is set to 1 if it is the first NAL unit of the array, otherwise num_nalus is incremented by 1. • A new set of nal_unit_length and nal_unit is added to the array as follows o nal_unit_length is set to nal_unit_length[ i ] of the CompactVvcDecoderConfigurationRecord, onal_unitis set tonal_unit[ i ]of the CompactVvcDecoderConfigurationRecord. • thearray_completenessis set equal to 1 inVvcDecoderConfigurationRecord. the other parameters are carried over as is, and repeated if needed In a variant of previous embodiment, the essential and non-essential NAL units are signaled in the same signaling loop. The minimal number of NAL units described in the processing loop is constrained to be equal to 2 by design of the decoder configuration record semantics. For example, the compact decoder configuration syntax may be the following: aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) lengthSizeMinusOne; unsigned int(6) num_nal_units_minus2; for (i=0; i< num_nal_units+2; i++) { unsigned int(8) nal_unit_length[ i ]; bit(8*nal_unit_length) nal_unit[ i ] ; } } The semantics of the lengthSizeMinusOne and nal_unit[ i ] syntax elements are the same as in previous variant. num_nal_units_minus2 plus 2 indicates the number of NAL units present in the decoder configuration. The encapsulation may constrain that when nal_unit_length[0] is greater than 0, the nal_unit[0] (the first NAL unit of the signaling loop) contains the SPS NAL unit; and that when nal_unit_length[1] is greater than 0, the nal_unit[1] (the second NAL unit of the signaling loop) contains the PPS NAL unit. To ensure that at least one Parameter Sets NAL unit provides essential information for single image, the encapsulation may constrain that only one of the nal_unit_length[0] or nal_unit_length[1] is equal to 0. The NAL units described after the two first NAL units correspond to non-essential information for a single image. In another variant, the length of the last NAL unit provided in the compact decoder configuration can be inferred from the total length of the compact decoder configuration record. This length can be provided for example prior to the signaling of this record. For example, it can be provided as a parameter (e.g. main_item_codec_config_size) in the MinimizedImageBox box. The length of the last NAL unit is equal to the length of the compact codec configuration (main_item_codec_config_size) minus the sum of the NAL unit lengths provided in the compact decoder configuration except the last one and minus the lengths of the other parameters provided in the compact decoder configuration unit. For example, the encapsulation compact decoder configuration may use the following syntax: aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) LengthSizeMinusOne; unsigned int(6) num_additional_nal_units; unsigned int(8) sps_nal_unit_length; bit(8*sps_nal_unit_length) sps_nal_unit; if (num_additional_nal_units > 0) unsigned int(8) pps_nal_unit_length; bit(8*pps_nal_unit_length) pps_nal_unit; for (i=0; i< num_additional_nal_units; i++) { if (i < num_additional_nal_units – 1) unsigned int(8) nal_unit_length[ i ]; bit(8* nal_unit_length) nal_unit[ i ]; } } The semantics of the parameters are the same as in previous variants with the following additional statements: When not present, pps_nal_unit_length is inferred equal to main_item_codec_config_size – 1 – sps_nal_unit_length – additional_units_length. In this formula 1 byte is subtracted to the length of the compact decoder configuration since corresponds to the length of lengthSizeMinusOne and num_additional_nal_units; The variable additional_units_length represents the sum of the lengths of the additional NAL units provided in the NAL unit signaling loop. Thus, additional_units_length is equal to the sum of nal_unit_length[ i ] for i in the range of 0 to num_additional_nal_units – 1, inclusive when num_additional_nal_units is greater than 1. Otherwise, when num_additional_nal_units is less than or equal to 1 additional_units_length is inferred equal to 0. Another possible syntax for the decoder configuration record may be the following: aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) lengthSizeMinusOne; unsigned int(6) num_nal_units_minus2; for (i=0; i< num_nal_units_minus2+2; i++) { if (i < num_nal_units_minus2 + 1) unsigned int(8) nal_unit_length[ i ]; bit(8*nal_unit_length) nal_unit[ i ] ; } } The semantics of the parameters are the same as in previous variant with the following additional statements: nal_unit_length[ num_nal_units_minus2 + 1 ] is inferred equal to main_item_codec_config_size – 1 – signaled_units_length. The variable signaled_units_length that is the sum of the lengths of the NAL units provided in the NAL unit signaling loop is equal to the sum of nal_unit_length[ i ] for i in the range of 0 to num_nal_units_minus2, inclusive when num_nal_units_minus2 is greater than 1. Otherwise, when num_nal_units_minus2 is less than or equal to 1 signaled_units_length is inferred equal to 0. In yet another variant of the previous embodiment, the compact decoder configuration syntax provides only essential information of a single image. As a result, any syntax element related to the signaling of non-essential information for a single image is removed from the compact decoder configuration record. For example, at most one SPS and one PPS NAL unit is provided in the decoder configuration record. Other non-essential information may be provided in the data of the image item. For example, the syntax of the compact decoder configuration record may be the following with similar semantics as in previous variants. aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) lengthSizeMinusOne; unsigned int(8) sps_nal_unit_length; bit(8*sps_nal_unit_length) sps_nal_unit; unsigned int(8) pps_nal_unit_length; / / optional bit(8*pps_nal_unit_length) pps_nal_unit; } As in one of the previous variants that avoids signaling the length of the last NAL unit of the record, which is the length of the PPS NAL unit represented by pps_nal_unit_length in the above syntax example , may be inferred from the length of the record. It can be deduced from length of configuration of record from the MinimizedImageBox (‘mini’ box) with the following formula pps_nal_unit_length = main_item_codec_config_size – 1 - sps_nal_unit_length. In another embodiment, the compact decoder configuration is further simplified and compressed by removing the lengths of the NAL units provided in the decoder configuration record. The set of NAL units provided as essential information for single image are Parameter Sets NAL units. The syntax of these NAL units is provided in codec specifications such as VVC, HEVC, AVC or any other MPEG codecs. As a result, a parser may follow the codec specification to determine the content of these essential NAL units and may infer the length of the NAL units by counting the number of parsed bytes. The encapsulation process may thus avoid signaling the lengths of the known NAL units from the decoder configuration. For example, considering VVC (but could apply to AVC, HEVC or any other MPEG codecs), the compact decoder configuration syntax may be the following with the same semantics as in previous embodiments and variants: aligned(8) class CompactVvcDecoderConfigurationRecord { unsigned int(2) lengthSizeMinusOne; VVC_SPSNALUnit sps_nal_unit; VVC_PPSNALUnit pps_nal_unit; } aligned(8) class VVC_SPSNALUnit { nal_unit_header( ) / / as per codec specification e.g. ISO / IEC 23090-3 seq_parameter_set_rbsp( ) / / as per codec specification e.g. ISO / IEC 23090-3 } aligned(8) class VVC_PPSNALUnit { nal_unit_header( ) / / as per codec specification e.g. ISO / IEC 23090-3 pps_parameter_set_rbsp( ) / / as per codec specification e.g. ISO / IEC 23090-3 } In a variant, the compact decoder configuration provides NAL units with a syntax that may be not supported by a parser. For example, it may correspond to SEI NAL units that are optional and may be discarded by a decoder. In such a case, the encapsulation process may signal these optional NAL units at the end of the decoder configuration. When a parser processes a compact decoder configuration that contains NAL units with an unknown syntax, it can skip all the remaining NAL units up to the end of the compact decoder configuration (that can be determined using the length of the compact decoder configuration provided for example as a parameter in a MinimizedImageBox). In a variant, the encapsulation may use a specific NAL unit as a delimiter between the essential and non-essential NAL units for a single image or between parse-able and optional NAL units that may not be known by any parser. When the parser encounters this specific delimiter NAL unit it may consider that remaining NAL units are optional (or non-essential). For example, the NAL unit used as a delimiter may be a NAL unit with a type that is usually not allowed in a regular decoder configuration. In case of MPEG codec, it can be for example an access unit delimiter or an EOB (end of bitstream) or EOS (end of slice) NAL unit. The advantage of EOB and EOS NAL units comes from the easy to parse syntax and small payload size. Any other similar NAL units different than Parameter Sets NAL unit or SEI NAL units not allowed in a regular decoder configuration may be used as delimiter. In another variant, the encapsulation process may signal the lengths of the NAL units only after the NAL unit that is used as delimiter. In that case, when the parser identified this NAL unit, it continues parsing the remaining data of the compact decoder configuration as a contiguous series of one NAL unit length followed by the NAL unit data. NAL units prior to the delimiter are not preceded by a NAL unit length. In another embodiment, the compact decoder configuration is not signaling the length in bytes of the NALUnitLength field in an image item associated with the decoder configuration. For example, it may correspond to the lengthSizeMinusOne of some of the previous embodiments. In that case, the parser infers the value of this parameter to a predetermined value for example equal to 3 bytes. In a variant, the value of this parameter is inferred from parameters defined in the box containing the compact decoder configuration. The MinimizedImageBox may define parameters or indicators that limit the range of some parameters coding the length in bytes of the image. For example, the MinimizedImageBox may provide a parameter (e.g. large_item_data_flag) that is equal to 0 when the image or item size is signaled using 15 bits and that is equal to 1 when the size is expressed on 28 bits. Since the length in bytes of the NALUnitLength is used to signal the length of NAL units encoding the item or image, the value of lengthSizeMinusOne may be determined accordingly to the MinimizedImageBox parameters. For example, lengthSizeMinusOne is infered to a value that allows specifying NAL units with a length in up to 2^28 bits when large_item_data_flag equal 1 (for example it can be 3 to indicate that 4 bytes are used for the NALUnitLength). Otherwise, when large_item_data_flag equal 0, lengthSizeMinusOne is infered equal to a value that can express a NAL unit length up to 2^15 bits (for example it can be 1 to indicate that 2 bytes are used for the NALUnitLength). In another variant, when compact decoder configuration is used in a MinimizedImageBox, the length of the NALUnitLength field in the item data has the same value as the parameter specifying the length in bytes of the item data in the MinimizedImageBox. For example, when large_item_data_flag is equal to 0 the length in bits of the NALUnitLength is 15 bits. When large_item_data_flag is equal to 1 the length in bits of the NALUnitLength is 28 bits. The syntax of the compact decoder configuration without the lengthSizeMinusOne may be the following with the same semantics as previous embodiments: aligned(8) class CompactVvcDecoderConfigurationRecord { VVC_SPSNALUnit sps_nal_unit; VVC_PPSNALUnit pps_nal_unit; } Figure 4 illustrates the main steps in a de-encapsulation process according to some embodiments of the invention. First, the media file is parsed in a step 401 to determine in the step 402 the presence of compact description in the media file. For example, this may be indicated by a specific brand in the FileTypeBox or by the presence of Minimized Image Box. When it is determined in step 402 that compact description is not present, the legacy or regular processing of the MetaBox is performed in a step 409. Otherwise, the processing applies successively each step 403 to 408. When the media file comprises compact description (checked in a step 405), it means that at least one Minimized Image Box is present in the media file. The Minimized Image Box is parsed in step 403 to determine the properties of the image described using compact description. The compact description of the image may comprise decoder configuration. In that case, it is determined in step 404 the type of the decoder configuration meaning if it is a decoder configuration with a regular or legacy syntax as per ISO / IEC 14496-15 (indicating that the compact description mode is inactive) or a compact codec configuration (indicating that the compact description mode is active). The compact codec configuration syntax may be determined using a version or a flag parameter indicated in the decoder configuration record. In a variant it may relies on the presence of a brand. In a variant, it can be determined using the type or a version or a flag of a box containing the decoder configuration record. For example, a VVC decoder configuration record using a regular or legacy syntax is generally provided in a box with “vvcC” type (i.e. the registered four character code as per ISO / IEC 14496-15, indicating the compact description mode is inactive). When compact syntax is used, the type of the box may be set to “vvcc” (i.e; to another four character code dedicated to this usage, the presence of these four-character code indicating the compact description mode is active). When the decoder configuration uses a compact syntax, the de-encapsulation processing infers the legacy decoder configuration (i.e. using legacy verbose syntax) in step 406. The de-encapsulation module reconstructs the Minimized Items selected by the application in step 407 by extracting the Minimized Item Data provided in the byte ranges as determined in 403. The decoder of the Minimized Item Data data structure can be initialized by the configuration data provided in the Minimized Item Data data structure when present or inferred from the major_brand of the FileTypeBox according to the value of the explicit_codec_type parameter. Data of each item described by the compact description may then be provided in step 408 to a decoder for generating the decoded version of the item data. When the compact decoder configuration is in use, the Minimized Image Box has the explicit_codec_type set to 0 and parser can deduce the codec to a predetermined type (for example VVC codec type) used for the main image item and alpha image item when present. This may also be deduced from a brand provided in the ISOBMFF file. For example, if one brand is related to VVC codec format the alpha_item (if present), gainmap_item (if present) and main_item have their decoder configuration information defined by CompactVvcDecoderConfigurationRecord and use VVC compression. Same principle may apply for any MPEG codec type. In another embodiment, the compact decoder configuration is used in sample entries of a track, or as an item property associated to one item described in meta box, the step 409 may include similar processing steps (not illustrated) as in steps 404 to 406 to parse the compact decoder configuration. As an alternative representation for compact decoder configuration, one or more control flags may be added in the syntax for decoder configuration to control the number of bits to represent a parameter. As an example, there may be one control flag that when set indicates that the decoder configuration is for a single image, and when not set that it is for multiple images. In this case, there may be a bit allocation depending on the value of this flag. For example, when the flag is set, the parameters corresponding to video are encoded on 0 bit, meaning inferred, and when the flag is set, the number of bits is the one as defined in ISO / IEC 14496-15. There may be control flags applying to compressed parameters (as in step 209) to indicate whether a reduced number of bits is in used or the expected number of bits from ISO / IEC 14496-15. For example, a flag called compact_representation, when set indicates to use the reduced number of bits, and when not set to use the non-compact representation. When used in FullBox, these flags could be values of the “flags” field of the FullBox. This allows to keep the same syntax for decoder configuration, whatever the codec, and makes easier the conversion from compact representation to expected representation from ISO / IEC 14496-15. It is to be noted that even if the examples of decoder configuration mainly mention MPEG video codecs, the invention may apply for other decoder configuration or metadata describing compressed bitstream and how to setup an appropriate decoder. Figure 5 is a schematic block diagram of a computing device 500 for implementation of one or more embodiments of the invention. The computing device 500 may be a device such as a micro-computer, a workstation or a light portable device. The computing device 500 comprises a communication bus connected to: - a central processing unit 501, such as a microprocessor, denoted CPU; - a random access memory 502, denoted RAM, for storing the executable code of the method of embodiments of the invention as well as the registers adapted to record variables and parameters necessary for implementing the method according to embodiments of the invention, the memory capacity thereof can be expanded by an optional RAM connected to an expansion port for example; - a read only memory 503, denoted ROM, for storing computer programs for implementing embodiments of the invention; - a network interface 504 is typically connected to a communication network over which digital data to be processed are transmitted or received. The network interface 504 can be a single network interface, or composed of a set of different network interfaces (for instance wired and wireless interfaces, or different kinds of wired or wireless interfaces). Data packets are written to the network interface for transmission or are read from the network interface for reception under the control of the software application running in the CPU 501; - a graphical user interface 505 may be used for receiving inputs from a user or to display information to a user; - a hard disk 506 denoted HD may be provided as a mass storage device; - an I / O module 507 may be used for receiving / sending data from / to external devices such as a video source or display. The executable code may be stored either in read only memory 503, on the hard disk 506 or on a removable digital medium such as for example a disk. According to a variant, the executable code of the programs can be received by means of a communication network, via the network interface 504, in order to be stored in one of the storage means of the communication device 500, such as the hard disk 506, before being executed. The central processing unit 501 is adapted to control and direct the execution of the instructions or portions of software code of the program or programs according to embodiments of the invention, which instructions are stored in one of the aforementioned storage means. After powering on, the CPU 501 is capable of executing instructions from main RAM memory 502 relating to a software application after those instructions have been loaded from the program ROM 503 or the hard-disc (HD) 506 for example. Such a software application, when executed by the CPU 501, causes the steps of the flowcharts of the invention to be performed. Any step of the algorithms of the invention may be implemented in software by execution of a set of instructions or program by a programmable computing machine, such as a PC (“Personal Computer”), a DSP (“Digital Signal Processor”) or a microcontroller; or else implemented in hardware by a machine or a dedicated component, such as an FPGA (“Field-Programmable Gate Array”) or an ASIC (“Application-Specific Integrated Circuit”). Although the present invention has been described hereinabove with reference to specific embodiments, the present invention is not limited to the specific embodiments, and modifications will be apparent to a skilled person in the art which lie within the scope of the present invention. Many further modifications and variations will suggest themselves to those versed in the art upon making reference to the foregoing illustrative embodiments, which are given by way of example only and which are not intended to limit the scope of the invention, that being determined solely by the appended claims. In particular the different features from different embodiments may be interchanged, where appropriate. Each of the embodiments of the invention described above can be implemented solely or as a combination of a plurality of the embodiments. Also, features from different embodiments can be combined where necessary or where the combination of elements or features from individual embodiments in a single embodiment is beneficial.
Claims
CLAIMS 1. A method for encapsulating an image in a media file, comprising the step of: - generating a media file including an image description and / or image header comprising a set of decoder configuration parameters, based on an encapsulation mode selected from at least a regular description mode and a compact description mode, wherein, with the compact description mode, fewer parameters and a more compact encoding of a parameter value for the set of decoder configuration parameters are used in comparison to the regular description mode.
2. The method according to claim 1, wherein, if the encapsulation mode indicates a compact description mode, a decoder configuration parameter associated with a reserved or predetermined value in the regular description mode is not present in the generated image description and / or image header.
3. The method according to any one of claims 1 or 2, wherein, if the encapsulation mode indicates a compact description mode, a decoder configuration parameter associated with multi-image content in the regular description mode is not present in the generated image description and / or image header.
4. The method according to any one of claims 1 to 3, wherein, if the encapsulation mode indicates a compact description mode, at least one decoder configuration parameter is encoded using a reduced predefined number in bits or bytes in comparison to the regular description mode.
5. The method according to any one of claims 1 to 4, wherein, if the encapsulation mode indicates a compact description mode, the set of decoder configuration parameters comprises a first indication indicating whether the set of decoder configuration parameters contain a length for at least one network abstraction layer unit mandatory for decoding the image.
6. The method according to claim 5, wherein, the set of decoder configuration parameters further comprises a second indication indicating whether the set of decoderconfiguration parameters contain a length for at least one additional network abstraction layer unit not mandatory for decoding the image.
7. The method according to any one of claims 1 to 6, in which at least one decoder configuration parameter is a number of network abstraction layer units present in a configuration record of the image.
8. The method according to any one of claims 1 to 7, wherein the generated media file further includes a parameter representative of the encapsulation mode.
9. The method according to any one of claims 1 to 8, wherein the image description and / or image header is comprised in a minimized image box and the set of decoder configuration parameters is comprised in a decoder configuration record structure in the minimized image box.
10. A method for de-encapsulating an image from a media file, comprising the steps of: - parsing an image description and / or image header comprising a set of decoder configuration parameters, based on determination of an encapsulation mode among one of at least a regular description mode and a compact description mode, wherein, with the compact description mode, fewer parameters and a more compact encoding of a parameter value for the set of decoder configuration parameters are parsed in comparison to the regular description mode.
11. The method according to claim 10, wherein, if the encapsulation mode indicates a compact description mode, the method further comprises the step of processing the set of decoder configuration parameters to obtain a set of decoder configuration parameters according to the regular description mode.
12. The method according to claim 11, wherein the obtained set of decoder configuration parameters according to the regular description mode comprises image dimension parameters determined from image dimension parameters in the image description and / or image header.
13. The method according to any one of claims 10 to 12, wherein the method comprises the step of obtaining, in the media file, a parameter representative of the encapsulation mode.
14. A device for encapsulating an image in a media file, comprising: - means for generating a media file including an image description and / or image header comprising a set of decoder configuration parameters, based on an encapsulation mode selected from at least a regular description mode and a compact description mode, wherein with the compact description mode, fewer parameters and a more compact encoding of a parameter value for the set of decoder configuration parameters are used in comparison to the regular description mode.
15. A device for de-encapsulating an image from a media file, comprising: - means of parsing an image description and / or image header comprising a set of decoder configuration parameters, based on selection of an encapsulation mode from at least a regular description mode and a compact description mode, wherein, with the compact description mode, fewer parameters and a more compact encoding of a parameter value for the set of decoder configuration parameters are parsed in comparison to the regular description mode.
16. A media file encapsulating an image comprising a parameter representative of an encapsulation mode selected from at least a regular description mode and a compact description mode, wherein, the compact description mode is used for generating an image description and / or image header comprising a set of decoder configuration parameters with fewer parameters and a more compact encoding of a parameter values for the set of decoder configuration parameters than a regular description mode.
17. A media file according to claim 16, in which the media file is compliant with an image encapsulation format according to the High Efficiency Image Format (HEIF) standard.
18. A computer program product, characterized in that it comprises instructions which upon execution by a computer cause the computer to execute the method according to any one of claims 1 to 13.
19. A computer-readable storage medium storing programming instructions which upon execution by a computer cause the computer to execute the method according to any one of claims 1 to 13.