Encapsulation method and device, decoding method and device, and storage medium

By generating metadata structures to describe the operation points and synthesis information in the VVC bit stream, the problem that the decoder in the prior art is unable to correctly decode the synthetic video data, and the effective decoding and playback of the VVC bit stream in the ISOBMFF file is realized.

CN114339239BActive Publication Date: 2025-07-11CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111151048.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-29
Filing Date
2021-09-29
Publication Date
2025-07-11
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

When encapsulating the VVC bitstream into an ISOBMFF file, the prior art cannot effectively describe the operation points and synthesis information in the filtered VVC bitstream, resulting in the decoder being unable to correctly decode the synthesized video data.

Method used

By generating metadata structures, including NAL units describing operation points and synthetic information, it is ensured that the decoder can correctly parse operation points and synthetic information in the VVC bitstream, avoiding decoding dependencies on encoding format-specific SEI NAL units.

Benefits of technology

The correct decoding and synthesis information of VVC bitstream is realized, ensuring effective parsing and playing of media files, and improving the compatibility and efficiency of the decoder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114339239B_ABST
    Figure CN114339239B_ABST
Patent Text Reader

Abstract

The present invention provides an encapsulation method and apparatus, a decoding method and apparatus, and a storage medium. It relates to a method of encapsulating a bitstream of encoded video data in a media file, the method comprising: obtaining a bitstream comprising a second plurality of operation points, the video data in the bitstream being organized into NAL units, the bitstream being filtered from an original bitstream comprising a first plurality of operation points, the first plurality of operation points comprising at least the second plurality of operation points; obtaining a first NAL unit from the bitstream that describes the first plurality of operation points; obtaining a second NAL unit from the bitstream that describes the second plurality of operation points; generating a media file comprising the bitstream and comprising at least one metadata structure that describes the second plurality of operation points; wherein the at least one metadata structure is generated based on the first NAL unit and the second NAL unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and apparatus for encapsulating video data into a file. In particular, it relates to encapsulating a VVC bitstream into an ISOBMFF data file. Background Art

[0002] The International Organization for Standardization Base Media File Format (ISO BMFF, ISO / IEC 14496-12) is a well-known flexible and extensible file format that encapsulates and describes encoded timed media data or encoded untimed media data for local storage or for transmission via a network or via other bitstream delivery mechanisms. Examples of extensions are ISO / IEC 14496-15, which describes encapsulation tools for video coding formats based on various NAL (Network Abstraction Layer) units. Examples of such coding formats are AVC (Advanced Video Coding), SVC (Scalable Video Coding), HEVC (High Efficiency Video Coding), L-HEVC (Layered HEVC), and VVC (Versatile Video Coding). Another example of a file format extension is ISO / IEC 23008-12, which describes encapsulation tools for still images or sequences of still images (such as HEVC still images, etc.). Another example of a file format extension is ISO / IEC 23090-2, which defines the Omnidirectional Media Application Format (OMAF). The ISO Base Media File Format is object-oriented. It consists of building blocks called boxes, which correspond to data structures characterized by a unique type identifier (usually a four-character code, also denoted as FourCC or 4CC). A full box is a data structure similar to a box that additionally includes version and flag value attributes. Hereinafter, the term box may specify both all boxes or boxes. These boxes or full boxes are organized hierarchically or sequentially in an ISOBMFF file, and parameters for describing encoded timed media data or encoded untimed media data, their structure, and timing (if any) are defined. All data (media data and metadata describing the media data) in the encapsulated media file is contained in boxes. There is no other data within the file. A file-level box is a box that is not contained within other boxes.

[0003] In the file format, an entire media presentation is called a movie. The movie is described by a movie box (with four - character code'moov') at the top level of the file. This movie box represents an initialization information container that holds a collection of various boxes describing the media presentation. The movie box is logically divided into tracks, represented by track boxes (with four - character code 'trak'). Each track (uniquely identified by a track_ID) represents a timed sequence of media data belonging to the presentation (e.g., frames of video or audio samples). Within each track, each timed unit of data is called a sample; this could be a frame of video, audio, or timed metadata. Samples are implicitly numbered in the decoding order sequence. Each track box contains a hierarchical structure of boxes that describe the samples of the track. For example, the sample table box ('stbl') contains all the time and data indices of the media samples in the track. The actual sample data is stored in a box called the media data box (with four - character code'mdat') or the identified media data box (with four - character code 'imda', similar to the media data box but containing additional identifiers) at the same level as the movie box. A movie can also be segmented in time and organized into a movie box that contains information for the entire presentation, followed by a list of media segments, i.e., a list that couples movie segments and media data boxes ('mdat' or 'imda'). Within a movie segment (a box with four - character code'moof'), there is a set of track segments (boxes with four - character code 'traf') that describe the tracks within the media segment, zero or more for each movie segment. A track segment in turn contains zero or more track run boxes ("trun"), and each track run box records the consecutive runs of samples of that track segment.

[0004] An ISOBMFF file can contain multiple encoded timed media data that form multiple tracks, or sub - parts of the encoded timed media data. When the sub - parts correspond to one or consecutive spatial parts of a video source taken over time (e.g., at least one rectangular region taken over time, sometimes called a 'tile' or'sub - picture'), the corresponding multiple tracks can be called tile tracks or sub - picture tracks. When the bitstream is a hierarchical bitstream (e.g., a scalability layer in SVC, HEVC, or VVC, or a layer for multi - view in MVC, HEVC, or VVC, or an independent layer as in VVC), one or more layers can be encapsulated into one track. These tracks can correspond to spatial parts of the video source.

[0005] The compression of video relies on block - based video coding in most coding systems (such as the HEVC standard, which stands for High Efficiency Video Coding; or the emerging VVC standard, which stands for Versatile Video Coding). This document focuses on (but is not limited to) VVC - coded bitstreams. In these coding systems, a video is composed of a sequence of frames or pictures or images or samples that can be displayed at several different times. In the case of multi - layer video (e.g., scalable, stereoscopic, 3D video), several frames can be decoded to synthesize a resulting image to be displayed at one moment.

[0006] According to the VVC standard, video data is organized into Network Abstraction Layer units or NAL units or NALUs. A NAL unit is a logical unit of data for encapsulating data into an encoded bitstream. NAL units are classified into VCL and non - VCL NAL units. A VCL NAL unit contains data representing the values of samples in a video picture, and a non - VCL NAL unit contains any associated additional information (such as parameter sets (important header data applicable to a large number of VCL NAL units), etc.) and supplementary enhancement information (timing information, and other supplementary data that can enhance the usability of the decoded video signal but is not necessary for decoding the values of samples in a video picture).

[0007] A layer corresponds to a set of VCL NAL units and associated non - VCL NAL units all having a specific value of nuh_layer_id. In VVC, a layer can be independent and then corresponds to all the necessary data for decoding a video, or the decoding of a given layer may require some data from other layers, which means that decoding the video corresponding to a layer may require decoding several related layers. In the latter case, these layers are said to be dependent.

[0008] In a VVC bitstream, layers are organized into an Output Layer Set (OLS) corresponding to a set of related layers. In an OLS, at least one layer must be marked as an output layer. An output layer is the layer intended to be output by the decoder. At decoding, the decoder selects the OLS and must decode all the output layers in the OLS to produce the output image of the output video sequence. There can be several output layers in an OLS. For example, this is the case when the output corresponds to a stereoscopic video sequence synthesized from two different video sequences (right video sequence and left video sequence). When the OLS includes several independent layers, these layers can all be marked as output layers.

[0009] Layers can be organized into sub-layers corresponding to the temporal scalability coding of the corresponding video data. In this case, the base sub-layer corresponds to the lowest temporal video data, e.g., a 30 Hz video sequence, while successive sub-layers allow for higher temporal versions of the decoded video, e.g., 60 Hz or 120 Hz. Each sub-layer is identified by a TemporalId (in the NAL unit header) indicating the hierarchical structure of the sub-layers in the temporal scalability coding scheme.

[0010] An operation point (OP) (also referred to as an operating point, the two terms being equivalent) corresponds to a temporal subset of an OLS, which is identified by an OLS index identifying the associated OLS. Each operation point is associated with a set of output layers, a maximum TemporalId value, and signaling of the profile, level, and layer. The highest value of the operation point and TemporalId can be identified. The operation point can thus be considered as an OLS that may be restricted in terms of the temporal solution.

[0011] In the VVC bitstream, a specific non-VCL NAL unit called the Video Parameter Set (VPS) provides information about the organization of the bitstream. Specifically, it describes the different layers, sub-layers, their dependencies, their organization in the OLS, and the list of operation points present in the bitstream. Some SEI messages can also provide further information about a specific OLS, a specific layer, or a specific set of sub-pictures, e.g., about scalable nested SEI messages.

[0012] When encapsulating the VVC bitstream into an ISOBMFF file, the different layers are encapsulated into one or several tracks reflecting the organization of the bitstream. Some structures present in the metadata section of the ISOBMFF file are used to describe the organization of the bitstream encapsulated into the file. Specifically, a box called 'vopi' is used to describe the different operation points in the bitstream, a box called 'linf' is used to describe the layers and sub-layers present in a given track, and a box called 'opeg' describes the mapping of tracks to operation points and profile level information. The content of these boxes is derived from the VPS NAL unit in the VVC bitstream. While 'vopi' and 'linf' apply to sample groups, i.e., they can change along the bitstream, the 'opeg' structure is static, i.e., it is a single instance for the entire bitstream. For a given bitstream, there is at most one track carrying the 'vopi' sample group. The 'vopi' information applies to sample groups from all tracks that reference this at most one track using the 'oref' track reference type. The 'linf' structure describes, for a given track, the list of layers and sub-layers carried by that track.

[0013] The VVC bitstream may be filtered. This means that the original VVC bitstream subject to the filtering process results in a filtered (or restricted) VVC bitstream in which some of the original layers and operation points or some sub - layers have been suppressed. An optional non - VCL NAL unit called OPI (standing for Operation Point Information) will be provided to the filtered VVC bitstream, which describes the operation points actually present in the filtered VVC bitstream. Thus, the presence of an OPI NAL unit in the VVC bitstream indicates first that the VVC bitstream is a filtered (or restricted) bitstream, and second provides information related to the operation points actually present in the filtered VVC bitstream. It should be noted that the VPS NAL unit of the filtered VVC bitstream still provides information related to the organization of the original VVC bitstream that describes all operation points (even those operation points that have been filtered out (or restricted) in the VVC filtered bitstream). Some specific restrictions or filtering can result in a filtered or restricted set of operation points that actually matches the initial set of operation points. In the first example, it can be the case that when the OPI NAL unit indicates an ols_idx of the layer with the highest layer ID in the bitstream (the layer having dependencies on other layers) and when the OPI NAL unit indicates the maximum temporal ID present in the bitstream as the highest temporal ID. In the second example, it can be the case that when the OPI NAL unit indicates an ols_idx of all layers of the bitstream and when the OPI NAL unit indicates the maximum temporal ID present in the bitstream as the highest temporal ID.

[0014] Thus, the encapsulation of the filtered VVC bitstream in an ISOBMFF file results in a metadata section that provides a description of the original VVC bitstream that describes some operation points and associated layers that may no longer be present in the file. Additionally, the metadata section providing a description of the original VVC bitstream may describe sub - layers that may no longer be present in the file. This may pose a problem for an ISOBMFF parser that should not be decoding NAL units and manipulating the file based on this metadata.

[0015] The VVC bitstream can correspond to synthesis. The synthesis bitstream can include at least two independent layers marked as output layers. Upon decoding, the at least two independent layers will be decoded and can be synthesized to generate an output image of the output image video sequence. To be able to synthesize the output image, the decoder needs information describing the synthesis. In particular, information such as the size of the synthesis and the position of each independent layer in the synthesis is needed. According to VVC, the synthesis information of the independent layers of the VVC bitstream can be provided in the bitstream as an SEI message. The SEI message is an optional non - VCL NAL unit, and the decoder may ignore these SEI NAL units during decoding. A decoder that ignores these SEI NAL units may not be able to render the synthesis.

[0016] When encapsulating a VVC bitstream corresponding to a composition into an ISOBMFF data file, there is no metadata structure that allows the description of the composition. The parser has to decode SEI NAL units to understand the composition and related parameters. It is advantageous to provide a description of the composition independent of the coding format. Such a description can be placed in the metadata section of the media file to allow the parser or reader to understand the composition without having to decode the coding format-specific SEI NAL units of the encapsulated bitstream. Summary of the Invention

[0017] The present invention aims to solve one or more of the foregoing problems.

[0018] According to a first aspect of the present invention, there is provided a method of encapsulating a bitstream of encoded video data in a media file, the method comprising:

[0019] obtaining a bitstream comprising a second plurality of operation points, wherein the video data in the bitstream is organized into NAL units; the bitstream is filtered from an original bitstream comprising a first plurality of operation points, the first plurality of operation points at least comprising the second plurality of operation points;

[0020] obtaining a first NAL unit from the bitstream, the first NAL unit describing the first plurality of operation points;

[0021] obtaining a second NAL unit from the bitstream, the second NAL unit describing the second plurality of operation points;

[0022] generating a media file comprising the bitstream and comprising at least one metadata structure describing the second plurality of operation points;

[0023] wherein the at least one metadata structure is generated based on both the first NAL unit and the second NAL unit.

[0024] In an embodiment, the first NAL unit is a video parameter set NAL unit.

[0025] In an embodiment, the second NAL unit is an operation point information NAL unit.

[0026] In an embodiment, the at least one metadata structure comprises a first metadata structure describing the first plurality of operation points and a different second metadata structure describing the second plurality of operation points.

[0027] In an embodiment, the at least one metadata structure comprises a metadata structure describing the second plurality of operation points and there is no metadata structure describing the first plurality of operation points.

[0028] In an embodiment, the at least one metadata structure includes a metadata structure describing the first plurality of operation points and signaling information indicating operation points belonging to the second plurality of operation points.

[0029] In an embodiment, the at least one metadata structure is included within metadata describing a dedicated track including the second plurality of operation points.

[0030] According to another aspect of the present invention, there is provided a method of encapsulating a bitstream of encoded video data in a media file, the method comprising:

[0031] obtaining a bitstream including a plurality of operation points, the plurality of operation points being associated with at least two independent layers of video data related to composition, the video data in the bitstream being organized into NAL units;

[0032] obtaining from the bitstream NAL units describing composition information applied to at least two independent layers of the video data;

[0033] generating a media file including the bitstream and at least one metadata structure describing the composition.

[0034] In an embodiment, the at least one metadata structure includes a structure dedicated to the description of each independent layer of the composition as a rectangular region.

[0035] In an embodiment, the at least one metadata structure includes a structure dedicated to the description of the layer, wherein the description of the composed layer includes the composition information.

[0036] In an embodiment, the at least one metadata structure further includes a transform for a part of the at least two independent layers to which the transform is to be applied.

[0037] In an embodiment, the at least one metadata structure describing the composition includes a transformation matrix.

[0038] In an embodiment, the at least one metadata structure describing the composition is associated with an output layer set including the at least two independent layers of the composition.

[0039] According to another aspect of the present invention, there is provided a method of encapsulating a bitstream of encoded video data in a media file, the method comprising:

[0040] obtaining a bitstream including a second plurality of operation points, the video data in the bitstream being organized into NAL units; the bitstream is filtered from an original bitstream including a first plurality of operation points, the first plurality of operation points at least including the second plurality of operation points; the second plurality of operation points being associated with at least two independent layers of video data related to composition;

[0041] Obtain a first NAL unit from the bitstream, the first NAL unit describing the first plurality of operating points;

[0042] Obtain a second NAL unit from the bitstream, the second NAL unit describing the second plurality of operating points;

[0043] Obtain a third NAL unit from the bitstream, the third NAL unit describing the synthesis information applied to at least two independent layers of video data;

[0044] Generate the media file, the media file including the bitstream and at least one metadata structure describing the second plurality of operating points and the synthesis;

[0045] wherein the at least one metadata structure is generated based on both the first NAL unit and the second NAL unit.

[0046] According to another aspect of the present invention, there is provided a method for decoding a media file including a bitstream of encoded video data, the method comprising:

[0047] Obtain the bitstream including the second plurality of operating points from the media file, the video data in the bitstream being organized into NAL units; the bitstream is filtered from an original bitstream including a first plurality of operating points, the first plurality of operating points at least including the second plurality of operating points;

[0048] Obtain at least one metadata structure describing the second plurality of operating points from the media file;

[0049] Select at least one layer associated with the second plurality of operating points;

[0050] Decode the selected layer of the video data.

[0051] According to another aspect of the present invention, there is provided a method for decoding a media file including a bitstream of encoded video data, the method comprising:

[0052] Obtain the bitstream including a plurality of operating points from the media file, the plurality of operating points being associated with at least two independent layers of video data related to synthesis, the video data in the bitstream being organized into NAL units;

[0053] Obtain at least one metadata structure describing the synthesis from the media file;

[0054] Decode at least two independent layers related to synthesis and generate a video sequence based on the synthesis information.

[0055] According to another aspect of the present invention, there is provided a method for decoding a media file of a bitstream including encoded video data, the method comprising:

[0056] Obtaining the bitstream including a second plurality of operation points from the media file, wherein the video data in the bitstream is organized into NAL units; the bitstream is filtered from an original bitstream including a first plurality of operation points, the first plurality of operation points at least including the second plurality of operation points; the second plurality of operation points is associated with at least two independent layers of video data related to synthesis;

[0057] Obtaining at least one metadata structure describing the second plurality of operation points from the media file;

[0058] Obtaining at least one metadata structure describing the synthesis from the media file;

[0059] Decoding the at least two independent layers related to the synthesis and generating a video sequence based on the synthesis information.

[0060] According to another aspect of the present invention, there is provided a computer program product for a programmable device, the computer program product including a sequence of instructions which, when loaded into and executed by the programmable device, implement the method according to the present invention.

[0061] According to another aspect of the present invention, there is provided a computer-readable storage medium storing instructions of a computer program for implementing the method according to the present invention.

[0062] According to another aspect of the present invention, there is provided an apparatus for encapsulating a bitstream of encoded video data in a media file, the apparatus including a processor configured to:

[0063] Obtain a bitstream including a second plurality of operation points, wherein the video data in the bitstream is organized into NAL units; the bitstream is filtered from an original bitstream including a first plurality of operation points, the first plurality of operation points at least including the second plurality of operation points;

[0064] Obtain a first NAL unit from the bitstream, the first NAL unit describing the first plurality of operation points;

[0065] Obtain a second NAL unit from the bitstream, the second NAL unit describing the second plurality of operation points;

[0066] Generate a media file including the bitstream and including at least one metadata structure describing the second plurality of operation points;

[0067] Wherein, the at least one metadata structure is generated based on both the first NAL unit and the second NAL unit.

[0068] According to another aspect of the present invention, there is provided an apparatus for encapsulating a bitstream of encoded video data in a media file, the apparatus comprising a processor configured to:

[0069] Obtain a bitstream including a plurality of operation points, the plurality of operation points being associated with at least two independent layers of video data related to synthesis, and the video data in the bitstream being organized into NAL units;

[0070] Obtain from the bitstream NAL units that describe synthesis information applied to at least two independent layers of video data;

[0071] Generate a media file including the bitstream and at least one metadata structure describing the synthesis.

[0072] According to another aspect of the present invention, there is provided an apparatus for encapsulating a bitstream of encoded video data in a media file, the apparatus comprising a processor configured to:

[0073] Obtain a bitstream including a second plurality of operation points, the video data in the bitstream being organized into NAL units; the bitstream is filtered from an original bitstream that includes a first plurality of operation points, the first plurality of operation points at least including the second plurality of operation points; the second plurality of operation points are associated with at least two independent layers of video data related to synthesis;

[0074] Obtain from the bitstream a first NAL unit that describes the first plurality of operation points;

[0075] Obtain from the bitstream a second NAL unit that describes the second plurality of operation points;

[0076] Obtain from the bitstream a third NAL unit that describes synthesis information applied to at least two independent layers of video data;

[0077] Generate the media file, the media file including the bitstream and at least one metadata structure describing the second plurality of operation points and the synthesis;

[0078] Wherein, the at least one metadata structure is generated based on both the first NAL unit and the second NAL unit.

[0079] The method according to the present invention can be implemented by a computer. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects, all of which are generally referred to herein as "circuits", "modules" or "systems". In addition, the present invention can take the form of a computer program product embodied in any tangible expression medium having computer-usable program code embodied therein.

[0080] Since the present invention can be implemented in software, the present invention can be embodied as computer-readable code provided to a programmable device on any suitable carrier medium. Tangible, non-transitory carrier media can include storage media such as floppy disks, CD-ROMs, hard disk drives, tape devices or solid state storage devices. Transient carrier media can include signals such as electrical, electronic, optical, acoustic, magnetic or electromagnetic signals (e.g., microwave or RF signals). BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Embodiments of the present invention will now be described by way of example only and with reference to the following drawings, in which:

[0082] Figure 1a and 1b show an example of a bitstream having multiple layers;

[0083] Figure 2 show a high-level view of an embodiment for encapsulating a filtered VVC bitstream in a media file;

[0084] Figure 3 show how to create a 'vopi' sample group according to an embodiment of the present invention;

[0085] Figure 4 show an example of a VVC bitstream;

[0086] Figure 5 show an example of an encapsulation process according to an embodiment of the present invention;

[0087] Figure 6 show an example of the main steps of a parsing process according to an embodiment of the present invention;

[0088] Figure 7 show an example of an ISOBMFF file according to an embodiment of the present invention;

[0089] Figure 8 is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. DETAILED DESCRIPTION

[0090] In a first aspect of the present invention, a method of encapsulating a filtered VVC bitstream is provided, the method generating an ISOBMFF file, wherein the metadata description of the bitstream is made to conform to the actual organization of the bitstream.

[0091] Figure 1a and 1b shows an example of a bitstream having multiple layers.

[0092] Figure 1a shows an example of a video bitstream, wherein video frames can be encoded into one or more independent layers (100, 110, and 120). Each layer can correspond, for example, to an alternative quality of the same video, or to different regions (e.g., spatial parts) of the same video. Each layer can contain one or more temporal sublayers (101, 102 or 111, 112 or 121, 122), for example, one temporal sublayer contains 30 frames per second and a second temporal sublayer allows 60 frames per second. In Figure 1a the example, the bitstream can be encoded with an indication, for example, of 3 output layer sets (OLSs) (one per layer) available in a video parameter set NAL unit. The VPS allows implicit indexing of the OLSs, as shown in 105, where the list of layers is included in each OLS. A decoder receiving such a bitstream with multiple OLSs must perform OLS selection to start the decoding process. When using VVC as the encoding format, the VVC specification defines the rules for a decoder to select such an OLS (decoded as a pair of a target output layer set and the highest temporal Id). The first rule relies on external means. For example, the application controlling the decoder or the user interface or configuration file controlling the decoder provides these two parameters. The second rule relies on an indication in the VVC bitstream provided by an operation point information NAL unit (OPI NALU), which can indicate the index of the output layer set and / or the value of the highest temporal ID for an encoded video sequence. The last rule is the default rule, which is standardized by VVC to select one OLS from the list of OLSs declared in the VPS NAL unit based on the number of layers and the output layer. This last rule only applies to the first access unit in the bitstream, while the two previous rules apply to the first access unit of the encoded video sequence. An access unit (AU) is defined as a set of picture units belonging to different layers and containing encoded pictures associated with the same time output from the decoded picture buffer.

[0093] Figure 1b shows another example of a video bitstream, wherein video frames can be encoded into one or more layers (150, 160, 170). There may be dependencies between the layers, as indicated by the arrows (171, 172) on the left. In VVC (also for HEVC), the layer dependencies are indicated in the VPS NAL unit. From the declarations in the VPS, it is possible to Figure 1bThe configuration on derives three sets of output layers (155). In VVC, OLSs are selected for decoding and follow the same three rules as described according to Figure 1a Each set of output layers in this example will contain (or refer to) other layers. For example, the OLS with index 3 contains layer 2, but also contains layers 1 and 0 due to dependencies. The dependencies come from the coding scheme, which may use reference frames from lower layers, for example. In Figure 1b we can see that each layer has a different number of sub-layers. As an example, the lowest sub-layer 151 contains NAL units with a temporal ID value equal to 0, while sub-layer 152 contains NAL units with a temporal ID equal to 1. Of course, as in Figure 1b layers may have the same number of sub-layers. Additionally, there may be layer references different from those shown by 171 and 172. For example, both layer 2 and layer 1 may only depend on layer 0, or only one of layer 1 or 2 may have dependencies on other layers and the other layer is independent.

[0094] Independent of Figure 1a or Figure 1b , the video bitstream may have a layer or sub-layer configuration that changes over time. This can be indicated by a new VPS NAL unit. Additionally, the bitstream may contain several OPI NAL units over time. The new OPI NAL units may appear when a new VPS occurs, but may also appear without indicating a new VPS, for example, to indicate that a given OLS is more relevant or recommended compared to other OLSs at a certain time interval, or that new filtering or restrictions apply to the bitstream.

[0095] It should be noted that the HEVC coding specification does not contain a NAL unit corresponding to the VVC OPI, even though it allows for the description of bitstreams with multiple sets of output layers. In HEVC, SEI non-VCL NAL units will be used to signal the filtered bitstream because the OPI non-VCL NAL unit is not defined for HEVC. Except for this difference, the proposed method can be used in the context of HEVC by replacing the OPI NAL unit with an SEI NAL unit, which may include equivalent filtering information related to the OPI actually present in the filtered bitstream. In the case of this mechanism in HEVC, the encapsulation of the filtered or restricted operating point list in the HEVC bitstream can be completed according to the embodiments described below.

[0096] Figure 2Shows a high - level view of the proposed method for encapsulating a filtered VVC bitstream in a media file. As shown, server 200 includes an encapsulation module 205 that is connected to communication network 210 via a network interface (not shown). The de - encapsulation module 215 of client 220 is also connected to the communication network 210 via a network interface (not shown).

[0097] Server 200 processes data (such as video and / or audio data) for streaming or for storage. To this end, server 200 obtains or receives data (referred to as source video) including, for example, scenes recorded by one or more cameras. The source video is received by the server as a raw picture sequence 225. The server uses a media encoder (e.g., a video encoder) (not shown) to encode the picture sequence in media data (i.e., a bitstream), and uses the encapsulation module 205 to encapsulate the media data into one or more media files or media segments 230. The media data or bitstream may contain multiple picture sequences 225. The encapsulation module 205 includes at least one of a writer or an encapsulator for encapsulating media data. The media encoder may be implemented within the encapsulation module 205 to encode the received data or may be separate from the encapsulation module 205. Once the picture sequence 225 is generated (real - time encoding) or offline, the encoding can be completed. The encapsulation can also be done in real - time or offline. The encapsulation can also include re - encapsulating a first media file 230 into a second media file. Here, the media file may include a single physical file or multiple media segment files or moving image segment files. Examples of re - encapsulation may include segmenting a media file, or conversely merging media segments into one media file. Another example of re - encapsulation may include filtering or restricting the operation points, tracks, layers, or sub - layers of a media file. It may also include editing the media with new tracks, layers, or operation points. Re - encapsulation means that instead of the picture sequence 225 as input, the encapsulation module 205 takes a first media file as input and produces a second media file 230 with different characteristics. The application of the embodiments below may include generating a media file as the second media file in which NAL units for some operation points have been removed from the first media file. To indicate this removal, filtering, or restriction, the proposed embodiments can be used in the second media file.

[0098] The client 220 is used to process data received from the communication network 210, for example, to process the media file 230. After the received data is unpacked in the unpacking module 215 (also called the parser), the unpacked data (or parsed data) corresponding to the media data bitstream is decoded to form audio and / or video data that can be stored, displayed, or output, for example. The media decoder can be implemented within the unpacking module 215 or can be separated from the unpacking module 215. The media decoder can be configured to decode one or more video bitstreams in parallel. For encoding and encapsulation, decoding and / or parsing can be done offline or in real time. The client or the server can be a user device, but can also be a network node that acts on the media file to be sent or stored.

[0099] It should be noted that the media file 230 can be communicated to the unpacking module 215 in different ways. Specifically, the encapsulation module 205 can generate a media file 230 with a media description (such as DASH MPD) and communicate (or stream) it directly to the unpacking module 215 when a request is received from the client 220. The media file 230 can also be downloaded by the client 220 and stored on the client 220.

[0100] For illustration purposes, the media file 230 can encapsulate media data (such as encoded audio or video) into boxes according to the ISO Base Media File Format (ISOBMFF, ISO / IEC 14496-12, and ISO / IEC 14496-15 standards). In this case, the media file 230 can correspond to one or more media files (indicated by the FileTypeBox 'ftyp'). According to ISOBMFF, the media file 230 can include two types of boxes, namely, "media data boxes" identified as'mdat' or 'imda' that contain media data, and "metadata boxes" (such as'moov' or'moof') that contain metadata defining the location and timing of the media data. In a specific embodiment, the picture sequence 225 is encoded or compressed according to the General Video Codec Specification ISO / IEC 23090-3 or the High Efficiency Video Codec Specification 23008-2.

[0101] Although the output layer sets, sub-layers, and their dependencies are described at the bitstream level in the VPS NALU, in the VVC file format, different structures allow the description of the operation points or output layer sets of the VVC bitstream and their association with layers ('vopi') or tracks ('opeg').

[0102] According to ISO / IEC 14496-15, "the storage of the VVC bitstream is supported by structures such as […] 'vopi', 'linf' or 'opeg' structures". These structures describe different operation points ('vopi'), the list of layers and sublayers carried in a track ('linf'), or the mapping of a track to operation points and profile-level information ('opeg') based on the information about layers and sublayers indicated in the VPS NAL unit of the VVC bitstream. In the case where the VVC bitstream contains an OPI NAL unit, this indicates that some NAL unit filtering (or restriction) has been done. Then, there may be a greater set of output layers or operation points described in the VPS NAL unit compared to what actually exists in the VVC bitstream (or encapsulated file). The structure for operation point description in ISO / IEC 14496-15 does not take into account this OPI NAL unit. When present in the bitstream and not considered in the encapsulation, there is a risk of exposing to the player or reader a set of output layers or operation points greater than what is actually possible in the media file (because layers or sublayers may have been removed when considering the specific OLS or sublayers indicated in the OPI NAL unit). Then, the question is about the "reliability" or "credibility" of the information contained in the structure that describes the operation points.

[0103] Then, in a first embodiment, we propose an additional structure for encapsulating the operation point information of multiple VVC layers. The rough idea is to ensure that the encapsulation does not describe operation points or sublayers for which the NAL units are no longer available through the 'linf', 'vopi' or 'opg' structures (also called descriptors). Adding the new structure does not modify the syntax of the mentioned file format structures. A second embodiment that avoids using other structures includes, for example, clarifying the semantics of the descriptor for operation points to take into account the OPI NALU by indicating that num_operating_points should correspond to what is derived from the OPI and VPS NALUs and not necessarily only from the VPS NALU, and that it must correspond to the number of operation points actually present in the bitstream encapsulated into the data file.

[0104] A third embodiment includes annotating the descriptor (or structure) that describes the operation points to indicate whether their corresponding NAL units are still available in the encapsulated media file.

[0105] New structure for filtered or recommended operation points

[0106] For VPS, there may be more than one OPI NAL unit in the VVC bitstream. As such, similar to 'vopi', the proposed structure is for handling potentially changing sample groups over time. In contrast, 'opeg' is a static structure that does not handle OPI changes well. Considering this, the 'vopi' structure is superior to the 'opeg' structure whenever the bitstream contains more than one OPI NAL unit, especially for real-time packaging and real-time encoding.

[0107] In a first variation, new sample groups are proposed to describe the recommended, restricted, or filtered output layer sets (or operating points) in a list of output layer sets (or operating points). A specific VisualSampleGroupEntry is defined to provide an indication of the index of the output layer set available in the set of operating points defined in the 'vopi' sample group. This VisualSampleGroupEntry is used in the SampleGroupDescriptionBox where the grouping_type is equal to 'ropi' (the four-character code here is just an example, and any reserved value that does not conflict with the existing 4ccs can be used).

[0108] Notify the application of the recommended, restricted, or filtered operating points provided by a given VVC bitstream by using the restricted operating point information sample group ('ropi'). The recommended, restricted, or filtered operating points are indicated as output layer set indices, possibly with an indication of the maximum TemporalId value. The index refers to one of the output layer set indices declared in the 'vopi' sample group. This sample group is associated with the 'vopi' sample group and restricts the set of possible operating points declared in the 'vopi' sample group. The information contained in this sample group can also provide the maximum temporal Id within the restricted operating points. Preferably, in multi-track packaging, the 'ropi' sample group is included in the track referenced by other tracks via the 'oref' track reference type. In the case where the media file contains more than one presentation or packages more than one VVC bitstream (each with its own set of operating points), there may be several 'ropi' sample groups in the media file, but each is related to the reference track that provides the 'vopi' or is indicated by a specific track reference type such as 'oref'.

[0109] The 'ropi' sample group description for a given sample group always indicates the restrictions applied to the operating points declared in the 'vopi' sample group description associated with that sample group. For example, there may be default sample groupings for both 'vopi' and 'ropi', in which case all track samples containing 'ropi' and all track samples referencing the tracks containing the 'ropi' sample group follow the same list of filtering or restrictions on the operating points. As another example, there may be a default 'vopi' sample grouping (the same list of operating points for all track samples), but several 'ropi' entries are listed in a SampleGroupDescriptionBox with grouping_type equal to 'ropi', thus indicating different restrictions are applied over time from one sample group to another. Another example includes 'vopi' and 'ropi' sample groups that are both non-default sample groupings. For example, a media file includes a program composed of moving images, advertisements, and TV shows. Different operating points are defined for each piece of content (moving image, advertisement, or TV show), and each operating point can be further constrained by an OPI NALU. In this example, there may be several sample groups associated with different 'vopi' sample group description entries and several sample groups associated with different 'ropi' sample group description entries.

[0110] The proposed syntax for the new VisualSampleGroupEntry is:

[0111] class RestrictedOperatingPointsInformation extendsVisualSampleGroupEntry('ropi'){

[0112] unsigned int(16) restricted_output_layer_set_idx;

[0113] unsigned int(3) restricted_highest_temporalID;

[0114] unsigned int(5) reserved;

[0115] }

[0116] Wherein, it has the following semantics:

[0117] The restricted_output_layer_set_idx is the zero-based index of the output layer set corresponding to the output layer set index (opi_ols_idx) indicated in the OPI NAL unit as defined in the VVC specification ISO / IEC 23090-3. A reserved or predefined value indicates that the restriction on the output layer set is not applied to the corresponding sample group. For example, this predefined or reserved value is set to the maximum value of restricted_output_layer_set_idx (or any higher value), i.e., one plus the maximum number of operating points. In a variant, this value is equal to two to the power of the syntax element length in bits minus one (2^len–1), where, in the above syntax example, len is equal to 16.

[0118] The restricted_highest_temporalID is the value of the maximum temporal ID present in the operating point, as indicated in the OPI NAL unit as defined in the VVC specification ISO / IEC 23090-3. If opi_htid_plus1 is greater than 0, its value is equal to opi_htid_plus1 minus one. Otherwise, the reserved or predefined value opi_htid_plus1 = 0 means that the restriction on the sublayer is not applied to the corresponding sample group. This value also provides the upper bound for the value max_TemporalId in the 'linf' sample group description entry associated with the sample group that defines such a restriction on the sublayer. The affected 'linf' sample group description entries are the sample group description entries contained in the track that references the track containing the 'rop' sample group by a specific track reference type (e.g., 'oref').

[0119] Although the above sample group description entries are simple, since the values from the OPI NAL unit are reflected at the file format level, default values must be set when the OPI NAL unit provides only one parameter (opi_ols_idx or opi_htid_plus1). For example, when only opi_htid_plus1 is provided in the OPI NALU, a predefined value is set in restricted_output_layer_set_idx to indicate no restriction on the output layer set. Similarly, when only opi_ols_idx is provided in the OPI NALU, a predefined value is set in restricted_highest_temporalID to indicate no restriction on the sublayer.

[0120] In a second variant, the structure of the visual sample group entry precisely reflects the payload of the OPI NALU, where control flags are used to indicate whether a restriction is applied to the layer or the sublayer or both.

[0121] The proposed syntax is:

[0122]

[0123]

[0124] Among them, it has the following semantics:

[0125] When output_layer_set_restriction_flag is set, it indicates that the restriction on the output layer set defined in the 'vopi' sample group is applied to the sample group. When it is not set, no restriction is applied to the sample group and any output layer set can be selected by the reader or parser for transmission, reconstruction, or display.

[0126] When subsub_restriction_flag is set, it indicates that the restriction on the temporal sublayer is defined for the sample group. When it is not set, no restriction is applied to the temporal sublayer of the list of operation points defined in the 'vopi' sample group, and any temporal layer can be selected by the analyzer or reader for transmission, reconstruction, or display.

[0127] restricted_output_layer_set_idx indicates the index of the output layer set from the list of operation points declared in 'vopi' available in the NAL unit in the media file. The parser or reader can select this output layer set for transmission, rendering, or display. When present in the VVC bitstream, it can correspond to the output layer set index indicated in the opi_ols_idx OPI NAL unit defined in the VVC specification ISO / IEC 23090-3.

[0128] restricted_highest_temporalID indicates the value of the maximum temporal ID allowed among the operation points listed in the 'vopi' sample group. It can correspond to the value indicated in the OPI NAL unit defined in the VVC specification ISO / IEC23090-3: if opi_htid_plus1 is greater than 0, the value of opi_htid_plus1 minus one. This value also provides the upper limit of the value max_TemporalId in the layer information sample group description entry ('linf') associated with the sample group that defines the restriction on the sublayer.

[0129] In the third variation, a new sample group is created to describe a list of restricted operating points. With this variation, the 'ropi' sample group no longer describes with reference to the 'vopi' sample group, but rather rewrites or replaces the operating point description of the 'vopi' sample group. This has implications for the parser and reader. When present in the media file, instead of the 'vopi' sample group, the 'ropi' sample group should be considered by the parser or reader to select the set of output layers to send or play. According to this embodiment, when both 'vopi' and 'ropi' coexist in the media file, 'vopi' should be considered unreliable and 'ropi' should be used instead.

[0130] It should be noted that most of the parameters in this sample group description entry are the same as those in 'vopi'. Only those in bold correspond to changes in the syntax and semantics of 'vopi'. In addition, some useless fields from 'vopi' have been removed (struck-through fields). In fact, repeating these fields from the VPS does not accomplish anything, since each operating point is described in detail below. It should be noted that following the same principle, the 'vopi' structure can also be simplified.

[0131]

[0132]

[0133] where num_restricted_operating_points gives the number of operating points that the information follows. The number of operating points can correspond to the number of sets of output layers derived from the VPS and OPI NAL units.

[0134] restricted_output_layer_set_idx is the index of the set of output layers that defines the operating point. After the filtering or restriction indicated in the OPI NAL unit for the set of output layers with index output_layer_set_idx, the mapping between output_layer_set_idx and the layer_id value should be the same as that specified by the VPS.

[0135] max_temporal_id gives the maximum TemporalId of the NAL unit of the operation point as indicated in the VPS after the filtering or restriction indicated in the OPI NAL unit. When no filtering or restriction is indicated in the OPI NAL unit (opi_htid_info_present_flag = 0 or opi_htid_plus1 = 0), it corresponds to the maximum temporal ID derived from the VPS. This value should correspond to opi_htid_plus1 minus one. This value also provides the upper limit of the value max_TemporalId in the 'linf' sample group description entry associated with the sample group that defines such a restriction for the sublayer.

[0136] Compared with the 'vopi' structure, the semantics of other fields remain unchanged. According to the second embodiment, this sample group is an alternative to the 'vopi' constructed by considering the VPS and the NAL unit. However, this embodiment makes media file editing easier by adding this new sample group every time an OPI NALU occurs in the bitstream and allows the 'vopi' structure to remain aligned with the VPS NAL unit.

[0137] For different variations in this first embodiment, when an existing moving picture file is filtered by an application that provides a set of target output layers and / or a target maximum temporal ID, a new set of samples for a restricted operation point can also be used. In cases where the original 'vopi' set of samples and the 'ropi' set of samples are used to indicate filtering or restrictions applied to a media file, the filtered media file can be stored without filtered NAL units having layer IDs not in the set of target output layers and / or without NAL units having temporal IDs greater than the target maximum temporal ID. An advanced parser can use such an indication to construct a VPS NALU that describes the filtered NAL units and provide it to the video decoder. A less advanced parser can push only the initial VPS plus an OPI NAL unit indicating the filtering to the video decoder. Some applications, such as a media file checker or a robust parser, can check the consistency of the media file based on the consistency between the operation point description and the actual NAL units present in the media file. For example, if the maximum temporal ID derivable by parsing the NAL unit headers is lower than the maximum temporal ID declared for any operation point in the file, the file checker can generate a structure for a restricted operation point (such as the 'ropi' set of samples) to indicate the restriction on the temporal sub-layer. Similarly for the layer ID, if the parsing of the NAL unit headers indicates a maximum layer ID different from the maximum layer ID declared for any operation point, the index of the corresponding set of output layers using that determined maximum layer ID can be declared in the restricted operation point information according to one of the above variations. A robust parser that identifies a similar issue regarding the maximum layer ID or the maximum temporal ID can produce a reliable bitstream in terms of the operation point declaration by adding an OPI NAL unit to the parsed bitstream. The OPI NAL unit will contain the determined value of the maximum temporal ID or the maximum layer ID actually used by the NAL units encapsulated in the media file and present in the parsed bitstream.

[0138] Changing the semantics of an existing operation point descriptor

[0139] Figure 3 Illustrates how to create a 'vopi' set of samples according to an embodiment of the proposed method. This embodiment of the proposed method is characterized in that the encapsulation module relies not only on the VPS information but also on the OPI NALU (when present in the VVC bitstream) to build a list of valid operation points. In other words, the encapsulation module filters a part of the set of output layers declared in the VPS according to the indication in the OPI NAL unit. This embodiment is suitable for the first encapsulation when the 'vopi' or the operation point descriptor structure has not been generated yet. For editing an already encapsulated media file, or when the 'vopi' or the operation point descriptor structure has already been generated, the first embodiment or the third embodiment may be preferred.

[0140] In step 300, the encapsulation module receives a bitstream for encapsulation. In step 301, it looks for the Video Parameter Set NAL unit. If there is no VPS, there are no multiple layers, then no description of the operation points (e.g., 'vopi' or 'opeg') is needed, and the encapsulation module considers the next NAL unit in 310. If a VPS NALU is detected in 301, the encapsulation module parses the VPS fields in 302. The encapsulation module obtains the number of layers in 303 and allocates tables to store the layers and their descriptions. During 304, the encapsulation module reads the dependencies of the individual layers from the payload of the VPS NAL unit and stores the reference layers in the allocated tables. In 305, the encapsulation module reads the information on which layers are output layers and records this information in the allocated tables. Then, the encapsulation module obtains the information on the number of sub-layers in 306 and records it in the allocated tables. In step 307, the encapsulation module derives the maximum number of operation points (or sets of output layers) based on the recorded information. This can depend on a combination of VPS flags such as vps_all_independent_layers_flag, vps_each_layer_is_an_ols_flag, vps_ols_mode_idc or as given in the list of vps_ols_output_layer_flag. Then, the encapsulation module obtains the next NAL unit in test 308 and checks whether it is an OPI NAL unit. If not found, it creates a 'vopi' in 309 to describe the number of operation points as the maximum number of operation points derived from the VPS NAL unit, and continues to process the bitstream in 310. If there is an OPI NAL unit, the encapsulation module checks in 311 whether there is an indication of the output layer set index and checks in 313 whether there is also an indication of the highest temporalID in the OPI NAL. When the first test 311 is true, the encapsulation module filters the list of the maximum number of operation points determined in 307 according to the indication in opi_ols_idx in 312 (in fact, some operation points can be a subset of the operation points with an index equal to opi_ols_idx and are thus also accessible). When test 313 is true, the maximum number of sub-layers is updated with the indication in the opi_htid_plus1 field of the OPI NAL unit in 314. Then, in 315, a 'vopi' sample group description entry is constructed with the filtered number of operation points (or sets of output layers). And the encapsulation module continues to parse the NAL units to establish tracks and sample descriptions. The filtering step 312 depends on the layer configuration, such as as Figure 1a and 1b shown. For example, in Figure 1aAbove, when each layer is independent and each layer is indicated as an output layer, there can be a one-to-one mapping between the OLS and the layer. Then filtering is easy, including keeping the layer and then the OLS (corresponding to the OLS index given in the OPI NAL unit). When there are dependencies between layers, such as in Figure 1b above, filtering can result in more than one OLS. For example, in Figure 1b in the case of the configuration of, an indication of opi_ols_idx = 2 will result in 2 OLSs: the OLS with index 2 includes layer 1 and layer 0, and the OLS with index 1 only includes layer 0. In the case of layer dependencies, the filtering step 312 explores the list of layer dependencies and collects the layers involved in the OLS. From this obtained set of layers, some subsets can also correspond to possible OLSs and are then inserted into the list of filtered OLSs. For sub-layer filtering in 314, the encapsulation module keeps the highest temporal ID given by the OPI NAL unit in the memory and uses it in the description of the operating point as the max_temporal ID.

[0141] According to one embodiment, the semantics of the VvcOperatingPointsRecord structure used in the 'vopi' sample group description entry are modified as follows:

[0142] num_operating_points: Gives the number of operating points following the information. The number of operating points can correspond to the number of sets of output layers derived from the VPS NAL unit or from the VPS and OPI NAL units when there is an OPI.

[0143] output_layer_set_idx is the index of the set of output layers that defines the operating point. The mapping between the output_layer_set_idx and the layer_id value should be the same as the mapping specified by the VPS for the set of output layers with index output_layer_set_idx.

[0144] max_temporal_id: Gives the maximum TemporalId of the NAL unit of this operating point as indicated in the VPS. When there is an OPI NAL unit, this value will correspond to opi_htid_plus1 minus one.

[0145] layer_count: This field indicates the number of necessary layers of this operating point as defined in ISO / IEC 23090-3. When there is an OPI NAL unit, layers other than those included in the OLS with OLS index equal to opi_ols_idx should not be counted.

[0146] max_layer_count: The count of all unique layers among all operation points related to the associated base track. When there is an OPINAL unit, layers other than those included in the OLS with an OLS index equal to opi_ols_idx should not be counted.

[0147] Figure 3 Describe the filtering of operation points at the start of the bitstream, but the encapsulation module within step 310 can check for the presence of new OPINALUs and can filter the maximum number of operation points again according to step 312 or 314 to generate new 'vopi' sample group description entries. Thus, when the VVC bitstream references several VPSs, or when the VVC bitstream contains several OPINAL units, it may be necessary to include several entries in the sample group description box with grouping_type 'vopi'. For the more common case of the presence of a single VPS or OPI, it is recommended to use the default sample group mechanism defined in ISO / IEC 14496-12 and include the operation point information sample group in the sample table box instead of including it in each track segment.

[0148] Furthermore, when the OPINALU contains an indication about the highest temporal ID, the 'linf' sample group may be affected by the presence of the OPINALU in the VVC bitstream. Then, when the VVC bitstream references several VPSs, or when the VVC bitstream contains several OPINAL units, it may be necessary to include several entries in the sample group description box with grouping_type 'linf'. For the more common case of the presence of a single VPS or OPI, it is recommended to use the default sample group mechanism defined in ISO / IEC 14496-12 and include the layer information sample group in the sample table box instead of including it in each track segment.

[0149] According to this embodiment, the 'vopi' sample group provides reliable information to the parser or reader because operation points for which the NAL unit is no longer available are no longer listed and described in the 'vopi' sample group. The parser and reader can safely select one of the operation points listed on 'vopi' for transmission, rendering, or display.

[0150] In addition to constructing a reliable 'vopi', when the encapsulation module has knowledge of the entire bitstream (e.g., offline encapsulation) and there is no multiple filtering of the output layer set in the bitstream, the 'opeg' structure can also be superior to 'vopi', and the same filtering mechanism (i.e., filtering the output layer set) can also be applied. The semantics of 'opeg' are then updated as follows:

[0151] num_operating_points: The number of operating points that the given information follows. The number of operating points may correspond to the number of output layer sets derived from the VPS NAL unit or from the VPS and OPI NAL units when an OPI NAL unit exists.

[0152] max_temporal_id: Gives the maximum TemporalId of the NAL unit of this operating point as indicated in the VPS. When an OPI NAL unit exists, this value will correspond to one less than opi_htid_plus1 (if greater than 0).

[0153] layer_count: This field indicates the number of necessary layers of this operating point as defined in ISO / IEC 23090-3. When an OPI NAL unit exists, layers other than those included in the OLS with an OLS index equal to opi_ols_idx should not be counted.

[0154] Changing the syntax and semantics of the existing point descriptor

[0155] The 'vopi' sample group notifies the application of different operating points provided by a given VVC bitstream. The metadata structure VvcOperatingPointsRecord included in the 'vopi' sample group description describes the operating points as profile levels, number of layers, maximum temporal ID, etc., thus reflecting the information from the VPS NALU. However, this description does not indicate to the application "preferred", "default", "recommended", or filtered operating points via the OPI NAL unit, even if this information exists in the VVC bitstream.

[0156] In this embodiment, compared with the existing 'vopi' structure, the modification in the metadata structure of the 'vopi' sample group indicates which one in the list of operating points corresponds to the recommended output layer set (when this information is available in the bitstream (bold change)).

[0157] In the first variant, the syntax can become:

[0158]

[0159]

[0160] wherein, the recommended_ols_flag, when set, indicates that the output layer set index corresponds to the target output layer set to be decoded, as indicated in the OPI NAL unit. This operating point can be selected as the target output layer set to be decoded. When not set, all NAL units of the corresponding output layer set may not exist in the media file and should not be selected as the target output layer set to be decoded.

[0161] The recommended_htid indicates the highest temporal sublayer to be decoded for this set of output layers, as indicated in the OPI NAL unit. When present, this value overrides the value of max_temporal_id to indicate the highest temporal sublayer available for the operation point.

[0162] Compared with the existing 'vopi', the semantics of other parameters remain unchanged.

[0163] In the second variant, the syntactic change is slightly different because two new flags are used to consider both OLS and temporal sublayer filtering or restriction, as indicated by the OPI NAL unit. This variant is expressed in the structure of the 'vopi' sample group as follows:

[0164]

[0165]

[0166] Among them, the restricted_ols_flag, when set, indicates that the output layer set index corresponds to the decodable target output layer set, as indicated in the OPI NAL unit (specifically the opi_ols_idx parameter). This operation point can be selected as the target output layer set for decoding by the reader or parser. When not set, the corresponding output layer set or operation point may not have all its NAL units present in the media file and should not be selected as the target output layer set for decoding by the reader or parser.

[0167] The restricted_sublayer_flag, when set, indicates that a restriction has been set on the temporal sublayer, as indicated by the OPINAL unit (specifically the opi_htid_plus1 parameter when it is greater than 0). When opi_htid_plus1 is equal to 0, the restricted_sublayer_flag is not set. When not set, no restriction is set on the sublayer for this operation point, and the max_temporal_id value can be used safely.

[0168] The restricted_htid indicates the highest temporal sublayer to be decoded for this set of output layers, as indicated in the OPI NAL unit. This value corresponds to the value of opi_htid_plus1 minus one. When present, this value overrides the value of max_temporal_id to indicate the highest temporal sublayer available for the operation point.

[0169] Compared with the existing structure of the 'vopi' sample group, the semantics of other parameters remain unchanged.

[0170] The above two variations show one or two control flags indicating some constraints or a constraint on some of the operation points in the initial list of operation points, and new values that may be considered for some parameters (e.g., maximum time ID) for the restricted operation points. This rewriting mechanism has the following advantages: when the original bitstream is filtered or sent, it is easy to edit or update the existing 'vopi'. For all and some of the restrictions that can be changed along the bitstream, the 'vopi' is described once and can be set / reset only through control flags and rewritten values. When the parameters of the operation points will be affected by the restrictions or filtering of the bitstream, additional variations can consider more control flags and rewritten values. It should be noted that the control flag indicating whether the 'vopi' contains restrictions can be placed, for example, higher up in the 'vopi' structure using reserved bits, i.e., before the loop over the operation points.

[0171] This embodiment and its variations allow for an explicit indication in the existing operation point descriptors of a subset of operation points or sets of output layers, where all NAL units can be used for the subset of the operation points or sets of output layers described above in one or more encapsulated tracks representing this (these) operation points or sets of output layers. The control flags added in the variation take into account the filtering or restrictions indicated by the OPI NAL unit. When an OPI NALU appears in the bitstream, a new VisualSampleGroupEntry of type 'vopi' is inserted into the sample group description box of type 'vopi', and this new sample group entry provides a new set of flags and may rewrite the values of the restrictions or filtering described in the OPI NALU. By doing so, the 'vopi' always contains a reliable indication for the parser or reader regarding the operation points and the availability of their corresponding NAL units.

[0172] The same modifications (control flags and rewritten values) can be applied to the 'opeg' structure. In the case where filtering of an already encapsulated file has been completed, the 'opeg' structure is modified by adding a restriction_flag to indicate whether the 'opeg' still applies to the media file:

[0173]

[0174]

[0175] where the restriction_flag, when set, indicates that the initial 'opeg' may be unreliable after filtering or restriction of the OLS (possibly indicated by the OPI NALU). This indicates that some of the operation points may not be reconstructed correctly. When not set, it indicates that all operation points can be selected and rendered.

[0176] For the above embodiments that describe a restricted or filtered list of operating points and a complete list of operating points in different structures of a media file, the encapsulation module may decide to keep the OPI NAL unit in the set non-VCL NAL unit. This can be done using the NAL unit array of the decoder configuration recorded in the 'vvcC' box, or in the samples of the parameter set track, or in the samples of the track (e.g., samples that declare 'vopi', 'opeg', or 'ropi'). In this case, it is recommended that the array be in the order of DCI, VPS, OPI, SPS, PPS, prefix APS, prefix SEI. This setting of the NAL unit array carries the initialization NAL unit. The NAL unit type is restricted to only indicate DCI, VPS, OPI, SPS, PPS, prefix APS, and prefix SEI NAL units.

[0177] In the above embodiments where the structure describing the operating points considers a filtered or restricted list of operating points, the encapsulation module may decide to skip the OPI NAL unit from the parameter set or non-VCL NAL unit. This can be used in a configuration where an application that obtains information from the media file can control the decoder to indicate the set of target output layers and the highest temporal ID to be decoded.

[0178] The above embodiments can also be applied to repackaging, for example, obtaining a list of initial operating points described in the input media file (e.g., reading from 'vopi' or 'opeg') from a user interface or from application settings. The user removes the NAL units corresponding to some operating points or sub-layers based on the layer ID or temporal ID field in the NAL unit header through the user interface or the application through predefined settings. For example, the NAL units corresponding to operating points where the application requires too high a profile or level or too high a frame rate can be filtered. The media file edited in this way is saved to the second media file 230, which contains an indication of the restricted or filtered operating points (e.g., 'ropi' or 'vopi' with additional syntax or semantics) in its metadata section.

[0179] Indication of the default track

[0180] The indication of the default track can be done at the track level.

[0181] When a bitstream containing multiple OLSs is encapsulated as a single track or when a track contains an OLS with an index equal to the index indicated in the OPI NAL unit, the track can be explicitly indicated as the default track for the media player. There should be at most one track marked as the default track for a given media processor. For a video bitstream, the media processor is 'vide' which indicates the video track. For such an indication, the track header box is extended with a new flag value:

[0182] track_is_default: When set, indicates that the track can be considered the default track to be presented. The flag value is 0x000010. For a given processor type, there should be only one track with this flag set. A track with this flag value set should also have track_enabled and track_in_movie also set. Any value can be used as long as it does not conflict with other track header flag values already in use.

[0183] It should be noted that when the bitstream contains an indication or when the encapsulation receives information about a recommended track in a set of tracks with the same media type through an external device (user interface, configuration, application, etc.), the track can have this new flag value set. A parser that encounters a media file with a track having this flag set can select that track as the default track for playback. Having such a flag enables the parser to quickly select a track without the need for further examination of the ISOBMFF structure in the file.

[0184] When the version or brand of an ISOBMFF file does not provide the track_is_default flag value, such a track can have its track_in_movie set to 1. For example, a track directly containing an operation point or a set of output layers authorized by an OPI NAL unit should have the flag track_in_movie set in its track header box. This allows a media parser or reader to easily identify the track to be selected and played, especially when no indication about track relationships or selection criteria is provided. Conversely, a track encapsulating a set of output layers indicated by an OPI NALU as filtered can have its track_in_movie flag set to 0, or even its track_enabled flag set to 0.

[0185] Optionally, or in addition to the new track header flags, a set of tracks providing alternative operation points of the same video can be indicated as part of the alternative tracks in their track header boxes by using the same value different from the reserved 0 in their alternate_group fields. To distinguish the tracks within the alternative group, track selection can be indicated for each track in the set as being part of a switching group by setting the value of the switch_group field to the value of the alternate_group of the track header. Additionally, the track selection box uses a new distinguishing attribute: the “operation point index” or “output layer set index” indicated by the reserved four-character code ‘olsi’ (where the 4cc is just an example and any other 4cc that does not conflict with the already registered 4ccs can be used for this purpose), the index of the operation point or output layer set to which the track corresponds. The set of descriptive attributes can also be extended with a new attribute: the “operation point information selection” or “output layer set selection” indicated by the reserved four-character code ‘olss’ (where the 4cc is just an example and any other 4cc that does not conflict with the already registered 4ccs can be used for this purpose). This new descriptive attribute indicates that the tracks in the switching group are different operation points of the same content. Track selection can include additional distinguishing attributes to indicate the differences between operation points using ‘cgsc’ or temporal scalability with ‘tesc’, such as scaling in terms of quality.

[0186] Indication of a track with multiple output layers

[0187] The VVC specification allows an output layer set to potentially contain more than one output layer. When encapsulating such a VVC bitstream in a single track, the purpose of these multiple output layers needs to be described.

[0188] In a first variant, the ‘vopi’ or ‘opeg’ structures can be used. For previous codecs such as MVC, SVC, HEVC, several sample entry types such as for example ‘mvc1’ for multi-view and ‘svc1’ or ‘lhv1’ for scalability were allowed to understand the purpose of the track. The VVC file format restricts the number of sample entry types and thus cannot help in identifying the purpose of multiple output layers. In fact, for example, a track with a single layer or a track containing an output layer set with multiple layers can have the sample entry ‘vvc1’. When a single track contains more than one output layer, the parser or reader may need to describe the structure of the output layer set to assist or inform the application about the use of such a track and the corresponding decoded pictures.

[0189] Then, we propose to modify ISO / IEC 14496-15 by adding a 'vopi' sample group or 'opg' structure (as proposed below) to such a VVC track: When a VVC bitstream is encapsulated as a VVC track and the output_layer_set_idx declared in the VVC decoder configuration record of that VVC track refers to an output layer set that includes more than one output layer, a 'vopi' sample group or 'opeg' entity group shall be present. In fact, in the absence of a 'vopi' or 'opeg' box, the VVC decoder configuration record box is not sufficient to determine the number of output layers present in the OLS described in the VVC decoder configuration record box.

[0190] As an alternative to having a 'vopi' or 'opeg' for a track containing at least one OLS with more than one output layer, the VVC decoder configuration record is modified (in bold) as follows:

[0191]

[0192] where num_output_layers indicates the number of output layers contained in the track.

[0193] When such a single-track encapsulation further encapsulates a bitstream that has been filtered or restricted in terms of output layer sets or sub-layers, the track may contain:

[0194] - In addition to 'vopi' or 'opeg', there is also a 'ropi' sample group

[0195] - A 'vopi' or 'opeg' constructed according to the steps in Figure 3

[0196] - A 'vopi' or 'opeg' with updated syntax and semantics indicating which operating points are filtered or restricted.

[0197] According to the second variation, a new sample entry type can be used. To explicitly indicate that a VVC track encapsulates more than one output layer, for a VVC track with multiple output layers, a new sample entry type is used, such as 'lvc1' or 'vvc1'. The samples of a track with the 'lvc1' (or 'vvc1') sample entry type include one or more layer components, which may result in more than one decoded picture being output when the corresponding NAL unit is given to the video decoder (setting their PictureOutputFlag to 1). Assuming that the SchemeTypeBox and SchemeInformationBox indicate what to do with these multiple components, these new sample entries can be declared within the OriginalFormatBox within the RestrictedSchemeInfoBox. If such a scheme is not defined, these new sample entries can be declared in the sample description box. Some examples of schemes can be combinations of stereoscopic pairs, spatial synthesis of decoded pictures...

[0198] According to the third variation, a limited set of sample entry types can be used. An alternative to the proliferation of sample entry types is to maintain a limited set for VVC tracks, but impose constraints on the mapping of layers to VVC tracks (especially the mapping of output layers to tracks). Constraints can be defined for VVC stores with multiple layers. ISO / IEC 14496-15 AMD 2 allows different mappings of layers to tracks.

[0199] For a separate track containing one or more layers, the proposed restrictions include avoiding or prohibiting a separate track from containing more than one output layer for a given OLS. In other words, a separate track containing one or more layers should contain at most one output layer for each OLS contained in that track.

[0200] Figure 4 An example of a VVC bitstream is shown. A video bitstream generated by a video encoder (usually a VVC encoder) can generate a bitstream including two or more independent layers. These independent layers can each have different coding parameters. For example, Figure 4 An exemplary bitstream composed of NAL units (e.g., 400 to 404) encoding two (or more) independent layers is shown. The bitstream includes a set of NAL units corresponding to each layer. For example, the NAL unit header can include an identifier of the layer that allows distinguishing the NAL units corresponding to a given layer. For example, the headers of NAL units 402 to 404 include a layer id corresponding to Figure 4 the nth layer 405 of the bitstream. The NAL units of another layer are after these NAL units and include another layer id corresponding to the n+1th layer 406.

[0201] The decoder is capable of decoding each of these independent layers separately. Specifically, the VVC bitstream may define a set of output layers including two or more independent layers, each independent layer being an output layer. Typically, the set of output layers is signaled in the VPS NAL unit 400 according to the VVC specification. Optionally, an OPI NAL unit (not shown) may follow the VPS NAL unit.

[0202] In this case, the VVC bitstream may also contain SEI messages that describe the composition of decoded pictures (e.g., Figure 4 401 in ). Typically, the SEI messages enable the decoder to determine the position of the layer in the picture resulting from the composition. For example, the syntax of JVET-S0213 "AHG8 / AHG9: Refinement of proposed positioning information SEImessage of output independent layers with example bitstreams, E.Thomas" or JVET-S0107 "AHG9 / AHG12: Recommended multi-layer composite picture SEI messages, J.Boyce" may be used. The SEI messages may define parameters regarding the overall composition (such as the size of the composite picture, etc.) or parameters regarding an individual layer (such as positioning information in the composition, etc.). In the former case, the layer identifier of the SEI message NAL unit may correspond to the lowest layer (e.g., the first layer in the decoding order) or a predetermined reserved value (e.g., 0). In the latter case, the layer identifier value corresponds to the layer of interest. For example, NAL 402 is an SEI message that provides composition information related to the nth layer.

[0203] In the VVC specification, the support for SEI messages is optional. For this reason, some compliant decoders may not consider the content of the composition described in the SEI messages. On the other hand, decoders that support these SEI messages can parse and process them to generate a composite picture from the set of decoded pictures corresponding to each independent layer of the set of output layers.

[0204] The proposed method addresses the encapsulation of such bitstreams in the ISOBMFF format and its extension for the carriage of NAL unit structured video. Specifically, methods for generating and parsing ISOBMFF files are proposed that allow the determination of the composition suggested by the SEI messages present in the VVC bitstream without having to parse the content of the SEI composition messages. A format-agnostic description of the composition information is provided, which avoids the media player or parser implementing format-specific parsing (to understand, for example, SEI messages).

[0205] It should be noted that the generated media file does not necessarily contain the tracks that will represent the composition, i.e., the base track that references several tracks of one or more layers describing the VVC bitstream.

[0206] Figure 5 An example of the encapsulation process according to an embodiment of the proposed method is shown. In the first step 500, a file format or media file writer (e.g., the encapsulation module 205) receives a video bitstream (e.g., the VVC bitstream as represented with reference to Figure 1a and 1b or Figure 4 ) for processing. The encapsulation process includes applying a processing loop to each access unit of the video bitstream. Specifically, the encapsulation module parses the NAL unit in step 501. The encapsulation module determines the set of output layers described in the bitstream in step 502. For a VVC bitstream, this information is typically encoded in the VPS NAL unit and possibly in the OPI NAL unit (when present). The writer derives the operation points from the OLSs described in the bitstream. For example, the writer can associate an operation point with each OLS described in the VPS. During this determination step 502, the writer also determines the presence of one or more independent layers in the OLSs and if they are output layers and applicable.

[0207] Step 503 determines whether the access unit includes composition information indicating that multiple independent layers belonging to the same OLS are part of the recommended composition. For VVC, this includes parsing the SEI message (when present), which signals the position of the decoded pictures of the layers in the composed picture. Additionally, these SEI messages can provide optional transformations (scaling, upsampling or downsampling, rotation, etc.) applied to each decoded picture or composed picture. It should be noted that different composed pictures can be defined for each OLS, and a single layer can be part of multiple OLSs. As a result, the writer associates the composition information determined from the SEI message to each OLS.

[0208] In step 504, the writer generates the composition information to be signaled in different boxes of the ISOBMFF output media file based on the composition information associated with the OLSs of the bitstream. Several embodiments are described to signal this information at different locations in the media file. In some embodiments, the writer associates the composition information with a specific operation point of the ISOBMFF file based on the information determined in step 503.

[0209] Finally, the process encapsulates the NAL units of the access unit in one or more tracks. Typically, when using more than one track (referred to as multi-track encapsulation), the encapsulation process does not necessarily define the track that will represent the composition (sometimes referred to as the base track).

[0210] Figure 6An example showing the main steps of the parsing process according to an embodiment of the proposed method is provided. According to this embodiment, the media file contains synthesis information for a given operating point.

[0211] As shown in the figure, the first step is to initialize the media player (such as 220) in step 600 to start reading the media file encapsulated according to the present invention. Then, in step 601, the media player determines the list of operating points described in the media file. Generally, the media player parses the VvcOperatingPointsRecord (‘vopi’) or OperatingPointGroupBox (‘opeg’) box. These operating points indicate the decoding possibilities of the video bitstream.

[0212] In one embodiment, the media file includes synthesis information in some ISOBMFF structures. This synthesis information is associated with the operating points and enables determination of whether the NAL units encoding the pictures in the operating points include information describing the synthesis of multiple independent layers. During step 602, the player determines the presence of such information. In some embodiments, the synthesis information is represented for each output layer and each set of output layers described in the VvcOperatingPointsRecord (‘vopi’), OperatingPointGroupBox (‘opeg’), or VvcDecoderConfigurationRecord box or structure.

[0213] Based on the information determined in steps 601 and 602, the media player can identify the operating points that include multiple independent layers with synthesis information encoded in non-VCL NAL units. Based on the decoding capabilities and the support for the processing of these non-VCL NAL units, in step 603, the player filters the list of selectable operating points. For example, when the video decoder responsible for decoding the bitstream cannot parse and process the synthesis information present in the non-VCL NAL units (usually SEI NAL units), the media player can ignore the operating points associated with the synthesis information (at the file format level). In another example, if the synthesis described in the synthesis information is not suitable for the media player display or the user's preference, the operating points can also be ignored, and then better operating point options can be selected in 603. On the other hand, if the decoder supports the non-VCL information describing the synthesis, the player can select the operating point.

[0214] The final stage 604 of the media player processing includes forming the reconstructed bitstream corresponding to the selected operating point. Specifically, the reconstructed bitstream can include non-VCL NAL units having synthesis information (such as contained in the NAL unit 401 according to Figure 4 )

[0215] Now, different embodiments of a method for describing composition in an ISOBMFF file are described.

[0216] Extension of the existing sample group entry box

[0217] In a first variation, an extension of the RectangularRegionGroupEntry (‘trif’) box defined in ISO / IEC 14496-15 is proposed for describing the composition of multiple layers. The RectangularRegionGroupEntry box can be extended to include composition information generated by the media file writer in step 504 and then used by the media file player in step 602. The RectangularRegionGroupEntry describes the rectangular region covered by the sample group. In this embodiment, the RectangularRegionGroupEntry is modified to include new composition information that indicates the position of the samples referenced by the RectangularRegionGroupEntry box in the composed picture. For example, the syntax of the modified RectangularRegionGroupEntry can be as follows (where the newly introduced syntax is marked with bold characters):

[0218]

[0219]

[0220] Wherein, the following semantics apply to the new syntax elements:

[0221] The composite_region_flag specifies (e.g., when equal to 1) that the decoded picture of the NAL unit associated with this rectangular region group entry is signaled in the bitstream as the rectangular region of the composed picture, and further information about the composite region is provided by subsequent fields in this rectangular region group entry. Another value (e.g., 0) specifies that the decoded picture of the NAL unit associated with this rectangular region group entry is not signaled as the rectangular region of the composed picture, and no further information about the region is provided in this rectangular region group entry.

[0222] The composite_horizontal_offset and composite_vertical_offset give the horizontal and vertical offsets, respectively, of the top-left pixel of the composite region covered by the NAL unit in each rectangular region associated with this rectangular region group entry in the luma samples relative to the top-left pixel of the base region. The base region used in the RectangularRegionGroupEntry is the composed picture when the composite_region_flag is equal to 1.

[0223] The composite_region_width and composite_region_height respectively give the width and height of the composite region in the composite picture covered by the NAL units in each of the rectangular regions associated with the rectangular region group entry in the luma samples.

[0224] As a result, in this embodiment, the positions of the NAL units (i.e., the NAL units of each layer) associated with the rectangular regions in the composite picture can be determined. The size of the composite picture is inferred from the composite region described at the origin of the composite picture (both composite_horizontal_offset and composite_vertical_offset are equal to 0) and the positions of the bottom-most and right-most pixels of the composite region in the composite picture. In an alternative, additional syntax elements specify the size of the composite picture in the RectangularRegionGroupEntry box for each region. To avoid repeating this information in all RectangularRegionGroupEntry boxes, an additional flag (e.g., composite_picture_present_flag) indicates that the RectangularRegionGroupEntry defines the size of the composite picture rather than the size of the composite region.

[0225] In a variation, for a single track encapsulation (i.e., all independent layers of the composite are in the same track), the size of the composite picture is inferred to be equal to the width and height of the track (signaled in the track header box). The composite information for each layer can be provided as a sample group entry of type 'nalm' (i.e., the NaluMapEntry box associated with the RectangularRegionGroupEntry box), one for each layer.

[0226] In a variation, for a track encapsulation where each track contains a single layer, the width and height of the composite region are inferred to be equal to the width and height of the track and are thus not signaled in the composite information (i.e., for any embodiment where the composite information is described in the sample group entry).

[0227] In a second variation, the composite information is associated with the samples of a given layer. A layer information sample group ('linf') conveys the information of the layers of a given track. Specifically, in one embodiment of the present invention, the composite information is signaled in a modified layer information sample group entry.

[0228] For example, the layer information sample group entry may have the following syntax (new fields are indicated in bold):

[0229]

[0230]

[0231] Among them, the added syntactic elements have the following semantics:

[0232] When composite_layer_flag equals 1, it specifies that the decoded picture (associated with the layer information sample group entry) of the NAL unit of the layer with nuh_layer_id (i.e., layer identifier) equal to layerID is signaled in the bitstream as the rectangular region of the composite picture, and further information about the composite region of this layer is provided by the subsequent fields in the layer information sample group entry. The value 0 specifies that the decoded picture (associated with the layer information sample group entry) of the NAL unit of the layer with nuh_layer_id (i.e., layer identifier) equal to layerID is not signaled as the rectangular region of the composite picture, and no further information about this region is provided in the rectangular region sample group entry.

[0233] composite_horizontal_offset and composite_vertical_offset respectively give the horizontal offset and vertical offset of the top-left pixel of the composite region covered by the NAL unit of the layer with nuh_layer_id (i.e., layer identifier) equal to layerID in the luminance samples, in each rectangular region associated with the layer information sample group entry, relative to the top-left pixel of the base region. The base region used in LayerInfoGroupEntry is the composite picture when composite_layer_flag equals 1.

[0234] composite_region_width and composite_region_height respectively give the width and height of the composite region in the composite picture covered by the NAL unit of the layer with nuh_layer_id (i.e., layer identifier) equal to layerID in the luminance samples, in each rectangular region associated with the layer information group entry.

[0235] New box 'crif' indicating composite information

[0236] To avoid mixing composite information with rectangular region information and thus simplify the parsing of the bitstream when composite information is not present in the bitstream, this embodiment includes signaling composite information in a dedicated sample group entry of the sample group description box. For this purpose, a new grouping type is defined: 'crif' for "composite region information". Here, the four-character code and name are only examples. Any other reserved 4cc that does not conflict with the existing 4cc can be used for this purpose of composite information description.

[0237] In the first variation, the composition is described as a rectangular region of width times height at positions hor_offset and ver_offset from the origin of the composite picture in the luma samples. For example, the new CompositeRegionGroupEntry sample group entry box may have the following syntax:

[0238] Definition:

[0239]

[0240] The semantics of the syntax elements may be similar to the previous embodiments:

[0241] A composite_region_flag equal to 1 specifies that the decoded picture of the NAL unit associated with this composite region group entry is signaled in the bitstream as a rectangular region of the composite picture, and further information about the composite region is provided by the subsequent fields in this composite region group entry. A value of 0 specifies that the decoded picture of the NAL unit associated with this composite region group entry is not signaled as a rectangular region of the composite picture, and no further information about the region is provided in this composite region group entry.

[0242] composite_horizontal_offset and composite_vertical_offset give the horizontal and vertical offsets, respectively, of the top-left pixel of the composite region covered by the NAL unit in each composite region associated with this composite region group entry in the luma samples relative to the top-left pixel of the base region. The base region used in CompositeRegionGroupEntry is the composite picture.

[0243] composite_region_width and composite_region_height give the width and height, respectively, of the composite region in the composite picture covered by the NAL unit in each composite region associated with this composite region group entry in the luma samples.

[0244] In a variation of any of the foregoing embodiments, instead of representing the position and size of the composite region in the luma samples, it is represented in arbitrary unit sizes. The arbitrary unit sizes are either pre-determined or represented as a new parameter of the composite information. For example, the composite information includes a composite_unit_size_flag flag, which being equal to 1 indicates that there is a syntax element in the luma samples that specifies the arbitrary unit size. This syntax element is, for example, an unsigned integer represented in 16-bit length. When the flag is equal to 0, it is inferred that the arbitrary unit is equal to one luma sample.

[0245] In the transformation, the composite region position and size parameters are normalized by the size of the composite picture. As a result, a value of 2^16 - 1 for composite_vertical_offset or composite_region_height corresponds to the height of the composite picture in the luminance samples. A value of 2^16 - 1 for composite_horizontal_offset or composite_region_width corresponds to the width of the composite picture in the luminance samples.

[0246] In a variation of any of the foregoing embodiments, the composite information includes additional syntax elements that allow representation of other kinds of transformations of the decoded pictures of the individual independent layers. For example, it specifies one or several of the syntax elements that represent parameters for a rotation transformation, and / or a scaling transformation on one or two axes, and / or a downsampling or upsampling operation, and / or a mirroring operation.

[0247] For example, the CompositeRegionGroupEntry syntax is as follows:

[0248]

[0249]

[0250] The syntax element that represents the rotation transformation of the decoded pixels of the composite region is composite_rotation_angle. It is a 16.16 fixed-point value that indicates the rotation angle in degrees. In the variation, composite_rotation_angle is encoded with a predetermined bit length, and each value corresponds to a predetermined rotation angle value. For example, composite_rotation_angle is encoded in 2 bits, and each value corresponds to a multiple of 90° rotation angle:

[0251] · A value equal to 0 corresponds to 0° rotation,

[0252] · A value equal to 1 corresponds to 90° rotation,

[0253] · A value equal to 2 corresponds to 180° rotation, and

[0254] · A value equal to 3 corresponds to 270° rotation.

[0255] In the variation, a flag can indicate whether the decoded picture of the composite region should be mirrored.

[0256] The syntax elements composite_scaling_horizontal_factor and composite_scaling_vertical_factor represent the horizontal and vertical magnification / upsampling (or reduction / downsampling) factors of the decoded pictures for the composite region. For example, 16.16 fixed-point representation is used to encode these two values. In the case of anamorphosis, when the upsampling factors are constrained to be equal in the composite information, the factor is signaled only once.

[0257] One advantage of the foregoing embodiments is that they allow the definition of dynamic composite information for one or more independent layers. The dynamic composite information allows the signaling of the characteristics of different composite arrangements or composites for different time ranges within the track encapsulating the VVC samples. Although the signaling is dynamic, static composite information can be signaled by defining the composite information in the default sample group according to the ISOBMFF specification, making the signaling of ISOBMFF more compact.

[0258] The composite matrix used to describe the composite

[0259] In some embodiments, the composite can be described using a composite matrix to be applied to the decoded pictures to place the pictures into the composite. The composite matrix defines the transformation to be applied to the decoded pictures.

[0260] In a first variation of providing composite information in the 'trif' box, the RectangularRegionGroupEntry box includes an array of nine 32-bit syntax elements, each representing a coefficient of the matrix. For the transformation matrix signaled in the TrackHeader Box (in ISOBMFF, the composite matrix is static for each track), the same syntax and semantics are used as described in ISOBMFF. By default, the inferred matrix corresponds to a unitary matrix, which is an identity transformation.

[0261] For example, RectangularRegionGroupEntry may include the following syntax elements:

[0262]

[0263] where the following semantics apply to the new syntax elements:

[0264] A composite_region_flag equal to 1 specifies that the decoded picture of the NAL unit associated with this rectangle region group entry is signaled in the bitstream as a rectangle region of the composite picture, and further information about the composite region is provided by subsequent fields in this rectangle region group entry. A value of 0 specifies that the decoded picture of the NAL unit associated with this rectangle region group entry is not signaled as a rectangle region of the composite picture, and no further information about the region is provided in this rectangle region group entry.

[0265] matrix provides the transform matrix for the decoded pixels of the rectangle region of the composite picture.

[0266] In the second variation of providing composite information in the 'crif' box, the CompositeRegionGroupEntry box includes an array of nine 32-bit syntax elements, each representing a coefficient of the matrix. The same syntax and semantics are used as described in ISOBMFF for the transform matrix signaled in the TrackHeader Box. By default, the inferred matrix corresponds to the unitary matrix, which is the identity transform. The syntax can be as follows:

[0267]

[0268] where the following semantics apply to the new syntax elements:

[0269] A composite_region_flag equal to 1 specifies that the decoded picture of the NAL unit associated with this composite region group entry is signaled in the bitstream as a rectangle region of the composite picture, and further information about the composite region is provided by subsequent fields in this composite region group entry. A value of 0 specifies that the decoded picture of the NAL unit associated with this composite region group entry is not signaled as a rectangle region of the composite picture, and no further information about the region is provided in this composite region group entry.

[0270] matrix provides the transform matrix for the decoded pixels of the composite region of the composite picture.

[0271] Association of composite information with operation points

[0272] In previous embodiments, the synthesis information is associated with samples of a track. The VVC bitstream may include at least two OLSs, each OLS including two or more independent layers. There is at least one independent layer in each of the at least two OLSs. In this case, the synthesis information associated with the samples of the independent layer may be related to the synthesis picture of one of the two OLSs. To address this issue, some embodiments of the proposed method not only associate the synthesis information with a set of samples, but also with an operation point. This enables determining in which decoding context (i.e., for which operation point) the synthesis information is applied.

[0273] According to a first variation, the synthesis information includes a list of operation points to which the synthesis information is applied. For example, a list of operation points involved in the synthesis may be provided in the 'crif' box. The synthesis information may include a list of operation points and associated synthesis parameters.

[0274] For example, when the synthesis information is signaled in the CompositeRegionGroupEntry box, it includes the following syntax elements:

[0275]

[0276] Among them, the following semantics apply to the new syntax elements:

[0277] num_operating_points: Gives the number of operation points followed by the information.

[0278] output_layer_set_idx is the index that defines the output layer set of the operation point. The mapping between output_layer_set_idx and the layer_id value should be the same as the mapping specified by the VPS for the output layer set with index output_layer_set_idx.

[0279] composite_region_flag being equal to 1 specifies that the decoded picture of the NAL unit associated with the composite region group entry is signaled in the bitstream (when decoded as part of the operation point) as a rectangular region of the synthesis picture, and further information about the synthesis region is provided by subsequent fields in the composite region group entry. A value of 0 specifies that the decoded picture of the NAL unit associated with the composite region group entry is not signaled as a rectangular region of the synthesis picture, and no further information about the region is provided in the composite region group entry.

[0280] composite_horizontal_offset and composite_vertical_offset respectively give (when decoded as part of this operation point) the horizontal and vertical offsets of the top-left pixel of the composite region covered by the NAL unit in each composite region associated with this composite region group entry in the luma samples relative to the top-left pixel of the base region. The base region used in CompositeRegionGroupEntry is the composite picture.

[0281] composite_region_width and composite_region_height respectively give (when decoded as part of this operation point) the width and height of the composite region in the composite picture covered by the NAL unit in each composite region associated with this composite region group entry in the luma samples.

[0282] In the second example, the composite information can be provided in the 'vopi' box. For each layer of the output layer set, the composite information related to the layer is described for this specific operation point. For example, VvcOperatingPointsRecord can have the following syntax:

[0283]

[0284]

[0285] where, for the new syntax elements, the following semantics apply:

[0286] composite_region_flag equal to 1 specifies that the decoded picture of the NAL unit of the layer with nuh_layer_id (i.e., layer identifier) equal to layerID is signaled in the bitstream (when decoded as part of this operation point, i.e., the i-th operation point described in the box) as a rectangular region of the composite picture, and further information about this composite region is provided by the subsequent fields in this box. A value of 0 specifies that the decoded picture of the NAL unit of the NAL unit of the layer with nuh_layer_id (i.e., layer identifier) equal to layerID is not signaled (when decoded as part of this operation point, i.e., the i-th operation point described in the box) as a rectangular region of the composite picture, and there is no further information about the composition of this layer in this operation point.

[0287] composite_horizontal_offset and composite_vertical_offset respectively give (when decoded as part of this operation point) the horizontal and vertical offsets of the top-left pixel of the composite region covered by the NAL units of the layer with nuh_layer_id (i.e., layer identifier) equal to layerID relative to the top-left pixel of the base region in the luma samples. The base region used is the composite picture.

[0288] composite_region_width and composite_region_height respectively give (when decoded as part of this operation point) the width and height of the composite region in the composite picture covered by the NAL units of the layer with nuh_layer_id (i.e., layer identifier) equal to layerID in the luma samples.

[0289] When a media file writer describes independent layers in separate tracks, it can specify the composite information for each layer and each operation point in an OperatingPointGroupBox (instead of a VvcOperatingPointsRecord with a syntax equal to that used above for VvcOperatingPointsRecord), as the two boxes share very similar syntax.

[0290] In the transform, the syntax of 'opeg' is modified as follows to provide the contribution of each track to the possible composition of each operation point, rather than the contribution of each layer to the possible composition of each operation point.

[0291]

[0292]

[0293] Among them, the following semantics apply to the new syntax elements:

[0294] composition_flag equal to 1 specifies that the decoded picture of the NAL units of the track with track_ID (i.e., track identifier) equal to entity_idx is signaled in the bitstream (when decoded as part of this operation point, i.e., the i-th operation point described in the box) as a rectangular region of the composite picture, and further information about the composite region is provided by the subsequent fields in this 'opeg'. A value of 0 specifies that the decoded picture of the NAL units of the track with track_ID (i.e., track identifier) equal to entity_idx is not signaled in the bitstream (when decoded as part of this operation point, i.e., the i-th operation point described in the box) as a rectangular region of the composite picture, and there is no further information following for the composition of this track.

[0295] composite_horizontal_offset and composite_vertical_offset give, respectively (when decoded as part of this operation point), the horizontal and vertical offsets of the top-left pixel of the composite region covered by the NAL units of the track with track_ID (i.e., track identifier) equal to entity_idx in the luma samples relative to the top-left pixel of the base region. The base region used is the composite picture.

[0296] composite_region_width and composite_region_height give, respectively (when decoded as part of this operation point), the width and height of the composite region in the composite picture covered by the NAL units of the track with track_ID (i.e., track identifier) equal to enity_idx in the luma samples.

[0297] composite_region_width and composite_region_height are optional. The parser may rely on the width and height of the track indicated by entity_idx.

[0298] In a second variant, the association between the composite information and the layers in an OLS is made different. Instead of listing the composite information in the description of the operation point (e.g., in 'vopi' or 'opeg'), an identifier is associated with each operation point. The composite information means that this identifier indicates in which operation point context the information is meaningful.

[0299] As a first example, the CompositeRegionGroupEntryBox provides the composition_ID syntax element represented as follows.

[0300]

[0301] This compositionID is the unique identifier for the composition described within this sample group entry. The value of the compositionID in the composite region group entry shall be greater than or equal to 0. When a sample group is associated with several SampleGroupDescriptionBoxes of type 'crif', all these SampleGroupDescriptionBoxes shall have different compositionID values. This last statement implies that the composite information with a given identifier provides the unique composite information for the sample groups within the track.

[0302] As another example, a media file writer may decide to encapsulate several independent layers in the same track. For example, this encapsulation makes sense when any or all of the independent layers of the track are not present in the set of operation points. These independent layers are expected to be played together.

[0303] In such a case, the composition information described in a sample group entry (e.g., CompositeRegionGroupEntry) applies to all NAL units of a given sample, which may limit the possibility of signaling composition. In fact, each sample of the track contains NAL units of several layers, and when the layers share the same position in the bitstream (possibly with a layer representing an alpha channel to provide some transparency mechanism), the composition information can apply.

[0304] To allow specifying different composition information for each layer in each sample, the NALUMapEntry box is signaled in association with a sample group of type 'nalm'. A sample group of type 'nalm' can set its grouping_type_parameter to 'crif' to indicate that within the 'nalm' sample group, NAL units are mapped to a sample group description box of type 'crif'. This NALUMapEntry will define different groupIDs (i.e., identifiers of groups of NAL units in a sample) for each group of NAL units corresponding to a layer.

[0305] In such a case, the composition information can refer to this groupID value to indicate which layer the composition information applies to. For example, the syntax of CompositeRegionGroupEntry is as follows:

[0306]

[0307] where the syntax elements have the following semantics:

[0308] groupID is the identifier of the composite region group entry described by this sample group entry. It is not the identifier for composition (for the purpose of compositionID). The value of groupID in the composite region group entry should be greater than 0. The value 0 is reserved for special purposes.

[0309] When there is a SampleToGroupBox of type 'nalm' and a grouping_type_parameter equal to 'crif', a SampleGroupDescriptionBox of type 'crif' should be present, and the following applies:

[0310] - The value of groupID in the composite region group entry shall be equal to the groupID in one of the entries of NALUMapEntry.

[0311] - Mapping a NAL unit to groupID 0 via NALUMapEntry means that the NAL unit is required to decode any composite region in the same composite picture as this NAL unit.

[0312] The compositionID is the unique identifier of the composition described by this sample group entry. The value of compositionID in the composite region group entry shall be greater than or equal to 0. When the sample group is associated with several SampleGroupDescriptionBoxes of type 'crif', all these SampleGroupDescriptionBoxes shall have different compositionID values.

[0313] Since groupID is an identifier used to distinguish different composite information of NAL units in a sample, and compositionID is an identifier used to distinguish composite pictures when NAL units are within different operation points, when the scope of a NAL unit is associated with several SampleGroupDescriptionBox entries of type 'crif', all these SampleGroupDescriptionBox entries shall have different compositionID values, but may have the same groupID value.

[0314] A composite_region_flag equal to 1 specifies that the decoded picture of the NAL unit associated with this composite region group entry is signaled in the bitstream as a rectangular region of the composite picture, and further information about the composite region is provided by the subsequent fields in this composite region group entry. A value of 0 specifies that the decoded picture of the NAL unit associated with this composite region group entry is not signaled as a rectangular region of the composite picture, and no further information about the region is provided in this composite region group entry.

[0315] When multiple layer bitstreams are carried in one or more tracks and each layer is a composite region of the same composition, for any two layers layerA and layerB of the bitstream, the following constraints apply: When the NAL unit of layerA is associated with a compositionID value cIdA and a groupId value gIdA where the composition_region_flag is equal to 1, and the NAL unit of layerB is associated with a compositionID value cIdB and a groupId value gIdB where the corresponding composition_region_flag is equal to 1, cIdA and cIdB shall be equal, and gIdA and gIdB may be equal.

[0316] When a track conveys one or more layers that can be used in more than one composition, the corresponding NAL units can be mapped to sample group description entries of type 'crif' (with different compositionIDs but should have the same groupID) to be referenced in the NALUMapEntry.

[0317] composite_horizontal_offset and composite_vertical_offset give the horizontal and vertical offsets of the top-left pixel of the composite region covered by the NAL unit in each composite region associated with the composite region group entry in the luma samples with respect to the top-left pixel of the base region. The base region used in the CompositeRegionGroupEntry is the composite picture.

[0318] composite_region_width and composite_region_height give the width and height of the composite region in the composite picture covered by the NAL unit in each composite region associated with the composite region group entry in the luma samples.

[0319] Note that parts of some layers of the composition may overlap. In the case of overlap, the layer with the higher ID shall be above the layer with the lower ID. The composition of layers may also leave gaps in the resulting output. In warping, the syntax elements of the composition information specify the order of overlap. The order of overlap is the increasing order of the values of this syntax element.

[0320] In a third variation, the VVC operation point record may refer to composition. The structure defining the operation point is extended to also indicate an identifier of the composition associated with the operation point. The composition may be signaled in the bitstream for encapsulation (e.g., a composition SEI message), but it may also correspond to a media file edited by a user or composition software. For the encapsulation of VVC NAL units in a file format, three boxes may indicate the operation points present in the bitstream. These boxes are VvcOperatingPointsRecord, OperatingPointGroupBox, or VvcDecoderConfigurationRecord. In one embodiment of the present invention, one or more of these boxes refer to the composition identifiers of one or more operation points defined in the box.

[0321] For example, the VvcOperatingPointsRecord may have the following syntax elements:

[0322]

[0323]

[0324]

[0325] The semantics of the new syntax elements added in the VvcOperatingPointsRecord box have the following semantics:

[0326] When recommended_composition_ols_flag equals 1, it specifies that the layers of the operation point include a number of output layers that are recommended to be presented as a composition. When recommended_composition_ols_flag equals 0, it specifies that the layers of the operation point are not associated with a composition recommendation.

[0327] composition_group_ID specifies the value of the composition region group entry associated with the set of output layers.

[0328] When recommended_composition_ols_flag equals 1 for an operation point, the requirements of the specification are that the following evaluations are all valid:

[0329] · The ISOBMFF file contains a composite region group entry (CompositeRegionGroupEntry) for the NAL units of the layers (where is_outputlayer equals 1 for the output layers of the operation point); and

[0330] · The composition_ID of the CompositeRegionGroupEntry should be equal to the composition_group_ID of the operating point as defined in the VVCOperatingPointRecord box.

[0331] These statements ensure that the composite information is defined in the sample group entry on the track that contains data of the operating point indicating that several output layers are recommended to be presented as a composite.

[0332] In a variation of this embodiment, the VVCOperatingPointRecord box may also define additional information for the composite picture. In particular, for each operating point, the VVCOperatingPointRecord signals the height and width of the composite picture. It may also signal whether a transform is applied to the decoded pictures of the individual layers of the operating point. For example, a flag may indicate whether an upsampling / downsampling operation is applied. Another flag may indicate that a rotation transform is required. Another flag may indicate whether the decoded pictures of at least two individual layers overlap in the composite picture. It may also indicate any unit size for signaling the position and size of the composite region in the composite picture. The VVCOperatingPointRecord may also signal that the composite picture should be scaled by specific horizontal and vertical scaling factors. Additionally, the translation of the resulting composite picture may be signaled.

[0333] In a fourth variation, the VVC operating point entities to be grouped may refer to a composite ID. When the composite identifier is associated with the operating point of the OperatingPointGroupBox, the syntax may be as follows, for example:

[0334]

[0335]

[0336]

[0337] wherein, it has the following semantics:

[0338] recommended_composition_ols_flag being equal to 1 specifies that the layers of the operating point include several output layers that are recommended to be presented as a composite. recommended_composition_ols_flag being equal to 0 specifies that the layers of the operating point are not associated with a composite recommendation.

[0339] composition_group_ID specifies the value of the composite region group entry associated with the set of output layers.

[0340] When recommended_composition_ols_flag is equal to 1 for an operation point, the requirements for the specification are that the following evaluations are all valid:

[0341] · The ISOBMFF file contains a CompositeRegionGroupEntry for the NAL units of the layer (which is the output layer of this operation point (is_outputlayer is equal to 1)); and

[0342] · The group_ID of the CompositeRegionGroupEntry shall be equal to the composition_group_ID of the operation point defined in the OperatingPointGroupBox box.

[0343] These statements ensure that the composite information is defined in the sample group entry on the track that contains data for the operation point indicating that several output layers are recommended to be presented as a composite.

[0344] In a variation of this embodiment, the OperatingPointGroupBox box can also define additional information for the composite picture. In particular, for each operation point, the OperatingPointGroupBox signals the height and width of the composite picture. It can also signal whether a transform is applied to the decoded pictures of the individual layers of the operation point. For example, a flag can indicate whether an upsampling / downsampling operation is applied. Another flag can indicate that a rotation transform is required. Another flag can indicate whether the decoded pictures of at least two individual layers overlap in the composite picture. It can also indicate any unit size for signaling the position and size of the composite region in the composite picture. The OperatingPointGroupBox can also signal that the composite picture should be scaled by specific horizontal and vertical scaling factors. Additionally, the translation of the resulting composite picture can be signaled.

[0345] In a fifth variation, the VvcDecoderConfigurationRecord box can refer to a composite ID. When the composite identifier is associated with an operation point of the VvcDecoderConfigurationRecord, the syntax can be, for example, as follows:

[0346]

[0347] where the following semantics apply to the new syntax elements:

[0348] When recommended_composition_ols_flag is equal to 1, it specifies that the output layer set with an index equal to output_layer_set_idx includes a number of output layers recommended to be presented as a composition. When recommended_composition_ols_flag is equal to 0, it specifies that the output layer set with an index equal to output_layer_set_idx is not associated with a composition recommendation.

[0349] composition_group_ID specifies the value of the identifier of the composition area group entry associated with the output layer set.

[0350] In a variation of this embodiment, the VvcDecoderConfigurationRecord box may also define additional information for the composition picture. In particular, for an operation point, the VvcDecoderConfigurationRecord signals the height and width of the composition picture. It may also signal whether the composition picture applies a transform to the decoded pictures of the individual independent layers of the operation point. For example, a flag may indicate whether an upsampling / downsampling operation is applied. Another flag may indicate that a rotation transform is required. Another flag may indicate whether the decoded pictures of at least two independent layers overlap in the composition picture. It may also indicate any unit size for signaling the position and size of the composition area in the composition picture. The VvcDecoderConfigurationRecord may also signal that the composition picture should be scaled by specific horizontal and vertical scaling factors. Additionally, the translation of the resulting composition picture may be signaled.

[0351] The combination of 'ropi' and 'crif'

[0352] In one embodiment, a media file writer or an intermediate node of a network filters or restricts (i.e., it removes the encoded data) an ISOBMFF media file that contains multiple independent layers with composition information. Compared with the list of operation points that may be declared in a 'vopi' or 'opeg' structure, this filtering or restriction results in a restricted set of OLSs present in the file. In such a case, a restricted operation point information sample group may be signaled to indicate the restricted set of OLSs accessible in the media file. The media file writer may signal a composition area group entry ('crif') to provide composition information for the independent layers. In such a case, the media file writer may remove composition area group entries that are not part of the restricted or filtered set of OLSs.

[0353] In the transform, the media file parser determines the OLS restriction information from the composition region group entry. When an operation point associated with the composition information signals a restriction, the media file parser determines whether the restriction applies to one of the layers or sublayers of the OLS. When a restriction is made for at least one layer of the output layer set (for an operation point, restricted_ols_flag equals 1), the signaled composition information is considered invalid and the operation point is ignored in step 603. On the other hand, if the restriction is for a sublayer (restricted_sublayer_flag equals 1 and restricted_ols_flag equals 0), the composition information is considered valid and the operation point can be selected in step 603. In practice, when the sublayer of an operation point has been filtered, since the size of the decoded picture remains the same, synthesis can still be performed, only the frame rate of the reconstructed image can be changed.

[0354] In the transform, the signaling of the composition information depends on the restriction information. For example, in one embodiment, the VvcOperatingPointsRecord syntax includes the following additional elements (bold):

[0355]

[0356] The semantics of the syntax elements are the same as in the previous embodiment, and the main difference is that the existence of the syntax element associating the composition information with the operation point depends on the restricted_ols_flag. For example, the composition_group_ID syntax element only exists for a valid OLS (i.e., when restricted_ols_flag is not equal to 0).

[0357] Figure 7Shows an example of an ISOBMFF file 705 according to an embodiment of the proposed method, where there are more than one independent layer in the same bitstream, and the layers are in an output layer set. The bitstream also provides syntax elements representing the composition of decoded pictures of multiple independent layers to non-VCL NAL units. In this example, the first track 730 of the media file 705 includes several independent layers. The composition information associated with the layers is specified in the CompositeRegionGroupEntry boxes 720-1 and 720-2. These CompositeRegionGroupEntry boxes include the same composition_ID value (i.e., the independent layers are part of the same operation point) and different groupID values. The groupId value is the identifier of the 'crif' entry, and the range of NAL units corresponding to each independent layer as described in the NALUMapEntry box 700 is mapped to this 'crif' entry. The sampleToGroupBox 710 defines the samples of the track for which the NAL unit mapping (700) is defined and to which the composition information (720-1, 720-2...) can be applied. As a result, the ISOBMFF generated by the media file writer allows signaling of the composition information (such as SEI messages) provided in the non-VCL NAL units of the VVC bitstream.

[0358] Figure 8 Is a schematic block diagram of a computing device 800 for implementing one or more embodiments of the present invention. The computing device 800 can be a device such as a microcomputer, a workstation, or a lightweight portable device.

[0359] The computing device 800 includes a communication bus that is connected to:

[0360] - A central processing unit 801 labeled CPU, such as a microprocessor;

[0361] - A random access memory 802 labeled RAM, for storing the executable code of the method of the embodiments of the present invention and registers suitable for recording variables and parameters that are necessary for implementing the method according to the embodiments of the present invention. The memory capacity of this random access memory can be extended, for example, by an optional RAM connected to an expansion port;

[0362] - A read-only memory 803 labeled ROM, for storing computer programs for implementing the embodiments of the present invention;

[0363] - A network interface (NET) 804, which is typically connected to a communication network and through which digital data to be processed is sent or received. The network interface 804 can be a single network interface or consist of a collection of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Under the control of a software application running in the CPU 801, data packets are written to the network interface for transmission, or data is read from the network interface for reception;

[0364] - A user interface (IHM) 805 can be used to receive input from a user or display information to the user;

[0365] - A hard disk (HD) 806 labeled HD can be set as a large storage device;

[0366] - An I / O module (IO) 807 can be used to receive / send data from / to external devices (such as a video source or a display, etc.).

[0367] The executable code can be stored either in the read-only memory 803, on the hard disk 806, or can be stored on, for example, a removable digital medium (such as a disk, etc.). According to a variant, the executable code of a program can be received via the network interface 804 by means of a communication network so that the executable code of the program is stored in one of the storage components (such as the hard disk 806, etc.) of the communication device 800 before being executed.

[0368] The central processing unit 801 is adapted to control and direct the execution of instructions or portions of software code of one or more programs according to an embodiment of the present invention, and these instructions are stored in one of the aforementioned storage components. After power-on, the CPU 801 is capable of executing these instructions after, for example, loading instructions related to the software application from the main RAM memory 802 from the program ROM 803 or the hard disk (HD) 806. When such a software application is executed by the CPU 801, the steps of the flowchart of the present invention are performed.

[0369] Any step of the algorithm of the present invention can be implemented in software by a programmable machine (such as a PC (“personal computer”), a DSP (“digital signal processor”) or a microcontroller, etc.) executing instructions or a set of programs; or implemented in hardware by a machine or a dedicated component (such as an FPGA (“field programmable gate array”) or an ASIC (“application specific integrated circuit”), etc.).

[0370] Although the present invention has been described above with reference to specific embodiments, the present invention is not limited to the specific embodiments, and modifications within the scope of the present invention will be apparent to those skilled in the art.

[0371] Many further modifications and suggestions for variations will occur to those skilled in the art in reference to the foregoing illustrative embodiments, which are given by way of example only and are not intended to limit the scope of the invention, which is determined only by the appended claims. In particular, different features from different embodiments may be interchanged where appropriate.

[0372] The various embodiments of the present invention described above may be implemented individually or as a combination of multiple embodiments. Additionally, features from different embodiments may be combined when necessary or combinations of elements or features from separate embodiments may be combined when beneficial in a single embodiment.

[0373] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are defined in mutually different dependent claims does not indicate that a combination of these features cannot be used advantageously.

Claims

1. A method for encapsulating a bitstream of encoded media data in an ISOBMFF media file, the method comprising: Obtaining the bitstream, wherein the media data in the bitstream is organized as NAL units; Obtaining NAL units from the bitstream that describe a plurality of operation points; Generating the ISOBMFF media file, the ISOBMFF media file including the bitstream and including at least one metadata box that describes the plurality of operation points, wherein the at least one metadata box includes information for indicating whether a part of the NAL units associated with at least one of the plurality of operation points is missing in the ISOBMFF media file.

2. The method according to claim 1, wherein The information indicates whether all of the plurality of operation points have all of their NAL units in the ISOBMFF media file.

3. The method according to claim 1, wherein The at least one metadata box includes information for each of the plurality of operation points indicating whether that operation point has all of its NAL units in the ISOBMFF media file.

4. The method according to claim 1, wherein, The at least one metadata box includes information for each of the plurality of operation points indicating whether the temporal sublayer of that operation point is missing.

5. The method according to claim 1, wherein The at least one metadata box includes information for each of the plurality of operation points indicating whether the layer of that operation point is missing.

6. The method according to claim 1, wherein The at least one metadata box includes information for each of the plurality of operation points indicating whether the layer and the temporal sublayer of that operation point are missing.

7. The method according to claim 1, wherein The metadata box represents a 'vopi' sample group.

8. The method according to claim 1, wherein, The metadata box is an 'opeg' box.

9. The method according to claim 1, wherein Obtaining the bitstream includes: Obtaining an original bitstream, wherein the video data in the original bitstream is organized as NAL units and includes an original plurality of operation points; Obtaining NAL units from the original bitstream that describe a second plurality of operation points, the original plurality of operation points including at least the second plurality of operation points; Obtaining the bitstream by filtering the original bitstream based on the NAL units that describe the second plurality of operation points.

10. The method according to claim 1, wherein The NAL units that describe the plurality of operation points are of the video parameter set type, i.e., VPS type.

11. The method according to claim 9, wherein, The NAL units that describe the second plurality of operation points are operation point information, i.e., OPI.

12. The method according to claim 9, wherein, The NAL units that describe the original plurality of operation points are of the video parameter set type, i.e., VPS type.

13. A method for decoding a bitstream of encoded media data in an ISOBMFF media file, the method comprising: Obtaining the bitstream from the ISOBMFF media file, wherein the media data in the bitstream is organized as NAL units; Obtaining at least one metadata box from the ISOBMFF media file that describes a plurality of operation points; Decoding the operation points in the bitstream, wherein: The at least one metadata box includes information for indicating whether a part of the NAL units associated with at least one of the plurality of operation points is missing in the ISOBMFF media file; and The decoded operation points are the operation points that have all of their NAL units in the ISOBMFF media file.

14. A non - transitory computer - readable storage medium storing a program for a programmable device, the program including a sequence of instructions which, when loaded into and executed by the programmable device, implement the method according to claim 1.

15. An apparatus for encapsulating a bitstream of encoded media data in an ISOBMFF media file, the apparatus including a processor configured to: obtain the bitstream, wherein the media data in the bitstream is organized as NAL units; obtain NAL units from the bitstream that describe a plurality of operation points; generate the ISOBMFF media file, the ISOBMFF media file including the bitstream and including at least one metadata box that describes the plurality of operation points, Among them, the at least one metadata box including information for indicating whether a portion of the NAL units associated with at least one of the plurality of operation points is missing in the ISOBMFF media file.

16. An apparatus for decoding a bitstream of encoded media data in an ISOBMFF media file, the apparatus including a processor configured to: obtain the bitstream from the ISOBMFF media file, wherein the media data in the bitstream is organized as NAL units; obtain at least one metadata box from the ISOBMFF media file that describes a plurality of operation points; decode the operation points in the bitstream, wherein: the at least one metadata box includes information for indicating whether a portion of the NAL units associated with at least one of the plurality of operation points is missing in the ISOBMFF media file; and the decoded operation points are those for which all NAL units are in the ISOBMFF media file.

Citation Information

Patent Citations

  • Method and apparatus for coding multilayer video, and method and apparatus for decoding multilayer video

    CN104620587A

  • Conformance and inoperability improvements in multi-layer video coding

    CN106464922A