Versatile video coding track coding

By redefining the VVC non-VCL track, allowing APS and Picture Header NAL units to be stored in separate tracks, and DCI signaling in the track-level box, the non-VCL elementary stream contains only non-VCL NAL units, thus solving the signaling notification and storage problems in the VVC video file format and achieving more efficient video file management and decoder operation.

CN114205599BActive Publication Date: 2026-02-13FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111090652.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-17
Filing Date
2021-09-17
Publication Date
2026-02-13
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

The existing VVC video file format has problems with signaling notification and storage, including unclear definitions of VVC reference tracks and non-VCL tracks, APS NAL units cannot be stored in multiple tracks, DCI NAL units are not considered in the video elementary stream, non-VCL elementary streams may contain VCL NAL units, decoder configuration information sample group signaling is complicated, and OPI NAL units are not allowed to be included in sample entries.

Method used

The VVC non-VCL track is redefined to contain only non-VCL NAL units, allowing APS and picture header NAL units to be stored in separate tracks. DCI NAL units are signaled in track-level boxes. The non-VCL elementary stream contains only non-VCL NAL units. The decoder configuration information sample group allows samples to share the same DCI. OPI NAL units can be included in the sample entry description.

Benefits of technology

It achieves optimized storage and signaling notification in the VVC video file format, supports post-banding of multi-layer bitstreams and efficient delivery of tracks, and simplifies decoder configuration and operation point management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114205599B_ABST
    Figure CN114205599B_ABST
Patent Text Reader

Abstract

This application relates to a general video coding track coding, describes a system, method and apparatus for encoding or decoding a file format storing one or more images. One example method includes performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a condition under which an item of control information is included in a non-video coding layer track of the visual media file, and wherein a presence of the non-video coding layer track in the visual media file is indicated by a particular track reference in a video coding layer track of the visual media file.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is made to promptly claim priority and interest in U.S. Provisional Patent Application No. 63 / 079,869, filed September 17, 2020, in accordance with applicable patent law and / or the rules of the Paris Convention. For all purposes under the law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to the generation, storage, and consumption of digital audio and video media information in file formats. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process video or image representations according to file formats.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a visual media file and a bitstream of visual media data according to format rules, wherein the format rules specify conditions for whether control information items are included in non-video codec layer tracks of the visual media file, and wherein the presence of non-video codec layer tracks in the visual media file is indicated by a specific track reference in the video codec layer tracks of the visual media file.

[0007] In another example, a video processing method is disclosed. This method includes performing a conversion between a visual media file and a bitstream of visual media data according to format rules, wherein the format rules specify the type of sample entries that determines whether a decoding capability information network abstraction layer unit is included in the sample entries of the video track in the visual media file or in both the sample entries of the video track and the sample entries of the video track in the visual media file.

[0008] In another example, a video processing method is disclosed. The method includes performing a conversion between visual media data and a file storing information corresponding to the visual media data according to format rules; wherein the format rules specify a first condition for identifying the non-video codec layer (VCL) track of the file and / or a second condition for identifying the VCL track of the file.

[0009] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0010] In yet another example aspect, a video decoder apparatus is disclosed. The video decoder comprises a processor configured to implement a method described above.

[0011] In yet another example aspect, a computer readable medium storing code is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.

[0012] In yet another example aspect, a computer readable medium storing a bitstream is disclosed. The bitstream is generated or processed using the methods described in this document.

[0013] These, additional, and other aspects are described throughout the present document. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 is a block diagram of an example video processing system.

[0015] Figure 2 is a block diagram of a video processing apparatus.

[0016] Figure 3 is a flowchart of an example method of video processing.

[0017] Figure 4 is a block diagram illustrating a video coding system according to some embodiments of the present disclosure.

[0018] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0019] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0020] Figure 7 An example of an encoder block diagram is shown.

[0021] Figures 8 to 9 is a flowchart of an example method of video processing. DETAILED DESCRIPTION

[0022] For ease of understanding, section headings are used in this document, and the applicability of techniques and embodiments disclosed in each section is not limited to that section only. Also, the use of H.266 terminology in some descriptions is merely for ease of understanding, and is not intended to limit the scope of the disclosed techniques. As such, the techniques described herein are applicable to other video codec protocols and designs as well. In this document, editorial changes to text are shown by left and right double brackets that indicate text between the double brackets is deleted (e.g., [[ ]]) and by bold italic text that indicates added text.

[0023] 1. Brief discussion

[0024] This document relates to video file formats. In particular, this document relates to signaling and storage of picture header (PH), adaptation parameter set (APS), decoding capability information (DCI), and operating point information (OPI) network abstraction layer (NAL) units for Versatile Video Coding (VVC) video bitstreams in ISO base media file format (ISOBMFF) based media files. These ideas can be applied to video bitstreams coded by any codec (e.g., the VVC standard) and to any video file format (e.g., the VVC video file format under development), either individually or in various combinations.

[0025] 2. Abbreviations

[0026] ACT adaptive color transform

[0027] ALF adaptive loop filter

[0028] AMVR adaptive motion vector resolution

[0029] APS adaptation parameter set

[0030] AU access unit

[0031] AUD access unit delimiter

[0032] AVC advanced video coding (Rec. ITU-T H.264 | ISO / IEC 14496-10)

[0033] B bi-prediction

[0034] BCW bi-prediction with CU-level weights

[0035] BDOF bi-directional optical flow

[0036] BDPCM block-based delta pulse code modulation

[0037] BP buffering period

[0038] CABAC context-based adaptive binary arithmetic coding

[0039] CB coded block

[0040] CBR constant bit rate

[0041] CCALF cross-component adaptive loop filter

[0042] CPB coded picture buffer

[0043] CRA clean random access

[0044] CRC cyclic redundancy check

[0045] CTB coded tree block

[0046] CTU coded tree unit

[0047] CU coding unit

[0048] CVS coded video sequence

[0049] DPB decoded picture buffer

[0050] DCI decoding capability information

[0051] DRAP dependent random access point

[0052] DU decoding unit

[0053] DUI decoding unit information

[0054] EG exponential Golomb

[0055] EGk k-th order exponential Golomb

[0056] EOB end of bitstream

[0057] EOS end of sequence

[0058] FD filler data

[0059] FIFO first in, first out

[0060] FL fixed length

[0061] GBR green, blue, and red

[0062] GCI general constraint information

[0063] GDR gradual decoding refresh

[0064] GPM geometric partition mode

[0065] HEVC high efficiency video coding (Rec. ITU-T H.265 | ISO / IEC 23008-2)

[0066] HRD hypothetical reference decoder

[0067] HSS hypothetical stream scheduler

[0068] I intra

[0069] IBC intra block copy

[0070] IDR instantaneous decoding refresh

[0071] ILRP inter-layer reference picture

[0072] IRAP intra random access point

[0073] LFNST low-frequency non-separable transform

[0074] LPS least probable symbol

[0075] LSB least significant bit

[0076] LTRP long-term reference picture

[0077] LMCS luma mapping with chroma scaling

[0078] MIP matrix-based intra prediction

[0079] MPS most probable symbol

[0080] MSB most significant bit

[0081] MTS multiple transform selection

[0082] MVP motion vector prediction

[0083] NAL network abstraction layer

[0084] OLS output layer set

[0085] OP operation point

[0086] OPI operation point information

[0087] P prediction

[0088] PH picture header

[0089] POC picture order count

[0090] PPS picture parameter set

[0091] PROF prediction refinement using optical flow

[0092] PT picture timing

[0093] PU picture unit

[0094] QP quantization parameter

[0095] RADL random access decodable leading (picture)

[0096] RASL random access skipped leading (picture)

[0097] RBSP raw byte sequence payload

[0098] RGB red, green, and blue

[0099] RPL reference picture list

[0100] SAO sample adaptive offset

[0101] SAR sample aspect ratio

[0102] SEI supplemental enhancement information

[0103] SH slice header

[0104] SLI subpicture level information

[0105] SODB string of data bits

[0106] SPS sequence parameter set

[0107] STRP short-term reference picture

[0108] STSA stepping time dermined sublayer access

[0109] TR truncated rice

[0110] VBR variable bit rate

[0111] VCL video coding layer

[0112] VPS video parameter set

[0113] VSEI general supplemental enhancement information (Rec. ITU-T H.274 | ISO / IEC 23002-7)

[0114] VUI video usability information

[0115] VVC versatile video coding (Rec. ITU-T H.266 | ISO / IEC 23090-3)

[0116] 3. Introduction to video coding

[0117] 3.1. Video coding standards

[0118] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure, where temporal prediction plus transform coding is utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and put into the reference software named Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project was officially started, the JVET was later renamed as Joint Video Expert Team (JVET). VVC is the new coding standard finalized by the JVET at its 19th meeting on July 1, 2020, with the goal of 50% bitrate reduction compared to HEVC.

[0119] The Versatile Video Coding (VVC) standard (Rec. ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplemental Enhancement Information (VSEI) standard (Rec. ITU-T H.274 | ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcast, video conferencing or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple coded video bitstreams, multi-view video, scalable layered coding and viewport-adaptive 360° immersive media.

[0120] 3.2. File format standards

[0121] Media streaming applications are typically based on IP, TCP and HTTP transport methods and typically rely on file formats such as ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). In order to use video formats with ISOBMFF and DASH, a video format specific file format specification (such as AVC File Format and HEVC File Format) is needed for encapsulating video content in ISOBMFF tracks and DASH representations and segments. Important information about the video bitstream (e.g. profiles, tiers and levels, and many other information) will need to be exposed as file format level metadata and / or DASH media presentation description (MPD) for content selection purposes, e.g. for selecting appropriate media segments for initialization at the start of a streaming session and for stream adaptation during the streaming session.

[0122] Similarly, in order to use image formats with ISOBMFF, image format specific file format specifications such as AVC Image File Format and HEVC Image File Format will be needed.

[0123] The VVC Video File Format (a file format for storing VVC video content based on ISOBMFF) is currently being developed by MPEG.

[0124] The VVC Image File Format (a file format for storing image content coded using VVC based on ISOBMFF) is currently being developed by MPEG.

[0125] 3.3. PH, APS, DCI and OPI NAL units in VVC

[0126] Some new types of NAL units have been introduced in VVC, including PH, APS, DCI and OPI NAL units.

[0127] 3.3.1. Adaptation Parameter Set (APS)

[0128] Adaptive parameter set (APS) conveys picture-level and / or slice-level information that can be shared across multiple slices of a picture and / or slices of different pictures, but can change frequently across pictures, and the total number of variants can be high, thus not suitable to be included into PPS. Three types of parameters are included in APS: adaptive loop filter (ALF) parameters, luma mapping with chroma scaling (LMCS) parameters, and scaling list parameters. APS can be carried in two different NAL unit types, either in front of or after the associated slices as a prefix or suffix. The latter can be helpful in ultra-low delay scenarios, e.g., allowing the encoder to send slices of a picture before generating ALF parameters based on the picture that will be used by subsequent pictures in decoding order.

[0129] 3.3.2. Picture header (PH)

[0130] A picture header (PH) structure exists for each PU. The PH is either in a separate PH NAL unit or included in the slice header (SH). If a PU consists of only one slice, the PH can only be included in the SH. To simplify the design, within a CLVS, the PH can only be all in PH NAL units or all in SHs. When the PH is in the SH, there is no PH NAL unit in the CLVS.

[0131] The PH is designed for two purposes. First, to help reduce the signaling overhead of the SHs of a picture that contains multiple slices by carrying all the parameters that have the same value for all slices of the picture, thus not repeating the same parameters in each SH. These include IRAP / GDR picture indication, inter / intra slice allowed flags, and information related to POC, RPL, deblocking filter, SAO, ALF, LMCS, scaling list, QP delta, weighted prediction, coding block partitioning, virtual boundary, collocated picture, etc. Second, to help the decoder identify the first slice of each coded picture that contains multiple slices. Since there is one and only one PH for each PU, when the decoder receives a PH NAL unit, it knows easily that the next VCL NAL unit is the first slice of the picture.

[0132] 3.3.3. Decoding capability information (DCI)

[0133] The DCI NAL unit contains bitstream-level PTL information. The DCI NAL unit includes one or more PTL syntax structures, which can be used during session negotiation between the sender and receiver of a VVC bitstream. When a DCI NAL unit is present in a VVC bitstream, each output layer set (OLS) in the CVS of the bitstream shall conform to the PTL information carried in at least one of the PTL structures in the DCI NAL unit.

[0134] In AVC and HEVC, the PTL information for session negotiation is available in SPS (for HEVC and AVC) and VPS (for HEVC layered extension). This design of conveying PTL information for session negotiation in HEVC and AVC has drawbacks, as the scope of SPS and VPS is within a CVS, not the whole bitstream. Therefore, sender-receiver session initiation can suffer from re-initiation during bitstream streaming of each new CVS. DCI solves this problem as it carries bitstream level information, thus, it can guarantee the decoding capability as indicated until the end of the bitstream.

[0135] 3.3.4. Operation point information (OPI)

[0136] The decoding process of HEVC and VVC has similar input variables to set the decoding operation point, i.e. the target OLS and the highest sublayer of the bitstream to be decoded, through the decoder API. However, in scenarios where the layers and / or sublayers of the bitstream are removed during transmission or the device does not expose the decoder API to the application, it can happen that the decoder cannot be correctly informed of the operation point of the decoder processing a given bitstream. Therefore, the decoder can not be able to conclude on the properties of the pictures in the bitstream, e.g. the correct buffer allocation of the decoded pictures and whether to output individual pictures or not. To solve this problem, VVC adds a mode in the bitstream to indicate these two variables through a newly introduced operation point information (OPI) NAL unit. In the AUs at the beginning of the bitstream and their individual CVSs, the OPI NAL unit informs the decoder of the target OLS and the highest sublayer of the bitstream to be decoded.

[0137] In case there is an OPI NAL unit and the operation point is also provided to the decoder via the decoder API information (e.g. the application can have more updated information on the target OLS and sublayers), the decoder API information takes precedence. In case there is no decoder API and any OPI NAL unit in the bitstream, a suitable fallback is specified in VVC to allow correct decoder operation.

[0138] 3.4. Some details on the VVC video file format

[0139] 3.4.1. Type of tracks

[0140] The VVC video file format specifies the following types of video tracks that carry VVC bitstreams in an ISOBMFF file:

[0141] a) VVC track:

[0142] A VVC track represents a VVC bitstream by including NAL units in its sample and sample entries, and possibly by referencing other VVC tracks containing other sublayers of the VVC bitstream, and possibly by referencing VVC subpicture tracks. When a VVC track references a VVC subpicture track, it is called a VVC reference track.

[0143] b) VVC non-VCL tracks:

[0144] APS carrying ALF, LMCS or scaling list parameters and other non-VCL NAL units can be stored in and transported through a track separate from the track containing VCL NAL units; this is the VVC non-VCL track.

[0145] c) VVC subpicture tracks:

[0146] A VVC subpicture track contains any of the following:

[0147] A sequence of one or more VVC subpictures.

[0148] A sequence of one or more complete slices forming a rectangular region.

[0149] A sample of a VVC subpicture track contains any of the following:

[0150] One or more complete subpictures that are consecutive in decoding order as specified in ISO / IEC 23090-3.

[0151] One or more complete slices forming a rectangular region and being consecutive in decoding order as specified in ISO / IEC 23090-3.

[0152] The VVC subpictures or slices included in any sample of a VVC subpicture track are consecutive in decoding order.

[0153] NOTE: VVC non-VCL tracks and VVC subpicture tracks enable optimal delivery of VVC video in streaming applications as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, the client can request the DASH representation containing the subset of VVC subpicture tracks and the DASH representation containing the non-VCL tracks piece-wise. In this way, redundant transport of APS and other non-VCL NAL units can be avoided.

[0154] 3.4.2. VVC elementary stream structure

[0155] Three types of elementary streams are defined for storing VVC content:

[0156] Video elementary stream, does not contain any parameter sets; all parameter sets are stored in the sample entry or sample entries;

[0157] Video and parameter set elementary stream, can contain parameter sets and can also store parameter sets in its sample entry or sample entries;

[0158] Non-VCL elementary stream, contains non-VCL NAL units that are synchronized with the elementary streams carried in the video track.

[0159] Note: VVC non-VCL tracks do not contain parameter sets in their sample entries.

[0160] 3.4.3. Decoder configuration information sample group

[0161] 3.4.3.1. Definition

[0162] The sample group description entry of this sample group contains the DCI NAL unit. All samples mapped to the same decoder configuration information sample group description entry belong to the same VVC bitstream.

[0163] This sample group indicates whether the same DCI NAL unit is used for different sample entries in the VVC track, i.e. whether the samples belonging to different sample entries belong to the same VVC bitstream. When the samples of two sample entries are mapped to the same decoder configuration information sample group description entry, the player can switch sample entries without reinitializing the decoder.

[0164] If any DCI NAL unit is present in any sample entry or in-band, this unit shall be exactly the same as the DCI NAL unit included in the decoder configuration information sample group.

[0165] 3.4.3.2. Syntax

[0166] class DecoderConfigurationInformation extends VisualSampleGroupEntry

[0167] ('dcfi'){

[0168] unsigned int(16) dciNalUnitLength;

[0169] bit(8 * nalUnitLength) dciNalUnit;

[0170] }

[0171] 3.4.3.3. Semantics

[0172] dciNalUnitLength indicates the byte length of the DCI NAL unit.

[0173] dciNalUnit contains a DCI NAL unit as specified in ISO / IEC 23090-3.

[0174] 4. Example technical problems solved by the disclosed technical solutions

[0175] The latest design of the signaling of PH, APS, DCI and OPI NAL units in VVC video file format has the following problems:

[0176] 1) Neither VVC base track nor VVC non-VCL track shall contain VCL NAL units. However, the current definition of VVC non-VCL track would also apply to VVC base track. Moreover, according to the current definition, a VVC non-VCL track always contains APS NAL units. However, this would not allow a non-VCL NAL unit to contain picture header NAL units and possibly other non-VCL NAL units, but not including APS NAL units. Allowing such a VVC non-VCL track would enable optimal storage of extractable sub-picture based single-layer bitstreams in a file for late-banding of sub-picture tracks when different sub-pictures use different APS sets, e.g., by having one PH track (as a non-VCL track, although it contains the same information as a VVC base track), multiple APS tracks (as VVC non-VCL tracks) and multiple VVC sub-picture tracks each containing a sequence of sub-pictures.

[0177] 2) APS NAL units are all stored in one VVC non-VCL track or VVC track. In other words, APS NAL units cannot be stored in more than one track. This applies to APS NAL units containing LMCS parameters (i.e., LMCS APS) or APS NAL units containing scaling list (SL) parameters (i.e., SL APS), but not to APS NAL units containing ALF parameters (i.e., ALF APS). Since different VVC sub-picture tracks can use different sets of ALF APS, it is desirable to enable multiple VVC non-VCL tracks to carry the ALF APS of a VVC bitstream.

[0178] 3) DCI NAL units are not considered in the definition of video elementary stream and video and parameter set elementary stream. Therefore, a video elementary stream does not contain parameter sets, but can contain DCI NAL units.

[0179] 4) The definition of non-VCL elementary stream does not exclude the possibility of containing VCL NAL units in a non-VCL elementary stream.

[0180] 5) The decoder configuration information sample group provides a mechanism for signaling of the DCI NAL unit. However, there are the following problems:

[0181] a. In the most common use case, all samples of a track will belong to the same bitstream (or share the same DCI regardless of the number of bitstreams). For this case, it is complex to figure out the applicable DCI through the sample group signaling.

[0182] b. It is stated that all samples that map to the same decoder configuration information sample group description entry belong to the same VVC bitstream. However, this does not allow samples that belong to multiple VVC bitstreams (e.g., determined by EOB NAL units) but are in the same track to share the same DCI NAL unit even when they can.

[0183] 6) OPI NAL units are not allowed to be included in the sample entry description. However, in many cases, when OPI NAL units are present in a VVC bitstream, the OPI NAL units should be similarly considered as parameter sets, so they should be allowed to be included in the sample entry description.

[0184] 5. Example solutions and embodiments

[0185] To solve the above problems and other problems, the following methods summarized below are disclosed. These items should be considered as examples to explain the general concept and should not be interpreted in a narrow sense. Furthermore, these items can be applied individually or in any combination.

[0186] 1) To solve problems 1 and 2, one or more of the following items are proposed:

[0187] a. A VVC non-VCL track is defined as a track containing only non-VCL NAL units and is referred to by a VVC track through the "vvcN" track reference.

[0188] b. It is specified that a VVC non-VCL track can contain APSs carrying ALF, LMCS, or scaling list parameters stored in and transported through a separate track from the track containing VCL NAL units, with or without other non-VCL NAL units.

[0189] c. It is specified that a VVC non-VCL track can also contain picture header NAL units stored in and transported through a separate track from the track containing VCL NAL units, with or without APS NAL units, and with or without other non-VCL NAL units.

[0190] d. It is specified that picture header NAL units of a video stream can be stored in the samples of a VVC track or the samples of a VVC non-VCL track, but not in both.

[0191] 2) To solve problem 3, one or more of the following are proposed:

[0192] a. A video elementary stream is defined as an elementary stream that contains VCL NAL units and does not contain any parameter set, DCI or OPI NAL units; all parameter set, DCI and OPI NAL units are stored in sample entries.

[0193] i. Alternatively, a video elementary stream is defined as an elementary stream that contains VCL NAL units and does not contain any parameter set or DCI NAL units; all parameter set and DCI NAL units are stored in sample entries.

[0194] b. DCI NAL units are treated exactly the same as parameter sets, i.e., DCI NAL units can be only in the sample entries of a video track (e.g., when the sample entry type name is “vvc1”), or can be in one or both of the samples and sample entries of a video track (e.g., when the sample entry type name is “vvi1”).

[0195] 3) To solve problem 4, it is specified that a non-VCL elementary stream is an elementary stream that contains only non-VCL NAL units, and these non-VCL NAL units are synchronized with the elementary streams carried in the video track.

[0196] 4) To solve problem 5, one or more of the following are proposed:

[0197] a. For the case that all samples of a track belong to the same bitstream (or share the same DCI regardless of the number of bitstreams), DCI NAL units can be signaled in a track level box (e.g., track header box, track level meta box or another track level box).

[0198] b. Allow samples that belong to multiple VVC bitstreams (e.g., determined by EOB NAL units) but in the same track to belong to the same decoder configuration information sample group, and thus share the same decoder configuration information sample group description entry.

[0199] 5) To solve problem 6, allow OPI NAL units to be included in a sample entry description, e.g., as one of the non-VCL NAL unit arrays in the decoder configuration record.

[0200] a. Alternatively, the OPI NAL unit is considered to be identical to the parameter set, i.e., the OPI NAL unit can be only in the sample entry of the video track (e.g., when the sample entry type name is “vvc1”), or can be in one or both of the sample and the sample entry of the video track (e.g., when the sample entry type name is “vvi1”).

[0201] 6. Embodiments

[0202] Below are some example embodiments of the inventive aspects summarized in Section 5 above, which can be applied to the standard specification of the VVC video file format. The changed text is based on the latest draft specification. Most of the relevant parts that have been added or modified are indicated by bold italic text, and some deleted parts are indicated by double square brackets (e.g., [ [ ] ), where the deleted text between the double square brackets indicates deleted or cancelled text. Some other changes can be of an editorial nature, and thus are not highlighted.

[0203] 6.1. First Embodiment

[0204] This embodiment is for item 1.

[0205] 6.1.1. Type of track

[0206] This specification specifies the following types of video tracks for carrying VVC bitstreams:

[0207] a) VVC track:

[0208]

[0209] b) VVC non-VCL track:

[0210]

[0211] c) VVC subpicture track:

[0212] A VVC subpicture track contains any of the following:

[0213] A sequence of one or more VVC subpictures.

[0214] A sequence of one or more complete slices forming a rectangular region.

[0215] The samples of a VVC subpicture track contain any of the following:

[0216] One or more complete subpictures in decoding order as specified in ISO / IEC 23090-3.

[0217] one or more complete slices that form a rectangular region and are consecutive in decoding order as specified in ISO / IEC 23090-3.

[0218] The VVC sub-pictures or slices included in any sample of a VVC sub-picture track are consecutive in decoding order.

[0219]

[0220] 6.2. Second embodiment

[0221] This embodiment is for item 4.b.

[0222] 6.2.1. Decoder [[configuration]] capability information sample group

[0223] 6.2.1.1. Definition

[0224] The sample group description entry of this sample group contains a DCI NAL unit. [[All samples that are mapped to the same decoder configuration information sample group description entry belong to the same VVC bitstream.]]

[0225] This sample group indicates whether the same DCI NAL unit is used for different sample entries in a VVC track [[i.e. whether the samples belonging to different sample entries belong to the same VVC bitstream]]. When the samples of two sample entries are mapped to the same decoder configuration information sample group description entry, the player can switch sample entries without re-initializing the decoder.

[0226] If any DCI NAL unit is present in any sample entry or in-band, this unit shall be identical to the DCI NAL unit included in the corresponding decoder configuration information sample group entry.

[0227] 6.2.1.2. Syntax

[0228]

[0229] 6.2.1.3. Semantics

[0230] dciNalUnitLength indicates the byte length of the DCI NAL unit.

[0231] dciNalUnit contains the DCI NAL unit as specified in ISO / IEC 23090-3.

[0232] 6.3. Third embodiment

[0233] This embodiment is for item 5.

[0234] 6.3.1. Definition of VVC decoder configuration record

[0235] This subsection specifies the decoder configuration information for ISO / IEC 23090-3 video content.

[0236] This record contains the size of the length field in each sample entry for indicating the length of the NAL units and parameter sets, DCI, OPI and SEI NAL units it contains, if stored in the sample entry. This record is externally scoped (its size is provided by the structure containing it).

[0237] This record contains a version field. This version of the specification defines version 1 of this record. Incompatible changes to the record will be indicated by a change in the version number. If the version number is not recognized, the reader shall not attempt to decode the record or the stream to which it applies.

[0238] Compatible extensions to this record shall extend it and shall not change the configuration version code. The reader should be prepared to ignore unrecognized data that extends beyond the data definitions they understand.

[0239] ...

[0241] There is an array set to carry initialization non-VCL NAL units. The NAL unit types are restricted to indicate only DCI, OPI, VPS, SPS, PPS, prefix APS and prefix SEI NAL units. NAL unit types reserved in ISO / IEC 23090-3 and this specification can acquire a definition in the future and the reader shall ignore the array with NAL unit types of reserved or unlicensed values.

[0242] NOTE 2: This "tolerant" behavior is designed to not produce an error, allowing the possibility of backward compatible extensions to these arrays in future specifications.

[0243] NOTE 3: NAL units carried in the sample entry are included in the access unit reconstructed from the first sample of the reference sample entry, immediately after the AUD and OPI NAL units, if any, or at the beginning of the access unit.

[0244] The array of suggested is in the order DCI, OPI, VPS, SPS, PPS, prefix APS, prefix SEI. ...

[0246] 6.3.2. Semantics of the VVC decoder configuration record ...

[0248] numArrays indicates the number of arrays of NAL units of the indicated type(s).

[0249] array_completeness is equal to 1 indicates that all NAL units of the given type are in the following array and none are in the stream; is equal to 0 indicates that additional NAL units of that indicated type can be in the stream; [[default and]] permitted values are constrained by the sample entry name.

[0250] NAL_unit_type indicates the type of the NAL units in the following array (which should be all of that type); it takes values as defined in ISO / IEC 23090-3; it is constrained to take one of the values indicating a DCI, OPI, VPS, SPS, PPS, prefix APS, or prefix SEI [[, or suffix SEI]] NAL unit.

[0251] numNalus indicates the number of NAL units of the indicated type included in the configuration record for the stream for which this configuration record applies. The SEI array shall contain only SEI messages of a "declarative" nature, i.e., those providing information about the stream as a whole. An example of such SEI can be user data SEI.

[0252] nalUnitLength indicates the byte length of the NAL unit.

[0253] nalUnit contains a DCI, OPI, VPS, SPS, PPS, APS, or declarative SEI NAL unit as specified in ISO / IEC 23090-3.

[0254] Figure 1 is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 1900. The system 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or can be in a compressed or encoded format. The input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0255] The system 1900 can include a codec component 1904 that can implement various coding or encoding methods described in this document. The codec component 1904 can reduce the average bitrate of video from the input 1902 to the output of the codec component 1904 to produce a coded representation of the video. The codec techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 can be stored, or transmitted via a communication connection as represented by the component 1906. The stored or communicated bitstream (or coded) representation of the video received at the input 1902 can be used by the component 1908 to generate pixel values or a displayable video to the display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Also, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that the coding tools or operations are used at an encoder, and corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.

[0256] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB), or High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices such as mobile telephones, laptop computers, smart phones, or other devices that are capable of performing digital data processing and / or video display.

[0257] Figure 2 is a block diagram of a video processing apparatus 3600. The apparatus 3600 can be used to implement one or more methods described herein. The apparatus 3600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 can be configured to implement one or more methods described in this document. The memory(ies) 604 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement, in hardware circuitry, some of the techniques described in this document. In some embodiments, the video processing hardware 3606 can be included at least in part in the processor 3602 (e.g., a graphics co-processor).

[0258] Figure 4 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.

[0259] As shown in Figure 4 FIG. 1, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 generates encoded video data, where the source device 110 can be referred to as a video encoding device. Destination device 120 can decode the encoded video data generated by source device 110, and the destination device 120 can be referred to as a video decoding device.

[0260] Source device 110 can include a video source 112, a video encoder 114, and an input / output (VO) interface 116.

[0261] Video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. VO interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to destination device 120 by VO interface 116 through network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by destination device 120.

[0262] Destination device 120 can include an VO interface 126, a video decoder 124, and a display device 122.

[0263] VO interface 126 can include a receiver and / or a modem. VO interface 126 can acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which is configured to interface with an external display device.

[0264] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.

[0265] Figure 5 is a block diagram illustrating an example of a video encoder 200, which can be Figure 4The video encoder 114 in the system 100 shown.

[0266] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0267] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0268] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0269] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for illustrative purposes, in Figure 5 The examples are shown separately.

[0270] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0271] The mode selection unit 203 can select one of the encoding / decoding modes (e.g., intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction modes (CIIP), where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the block's motion vector (e.g., sub-pixel or integer pixel precision).

[0272] To perform inter prediction on a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0273] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0274] In some examples, motion estimation unit 204 can perform uni-prediction on a current video block, and motion estimation unit 204 can search for a reference video block for the current video block in a reference picture of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0275] In other examples, motion estimation unit 204 can perform bi-prediction on a current video block, motion estimation unit 204 can search for a reference video block for the current video block in a reference picture of list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 204 can then generate a reference index indicating the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector indicating a spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and the motion vector for the current video block as the motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.

[0276] In some examples, motion estimation unit 204 can output a full set of motion information for a current video block for decoding processing at a decoder.

[0277] In some examples, motion estimation unit 204 can not output a full set of motion information for a current video block. Instead, motion estimation unit 204 can signal the motion information for the current video block with reference to motion information of another video block. For example, motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.

[0278] In one example, the motion estimation unit 204 can indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0279] In another example, the motion estimation unit 204 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0280] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0281] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0282] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.

[0283] In other examples, such as in skip mode, there can be no residual data for the current video block for the current video block, and the residual generation unit 207 can not perform the subtraction operation.

[0284] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0285] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0286] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to corresponding samples of one or more prediction video blocks generated from prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.

[0287] After reconstruction unit 212 reconstructs a video block, in-loop filtering operations can be performed to reduce video block artifacts in the video block.

[0288] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.

[0289] Figure 6 FIG. 3 is a block diagram illustrating an example of a video decoder 300 that can be Figure 4 the video decoder 114 in the system 100 shown.

[0290] Video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 6 example, video decoder 300 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0291] In Figure 6 example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to video encoder 200 Figure 5 ) above.

[0292] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data and, from the entropy decoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precisions, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes.

[0293] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0294] The motion compensation unit 302 can use an interpolation filter, such as that used by the video encoder 200 during the encoding of a video block, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0295] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0296] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0297] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 307 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates the decoded video for presentation on the display device.

[0298] The following is a list of preferred solutions for some embodiments.

[0299] The following solutions illustrate example embodiments of the techniques discussed in the previous sections (e.g., items 1 through 4).

[0300] 1. A method for processing visual media (e.g., Figure 3 The method described in the text (3000) includes: performing a conversion between visual media data and a file storing information corresponding to the visual media data according to format rules (3002); wherein the format rules specify a first condition for identifying the non-video codec layer (VCL) track of the file and / or a second condition for identifying the VCL track of the file.

[0301] 2. The method according to solution 1, wherein the first condition specifies that the non-VCL track contains only non-VCL network abstraction layer units and is identified in the VCL track by a specific track reference.

[0302] 3. The method according to solutions 1-2, wherein the first condition specifies that the non-VCL track contains an adaptation parameter set (APS) corresponding to the VCL track.

[0303] 4. The method according to any of solutions 1-3, wherein the second condition for the VCL track specifies that the VCL track is not allowed to include decoding capability information (DCI) or operation point information (OPI) network abstraction units.

[0304] 5. The method according to solution 1, wherein the first condition specifies that the non-VCL track includes one or more elementary streams containing non-VCL network abstraction layer units, and wherein the non-VCL network abstraction layer units are synchronized with the elementary streams in the VCL track.

[0305] 6. The method according to any of solutions 1-5, wherein the conversion comprises generating a bitstream representation of the visual media data, and storing the bitstream representation into a file according to the format rule.

[0306] 7. The method according to any of solutions 1-5, wherein the conversion comprises parsing a file according to the format rule to recover the visual media data.

[0307] 8. A video decoding apparatus comprising a processor configured to implement a method according to one or more of solutions 1 to 7.

[0308] 9. A video encoding apparatus comprising a processor configured to implement a method according to one or more of solutions 1 to 7.

[0309] 10. A computer program product having computer code stored therein, the code, which when executed by a processor, causes the processor to implement a method according to any of solutions 1 to 7.

[0310] 11. A computer readable medium, a bitstream representation on the computer readable medium being in accordance with a file format generated according to any of solutions 1 to 7.

[0311] 12. A method, apparatus or system described in the present document.

[0312] In the solutions described herein, an encoder can conform to a format rule by generating a coded representation according to the format rule. In the solutions described herein, a decoder can use a format rule to parse syntax elements in a coded representation knowing the presence and absence of syntax elements according to the format rule to generate a decoded video.

[0313] TECHNIQUE 1. A method of processing visual media data (e.g., method 8000 depicted in FIG. 8), comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule (8002), wherein the format rule specifies a condition whether a control information item is included in a non- video coded layer track of the visual media file, and wherein a presence of the non- video coded layer track in the visual media file is indicated by a specific track reference in a video coded layer track of the visual media file. Figure 8

[0314] TECHNIQUE 2. The method according to technique 1, wherein the condition specifies that the non-video coded layer track includes only non-video coded layer network abstraction layer units as the information item.

[0315] TECHNIQUE 3. The method according to any of techniques 1-2, wherein the condition specifies that the non-video coded layer track includes an adaptation parameter set as the information item, wherein the adaptation parameter set includes an adaptive loop filter parameter, a luma mapping parameter with chroma scaling, or a scaling list parameter, and wherein the condition specifies that the adaptation parameter set is stored in and transported through a track separate from another track that includes video coded layer network abstraction layer units.

[0316] TECHNIQUE 4. The method according to technique 3, wherein the condition allows the non- video coded layer track to additionally include non-video coded layer network abstraction layer units of other types.

[0317] TECHNIQUE 5. The method according to technique 3, wherein the condition does not allow the non-video coded layer track to additionally include non-video coded layer network abstraction layer units of other types.

[0318] TECHNIQUE 6. The method according to any of techniques 1-2, wherein the condition specifies that the non-video coded layer track includes a picture header network abstraction layer unit as the information item, and wherein the condition specifies that the picture header network abstraction layer unit is stored in and transported through a track separate from another track that includes video coded layer network abstraction layer units.

[0319] TECHNIQUE 7. The method according to technique 6, wherein the condition allows the non- video coded layer track to additionally include non-video coded layer network abstraction layer units of other types.

[0320] ​TECHNIQUE 8. The method according to TECHNIQUE 6, wherein the condition disallows a non- video coded layer track to additionally include other types of non- video coded layer network abstraction layer units.

[0321] TECHNIQUE 9. The method according to TECHNIQUE 6, wherein the condition allows a non- video coded layer track to additionally include an adaptation parameter set network abstraction layer unit.

[0322] TECHNIQUE 10. The method according to TECHNIQUE 6, wherein the condition disallows a non- video coded layer track to additionally include an adaptation parameter set network abstraction layer unit.

[0323] TECHNIQUE 11. The method according to any of TECHNIQUES 1-2, wherein the condition specifies that a picture header network abstraction layer unit of the video stream is an item of information that is stored in either a first set of sample of a track containing a video coded layer network abstraction layer unit or a second set of sample of a non- video coded layer track but not in both.

[0324] TECHNIQUE 12. The method according to any of TECHNIQUES 1-11, wherein the converting includes generating a visual media file according to a format rule and storing the bitstream to the visual media file.

[0325] TECHNIQUE 13. The method according to any of TECHNIQUES 1-11, wherein the converting includes generating a visual media file, and the method further includes storing the visual media file in a non-transitory computer readable recording medium.

[0326] TECHNIQUE 14. The method according to any of TECHNIQUES 1-11, wherein the converting includes parsing a visual media file according to a format rule to reconstruct the bitstream.

[0327] TECHNIQUE 15. The method according to any of TECHNIQUES 1 to 14, wherein the visual media file is processed by Versatile Video Coding (VVC), and wherein the non- video coded layer track or the video coded layer track is a VVC track.

[0328] TECHNIQUE 16. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement a method recited in any one or more of TECHNIQUES 1 to 15.

[0329] TECHNIQUE 17. A non-transitory computer readable storage medium storing instructions that cause a processor to implement a method recited in any one or more of TECHNIQUES 1 to 15.

[0330] TECHNIQUE 18. A video decoding apparatus comprising a processor configured to implement a method recited in one or more of TECHNIQUES 1 to 15.

[0331] Technology 16. A video encoding apparatus, including a processor configured to implement the method according to one or more of technologies 1 to 15.

[0332] Technology 17. A computer program product storing computer code, which, when executed by a processor, causes the processor to perform the method according to any one of technologies 1 to 15.

[0333] Technology 18. A computer-readable medium on which a visual media file conforms to a file format generated according to any one of technologies 1 to 15.

[0334] Technology 19. A non-transitory computer-readable recording medium for storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method is described in any one of technologies 1 to 15.

[0335] Technique 20. A method for generating a visual media file, comprising: generating a visual media file according to any one of Techniques 1 to 15, and storing the visual media file on a computer-readable program medium.

[0336] Implementation 1. A method for processing visual media data (e.g., Figure 9 The method described in the text (9000) includes: performing a conversion between a visual media file and a bitstream of visual media data according to format rules (9002), wherein the format rules specify the type of sample entry to determine whether the decoding capability information network abstraction layer unit is included in the sample entry of the video track in the visual media file or in the sample of both the video track and the video track in the visual media file.

[0337] Implementation Method 2. The method according to claim 1, wherein the format rule specifies that, in response to the sample entry type being vvc1, the decoding capability information network abstraction layer unit is included in the sample entry of the video track.

[0338] Implementation Method 3. According to the method of Implementation Method 1, wherein the format rule specifies that, in response to the sample entry type being vvi1, the decoding capability information network abstraction layer unit is included in the sample of the video track and the sample entry of the video track.

[0339] Implementation 4. According to the method of Implementation 1, wherein the format rules specify that the video basic stream in the visual media file includes a video codec layer network abstraction layer unit, wherein the format rules specify that the video basic stream in the visual media file is not allowed to include a parameter set or decoding capability information network abstraction unit, and wherein the format rules specify that the sample entries in the visual media file store the parameter set and the decoding capability information network abstraction unit.

[0340] Embodiment 5. The method according to embodiment 4, wherein the format rule specifies that a video elementary stream in the visual media file is not allowed to include parameter sets, decoding capability information network abstraction units, or operation point information network abstraction units, and wherein the format rule specifies that a sample entry in the visual media file stores parameter sets, decoding capability information network abstraction units, and operation point information network abstraction units.

[0341] Embodiment 6. A method of processing visual media data, comprising performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that, in response to a sample in the visual media file belonging to multiple general video coding bitstreams and in response to the sample being included in a same track, the sample is allowed to belong to a same decoder capability information sample group, and wherein the format rule specifies that all samples belonging to the same decoder capability information sample group share a same decoder capability information sample group description entry. In some embodiments, the format rule specifies that, in response to a sample in the visual media file belonging to multiple general video coding bitstreams and in response to the sample being included in a same track, the sample is allowed to belong to a same decoder capability information sample group, and wherein the format rule specifies that all samples belonging to the same decoder capability information sample group share a same decoder capability information sample group description entry.

[0342] Embodiment 7. The method according to embodiment 6, wherein the format rule specifies that, in response to all samples of a track belonging to a same bitstream or in response to all samples sharing a same decoding capability information regardless of a number of bitstreams, a decoding capability information network abstraction layer unit is indicated in a track level box in the visual media file. In some embodiments, the format rule specifies that, in response to all samples of a track belonging to a same bitstream or in response to all samples sharing a same decoding capability information regardless of a number of bitstreams, a decoding capability information network abstraction layer unit is indicated in a track level box in the visual media file.

[0343] Embodiment 8. The method according to embodiment 7, wherein the track level box is a track header box, a track level meta box, or another track level box.

[0344] Embodiment 9. A method of processing visual media data, comprising performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that an operation point information network abstraction layer unit is allowed to be included in the visual media file in a sample entry description as one of a plurality of non- video coded layer network abstraction layer unit arrays in a decoder configuration record. In some embodiments, the format rule specifies that an operation point information network abstraction layer unit is allowed to be included in the visual media file in a sample entry description as one of a plurality of non- video coded layer network abstraction layer unit arrays in a decoder configuration record.

[0345] Embodiment 10. A method of visual media processing, comprising performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that a type of a sample entry determines whether an operation point information network abstraction layer unit is included in: (1) a sample entry of a video track in the visual media file, or (2) a sample of the video track in the visual media file or the sample entry of the video track in the visual media file or both. In some embodiments, the format rule specifies that a second type of a second sample entry determines whether an operation point information network abstraction layer unit is included in: (1) the second sample entry of the video track in the visual media file, or (2) a sample of the video track in the visual media file or the second sample entry of the video track in the visual media file or both.

[0346] Embodiment 11. The method of embodiment 10, wherein the format rule specifies that the operation point information network abstraction layer unit is included in the sample entry of the video track in response to the type of the sample entry being vvc1. In some embodiments, the format rule specifies that the operation point information network abstraction layer unit is included in the second sample entry of the video track in response to the second type of the second sample entry being vvc1.

[0347] Embodiment 12. The method of embodiment 10, wherein the format rule specifies that the operation point information network abstraction layer unit is included in the sample of the video track or the sample entry of the video track or both in response to the type of the sample entry being vvi1. In some embodiments, the format rule specifies that the operation point information network abstraction layer unit is included in the sample of the video track or the second sample entry of the video track or both in response to the second type of the second sample entry being vvi1.

[0348] Embodiment 13. A method of processing visual media data, comprising performing conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that a non- video coded layer elementary stream in the visual media file is not allowed to include a video coded layer network abstraction layer unit, and wherein the format rule specifies that the non- video coded layer network abstraction layer unit is synchronized with an elementary stream carried in a video track. In some embodiments, the format rule specifies that a non- video coded layer elementary stream in the visual media file is not allowed to include a video coded layer network abstraction layer unit, and wherein the format rule specifies that the non- video coded layer network abstraction layer unit is synchronized with an elementary stream carried in a video track.

[0349] Embodiment 14. The method according to any of embodiments 1-13, wherein the conversion comprises generating the visual media file and storing the bitstream to the visual media file according to the format rule.

[0350] Embodiment 15. The method according to any of embodiments 1-13, wherein the conversion comprises generating the visual media file, and the method further comprises storing the visual media file in a non-transitory computer readable recording medium.

[0351] Embodiment 16. The method according to any of embodiments 1-13, wherein the conversion comprises parsing the visual media file to reconstruct the bitstream according to the format rule.

[0352] Embodiment 17. The method according to any of embodiments 1-16, wherein the visual media file is processed by Versatile Video Coding (VVC), and the video track is a VVC track.

[0353] Embodiment 18. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method according to one or more of embodiments 1-17.

[0354] Embodiment 19. A non-transitory computer readable storage medium storing instructions that cause a processor to implement the method according to any of embodiments 1-17.

[0355] Embodiment 20. A video decoding apparatus comprising a processor configured to implement the method according to one or more of embodiments 1-17.

[0356] Embodiment 21. A video encoding apparatus comprising a processor configured to implement the method according to one or more of embodiments 1-17.

[0357] Embodiment 22. A computer program product storing computer code which, when executed by a processor, causes the processor to implement the method according to any one of Embodiments 1 to 17.

[0358] Embodiment 23. A computer readable medium on which a visual media file is stored, the file format of which is generated according to any one of Embodiments 1 to 17.

[0359] Embodiment 24. A non-transitory computer readable recording medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method is recited in any one of Embodiments 1 to 17.

[0360] Embodiment 25. A method of visual media file generation, comprising generating a visual media file according to the method recited in any one of Embodiments 1 to 17, and storing the visual media file on a computer readable program medium.

[0361] In this document, the term "video processing" can refer to video encoding, video decoding, video compression or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block can for example correspond to collocated or scattered bits within the bitstream, as defined by the syntax. For example, a macroblock can be encoded according to a transform and coded error residual values and also using bits in headers and other fields in the bitstream. Furthermore, during the conversion, a decoder can parse the bitstream based on the determination, knowing that some fields can or can not be present, as described by the above solutions. Similarly, an encoder can determine to include or not include certain syntax fields, and generate the coded representation accordingly by including the syntax fields or excluding the syntax fields from the coded representation.

[0362] The disclosed and other solutions, examples, embodiments, modules and functional operations described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of devices including one or more of them, or a combination of one or more of them. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated for the purpose of encoding information for transmission to suitable receiver apparatus.

[0363] A computer program (which can also be referred to or referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be run on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.

[0364] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0365] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks, or a computer program product suitable for storing a computer program and data. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0366] Although the patent document contains many details, these should not be construed as limiting the subject matter or the scope of any patent that may issue on this patent document, but as a description of features that are specific to particular embodiments of the specific technology. Certain features described in the context of separate embodiments in this patent document can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.

[0367] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order nor that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0368] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method of processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a condition whether a control information item is included in a non- video coded layer track of the visual media file, wherein a presence of the non- video coded layer track in the visual media file is indicated by a specific track reference in a video coded layer track of the visual media file, wherein the condition specifies that a picture header network abstraction layer unit of a video stream is an information item stored in a first sample set of a track containing video coded layer network abstraction layer units or in a second sample set of a non- video coded layer track but not in both the first sample set and the second sample set.

2. The method according to claim 1, wherein the condition specifies that the non- video coded layer track only includes a non- video coded layer network abstraction layer unit as the information item.

3. The method according to claim 1 or 2, wherein the condition specifies that the non- video coded layer track includes an adaptation parameter set as the information item, wherein the adaptation parameter set includes an adaptive loop filter parameter, a luma mapping parameter with chroma scaling, or a scaling list parameter, and wherein the condition specifies that the adaptation parameter set is stored in and transmitted by a track separate from another track including video coded layer network abstraction layer units.

4. The method of claim 3, wherein, the condition allows the non- video coded layer track to additionally include a non- video coded layer network abstraction layer unit of other types.

5. The method of claim 3, wherein, the condition does not allow the non- video coded layer track to additionally include a non- video coded layer network abstraction layer unit of other types.

6. The method according to claim 1 or 2, wherein the condition specifies that the non- video coded layer track includes a picture header network abstraction layer unit as the information item, and wherein the condition specifies that the picture header network abstraction layer unit is stored in and transmitted by a track separate from another track including video coded layer network abstraction layer units.

7. The method of claim 6, wherein, the condition allows the non- video coded layer track to additionally include a non- video coded layer network abstraction layer unit of other types.

8. The method of claim 6, wherein, the condition does not allow the non- video coded layer track to additionally include a non- video coded layer network abstraction layer unit of other types.

9. The method of claim 6, wherein, the condition allows the non- video coded layer track to additionally include an adaptation parameter set network abstraction layer unit.

10. The method of claim 6, wherein, the condition does not allow the non- video coded layer track to additionally include an adaptation parameter set network abstraction layer unit.

11. The method of any one of claims 1-10, wherein, the conversion includes generating the visual media file and storing the bitstream to the visual media file according to the format rule.

12. The method of any one of claims 1-10, wherein, the conversion includes parsing the visual media file to reconstruct the bitstream according to the format rule.

13. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to implement the method according to any one of claims 1-12.

14. A non-transitory computer-readable storage medium storing instructions that cause a processor to implement the method of any one of claims 1-12.