Decoder configuration record in coded video

By improving the syntax and semantics of the decoder configuration record and sub-picture entity group in the VVC video file format, the inaccuracy and compatibility issues of the decoder configuration information are resolved, more accurate video stream decoding and merging are achieved, and the processing efficiency of the decoder is improved.

CN114205601BActive Publication Date: 2025-10-24FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111090725.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-17
Filing Date
2021-09-17
Publication Date
2025-10-24
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

In the existing VVC video file format design, the decoder configuration information and sub-picture entity group signaling notification have semantic ambiguity and compatibility issues, leading to decoder configuration errors and information loss.

Method used

By modifying the syntax and semantics of the decoder configuration record and sub-picture entity group of the VVC video file format, the accuracy of information such as profile, layer, level and chroma format is ensured, fewer bits are used for encoding and decoding, the optionality and compatibility of fields are increased, and the conditions for signaling notification and the definition of fields are improved.

Benefits of technology

Improves the accuracy and compatibility of decoder configuration information for the VVC video file format, avoids erroneous signaling notifications, ensures correct decoding and merging of video streams, and simplifies the decoder processing flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114205601B_ABST
    Figure CN114205601B_ABST
Patent Text Reader

Abstract

This relates to decoder configuration recording in coded video. Systems, methods, and apparatuses for encoding or decoding a file format storing one or more images are described. An example method includes performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a characteristic of a syntax element in the visual media file, wherein the syntax element has a value indicating a number of bytes for indicating constraint information associated with the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is filed under applicable patent law and / or under the rules of the Paris Convention to timely claim priority to and the benefit of U.S. Provisional Patent Application No. 63 / 079,892, filed on September 17, 2020. The entire disclosure of the above application is incorporated by reference as a part of the disclosure of this application for all purposes of law. Technical Field

[0003] This patent document relates to the generation, storage, and consumption of digital audio-visual media information in a file format. Background Art

[0004] Digital video accounts for the largest usage of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process codec representations of videos or images according to file formats.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a visual media file and a bitstream of visual media data according to format rules, wherein the bitstream includes one or more output layer sets and one or more parameter sets, the one or more parameter sets including one or more profile layer-level syntax structures, wherein at least one of the profile layer-level syntax structures includes a general constraint information syntax structure, wherein the format rules specify that a syntax element be included in a configuration record in the visual media file, and wherein the syntax element indicates a profile, layer, or level to which an output layer set identified by an output layer set index indicated in the configuration record conforms.

[0007] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies characteristics of a syntax element in the visual media file, wherein the syntax element has a value indicating a number of bytes used to indicate constraint information associated with the bitstream.

[0008] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a property of a syntax element in the visual media file, and wherein the format rule specifies that a syntax element having a value indicative of a level identification is coded using octets in either or both of a subpicture common group box or a subpicture multiple group box.

[0009] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a property of a syntax element in the visual media file, and wherein the format rule specifies that a syntax element having a value indicative of a level identification is coded using octets in either or both of a subpicture common group box or a subpicture multiple group box.

[0010] In yet another example aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement a method described above.

[0011] In yet another example aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement a method described above.

[0012] In yet another example aspect, a computer readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.

[0013] In yet another example aspect, a computer readable medium having a bitstream stored thereon is disclosed. The bitstream is generated or processed using the methods presented in this document.

[0014] These and other features are presented throughout this document. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 is a block diagram of an example video processing system.

[0016] Figure 2 is a block diagram of a video processing apparatus.

[0017] Figure 3 is a flowchart of an example method of video processing.

[0018] Figure 4 is a block diagram illustrating a video coding system according to some embodiments of the present disclosure.

[0019] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0020] Figure 6is a block diagram illustrating a decoder according to some embodiments of the disclosure.

[0021] Figure 7 An example of an encoder block diagram is shown.

[0022] Figures 8 to 10 is a flowchart of, for example, a video processing method. DETAILED DESCRIPTION

[0023] For ease of understanding, section headings have been used in this document, and the teachings and embodiments disclosed in each section are not intended to be limited only to that section. Furthermore, the use of H.266 terminology in some descriptions is only for ease of understanding and is not intended to limit the scope of the disclosed technology. Likewise, the technology described herein is also applicable to other video codec protocols and designs. In this document, editorial changes to text are shown by opening and closing double brackets (e.g., [[]]) indicating that the text between the double brackets is cancelled, and by bold italic text indicating added text, with respect to the current draft of the VVC specification or the ISOBMFF file format specification.

[0024] 1. Brief discussion

[0025] This document is related to video file formats. In particular, it is concerned with the signaling of decoder configuration information and subpicture entity groups in media files carrying Versatile Video Coding (VVC) video bitstreams based on the ISO Base Media File Format (ISOBMFF). These ideas can be applied individually or in various combinations for video bitstreams coded by any codec (e.g., the VVC standard), and for any video file format (e.g., the VVC video file format that is being developed).

[0026] 2. Abbreviations

[0027] ACT Adaptive Color Transform

[0028] ALF Adaptive Loop Filter

[0029] AMVR Adaptive Motion Vector Resolution

[0030] APS Adaptation Parameter Set

[0031] AU Access Unit

[0032] AUD Access Unit Delimiter

[0033] AVC Advanced Video Coding (Rec. ITU-T H.264 | ISO / IEC 14496-10)

[0034] B Bi-prediction

[0035] BCW bi-directional prediction with CU-level weights

[0036] BDOF bi-directional optical flow

[0037] BDPCM block-based delta pulse code modulation

[0038] BP buffering period

[0039] CABAC context-based adaptive binary arithmetic coding

[0040] CB coded block

[0041] CBR constant bit rate

[0042] CCALF cross-component adaptive loop filter CPB coded picture buffer CRA clean random access CRC cyclic redundancy check CTB coded tree block CTU coded tree unit CU coding unit CVS coded video sequence DPB decoded picture buffer DCI decoding capability information DRAP dependent random access point DU decoding unit DUI decoding unit information EG exponential Golomb k-th order exponential Golomb EOB end of bitstream EOS end of sequence FD filler data FIFO first-in-first-out FL fixed length GBR green, blue, red GCI general constraint information CGR gradual decoding refresh GPM geometric partition mode HEVC high efficiency video coding (Rec. ITU-T H.265 | ISO / IEC 23008-2) HRD hypothetical reference decoder HSS hypothetical stream scheduler I intra frame IBC intra block copy IDR instantaneous decoding refresh ILRP inter-layer reference picture IRAP intra random access point

[0043] LFNST low-frequency non-separable transform

[0044] LPS least probable symbol

[0045] LSB least significant bit

[0046] LTRP long-term reference picture

[0047] MCS luma mapping with chroma scaling

[0048] MIP matrix-based intra prediction

[0049] MPS most probable symbol

[0050] MSB most significant bit

[0051] MTS multiple transform selection

[0052] MVP motion vector prediction

[0053] NAL network abstraction layer

[0054] OLS output layer set

[0055] OP operation point

[0056] OPI operation point information

[0057] P predictive

[0058] PH picture header

[0059] POC picture order count

[0060] PPS picture parameter set

[0061] PROF prediction refinement with optical flow

[0062] PT picture timing

[0063] PU picture unit

[0064] QP quantization parameter

[0065] RADL random access decodable leading (picture)

[0066] RASL random access skipped leading (picture)

[0067] RBSP raw byte sequence payload

[0068] RGB red, green, blue

[0069] RPL reference picture list

[0070] SAO sample adaptive offset

[0071] SAR sample aspect ratio

[0072] SEI supplemental enhancement information

[0073] SH slice header

[0074] SLI subpicture level information

[0075] SODB string of data bits

[0076] SPS sequence parameter set

[0077] STRP short-term reference picture

[0078] STSA stepping temporal sublayer access

[0079] TR truncated rice code

[0080] VBR variable bit rate

[0081] VCL video coding layer

[0082] VPS video parameter set

[0083] VSEI versatile supplemental enhancement information (Rec. ITU-T H.274 ISO / IEC 23002-7)

[0084] VUI video usability information

[0085] VVC versatile video coding (Rec. ITU-T H.266 ISO / IEC 23090-3)

[0086] 3. Introduction to video coding

[0087] 3.1. Video coding standards

[0088] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) standards and the H.265 / HEVC standard. Since H.262, the video coding standards are based on the hybrid video coding structure wherein temporal prediction plus transform coding is used. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by the JVET and brought into the reference software named Joint Exploration Model (JEM). The JVET was later renamed to the Joint Video Team (JVT) when the Versatile Video Coding (VVC) project officially started. VVC is the new coding standard targeting at 50% bitrate reduction compared to HEVC, which has been finalized by the JVET at its 19th meeting, July 1, 2020.

[0089] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) are designed for the broadest range of applications, including traditional uses such as television broadcast, video conferencing or playback from storage media, as well as newer and more advanced usage scenarios such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple coded video bitstreams, multi-view video, scalable layered coding and viewport-adaptive 360° immersive media.

[0090] 3.2. File format standards

[0091] Media streaming applications are typically based on IP, TCP and HTTP transport methods and often rely on file formats such as the ISO Base Media File Format (ISOBMFF). One such streaming system is dynamic adaptive streaming over HTTP (DASH). For video formats using with ISOBMFF and DASH, video format specific file format specifications such as AVC file format and HEVC file format will be needed to encapsulate video content in ISOBMFF tracks and in DASH representations and segments. Important information about the video bitstream, e.g. profile, tier and level and many other information, will need to be exposed as file format level metadata and / or DASH media presentation description (MPD) for content selection purposes, e.g. selection of appropriate media segments both for initialization at the beginning of a streaming session and for stream adaptation during a streaming session.

[0092] Similarly, for image formats using with ISOBMFF, image format specific file format specifications such as AVC image file format and HEVC image file format are needed.

[0093] The VVC video file format, a file format for storing VVC video content based on ISOBMFF, is currently being developed by MPEG.

[0094] The VVC image file format, a file format for storing image content coded using VVC based on ISOBMFF, is currently being developed by MPEG.

[0095] 3.3. Some details of the VVC video file format

[0096] 3.3.1. Decoder configuration information

[0097] 3.3.1.1.VVC decoder configuration record

[0098] 3.3.1.1.1. Definition

[0099] This subclause specifies decoder configuration information for ISO / IEC 23090-3 video content.

[0100] This record contains the size of the length field used in each sample to indicate the length of the NAL unit it contains and the parameter set (if stored in the sample entry). This record is externally formulated (its size is provided by the structure that contains it).

[0101] This record contains a version field. This version of the specification defines version 1 of this record. Incompatible changes to the record will be indicated by changes in the version number. If the version number is not recognized, readers must not attempt to decode this record or the stream to which it applies.

[0102] Compatible extensions to this record will extend it and will not change the configuration version code. Readers should be prepared to ignore unrecognized data that is beyond their understanding of the data definition.

[0103] When a track contains a VVC bitstream itself or by parsing a 'subp' track reference, a VvcPtlRecord shall be present in the decoder configuration record. If the ptl_present_flag in the decoder configuration record of a track is equal to zero, the track shall have an 'oref' track reference.

[0104] The values ​​of the VvcPTLRecord, chroma_format_idc, and bit_depth_minus8 syntax elements shall be valid for all parameter sets that are activated when decoding the stream described by this record (referred to as "all parameter sets" in the following sentences of this paragraph). Specifically, the following restrictions apply:

[0105] Profile indication general_profile_idc shall indicate the profile to which the stream associated with this configuration record conforms.

[0106] Note 1: If the SPS is marked with different profiles, the stream may need to be checked to determine which profile (if any) the entire stream complies with. If the entire stream is not checked, or the check shows that there is no profile that the entire stream complies with, then the entire stream should be split into two or more sub-streams with separate configuration records in which these rules can be met.

[0107] The layer indication general_tier_flag shall indicate a layer equal to or greater than the highest layer indicated in all parameter sets.

[0108] Each bit in general_constraint_info can only be set if all parameter sets have that bit set.

[0109] The general_level_idc shall indicate a capability level that is equal to or greater than the highest level indicated for all parameter sets.

[0110] The following constraint applies to chroma_format_idc:

[0111] - If the value of sps_chroma_format_idc as defined in ISO / IEC 23090-3 is the same in all SPS referred to by the NAL units of the track, chroma_format_idc shall be equal to sps_chroma_format_idc.

[0112] - Otherwise, if ptl_present_flag is equal to 1, chroma_format_idc shall be equal to vps_ols_dpb_chroma_format[ output_layer_set_idx ], as defined in ISO / IEC 23090-3.

[0113] - Otherwise, chroma_format_idc shall not be present.

[0114] The following constraint applies to bit_depth_minus8:

[0115] - If the value of sps_bitdepth_minus8 as defined in ISO / IEC 23090-3 is the same in all SPS referred to by the NAL units of the track, bit_depth_minus8 shall be equal to sps_bitdepth_minus8.

[0116] - Otherwise, if ptl_present_flag is equal to 1, bit_depth_minus8 shall be equal to vps_ols_dpb_bitdepth_minus8[ output_layer_set_idx ], as defined in ISO / IEC 23090-3.

[0117] - Otherwise, bit_depth_minus8 shall not be present.

[0118] An explicit indication about chroma format and bit depth is provided in the VVC decoder configuration record, along with other important format information used by VVC video elementary streams. Two different VVC sample entries are also needed if two sequences differ in their color space indication in the VUI information.

[0119] There are array groups to carry initialization NAL units. NAL unit types are restricted to indicate DCI, VPS, SPS, PPS, prefix APS, and prefix SEI NAL units. NAL unit types reserved in ISO / IEC 23090-3 and this specification can be defined in the future, and the reader should ignore arrays with reserved or disallowed NAL unit type values.

[0120] NOTE 2: This ‘tolerant’ design is to not raise errors, allowing backward-compatible extensions to these arrays in future specifications.

[0121] NOTE 3: NAL units carried in the sample entry are included immediately after the AUD and OPINAL units (if any) in the access unit that the sample entry refers to or at the beginning of the access unit.

[0122] The order of the suggested arrays is DCI, VPS, SPS, PPS, prefix APS, prefix SEI.

[0123] 3.3.1.1.2. Syntax

[0124]

[0125]

[0126]

[0127] 3.3.1.1.3. Semantics

[0128] general_profile_idc, general_tier_flag, general_sub_profile_idc, general_constraint_info, general_level_idc, ptl_frame_only_constraint_flag, ptl_multilayer_enabled_flag, sublayer_level_idc[i] contain matching values of the bits in the fields general_profile_idc, general_tier_flag, general_sub_profile_idc, general_constraint_info(), general_level_idc, ptl_multilayer_enabled_flag, ptl_frame_only_constraint_flag, sublayer_level_present, and sublayer_level_idc[i] as defined for the stream to which this configuration record applies in ISO / IEC 23090-3.

[0129] avgFrameRate gives the average frame rate of the stream to which this configuration record applies in frames per (256 seconds). The value 0 indicates that the average frame rate is not specified.

[0130] constantFrameRate equal to 1 indicates that the stream to which this configuration record applies has a constant frame rate. The value 2 indicates that the representations of each temporal layer in the stream are constant frame rate. The value 0 indicates that the stream can or can not be constant frame rate.

[0131] numTemporalLayers greater than 1 indicates that the track to which this configuration record applies is scalable in the temporal domain and contains a number of temporal layers (also referred to as temporal sub-layers or sub-layers in ISO / IEC 23090-3) equal to numTemporalLayers. The value 1 indicates that the track to which this configuration record applies is not scalable in the temporal domain. The value 0 indicates that it is unknown whether the track to which this configuration record applies is scalable in the temporal domain.

[0132] lengthSizeMinusOne plus 1 indicates the length (in bytes) of the NALUnitLength field in the VVC video stream samples in the stream to which this configuration record applies. For example, a size of one byte is indicated with the value 0. The value of this field shall be one of 0, 1, or 3, which correspond to lengths coded with 1, 2, or 4 bytes, respectively.

[0133] ptl_present_flag equal to 1 specifies that the track contains a VVC bitstream corresponding to a particular output layer set. ptl_present_flag equal to 0 specifies that the track can not contain a VVC bitstream corresponding to a particular output layer set, but can contain one or more individual layers that do not form an output layer set or individual sub-layers that do not include a sub-layer with Temporalld equal to 0.

[0134] num_sub_profiles defines the number of sub-profiles indicated in the decoder configuration record.

[0135] track_ptl specifies the profile, tier, and level of the output layer set represented by the VVC bitstream contained in the track.

[0136] output_layer_set_idx specifies the output layer set index of the output layer set represented by the VVC bitstream contained in the track. The value of output_layer_set_idx can be used as the value of the TargetOlsIdx variable provided by an external component to a VVC decoder, as specified in ISO / IEC 23090-3 for decoding the bitstream contained in the track.

[0137] chroma_format_present_flag equal to 0 specifies that chroma_format_idc is not present. chroma_format_present_flag equal to 1 specifies that chroma_format_idc is present.

[0138] chroma_format_idc indicates the chroma format applied to this track. The following constraints apply to chroma_format_idc:

[0139] - If the value of sps_chroma_format_idc as defined in ISO / IEC 23090-3 is the same in all SPS referred to by the NAL units of the track, chroma_format_idc shall be equal to sps_chroma_format_idc.

[0140] - Otherwise, if ptl_present_flag is equal to 1, chroma_format_idc shall be equal to vps_ols_dpb_chroma_format[output_layer_set_idx] as defined in ISO / IEC 23090-3.

[0141] - Otherwise, chroma_format_idc shall not be present.

[0142] bit_depth_present_flag equal to 0 specifies that bit_depth_minus8 is not present. bit_depth_present_flag equal to 1 specifies that bit_depth_minus8 is present.

[0143] bit_depth_minus8 indicates the bit depth that should be used for this track. The following constraints apply to bit_depth_minus8:

[0144] - If the value of sps_bitdepth_minus8 defined in ISO / IEC 23090-3 is the same in all SPSs referenced by NAL units of a track, bit_depth_minus8 shall be equal to sps_bitdepth_minus8.

[0145] Otherwise, if ptl_present_flag is equal to 1, bit_depth_minus8 shall be equal to vps_ols_dpb_bitdepth_minus8[output_layer_set_idx], as defined in ISO / IEC 23090-3.

[0146] - Otherwise, bit_depth_minus8 will not exist.

[0147] numArrays indicates the number of arrays of NAL units of the indicated type.

[0148] When array_completeness is equal to 1, it indicates that all NAL units of the given type are in the array below and none are in the stream; when array_completeness is equal to 0, it indicates that additional NAL units of the indicated type may be in the stream; the default and allowed values ​​are subject to the example entry name.

[0149] NAL_unit_type indicates the type of NAL unit in the following array (the array should all be of this type); it uses the values ​​defined in ISO / IEC 23090-2; it is restricted to using one of the values ​​indicating a DCI, VPS, SPS, PPS, APS, prefix SEI, or suffix SEI NAL unit.

[0150] numNalus indicates the number of NAL units of the indicated type included in the configuration record of the stream to which this configuration record applies. The SEI array will only contain 'declarative' SEI messages, i.e., those SEI messages that provide information about the entire stream. An example of such an SEI is the user data SEI.

[0151] nalUnitLength indicates the length of the NAL unit in bytes.

[0152] The nalUnit contains a DCI, VPS, SPS, PPS, APS, or declarative SEI NAL unit as specified in ISO / IEC 23090-3.

[0153] 3.3.2. Subpicture Entity Groups

[0154] 3.3.2.1. Overview

[0155] Subpicture Entity Groups are defined to provide level information indicating the level of conformance of a bitstream merged from multiple VVC subpicture tracks.

[0156] NOTE: The VVC base track provides another mechanism for merging VVC subpicture tracks.

[0157] The implicit reconstruction process requires modification of parameter sets. Subpicture Entity Groups give guidance to simplify the generation of parameter sets for reconstructing a bitstream.

[0158] When the coded subpictures within the group to be jointly decoded are interchangeable, i.e., the player selects several active tracks from one group of subpictures with the same level contribution in terms of samples, the SubpicCommonGroupBox indicates the combination rule when jointly decoded and the level_idc of the resulting combination.

[0159] When there are coded subpictures with different properties (e.g., different resolutions) selected to be jointly decoded, the SubpicMultipleGroupsBox indicates the combination rule when jointly decoded and the level_idc of the resulting combination.

[0160] All entity_id values included in a Subpicture Entity Group shall identify VVC subpicture tracks. When present, the SubpicCommonGroupBox and SubpicMultipleGroupsBox shall be included in the GroupsListBox in the movie-level MetaBox and shall not be included in the file-level or track-level MetaBox.

[0161] 3.3.2.2. Syntax of Subpicture Common Group Box

[0162]

[0163] 3.3.2.3. Semantics of Subpicture Common Group Box

[0164] level_idc specifies the level that any selection of num_active_tracks entities in the entity group conforms to.

[0165] num_active_tracks specifies the number of tracks for which the level_idc value is provided.

[0166] 3.3.2.4. Syntax of subpicture multi-group box

[0167]

[0168] 3.3.2.5. Semantics

[0169] Level_idc specifies the level that the combination of any num_active_tracks[i] tracks in the sub-group with ID equal to i conforms to for all i values in the range of 0 to num_subgroup_ids-1, inclusive.

[0170] num_subgroup_ids specifies the number of separate sub-groups, each identified by the same value of track_subgroup_id[i]. Different sub-groups are identified by different values of track_subgroup_id[i].

[0171] track_subgroup_id[i] specifies the sub-group ID of the i-th track in this entity group. The sub-group ID value is in the range of 0 to num_subgroup_ids-1, inclusive.

[0172] num_active_tracks[i] specifies the number of tracks in the sub-group with ID equal to i as recorded in level_idc.

[0173] 4. Examples of technical problems solved by the disclosed technical solutions

[0174] The latest design of VVC video file format on signaling of decoder configuration information and subpicture entity group information has the following problems:

[0175] 1) It is specified that the profile indication general_profile_idc shall indicate the profile to which the stream associated with this configuration record conforms. However, the stream can correspond to multiple output layer sets, so this semantics can allow an incorrect value of general_profile_idc to be signaled in the configuration record.

[0176] 2) The specification states that the general_tier_flag shall indicate the tier equal to or greater than the highest tier indicated in all parameter sets. However, there can be profile_tier_level() structures signaled in parameter sets and applied to OLSs not within the scope of the current configuration record, so this semantics can allow an erroneous value of this field to be signaled in the configuration record. In addition, there can be profile_tier_level() structures signaled in parameter sets and not referenced, and the VPS can include PTL structures applied to OLSs not within the scope of the current configuration record.

[0177] 3) The specification states that each bit in general_constraint_info can only be set if all parameter sets have that bit set. However, there can be profile_tier_level() structures signaled in parameter sets and applied to OLSs not within the scope of the current configuration record, so this semantics can allow an erroneous value of this field to be signaled in the configuration record.

[0178] 4) The specification states that the general_level_idc shall indicate the level equal to or greater than the highest level indicated for the highest tier in all parameter sets. However, there can be profile_tier_level() structures signaled in parameter sets and applied to OLSs not within the scope of the current configuration record. In addition, the highest level can have a highest level lower than the highest level of the lowest tier, and the level determines the maximum picture width, height, etc., which are essential to determine the required decoding capability. So this semantics can allow an erroneous value of this field to be signaled in the configuration record.

[0179] 5) In the syntax and semantics of the VvcPTLRecord() syntax structure, the issues related to the fields num_bytes_constraint_info and general_constraint_info are as follows:

[0180] a. The field num_bytes_constraint_info is coded using 8 bits. However, the maximum number of bits in the general_constraint_info() syntax structure defined in the VVC specification is 336 bits, i.e., 42 bytes, so 6 bits are sufficient.

[0181] b. In addition, the semantics of the field num_bytes_constraint_info is missing.

[0182] c. The condition of the field general_constraint_info is "if (num_bytes_constraint_info > 0)". However, in the profile_tier_level() syntax structure defined in the VVC specification, the general_constraint_info() syntax structure is present whenever the profile, tier and level are present, while even if the first syntax element gci_present_flag in the general_constraint_info() syntax structure is equal to 0, the length of the general_constraint_info() syntax structure is still one byte, not zero byte. Therefore, the condition should be changed to "if (num_bytes_constraint_info > 1)", i.e. when the gci_present_flag of the general_constraint_info() syntax structure is equal to 0, the field general_constraint_info is not signaled.

[0183] d. The field general_constraint_info is coded using (8 * num_bytes_constraint_info - 2) bits. However, the length of general_constraint_info(), i.e. the general_constraint_info() syntax structure defined in the VVC specification, is an integer number of bytes.

[0184] 6) In VvcDecoderConfigurationRecord, the field output_layer_set_idx is always signaled when ptl_present_flag is equal to 1, i.e. when the track_ptl field is signaled. However, if the VVC bitstream carried by the VVC track (after parsing the referenced VVC track or VVC subpicture track, if any) is a single-layer bitstream, the value of the OLS index is usually not needed to know, even if it is useful to know, it can be easily derived that it is the OLS index of the OLS containing only that layer.

[0185] 7) The NAL_unit_type field in VvcDecoderConfigurationRecord uses 6 bits. However, 5 bits are enough.

[0186] 8) The semantics of ptl_present_flag is specified as follows: ptl_present_flag equal to 1 specifies that the track contains a VVC bitstream corresponding to a particular output layer set. ptl_present_flag equal to 0 specifies that the track can not contain a VVC bitstream corresponding to a particular output layer set, but can contain one or more individual layers that do not form an output layer set or individual sub-layers that do not include a sub-layer with Temporalld equal to 0.

[0187] However, the case where a track contains VVC bitstreams corresponding to multiple output layer sets is not covered.

[0188] 9) The levl_idc field in SubpicCommonGroupBox and SubpicMultipleGroupsBox is coded using 32 bits. However, 8 bits are sufficient.

[0189] 10) The num_active_tracks field in SubpicCommonGroupBox and the num_subgroup_ids field and num_active_tracks[i] fields in SubpicMultipleGroupsBox are all coded using 32 bits. However, 16 bits are sufficient for all of them.

[0190] 5. List of solutions

[0191] To solve the above problems and other problems, the methods described below are disclosed. These items should be considered as examples to explain the general concept and should not be interpreted in a narrow way. Furthermore, these items can be applied individually or combined in any way.

[0192] 1) To solve problem 1, it is specified that the general_profile_idc shall indicate the profile to which the output layer set identified by output_layer_set_idx in this configuration record conforms.

[0193] 2) To solve problem 2, it is specified that the general_tier_flag shall indicate one layer that is equal to or greater than the highest layer indicated in all profile_tier_level() syntax structures (in all parameter sets) of the profile_tier_level() syntax structures (in all parameter sets) indicated by the output_layer_set_idx in this configuration record.

[0194] a. Alternatively, the tier indication specifies that the general_tier_flag shall indicate the highest tier indicated in all profile_tier_level() syntax structures (in all parameter sets) that the output layer set identified by output_layer_set_idx in this configuration record conforms to.

[0195] b. Alternatively, the tier indication specifies that the general_tier_flag shall indicate the highest tier that the stream associated with this configuration record conforms to.

[0196] c. Alternatively, the tier indication specifies that the general_tier_flag shall indicate one tier that the stream associated with this configuration record conforms to.

[0197] 3) To address issue 3, it is proposed that each bit in general_constraint_info can only be set if the bit is set in all general_constraints_info() syntax structures in all profile_tier_level() syntax structures (in all parameter sets) that the output layer set identified by output_layer_set_idx in this configuration record conforms to.

[0198] 4) To address issue 4, it is proposed that the level indication specifies that the general_level_idc shall indicate the capability level that is equal to or greater than the highest level in all profile_tier_level() syntax structures (in all parameter sets) that the output layer set identified by output_layer_set_idx in this configuration record conforms to.

[0199] 5) To address issue 5, one or more of the following is proposed:

[0200] a. The field num_bytes_constraint_info is coded using 6 bits.

[0201] b. The field num_bytes_constraint_info is coded immediately after the ptl_multilayer_enabled_flag field.

[0202] c. The semantics of the field num bytes constraint info are specified as follows: num bytes constraint info specifies the number of bytes in the general constraint info() syntax structure defined in ISO / IEC 23090-3. The value equal to 1 indicates that gci present flag in the general constraint info() syntax structure is equal to 0 and no general_constraint_info field is signaled in this VvcPTLRecord.

[0203] d. The signaling condition of the field general_constraint_info is changed from "if (num bytes constraint info > 0)" to "if (num bytes constraint info > 1)".

[0204] e. The general_constraint_info field is coded using 8 * num bytes constraint info bits instead of (8 * num bytes constraint info - 2) bits.

[0205] 6) To solve problem 6, the signaling of the field output layer set idx in VvcDecoderConfigurationRecord is optional even when ptl present flag is equal to 1, e.g., conditioned on "if (track_ptl.ptl_multilayer_enabled_flag)", which indicates that the VVC bitstream contains only one layer carried in the VVC track (after resolving the referenced VVC track or VVC subpicture track (if any)).

[0206] a. Alternatively, when ptl present flag is equal to 1 and output layer set idx is not present, its value is inferred to be equal to the OLS index of the OLS containing only the layer carried in the VVC track (after resolving the referenced VVC track or VVC subpicture track (if any)).

[0207] 7) To solve problem 7, the NAL unit type field in VvcDecoderConfigurationRecord uses 5 bits instead of 6 bits.

[0208] 8) To solve problem 8, the semantics of ptl_present_flag is specified as follows: ptl_present_flag equal to 1 specifies that the track contains VVC bitstream corresponding to a particular output layer set. ptl_present_flag equal to 0 specifies that the track can not contain VVC bitstream corresponding to a particular output layer set, but can contain VVC bitstream corresponding to multiple output layer sets or contain one or more individual layers that do not form an output layer set or contain individual sub-layers that do not include sub-layers with Temporalld equal to 0.

[0209] 9) To solve problem 9, the coding of the level_idc field in one or both of SubpicCommonGroupBox and SubpicMultipleGroupsBox is changed to use 8 bits.

[0210] a. Alternatively, the subsequent 24 bits after the level_idc field can also be specified as reserved bits.

[0211] b. Alternatively, the subsequent 8 bits after the level_idc field can also be specified as reserved bits.

[0212] c. Alternatively, the subsequent zero bits after the level_idc field can also be specified as reserved bits.

[0213] 10) To solve problem 10, the coding of one or more of the num_active_tracks field in SubpicCommonGroupBox and the num_subgroup_ids field and the num_active_tracks[i] fields in SubpicMultipleGroupsBox is changed to use 16 bits.

[0214] a. Alternatively, the subsequent 16 bits after one or more of the above fields can also be specified as reserved bits.

[0215] b. Alternatively, the subsequent zero bits after one or more of the above fields can also be specified as reserved bits.

[0216] 6. Embodiments

[0217] Below are some example embodiments for some aspects of the present invention summarized above in section 5, which can be applied to the standard specification of VVC video file format. The modified text is based on the latest draft specification. Most of the added or modified relevant parts are indicated in bold italic text, some deleted parts are indicated in open and close brackets (e.g. [[]]), the text between the two brackets indicates the deleted or cancelled text. There can be other editorial nature changes, thus not highlighted.

[0218] 6.1. First embodiment

[0219] This embodiment is for items 1, 2, 3, 4, 5a, 5b, 5c, 5d, 5e, 6, 6a, 7 and 8.

[0220] 6.1.1. Decoder configuration information

[0221] 6.1.1.1. VVC decoder configuration record

[0222] 6.1.1.1.1. Definition

[0223] This subclause specifies the decoder configuration information for ISO / IEC 23090-3 video content.

[0224] This record contains the size of the length field used in each sample to indicate the length of the NAL units it contains and parameter sets, (if stored in the sample entry). This record is externally specified (its size is provided by the structure containing it).

[0225] This record contains the version field. Version 1 of this record is defined by this specification. Incompatible changes to the record will be indicated by a change in the version number. If the version number cannot be recognized, the reader shall not attempt to decode the stream to which this record applies or it applies to.

[0226] Compatible extensions to this record shall extend it and shall not change the configuration version code. The reader shall be prepared to ignore unrecognized data that extends beyond the data definition they understand.

[0227] When the track itself contains a VVC bitstream or contains a VVC bitstream by resolving the ‘subp’ track reference, the VvcPtlRecord shall be present in the decoder configuration record, If the ptl_present_flag in the decoder configuration record of the track is equal to zero, the track shall have an ‘oref’ track reference.

[0228] The values of the VvcPTLRecord, chroma_format_idc, and bit_depth_minus8 syntax elements are valid for all parameter sets (referred to as “all parameter sets” in the sentences below this paragraph) that are [[activated]] when decoding the stream that this record describes. Specifically, the following restrictions apply:

[0229] The profile indication general_profile_idc shall indicate the profile to which the [[associated stream]] of this configuration record conforms.

[0230] NOTE 1: If [[the SPS is with]] a different profile The stream can need to be examined to determine which profile (if any) the entire stream conforms to. If no examination of the entire stream is made, or if the examination shows that no profile is conformed to by the entire stream, the entire stream needs to be split into two or more sub-streams with separate configuration records in which these rules can be met.

[0231] The tier indication general_tier_flag shall indicate one tier that is equal to or greater than the highest tier indicated in all [[parameter sets]] profile_tier_level( ) syntax structure (in all parameter sets) of the [[associated stream]].

[0232] Each bit in general_constraints_info can only be set if [[all parameter sets set that bit]]

[0233] The level indication general_level_idc shall indicate a level that is equal to or greater than the highest level [[indicated for the highest tier]] in all [[parameter sets]] of the [[associated stream]].

[0234] The following constraint applies to chroma_format_idc:

[0235] - If the value of sps_chroma_format_idc as defined in ISO / IEC 23090-3 is the same in all SPS to which the NAL units of the track refer, chroma_format_idc shall be equal to sps_chroma_format_idc.

[0236] ​​​- Otherwise, if ptl_present_flag is equal to 1, chroma_format_idc shall be equal to vps_ols_dpb_chroma_format[ output_layer_set_idx ], as defined in ISO / IEC 23090-3.

[0237] - Otherwise, chroma_format_idc shall not be present.

[0238] The following constraint applies to bit_depth_minus8:

[0239] - If the value of sps_bitdepth_minus8 as defined in ISO / IEC 23090-3 is the same in all SPS referred to by the NAL units of the track, bit_depth_minus8 shall be equal to sps_bitdepth_minus8.

[0240] - Otherwise, if ptl_present_flag is equal to 1, bit_depth_minus8 shall be equal to vps_ols_dpb_bitdepth_minus8[ output_layer_set_idx ], as defined in ISO / IEC 23090-3.

[0241] - Otherwise, bit_depth_minus8 shall not be present.

[0242] Display indications about chroma format and bit depth are provided in the VVC decoder configuration record, along with other important format information used by VVC video elementary streams. If two sequences have different color spaces indicated in their VUI information, then two different VVC sample entries are also needed.

[0243] There are array groups to carry initialization NAL units. NAL unit types are restricted to indicate DCI, VPS, SPS, PPS, prefix APS, and prefix SEI NAL units. NAL unit types reserved in ISO / IEC 23090-3 and this specification can be defined in the future, and the reader should ignore arrays with reserved or disallowed NAL unit type values.

[0244] NOTE 2: This design of ‘tolerating’ behavior is to not raise errors, allowing backward-compatible extensions to these arrays in future specifications.

[0245] NOTE 3: The NAL units carried in the sample entry are included in the access unit reconstructed from the first sample referring to this sample entry or at the beginning of the access unit, immediately following the AUD and OPINAL units, if any.

[0246] The order of the suggested array is DCI, VPS, SPS, PPS, prefix APS, prefix SEI.

[0247] 6.1.1.1.2. Syntax

[0248]

[0249]

[0250]

[0251] 6.1.1.1.3. Semantics

[0252] general_profile_idc, general_tier_flag, general_level_idc, ptl_frame_only_constraint_flag, ptl_multilayer_enabled_flag, general_constraint_info(), ptl_sublayer_level_present[i], sublayer_level_idc[i], tl_num_sub_profiles, and general_sub_profile_idc[j] contain the matching values of the fields or syntax structures general_profile_idc, general_tier_flag, general_level_idc, ptl_frame_only_constraint_flag, ptl_multilayer_enabled_flag, general_constraint_info(), ptl_sublayer_level_present[i], sublayer_level_idc[i], tl_num_sub_profiles, and general_sub_profile_idc[j] as defined in ISO / IEC 23090-3 for the stream to which this configuration record applies.

[0253]

[0254] avgFrameRate gives the average frame rate of the stream to which this configuration record applies in frames per (256 seconds). The value 0 indicates that the average frame rate is not specified.

[0255] constantFrameRate equal to 1 indicates that the stream to which this configuration record applies has a constant frame rate. Value 2 indicates that the representations of each temporal layer in that stream are constant frame rate. Value 0 indicates that the stream can or can not be constant frame rate.

[0256] numTemporalLayers greater than 1 indicates that the track to which this configuration record applies is scalable in the temporal domain and contains a number of temporal layers (also referred to as temporal sub-layers or sub-layers in ISO / IEC 23090-3) equal to numTemporalLayers. Value 1 indicates that the track to which this configuration record applies is not scalable in the temporal domain. Value 0 indicates that it is unknown whether the track to which this configuration record applies is scalable in the temporal domain.

[0257] lengthSizeMinusOne plus 1 indicates the length (in bytes) of the NALUnitLength field in VVC video stream samples in the stream to which this configuration record applies. For example, a size of one byte is indicated with a value of 0. The value of this field shall be one of 0, 1, or 3, which correspond to lengths coded with 1, 2, or 4 bytes, respectively.

[0258] ptl_present_flag equal to 1 specifies that the track contains a VVC bitstream corresponding to a particular output layer set. ptl_present_flag equal to 0 specifies that the track can not contain a VVC bitstream corresponding to a particular output layer set, and either contains one or more individual layers that do not form an output layer set or contains individual sub-layers that do not include a sub-layer with Temporalld equal to 0.

[0259] track_ptl specifies the profile, tier, and level of the output layer set represented by the VVC bitstream contained in the track.

[0260] output_layer_set_idx specifies the output layer set index of the output layer set represented by the VVC bitstream contained in the track. The value of output_layer_set_idx can be used as the value of the TargetOlsIdx variable provided by an external component to a VVC decoder as specified in ISO / IEC 23090-3 for decoding the bitstream contained in the track.

[0261] chroma_format_present_flag equal to 0 specifies that chroma_format_idc is not present. chroma_format_present_flag equal to 1 specifies that chroma_format_idc is present.

[0262] chroma_format_idc indicates the chroma format applied to this track.

[0263] bit_depth_present_flag equal to 0 specifies that bit_depth_minus8 is not present. bit_depth_present_flag equal to 1 specifies that bit_depth_minus8 is present.

[0264] bit_depth_minus8 indicates the bit depth applied to this track.

[0265] numArrays indicates the number of arrays of NAL units of the indicated type.

[0266] array_completeness equal to 1 indicates that all NAL units of the given type are in the following array, none in the stream; array_completeness equal to 0 indicates that additional NAL units of the indicated type can be in the stream; the value is allowed to be constrained by the example entry name.

[0267] NAL_unit_type indicates the type of NAL units in the following array (which shall all be of this type); it takes values defined in ISO / IEC 23090-3; it is constrained to take one of the values indicating a DCI, VPS, SPS, PPS, prefix APS, or suffix SEI NAL unit.

[0268] numNalus indicates the number of NAL units of the indicated type included in the configuration record for the stream to which this configuration record applies. The SEI array shall contain only 'declarative' SEI messages, i.e. those providing information about the entire stream. One example of such SEI is the user data SEI.

[0269] nalUnitLength indicates the length of the NAL unit in bytes.

[0270] nalUnit contains a DCI, VPS, SPS, PPS, APS, or declarative SEI NAL unit as specified in ISO / IEC 23090-3.

[0271] Figure 1is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of system 1900. System 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8-bit or 10-bit multi-component pixel values, or can be received in a compressed or encoded format. Input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0272] System 1900 can include a codec 1904, which can implement various coding or encoding methods described in this document. Codec 1904 can reduce the average bitrate of a video from input 1902 to the output of codec 1904 to produce a coded representation of the video. Thus, coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec 1904 can be stored, or transmitted via a communication, as represented by component 1906. The stored or communicated bitstream (or encoded) representation of the video received at input 1902 can be used by component 1908 to generate pixel values or displayable video that is sent to a display interface 1910. The process of generating user-viewable video from a bitstream representation is sometimes referred to as video decompression. Furthermore, although certain video processing operations are referred to as “encoding” operations or tools, it should be understood that encoding tools or operations are used at an encoder, and corresponding decoding tools or operations that reverse the encoding results will be performed by a decoder.

[0273] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices, such as mobile telephones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.

[0274] Figure 2is a block diagram of a video processing device 3600. The device 3600 can be used to implement one or more of the methods described herein. The device 3600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 can be configured to implement one or more methods described in the present document. The memory(ies) 3604 can be used for storing data and code used during operation of the present techniques described herein. The video processing hardware 3606 can be used to implement, in hardware circuitry, some of the techniques described in the present document. In some embodiments, the video processing hardware 3606 can be included at least in part in the processor 3602, e.g., as a graphics co-processor.

[0275] Figure 4 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.

[0276] As Figure 4 shown, the video coding system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, and the destination device 120 can be referred to as a video decoding device.

[0277] The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0278] The video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data can comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to the destination device 120 by the network 130a via the I / O interface 116. The encoded video data can also be stored onto a storage medium / server 130b for access by the destination device 120.

[0279] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.

[0280] The I / O interface 126 can include a receiver and / or a modem. The I / O interface 126 can obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to a user. The display device 122 can be integrated with the destination device 120, or can be external to the destination device 120 which is configured to interface with an external display device.

[0281] The video encoder 114 and the video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.

[0282] Figure 5 is a block diagram illustrating an example of a video encoder 200 that can be Figure 4 the video encoder 114 in the system 100 shown.

[0283] The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 5 examples, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared amongst the various components of the video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0284] The functional components of the video encoder 200 can include a partition unit 201, a prediction unit 202, which can include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.

[0285] In other examples, the video encoder 200 can include more, less, or different functional components. In one example, the prediction unit 202 can include an intra block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0286] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but are represented separately for illustrative purposes. Figure 5 in examples.

[0287] Partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.

[0288] Mode selection unit 203 can select one of the intra or inter coding modes, e.g., based on error results, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 can select a combination of intra and inter prediction (CIIP) mode, where the prediction is based on both inter prediction signals and intra prediction signals. Mode selection unit 203 can also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision) in the case of inter prediction.

[0289] To perform inter prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0290] Motion estimation unit 204 and motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0291] In some examples, motion estimation unit 204 can perform single prediction for a current video block, and motion estimation unit 204 can search a reference picture in List 0 or List 1 for a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference picture in List 0 or List 1 that contains the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0292] In other examples, the motion estimation unit 204 can perform bi-prediction for the current video block, the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0 and also search for another reference video block for the current video block in a reference picture in list 1. The motion estimation unit 204 can then generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector that indicates a spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 can output the reference index and the motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.

[0293] In some examples, the motion estimation unit 204 can output a complete set of motion information for a decoding process of a decoder.

[0294] In some examples, the motion estimation unit 204 can not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 can signal the motion information for the current video block with reference to motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.

[0295] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.

[0296] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0297] As described above, the video encoder 200 can predictively signal the motion vector. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0298] The intra prediction unit 206 can perform intra prediction for the current video block. When the intra prediction unit 206 performs intra prediction for the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0299] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the predicted video block for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0300] In other examples, such as in a skip mode, there can be no residual data for the current video block for the current video block, and residual generation unit 207 can not perform a subtraction operation.

[0301] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0302] Quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block after transform processing unit 208 generates the transform coefficient video block associated with the current video block.

[0303] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to a transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.

[0304] After reconstruction unit 212 reconstructs a video block, in-loop filtering operations can be performed to reduce video block artifacts in the video block.

[0305] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.

[0306] Figure 6 is a block diagram illustrating an example of a video decoder 300 that can be Figure 4 the video decoder 114 in the system 100 shown.

[0307] Video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 6 examples, video decoder 300 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0308] In Figure 6 In the example of FIG. 3, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, video decoder 300 can perform decoding passes substantially reciprocal to the encoding passes described with respect to video encoder 200. Figure 5

[0309] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., coded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data and motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference picture list indices, and other motion information, from the entropy decoded video data. Motion compensation unit 302 can determine such information, for example, by performing AMVP and merge modes.

[0310] Motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter to use at sub-pixel precision can be included in the syntax elements.

[0311] Motion compensation unit 302 can use the interpolation filter used by video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from the received syntax information and use the interpolation filter to generate the prediction block.

[0312] Motion compensation unit 302 can use some of the syntax information to determine the size of blocks used to encode frames and / or slices of the coded video sequence, partition information describing how each macroblock of a picture of the coded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used in decoding the coded video sequence.

[0313] Intra-prediction unit 303 can use intra-prediction modes, e.g., received in the bitstream, to form prediction blocks from spatially neighboring blocks. Inverse quantization unit 303 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transformation unit 303 applies an inverse transform.

[0314] ​The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and also produces decoded video for presentation on a display device.

[0315] Next, a list of preferred solutions for some embodiments is provided.

[0316] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., items 1-4).

[0317] 1. A method of visual media processing (e.g., the method 3000 described in Figure 3 Solution 1), the method comprising performing (3002) a conversion between visual media data and a file storing a bitstream representation of the visual media data according to a format rule; wherein the format rule specifies a constraint on information included in the file in relation to a profile, tier, constraint, or layer associated with the bitstream representation identified in the file.

[0318] 2. The method according to solution 1, wherein the format rule specifies that the file includes an identification of a profile to which an output layer set of the bitstream representation identified in the file conforms.

[0319] 3. The method according to any of solutions 1 to 2, wherein the format rule specifies that a layer identified in the file is equal to or higher than a highest layer indicated in all syntax structures included in an output layer set included in the file to which the output layer set conforms.

[0320] 4. The method according to any of solutions 1 to 3, wherein the format rule specifies that a constraint identified in the file aligns with corresponding values indicated by one or more constraint fields of a syntax structure indicating a constraint to which an output layer set in the file conforms.

[0321] 5. The method according to any of solutions 1 to 4, wherein the format rule specifies that a tier identified in the file aligns with corresponding values indicated by one or more constraint fields of a syntax structure indicating a tier to which an output layer set in the file conforms.

[0322] 6. The method according to any of solutions 1 to 5, wherein the conversion comprises generating a bitstream representation of the visual media data and storing the bitstream representation into the file according to the format rule.

[0323] 7. The method according to any one of solutions 1 to 5, wherein the converting comprises parsing the file according to the format rules to recover the visual media data.

[0324] 8. A video decoding device comprising a processor configured to implement the method described in one or more of solutions 1 to 7.

[0325] 9. A video encoding and decoding device, comprising a processor, wherein the processor is configured to implement the method described in one or more of solutions 1 to 7.

[0326] 10. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of solutions 1 to 7.

[0327] 11. A computer readable medium having thereon a bitstream representation conforming to a file format generated according to any one of solutions 1 to 7.

[0328] 12. The methods, apparatus, or systems described in this document.

[0329] In the solution described herein, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder can use the format rules to parse syntax elements in the codec representation to produce decoded video, knowing the presence and absence of syntax elements according to the format rules.

[0330] Technology 1. A method for processing visual media data (e.g., Figure 8 Method 8000 as shown) comprises: performing (8002) conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the bitstream comprises one or more output layer sets and one or more parameter sets, the one or more parameter sets comprising one or more profile layer level syntax structures, wherein at least one of the profile layer level syntax structures comprises a general constraint information syntax structure, wherein the format rule specifies that a syntax element is included in a configuration record in the visual media file, and wherein the syntax element indicates a profile, layer or level to which the output layer set identified by an output layer set index indicated in the configuration record conforms.

[0331] Technique 2. The method of Technique 1, wherein the syntax element is a general profile indicator syntax element that indicates the profile to which the output layer set identified by the output layer set index conforms.

[0332] TECHNIQUE 3. The method according to technique 1, wherein the syntax element is a general layer syntax element that indicates a highest layer indicated in all tier level syntax structures to which the output layer set identified by the output layer set index conforms.

[0333] TECHNIQUE 4. The method according to technique 1, wherein the syntax element is a general layer syntax element that indicates a highest layer indicated in all tier level syntax structures to which the output layer set identified by the output layer set index conforms.

[0334] TECHNIQUE 5. The method according to technique 1, wherein the syntax element is a general layer syntax element that indicates a highest layer to which a stream associated with the configuration record conforms.

[0335] TECHNIQUE 6. The method according to technique 1, wherein the syntax element is a general layer syntax element that indicates a layer to which a stream associated with the configuration record conforms.

[0336] TECHNIQUE 7. The method according to technique 1, wherein the configuration record includes a general constraint information syntax element, wherein the format rule specifies that a first bit in the general constraint information syntax element corresponds to a second bit in all general constraint information syntax structures in all tier level syntax structures to which the output layer set identified by the output layer set index conforms, and wherein the format rule specifies that the first bit is set to one only when the second bit in all general constraint information syntax structures is set to one.

[0337] TECHNIQUE 8. The method according to technique 1, wherein the syntax element is a general level syntax element whose value indicates a highest level indicated in all tier level syntax elements to which the output layer set identified by the output layer set index conforms.

[0338] TECHNIQUE 9. The method according to technique 1, wherein the format rule specifies that the syntax element is not allowed to be associated with one or more other output layer sets included in a stream stored in the visual media file.

[0339] TECHNIQUE 10. The method according to any of techniques 1 to 9, wherein the converting includes generating the visual media file and storing the bitstream into the visual media file according to the format rule.

[0340] TECHNIQUE 11. The method according to any of techniques 1 to 9, wherein the converting includes generating the visual media file, and the method further includes storing the visual media file in a non-transitory computer readable recording medium.

[0341] TECHNIQUE 12. The method according to any of techniques 1 to 9, wherein the converting includes parsing the visual media file to reconstruct the bitstream according to the format rule.

[0342] Technical 13. The method according to any of techniques 1 to 12, wherein the visual media file is processed by Versatile Video Coding (VVC).

[0343] Technical 14. An apparatus for processing visual media data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method in one or more of techniques 1 to 13.

[0344] Technical 15. A non-transitory computer-readable storage medium storing instructions that cause a processor to implement the method of any of techniques 1 to 13.

[0345] Technical 16. A video decoding apparatus comprising a processor configured to implement the method recited in one or more of techniques 1 to 13.

[0346] Technical 17. A video encoding apparatus comprising a processor configured to implement the method recited in one or more of techniques 1 to 13.

[0347] Technical 18. A computer program product having stored thereon computer code which, when executed by a processor, causes the processor to implement the method recited in any of techniques 1 to 13.

[0348] Technical 19. A computer-readable medium having thereon a visual media file conforming to a file format generated according to any of techniques 1 to 13.

[0349] Technical 20. A visual media file generation method comprising: generating a visual media file according to the method of any of techniques 1 to 13, and storing the visual media file on a computer-readable program medium.

[0350] Technical 21. A non-transitory computer-readable recording medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method is recited in any of techniques 1 to 13. In some embodiments, a non-transitory computer-readable recording medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method comprises: generating a visual media file based on visual media data according to a format rule, wherein the bitstream comprises one or more output layer sets and one or more parameter sets, the one or more parameter sets comprising one or more profile tier level syntax structures, wherein at least one of the profile tier level syntax structures comprises a general constraint information syntax structure, wherein the format rule specifies that a syntax element is included in a configuration record in the visual media file, and wherein the syntax element indicates a profile, tier, or level to which an output layer set identified by an output layer set index indicated in the configuration record conforms.

[0351] Implementation 1. A method of processing visual media data (e.g., the method 9000 described in Figure 9 Implementation 1), the method comprising performing (9002) a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a characteristic of a syntax element in the visual media file, wherein the syntax element has a value that indicates a number of bytes used to indicate constraint information associated with the bitstream.

[0352] Implementation 2. The method of claim 1, wherein the format rule specifies that the syntax element is coded in the visual media file using six bits.

[0353] Implementation 3. The method of implementation 1, wherein the format rule specifies that the syntax element is coded in the visual media file immediately after a tier level level-idc syntax element in the visual media file.

[0354] Implementation 4. The method of implementation 1, wherein the format rule specifies that the syntax element is coded in the visual media file, specifies a number of bytes in a general constraint information syntax element in the visual media file, and wherein the format rule specifies that a value of the syntax element equal to one indicates that a general constraint information flag in the general constraint information syntax element is equal to zero and the general constraint information syntax element is prohibited from being included in a tier level record in the visual media file.

[0355] Implementation 5. The method of implementation 1, wherein the format rule specifies that a condition for including a general constraint information syntax element in the visual media file depends on whether a value indicated by the syntax element is greater than one.

[0356] Implementation 6. The method of implementation 1, wherein the format rule specifies that a number of bits used to code a general constraint information syntax element into the visual media file is a result of eight multiplied by a value that indicates a number of bytes used to indicate the constraint information, and wherein the format rule specifies that the result of eight multiplied by the value that indicates the number of bytes used to indicate the constraint information is not subtracted by two.

[0357] Implementation 7. A method of processing visual media data, the method comprising performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that five bits are used for a syntax element in the visual media file, wherein the syntax element has a value that indicates a network abstraction layer unit type in a decoder configuration record in the visual media file. In some embodiments, wherein the format rule specifies that five bits are used for another syntax element in the visual media file, and wherein the another syntax element has another value that indicates the network abstraction layer unit type in the decoder configuration record in the visual media file.

[0358] Implementation 8. A method of processing visual media data, the method comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein a track in the visual media file comprises a video bitstream, the video bitstream comprising one or more output layer sets; and wherein the format rule specifies that a syntax element is indicated for the track, wherein the syntax element indicates whether the track comprises a video bitstream corresponding to a particular output layer set from the one or more output layer sets. In some embodiments, wherein a track in the visual media file comprises a video bitstream, the video bitstream comprising one or more output layer sets, wherein the format rule specifies that another syntax element is indicated for the track, and wherein the another syntax element indicates whether the track comprises a video bitstream corresponding to a particular output layer set from the one or more output layer sets.

[0359] Implementation 9. The method of implementation 8, wherein the syntax element indicates that the track comprises a video bitstream corresponding to a plurality of output layer sets. In some embodiments, the another syntax element indicates that the track comprises a video bitstream corresponding to a plurality of output layer sets.

[0360] Implementation 10. The method of implementation 8, wherein the syntax element indicates that the track comprises a video bitstream that does not correspond to a particular output layer set from the one or more output layer sets. In some embodiments, the another syntax element indicates that the track comprises a video bitstream that does not correspond to a particular output layer set from the one or more output layer sets.

[0361] Implementation 11. A method of processing visual media data, the method comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies whether the visual media file comprises a syntax element whose value indicates an output layer set index for indicating an output layer set. In some embodiments, the format rule specifies whether the visual media file comprises another syntax element whose value indicates an output layer set index for indicating an output layer set.

[0362] Implementation 12. The method of implementation 11, wherein the format rule specifies that, in response to another value of a profile layer presence flag syntax element in the visual media file being equal to one, or in response to a tier level multi-layer enabled flag being equal to one, the visual media file selectively indicates the syntax element whose value indicates an output layer set index in a decoder configuration record. In some embodiments, the format rule specifies that, in response to another value of a profile layer presence flag syntax element in the visual media file being equal to one, or in response to a tier level multi-layer enabled flag being equal to one, the visual media file selectively indicates the another syntax element whose value indicates an output layer set index in a decoder configuration record.

[0363] Implementation 13. The method according to implementation 11, wherein the format rule specifies that the visual media file is not allowed to include a syntax element whose value indicates the output layer set index, and wherein the format rule specifies that in response to a profile layer present flag syntax element in the visual media file being equal to one, a value of the output layer set index is inferred to be equal to a second value of a second output layer index of a second output layer set that includes only layers carried in the track. In some embodiments, the format rule specifies that the visual media file is not allowed to include another syntax element whose value indicates the output layer set index, and wherein the format rule specifies that in response to the profile layer present flag syntax element in the visual media file being equal to one, the value of the output layer set index is inferred to be equal to the second value of the second output layer index of the second output layer set that includes only layers carried in the track.

[0364] Implementation 14. The method according to any of implementations 1 to 13, wherein the converting comprises generating the visual media file and storing the bitstream into the visual media file according to the format rule.

[0365] Implementation 15. The method according to any of implementations 1 to 13, wherein the converting comprises generating the visual media file, and the method further comprises storing the visual media file in a non-transitory computer readable recording medium.

[0366] Implementation 16. The method according to any of implementations 1 to 13, wherein the converting comprises parsing the visual media file to reconstruct the bitstream according to the format rule.

[0367] Implementation 17. The method according to any of implementations 1 to 16, wherein the visual media file is processed by Versatile Video Coding (VVC).

[0368] Implementation 18. An apparatus for processing visual media data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method in one or more of implementations 1 to 17.

[0369] Implementation 19. A non-transitory computer readable storage medium storing instructions that cause a processor to implement the method of any of implementations 1 to 17.

[0370] Implementation 20. A video decoding apparatus comprising a processor configured to implement the method recited in one or more of implementations 1 to 17.

[0371] Implementation 21. A video coding apparatus comprising a processor configured to implement the method recited in one or more of implementations 1 to 17.

[0372] Implementation 22. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement the method recited in any of implementations 1 to 17.

[0373] Implementation 23. A computer readable medium having stored thereon a visual media file conforming to a file format generated according to any of implementations 1 to 17.

[0374] Implementation 24. A visual media file generation method, comprising: generating a visual media file according to the method of any of implementations 1 to 17, and storing the visual media file on a computer readable program medium.

[0375] Implementation 25. A non-transitory computer readable recording medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method is recited in any of implementations 1 to 17. A non-transitory computer readable recording medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method comprises: generating the visual media file based on visual media data according to a format rule, wherein the format rule specifies a property of a syntax element in the visual media file, and wherein the syntax element has a value indicating a number of bytes used to indicate constraint information associated with the bitstream.

[0376] Operation 1. A method of processing visual media data (e.g., Figure 9 The method 10000 shown, comprising: performing (10002) a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a property of a syntax element in the visual media file, and wherein the format rule specifies that a syntax element having a value indicating a level identification is used in either or both of a subpicture common group box or a subpicture multiple group box using octet coding.

[0377] Operation 2. The method of operation 1, wherein the format rule specifies absence of reserved bits immediately following the syntax element whose value indicates the level identification.

[0378] Operation 3. The method of operation 1, wherein the format rule specifies that 24 bits immediately following the syntax element whose value indicates the level identification are reserved bits.

[0379] Operation 4. The method of operation 1, wherein the format rule specifies that 8 bits immediately following the syntax element whose value indicates the level identification are reserved bits.

[0380] Operation 5. A method of processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a property related to a first syntax element, a second syntax element, or a third syntax element set in the visual media file, wherein the first syntax element has a first value indicating a number of active tracks in the visual media file, wherein the second syntax element has a second value indicating a number of subgroup identifiers in the visual media file, and wherein each syntax element in the third syntax element set has a third value indicating a number of active tracks in the visual media file. In some embodiments, wherein the format rule specifies a property related to a first syntax element, a second syntax element, or a third syntax element set in the visual media file, wherein the first syntax element has a first value indicating a number of active tracks in the visual media file, wherein the second syntax element has a second value indicating a number of subgroup identifiers in the visual media file, and wherein each syntax element in the third syntax element set has a third value indicating a number of active tracks in the visual media file.

[0381] Operation 6. The method of operation 5, wherein the format rule specifies 16 bits for indicating the first syntax element having the first value indicating the number of active tracks in the subpicture common group box in the visual media file.

[0382] Operation 7. The method of operation 5, wherein the format rule specifies 16 bits for indicating the second syntax element having the second value indicating the number of subgroup identifiers in the subpicture multiple group box in the visual media file, and wherein the format rule specifies 16 bits for indicating each syntax element in the third syntax element set having the third value indicating the number of active tracks in the subpicture multiple group box in the visual media file.

[0383] Operation 8. The method of operation 5, wherein the format rule specifies that 16 bits immediately following the first syntax element having the first value indicating the number of active tracks, the second syntax element having the second value indicating the number of subgroup identifiers, or each syntax element in the third syntax element set having the third value indicating the number of active tracks are reserved.

[0384] Operation 9. The method of operation 5, wherein the format rule specifies a lack of reserved bits immediately following the first syntax element having the first value indicating the number of active tracks, the second syntax element having the second value indicating the number of subgroup identifiers, or each syntax element in the third syntax element set having the third value indicating the number of active tracks.

[0385] Operation 10. The method according to any one of operations 1 to 9, wherein the converting comprises generating the visual media file and storing the bitstream into the visual media file according to the format rule.

[0386] Operation 11. The method according to any one of operations 1 to 9, wherein the converting comprises generating the visual media file, and the method further comprises storing the visual media file in a non-transitory computer readable recording medium.

[0387] Operation 12. The method according to any one of operations 1 to 9, wherein the converting comprises parsing the visual media file according to the format rule to reconstruct the bitstream.

[0388] Operation 13. The method according to any one of operations 1 to 12, wherein the visual media file is processed by Versatile Video Coding (VVC).

[0389] Operation 14. An apparatus for processing visual media data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method in one or more of operations 1 to 13.

[0390] Operation 15. A non-transitory computer readable storage medium storing instructions, which cause a processor to implement the method of any one of operations 1 to 13.

[0391] Operation 16. A video decoding apparatus comprising a processor configured to implement the method recited in one or more of operations 1 to 13.

[0392] Operation 17. A video encoding apparatus comprising a processor configured to implement the method recited in one or more of operations 1 to 13.

[0393] Operation 18. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement the method recited in any one of operations 1 to 13.

[0394] Operation 19. A computer readable medium having a visual media file in a file format generated according to any one of operations 1 to 13.

[0395] Operation 20. A visual media file generation method comprising: generating a visual media file according to the method of any one of operations 1 to 13, and storing the visual media file on a computer readable program medium.

[0396] Operation 21. A non-transitory computer-readable recording medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method is described in any of operations 1 to 13. In some embodiments, a non-transitory computer-readable recording medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method comprises: generating the visual media file based on visual media data according to a format rule, wherein the format rule specifies a property of a syntax element in the visual media file, and wherein the format rule specifies that a syntax element having a value indicative of a level identification is signaled using octets in either or both of a subpicture common group box or a subpicture multiple group box.

[0397] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, a bitstream representation of a current video block can correspond to bits that are collocated or spread out in different locations within a bitstream as defined by syntax. For example, a macroblock can be encoded according to transformed and coded error residual values, and also using bits in a header and other fields in the bitstream. Furthermore, during conversion, a decoder can parse a bitstream based on a decision as described in the above-described solutions, with knowledge that some fields can or can not be present. Similarly, an encoder can determine that certain syntax fields are included or not included, and generate a coded representation accordingly by including or excluding syntax fields from the coded representation.

[0398] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine- readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for said program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.

[0399] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.

[0400] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0401] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0402] Although this patent document contains many details, these should not be construed as limiting the scope of any subject matter or of any embodiment, but as merely describing features that can be present and combinations thereof. Certain features that are described in the context of separate embodiments can also be implemented in combination with each other. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any appropriate subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.

[0403] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to achieve desirable results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0404] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method of processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a property of a syntax element in the visual media file, wherein the syntax element has a value indicative of a number of bytes used to indicate constraint information associated with the bitstream, wherein the format rule specifies that the syntax element indicates the number of bytes in a general constraint information syntax element in the visual media file and is coded in the visual media file using N bits, wherein N is specified to be equal to 6 only, wherein a track in the visual media file comprises a video bitstream, the video bitstream comprising one or more output layer sets, wherein the format rule specifies that a first bit having a particular bit index in the general constraint information syntax element corresponds to a second bit having the particular bit index in all general constraint information syntax structures in all tier level syntax structures, the one or more output layer sets conforming to the tier level syntax structures, and wherein the format rule specifies that the first bit is set to one only when the second bits in the all general constraint information syntax structures are set to one.

2. The method of claim 1, wherein wherein the format rule specifies that the syntax element is coded in the visual media file immediately after a tier level multi-layer enabled flag syntax element in the visual media file.

3. The method of claim 1, wherein, wherein the format rule specifies that the syntax element is coded in the visual media file, specifies a number of bytes in a general constraint information syntax element in the visual media file, and wherein the format rule specifies that a value of the syntax element equal to one indicates that a general constraint information flag in the general constraint information syntax element is equal to zero and disallows the general constraint information syntax element to be included in a tier level record in the visual media file.

4. The method of claim 1, wherein wherein the format rule specifies that a condition for including a general constraint information syntax element in the visual media file depends on whether the value indicated by the syntax element is greater than one.

5. The method of claim 1, wherein wherein the format rule specifies that five bits are used for another syntax element in the visual media file, and wherein the another syntax element has another value, the another value indicating a network abstraction layer unit type in a decoder configuration record in the visual media file.

6. The method of claim 1, wherein, wherein a track in the visual media file comprises a video bitstream, the video bitstream comprising one or more output layer sets, wherein the format rule specifies that another syntax element is indicated for the track, and wherein the another syntax element indicates whether the track comprises the video bitstream corresponding to a particular output layer set from the one or more output layer sets.

7. The method of claim 6, wherein, wherein the another syntax element indicates that the track comprises the video bitstream corresponding to a plurality of output layer sets.

8. The method of claim 6, wherein, The other syntax element indicates that the track includes the video bitstream that does not correspond to the particular output layer set from the one or more output layer sets.

9. The method of claim 1, wherein, The conversion includes generating the visual media file and storing the bitstream into the visual media file according to the format rule.

10. The method of claim 1, wherein, The conversion includes generating the visual media file and the method further includes storing the visual media file in a non-transitory computer-readable recording medium.

11. The method of claim 1, wherein, The conversion includes parsing the visual media file according to the format rule to reconstruct the bitstream.

12. The method of claim 1, wherein, The visual media file is processed by a versatile video coding (VVC).

13. An apparatus for processing visual media data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein, The instructions, when executed by the processor, cause the processor to implement the method recited in one or more of claims 1-12.

14. A non-transitory computer-readable storage medium storing instructions that cause a processor to implement the method recited in any of claims 1-12.

15. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein, the format rule specifies a property of a syntax element in the visual media file, wherein the syntax element has a value indicating a number of bytes used to indicate constraint information associated with the bitstream, wherein the format rule specifies that the syntax element indicates the number of bytes in a general constraint information syntax element in the visual media file and the syntax element is coded in the visual media file using N bits, wherein it is specified that N is only equal to 6, wherein a track in the visual media file includes a video bitstream, the video bitstream including one or more output layer sets, wherein the format rule specifies that a first bit having a particular bit index in the general constraint information syntax element corresponds to a second bit having the particular bit index in all general constraint information syntax structures in all tier level syntax structures, the one or more output layer sets conforming to the tier level syntax structures, and wherein the format rule specifies that the first bit is set to 1 only when the second bit in the all general constraint information syntax structures is set to 1.

16. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies a property of a syntax element in the visual media file, wherein the syntax element has a value indicating a number of bytes used to indicate constraint information associated with the bitstream, wherein the format rule specifies that the syntax element indicates the number of bytes in a general constraint information syntax element in the visual media file and the syntax element is coded in the visual media file using N bits, wherein it is specified that N is only equal to 6, wherein a track in the visual media file includes a video bitstream, the video bitstream including one or more output layer sets, wherein the format rule specifies that a first bit having a particular bit index in the general constraint information syntax element corresponds to a second bit having the particular bit index in all general constraint information syntax structures in all tier level syntax structures to which the one or more output layer sets conform, and wherein the format rule specifies that the first bit is set to one only when the second bit in all the general constraint information syntax structures is set to one.

17. A non-transitory computer-readable recording medium storing a visual media file generated by a method performed by a video processing apparatus, wherein, The method comprises: generating the visual media file based on visual media data according to a format rule, wherein the format rule specifies a property of a syntax element in the visual media file, wherein the syntax element has a value indicating a number of bytes used to indicate constraint information associated with a bitstream, wherein the format rule specifies that the syntax element indicates a number of bytes in a general constraint information syntax element in the visual media file, and the syntax element is coded in the visual media file using N bits, wherein N is specified to be equal to 6 only, wherein a track in the visual media file comprises a video bitstream, the video bitstream comprising one or more output layer sets, wherein the format rule specifies that a first bit having a particular bit index in the general constraint information syntax element corresponds to a second bit having the particular bit index in all general constraint information syntax structures in all tier level syntax structures to which the one or more output layer sets conform, and wherein the format rule specifies that the first bit is set to one only when the second bit in all the general constraint information syntax structures is set to one.

18. A method of storing a visual media file, comprising: generating the visual media file based on visual media data according to a format rule, storing the visual media file in a non-transitory computer-readable recording medium; wherein the format rule specifies a property of a syntax element in the visual media file, wherein the syntax element has a value indicating a number of bytes used to indicate constraint information associated with a bitstream, wherein the format rule specifies that the syntax element indicates a number of bytes in a general constraint information syntax element in the visual media file, and the syntax element is coded in the visual media file using N bits, wherein N is specified to be equal to 6 only, wherein a track in the visual media file comprises a video bitstream, the video bitstream comprising one or more output layer sets, wherein the format rule specifies that a first bit having a particular bit index in the general constraint information syntax element corresponds to a second bit having the particular bit index in all general constraint information syntax structures in all tier level syntax structures to which the one or more output layer sets conform, and wherein the format rule specifies that the first bit is set to one only when the second bit in all the general constraint information syntax structures is set to one.